Public scan — anyone with this URL can view this analysis. Sign up to track your own repos privately, run scheduled re-scans, and get AI fix prompts via your dashboard.

benjibrcz/probing-silent-reward-hacking

https://github.com/benjibrcz/probing-silent-reward-hacking · scanned 2026-06-17 01:38 UTC (1 month, 2 weeks ago)

8 raw signals (0 security + 8 graph)

UNIFIED Repobility · multi-layer engine · AI coders

Complete repo analysis

Last scanned 1 month, 2 weeks ago · v2 · last Δ +14.9 (diff) · 8 actionable findings from 1 signal source. Security checks, system graph analysis, and verified AI-agent feedback are merged into one review queue.

JSON
Severity distribution — click a segment to filter
Active filters: excluding tests × Reset all

All 79 nodes from the latest scan, grouped by kind. Each node is a unit the engine identified (file, function, endpoint, table…). Most users won't need this view — it's primarily for debugging the engine's graph extraction or for AI agents that want to enumerate the project structure.

LabelLayerStatusPath
RESUME.md software healthy RESUME.md
setup.sh software healthy setup.sh
SUBMISSION.md software healthy SUBMISSION.md
README.md software healthy README.md
RED_TEAM.md software healthy RED_TEAM.md
WRITEUP.md software healthy WRITEUP.md
SUBMISSION_COVER.md software healthy SUBMISSION_COVER.md
NOTES.md software healthy NOTES.md
SUBMISSION.html software healthy SUBMISSION.html
generate_sdf.py software healthy src/generate_sdf.py
analyze_sdf.py software healthy src/analyze_sdf.py
controls.py software healthy src/controls.py
analyze_auc.py software healthy src/analyze_auc.py
transfer.py software healthy src/transfer.py
fig_twoorganism.py software healthy src/fig_twoorganism.py
monitor_v2.py software healthy src/monitor_v2.py
aisi_prompts.py software healthy src/aisi_prompts.py
llm_monitor.py software healthy src/llm_monitor.py
extract.py software healthy src/extract.py
build_dataset.py software healthy src/build_dataset.py
llm_monitor_scores.json software healthy results/llm_monitor_scores.json
llm_monitor_b00.json software healthy results/llm_monitor_b00.json
mon_b002_gpt41_zs_scores.json software healthy results/mon_b002_gpt41_zs_scores.json
auc_analysis.json software healthy results/auc_analysis.json
llm_monitor.json software healthy results/llm_monitor.json
llm_monitor_b00_scores.json software healthy results/llm_monitor_b00_scores.json
monitor_2x2.json software healthy results/monitor_2x2.json
sdf_analysis.json software healthy results/sdf_analysis.json
mon_sdf_fs6_scores.json software healthy results/mon_sdf_fs6_scores.json
gpt_critique.md software healthy results/gpt_critique.md
llm_monitor_sdf_scores.json software healthy results/llm_monitor_sdf_scores.json
transfer.json software healthy results/transfer.json
mon_sdf_gpt41_zs.json software healthy results/mon_sdf_gpt41_zs.json
llm_monitor_sdf.json software healthy results/llm_monitor_sdf.json
mon_b002_fs6.json software healthy results/mon_b002_fs6.json
mon_sdf_gpt41_zs_scores.json software healthy results/mon_sdf_gpt41_zs_scores.json
controls.json software healthy results/controls.json
mon_b002_gpt41_zs.json software healthy results/mon_b002_gpt41_zs.json
mon_b002_fs6_scores.json software healthy results/mon_b002_fs6_scores.json
mon_sdf_fs6.json software healthy results/mon_sdf_fs6.json

LabelLayerStatusPath
split_think software healthy src/generate_sdf.py:36
knowledge_check software healthy src/generate_sdf.py:44
main software healthy src/generate_sdf.py:61
load software healthy src/analyze_sdf.py:28
grouped_oof software healthy src/analyze_sdf.py:36
boot_ci software healthy src/analyze_sdf.py:48
main software healthy src/analyze_sdf.py:55
load software healthy src/controls.py:31
oof_proba software healthy src/controls.py:41
recall_at_fpr software healthy src/controls.py:53
main software healthy src/controls.py:62
load software healthy src/analyze_auc.py:26
cv_auc software healthy src/analyze_auc.py:40
boot_ci software healthy src/analyze_auc.py:54
main software healthy src/analyze_auc.py:62
load software healthy src/transfer.py:31
indist_cv software healthy src/transfer.py:41
transfer software healthy src/transfer.py:54
direction software healthy src/transfer.py:61
best_layer software healthy src/transfer.py:69
main software healthy src/transfer.py:80
main software healthy src/monitor_v2.py:25
_format_hack_hints software healthy src/aisi_prompts.py:167
_build_prompt software healthy src/aisi_prompts.py:244
build_shuffled_prompt software dead src/aisi_prompts.py:268
main software healthy src/llm_monitor.py:34
recall_at_fpr software healthy src/llm_monitor.py:73
build_inputs software healthy src/extract.py:44
extract_run software healthy src/extract.py:82
main software healthy src/extract.py:117
parse_dump software healthy src/build_dataset.py:22
at software healthy src/build_dataset.py:33
grab software healthy src/build_dataset.py:46
assign_splits software healthy src/build_dataset.py:80
main software healthy src/build_dataset.py:88

LabelLayerStatusPath
src software healthy src
results software healthy results

LabelLayerStatusPath
benjibrcz__probing-silent-reward-hacking software healthy /data/fable5_failed_archive/benjibrcz__probing-silent-rewar…

LabelLayerStatusPath
gpu (detected) hardware healthy setup.sh
For AI agents: Voting guide (TP/FP) MCP manifest Stdio wrapper SARIF Integrate Findings queue Vote TP/FP on findings to calibrate the engine.
For AI agents + API integrations
Email me when this repo regresses
Free. We re-scan periodically; new criticals → your inbox. No signup required for the scan itself.
API access

This page is publicly accessible at: https://repobility.com/scan/e98a212a-70fb-4baf-b67a-911d467b0f32/

To check status programmatically (no auth required):

curl -s https://repobility.com/api/v1/public/scan/e98a212a-70fb-4baf-b67a-911d467b0f32/

Important — please don't re-submit the same URL repeatedly. The submission endpoint is idempotent: re-submitting the same git URL returns this same scan_token, not a new one. To re-scan this repo, sign up free and use the dashboard.