Simon Sure · Sarthak Munshi · Rishabh Sinha · Zhun Wang · Dawn Song · ICLR 2027 submission
A composable framework separating attackers, systems under test, security claims,
and controllers. Security domains make adversary trust boundaries explicit,
enforced, and sweepable across experiments.
Rishabh Sinha · UC Berkeley · Agents in the Wild @ NeurIPS 2026
Separates injection exposure from behavioral commitment, in both the
evaluation and the training population. Monitors trained on the naive
contrast mostly learn exposure; retrained on exposed cases only, the same
probes predict the agent’s own unsafe tool call up to two steps early.
0.85 – 0.91
commitment AUROC, three models
2 steps
lead time, trajectory maximum
Under review
A Nonzero Final Learning Rate Floor Improves Transfer of Agent-Discovered Schedule Tweaks
Rishabh Sinha · NeurIPS 2026 main conference
Tests whether an AutoResearch-discovered schedule change survives controlled
decomposition, five-seed reruns, and nearby depth transfer. A 0.05 final
learning-rate floor—not warmup—was the robust component in the fixed-budget
single-GPU harness.
4 / 5
base seeds improved
5 / 5
depth-10 seeds improved
0.000877
base mean BPB gain
0.001570
depth-10 mean BPB gain
05 / Education
Education
Computer science, mathematics, data systems, and computational finance.