Simon Sure · Sarthak Munshi · Rishabh Sinha · Zhun Wang · Dawn Song · ICLR 2027 submission
A composable framework separating attackers, systems under test, security claims,
and controllers. Security domains make adversary trust boundaries explicit,
enforced, and sweepable across experiments.
Adds activation-derived evidence immediately before agent tool use. Learned
NLA-text and raw-activation monitors flagged all 22 generated unsafe actions
before execution in the primary AgentDojo-grounded evaluation.
22 / 22
unsafe actions flagged
2.0
mean steps of lead time
NeurIPS submission
A Nonzero Final Learning Rate Floor Improves Transfer of Agent-Discovered Schedule Tweaks
Rishabh Sinha · NeurIPS 2026 submission
Tests whether an AutoResearch-discovered schedule change survives controlled
decomposition, five-seed reruns, and nearby depth transfer. A 0.05 final
learning-rate floor—not warmup—was the robust component in the fixed-budget
single-GPU harness.
4 / 5
base seeds improved
5 / 5
depth-10 seeds improved
0.000877
base mean BPB gain
0.001570
depth-10 mean BPB gain
04 / Education
Education
Computer science, mathematics, data systems, and computational finance.
2026
University of California, Berkeley
Master of Information and Data Science
Part-time professional programGPA 4.0
2025
University of Maryland, College Park
Bachelor’s in Computer Science and Mathematics
Minor in Computational FinanceGPA 3.98
05 / Toolkit
Skills
A focused toolkit for production ML, distributed infrastructure, and performance work.