A one-day, invite-only summit providing a first look at the
benchmarks and research that will shape the frontier.
October 8, 2026
San Francisco
Topics
Environment complexity
Evaluating agents in realistic, dynamic, tool-rich environments.
Autonomy horizon
Measuring how far agents can act independently over long horizons.
Output complexity
Scoring sophisticated, verifiable, multi-artifact deliverables.
Speakers
Niko Grupen
Head of Applied Research
Harvey
Alex Ratner
CEO and Founder
Snorkel AI
Russell Yang
Applied Scientist II
Microsoft
Project Lead
JudgmentBench
Yiyou Sun
Postdoc
UC Berkeley
Project Lead
AgentLE

Steven Dillmann
Project Lead
Terminal-Bench-Science
Annas Bin Adil
CTO
Atella.ai
Co-Lead
STELLA-Bench
Gabe Orlanski
PhD Student
University of Wisconsin-Madison
Project Lead
SlopCodeBench
Nicholas Roberts
Postdoc Fellow
Princeton University
Project Lead
Train-to-Test (T²) Scaling Laws
Kelly Buchanan
Postdoctoral Researcher
Stanford University
Project Lead
Terminal Bench 2.1
Virginia Smith
Associate Professor
Carnegie Mellon University
Faculty Lead
CollusionBench
Vincent Sunn Chen
Founding Team
Snorkel AI
Shaping the frontier of AI
