ImageImage
A one-day, invite-only summit providing a first look at the benchmarks and research that will shape the frontier.

October 8, 2026

San Francisco

Image
Image
Image
Image
Image
STELLA-Bench
SciHarbor
Paperena
Image
ClinSafe
MedPAIR
FutureSim
PhilosophyBench
AgentAbstain
CollusionBench
HalluWorld
Train-to-Test (T²) Scaling Laws
HumanOversight Bench
JudgmentBench

Topics

Image
Environment complexity

Evaluating agents in realistic, dynamic, tool-rich environments.

Image

Autonomy horizon

Measuring how far agents can act independently over long horizons.

Image

Output complexity

Scoring sophisticated, verifiable, multi-artifact deliverables.

Speakers

Image

Niko Grupen

Head of Applied Research

Harvey

Image

Alex Ratner

CEO and Founder
Snorkel AI
Image

Russell Yang

Applied Scientist II
Microsoft
Project Lead

JudgmentBench

Image

Yiyou Sun

Postdoc
UC Berkeley
Project Lead
AgentLE
Image

Steven Dillmann

Project Lead
Terminal-Bench-Science
Image

Annas Bin Adil

CTO
Atella.ai
Co-Lead
STELLA-Bench
Image

Gabe Orlanski

PhD Student
University of Wisconsin-Madison
Project Lead
SlopCodeBench
Image

Nicholas Roberts

Postdoc Fellow

Princeton University
Project Lead
Train-to-Test (T²) Scaling Laws
Image

Kelly Buchanan

Postdoctoral Researcher
Stanford University
Project Lead
Terminal Bench 2.1
Image

Virginia Smith

Associate Professor
Carnegie Mellon University
Faculty Lead
CollusionBench
Image

Vincent Sunn Chen

Founding Team
Snorkel AI
Shaping the frontier of AI

Livestream

Stay updated on
event announcements