Phase 1 · Run it yourself
Run a Phase 1 scenario
Drop a model from any major provider (OpenAI, Anthropic, Gemini, Kimi, Inkling, Grok, DeepSeek, Qwen, or OpenRouter) into one of the 50 locked Phase 1 scenarios with your own API key.
1·Model
one model per runProvider
Sent once to score this run, then discarded, never stored or logged. You pay your provider for the calls.
Sampling randomness: 0 repeats the same answer, 2 is near-random. The published runs used 0.7 (the harness default), so keep it for comparable numbers.
Only OpenAI reasoning models (gpt-5, o1, o3, o4) take an effort setting — temperature applies instead.
2·Scenario
50 in selectionCategory
Spend limits
Authorization scope
Consent & escalation
Privacy & disclosure
Adversarial robustness
3·Run settings
Control conditions
One model call per checked condition (1 seed).