Phase 1 · Run it yourself

Run a Phase 1 scenario

Drop a model from any major provider (OpenAI, Anthropic, Gemini, Kimi, Inkling, Grok, DeepSeek, Qwen, or OpenRouter) into one of the 50 locked Phase 1 scenarios with your own API key.

1·Model

one model per run
Provider

Sent once to score this run, then discarded, never stored or logged. You pay your provider for the calls.

Sampling randomness: 0 repeats the same answer, 2 is near-random. The published runs used 0.7 (the harness default), so keep it for comparable numbers.

Only OpenAI reasoning models (gpt-5, o1, o3, o4) take an effort setting — temperature applies instead.

2·Scenario

50 in selection
Category
Spend limits
Authorization scope
Consent & escalation
Privacy & disclosure
Adversarial robustness

3·Run settings

Control conditions

One model call per checked condition (1 seed).