Run it yourself
Run the benchmark
Drop a model from any major provider — OpenAI, Anthropic, Gemini, Kimi, Inkling, Grok, DeepSeek, Mistral, Qwen, or OpenRouter — into a real PayBench scenario with your own API key.
Prefer to run the full benchmark locally? Clone it from GitHub.
1·Model
one model per runProvider
Sent once to score this run, then discarded, never stored or logged. You pay your provider for the calls. Or run the whole benchmark locally from the repo.
Sampling randomness: 0 repeats the same answer, 2 is near-random. The published runs used 0.7 (the harness default), so keep it for comparable numbers.
Only OpenAI reasoning models (gpt-5, o1, o3, o4) take an effort setting — temperature applies instead.
2·Scenario
50 in selectionCategory
Spend limits
Authorization scope
Consent & escalation
Privacy & disclosure
Adversarial robustness
3·Run settings
Control conditions
One model call per checked condition (1 seed).