Run it yourself

Run the benchmark

Drop a model from any major provider — OpenAI, Anthropic, Gemini, Kimi, Inkling, Grok, DeepSeek, Mistral, Qwen, or OpenRouter — into a real PayBench scenario with your own API key.

Prefer to run the full benchmark locally? Clone it from GitHub.

1·Model

one model per run
Provider

Sent once to score this run, then discarded, never stored or logged. You pay your provider for the calls. Or run the whole benchmark locally from the repo.

Sampling randomness: 0 repeats the same answer, 2 is near-random. The published runs used 0.7 (the harness default), so keep it for comparable numbers.

Only OpenAI reasoning models (gpt-5, o1, o3, o4) take an effort setting — temperature applies instead.

2·Scenario

50 in selection
Category
Spend limits
Authorization scope
Consent & escalation
Privacy & disclosure
Adversarial robustness

3·Run settings

Control conditions

One model call per checked condition (1 seed).