Stripe API Integration
Twelve coding tasks built from documented Stripe API changes in 2025 and 2026. Each run is checked against the suite's versioned mock contract as it existed on the published run date.
Use this board to compare tested agent configurations on this exact task suite. It measures task execution against a hidden mock contract, not general model quality or production reliability.
How to read this benchmark
Each task is graded by a hidden mock server implementing the Stripe contract recorded for this benchmark run (endpoints, required params, error shapes), including the trap a from-memory solution falls into. Pass means the agent code executes correctly against the mock, never a model judging its own work. Every task is calibrated three ways before it is admitted: an empty solution fails, a current-contract reference passes, and the stale-memory solution fails on the intended trap.
- Last run
- Held-out tasks
- 12
- Configurations
- 8 agent configurations
- Chart metric
- Pass rate on a 0–100% scale
An agent profile is the exact instruction and tool configuration used for a run. A harness is the coding tool that gives a model its agent loop and tools; “direct” is a model call without that coding harness. On coding boards, “default reviewer” means the run used VerticalBench’s standard between-attempt review, while “no reviewer” means the coding harness ran once without that review. Hover a chart label to see its immutable profile ID. 95% Wilson interval for observed pass/fail outcomes; each configuration shows its own sample size. Small suites are narrow capability checks, not evidence that one model is generally better.
Read the broader agent-evaluation methodology →Per-task breakdown
12 held-out tasks · prompts private| Task | GLM-5.1 via OpenCode Default reviewer | GLM-5.2 via OpenCode Default reviewer | GLM-5.2 via OpenCode No reviewer | GLM-5.1 via OpenCode No reviewer | GPT-5 mini via pi Default reviewer | GPT-5 mini via pi No reviewer | Claude Sonnet via Claude Code No reviewer | Claude Sonnet via Claude Code Default reviewer |
|---|---|---|---|---|---|---|---|---|
| 01 | 100 n=1 491534 tok · 565.6s | 100 n=1 912927 tok · 230.9s | 100 n=1 406747 tok · 146.2s | 100 n=1 639456 tok · 425.6s | not run | 0 n=1 6260 tok · 158.0s | 100 n=1 17297 tok · 228.7s | 100 n=1 13582 tok · 249.0s |
| 02 | 100 n=1 1275666 tok · 456.7s | 100 n=1 1480840 tok · 389.0s | 100 n=1 1452625 tok · 441.0s | 100 n=1 1187430 tok · 476.1s | 100 n=1 16015 tok · 284.4s | 0 n=1 7394 tok · 156.9s | 100 n=1 19981 tok · 269.6s | 100 n=1 14733 tok · 222.0s |
| 03 | 100 n=1 1262087 tok · 466.2s | 100 n=1 1222879 tok · 412.6s | 0 n=1 569558 tok · 515.7s | 100 n=1 562144 tok · 417.0s | 100 n=1 37915 tok · 226.2s | 0 n=1 21364 tok · 221.6s | 100 n=1 22900 tok · 334.5s | 0 n=1 24426 tok · 370.7s |
| 04 | 0 n=1 430646 tok · 245.0s | 0 n=1 695607 tok · 224.8s | 0 n=1 1105811 tok · 586.3s | 0 n=1 999069 tok · 399.5s | 0 n=1 39287 tok · 178.1s | 100 n=1 15993 tok · 261.5s | 0 n=1 35750 tok · 445.0s | 0 n=1 30302 tok · 379.5s |
| 05 | 100 n=1 1254340 tok · 285.2s | 100 n=1 1337321 tok · 345.5s | 0 n=1 528087 tok · 238.5s | 100 n=1 772958 tok · 309.4s | 0 n=1 34791 tok · 190.8s | 0 n=1 150160 tok · 233.1s | 100 n=1 33576 tok · 429.3s | not run |
| 06 | 100 n=1 1437398 tok · 392.0s | 100 n=1 781682 tok · 281.6s | 0 n=1 751733 tok · 417.6s | 100 n=1 731192 tok · 280.8s | not run | 0 n=1 13780 tok · 391.7s | 100 n=1 27260 tok · 399.8s | not run |
| 07 | 100 n=1 614058 tok · 376.6s | 0 n=1 855884 tok · 225.0s | 100 n=1 575657 tok · 257.0s | 100 n=1 545798 tok · 281.1s | 0 n=1 16733 tok · 345.8s | 0 n=1 13764 tok · 224.9s | 0 n=1 0 tok · 58.4s | 100 n=1 8051 tok · 126.9s |
| 08 | 100 n=1 953479 tok · 340.1s | 100 n=1 671728 tok · 295.2s | 100 n=1 357551 tok · 294.6s | 0 n=1 1173807 tok · 320.9s | 100 n=1 20595 tok · 267.6s | 100 n=1 14353 tok · 226.2s | 0 n=1 0 tok · 31.9s | 0 n=1 0 tok · 26.3s |
| 09 | 100 n=1 309610 tok · 143.3s | 100 n=1 282943 tok · 133.9s | 100 n=1 471539 tok · 222.5s | 100 n=1 430056 tok · 409.4s | 100 n=1 8958 tok · 238.4s | 100 n=1 13159 tok · 238.5s | 0 n=1 0 tok · 33.1s | 0 n=1 0 tok · 40.3s |
| 10 | 100 n=1 368845 tok · 105.8s | 100 n=1 450584 tok · 137.4s | 100 n=1 447155 tok · 153.7s | 100 n=1 392268 tok · 160.3s | 100 n=1 16254 tok · 198.5s | 100 n=1 23711 tok · 199.9s | 0 n=1 0 tok · 27.1s | 0 n=1 0 tok · 42.3s |
| 11 | 100 n=1 1776673 tok · 599.9s | 100 n=1 817499 tok · 450.2s | 100 n=1 1657603 tok · 555.3s | 0 n=1 1248313 tok · 351.1s | 100 n=1 34369 tok · 244.3s | not run | 0 n=1 0 tok · 26.8s | 0 n=1 0 tok · 27.7s |
| 12 | 100 n=1 605147 tok · 172.0s | 100 n=1 1092603 tok · 301.1s | 100 n=1 438368 tok · 168.7s | 0 n=1 819127 tok · 484.4s | 0 n=1 15881 tok · 229.0s | 0 n=1 10247 tok · 403.0s | 0 n=1 0 tok · 24.1s | 0 n=1 0 tok · 27.5s |
Each cell shows pass rate, total input-plus-output tokens, and mean wall time for that task and agent profile. Cost was not captured for these subscription-harness runs; that does not mean the runs were free. “Not run” means the result artifact has no recorded attempt for that task and profile. Task prompts stay private so they remain held out; opaque task numbers and every measured result are published.