← All benchmarks
Coding benchmark · by Tangle VerticalBench

Standard Webhooks (Svix) Signatures

Two coding tasks that implement a verifier for the Standard Webhooks signature scheme used by Svix, which differs from the far more common Stripe scheme in signed content, secret decoding, and signature format.

Use this board to compare tested agent configurations on this exact task suite. It measures task execution against a hidden mock contract, not general model quality or production reliability.

How to read this benchmark

Each task is graded by a hidden test suite that signs fresh deliveries with its own random key each run and requires the verifier to accept authentic deliveries and reject tampered payloads, wrong keys, stale timestamps, and rotated multi-signature headers. Pass means the agent code executes correctly, never a model judging its own work. Every task is calibrated three ways before it is admitted: an empty solution fails, a current-scheme reference passes, and the Stripe-scheme solution fails on the intended trap. The prompts and calibration fixtures remain private to keep the test held out; this page publishes every scored attempt. The chart is the mean of the two task cells. Five profiles passed both tasks: Claude Sonnet with either reviewer setting, GLM-5.2 with either reviewer setting, and GLM-5.1 with no reviewer. Three profiles passed task 01 and failed task 02: the two pi harness profiles and GLM-5.1 with the default reviewer.

Last run
Held-out tasks
2
Configurations
8 agent configurations
Chart metric
Pass rate on a 0–100% scale

An agent profile is the exact instruction and tool configuration used for a run. A harness is the coding tool that gives a model its agent loop and tools; “direct” is a model call without that coding harness. On coding boards, “default reviewer” means the run used VerticalBench’s standard between-attempt review, while “no reviewer” means the coding harness ran once without that review. Hover a chart label to see its immutable profile ID. 95% Wilson interval for observed pass/fail outcomes; each configuration shows its own sample size. Small suites are narrow capability checks, not evidence that one model is generally better.

Read the broader agent-evaluation methodology →
95% CI
Pass rate, %
0
20
40
60
80
100
100.0
100.0
100.0
100.0
100.0
50.0
50.0
50.0
Claude Sonnetvia Claude CodeDefault reviewern=2
Claude Sonnetvia Claude CodeNo reviewern=2
GLM-5.2via OpenCodeDefault reviewern=2
GLM-5.1via OpenCodeNo reviewern=2
GLM-5.2via OpenCodeNo reviewern=2
GPT-5 minivia piDefault reviewern=2
GPT-5 minivia piNo reviewern=2
GLM-5.1via OpenCodeDefault reviewern=2

Same 2 held-out cases for every bar.

Per-task breakdown

2 held-out tasks · prompts private
Task GPT-5 mini via pi Default reviewer Claude Sonnet via Claude Code Default reviewer Claude Sonnet via Claude Code No reviewer GLM-5.2 via OpenCode Default reviewer GLM-5.1 via OpenCode No reviewer GLM-5.2 via OpenCode No reviewer GPT-5 mini via pi No reviewer GLM-5.1 via OpenCode Default reviewer
01 100 n=1 25094 tok · 188.8s 100 n=1 10610 tok · 201.2s 100 n=1 17733 tok · 267.3s 100 n=1 435636 tok · 189.9s 100 n=1 558881 tok · 342.8s 100 n=1 1106389 tok · 401.7s 100 n=1 20028 tok · 227.7s 100 n=1 458801 tok · 365.4s
02 0 n=1 22502 tok · 219.6s 100 n=1 22305 tok · 301.4s 100 n=1 31711 tok · 459.2s 100 n=1 747194 tok · 329.3s 100 n=1 704842 tok · 317.6s 100 n=1 693853 tok · 282.9s 0 n=1 20525 tok · 190.9s 0 n=1 0 tok · 0.5s

Each cell shows pass rate, total input-plus-output tokens, and mean wall time for that task and agent profile. Cost was not captured for these subscription-harness runs; that does not mean the runs were free. “Not run” means the result artifact has no recorded attempt for that task and profile. Task prompts stay private so they remain held out; opaque task numbers and every measured result are published.