Skip to content

Benchmark Catalog

5 published experiments with full methodology documentation.

Filter Benchmarks

Coding medium

Python Bug Fix — Null Reference in Async Handler

Four frontier models attempt to fix a subtle async/await bug in a FastAPI endpoint. Evaluated on correctness, test pass rate, and minimal diff size.

Models
4
Top Score
96%
Leader
Claude 3.5 Sonnet
Published
June 12, 2025
Writing medium

Technical Blog Post — Kubernetes Networking Explainer

Models write a 1,200-word technical blog post explaining Kubernetes networking to intermediate developers. Evaluated on accuracy, structure, and clarity.

Models
3
Top Score
88%
Leader
Claude 3.5 Sonnet
Published
June 1, 2025