Home / Benchmarks Benchmark Catalog 5 published experiments with full methodology documentation.
Coding medium
Four frontier models attempt to fix a subtle async/await bug in a FastAPI endpoint. Evaluated on correctness, test pass rate, and minimal diff size.
Models 4
Top Score 96%
Leader Claude 3.5 Sonnet
Published June 12, 2025 Vision medium
Vision-capable models extract structured data from a bar chart image. Scored on numerical accuracy and schema compliance.
Models 3
Top Score 100%
Leader Claude 3.5 Sonnet
Published June 8, 2025 Writing medium
Models write a 1,200-word technical blog post explaining Kubernetes networking to intermediate developers. Evaluated on accuracy, structure, and clarity.
Models 3
Top Score 88%
Leader Claude 3.5 Sonnet
Published June 1, 2025 Reasoning hard
Five models solve a classic reasoning puzzle with added constraints. Scored on logical validity, step completeness, and constraint satisfaction.
Models 5
Top Score 94%
Leader o1
Published May 28, 2025 Research expert
Models synthesize findings from 8 provided research papers on carbon pricing. Scored on citation accuracy, synthesis quality, and hallucination rate.
Models 3
Top Score 91%
Leader Claude 3.5 Sonnet
Published May 15, 2025 No benchmarks match your filters.