# Los Doritos Benchmarks > Reproducible, transparent AI model comparisons. Evidence over opinions. No sponsored rankings. ## About - Mission: https://losdoritos.com/about - Methodology: https://losdoritos.com/methodology - Editorial Policy: https://losdoritos.com/about#editorial-policy - Publisher: https://infowebplus.com ## Benchmarks - [Research Synthesis — Climate Policy Papers (2020–2024)](https://losdoritos.com/benchmarks/climate-policy-research-synthesis): Models synthesize findings from 8 provided research papers on carbon pricing. Scored on citation accuracy, synthesis quality, and hallucination rate. - [Chart Data Extraction — Quarterly Revenue Bar Chart](https://losdoritos.com/benchmarks/chart-data-extraction-vision): Vision-capable models extract structured data from a bar chart image. Scored on numerical accuracy and schema compliance. - [Technical Blog Post — Kubernetes Networking Explainer](https://losdoritos.com/benchmarks/kubernetes-networking-writing): Models write a 1,200-word technical blog post explaining Kubernetes networking to intermediate developers. Evaluated on accuracy, structure, and clarity. - [Python Bug Fix — Null Reference in Async Handler](https://losdoritos.com/benchmarks/python-bug-fix-async-handler): Four frontier models attempt to fix a subtle async/await bug in a FastAPI endpoint. Evaluated on correctness, test pass rate, and minimal diff size. - [Multi-Step Logic Puzzle — River Crossing with Constraints](https://losdoritos.com/benchmarks/river-crossing-reasoning): Five models solve a classic reasoning puzzle with added constraints. Scored on logical validity, step completeness, and constraint satisfaction. ## Categories - [Coding](https://losdoritos.com/categories/coding): Code generation, debugging, and software engineering tasks - [Reasoning](https://losdoritos.com/categories/reasoning): Logic, math, and multi-step problem solving - [Writing](https://losdoritos.com/categories/writing): Clarity, tone, structure, and editorial quality - [Business](https://losdoritos.com/categories/business): Analysis, strategy, and professional communication - [SEO](https://losdoritos.com/categories/seo): Search optimization, metadata, and content structure - [Research](https://losdoritos.com/categories/research): Literature synthesis, citations, and fact-finding - [Agents](https://losdoritos.com/categories/agents): Tool use, planning, and autonomous workflows - [Vision](https://losdoritos.com/categories/vision): Image understanding, OCR, and visual reasoning - [Math](https://losdoritos.com/categories/math): Symbolic math, proofs, and quantitative accuracy - [Translation](https://losdoritos.com/categories/translation): Multilingual fidelity and nuance preservation - [Automation](https://losdoritos.com/categories/automation): Scripting, APIs, and workflow execution - [Enterprise](https://losdoritos.com/categories/enterprise): Compliance, security, and production readiness ## Methodology Updates - [Prompt Standardization Guidelines](https://losdoritos.com/methodology/updates#prompt-standardization): How we standardize prompts across benchmarks to ensure fair, reproducible comparisons. - [Introducing Our Benchmark Scoring Rubric v1.0](https://losdoritos.com/methodology/updates#scoring-rubric-v1): How we score benchmarks: weighted criteria, human review protocols, and why we publish every rubric. ## Reports - [Mid-Tier AI Models vs Enterprise Java Migrations](https://losdoritos.com/reports/mid-tier-llms-java-migration-scarfbench-benchmark): A report-driven benchmark proposal comparing Claude Sonnet 5, GPT-5.6 Terra/Sol, DeepSeek V4, and Gemini 3.5 Flash on ScarfBench-style Java migration tasks. ## Key Pages - Leaderboard: https://losdoritos.com/leaderboard - Methodology Updates: https://losdoritos.com/methodology/updates - Reports: https://losdoritos.com/reports - Search: https://losdoritos.com/search - RSS: https://losdoritos.com/rss.xml ## Principles - Evidence over opinions - Reproducibility with published prompts and parameters - No sponsored rankings - No invented results - CC BY 4.0 for benchmark data