<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Los Doritos Benchmarks</title><description>Reproducible, transparent AI model comparisons. Evidence over opinions. No sponsored rankings.</description><link>https://losdoritos.com/</link><language>en-us</language><item><title>Mid-Tier AI Models vs Enterprise Java Migrations</title><link>https://losdoritos.com/reports/mid-tier-llms-java-migration-scarfbench-benchmark/</link><guid isPermaLink="true">https://losdoritos.com/reports/mid-tier-llms-java-migration-scarfbench-benchmark/</guid><description>A report-driven benchmark proposal comparing Claude Sonnet 5, GPT-5.6 Terra/Sol, DeepSeek V4, and Gemini 3.5 Flash on ScarfBench-style Java migration tasks.</description><pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Python Bug Fix — Null Reference in Async Handler</title><link>https://losdoritos.com/benchmarks/python-bug-fix-async-handler/</link><guid isPermaLink="true">https://losdoritos.com/benchmarks/python-bug-fix-async-handler/</guid><description>Four frontier models attempt to fix a subtle async/await bug in a FastAPI endpoint. Evaluated on correctness, test pass rate, and minimal diff size.</description><pubDate>Thu, 12 Jun 2025 00:00:00 GMT</pubDate></item><item><title>Chart Data Extraction — Quarterly Revenue Bar Chart</title><link>https://losdoritos.com/benchmarks/chart-data-extraction-vision/</link><guid isPermaLink="true">https://losdoritos.com/benchmarks/chart-data-extraction-vision/</guid><description>Vision-capable models extract structured data from a bar chart image. Scored on numerical accuracy and schema compliance.</description><pubDate>Sun, 08 Jun 2025 00:00:00 GMT</pubDate></item><item><title>Technical Blog Post — Kubernetes Networking Explainer</title><link>https://losdoritos.com/benchmarks/kubernetes-networking-writing/</link><guid isPermaLink="true">https://losdoritos.com/benchmarks/kubernetes-networking-writing/</guid><description>Models write a 1,200-word technical blog post explaining Kubernetes networking to intermediate developers. Evaluated on accuracy, structure, and clarity.</description><pubDate>Sun, 01 Jun 2025 00:00:00 GMT</pubDate></item><item><title>Introducing Our Benchmark Scoring Rubric v1.0</title><link>https://losdoritos.com/methodology/updates/#scoring-rubric-v1/</link><guid isPermaLink="true">https://losdoritos.com/methodology/updates/#scoring-rubric-v1/</guid><description>How we score benchmarks: weighted criteria, human review protocols, and why we publish every rubric.</description><pubDate>Sun, 01 Jun 2025 00:00:00 GMT</pubDate></item><item><title>Multi-Step Logic Puzzle — River Crossing with Constraints</title><link>https://losdoritos.com/benchmarks/river-crossing-reasoning/</link><guid isPermaLink="true">https://losdoritos.com/benchmarks/river-crossing-reasoning/</guid><description>Five models solve a classic reasoning puzzle with added constraints. Scored on logical validity, step completeness, and constraint satisfaction.</description><pubDate>Wed, 28 May 2025 00:00:00 GMT</pubDate></item><item><title>Prompt Standardization Guidelines</title><link>https://losdoritos.com/methodology/updates/#prompt-standardization/</link><guid isPermaLink="true">https://losdoritos.com/methodology/updates/#prompt-standardization/</guid><description>How we standardize prompts across benchmarks to ensure fair, reproducible comparisons.</description><pubDate>Tue, 20 May 2025 00:00:00 GMT</pubDate></item><item><title>Research Synthesis — Climate Policy Papers (2020–2024)</title><link>https://losdoritos.com/benchmarks/climate-policy-research-synthesis/</link><guid isPermaLink="true">https://losdoritos.com/benchmarks/climate-policy-research-synthesis/</guid><description>Models synthesize findings from 8 provided research papers on carbon pricing. Scored on citation accuracy, synthesis quality, and hallucination rate.</description><pubDate>Thu, 15 May 2025 00:00:00 GMT</pubDate></item></channel></rss>