AI Coding Benchmarks

154 articles · Page 1 of 8
Which AI model writes the best code? HumanEval and MBPP are common benchmarks, but they only test function-level completion — far from real-world development. The YZ Index Execution dimension runs model-generated programs in isolated sandboxes, verifying compilation, runtime correctness, and edge-case handling. It is one of the few independent benchmarks using real code execution verification rather than model-as-judge scoring. This topic tracks coding capability rankings, programming tool updates, and AI-assisted development practices.

In-depth Guides

Review Gemini 2.5 Pro Integrity Rating Fails: Probe 25, Main Leaderboard 18.37, Code Execution 23.50
Gemini 2.5 Pro received a failing integrity rating in the 2026-08-21 Run#288 smoke test, with a probe score of only 25.00, a main leaderboard score of
Aug 21, 2026
Review GPT-5.5 Smoke Evaluation Main Leaderboard Plummets 10.1 Points; Material Constraints Drop 16.3 in Single Day
GPT-5.5's main leaderboard score fell from 98.35 to 88.27 in today's Smoke evaluation, a 10.1-point decline driven primarily by the material constrain
Aug 20, 2026
Review Claude Opus 4.7 Plunges 8.2 Points on Main Leaderboard; Material Constraint Drops 12.1 in a Single Day
Claude Opus 4.7's main leaderboard score fell 8.2 points to 90.16 in today's Smoke evaluation, driven mainly by a 12.1-point single-day plunge in mate
Aug 20, 2026
Cursor capitalizes on Github frustration, launches rival hosting platform
Cursor, known for its AI Code Editor, is launching a new code hosting platform to rival developers' long preferred favorite, Github.
Aug 19, 2026
Review Doubao Pro Material Constraint Drops 40.9 Points in a Single Day; Code Execution Up 25 Points; Main Leaderboard Slips 4.7
Doubao Pro's material constraint score fell from 90.90 to 50.00 in today's Smoke evaluation, while code execution rose from 75.00 to 100.00, dragging
Aug 19, 2026
Review Doubao Pro Main Leaderboard Plummets 12.6 Points, Code Execution Drops 25 Points in a Single Day
Doubao Pro's main leaderboard score in today's Smoke evaluation fell from 94.74 to 82.16, with the code execution dimension dropping from 100.00 to 75
Aug 18, 2026
Review DeepSeek V4 Pro Code Execution Plunges 41.7 Points; Main Leaderboard Down 19.2 in a Day
DeepSeek V4 Pro's main leaderboard score dropped from 70.44 to 51.24 in today's Smoke evaluation, a 19.2-point decline driven primarily by the code ex
Aug 17, 2026
Review Claude Sonnet 4.6 Code Execution Drops from 100 to 75 Points, Main Leaderboard Falls 5.5 Points
In today's Smoke evaluation, Claude Sonnet 4.6's code execution score fell from 100.00 to 75.00, a 25-point decline, which pulled the main leaderboard
Aug 16, 2026
Review Claude Opus 4.7 Scores 75.00 in Code Execution, 100.00 in Material Constraint, Main Leaderboard Rises 3.7
In today's Smoke evaluation, Claude Opus 4.7's code execution score dropped from 99.60 to 75.00, while its material constraint score jumped from 61.70
Aug 16, 2026
SpaceX officially closes its Cursor acquisition
AI coding startup Cursor is now officially a part of SpaceX.
Aug 16, 2026
AI coding startup Cognition reportedly already in talks to raise at $40B valuation
Cognition may be looking to raise another mega round just a few months after raising $1 billion at a $26 billion valuation.
Aug 13, 2026
Review Gemini 2.5 Pro Drops 9.1 Points on Main Leaderboard; Code Execution Plunges 29.6 in Single Day
Gemini 2.5 Pro's main leaderboard score fell 9.1 points to 79.41 in today's Smoke evaluation, driven by a 29.6-point plunge in the code execution dime
Aug 13, 2026
Lovable confirms new $13.3B valuation, raises another $400M
This new funding comes after Lovable hit $500 million in annualized run rate revenue in June, the startup told TechCrunch.
Aug 13, 2026
Review GLM-4.6 Code Execution Drops 12.5 Points, Perfect Material Constraint Score Lifts Main Leaderboard by 5.4
GLM-4.6 scored 79.38 on today's Smoke evaluation main leaderboard, with code execution falling 12.5 points to 62.50 while material constraints reached
Aug 12, 2026
Anthropic is turning Claude Code’s auto mode on by default
Programming with Claude Code will soon require even less human oversight.
Aug 10, 2026
Review Claude Sonnet 4.6 Code Execution Plunges 19.5 Points While Leaderboard Score Rises 13.8 Points
In today's Smoke evaluation, Claude Sonnet 4.6's code execution score dropped from 94.50 to 75.00, a decline of 19.5 points, while its material constr
Aug 8, 2026
Review Claude Opus 4.7 Code Execution Plunges 30.5 Points, Main Leaderboard Drops Only 6.4 Points
Claude Opus 4.7's code execution score fell from 100.00 to 69.50 in today's Smoke evaluation, while the main leaderboard dropped only 6.4 points. The
Aug 7, 2026
Review Grok 4 Material Constraint Drops 17.6 Points; Main Leaderboard Falls Just 1.8 Points
Grok 4's material constraint score fell from 82.60 to 65.00 in today's Smoke evaluation, while its overall main leaderboard score only declined from 8
Aug 7, 2026
Cloudflare open-sources vibe-coding platform for people who aren't coders
Cloudflare built an AI agent workspace for its employees. Now it’s open source.
Aug 7, 2026
Review GLM-4.6 Smoke Evaluation: Main Leaderboard Score 74, Code Execution 82.3, Material Constraint 95, API Failure Leaves Dimensions Missing
GLM-4.6 scored 74.00 on the main leaderboard in today's Smoke evaluation, with 82.30 on code execution and 95.00 on material constraint. Two dimension
Aug 2, 2026