AI Coding Benchmarks

202 articles · Page 1 of 11
Which AI model writes the best code? HumanEval and MBPP are common benchmarks, but they only test function-level completion — far from real-world development. The YZ Index Execution dimension runs model-generated programs in isolated sandboxes, verifying compilation, runtime correctness, and edge-case handling. It is one of the few independent benchmarks using real code execution verification rather than model-as-judge scoring. This topic tracks coding capability rankings, programming tool updates, and AI-assisted development practices.

In-depth Guides

Review Claude Sonnet 4.6 Code Execution Plunges 27.4 Points, While Material Constraints Soar 37.7 Points
Claude Sonnet 4.6's Code Execution score fell 27.4 points in today's Smoke evaluation, while Material Constraints rose 37.7 points, leaving the main l
Oct 5, 2026
Review Grok 4 Code Execution Plunges 31.3 Points as Material Constraints Jumps 34.7
In today's Smoke evaluation, Grok 4's code execution score fell 31.3 points while its material constraints score rose 34.7 points, leaving the main le
Oct 5, 2026
Review GPT-6.1 Sol Main Leaderboard Plunges 9.2 Points, Code Execution Falls 16.7 Points in a Single Day
GPT-6.1 Sol's main leaderboard score dropped from 86.25 to 77.07 in today's Smoke evaluation, dragged down entirely by a 16.7-point single-day decline
Oct 3, 2026
Review Gemini 2.5 Pro Material Constraints Plunge 28 Points, Code Execution Rises to 100
In today's Smoke evaluation, Gemini 2.5 Pro's Material Constraints score fell from 100.00 to 72.00, while Code Execution rose from 71.90 to 100.00, li
Oct 3, 2026
Review GLM-4.6 Scores 100 on Material Constraints, 41.70 on Code Execution, Integrity Probe Only 15
In the 2026-10-02 YZ Index Smoke quick test, GLM-4.6 scored 67.94 on the main leaderboard, 41.70 on code execution, and 100.00 on Material Constraints
Oct 2, 2026
Google releases Gemini 4 Argon, called its most powerful model yet
Google has released its latest Gemini model, marketing it as a workhorse for coding and cybersecurity work.
Oct 1, 2026
Review GPT-5.5 Code Execution Plummets 22 Points While Material Constraints Surge 25.9 Points
In today's Smoke evaluation, GPT-5.5's code execution score fell from 72.00 to 50.00 while its material constraints score rose from 69.10 to 95.00, le
Oct 1, 2026
Review Qwen3 Max Code Execution Plunges 24 Points, Main Ranking Falls 7.4 Points
In today's Smoke evaluation, Qwen3 Max's code execution score dropped from 94.00 to 70.00, pulling the main ranking down from 84.51 to 77.16.
Oct 1, 2026
Review GPT-o3 Code Execution Plummets 21.8 Points; Material Constraints Rise 33.5 Points; Main Leaderboard Edges Up
In today's Smoke evaluation, GPT-o3's code execution score fell 21.8 points while its material constraints score rose 33.5 points, and its main leader
Sep 30, 2026
Review Claude Opus 4.7 Scores 72 in Code Execution, 85 in Material Constraints; Smoke Evaluation Main Leaderboard Slips 2.4 Points
In today's Smoke evaluation, Claude Opus 4.7's code execution score fell from 100.00 to 72.00 while material constraints rose from 56.00 to 85.00, pus
Sep 30, 2026
Review GPT-o3 Code Execution Plunges 24.5 Points, Main Leaderboard Falls 15.2: Random Draw or Genuine Degradation?
In today's Smoke evaluation, GPT-o3's main leaderboard score fell from 96.01 to 80.78, with code execution plunging 24.5 points. The article argues th
Sep 28, 2026
Review Claude Opus 4.7 Falls 13.9 Points on Smoke Benchmark Main Leaderboard, Code Execution Plunges 22 Points in One Day
Claude Opus 4.7 dropped from 96.01 to 82.16 on the Smoke benchmark main leaderboard today, with the code execution dimension falling 22 points in a si
Sep 28, 2026
Review GPT-5.5 Smoke Evaluation Main Leaderboard Plunges 12.5 Points as Code Execution Falls 28 in a Single Day
In today's Smoke evaluation, GPT-5.5's main leaderboard score fell 12.5 points to 78.80, driven by a 28-point single-day drop in the code execution di
Sep 27, 2026
Some Supabase customers are publicly exposing reams of people’s data to the web
The findings highlight how AI-generated and vibe-coded apps can spill and expose users' data when not configured or secured properly.
Sep 26, 2026
Lovable’s annualized revenue crosses $600M as vibe coding takes off
Lovable co-founder Fabian Hedin said that apps created on the platform are getting nearly a billion monthly views each month.
Sep 25, 2026
Grok 4.7 Released: 71% on DeepSWE and the Real Cost Behind the Half-Price Pay-as-You-Go Paradox
xAI released Grok 4.7 with 71.0% on DeepSWE v1.1 xHigh and unchanged $2/$6 per million tokens, but 125% higher output token consumption undermines the
Sep 23, 2026
Review GPT-o3 Code Execution Plunges 32.5 Points as Main Leaderboard Falls from 96.01 to 80.48
GPT-o3's main leaderboard score in today's Smoke evaluation dropped to 80.48 from 96.01 yesterday, driven by a 32.5-point collapse in code execution,
Sep 23, 2026
Review Grok 4 Code Execution Falls from 96.30 to 69.50; Main Leaderboard Drops 10 Points in a Single Day
In today's Smoke evaluation, Grok 4's main leaderboard score fell from 90.86 to 80.89, with code execution plunging from 96.30 to 69.50 while material
Sep 23, 2026
Review Grok 4 Material Constraint Score Plummets 15.8 Points, Main Leaderboard Slips from 95.44 to 90.86
In today's Smoke evaluation, Grok 4's material constraint score dropped from 100.00 to 84.20, a 15.8-point decline that pulled its main leaderboard sc
Sep 22, 2026
Review Claude Sonnet 4.6 Material Constraint Plummets 15.8 Points, Code Execution Rebounds 47 Points
In today's Smoke evaluation, Claude Sonnet 4.6's material constraint score dropped from 100.00 to 84.20, while code execution rose from 50.00 to 97.00
Sep 22, 2026