YZ Index
Weekly Report
Weekly model performance changes and trend analysis.
Baseline: Run #333 · Formula v7 · Judge v6.4 · Benchmark v7 · 2026-09-21 05:05 SGT
Current: Run #342 · Formula v7 · Judge v6.4 · Benchmark v7 · 2026-09-28 05:02 SGT
Overall Score Changes Ranked by absolute change magnitude
Doubao Pro
+9.5
71.0 → 80.5
DeepSeek V4 Pro
+5.8
69.4 → 75.2
Grok 4
+5.1
78.2 → 83.3
Gemini 3.1 Pro
+3.0
72.5 → 75.5
Gemini 2.5 Pro
+2.1
71.9 → 74.0
GPT-5.5
+1.8
76.4 → 78.2
Claude Opus 4.7
+1.1
82.8 → 83.9
GLM-4.6
-32.5
57.8 → 25.4
Qwen3 Max
-4.6
73.3 → 68.7
GPT-o3
-3.6
81.3 → 77.7
Claude Sonnet 4.6
-1.0
77.3 → 76.4
Side Dimension Changes Communication and Judgment changes
Grok 4
+12.1
Judgment: 55.6 → 67.7
Gemini 2.5 Pro
+9.6
Judgment: 63.3 → 72.9
Claude Sonnet 4.6
+7.4
Judgment: 83.6 → 91.0
Gemini 2.5 Pro
+2.5
Communication: 65.8 → 68.3
Gemini 3.1 Pro
+2.2
Judgment: 77.8 → 80.0
Qwen3 Max
+1.7
Judgment: 40.0 → 41.7
GLM-4.6
-47.9
Judgment: 75.0 → 27.1
GLM-4.6
-41.3
Communication: 78.8 → 37.5
Claude Sonnet 4.6
-12.5
Communication: 87.4 → 74.9
DeepSeek V4 Pro
-7.5
Communication: 73.6 → 66.1
Grok 4
-7.1
Communication: 87.9 → 80.8
Claude Opus 4.7
-5.0
Communication: 70.3 → 65.3
DeepSeek V4 Pro
-5.0
Judgment: 77.8 → 72.8
Claude Opus 4.7
-3.9
Judgment: 87.2 → 83.3
Qwen3 Max
-2.5
Communication: 50.3 → 47.8
GPT-o3
-0.9
Judgment: 82.4 → 81.5
Operational Signal Changes Stability and Availability changes
Doubao Pro
+13.6
Stability: 35.0 → 48.6
Grok 4
+12.5
Stability: 25.9 → 38.4
DeepSeek V4 Pro
+11.5
Stability: 40.4 → 51.9
GPT-5.5
+5.9
Stability: 33.2 → 39.1
GLM-4.6
+4.2
Stability: 40.1 → 44.3
Gemini 2.5 Pro
+2.8
Stability: 35.2 → 38.0
DeepSeek V4 Pro
+2.4
Value: 42.1 → 44.5
Grok 4
+2.3
Value: 24.1 → 26.4
Doubao Pro
+1.9
Value: 93.6 → 95.5
Gemini 2.5 Pro
+1.7
Value: 36.3 → 38.0
Doubao Pro
+1.0
Availability: 97.9 → 98.9
Gemini 3.1 Pro
+0.9
Value: 24.9 → 25.8
GPT-5.5
+0.6
Value: 18.8 → 19.4
GLM-4.6
-38.8
Availability: 80.0 → 41.2
GLM-4.6
-17.0
Value: 36.0 → 19.0
DeepSeek V4 Pro
-2.5
Availability: 89.5 → 87.0
GPT-o3
-2.0
Availability: 99.0 → 97.0
GPT-o3
-1.7
Stability: 37.3 → 35.6
Claude Sonnet 4.6
-1.6
Stability: 36.7 → 35.1
Qwen3 Max
-1.4
Stability: 23.8 → 22.4
Qwen3 Max
-1.1
Value: 46.4 → 45.3
Gemini 3.1 Pro
-0.8
Stability: 28.3 → 27.5
Show legacy dimension changes
4
Up
7
Down
0
Stable
11
models
Significant Increases
Significant Decreases
GLM-4.6
-23.8
GLM-4.6: Engineering Judgment -23.8
judgment_raw
Gemini 2.5 Pro
-8.1
Gemini 2.5 Pro: Code Execution -8.1
execution_raw
Claude Sonnet 4.6
-7.9
Claude Sonnet 4.6: Code Execution -7.9
execution_raw
Claude Opus 4.7
-6.4
Claude Opus 4.7: Code Execution -6.4
execution_raw
GPT-o3
-4.4
GPT-o3: Code Execution -4.4
execution_raw
DeepSeek V4 Pro
-2.9
DeepSeek V4 Pro: Grounding -2.9
grounding_raw
GPT-5.5
-2.2
GPT-5.5: Code Execution -2.2
execution_raw