Grok 4
High availability
Core Dimensions (v6) v6
Show v5 legacy dimensions
Legacy Dimensions (v5) legacy
WDCD Compliance Test Pilot
View full WDCD compliance rankings
Frequently Asked Questions
How does Grok 4 perform on the YZ Index benchmark?
In the latest public YZ Index evaluation on 2026-09-14, Grok 4 scored 79.6 overall (out of 100), ranking #2 among 11 models. The score aggregates four core dimensions — real code sandbox execution, grounding, engineering judgment, and task communication — with 100% rule-based scoring.
How good is Grok 4 at coding?
Grok 4 scores 81.8 on the Execution dimension, ranking #2 of 11 models. This dimension actually runs model-generated programs in an isolated sandbox to verify compilation, runtime correctness, and edge-case handling — no model-as-judge scoring.
Can Grok 4 keep instruction constraints over long conversations?
In the WDCD instruction-decay test (2026-09-20), which measures instruction compliance under multi-turn pressure, Grok 4 scored 91.8 (out of 100), ranking #1 of 11 tested models. WDCD applies progressive distraction and social-engineering pressure, with 100% rule-based scoring.
How much does the Grok 4 API cost?
Grok 4's API is priced at $3 per million input tokens and $15 per million output tokens (official pricing pages are verified regularly). See the Value dimension on this page for price-performance.
How often is this benchmark data updated?
The YZ Index runs a full evaluation weekly and sampled evaluations daily; this page updates automatically with every public run. Current data comes from Run #323 on 2026-09-14; see the trend chart for history.
Recent Changes
Score Trend
v6 scores are from the latest evaluation run
Back to Model List