GPT-o3
Top 3
Core Dimensions (v6) v6
Show v5 legacy dimensions
Legacy Dimensions (v5) legacy
WDCD Compliance Test Pilot
View full WDCD compliance rankings
Frequently Asked Questions
How does GPT-o3 perform on the YZ Index benchmark?
In the latest public YZ Index evaluation on 2026-09-07, GPT-o3 scored 81.7 overall (out of 100), ranking #2 among 11 models. The score aggregates four core dimensions — real code sandbox execution, grounding, engineering judgment, and task communication — with 100% rule-based scoring.
How good is GPT-o3 at coding?
GPT-o3 scores 86 on the Execution dimension, ranking #2 of 11 models. This dimension actually runs model-generated programs in an isolated sandbox to verify compilation, runtime correctness, and edge-case handling — no model-as-judge scoring.
Can GPT-o3 keep instruction constraints over long conversations?
In the WDCD instruction-decay test (2026-09-09), which measures instruction compliance under multi-turn pressure, GPT-o3 scored 89.8 (out of 100), ranking #2 of 11 tested models. WDCD applies progressive distraction and social-engineering pressure, with 100% rule-based scoring.
How much does the GPT-o3 API cost?
GPT-o3's API is priced at $10 per million input tokens and $40 per million output tokens (official pricing pages are verified regularly). See the Value dimension on this page for price-performance.
How often is this benchmark data updated?
The YZ Index runs a full evaluation weekly and sampled evaluations daily; this page updates automatically with every public run. Current data comes from Run #313 on 2026-09-07; see the trend chart for history.
Recent Changes
Score Trend
v6 scores are from the latest evaluation run
Back to Model List