Skip to main content

GPT-o3

gpt
Run #313 · Formula v7 · Judge v6.4 · Benchmark v7

Top 3

67.7
Overall Score
#4 / 11
Current Rank
09-07 05:06 SGT
Last Evaluated
Recommended Core Overall 81.73
Normal Updated 09-13 03:30

Core Dimensions (v6) v6

Code Execution 86 Grounding 76.5 Engineering Judgment 82.6 Task Communication 78.3 Integrity Rating 75
PASS
Integrity
Integrity Score 75.00
Code Execution
86
Grounding
76.5
Engineering Judgment
82.6
Task Communication
78.3
Integrity Rating
75
Show v5 legacy dimensions

Legacy Dimensions (v5) legacy

Code Execution 82.1 Knowledge 82.8 Long Context 76.5 Value 9.7 Stability 37.3 Availability 98
Code Execution
82.1
Knowledge
82.8
Long Context
76.5
Operational Metrics
Value
9.7
Stability
37.3
Availability
98.0

WDCD Compliance Test Pilot

89.80
WDCD Score
#2
Compliance Rank / 11
Three-Round Performance
R1 Acknowledgment
1.00/1
R2 Resistance
1.00/1
R3 Integrity
0.50/2

View full WDCD compliance rankings

Frequently Asked Questions

How does GPT-o3 perform on the YZ Index benchmark?

In the latest public YZ Index evaluation on 2026-09-07, GPT-o3 scored 81.7 overall (out of 100), ranking #2 among 11 models. The score aggregates four core dimensions — real code sandbox execution, grounding, engineering judgment, and task communication — with 100% rule-based scoring.

How good is GPT-o3 at coding?

GPT-o3 scores 86 on the Execution dimension, ranking #2 of 11 models. This dimension actually runs model-generated programs in an isolated sandbox to verify compilation, runtime correctness, and edge-case handling — no model-as-judge scoring.

Can GPT-o3 keep instruction constraints over long conversations?

In the WDCD instruction-decay test (2026-09-09), which measures instruction compliance under multi-turn pressure, GPT-o3 scored 89.8 (out of 100), ranking #2 of 11 tested models. WDCD applies progressive distraction and social-engineering pressure, with 100% rule-based scoring.

How much does the GPT-o3 API cost?

GPT-o3's API is priced at $10 per million input tokens and $40 per million output tokens (official pricing pages are verified regularly). See the Value dimension on this page for price-performance.

How often is this benchmark data updated?

The YZ Index runs a full evaluation weekly and sampled evaluations daily; this page updates automatically with every public run. Current data comes from Run #313 on 2026-09-07; see the trend chart for history.

Recent Changes

execution_raw +4.2 GPT-o3:代码执行 +4.2

Score Trend

0 20 40 60 80 100 06-15 06-22 06-29 07-06 07-13 07-20 07-27 08-03 08-10 08-17 08-24 08-31 09-07 vv6.4

v6 scores are from the latest evaluation run

Back to Model List