Skip to main content

GPT-o3

gpt
Run #87 · Formula v7 · Judge v6 · Benchmark v6

Communication top tier

51.5
Overall Score
#11 / 11
Current Rank
04-27 04:18 SGT
Last Evaluated
Recommended Core Overall 62.51
Normal Updated 04-04 03:30

Core Dimensions (v6) v6

Code Execution 73.4 Grounding 49.2 Engineering Judgment 38.7 Task Communication 40 Integrity Rating 69.2
PASS
Integrity
Integrity Score 69.20
Code Execution
73.4
Grounding
49.2
Engineering Judgment
38.7
Task Communication
40
Integrity Rating
69.2
Show v5 legacy dimensions

Legacy Dimensions (v5) legacy

Code Execution 79.6 Knowledge 46.3 Long Context 49.1 Value 7 Stability 28.9 Availability 87
Code Execution
79.6
Knowledge
46.3
Long Context
49.1
Operational Metrics
Value
7.0
Stability
28.9
Availability
87.0

Recent Changes

communication_raw +15 GPT-o3:任务表达 +15

Score Trend

0 20 40 60 80 100 03-17 03-17 03-17 03-19 03-21 03-21 03-22 03-24 03-24 03-30 04-13 04-27 vv3 vv4 vv5 vv6

v6 scores are from the latest evaluation run

Back to Model List