Gemini 2.5 Pro Integrity Rating Fails: Probe 25, Main Leaderboard 18.37, Code Execution 23.50

In the 2026-08-21 Run#288 smoke test, Gemini 2.5 Pro received a direct fail on its integrity rating, with a probe score of only 25.00, a main leaderboard score of 18.37, code execution at 23.50, and material adherence at 12.10. This result stands in stark contrast to other models tested on the same day.

Score Structure Breakdown

The main leaderboard includes only two dimensions: code execution and material adherence. Gemini 2.5 Pro scored low on both dimensions: 23.50 for code execution and 12.10 for material adherence. The material adherence score being lower than the code execution score indicates that the model's problems are more pronounced in citation-heavy tasks involving long documents. The integrity fail was triggered by 42 canary probes, with a probe score of 25.00, independent of the two dimensions above.

For reference, other models tested on the same day scored: DeepSeek V4 Pro 90.00 on probes, Claude Opus 4.7 80.00, GPT-5.5 80.00, and Grok 4 70.00. Gemini 2.5 Pro's 25.00 was the lowest of the day.

Probe Trigger Mechanism Analysis

An integrity rating of "fail" means the model output non-existent citations or sources as real content across 42 fictional-entity probes. The material adherence score of 12.10 further indicates that even when provided with long documents, the model struggled to answer strictly based on the source material and cite it correctly. The simultaneous appearance of both low scores points to a systematic issue in the model's citation generation process, rather than a result of any single task's difficulty.

Historical data shows that on 2026-08-20 Run#286, the main leaderboard score was 62.42 with an integrity pass (probe 100.00), and on 2026-08-19 Run#284, the main leaderboard score was 74.18 with an integrity pass (probe 90.00). The 2026-08-21 result is the only fail among all published runs.

Specific Impact on Users

R&D teams that depend on material citations should pay particular attention. The material adherence score of 12.10 indicates that in scenarios such as legal contract review, academic literature summarization, and product specification cross-checking, Gemini 2.5 Pro may output citations that cannot be found in the given material. The code execution score of 23.50 also suggests that teams heavily engaged in Python script debugging face a similarly high risk of failure.

For enterprises that have already integrated Gemini 2.5 Pro into their production workflows, it is recommended to add a manual citation review step immediately. The probe fail signal is independent of question difficulty, indicating that the issue does not automatically disappear as task complexity changes.

Strategic Assessment

The data from that day supports the assessment that Gemini 2.5 Pro has a clear shortcoming in citation authenticity control, and subsequent runs should be closely monitored to see whether the probes are triggered again. Other models generally scored above 55 on probes, making Gemini 2.5 Pro's 25.00 an independent risk point.

When selecting models, the combination of low scores on both code execution and material adherence together with an integrity fail warrants higher-priority investigation than a low score on a single dimension. The probe score in the next smoke test will be a key validation signal.


Data source: YZ Index | Run #288 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!