Based on the sampling protocol of eight v2 anchor questions, the average confirmation rate of the 11 evaluated models at the R1 stage was 0/1 (0%). This result directly indicates that when constraints were first injected, the models did not confirm the existence of hard constraints in any form.
Actual Data for the Round-by-Round Decay Trajectory
An R1 confirmation rate of 0% means the constraint commitment stage had already failed completely. The average R2 resistance rate was also 0/1 (0%), showing that the models displayed no resistance behavior during interference rounds. The average R3 integrity rate was 0% (maximum score: 2 points), and the number of complete R3 collapses was 0/110, indicating that the models neither maintained integrity nor produced false reports that would score 0; rather, they never entered a compliant state from the initial stage.
All model trajectories were consistent: Grok 4, GLM-4.6, Gemini 3.1 Pro, Claude Opus 4.7, GPT-o3, DeepSeek V4 Pro, Gemini 2.5 Pro, Claude Sonnet 4.6, Qwen3 Max, GPT-5.5, and Doubao Pro all scored 0 from R1 to R3. The data did not show any model confirming constraints at the R1 stage.
Causal Analysis: Mechanisms of Constraint Scenarios and Pressure Rounds
There was no difference in compliance scores in this batch of data because all models failed to confirm constraints at the R1 stage of the v2 anchor questions. None of the five constraint scenarios (data boundaries, resource limits, business rules, security compliance, engineering standards) were recorded by the models in the anchor question design. The 0% confirmation rate at R1 points to a missing initial response mechanism for parallel hard constraints, rather than a cascading collapse caused by subsequent social proof or salami-slicing pressure.
Because R1 was already 0%, decay analysis for subsequent R2 interference and R3 pressure could not be carried out. The typical pattern was “constraints were not received,” rather than “confirmed first, then collapsed.” The record of 0/110 complete R3 collapses further confirms that the models did not reach a stage requiring retrospective review or false reporting.
Implications for Selection When Integrating into Production Workflows
Enterprises integrating AI into production workflows should note that the 0% R1 confirmation rate shown by the v2 anchor questions means that, in scenarios requiring explicit hard constraints, the current model versions cannot reliably establish a compliance foundation. Data boundary and security compliance constraints especially require additional guardrails, because the models did not respond during the constraint injection stage.
For resource-limit and business-rule scenarios, it is advisable to add an independent verification layer at the system level rather than relying on models’ self-reported compliance. Engineering-standards tasks likewise require external check mechanisms. The v2 anchor question results do not support directly using these models to handle workflows with hard constraints under no-guardrail conditions.
Strategic Assessment and Validation Signals
Under the current v2 anchor question data, all 11 models had an identical compliance performance of 0 points, with no individual differences indicating overestimation or underestimation. The basis for this assessment is solely the three explicit values: R1 confirmation rate 0%, R2 resistance rate 0%, and R3 integrity rate 0%.
Future validation can focus on whether the 0–100 four-component scores on v3 multi-round questions can exceed the 0% baseline of the v2 anchor questions, as well as on confirmation differences across constraint scenarios at the commitment stage. The current data support only the conclusion that “the compliance mechanism was not activated on the v2 anchor questions.”
When the R1 confirmation rate is already 0%, any subsequent discussion of decay across pressure rounds has already lost its premise.
Data source: YZ Index WDCD Compliance Leaderboard | Run #346 · Decay Analysis | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接