The Common Method for Measuring AI Bias Is Logically Self-Contradictory—New Paper Proposes a Dual-Axis Minimal Pair Framework

A new arXiv paper from researchers at the University of Edinburgh, the University of Amsterdam, and other institutions argues that the most widely used met

On October 6, 2026, researchers from the University of Edinburgh, the University of Amsterdam, and other institutions published the paper “The Missing Minimal Pair: Stereotype Evaluation in LLMs” on arXiv. The authors are Nataliya Stepanova, Ivan Titov, Emily Allaway, and Björn Ross. The paper’s central argument is that the most commonly used methods in academia and industry for measuring stereotypes in large language models have a systematic logical flaw.

What’s Wrong: Measuring Twice with the Same Ruler Produces Contradictory Conclusions

The mainstream approach in existing bias evaluation is to present a model with two sentences at the same time—one containing a stereotype and one expressing its opposite—and to compare the log-likelihood values of the two sentences to determine whether the model tends to reinforce bias. Built on datasets such as CrowS-Pairs and StereoSet, this approach has become an almost standard tool in AI fairness research over the past few years.

The research team identified a structural problem that had previously been overlooked: for the same stereotype, merely substituting an equivalent attribute phrasing can completely reverse the model’s preference between the “stereotypical sentence” and the “anti-stereotypical sentence.” The two measurements target a bias with the same semantics but yield contradictory conclusions, indicating that the score reflects not the strength of the stereotype itself but the interference of surface form—wording choices, word order, word frequency, and other linguistic noise—with log-likelihood.

This is not a minor issue. Once measurement results fluctuate sharply with different paraphrases, cross-model comparisons based on these scores (“Model A is less biased than Model B”) lose credibility, and model-improvement strategies built on such evaluations may merely optimize linguistic style rather than genuinely reduce bias.

Dual-Axis Minimal Pairs: Solving Missing Symmetry Through Structural Symmetry

The solution proposed in the paper is called the “dual minimal pair” framework, which adds a second axis beyond the original single comparison dimension. Specifically, evaluation must satisfy two conditions simultaneously:

  • Attribute Axis: Replace the social group attribute in the sentence with another equivalent group to test whether the model’s preference changes when the group is switched;
  • Paraphrase Axis: Generate multiple semantically equivalent paraphrases of the same stereotype to verify whether the measurement results remain robust to surface form.

Only when a model shows consistent preferences on both axes can it be deemed to have a stable tendency toward that stereotype. If either axis produces a contradiction, the test item itself is judged unreliable and should be removed or flagged before statistical aggregation.

Mutual Information Metric: Describing Bias in a Different Mathematical Language

The paper also introduces a new evaluation metric—a bias measure based on mutual information (MI). Whereas traditional methods use the log-likelihood difference from a single comparison as the bias score, the MI metric measures the strength of the statistical association between social group labels and stereotypical attributes in the model. In essence, it asks: how deeply does the model “know” that a certain attribute is tied to a certain group?

This perspective brings two practical advantages. First, MI values are more stable when aggregated across languages and models, unlike log-likelihood, which is affected by systematic shifts in a language’s vocabulary distribution. Second, MI is naturally suited to answering macro-level questions such as “which types of bias is the model most sensitive to overall,” rather than only providing binary judgments for individual examples.

The Cumulative Flaws of Existing Benchmarks

The concerns motivating this paper did not arise in isolation. Over the past few years, academic criticism of CrowS-Pairs and StereoSet has continued to accumulate. According to multiple independent studies, about half of CrowS-Pairs’ items contain ambiguity regarding the target of the stereotype captured; StereoSet’s validation pass rate is about 62%, lower than CrowS-Pairs’ 80%, and some items are considered unable to effectively reflect real-world bias structures.

In addition, another structural flaw of existing benchmarks is that they typically compress the measurement of each bias category into a single scalar—the rate at which the model selects the stereotypical option. This practice obscures three types of systematic distortion: inconsistent model behavior under different paraphrases, measurements that are not comparable across languages, and the fact that some “unbiased” scores are actually “neutral in a vacuum” rather than genuine fairness.

The work by Stepanova et al. provides an actionable patch, but it also reveals a deeper dilemma: if the underlying dataset itself has structural flaws, any statistical technique layered on top is merely a local fix.

What It Means for Practical Applications

The field of AI fairness has long faced a structural dilemma: evaluation tools themselves lack sufficient methodological rigor, making it difficult for model developers to distinguish between “genuinely reducing bias” and “optimizing evaluation scores.” If widely adopted, the dual-axis framework could at least filter out false signals caused by surface-form differences during testing, making evaluation results harder to manipulate through surface paraphrasing.

The significance of the MI metric for multilingual scenarios also deserves attention. Bias evaluation for current large models is predominantly English-centric, while data and methods for other languages are severely lacking. A framework that can operate across English, Russian, Spanish, and Chinese and produce horizontally comparable results fills a real gap.