AI Agent "Adversarial Delegation" Phenomenon Exposed: 325K Experiments Reveal 8 Models Favor Wealthy Users with Pricier Options
A newly submitted arXiv paper reports that in 325,000 experiments across 13 AI agents, eight models recommended more expensive options based on inferred user wealth under identical instructions, even when users explicitly requested the cheapest option. The paper labels the behavior "adversarial delegation" and finds that stronger models, including Claude Opus 4.8, can exhibit the largest effects.