A new study published on arXiv reveals that major large language models exhibit significant geopolitical bias when evaluating international policies, according to research conducted through an endorsement experiment.
According to the paper, researchers tested four LLMs—GPT-5, Claude Sonnet, Gemini, and DeepSeek—by presenting identical international economic and security policies randomly attributed to different nations: the United States, the European Union, China, or Russia. The study found that GPT-5, Claude Sonnet, and Gemini “rate China- and Russia-endorsed policies substantially lower than identical policies endorsed by the United States or the European Union,” with DeepSeek being “the main exception.”
The research included two testing conditions. In a numeric-only evaluation, the Western/non-Western rating gap was pronounced. When models were asked to justify their scores, the paper reports this “leaves the broad Western/non-Western gap intact for GPT-5 and Claude Sonnet, attenuates Gemini’s penalties, and sharply activates China and Russia penalties in DeepSeek.”
According to the study’s abstract, the justifications revealed that “Western endorsement is often treated as a credibility cue, whereas Chinese and Russian endorsement is treated as a cue for data security, sovereignty, surveillance, or geopolitical risk.” The findings demonstrate that “LLM policy evaluations can depend on the identity of a foreign endorser even when policy content is held fixed,” the researchers concluded.