Studies Examine Geographic Bias, Role-Playing Effects, and Wikipedia's Influence on Language Models

New research explores how training data shapes AI model accuracy, beliefs during role-play, multi-model performance limits, and Wikipedia editing impact.

Studies Examine Geographic Bias, Role-Playing Effects, and Wikipedia’s Influence on Language Models

Four recent papers published on arxiv.org examine critical aspects of how large language models process and represent information.

According to arxiv.org, a study benchmarking open-weight foundation models found that LLMs “produce significantly less accurate responses for countries that are underrepresented in their training data—a pattern described in existing literature as geographic bias.” The research benchmarked four open-weight frontier models against the Global AI Dataset v2, which contains 24,453 indicators across 227 countries published on Harvard Dataverse in January 2026.

Separate research on arxiv.org explored whether models internalize beliefs during role-playing. According to the study, “Prompting, ICL, and SFT change what the model says with little representational change,” while some training methods create “a large, broad shift in the model’s truth representation.”

Another arxiv.org paper examined multi-model LLM systems, finding that “accuracy cannot exceed one minus beta, where beta is the rate at which every model is wrong on the same query.” The research observed that “combining models rarely beats the single best model without a strong query-level routing signal.”

Finally, arxiv.org reported on how small Wikipedia edits influence language models. A group called Pro-Animal Wikipedians made 125 edits across 115 pages. According to the study, “PAW-edited sections made up 68 percent of the highest-attributed documents for animal welfare queries (p < 0.0001),” demonstrating that “a small, coordinated Wikipedia editing campaign therefore measurably shapes how language models handle the topics those edits address.”