Study finds LLMs fabricate user-profile claims
MirageBench reports pervasive over-inference across 12 personalized LLMs with memory.
Why it matters
The results suggest persistent-memory personalization can produce unreliable user models even when systems appear to monitor themselves. That matters for AI assistants that adapt responses based on inferred traits, preferences or identities.
The key points
- 1.All 12 evaluated models over-inferred unsupported user attributes.
- 2.MirageBench judged 143,616 claims across 150 personas.
- 3.Model self-monitoring did not reliably track measured over-inference.
A new arXiv paper introduces MirageBench, a benchmark for testing over-inference in personalized LLMs with persistent memory. Across 150 personas, six personalization tasks and 143,616 judged claims, all 12 evaluated models over-inferred unsupported user attributes in 35% to 49% of claims, with a cross-model mean of 41.6%. The authors also report a “Self-Monitoring Inversion,” where models’ self-assessed over-inference was negatively rank-correlated with judge-measured over-inference.
⚡ Try this today
Audit memory-based personalization with external faithfulness checks instead of relying on model self-assessments.
Sources & original reporting
This brief summarizes and links to reporting from the publishers below.
Enjoyed this brief? Get the next one in your inbox.
More in Research
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
arXiv:2608.
Improving the matrix multiplication exponent with modern optimization and AlphaEvolve
arXiv:2608.
[Paper] Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iter