Yunze (Lorenzo) Xiao

MS in Language Technology · Carnegie Mellon University
My research asks how AI systems should represent individuals, simulate social dynamics, and govern interactions.
I am developing a research agenda on AI-mediated societies: how information boundaries and institutional rules shape interactions among people and AI agents.
- Private Personal Intelligence
- How can personal agents use memory, preferences, and context while keeping people in control of what they disclose?
- Experimental Social Simulation
- How can simulations grounded in observations of people help us compare interventions and evaluate their effects on collective behavior?
- Institutional Agent Safety
- How can permissions, incentives, audit mechanisms, and scalable oversight govern interacting agents and protect human agency?
Selected publications
All publications & preprintsMy work spans anthropomorphism and human-centered evaluation; persona and multi-agent social simulation; and the evidential foundations of AI welfare claims.
* Equal contribution.
Human-centered evaluation
How do we evaluate human-like AI?

-
Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design [P11]
We propose evaluating anthropomorphic cues against user goals and context, treating human-like design as something to assess rather than simply maximize.
2025. EMNLP 2025. Oral.
-
Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens [P7]
An audit of 20 cultural benchmarks identifies six recurring methodological pitfalls in how cultural knowledge and behavior are evaluated.
2026. Findings of EACL 2026.
-
Position: AI Welfare Is Bullshit [P5]
We argue that AI welfare claims need independent validation beyond steerable behavioral indicators, and that governance should prioritize verifiable harms.
2026. ICML 2026, Position Paper Track.
People and social simulation
When can simulated people tell us something about society?

-
The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models [W4]
High persona fidelity can coexist with a homogeneous simulated population: convincing individuals do not guarantee population diversity.
2026. arXiv:2604.24698.
Study overview · Co-first author.
-
Sentipolis: Emotion-Aware Agents for Social Simulations [P2]
Persistent emotion–memory coupling improves emotional continuity in the evaluated simulations; gains in believability depend on the model.
2026. Findings of ACL 2026.
Study overview · Co-first author.
Trustworthy multi-agent systems
How can we make multi-agent systems trustworthy?

-
TartanMaroon: Multi-Agent Academic Advising with Iterative Negotiation and Transparent Collaboration [P4]
Iterative proposal–critique negotiation improves constrained degree planning, with greater benefits on complex tasks than on simple factual queries.
2026. ACL 2026, System Demonstrations.
Study overview · Co-author; research mentor to Peidi Dong at Carnegie Mellon University Qatar.
Background
I am pursuing an MS in Language Technology at Carnegie Mellon University’s Language Technology Institute, advised by Prof. Mona Diab. I received my BS in Computer Science from Carnegie Mellon University Qatar in May 2025, with a minor in Computational Ethics, working with Prof. Houda Bouamor and Prof. Kemal Oflazer.
News
| Jul. 2026 | Our ACL 2026 papers include Sentipolis, TartanMaroon, The Confidence Dichotomy, and student difficulty estimation. |
|---|---|
| May 2026 | I presented Persona Collapse: Measuring Diversity Failures in LLM-Simulated Populations at the STAMINA-WG Research Talk Series. |
| Apr 26, 2026 | New blog post & interactive microsite — The Chameleon’s Limit summarizes our preprint on persona collapse in LLMs (explore the microsite). |
Contact
GHC 5418, Carnegie Mellon University
4902 Forbes Ave · Pittsburgh, PA 15213
Schedule a conversationView my calendar
Times shown in Eastern Time (Pittsburgh). You can also email me.