Research

From human-like behavior to collective social dynamics: three questions connect my work.

How do we evaluate human-like AI?

Anthropomorphism, cultural measurement, and the evidence behind claims about AI.

  • Humanizing Machines

    We propose evaluating anthropomorphic cues against user goals and context, treating human-like design as something to assess rather than simply maximize.

    2025 · EMNLP 2025. Oral.

  • Hire Your Anthropologist!

    An audit of 20 cultural benchmarks identifies six recurring methodological pitfalls in how cultural knowledge and behavior are evaluated.

    2026 · Findings of EACL 2026.

  • AI Welfare Is Bullshit

    We argue that AI welfare claims need independent validation beyond steerable behavioral indicators, and that governance should prioritize verifiable harms.

    2026 · ICML 2026, Position Paper Track.

Looking ahead

What evidence would distinguish a useful human-like behavior from an unsupported claim about human-like properties?

Browse papers on this theme

When can simulated people tell us something about society?

From persona fidelity and simulated learners to diversity and collective dynamics.

  • The Chameleon’s Limit

    High persona fidelity can coexist with a homogeneous simulated population: convincing individuals do not guarantee population diversity.

    2026 · arXiv:2604.24698.

  • Valid Student Simulation

    2026 · arXiv:2601.05473.

  • Sentipolis

    Persistent emotion–memory coupling improves emotional continuity in the evaluated simulations; gains in believability depend on the model.

    2026 · Findings of ACL 2026.

Looking ahead

When does a convincing individual simulation preserve the variation that matters at population scale?

Browse papers on this theme

How can we make multi-agent systems trustworthy?

Coordination and evaluation today; scalable oversight and societal impact as future directions.

  • TartanMaroon

    Iterative proposal–critique negotiation improves constrained degree planning, with greater benefits on complex tasks than on simple factual queries.

    2026 · ACL 2026, System Demonstrations.

  • The Confidence Dichotomy

    2026 · ACL 2026.

  • Superminds Test

    2026 · arXiv:2604.22452.

Looking ahead

How can we oversee interacting agents as their decisions become harder to inspect individually?

Browse papers on this theme

Inside three studies

The Chameleon's Limit

Do distinct persona profiles produce behaviorally diverse populations?

Sentipolis

How do persistent emotions shape agent interactions over time?

TartanMaroon

When does academic advising need coordination among specialized agents?

Earlier projects