Publications

Human–AI interaction, human-centered evaluation, and multi-agent social simulation.

30 papers · 19 published · 11 preprints · 1,616 citations · Google Scholar

Google Scholar profile total · Updated 2026-09-10 · Individual counts verified for 30/30 listed papers.

Download all BibTeX · CV (PDF)

* Equal contribution in author lists. Citation counts retain Google Scholar’s own * markers.

2026

19 papers · Published work and preprints

  1. Weihao Xuan, Qingcheng Zeng, Heli Qi, Yunze Xiao, Junjue Wang, and Naoto Yokoya.

    2026 · ACL 2026. [P1]

    12 citations · Google Scholar (2026-09-10)

    Research figure: The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents
    BibTeX
    @inproceedings{2026-acl-long-520,
      title = {{The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents}},
      author = {Weihao Xuan and Qingcheng Zeng and Heli Qi and Yunze Xiao and Junjue Wang and Naoto Yokoya},
      year = {2026},
      booktitle = {ACL 2026},
      url = {https://aclanthology.org/2026.acl-long.520/}
    }
    
    Download .bib
  2. Chiyuan Fu*, Lyuhao Chen*, Yunze Xiao*, Weihao Xuan, Carlos Busso, and Mona Diab.

    Persistent emotion–memory coupling improves emotional continuity in the evaluated simulations; gains in believability depend on the model.

    2026 · Findings of ACL 2026. [P2]

    2 citations · Google Scholar (2026-09-10)

    Research figure: Sentipolis: Emotion-Aware Agents for Social Simulations
    BibTeX
    @inproceedings{2026-findings-acl-368,
      title = {{Sentipolis: Emotion-Aware Agents for Social Simulations}},
      author = {Chiyuan Fu and Lyuhao Chen and Yunze Xiao and Weihao Xuan and Carlos Busso and Mona Diab},
      year = {2026},
      booktitle = {Findings of ACL 2026},
      url = {https://aclanthology.org/2026.findings-acl.368/}
    }
    
    Download .bib
  3. Ming Li*, Han Chen*, Yunze Xiao, Jian Chen, Hong Jiao, and Tianyi Zhou.

    2026 · Findings of ACL 2026. [P3]

    12 citations · Google Scholar (2026-09-10)

    BibTeX
    @inproceedings{2026-findings-acl-1270,
      title = {{Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction}},
      author = {Ming Li and Han Chen and Yunze Xiao and Jian Chen and Hong Jiao and Tianyi Zhou},
      year = {2026},
      booktitle = {Findings of ACL 2026},
      url = {https://aclanthology.org/2026.findings-acl.1270/}
    }
    
    Download .bib
  4. Peidi Dong, Houda Bouamor, Yunze Xiao, and Devi G Kurup.

    Iterative proposal–critique negotiation improves constrained degree planning, with greater benefits on complex tasks than on simple factual queries.

    2026 · ACL 2026, System Demonstrations. [P4]

    0 citations · Google Scholar (2026-09-10)

    Research figure: TartanMaroon: Multi-Agent Academic Advising with Iterative Negotiation and Transparent Collaboration
    BibTeX
    @inproceedings{2026-acl-demo-83,
      title = {{TartanMaroon: Multi-Agent Academic Advising with Iterative Negotiation and Transparent Collaboration}},
      author = {Peidi Dong and Houda Bouamor and Yunze Xiao and Devi G Kurup},
      year = {2026},
      booktitle = {ACL 2026, System Demonstrations},
      url = {https://aclanthology.org/2026.acl-demo.83/}
    }
    
    Download .bib
  5. Yunze Xiao*, Gordon Dai*, Shahan Ali Memon*, Jen-tse Huang, Maarten Sap, and Mona T. Diab.

    We argue that AI welfare claims need independent validation beyond steerable behavioral indicators, and that governance should prioritize verifiable harms.

    2026 · ICML 2026, Position Paper Track. [P5]

    1 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{ai-welfare,
      title = {{Position: AI Welfare Is Bullshit}},
      author = {Yunze Xiao and Gordon Dai and Shahan Ali Memon and Jen-tse Huang and Maarten Sap and Mona T. Diab},
      year = {2026},
      howpublished = {ICML 2026, Position Paper Track},
      url = {https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6574439}
    }
    
    Download .bib
  6. Yunze Xiao*, Tingyu He*, Lionel Z. Wang*, Yiming Ma, Xingyu Song, Xiaohang Xu, Mona Diab, Irene Li, and Ka Chung Ng.

    2026 · EACL 2026. Oral. [P6]

    6 citations · Google Scholar (2026-09-10)

    BibTeX
    @inproceedings{2026-eacl-long-23,
      title = {{JiraiBench: A Cross-lingual Benchmark for Evaluating Large Language Models' Detection of Human Risky Health Behavior Content in Jirai Community}},
      author = {Yunze Xiao and Tingyu He and Lionel Z. Wang and Yiming Ma and Xingyu Song and Xiaohang Xu and Mona Diab and Irene Li and Ka Chung Ng},
      year = {2026},
      booktitle = {EACL 2026. Oral},
      url = {https://aclanthology.org/2026.eacl-long.23/}
    }
    
    Download .bib
  7. Mai Alkhamissi*, Yunze Xiao*, Badr AlKhamissi, and Mona Diab.

    An audit of 20 cultural benchmarks identifies six recurring methodological pitfalls in how cultural knowledge and behavior are evaluated.

    2026 · Findings of EACL 2026. [P7]

    9 citations · Google Scholar (2026-09-10)

    Research figure: Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens
    BibTeX
    @inproceedings{culture,
      title = {{Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens}},
      author = {Mai Alkhamissi and Yunze Xiao and Badr AlKhamissi and Mona Diab},
      year = {2026},
      booktitle = {Findings of EACL 2026},
      url = {https://aclanthology.org/2026.findings-eacl.63/}
    }
    
    Download .bib
  8. Lynnette Hui Xian Ng*, Yunze Xiao*, Lionel Z. Wang, Weihao Xuan, and Mona Diab.

    2026 · Journal of Ambient Intelligence and Humanized Computing, 17, 1567–1578. [P8]

    0 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{2607-13924,
      title = {{ExpressionCueLens: a cross-cultural analysis of human-AI companion conversations on social media}},
      author = {Lynnette Hui Xian Ng and Yunze Xiao and Lionel Z. Wang and Weihao Xuan and Mona Diab},
      year = {2026},
      howpublished = {Journal of Ambient Intelligence and Humanized Computing, 17, 1567–1578},
      url = {https://link.springer.com/article/10.1007/s12652-026-05109-z}
    }
    
    Download .bib
  9. Shu Yang, Junchao Wu, Xin Chen, Yunze Xiao, Xinyi Yang, Derek F. Wong, and Di Wang.

    2026 · TACL. [P9]

    47 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{2504-02956,
      title = {{Understanding Aha Moments: From External Observations to Internal Mechanisms}},
      author = {Shu Yang and Junchao Wu and Xin Chen and Yunze Xiao and Xinyi Yang and Derek F. Wong and Di Wang},
      year = {2026},
      howpublished = {TACL},
      url = {https://arxiv.org/abs/2504.02956}
    }
    
    Download .bib
  10. Center for AI Safety, Scale AI, and HLE Contributors Consortium (including Yunze Xiao).

    2026 · (Humanity's Last Exam). Nature, 649, 1139–1146. [P10]

    869* citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{10-1038-s41586-025-09962-4,
      title = {{A benchmark of expert-level academic questions to assess AI capabilities}},
      author = {Center for AI Safety and Scale AI and HLE Contributors Consortium (including Yunze Xiao)},
      year = {2026},
      howpublished = {(Humanity's Last Exam). Nature, 649, 1139–1146},
      url = {https://www.nature.com/articles/s41586-025-09962-4}
    }
    
    Download .bib
  11. Haokai Zhao, Yunze Xiao, Weihao Xuan, Flora Salim, Benjamin Tag, and Aditya Joshi.

    2026 · arXiv:2608.11528. [W1]

    0 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{2608-11528,
      title = {{Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment}},
      author = {Haokai Zhao and Yunze Xiao and Weihao Xuan and Flora Salim and Benjamin Tag and Aditya Joshi},
      year = {2026},
      howpublished = {arXiv:2608.11528},
      url = {https://arxiv.org/abs/2608.11528}
    }
    
    Download .bib
  12. Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja, Jiayi Geng, Yunze Xiao, Daniel Lee, Aditya Bharat Soni, Vincent Lo, Xiang Yue, and Graham Neubig.

    2026 · arXiv:2607.02032. [W2]

    0 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{2607-02032,
      title = {{PACE: A Proxy for Agentic Capability Evaluation}},
      author = {Yueqi Song and Lintang Sutawika and Jiarui Liu and Lindia Tjuatja and Jiayi Geng and Yunze Xiao and Daniel Lee and Aditya Bharat Soni and Vincent Lo and Xiang Yue and Graham Neubig},
      year = {2026},
      howpublished = {  arXiv:2607.02032},
      url = {https://arxiv.org/abs/2607.02032}
    }
    
    Download .bib
  13. Weijia Zhang, Ruiqi Chen, Yunze Xiao, and Weihao Xuan.

    2026 · arXiv:2606.11232. [W3]

    1 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{2606-11232,
      title = {{Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs}},
      author = {Weijia Zhang and Ruiqi Chen and Yunze Xiao and Weihao Xuan},
      year = {2026},
      howpublished = {arXiv:2606.11232},
      url = {https://arxiv.org/abs/2606.11232}
    }
    
    Download .bib
  14. Yunze Xiao*, Vivienne J. Zhang*, Chenghao Yang, Ningshan Ma, Weihao Xuan, and Jen-tse Huang.

    High persona fidelity can coexist with a homogeneous simulated population: convincing individuals do not guarantee population diversity.

    2026 · arXiv:2604.24698. [W4]

    3 citations · Google Scholar (2026-09-10)

    Research figure: The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models
    BibTeX
    @misc{chameleon,
      title = {{The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models}},
      author = {Yunze Xiao and Vivienne J. Zhang and Chenghao Yang and Ningshan Ma and Weihao Xuan and Jen-tse Huang},
      year = {2026},
      howpublished = {  arXiv:2604.24698},
      url = {https://arxiv.org/abs/2604.24698}
    }
    
    Download .bib
  15. Xirui Li, Ming Li, Yunze Xiao, Ryan Wong, Dianqi Li, Timothy Baldwin, and Tianyi Zhou.

    2026 · arXiv:2604.22452. [W5]

    2 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{2604-22452,
      title = {{Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents}},
      author = {Xirui Li and Ming Li and Yunze Xiao and Ryan Wong and Dianqi Li and Timothy Baldwin and Tianyi Zhou},
      year = {2026},
      howpublished = {arXiv:2604.22452},
      url = {https://arxiv.org/abs/2604.22452}
    }
    
    Download .bib
  16. Yunze Xiao, Wenkai Li, Xiaoyuan Wu, Ningshan Ma, Yueqi Song, and Weihao Xuan.

    2026 · arXiv:2604.06409. [W6]

    0 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{2604-06409,
      title = {{Say Something Else: Rethinking Contextual Privacy as Information Sufficiency}},
      author = {Yunze Xiao and Wenkai Li and Xiaoyuan Wu and Ningshan Ma and Yueqi Song and Weihao Xuan},
      year = {2026},
      howpublished = {arXiv:2604.06409},
      url = {https://arxiv.org/abs/2604.06409}
    }
    
    Download .bib
  17. Xuan Liu, Haoyang Shang, Zizhang Liu, Xinyan Liu, Yunze Xiao, Yiwen Tu, and Haojian Jin.

    2026 · arXiv:2602.00685. [W7]

    4 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{2602-00685,
      title = {{HumanStudy-Bench: Towards AI Agent Design for Participant Simulation}},
      author = {Xuan Liu and Haoyang Shang and Zizhang Liu and Xinyan Liu and Yunze Xiao and Yiwen Tu and Haojian Jin},
      year = {2026},
      howpublished = {arXiv:2602.00685},
      url = {https://arxiv.org/abs/2602.00685}
    }
    
    Download .bib
  18. Zhihao Yuan*, Yunze Xiao*, Ming Li*, Weihao Xuan, Richard Tong, Mona Diab, and Tom Mitchell.

    2026 · arXiv:2601.05473. [W8]

    13 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{student,
      title = {{Towards Valid Student Simulation with Large Language Models}},
      author = {Zhihao Yuan and Yunze Xiao and Ming Li and Weihao Xuan and Richard Tong and Mona Diab and Tom Mitchell},
      year = {2026},
      howpublished = {arXiv:2601.05473},
      url = {https://arxiv.org/abs/2601.05473}
    }
    
    Download .bib
  19. Rui Yang, Huitao Li, Weihao Xuan, Heli Qi, Xin Li, Kunyu Yu, Yingjian Chen, Rongrong Wang, Jacques Behmoaras, Tianxi Cai, Bibhas Chakraborty, Qingyu Chen, Lionel Tim-Ee Cheng, Marie-Louise Damwanza, Chido Dzinotyiwei, Aosong Feng, Chuan Hong, Yusuke Iwasawa, Yuhe Ke, Linah Kitala, Taehoon Ko, Jisan Lee, Irene Li, Jonathan Chong Kai Liew, Hongfang Liu, Lian Leng Low, Edison Marrese-Taylor, Yutaka Matsuo, Isheanesu Misi, Yilin Ning, Jasmine Chiat Ling Ong, Marcus Eng Hock Ong, Enrico Petretto, Hossein Rouhizadeh, Abiram Sandralegar, Oren Schreier, Iain Bee Huat Tan, Patrick Tan, Daniel Shu Wei Ting, Junjue Wang, Chunhua Weng, Matthew Yu Heng Wong, Fang Wu, Yunze Xiao, Xuhai Xu, Qingcheng Zeng, Zhuo Zheng, Yifan Peng, Douglas Teodoro, and Nan Liu.

    2026 · arXiv:2601.02186. [W9]

    5 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{2601-02186,
      title = {{Toward Global Large Language Models in Medicine}},
      author = {Rui Yang and Huitao Li and Weihao Xuan and Heli Qi and Xin Li and Kunyu Yu and Yingjian Chen and Rongrong Wang and Jacques Behmoaras and Tianxi Cai and Bibhas Chakraborty and Qingyu Chen and Lionel Tim-Ee Cheng and Marie-Louise Damwanza and Chido Dzinotyiwei and Aosong Feng and Chuan Hong and Yusuke Iwasawa and Yuhe Ke and Linah Kitala and Taehoon Ko and Jisan Lee and Irene Li and Jonathan Chong Kai Liew and Hongfang Liu and Lian Leng Low and Edison Marrese-Taylor and Yutaka Matsuo and Isheanesu Misi and Yilin Ning and Jasmine Chiat Ling Ong and Marcus Eng Hock Ong and Enrico Petretto and Hossein Rouhizadeh and Abiram Sandralegar and Oren Schreier and Iain Bee Huat Tan and Patrick Tan and Daniel Shu Wei Ting and Junjue Wang and Chunhua Weng and Matthew Yu Heng Wong and Fang Wu and Yunze Xiao and Xuhai Xu and Qingcheng Zeng and Zhuo Zheng and Yifan Peng and Douglas Teodoro and Nan Liu},
      year = {2026},
      howpublished = {arXiv:2601.02186},
      url = {https://arxiv.org/abs/2601.02186}
    }
    
    Download .bib

2025

4 papers · Published work and preprints

  1. Yunze Xiao*, Lynnette Hui Xian Ng*, Jiarui Liu, and Mona Diab.

    We propose evaluating anthropomorphic cues against user goals and context, treating human-like design as something to assess rather than simply maximize.

    2025 · EMNLP 2025. Oral. [P11]

    35 citations · Google Scholar (2026-09-10)

    Research figure: Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design
    BibTeX
    @inproceedings{humanizing,
      title = {{Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design}},
      author = {Yunze Xiao and Lynnette Hui Xian Ng and Jiarui Liu and Mona Diab},
      year = {2025},
      booktitle = {EMNLP 2025. Oral},
      url = {https://aclanthology.org/2025.emnlp-main.164/}
    }
    
    Download .bib
    Abstract

    Large Language Models (LLMs) increasingly exhibit anthropomorphism characteristics – human-like qualities portrayed across their outlook, language, behavior, and reasoning functions. Such characteristics enable more intuitive and engaging human-AI interactions. However, current research on anthropomorphism remains predominantly risk-focused, emphasizing over-trust and user deception while offering limited design guidance. We argue that anthropomorphism should instead be treated as a concept of design that can be intentionally tuned to support user goals. Drawing from multiple disciplines, we propose that the anthropomorphism of an LLM-based artifact should reflect the interaction between artifact designers and interpreters. This interaction is facilitated by cues embedded in the artifact by the designers and the (cognitive) responses of the interpreters to the cues. Cues are categorized into four dimensions: perceptive, linguistic, behavioral, and cognitive. By analyzing the manifestation and effectiveness of each cue, we provide a unified taxonomy with actionable levers for practitioners. Consequently, we advocate for function-oriented evaluations of anthropomorphic design.

  2. Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, and Irene Li.

    2025 · EMNLP 2025. [P12]

    134 citations · Google Scholar (2026-09-10)

    BibTeX
    @inproceedings{2025-emnlp-main-79,
      title = {{MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation}},
      author = {Weihao Xuan and Rui Yang and Heli Qi and Qingcheng Zeng and Yunze Xiao and Aosong Feng and Dairui Liu and Yun Xing and Junjue Wang and Fan Gao and Jinghui Lu and Yuang Jiang and Huitao Li and Xin Li and Kunyu Yu and Ruihai Dong and Shangding Gu and Yuekang Li and Xiaofei Xie and Felix Juefei-Xu and Foutse Khomh and Osamu Yoshie and Qingyu Chen and Douglas Teodoro and Nan Liu and Randy Goebel and Lei Ma and Edison Marrese-Taylor and Shijian Lu and Yusuke Iwasawa and Yutaka Matsuo and Irene Li},
      year = {2025},
      booktitle = {EMNLP 2025},
      url = {https://aclanthology.org/2025.emnlp-main.79/}
    }
    
    Download .bib
    Abstract

    Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities. This dual limitation makes it challenging to assess LLMs’ performance in the multilingual setting comprehensively. To fill this gap, we introduce MMLU-ProX, a comprehensive benchmark covering 29 languages, built on an English benchmark. Each language version consists of 11,829 identical questions, enabling direct cross-lingual comparisons. Additionally, to meet efficient evaluation needs, we provide a lite version containing 658 questions per language. To ensure the high quality of MMLU-ProX, we employ a rigorous development process that involves multiple powerful LLMs for translation, followed by expert review to ensure accurate expression, consistent terminology, and cultural relevance. Building on this, we systematically evaluate 36 state-of-the-art LLMs, including reasoning-enhanced and multilingual-optimized LLMs. The results reveal significant disparities in the multilingual capabilities of LLMs: While they perform well in high-resource languages, their performance declines markedly in low-resource languages, particularly for African languages. Through MMLU-ProX, we aim to advance the development of more inclusive AI systems and promote equitable access to technology across global contexts.

  3. Jiarui Liu, Yueqi Song, Yunze Xiao, Mingqian Zheng, Lindia Tjuatja, Jana Schaich Borg, Mona Diab, and Maarten Sap.

    2025 · EMNLP 2025. [P13]

    28 citations · Google Scholar (2026-09-10)

    BibTeX
    @inproceedings{2025-emnlp-main-831,
      title = {{Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics}},
      author = {Jiarui Liu and Yueqi Song and Yunze Xiao and Mingqian Zheng and Lindia Tjuatja and Jana Schaich Borg and Mona Diab and Maarten Sap},
      year = {2025},
      booktitle = {EMNLP 2025},
      url = {https://aclanthology.org/2025.emnlp-main.831/}
    }
    
    Download .bib
    Abstract

    As large language models (LLMs) are increasingly used in morally sensitive domains, it is crucial to understand how persona traits affect their moral reasoning and persuasive behavior. We present the first large-scale study of multi-dimensional persona effects in AI-AI debates over real-world moral dilemmas. Using a 6-dimensional persona space (age, gender, country, social class, ideology, and personality), we simulate structured debates between AI agents over 131 relationship-based cases. Our results show that personas affect initial moral stances and debate outcomes, with political ideology and personality traits exerting the strongest influence. Persuasive success varies across traits, with liberal and open personalities reaching higher consensus. While logit-based confidence grows during debates, emotional and credibility-based appeals diminish, indicating more tempered argumentation over time. These trends mirror findings from psychology and cultural studies, reinforcing the need for persona-aware evaluation frameworks for AI moral reasoning.

  4. Gordon Dai* and Yunze Xiao*.

    2025 · NeurIPS 2025, Position Paper Track. [P14]

    9 citations · Google Scholar (2026-09-10)

    BibTeX
    @inproceedings{2505-18139,
      title = {{Embracing Contradiction: Theoretical Inconsistency Will Not Impede the Road of Building Responsible AI Systems}},
      author = {Gordon Dai and Yunze Xiao},
      year = {2025},
      booktitle = {NeurIPS 2025, Position Paper Track},
      url = {https://proceedings.neurips.cc/paper_files/paper/2025/file/dd4a4bc7a7ba0b197b679c0025cb2df8-Paper-Position_Paper_Track.pdf}
    }
    
    Download .bib

2024

5 papers · Published work and preprints

  1. Yunze Xiao*, Yujia Hu*, Kenny Tsu Wei Choo, and Roy Ka-Wei Lee.

    2024 · EMNLP 2024. [P15]

    59 citations · Google Scholar (2026-09-10)

    BibTeX
    @inproceedings{toxicloak,
      title = {{ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations}},
      author = {Yunze Xiao and Yujia Hu and Kenny Tsu Wei Choo and Roy Ka-Wei Lee},
      year = {2024},
      booktitle = {EMNLP 2024},
      url = {https://aclanthology.org/2024.emnlp-main.345/}
    }
    
    Download .bib
    Abstract

    Detecting hate speech and offensive language is essential for maintaining a safe and respectful digital environment. This study examines the limitations of state-of-the-art large language models (LLMs) in identifying offensive content within systematically perturbed data, with a focus on Chinese, a language particularly susceptible to such perturbations. We introduce ToxiCloakCN, an enhanced dataset derived from ToxiCN, augmented with homophonic substitutions and emoji transformations, to test the robustness of LLMs against these cloaking perturbations. Our findings reveal that existing models significantly underperform in detecting offensive content when these perturbations are applied. We provide an in-depth analysis of how different types of offensive content are affected by these perturbations and explore the alignment between human and model explanations of offensiveness. Our work highlights the urgent need for more advanced techniques in offensive language detection to combat the evolving tactics used to evade detection mechanisms.

  2. Xintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, and Yanghua Xiao.

    2024 · ACL 2024. [P16]

    308* citations · Google Scholar (2026-09-10)

    BibTeX
    @inproceedings{2024-acl-long-102,
      title = {{InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews}},
      author = {Xintao Wang and Yunze Xiao and Jen-tse Huang and Siyu Yuan and Rui Xu and Haoran Guo and Quan Tu and Yaying Fei and Ziang Leng and Wei Wang and Jiangjie Chen and Cheng Li and Yanghua Xiao},
      year = {2024},
      booktitle = {ACL 2024},
      url = {https://aclanthology.org/2024.acl-long.102/}
    }
    
    Download .bib
    Abstract

    Role-playing agents (RPAs), powered by large language models, have emerged as a flourishing field of applications. However, a key challenge lies in assessing whether RPAs accurately reproduce the personas of target characters, namely their character fidelity. Existing methods mainly focus on the knowledge and linguistic patterns of characters. This paper, instead, introduces a novel perspective to evaluate the personality fidelity of RPAs with psychological scales. Overcoming drawbacks of previous self-report assessments on RPAs, we propose InCharacter, namely **In**terviewing **Character** agents for personality tests. Experiments include various types of RPAs and LLMs, covering 32 distinct characters on 14 widely used psychological scales. The results validate the effectiveness of InCharacter in measuring RPA personalities. Then, with InCharacter, we show that state-of-the-art RPAs exhibit personalities highly aligned with the human-perceived personalities of the characters, achieving an accuracy up to 80.7%.

  3. David R. Mortensen*, Valentina Izrailevitch*, Yunze Xiao, Hinrich Schütze, and Leonie Weissweiler.

    2024 · LREC-COLING 2024. [P17]

    6 citations · Google Scholar (2026-09-10)

    BibTeX
    @inproceedings{2024-lrec-main-1508,
      title = {{Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs}},
      author = {David R. Mortensen and Valentina Izrailevitch and Yunze Xiao and Hinrich Schütze and Leonie Weissweiler},
      year = {2024},
      booktitle = {LREC-COLING 2024},
      url = {https://aclanthology.org/2024.lrec-main.1508/}
    }
    
    Download .bib
    Abstract

    Lexical-syntactic flexibility, in the form of conversion (or zero-derivation) is a hallmark of English morphology. In conversion, a word with one part of speech is placed in a non-prototypical context, where it is coerced to behave as if it had a different part of speech. However, while this process affects a large part of the English lexicon, little work has been done to establish the degree to which language models capture this type of generalization. This paper reports the first study on the behavior of large language models with reference to conversion. We design a task for testing lexical-syntactic flexibility—the degree to which models can generalize over words in a construction with a non-prototypical part of speech. This task is situated within a natural language inference paradigm. We test the abilities of five language models—two proprietary models (GPT-3.5 and GPT-4), three open source model (Mistral 7B, Falcon 40B, and Llama 2 70B). We find that GPT-4 performs best on the task, followed by GPT-3.5, but that the open source language models are also able to perform it and that the 7-billion parameter Mistral displays as little difference between its baseline performance on the natural language inference task and the non-prototypical syntactic category task, as the massive GPT-4.

  4. Qingyang Wu, Ying Xu, Tingsong Xiao, Yunze Xiao, Yitong Li, Tianyang Wang, Yichi Zhang, Shanghai Zhong, Yuwei Zhang, Wei Lu, and Yifan Yang.

    2024 · arXiv:2404.13885. [W10]

    22 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{2404-13885,
      title = {{Surveying Attitudinal Alignment Between Large Language Models Vs. Humans Towards 17 Sustainable Development Goals}},
      author = {Qingyang Wu and Ying Xu and Tingsong Xiao and Yunze Xiao and Yitong Li and Tianyang Wang and Yichi Zhang and Shanghai Zhong and Yuwei Zhang and Wei Lu and Yifan Yang},
      year = {2024},
      howpublished = {arXiv:2404.13885},
      url = {https://arxiv.org/abs/2404.13885}
    }
    
    Download .bib
  5. Yunze Xiao, Houda Bouamor, and Wajdi Zaghouani.

    2024 · arXiv:2403.18314. [W11]

    17 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{survey,
      title = {{Chinese Offensive Language Detection: Current Status and Future Directions}},
      author = {Yunze Xiao and Houda Bouamor and Wajdi Zaghouani},
      year = {2024},
      howpublished = {arXiv:2403.18314},
      url = {https://arxiv.org/abs/2403.18314}
    }
    
    Download .bib

2023

1 papers · Published work and preprints

  1. Yunze Xiao and Firoj Alam.

    2023 · ArabicNLP 2023. [P18]

    4 citations · Google Scholar (2026-09-10)

    BibTeX
    @inproceedings{nexus,
      title = {{Nexus at ArAIEval Shared Task: Fine-Tuning Arabic Language Models for Propaganda and Disinformation Detection}},
      author = {Yunze Xiao and Firoj Alam},
      year = {2023},
      booktitle = {ArabicNLP 2023},
      url = {https://aclanthology.org/2023.arabicnlp-1.58/}
    }
    
    Download .bib

2022

1 papers · Published work and preprints

  1. Yunze Xiao.

    2022 · ICCRD 2022, 167–170. [P19]

    7 citations · Google Scholar (2026-09-10)

    BibTeX
    @misc{10-1109-ICCRD54409-2022-9730454,
      title = {{A Transformer-based Attention Flow Model for Intelligent Question and Answering Chatbot}},
      author = {Yunze Xiao},
      year = {2022},
      howpublished = {ICCRD 2022, 167–170},
      url = {https://doi.org/10.1109/ICCRD54409.2022.9730454}
    }
    
    Download .bib