Hongchuan Zeng

dblp:375/0990 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 38% Trustworthy machine learning · 36% Knowledge representation and reasoning · 20%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
1.012026
XToM: Exploring the Multilingual Theory of Mind for Large Language Models · ACL (1) 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
theory of mind
1.012026
XToM: Exploring the Multilingual Theory of Mind for Large Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning › language model interpretability
attention head analysis
0.912025
Heads up! Large Language Models Can Perform Tasks Without Your Instruction via Selective Attention Head Masking · ICML 2025
Natural language and speech › Language models and text generation
instruction following
0.912025
Heads up! Large Language Models Can Perform Tasks Without Your Instruction via Selective Attention Head Masking · ICML 2025
Machine learning › Trustworthy machine learning
language model interpretability
0.912025
Heads up! Large Language Models Can Perform Tasks Without Your Instruction via Selective Attention Head Masking · ICML 2025
Machine learning › Deep learning architectures and training
transformer
0.312025
Heads up! Large Language Models Can Perform Tasks Without Your Instruction via Selective Attention Head Masking · ICML 2025

Methods — techniques the papers use, named apart from their topics

multilingual benchmarking · 1.0inference-time intervention · 0.9attention head masking · 0.9
YearPublicationVenuePosition
2026 XToM: Exploring the Multilingual Theory of Mind for Large Language Models
abstract
Theory of Mind (ToM), the ability to infer mental states in others, is pivotal for human social cognition. Existing evaluations of ToM in LLMs are largely limited to English, neglecting the linguistic diversity that shapes human cognition. This limitation raises a critical question: can LLMs exhibit Multilingual Theory of Mind, which is the capacity to reason about mental states across diverse linguistic contexts? To address this gap, we present XToM, a rigorously validated multilingual benchmark that evaluates ToM across five languages and incorporates diverse, contextually rich task scenarios. Using XToM, we systematically evaluate LLMs (e.g., DeepSeek R1), revealing a pronounced dissonance: while models excel in multilingual language understanding, their ToM performance varies across languages. Our findings expose limitations in LLMs' ability to replicate human-like mentalizing across linguistic contexts.
Chunkit Chan, Yauwai Yim, Hongchuan Zeng, Zhiying Zou, Xinyuan Cheng, Zhifan Sun, Zheye Deng, Kawai Chung, Yuzhuo Ao, Yixiang Fan, Cheng Jiayang, Ercong Nie, Ginny Y. Wong, Helmut Schmid, Hinrich Schütze, Simon See, Yangqiu Song
ACL (1)3
2025 Converging to a Lingua Franca: Evolution of Linguistic Regions and Semantics Alignment in Multilingual Large Language Models
abstract
Large language models (LLMs) have demonstrated remarkable performance, particularly in multilingual contexts. While recent studies suggest that LLMs can transfer skills learned in one language to others, the internal mechanisms behind this ability remain unclear. We observed that the neuron activation patterns of LLMs exhibit similarities when processing the same language, revealing the existence and location of key linguistic regions. Additionally, we found that neuron activation patterns are similar when processing sentences with the same semantic meaning in different languages. This indicates that LLMs map semantically identical inputs from different languages into a “Lingua Franca”, a common semantic latent space that allows for consistent processing across languages. This semantic alignment becomes more pronounced with training and increased model size, resulting in a more language-agnostic activation pattern. Moreover, we found that key linguistic neurons are concentrated in the first and last layers of LLMs, becoming denser in the first layers as training progresses. Experiments on BLOOM and LLaMA2 support these findings, highlighting the structural evolution of multilingual LLMs during training and scaling up. This paper provides insights into the internal workings of LLMs, offering a foundation for future improvements in their cross-lingual capabilities.
Hongchuan Zeng, Senyu Han, Lu Chen 0002, Kai Yu 0004
COLING1
2025 Heads up! Large Language Models Can Perform Tasks Without Your Instruction via Selective Attention Head Masking
abstract
Large language models (LLMs) consist of numerous Transformer modules, and while the models can perform various functions, it remains an open question of how these modules are combined to elicit distinct inherent functionalities. In this paper, we investigate the modules inside LLMs and demonstrate that, by simply masking or retaining specific attention heads during inference, LLMs can exhibit specific task functionalities without requiring explicit instructions or modifications to the model parameters. Experiments across various models and tasks reveal that LLMs inherently encode “functional pathways”, the structured groups of interdependent attention heads that are crucial for executing specific tasks. These pathways not only govern the model’s functional behaviors but also enhance parameter efficiency, as suppressing attention heads outside the pathway can improve task performance. The code is available in this repository: https://github.com/OpenDFM/HeadsUp.
Senyu Han, Hongchuan Zeng, Kai Yu 0004, Lu Chen 0002
ICML2
2024 Multilingual Brain Surgeon: Large Language Models Can Be Compressed Leaving No Language behind
abstract
Large Language Models (LLMs) have ushered in a new era in Natural Language Processing, but their massive size demands effective compression techniques for practicality. Although numerous model compression techniques have been investigated, they typically rely on a calibration set that overlooks the multilingual context and results in significant accuracy degradation for low-resource languages. This paper introduces Multilingual Brain Surgeon (MBS), a novel calibration data sampling method for multilingual LLMs compression. MBS overcomes the English-centric limitations of existing methods by sampling calibration data from various languages proportionally to the language distribution of the model training datasets. Our experiments, conducted on the BLOOM multilingual LLM, demonstrate that MBS improves the performance of existing English-centric compression methods, especially for low-resource languages. We also uncover the dynamics of language interaction during compression, revealing that the larger the proportion of a language in the training set and the more similar the language is to the calibration language, the better performance the language retains after compression. In conclusion, MBS presents an innovative approach to compressing multilingual LLMs, addressing the performance disparities and improving the language inclusivity of existing compression techniques. Keywords: Large Language Model, Multilingual Model Compression
Hongchuan Zeng, Hongshen Xu, Lu Chen 0002, Kai Yu 0004
LREC/COLING1