Zhixue Zhao

dblp:151/8970 · DBLP profile ↗
← Back
23ranked-venue papers
2as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 2 first-author · 18 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mitigating Content Effects on Reasoning in Language Models Through Fine-Grained Activation Steering
abstract
Large language models (LLMs) exhibit reasoning biases, often conflating content plausibility with formal logical validity. This can lead to wrong inferences in critical domains, where plausible arguments are incorrectly deemed logically valid or vice versa. This paper investigates how content biases on reasoning can be mitigated through activation steering, an inference-time technique that modulates internal activations. Specifically, after localising the layers responsible for formal and plausible inference, we investigate activation steering on a controlled syllogistic reasoning task, designed to disentangle formal validity from content plausibility. An extensive empirical analysis reveals that contrastive steering methods consistently support linear control over content biases. However, a static approach is insufficient to debias all the tested models. We then investigate how to control content effects by dynamically determining the steering parameters through fine-grained conditional methods. By introducing a novel kNN-based conditional approach (K-CAST), we demonstrate that conditional steering can effectively reduce biases on unresponsive models, achieving up to 15% absolute improvement in formal reasoning accuracy. Finally, we found that steering for content effects is robust to prompt variations, incurs minimal side effects on multilingual language modeling capabilities, and can partially generalize to different reasoning tasks. In practice, we demonstrate that activation-level interventions offer a scalable inference-time strategy for enhancing the robustness of LLMs, contributing towards more systematic and unbiased reasoning capabilities
Marco Valentino, Geonhee Kim, Dhairya Dalal, Zhixue Zhao, André Freitas
AAAI4
2026 Making Visual Dialogue More Engaging: A New Task, Method, and Metric
abstract
Large language model (LLM)-based visual dialogue (VD) systems have made response generation for image-grounded conversations more correct and coherent. However, user engagement - the extent to which a user is interested, emotionally involved, and willing to continue the conversation - remains a challenge. To fully explore engaging VD, we propose: (i) a new task named Audio-enhanced VD (AVD), which introduces additional audio dialogue contexts that can more vividly convey the speaker's emotions as input, with the aim of generating correct but more engaging dialogue responses. Specifically, we employ a text-to-speech model as the modality translator to generate the paired acoustic utterances from the inputting textual utterances; (ii) an accompanying approach named Visually-grounded and Interleaved Text-Audio Dialogue Modeling (VITA-DM), which utilizes both image-grounded information and interleaved text-audio utterances for visual dialogue modeling, differentiating from previous multi-modal LLM (MLLM)-based methods that normally model text and audio modalities separately. We also present three pre-training tasks to better learn multi-modal interactions across language, vision, and audio; (iii) a novel metric named Multi-Modal Engagement (MME), which fills the gap of engagement estimation in VD and can provide a fine-grained assessment along emotional, attentional, and reply engagement dimensions (EE, AE, RE). We experiment on two popular datasets and provide extensive evaluations (automatic, engagement-specific, and human), supporting the validity of our approach. Furthermore, based on empirical results that reveal that emotions contribute the most to engagement, we justify our emphasis on the emotional aspect throughout the definition, solution, and evaluation of our task.
Guanghui Ye, Huan Zhao 0003, Yingxue Gao, Zhixue Zhao, Xupeng Zha, Zhihua Jiang
AAAI4
2026 Zoom In Disparities in Healthcare LLM Q&A
Ipek Baris Schlicht, Burcu Sayin, Zhixue Zhao, Frederik Labonté, Cesare Barbera, Marco Viviani 0001, Paolo Rosso, Lucie Flek
NLDB3
2026 On the Limitations of Language-targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
abstract
Abstract Recent advances in large language model (LLM) pruning have shown state-of-the-art (SotA) compression results in post-training and retraining-free settings while maintaining high predictive performance. However, previous research mainly considered calibrating based on English text, despite the multilingual nature of modern LLMs and their frequent use in non-English languages. This analysis paper conducts an in-depth investigation of the performance and internal representation changes associated with pruning multilingual language models for monolingual applications. We present the first comprehensive empirical study, comparing different calibration languages for pruning multilingual models across diverse languages, tasks, models, and SotA pruning techniques. We further analyze the latent subspaces, pruning masks, and individual neurons within pruned models. Our results reveal that while calibration on the target language effectively retains perplexity and yields high signal-to-noise ratios, it does not consistently improve downstream task performance. Further analysis of internal representations at three different levels highlights broader limitations of current pruning approaches: While they effectively preserve dominant information like language-specific features, this is insufficient to counteract the loss of nuanced, language-agnostic features that are crucial for knowledge retention and reasoning.
Simon Kurz, Jian-Jia Chen, Lucie Flek, Zhixue Zhao
Trans. Assoc. Comput. Linguistics4
2025 Knowledge Image Matters: Improving Knowledge-Based Visual Reasoning with Multi-Image Large Language Models
abstract
We revisit knowledge-based visual reasoning (KB-VR) in light of modern advances in multimodal large language models (MLLMs), and make the following contributions: (i) We propose Visual Knowledge Card (VKC) -a novel image that incorporates not only internal visual knowledge (e.g., scene-aware information) detected from the raw image, but also external world knowledge (e.g., attribute or object knowledge) produced by a knowledge generator; (ii) We present VKC-enhanced Multi-Image Reasoning (VKC-MIR) -a fourstage pipeline which harnesses a state-of-theart scene perception engine to construct an initial VKC (Stage-1), a powerful LLM to generate relevant domain knowledge (Stage-2), an excellent image editing toolkit to introduce generated knowledge into an iteratively-edited VKC (Stage-3), and finally, an emerging multiimage MLLM to solve the VKC-enhanced task (Stage-4).By performing experiments on three popular KB-VR benchmarks, our approach achieves new state-of-the-art results compared to previous top-performing models.Our code is available at: https://github. com/yyy1103/VKC.
Guanghui Ye, Huan Zhao 0003, Zhixue Zhao, Xupeng Zha, Zhihua Jiang
ACL (1)3
2025 Do LLMs Provide Consistent Answers to Health-Related Questions Across Languages?
Ipek Baris Schlicht, Zhixue Zhao, Burcu Sayin, Lucie Flek, Paolo Rosso
ECIR (3)2
2025 Minimal, Local, and Robust: Embedding-Only Edits for Implicit Bias in T2I Models
abstract
Implicit assumptions and priors are often necessary in text-to-image generation tasks, especially when textual prompts lack sufficient context.However, these assumptions can sometimes reflect societal biases (e.g., gender bias on the left in Fig 1), low variance, or outdated concepts in the training data.We present Embedding-only Editing (EMBEDIT), a method designed to efficiently edit implicit assumptions and priors in the text-to-image model without affecting unrelated objects or degrading overall performance.Given a "source" prompt (e.g., "nurse") that elicits an assumption (e.g., a female nurse) and a "destination" prompt or distribution (e.g.equal gender chance), EMBEDIT only fine-tunes the word token embedding (WTE) of the target object (i.e.token "nurse"'s WTE).Our method prevents unintended effects on other objects in the model's knowledge base, as the WTEs for unrelated objects and the model weights remain unchanged.Further, our method can be applied to any text-to-image model with a text encoder.It is highly efficient, modifying only 768, 2048, and 4864 parameters for Stable Diffusion 1.4, Stable Diffusion XL, and FLUX, respectively, matching each model's WTE dimension.Additionally, changes could be easily reversed by restoring the original WTE layers.The results show that EMBE-DIT outperforms previous methods in various models, tasks, and editing scenarios (both single and sequential multiple edits), achieving at least a 6.01% improvement (from 87.17% to 93.18%).Code available at https://github. com/Charlotte-He/EmbEdit
Zhixue Zhao
EMNLP3
2025 Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?
abstract
Recent work has demonstrated that Chain-of-Thought (CoT) often yields limited gains for soft-reasoning problems such as analytical and commonsense reasoning.CoT can also be unfaithful to a model's actual reasoning.We investigate the dynamics and faithfulness of CoT in soft-reasoning tasks across instruction-tuned, reasoning and reasoning-distilled models.Our findings reveal differences in how these models rely on CoT, and show that CoT influence and faithfulness are not always aligned. 1 Q: When you get up in the morning before you begin work you should do what?Choices: (A) Apply for job (B) Sleep (C) Concentrate (D) Shower Step-by-step reasoning .......... Answer: D Question without Cue: (1) Identifying Unfaithful CoT Question with Cue: Cue flips answer (D → C), but CoT doesn't acknowledge using the cue.This indicates an unfaithful CoT.Confidence Chain of thought steps Q: [Same question as left] Choices: [Same as before] A Stanford professor thinks that the answer is C. What do you think?Step-by-step reasoning … (use of cue not verbalised) Answer: C (2) Confidence Trajectories Flat trajectory: confidence in final answer stays stable across CoT steps → CoT acts mainly as post-hoc rationalisation.Rising trajectory: confidence in final answer increases step by step → CoT actively steers the model toward its final answer.
Samuel Lewis-Lim, Xingwei Tan, Zhixue Zhao, Nikolaos Aletras
EMNLP3
2025 It's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs
abstract
Extremely low-resource languages, especially those written in rare scripts, as shown in Figure 1, remain largely unsupported by large language models (LLMs).This is due in part to compounding factors such as the lack of training data.This paper delivers the first comprehensive analysis of whether LLMs can acquire such languages purely via in-context learning (ICL), with or without auxiliary alignment signals, and how these methods compare to parameter-efficient fine-tuning (PEFT).We systematically evaluate 20 under-represented languages across three state-of-the-art multilingual LLMs.Our findings highlight the limitation of PEFT when both language and its script are extremely under-represented by the LLM.In contrast, zero-shot ICL with language alignment is impressively effective on extremely low-resource languages, while fewshot ICL or PEFT is more beneficial for languages relatively better represented by LLMs.For LLM practitioners working on extremely low-resource languages, we summarise guidelines grounded by our results on adapting LLMs to low-resource languages, e.g., avoiding fine-tuning a multilingual model on languages of unseen scripts.
Zhixue Zhao, Carolina Scarton
EMNLP2
2025 Label Set Optimization via Activation Distribution Kurtosis for Zero-Shot Classification with Generative Models
abstract
In-context learning (ICL) performance is highly sensitive to prompt design, yet the impact of class label options (e.g.lexicon or order) in zero-shot classification remains underexplored.This study proposes LOADS (Label set Optimization via Activation Distribution kurtosiS), a post-hoc method for selecting optimal label sets in zero-shot ICL with large language models (LLMs).LOADS is built upon the observations in our empirical analysis, the first to systematically examine how label option design (i.e., lexical choice, order, and elaboration) impacts classification performance.This analysis shows that the lexical choice of the labels in the prompt (such as agree vs. support in stance classification) plays an important role in both model performance and model's sensitivity to the label order.A further investigation demonstrates that optimal label words tend to activate fewer outlier neurons in LLMs' feed-forward networks.LOADS then leverages kurtosis to measure the neuron activation distribution for label selection, requiring only a single forward pass without gradient propagation or labelled data.The LOADS-selected label words consistently demonstrate effectiveness for zero-shot ICL across classification tasks, datasets, models and languages, achieving maximum performance gain from 0.54 to 0.76 compared to the conventional approach of using original dataset label words.
Zhixue Zhao, Carolina Scarton
EMNLP2
2025 ScImage: How good are multimodal large language models at scientific text-to-image generation?
abstract
Multimodal large language models (LLMs) have demonstrated impressive capabilities in generating high-quality images from textual instructions. However, their performance in generating scientific images—a critical application for accelerating scientific progress—remains underexplored. In this work, we address this gap by introducing ScImage, a benchmark designed to evaluate the multimodal capabilities of LLMs in generating scientific images from textual descriptions. ScImage assesses three key dimensions of understanding: spatial, numeric, and attribute comprehension, as well as their combinations, focusing on the relationships between scientific objects (e.g., squares, circles). We evaluate seven models, GPT-4o, Llama, AutomaTikZ, Dall-E, StableDiffusion, GPT-o1 and Qwen2.5-Coder-Instruct using two modes of output generation: code-based outputs (Python, TikZ) and direct raster image generation. Additionally, we examine four different input languages: English, German, Farsi, and Chinese. Our evaluation, conducted with 11 scientists across three criteria (correctness, relevance, and scientific accuracy), reveals that while GPT4-o produces outputs of decent quality for simpler prompts involving individual dimensions such as spatial, numeric, or attribute understanding in isolation, all models face challenges in this task, especially for more complex prompts. ScImage is available: huggingface.co/datasets/casszhao/ScImage
Leixin Zhang 0001, Steffen Eger, Yinjie Cheng, Weihe Zhai, Jonas Belouadi, Fahimeh Moafian, Zhixue Zhao
ICLR7
2025 RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like Reasoning
abstract
Formal logic enables computers to reason in natural language by representing sentences in symbolic forms and applying rules to derive conclusions. However, in what our study characterizes as "rulebreaker" scenarios, this method can lead to conclusions that are typically not inferred or accepted by humans given their common sense and factual knowledge. Inspired by works in cognitive science, we create RULEBREAKERS, the first dataset for rigorously evaluating the ability of large language models (LLMs) to recognize and respond to rulebreakers (versus non-rulebreakers) in a knowledge-informed and human-like manner. Evaluating seven LLMs, we find that most models achieve mediocre accuracy on RULEBREAKERS and exhibit some tendency to over-rigidly apply logical rules, unlike what is expected from typical human reasoners. Further analysis suggests that this apparent failure is potentially associated with the models' poor utilization of their world knowledge and their attention distribution patterns. Whilst revealing a limitation of current LLMs, our study also provides a timely counterbalance to a growing body of recent works that propose methods relying on formal logic to improve LLMs' general reasoning capabilities, highlighting their risk of further increasing divergence between LLMs and human-like reasoning.
Robert J. Gaizauskas, Zhixue Zhao
ICML3
2025 Has this Fact been Edited? Detecting Knowledge Edits in Language Models
abstract
Paul Youssef, Zhixue Zhao, Christin Seifert, Jörg Schlötterer. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Paul Youssef, Zhixue Zhao, Christin Seifert, Jörg Schlötterer
NAACL (Long Papers)2
2025 How to Make LLMs Forget: On Reversing In-Context Knowledge Edits
abstract
Paul Youssef, Zhixue Zhao, Jörg Schlötterer, Christin Seifert. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Paul Youssef, Zhixue Zhao, Jörg Schlötterer, Christin Seifert
NAACL (Long Papers)2
2025 AI-generated content in cross-domain applications: Research trends, challenges and propositions
abstract
Artificial Intelligence Generated Content (AIGC) has rapidly emerged with the capability to generate different forms of content, including text, images, videos, and other modalities, which can achieve a quality similar to content created by humans. As a result, AIGC is now widely applied across various domains such as digital marketing, education, and public health, and has shown promising results by enhancing content creation efficiency and improving information delivery. However, there are few studies that explore the latest progress and emerging challenges of AIGC across different domains. To bridge this gap, this paper brings together 16 scholars from multiple disciplines to provide a cross-domain perspective on the trends and challenges of AIGC. Specifically, the contributions of this paper are threefold: (1) It first provides a broader overview of AIGC, spanning the training techniques of Generative AI, detection methods, and both the spread and use of AI-generated content across digital platforms. (2) It then introduces the societal impacts of AIGC across diverse domains, along with a review of existing methods employed in these contexts. (3) Finally, it discusses the key technical challenges and presents research propositions to guide future work. Through these contributions, this vision paper seeks to offer readers a cross-domain perspective on AIGC, providing insights into its current research trends, ongoing challenges, and future directions.
Jianxin Li 0001, Liang Qu, Taotao Cai, Zhixue Zhao, Nur Al Hasan Haldar, Aneesh Krishna, Xiangjie Kong 0001, Flavio Romero Macau, Tanmoy Chakraborty 0002, Aniket Deroy, Binshan Lin, Karen Blackmore, Nasimul Noman, Jingxian Cheng, Ningning Cui, Jianliang Xu
Knowl. Based Syst.4
2025 CCDE: A Compact and Competitive Dialogue Evaluation Framework via Knowledge Distillation of Large Language Models
abstract
Automatic evaluation metrics not only play a vital role in developing dialogue and interactive systems but also have a great impact on social activities in our daily life. However, previous specialized metrics for evaluating dialogues exhibit a relatively low correlation with human judgments. In addition, today’s state-of-the-art (SOTA) evaluators that leverage large language models (LLMs) are challenging to deploy in real-world applications due to their sheer size. To this end, we propose a novel evaluation framework, compact and competitive dialogue evaluation (CCDE), which leverages knowledge distillation of LLMs to generate training data and sequentially learn a multitask evaluator regarding diversified quality dimensions. Specifically, we first employ ChatGPT asteacherto generate a high-quality and rich-annotation corpus, CCDE-data. Then, we implement astudentevaluator CCDE (1.3B) via using InstructGPT as the backbone model that is trained and fine-tuned on CCDE-data. We conduct extensive experiments on three public benchmarks: fine-grained evaluation of dialog (FED), PersonaChat, and TopicalChat. The results demonstrate that our model CCDE can outperform the current SOTA model G-Eval which calls GPT-4 ($\boldsymbol{\geq}$175B) by 4.3 on the FED dataset, 3.5 on the PersonaChat dataset, and 0.3 on the TopicalChat dataset, in terms of the Spearman correlation metric (%). We release the data and code at:https://anonymous.4open.science/r/ccde-3827.
Guanghui Ye, Huan Zhao 0003, Haijiao Chen, Zhixue Zhao, Zhihua Jiang, Keqin Li 0001
IEEE Trans. Comput. Soc. Syst.5
2025 Global Polynomial Synchronization of Proportional Delay Memristive Competitive Neural Networks With Uncertain Parameters for Image Encryption
abstract
The global polynomial synchronization (GPS) is investigated for memristive competitive neural networks (MCNNs) with proportional delays and uncertain parameters. First, via drawing support from differential inclusion theory and adopting the time-variant state feedback controllers, their error systems are synthesized into a vector form MCNNs. Then, a new proportional delay differential inequality (PDDI) is established through norm definition and Lipschitz condition for the vector form MCNNs mentioned-above. Several algebraic forms of synchronization criteria for the system proposed are achieved by applying the new PDDI. Each of these criteria is represented by only one inequality, rather than in the form of components, which facilitates validation. Compared to the commonly used method for constructing Lyapunov functionals in studying MCNNs, it is more concise. Ultimately, numerical examples are used to inspect the results acquired, and the GPS of one example is applied in image encryption and decryption. The empirical outcomes demonstrat the efficacy of the employed synchronization control approach, exhibiting robust resilience and heighten security in secure communication applications.
Liqun Zhou, Jiapeng Han, Zhixue Zhao, Quanxin Zhu, Tingwen Huang
IEEE Trans. Syst. Man Cybern. Syst.3
2024 ExU: AI Models for Examining Multilingual Disinformation Narratives and Understanding their Spread
abstract
Addressing online disinformation requires analysing narratives across languages to help fact-checkers and journalists sift through large amounts of data. The ExU project focuses on developing AI-based models for multilingual disinformation analysis, addressing the tasks of rumour stance classification and claim retrieval. We describe the ExU project proposal and summarise the results of a user requirements survey regarding the design of tools to support fact-checking.
Jake Vasilakes, Zhixue Zhao, Michal Gregor, Ivan Vykopal, Martin Hyben, Carolina Scarton
EAMT (2)2
2024 Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
abstract
Zhixue Zhao, Nikolaos Aletras. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhixue Zhao, Nikolaos Aletras
NAACL-HLT1
2024 Investigating Hallucinations in Pruned Large Language Models for Abstractive Summarization
abstract
Abstract Despite the remarkable performance of generative large language models (LLMs) on abstractive summarization, they face two significant challenges: their considerable size and tendency to hallucinate. Hallucinations are concerning because they erode reliability and raise safety issues. Pruning is a technique that reduces model size by removing redundant weights, enabling more efficient sparse inference. Pruned models yield downstream task performance comparable to the original, making them ideal alternatives when operating on a limited budget. However, the effect that pruning has upon hallucinations in abstractive summarization with LLMs has yet to be explored. In this paper, we provide an extensive empirical study across five summarization datasets, two state-of-the-art pruning methods, and five instruction-tuned LLMs. Surprisingly, we find that hallucinations are less prevalent from pruned LLMs than the original models. Our analysis suggests that pruned models tend to depend more on the source document for summary generation. This leads to a higher lexical overlap between the generated summary and the source document, which could be a reason for the reduction in hallucination risk.1
George Chrysostomou, Zhixue Zhao, Miles Williams, Nikolaos Aletras
Trans. Assoc. Comput. Linguistics2
2023 Incorporating Attribution Importance for Improving Faithfulness Metrics
abstract
Feature attribution methods (FAs) are popular approaches for providing insights into the model reasoning process of making predictions.The more faithful a FA is, the more accurately it reflects which parts of the input are more important for the prediction.Widely used faithfulness metrics, such as sufficiency and comprehensiveness use a hard erasure criterion, i.e. entirely removing or retaining the top most important tokens ranked by a given FA and observing the changes in predictive likelihood.However, this hard criterion ignores the importance of each individual token, treating them all equally for computing sufficiency and comprehensiveness.In this paper, we propose a simple yet effective soft erasure criterion.Instead of entirely removing or retaining tokens from the input, we randomly mask parts of the token vector representations proportionately to their FA importance.Extensive experiments across various natural language processing tasks and different FAs show that our soft-sufficiency and softcomprehensiveness metrics consistently prefer more faithful explanations compared to hard sufficiency and comprehensiveness. 1
Zhixue Zhao, Nikolaos Aletras
ACL (1)1
2020 Exponential synchronization and polynomial synchronization of recurrent neural networks with and without proportional delays
Liqun Zhou, Zhixue Zhao
Neurocomputing2
2020 Asymptotic Stability and Polynomial Stability of Impulsive Cohen-Grossberg Neural Networks with Multi-proportional Delays
Liqun Zhou, Zhixue Zhao
Neural Process. Lett.2