VLDB 2026 Research / reviewers in the wild / expert
Chenghao Xiao
dblp:325/2555
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0001-7623-8232ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Understanding the Behaviors of Environment-aware Information RetrievalabstractRuifeng Yuan, Chaohao Yuan, David Dai, Yu Rong, Hong Cheng, Hou Pong Chan, Chenghao Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ruifeng Yuan, Chaohao Yuan, David Dai, Yu Rong 0001, Hong Cheng 0001, Hou Pong Chan, Chenghao Xiao |
ACL (1) | 7 |
| 2026 | Reevaluating zero-shot information extraction: Sampling bias, prompting transferability and sensitivity in large language models
Chenghao Xiao, Noura Al Moubayed |
Inf. Process. Manag. | 2 |
| 2026 | Beyond One-Size-Fits-All : Inversion Learning for Highly Effective NLG Evaluation PromptsabstractAbstract Evaluating natural language generation systems is challenging due to the diversity of valid outputs. While human evaluation is the gold standard, it suffers from inconsistencies, lack of standardization, and demographic biases, limiting reproducibility. LLM-based evaluators offer a scalable alternative but are highly sensitive to prompt design, where small variations can lead to significant discrepancies. In this work, we propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions, enabling the automatic generation of highly effective, model-specific evaluation prompts. Our method requires only a single evaluation sample and eliminates the need for time-consuming manual prompt engineering, thereby improving both efficiency and robustness. Our work contributes toward a new direction for more robust and efficient LLM-based evaluation. Hanhua Hong, Chenghao Xiao, Yang Wang 0015, Wenge Rong, Chenghua Lin 0002 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal RepresentationsabstractWhile understanding the knowledge boundaries of LLMs is crucial to prevent hallucination, research on the knowledge boundaries of LLMs has predominantly focused on English. In this work, we present the first study to analyze how LLMs recognize knowledge boundaries across different languages by probing their internal representations when processing known and unknown questions in multiple languages. Our empirical studies reveal three key findings: 1) LLMs' perceptions of knowledge boundaries are encoded in the middle to middle-upper layers across different languages. 2) Language differences in knowledge boundary perception follow a linear structure, which motivates our proposal of a training-free alignment method that effectively transfers knowledge boundary perception ability across languages, thereby helping reduce hallucination risk in low-resource languages; 3) Fine-tuning on bilingual question pair translation further enhances LLMs' recognition of knowledge boundaries across languages. Given the absence of standard testbeds for cross-lingual knowledge boundary analysis, we construct a multilingual evaluation suite comprising three representative types of knowledge boundary data. Our code and datasets are publicly available at https://github.com/DAMO-NLP-SG/ LLM-Multilingual-Knowledge-Boundaries. Chenghao Xiao, Hou Pong Chan, Hao Zhang 0048, Mahani Aljunied, Lidong Bing, Noura Al Moubayed, Yu Rong 0001 |
ACL (1) | 1 |
| 2025 | ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical ReasoningabstractYu Sun, Xingyu Qian, Weiwen Xu, Hao Zhang, Chenghao Xiao, Long Li, Deli Zhao, Wenbing Huang, Tingyang Xu, Qifeng Bai, Yu Rong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xingyu Qian, Weiwen Xu, Hao Zhang 0048, Chenghao Xiao, Deli Zhao, Wenbing Huang 0001, Tingyang Xu, Qifeng Bai, Yu Rong 0001 |
EMNLP | 5 |
| 2025 | Drivel-ology: Challenging LLMs with Interpreting Nonsense with DepthabstractWe introduce Drivelology, a unique linguistic phenomenon characterised as "nonsense with depth" -utterances that are syntactically coherent yet pragmatically paradoxical, emotionally loaded, or rhetorically subversive.While such expressions may resemble surface-level nonsense, they encode implicit meaning requiring contextual inference, moral reasoning, or emotional interpretation.We find that current large language models (LLMs), despite excelling at many natural language processing (NLP) tasks, consistently fail to grasp the layered semantics of Drivelological text.To investigate this, we construct a benchmark dataset of over 1,200+ meticulously curated and diverse examples across English, Mandarin, Spanish, French, Japanese, and Korean.Each example underwent careful expert review to verify its Drivelological characteristics, involving multiple rounds of discussion and adjudication to address disagreements.Using this dataset, we evaluate a range of LLMs on classification, generation, and reasoning tasks.Our results reveal clear limitations of LLMs: models often confuse Drivelology with shallow nonsense, produce incoherent justifications, or miss implied rhetorical functions altogether.These findings highlight a deep representational gap in LLMs' pragmatic understanding and challenge the assumption that statistical fluency implies cognitive comprehension.We release our dataset 1 and code 2 to facilitate further research in modelling linguistic depth beyond surface-level coherence. Yang Wang 0015, Chenghao Xiao, Chia-Yi Hsiao, Zi Yan Chang, Chi-Li Chen, Tyler Loakman, Chenghua Lin 0002 |
EMNLP | 2 |
| 2025 | Everything is a Video: Unifying Modalities Through Next-Frame PredictionabstractMultimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual question answering, cross-modal retrieval, and caption generation. Traditional approaches rely on modality-specific encoders and late fusion techniques, which can hinder scalability and flexibility when adapting to new tasks or modalities. To address these limitations, we introduce a novel framework that extends the concept of task reformulation beyond natural language processing (NLP) to multimodal learning. We propose to reformulate diverse multimodal tasks into a unified next-frame prediction problem, allowing a single model to handle different modalities without modality-specific components. This method treats all inputs and outputs as sequential frames in a video, enabling seamless integration of modalities and effective knowledge transfer across tasks. Our approach is evaluated on a range of tasks, including text-to-text, image-to-text, video-to-video, video-to-text, and audio-to-text, demonstrating the model's ability to generalize across modalities with minimal adaptation. We show that task reformulation can significantly simplify multimodal model design across various tasks, laying the groundwork for more generalized multimodal foundation models. G. Thomas Hudson, Dean L. Slack, Thomas Winterbottom, Jamie Sterling, Chenghao Xiao, Junjie Shentu, Noura Al Moubayed |
ICCV | 5 |
| 2025 | Mieb: Massive Image Embedding BenchmarkabstractImage representations are often evaluated through disjointed, task-specific protocols, leading to a fragmented understanding of model capabilities. For instance, it is unclear whether an image embedding model adept at clustering images is equally good at retrieving relevant images given a piece of text. We introduce the Massive Image Embedding Benchmark (MIEB) to evaluate the performance of image and image-text embedding models across the broadest spectrum to date. MIEB spans 38 languages across 130 individual tasks, which we group into 8 high-level categories. We benchmark 50 models across our benchmark, finding that no single method dominates across all task categories. We reveal hidden capabilities in advanced vision models such as their accurate visual representation of texts, and their yet limited capabilities in interleaved encodings and matching images and texts in the presence of confounders. We also show that the performance of vision encoders on MIEB correlates highly with their performance when used in multimodal large language models. Our code, dataset, and leaderboard are publicly available at https://github.com/embeddings-benchmark/mteb. Chenghao Xiao, Isaac Chung, Imene Kerboua, Jamie Stirling, Xin Zhang 0097, Márton Kardos, Roman Solomatin, Noura Al Moubayed, Kenneth C. Enevoldsen, Niklas Muennighoff |
ICCV | 1 |
| 2025 | CAST: Corpus-Aware Self-similarity Enhanced Topic modellingabstractYanan Ma, Chenghao Xiao, Chenhan Yuan, Sabine N Van Der Veer, Lamiece Hassan, Chenghua Lin, Goran Nenadic. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Chenghao Xiao, Chenhan Yuan, Sabine van der Veer, Lamiece Hassan, Goran Nenadic |
NAACL (Long Papers) | 2 |
| 2025 | Scaling Language-centric Omnimodal Representation LearningabstractRecent multimodal embedding approaches leveraging multimodal large language models (MLLMs) fine-tuned with contrastive learning (CL) have shown promising results, yet the underlying reasons behind their superiority remain underexplored. This work argues that a crucial advantage of MLLM-based approaches stems from implicit cross-modal alignment achieved during generative pretraining, where the language decoder learns to exploit multimodal signals within a shared representation space for generating unimodal outputs. Through analysis of anisotropy and kernel similarity structure, we empirically confirm that latent alignment emerges within MLLM representations, allowing CL to serve as a lightweight refinement stage. Leveraging this insight, we propose a Language-Centric Omnimodal Embedding framework, termed LCO-Embed. Extensive experiments across diverse backbones and benchmarks demonstrate its effectiveness, achieving state-of-the-art performance across modalities. Furthermore, we identify a Generation-Representation Scaling Law (GRSL), showing that the representational capabilities gained through contrastive refinement scale positively with the MLLM's generative capabilities. This suggests that improving generative abilities evolves as an effective paradigm for enhancing representation quality. We provide a theoretical explanation of GRSL, which formally links the MLLM's generative quality to the upper bound on its representation performance, and validate it on a challenging, low-resource visual-document retrieval task, showing that continual generative pretraining before CL can further enhance the potential of a model's embedding capabilities. Codes, models, and resources are available at https://github.com/LCO-Embedding/LCO-Embedding. Chenghao Xiao, Hou Pong Chan, Hao Zhang 0048, Weiwen Xu, Mahani Aljunied, Yu Rong 0001 |
NeurIPS | 1 |
| 2025 | Adversarial Defense without Adversarial Defense : Enhancing Language Model Robustness via Instance-level Principal Component RemovalabstractPre-trained language models (PLMs) have driven substantial progress in natural language processing but remain vulnerable to adversarial attacks, raising concerns about their robustness in real-world applications. Previous studies have sought to mitigate the impact of adversarial attacks by introducing adversarial perturbations into the training process, either implicitly or explicitly. While both strategies enhance robustness, they often incur high computational costs. In this work, we propose a simple yet effective add-on module that enhances the adversarial robustness of PLMs by removing instance-level principal components, without relying on conventional adversarial defenses or perturbing the original training data. Our approach transforms the embedding space to approximate Gaussian properties, thereby reducing its susceptibility to adversarial perturbations while preserving semantic relationships. This transformation aligns embedding distributions in a way that minimizes the impact of adversarial noise on decision boundaries, enhancing robustness without requiring adversarial examples or costly training-time augmentation. Evaluations on eight benchmark datasets show that our approach improves adversarial robustness while maintaining comparable before-attack accuracy to baselines, achieving a balanced trade-off between robustness and generalization. Yang Wang 0015, Chenghao Xiao, Stuart E. Middleton, Noura Al Moubayed, Chenghua Lin 0002 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2024 | Effective Distillation of Table-based Reasoning Ability from LLMsabstractLarge Language Models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing tasks. However, their enormous parameter size and extremely high requirements for compute power pose challenges for their practical deployment. Recent research has revealed that specific capabilities of LLMs, such as numerical reasoning, can be transferred to smaller models through distillation. Some studies explore the potential of leveraging LLMs to perform table-based reasoning. However, there has been no prior work focusing on table reasoning skills in smaller models specifically tailored for scientific table-to-text generation tasks. In this paper, we propose a novel table-based reasoning distillation approach, with the aim of distilling LLMs into tailored smaller models. Our experimental results have shown that a 220 million parameter model (Flan-T5-base) fine-tuned using distilled data, not only achieves a significant improvement compared to traditionally fine-tuned baselines, but also surpasses specific LLMs on a scientific table-to-text generation dataset. Our code is available at https://github.com/Bernard-Yang/DistillTableCoT. Bohao Yang, Kun Zhao 0007, Chenghao Xiao, Chenghua Lin 0002 |
LREC/COLING | 4 |
| 2024 | MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingabstractSelf-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech.
Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored. This is partially due to the distinctive challenges associated with modelling musical knowledge, particularly tonal and pitched characteristics of music.
To address this research gap, we propose an acoustic **M**usic und**ER**standing model with large-scale self-supervised **T**raining (**MERT**), which incorporates teacher models to provide pseudo labels in the masked language modelling (MLM) style acoustic pre-training.
In our exploration, we identified an effective combination of teacher models, which outperforms conventional speech and audio approaches in terms of performance.
This combination includes an acoustic teacher based on Residual Vector Quantization - Variational AutoEncoder (RVQ-VAE) and a musical teacher based on the Constant-Q Transform (CQT).
Furthermore, we explore a wide range of settings to overcome the instability in acoustic language model pre-training, which allows our designed paradigm to scale from 95M to 330M parameters.
Experimental results indicate that our model can generalise and perform well on 14 music understanding tasks and attain state-of-the-art (SOTA) overall scores. Ruibin Yuan, Ge Zhang 0009, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin 0002, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, Roger B. Dannenberg, Ruibo Liu, Wenhu Chen, Gus Xia, Yemin Shi 0001, Wenhao Huang 0001, Yike Guo, Jie Fu 0001 |
ICLR | 7 |
| 2023 | Length is a Curse and a Blessing for Document-level SemanticsabstractIn recent years, contrastive learning (CL) has been extensively utilized to recover sentence and document-level encoding capability from pre-trained language models.In this work, we question the length generalizability of CLbased models, i.e., their vulnerability towards length-induced semantic shift.We verify not only that length vulnerability is a significant yet overlooked research gap, but we can devise unsupervised CL methods solely depending on the semantic signal provided by document length.We first derive the theoretical foundations underlying length attacks, showing that elongating a document would intensify the high intra-document similarity that is already brought by CL.Moreover, we found that isotropy promised by CL is highly dependent on the length range of text exposed in training.Inspired by these findings, we introduce a simple yet universal document representation learning framework, LA(SER) 3 : length-agnostic self-reference for semantically robust sentence representation learning, achieving state-of-theart unsupervised performance on the standard information retrieval benchmark.Our code is publicly available. Chenghao Xiao, G. Thomas Hudson, Chenghua Lin 0002, Noura Al Moubayed |
EMNLP | 1 |
| 2022 | Fine-grained Main Ideas Extraction and Clustering of Online Course Reviews
Chenghao Xiao, Lei Shi 0003, Alexandra I. Cristea, Zhaoxing Li |
AIED (1) | 1 |