EDBT 2026 Demo / reviewers in the wild / expert
Wenxuan Zhou 0002
dblp:78/9975-2
· DBLP profile ↗
18ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0003-1199-885XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RedCoder: Automated Multi-Turn Red Teaming for Code LLMsabstractWenjie Jacky Mo, Qin Liu, Xiaofei Wen, Dongwon Jung, Hadi Askari, Wenxuan Zhou, Zhe Zhao, Muhao Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wenjie Mo 0001, Qin Liu 0010, Xiaofei Wen, Dongwon Jung, Hadi Askari, Wenxuan Zhou 0002, Muhao Chen 0001 |
ACL (1) | 6 |
| 2025 | Code Execution as Grounded Supervision for LLM ReasoningabstractTraining large language models (LLMs) with chain-of-thought (CoT) supervision has proven effective for enhancing their reasoning abilities.However, obtaining reliable and accurate reasoning supervision remains a significant challenge.We propose a scalable method for generating a high-quality CoT supervision dataset by leveraging the determinism of program execution.Unlike existing reasoning dataset generation methods that rely on costly human annotations or error-prone LLM-generated CoT, our approach extracts verifiable, step-by-step reasoning traces from code execution and transforms them into a natural language CoT reasoning.Experiments on reasoning benchmarks across various domains show that our method effectively equips LLMs with transferable reasoning abilities across diverse tasks.Furthermore, the ablation studies validate that our method produces highly accurate reasoning data and reduces overall token length during inference by reducing meaningless repetition and overthinking.1 Dongwon Jung, Wenxuan Zhou 0002, Muhao Chen 0001 |
EMNLP | 2 |
| 2025 | MuirBench: A Comprehensive Benchmark for Robust Multi-image UnderstandingabstractWe introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tasks (e.g., scene understanding, ordering) that involve 10 categories of multi-image relations (e.g., multiview, temporal relations). Comprising 11,264 images and 2,600 multiple-choice questions, MuirBench is created in a pairwise manner, where each standard instance is paired with an unanswerable variant that has minimal semantic differences, in order for a reliable assessment. Evaluated upon 20 recent multi-modal LLMs, our results reveal that even the best-performing models like GPT-4o and Gemini Pro find it challenging to solve MuirBench, achieving 68.0% and 49.3% in accuracy. Open-source multimodal LLMs trained on single images can hardly generalize to multi-image questions, hovering below 33.3% in accuracy. These results highlight the importance of MuirBench in encouraging the community to develop multimodal LLMs that can look beyond a single image, suggesting potential pathways for future improvements. Fei Wang 0060, James Y. Huang, Zekun Li 0007, Qin Liu 0010, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu 0014, Wenxuan Zhou 0002, Kai Zhang 0008, Tianyi Lorena Yan, Wenjie Mo 0001, Hsiang-Hui Liu, Pan Lu, Chunyuan Li, Chaowei Xiao, Kai-Wei Chang 0001, Dan Roth 0001, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
ICLR | 9 |
| 2024 | mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsabstractDirect preference optimization (DPO) has shown to be an effective method for large language model (LLM) alignment.Recent works have attempted to apply DPO to multimodal scenarios but have found it challenging to achieve consistent improvement.Through a comparative experiment, we identify the unconditional preference problem in multimodal preference optimization, where the model overlooks the image condition.To address this problem, we propose MDPO, a multimodal DPO objective that prevents the over-prioritization of language-only preferences by also optimizing image preference.Moreover, we introduce a reward anchor that forces the reward to be positive for chosen responses, thereby avoiding the decrease in their likelihood-an intrinsic problem of relative preference optimization.Experiments on two multimodal LLMs of different sizes and three widely used benchmarks demonstrate that MDPO effectively addresses the unconditional preference problem in multimodal preference optimization and significantly improves model performance, particularly in reducing hallucination. Fei Wang 0060, Wenxuan Zhou 0002, James Y. Huang, Nan Xu 0014, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
EMNLP | 2 |
| 2024 | UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionabstractLarge language models (LLMs) have demonstrated remarkable generalizability, such as understanding arbitrary entities and relations. Instruction tuning has proven effective for distilling LLMs into more cost-efficient models such as Alpaca and Vicuna. Yet such student models still trail the original LLMs by large margins in downstream applications. In this paper, we explore targeted distillation with mission-focused instruction tuning to train student models that can excel in a broad application class such as open information extraction. Using named entity recognition (NER) for case study, we show how ChatGPT can be distilled into much smaller UniversalNER models for open NER. For evaluation, we assemble the largest NER benchmark to date, comprising 43 datasets across 9 diverse domains such as biomedicine, programming, social media, law, finance. Without using any direct supervision, UniversalNER attains remarkable NER accuracy across tens of thousands of entity types, outperforming general instruction-tuned models such as Alpaca and Vicuna by over 30 absolute F1 points in average. With a tiny fraction of parameters, UniversalNER not only acquires ChatGPT's capability in recognizing arbitrary entity types, but also outperforms its NER accuracy by 7-9 absolute F1 points in average. Remarkably, UniversalNER even outperforms by a large margin state-of-the-art multi-task instruction-tuned systems such as InstructUIE, which uses supervised NER examples. We also conduct thorough ablation studies to assess the impact of various components in our distillation approach. We release the distillation recipe, data, and UniversalNER models to facilitate future research on targeted distillation. Wenxuan Zhou 0002, Sheng Zhang 0012, Yu Gu 0017, Muhao Chen 0001, Hoifung Poon |
ICLR | 1 |
| 2023 | Continual Contrastive Finetuning Improves Low-Resource Relation ExtractionabstractRelation extraction (RE), which has relied on structurally annotated corpora for model training, has been particularly challenging in lowresource scenarios and domains.Recent literature has tackled low-resource RE by selfsupervised learning, where the solution involves pretraining the entity pair embedding by RE-based objective and finetuning on labeled data by classification-based objective.However, a critical challenge to this approach is the gap in objectives, which prevents the RE model from fully utilizing the knowledge in pretrained representations.In this paper, we aim at bridging the gap and propose to pretrain and finetune the RE model using consistent objectives of contrastive learning.Since in this kind of representation learning paradigm, one relation may easily form multiple clusters in the representation space, we further propose a multi-center contrastive loss that allows one relation to form multiple clusters to better align with pretraining.Experiments on two document-level RE datasets, BioRED and Re-DocRED, demonstrate the effectiveness of our method.Particularly, when using 1% end-task training data, our method outperforms PLMbased RE classifier by 10.5% and 6.1% on the two datasets, respectively. Wenxuan Zhou 0002, Sheng Zhang 0012, Tristan Naumann, Muhao Chen 0001, Hoifung Poon |
ACL (1) | 1 |
| 2023 | How Fragile is Relation Extraction under Entity Replacements?abstractYiwei Wang, Bryan Hooi, Fei Wang, Yujun Cai, Yuxuan Liang, Wenxuan Zhou, Jing Tang, Manjuan Duan, Muhao Chen. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). 2023. Yiwei Wang 0001, Bryan Hooi, Fei Wang 0060, Yujun Cai, Yuxuan Liang 0002, Wenxuan Zhou 0002, Jing Tang 0004, Manjuan Duan, Muhao Chen 0001 |
CoNLL | 6 |
| 2023 | Parameter-Efficient Tuning with Special Token AdaptationabstractParameter-efficient tuning aims at updating only a small subset of parameters when adapting a pretrained model to downstream tasks.In this work, we introduce PASTA, in which we only modify the special token representations (e.g., [SEP] and [CLS] in BERT) before the self-attention module at each layer in Transformer-based models.PASTA achieves comparable performance to full finetuning in natural language understanding tasks including text classification and NER with up to only 0.029% of total parameters trained.Our work not only provides a simple yet effective way of parameter-efficient tuning, which has a wide range of practical applications when deploying finetuned models for multiple tasks, but also demonstrates the pivotal role of special tokens in pretrained language models. 1 Xiaocong Yang, James Y. Huang, Wenxuan Zhou 0002, Muhao Chen 0001 |
EACL | 3 |
| 2023 | GeoLM: Empowering Language Models for Geospatially Grounded Language UnderstandingabstractHumans subconsciously engage in geospatial reasoning when reading articles.We recognize place names and their spatial relations in text and mentally associate them with their physical locations on Earth.Although pretrained language models can mimic this cognitive process using linguistic context, they do not utilize valuable geospatial information in large, widely available geographical databases, e.g., OpenStreetMap.This paper introduces GEOLM ( ), a geospatially grounded language model that enhances the understanding of geo-entities in natural language.GEOLM leverages geo-entity mentions as anchors to connect linguistic information in text corpora with geospatial information extracted from geographical databases.GEOLM connects the two types of context through contrastive learning and masked language modeling.It also incorporates a spatial coordinate embedding mechanism to encode distance and direction relations to capture geospatial context.In the experiment, we demonstrate that GEOLM exhibits promising capabilities in supporting toponym recognition, toponym linking, relation extraction, and geo-entity typing, which bridge the gap between natural language processing and geospatial sciences.The code is publicly available at https://github.com/ knowledge-computing/geolm. Zekun Li 0007, Wenxuan Zhou 0002, Yao-Yi Chiang, Muhao Chen 0001 |
EMNLP | 2 |
| 2022 | Prix-LM: Pretraining for Multilingual Knowledge Base ConstructionabstractKnowledge bases (KBs) contain plenty of structured world and commonsense knowledge.As such, they often complement distributional text-based information and facilitate various downstream tasks.Since their manual construction is resource-and timeintensive, recent efforts have tried leveraging large pretrained language models (PLMs) to generate additional monolingual knowledge facts for KBs.However, such methods have not been attempted for building and enriching multilingual KBs.Besides wider application, such multilingual KBs can provide richer combined knowledge than monolingual (e.g., English) KBs.Knowledge expressed in different languages may be complementary and unequally distributed: this implies that the knowledge available in high-resource languages can be transferred to low-resource ones.To achieve this, it is crucial to represent multilingual knowledge in a shared/unified space.To this end, we propose a unified representation model, Prix-LM , for multilingual KB construction and completion.We leverage two types of knowledge, monolingual triples and cross-lingual links, extracted from existing multilingual KBs, and tune a multilingual language encoder XLM-R via a causal language modeling objective.Prix-LM integrates useful multilingual and KB-based factual knowledge into a single model.Experiments on standard entity-related tasks, such as link prediction in multiple languages, cross-lingual entity linking and bilingual lexicon induction, demonstrate its effectiveness, with gains reported over strong task-specialised baselines. Wenxuan Zhou 0002, Fangyu Liu 0001, Ivan Vulic, Nigel Collier, Muhao Chen 0001 |
ACL (1) | 1 |
| 2022 | Should We Rely on Entity Mentions for Relation Extraction? Debiasing Relation Extraction with Counterfactual AnalysisabstractYiwei Wang, Muhao Chen, Wenxuan Zhou, Yujun Cai, Yuxuan Liang, Dayiheng Liu, Baosong Yang, Juncheng Liu, Bryan Hooi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yiwei Wang 0001, Muhao Chen 0001, Wenxuan Zhou 0002, Yujun Cai, Yuxuan Liang 0002, Dayiheng Liu, Baosong Yang, Bryan Hooi |
NAACL-HLT | 3 |
| 2022 | Answer Consolidation: Formulation and BenchmarkingabstractWenxuan Zhou, Qiang Ning, Heba Elfardy, Kevin Small, Muhao Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Wenxuan Zhou 0002, Qiang Ning, Heba Elfardy, Kevin Small, Muhao Chen 0001 |
NAACL-HLT | 1 |
| 2021 | Document-Level Relation Extraction with Adaptive Thresholding and Localized Context PoolingabstractDocument-level relation extraction (RE) poses new challenges compared to its sentence-level counterpart. One document commonly contains multiple entity pairs, and one entity pair occurs multiple times in the document associated with multiple possible relations. In this paper, we propose two novel techniques, adaptive thresholding and localized context pooling, to solve the multi-label and multi-entity problems. The adaptive thresholding replaces the global threshold for multi-label classification in the prior work with a learnable entities-dependent threshold. The localized context pooling directly transfers attention from pre-trained language models to locate relevant context that is useful to decide the relation. We experiment on three document-level RE benchmark datasets: DocRED, a recently released large-scale RE dataset, and two datasets CDRand GDA in the biomedical domain. Our ATLOP (Adaptive Thresholding and Localized cOntext Pooling) model achieves an F1 score of 63.4, and also significantly outperforms existing models on both CDR and GDA. We have released our code at https://github.com/wzhouad/ATLOP. Wenxuan Zhou 0002, Kevin Huang 0002, Tengyu Ma 0001, Jing Huang 0019 |
AAAI | 1 |
| 2021 | IsoBN: Fine-Tuning BERT with Isotropic Batch NormalizationabstractFine-tuning pre-trained language models (PTLMs), such as BERT and its better variant RoBERTa, has been a common practice for advancing performance in natural language understanding (NLU) tasks. Recent advance in representation learning shows that isotropic (i.e., unit-variance and uncorrelated) embeddings can significantly improve performance on downstream tasks with faster convergence and better generalization. The isotropy of the pre-trained embeddings in PTLMs, however, is relatively under-explored. In this paper, we analyze the isotropy of the pre-trained [CLS] embeddings of PTLMs with straightforward visualization, and point out two major issues: high variance in their standard deviation, and high correlation between different dimensions. We also propose a new network regularization method, isotropic batch normalization (IsoBN) to address the issues, towards learning more isotropic representations in fine-tuning by dynamically penalizing dominating principal components. This simple yet effective fine-tuning method yields about 1.0 absolute increment on the average of seven NLU tasks. Wenxuan Zhou 0002, Bill Y. Lin, Xiang Ren 0001 |
AAAI | 1 |
| 2021 | Contrastive Out-of-Distribution Detection for Pretrained TransformersabstractPretrained Transformers achieve remarkable performance when training and test data are from the same distribution.However, in realworld scenarios, the model often faces out-ofdistribution (OOD) instances that can cause severe semantic shift problems at inference time.Therefore, in practice, a reliable model should identify such instances, and then either reject them during inference or pass them over to models that handle another distribution.In this paper, we develop an unsupervised OOD detection method, in which only the indistribution (ID) data are used in training.We propose to fine-tune the Transformers with a contrastive loss, which improves the compactness of representations, such that OOD instances can be better differentiated from ID ones.These OOD instances can then be accurately detected using the Mahalanobis distance in the model's penultimate layer.We experiment with comprehensive settings and achieve near-perfect OOD detection performance, outperforming baselines drastically.We further investigate the rationales behind the improvement, finding that more compact representations through margin-based contrastive learning bring the improvement.We release our code to the community for future research 1 . Wenxuan Zhou 0002, Fangyu Liu 0001, Muhao Chen 0001 |
EMNLP (1) | 1 |
| 2021 | Learning from Noisy Labels for Entity-Centric Information ExtractionabstractRecent information extraction approaches have relied on training deep neural models.However, such models can easily overfit noisy labels and suffer from performance degradation.While it is very costly to filter noisy labels in large learning resources, recent studies show that such labels take more training steps to be memorized and are more frequently forgotten than clean labels, therefore are identifiable in training.Motivated by such properties, we propose a simple co-regularization framework for entity-centric information extraction, which consists of several neural models with identical structures but different parameter initialization.These models are jointly optimized with the task-specific losses and are regularized to generate similar predictions based on an agreement loss, which prevents overfitting on noisy labels.Extensive experiments on two widely used but noisy benchmarks for information extraction, TACRED and CoNLL03, demonstrate the effectiveness of our framework.We release our code to the community for future research 1 . Wenxuan Zhou 0002, Muhao Chen 0001 |
EMNLP (1) | 1 |
| 2020 | Learning from Explanations with Neural Execution Tree
Ziqi Wang 0003, Yujia Qin, Wenxuan Zhou 0002, Jun Yan 0012, Qinyuan Ye, Leonardo Neves, Zhiyuan Liu 0001, Xiang Ren 0001 |
ICLR | 3 |
| 2020 | NERO: A Neural Rule Grounding Framework for Label-Efficient Relation ExtractionabstractDeep neural models for relation extraction tend to be less reliable when perfectly labeled data is limited, despite their success in label-sufficient scenarios. Instead of seeking more instance-level labels from human annotators, here we propose to annotate frequent surface patterns to form labeling rules. These rules can be automatically mined from large text corpora and generalized via a soft rule matching mechanism. Prior works use labeling rules in an exact matching fashion, which inherently limits the coverage of sentence matching and results in the low-recall issue. In this paper, we present a neural approach to ground rules for RE, named Nero, which jointly learns a relation extraction module and a soft matching module. One can employ any neural relation extraction models as the instantiation for the RE module. The soft matching module learns to match rules with semantically similar sentences such that raw corpora can be automatically labeled and leveraged by the RE module (in a much better coverage) as augmented supervision, in addition to the exactly matched sentences. Extensive experiments and analysis on two public and widely-used datasets demonstrate the effectiveness of the proposed Nero framework, comparing with both rule-based and semi-supervised methods. Through user studies, we find that the time efficiency for a human to annotate rules and sentences are similar (0.30 vs. 0.35 min per label). In particular, Nero’s performance using 270 rules is comparable to the models trained using 3,000 labeled sentences, yielding a 9.5x speedup. Moreover, Nero can predict for unseen relations at test time and provide interpretable predictions. We release our code1 to the community for future research. Wenxuan Zhou 0002, Bill Y. Lin, Ziqi Wang 0003, Junyi Du, Leonardo Neves, Xiang Ren 0001 |
WWW | 1 |