Wangchunshu Zhou

dblp:245/8640 · DBLP profile ↗
← Back
40ranked-venue papers
12as first author
33since 2021 · last 2026
0000-0003-4668-3348ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 12 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SupReMix: Supervised contrastive learning for medical imaging regression with mixup
abstract
In medical image analysis, regression plays a critical role in computer-aided diagnosis. It enables quantitative measurements such as age prediction from structural imaging, cardiac function quantification, and molecular measurement from PET scans. While deep learning has shown promise for these tasks, most approaches focus solely on optimizing regression loss or model architecture, neglecting the quality of learned feature representations which are crucial for robust clinical predictions. Directly applying representation learning techniques designed for classification to regression often results in fragmented representations in the latent space, yielding sub-optimal performance. In this paper, we argue that the potential of contrastive learning for medical image regression has been overshadowed due to the neglect of two crucial aspects: ordinality-awareness and hardness. To address these challenges, we propose Supervised Contrastive Learning for Medical Imaging Regression with Mixup (SupReMix). It takes anchor-inclusive mixtures (mixup of the anchor and a distinct negative sample) as hard negative pairs and anchor-exclusive mixtures (mixup of two distinct negative samples) as hard positive pairs at the embedding level. This strategy formulates harder contrastive pairs by integrating richer ordinal information. Through theoretical analysis and extensive experiments on six datasets spanning MRI, X-ray, ultrasound, and PET modalities, we demonstrate that SupReMix fosters continuous ordered representations, significantly improving regression performance.
Yilei Wu, Zijian Dong 0001, Chongyao Chen, Wangchunshu Zhou, Juan Helen Zhou
Medical Image Anal.4
2025 OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use
abstract
Xueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu, Shenzhi Wang, Xinchen Xu, Shuofei Qiao, Zhaokai Wang, Kun Kuang, Tieyong Zeng, Liang Wang, Jiwei Li, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang, Keting Yin, Zhou Zhao, Hongxia Yang, Fan Wu, Shengyu Zhang, Fei Wu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xueyu Hu, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen 0004, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao 0001, Yuhuai Li, Shengze Xu, Shenzhi Wang, Shuofei Qiao, Zhaokai Wang, Kun Kuang 0001, Tieyong Zeng, Liang Wang 0001, Jiwei Li 0001, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang 0002, Keting Yin, Zhou Zhao 0001, Hongxia Yang, Fan Wu 0006, Shengyu Zhang 0001, Fei Wu 0001
ACL (1)22
2025 PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment
abstract
Zekun Moore Wang, Shenzhi Wang, King Zhu, Jiaheng Liu, Ke Xu, Jie Fu, Wangchunshu Zhou, Wenhao Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zekun Moore Wang, Shenzhi Wang, King Zhu, Ke Xu 0001, Jie Fu 0001, Wangchunshu Zhou, Wenhao Huang 0001
ACL (1)7
2025 MIO: A Foundation Model on Multimodal Tokens
abstract
Zekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jessie Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, Wenhao Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jessie Jiashuo Wang, Ning Shi, Haoran Que, Zhaoxiang Zhang 0001, Yuanxing Zhang, Ge Zhang 0009, Ke Xu 0001, Jie Fu 0001, Wenhao Huang 0001
EMNLP4
2025 ChemAgent: Self-updating Memories in Large Language Models Improves Chemical Reasoning
abstract
Chemical reasoning usually involves complex, multi-step processes that demand precise calculations, where even minor errors can lead to cascading failures. Furthermore, large language models (LLMs) encounter difficulties handling domain-specific formulas, executing reasoning steps accurately, and integrating code ef- effectively when tackling chemical reasoning tasks. To address these challenges, we present ChemAgent, a novel framework designed to improve the performance of LLMs through a dynamic, self-updating library. This library is developed by decomposing chemical tasks into sub-tasks and compiling these sub-tasks into a structured collection that can be referenced for future queries. Then, when presented with a new problem, ChemAgent retrieves and refines pertinent information from the library, which we call memory, facilitating effective task decomposition and the generation of solutions. Our method designs three types of memory and a library-enhanced reasoning component, enabling LLMs to improve over time through experience. Experimental results on four chemical reasoning datasets from SciBench demonstrate that ChemAgent achieves performance gains of up to 46% (GPT-4), significantly outperforming existing methods. Our findings suggest substantial potential for future applications, including tasks such as drug discovery and materials science. Our code can be found at https://github.com/gersteinlab/ChemAgent.
Xiangru Tang, Muyang Ye, Yanjun Shao, Xunjian Yin, Siru Ouyang, Wangchunshu Zhou, Pan Lu, Zhuosheng Zhang 0001, Yilun Zhao 0001, Arman Cohan, Mark Gerstein
ICLR7
2025 M+: Extending MemoryLLM with Scalable Long-Term Memory
abstract
Equipping large language models (LLMs) with latent-space memory has attracted increasing attention as they can extend the context window of existing language models. However, retaining information from the distant past remains a challenge. For example, MemoryLLM (Wang et al., 2024a), as a representative work with latent-space memory, compresses past information into hidden states across all layers, forming a memory pool of 1B parameters. While effective for sequence lengths up to 16k tokens, it struggles to retain knowledge beyond 20k tokens. In this work, we address this limitation by introducing M+, a memory-augmented model based on MemoryLLM that significantly enhances long-term information retention. M+ integrates a long-term memory mechanism with a co-trained retriever, dynamically retrieving relevant information during text generation. We evaluate M+ on diverse benchmarks, including long-context understanding and knowledge retention tasks. Experimental results show that M+ significantly outperforms MemoryLLM and recent strong baselines, extending knowledge retention from under 20k to over 160k tokens with similar GPU memory overhead.
Yu Wang 0170, Dmitry Krotov, Yifan Gao 0001, Wangchunshu Zhou, Julian J. McAuley, Dan Gutfreund, Rogério Feris, Zexue He
ICML5
2025 KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
abstract
Recent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM’s general reasoning potential. To address this limitation, we introduce the **Knowledge Orthogonal Reasoning Gymnasium (KORGym)**, a dynamic evaluation platform inspired by KOR-Bench and Gymnasium. KORGym offers over fifty games in either textual or visual formats and supports interactive, multi-turn assessments with reinforcement learning scenarios. Using KORGym, we conduct extensive experiments on 19 LLMs and 8 VLMs, revealing consistent reasoning patterns within model families and demonstrating the superior performance of closed-source models. Further analysis examines the effects of modality, reasoning strategies, reinforcement learning techniques, and response length on model performance. We expect KORGym to become a valuable resource for advancing LLM reasoning research and developing evaluation methodologies suited to complex, interactive environments.
Jiajun Shi, Jian Yang 0037, Xingyuan Bu, Jiangjie Chen, Junting Zhou, Kaijing Ma, Zhoufutu Wen, Bingli Wang, Yancheng He, Hualei Zhu, Wei Zhang 0021, Ruibin Yuan, Yunli Wang, Siyuan Fang, Qianyu He, Robert Tang, Yingshui Tan, Wangchunshu Zhou, Zhaoxiang Zhang 0001, Zhoujun Li 0001, Wenhao Huang 0001, Ge Zhang 0009
NeurIPS25
2025 Training-free Periodic Interest Augmentation in Incremental Recommendation
abstract
Industrial recommender systems usually train models incrementally to grasp recent interests of users. However, a fundamental issue of these incremental updated models is their tendency to overfit current data while neglecting past information. Specifically, we have observed that the data distribution of real systems exhibits periodic drifts, leading to periodic fluctuations of prediction bias. To alleviate the above bias fluctuations while minimizing the loss of recent interests, we propose TPIA, a Training-free approach for Periodic Interest Augmentation in incremental recommendation. Specifically, after the latest model is trained, we first calculate the importance score of each model in the previous period. Then, we merge these models based on the importance scores. To minimize information loss due to interference of parameters during model merging, we further develop a method for trimming redundant and abnormal parameters. Offline experiments on both public and private datasets demonstrate the effectiveness of TPIA. It has also been deployed on a large-scale industrial recommender system, and has shown a notable 1.61% increase in CVR and a 1.97% increase in CPM, along with enhanced stability in prediction bias.
Heyuan Huang, Xingyu Lou, Changwang Zhang, Chaochao Chen 0001, Kuiyao Dong, Han Lei, Yihao Wang 0007, Wangchunshu Zhou, Jun Wang 0020
SIGIR9
2024 AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning
abstract
Shuofei Qiao, Ningyu Zhang, Runnan Fang, Yujie Luo, Wangchunshu Zhou, Yuchen Jiang, Chengfei Lv, Huajun Chen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shuofei Qiao, Ningyu Zhang 0001, Runnan Fang, Wangchunshu Zhou, Yuchen Eleanor Jiang, Chengfei Lv, Huajun Chen
ACL (1)5
2024 SmartTrim: Adaptive Tokens and Attention Pruning for Efficient Vision-Language Models
abstract
Despite achieving remarkable performance on various vision-language tasks, Transformer-based Vision-Language Models (VLMs) suffer from redundancy in inputs and parameters, significantly hampering their efficiency in real-world applications. Moreover, the degree of redundancy in token representations and model parameters, such as attention heads, varies significantly for different inputs. In light of the challenges, we propose SmartTrim, an adaptive acceleration framework for VLMs, which adjusts the computational overhead per instance. Specifically, we integrate lightweight modules into the original backbone to identify and prune redundant token representations and attention heads within each layer. Furthermore, we devise a self-distillation strategy to enhance the consistency between the predictions of the pruned model and its fully-capacity counterpart. Experimental results across various vision-language tasks consistently demonstrate that SmartTrim accelerates the original model by 2-3 times with minimal performance degradation, highlighting the effectiveness and efficiency compared to previous approaches. Code will be available at https://github.com/kugwzk/SmartTrim.
Zekun Wang 0001, Jingchang Chen, Wangchunshu Zhou, Jiafeng Liang, Liping Shan, Ming Liu 0004, Dongliang Xu, Qing Yang 0033, Bing Qin 0001
LREC/COLING3
2024 How Many Are in This Image A Safety Evaluation Benchmark for Vision LLMs
Haoqin Tu, Chenhang Cui, Yiyang Zhou, Bingchen Zhao, Junlin Han, Wangchunshu Zhou, Huaxiu Yao, Cihang Xie
ECCV (51)7
2024 OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
abstract
To help the open-source community have a better understanding of Mixture-of-Experts (MoE) based large language models (LLMs), we train and release OpenMoE, a series of fully open-sourced and reproducible decoder-only MoE LLMs, ranging from 650M to 34B parameters and trained on up to over 1T tokens. Our investigation confirms that MoE-based LLMs can offer a more favorable cost-effectiveness trade-off than dense LLMs, highlighting the potential effectiveness for future LLM development. One more important contribution of this study is an in-depth analysis of the routing mechanisms within our OpenMoE models, leading to three significant findings: Context-Independent Specialization, Early Routing Learning, and Drop-towards-the-End. We discovered that routing decisions in MoE models are predominantly based on token IDs, with minimal context relevance. The token-to-expert assignments are determined early in the pre-training phase and remain largely unchanged. This imperfect routing can result in performance degradation, particularly in sequential tasks like multi-turn conversations, where tokens appearing later in a sequence are more likely to be dropped. Finally, we rethink our design based on the above-mentioned observations and analysis. To facilitate future MoE LLM development, we propose potential strategies for mitigating the issues we found and further improving off-the-shelf MoE LLM designs.
Fuzhao Xue, Zian Zheng 0001, Jinjie Ni, Zangwei Zheng, Wangchunshu Zhou, Yang You 0001
ICML6
2024 CLUES: Collaborative Private-domain High-quality Data Selection for LLMs via Training Dynamics
abstract
Recent research has highlighted the importance of data quality in scaling large language models (LLMs). However, automated data quality control faces unique challenges in collaborative settings where sharing is not allowed directly between data silos. To tackle this issue, this paper proposes a novel data quality control technique based on the notion of data influence on the training dynamics of LLMs, that high quality data are more likely to have similar training dynamics to the anchor dataset. We then leverage the influence of the training dynamics to select high-quality data from different private domains, with centralized model updates on the server side in a collaborative training fashion by either model merging or federated learning. As for the data quality indicator, we compute the per-sample gradients with respect to the private data and the anchor dataset, and use the trace of the accumulated inner products as a measurement of data quality. In addition, we develop a quality control evaluation tailored for collaborative settings with heterogeneous medical domain data. Experiments show that training on the high-quality data selected by our method can often outperform other data selection methods for collaborative fine-tuning of LLMs, across diverse private domain datasets, in medical, multilingual and financial settings.
Wanru Zhao, Hongxiang Fan, Shell Xu Hu, Wangchunshu Zhou, Nicholas D. Lane
NeurIPS4
2024 X$^{2}$2-VLM: All-in-One Pre-Trained Model for Vision-Language Tasks
abstract
Vision language pre-training aims to learn alignments between vision and language from a large amount of data. Most existing methods only learn image-text alignments. Some others utilize pre-trained object detectors to leverage vision language alignments at the object level. In this paper, we propose to learn multi-grained vision language alignments by a unified pre-training framework that learns multi-grained aligning and multi-grained localization simultaneously. Based on it, we present X2-VLM, an all-in-one model with a flexible modular architecture, in which we further unify image-text pre-training and video-text pre-training in one model. X2-VLM is able to learn unlimited visual concepts associated with diverse text descriptions. Experiment results show that X2-VLM performs the best on base and large scale for both image-text and video-text tasks, making a good trade-off between performance and model scale. Moreover, we show that the modular design of X2-VLM results in high transferability for it to be utilized in any language or domain. For example, by simply replacing the text encoder with XLM-R, X2-VLM outperforms state-of-the-art multilingual multi-modal pre-trained models without any multilingual pre-training. The code and pre-trained models are available athttps://github.com/zengyan-97/X2-VLM.
Yan Zeng 0003, Xinsong Zhang, Hang Li 0001, Wangchunshu Zhou
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-training
abstract
In this paper, we introduce Cross-View Language Modeling, a simple and effective pretraining framework that unifies cross-lingual and cross-modal pre-training with shared architectures and objectives.Our approach is motivated by a key observation that cross-lingual and cross-modal pre-training share the same goal of aligning two different views of the same object into a common semantic space.To this end, the cross-view language modeling framework considers both multi-modal data (i.e., image-caption pairs) and multi-lingual data (i.e., parallel sentence pairs) as two different views of the same object, and trains the model to align the two views by maximizing the mutual information between them with conditional masked language modeling and contrastive learning.We pre-train CCLM, a Crosslingual Cross-modal Language Model, with the cross-view language modeling framework.Empirical results on IGLUE, a multi-lingual multi-modal benchmark, and two multi-lingual image-text retrieval datasets show that while conceptually simpler, CCLM significantly outperforms the prior state-of-the-art with an average absolute improvement of over 10%.Moreover, CCLM is the first multi-lingual multimodal pre-trained model that surpasses the translate-test performance of representative English vision-language models by zero-shot cross-lingual transfer.1
Yan Zeng 0003, Wangchunshu Zhou, Ao Luo, Ziming Cheng, Xinsong Zhang
ACL (1)2
2023 Automatic Educational Question Generation with Difficulty Level Controls
Ying Jiao, Kumar Shridhar, Peng Cui 0006, Wangchunshu Zhou, Mrinmaya Sachan
AIED4
2023 Poor Man's Quality Estimation: Predicting Reference-Based MT Metrics Without the Reference
abstract
Vilém Zouhar, Shehzaad Dhuliawala, Wangchunshu Zhou, Nico Daheim, Tom Kocmi, Yuchen Eleanor Jiang, Mrinmaya Sachan. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Vilém Zouhar, Shehzaad Dhuliawala, Wangchunshu Zhou, Nico Daheim, Tom Kocmi, Yuchen Eleanor Jiang, Mrinmaya Sachan
EACL3
2023 Doolittle: Benchmarks and Corpora for Academic Writing Formalization
abstract
Shizhe Diao, Yongyu Lei, Liangming Pan, Tianqing Fang, Wangchunshu Zhou, Sedrick Keh, Min-Yen Kan, Tong Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Shizhe Diao, Yongyu Lei, Liangming Pan, Tianqing Fang, Wangchunshu Zhou, Sedrick Keh, Min-Yen Kan, Tong Zhang 0001
EMNLP5
2023 Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models
abstract
Yifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosselut, Mrinmaya Sachan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Jiaoda Li, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosselut, Mrinmaya Sachan
EMNLP5
2023 Evaluating Large Language Models on Controlled Generation Tasks
abstract
Jiao Sun, Yufei Tian, Wangchunshu Zhou, Nan Xu, Qian Hu, Rahul Gupta, John Wieting, Nanyun Peng, Xuezhe Ma. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Jiao Sun, Yufei Tian, Wangchunshu Zhou, Rahul Gupta 0001, John Wieting, Nanyun Peng 0001, Xuezhe Ma
EMNLP3
2023 Write and Paint: Generative Vision-Language Models are Unified Modal Learners
Shizhe Diao, Wangchunshu Zhou, Xinsong Zhang
ICLR2
2023 Controlled Text Generation with Natural Language Instructions
abstract
Large language models can be prompted to pro- duce fluent output for a wide range of tasks without being specifically trained to do so. Nevertheless, it is notoriously difficult to control their generation in such a way that it satisfies user-specified constraints. In this paper, we present InstructCTG, a simple controlled text generation framework that incorporates different constraints by verbalizing them as natural language instructions. We annotate natural texts through a combination of off-the-shelf NLP tools and simple heuristics with the linguistic and extra-linguistic constraints they satisfy. Then, we verbalize the constraints into natural language instructions to form weakly supervised training data, i.e., we prepend the natural language verbalizations of the constraints in front of their corresponding natural language sentences. Next, we fine-tune a pre-trained language model on the augmented corpus. Compared to existing methods, InstructCTG is more flexible in terms of the types of constraints it allows the practitioner to use. It also does not require any modification of the decoding procedure. Finally, InstructCTG allows the model to adapt to new constraints without re-training through the use of in-context learning.
Wangchunshu Zhou, Yuchen Eleanor Jiang, Ethan Wilcox, Ryan Cotterell, Mrinmaya Sachan
ICML1
2023 To Repeat or Not To Repeat: Insights from Scaling LLM under Token-Crisis
abstract
Recent research has highlighted the importance of dataset size in scaling language models. However, large language models (LLMs) are notoriously token-hungry during pre-training, and high-quality text data on the web is likely to be approaching its scaling limit for LLMs. To further enhance LLMs, a straightforward approach is to repeat the pre-training data for additional epochs. In this study, we empirically investigate three key aspects under this approach. First, we explore the consequences of repeating pre-training data, revealing that the model is susceptible to overfitting, leading to multi-epoch degradation. Second, we examine the key factors contributing to multi-epoch degradation, finding that significant factors include dataset size, model parameters, and training objectives, while less influential factors consist of dataset quality and model FLOPs. Finally, we explore whether widely used regularization can alleviate multi-epoch degradation. Most regularization techniques do not yield significant improvements, except for dropout, which demonstrates remarkable effectiveness but requires careful tuning when scaling up the model size. Additionally, we discover that leveraging mixture-of-experts (MoE) enables cost-effective and efficient hyper-parameter tuning for computationally intensive dense LLMs with comparable trainable parameters, potentially impacting efficient LLM development on a broader scale.
Fuzhao Xue, Wangchunshu Zhou, Zangwei Zheng, Yang You 0001
NeurIPS3
2023 Scaling-up medical vision-and-language representation learning with federated learning
Tianlin Liu, Wangchunshu Zhou
Eng. Appl. Artif. Intell.4
2022 Contextual Representation Learning beyond Masked Language Modeling
abstract
How do masked language models (MLMs) such as BERT learn contextual representations?In this work, we analyze the learning dynamics of MLMs.We find that MLMs adopt sampled embeddings as anchors to estimate and inject contextual semantics to representations, which limits the efficiency and effectiveness of MLMs.To address these issues, we propose TACO, a simple yet effective representation learning approach to directly model global semantics.TACO extracts and aligns contextual semantics hidden in contextualized representations to encourage models to attend global semantics when generating contextualized representations.Experiments on the GLUE benchmark show that TACO achieves up to 5x speedup and up to 1.2 points average improvement over existing MLMs.The code is available at https:// github.com/FUZHIYI/TACO.
Zhiyi Fu, Wangchunshu Zhou, Jingjing Xu 0001, Hao Zhou 0012, Lei Li 0005
ACL (1)2
2022 BERT Learns to Teach: Knowledge Distillation with Meta Learning
abstract
We present Knowledge Distillation with Meta Learning (MetaDistil), a simple yet effective alternative to traditional knowledge distillation (KD) methods where the teacher model is fixed during training.We show the teacher network can learn to better transfer knowledge to the student network (i.e., learning to teach) with the feedback from the performance of the distilled student network in a meta learning framework.Moreover, we introduce a pilot update mechanism to improve the alignment between the inner-learner and meta-learner in meta learning algorithms that focus on an improved inner-learner.Experiments on various benchmarks show that MetaDistil can yield significant improvements compared with traditional KD algorithms and is less sensitive to the choice of different student capacity and hyperparameters, facilitating the use of KD on different tasks and models. 1
Wangchunshu Zhou, Canwen Xu, Julian J. McAuley
ACL (1)1
2022 Efficiently Tuned Parameters Are Task Embeddings
abstract
Intermediate-task transfer can benefit a wide range of NLP tasks with properly selected source datasets.However, it is computationally infeasible to experiment with all intermediate transfer combinations, making choosing a useful source task a challenging problem.In this paper, we anticipate that task-specific parameters updated in parameter-efficient tuning methods are likely to encode task-specific information.Therefore, such parameters can be predictive for inter-task transferability.Thus, we propose to exploit these efficiently tuned parameters as off-the-shelf task embeddings for the efficient selection of source datasets for intermediate-task transfer.We experiment with 11 text classification tasks and 11 question answering tasks.Experimental results show that our approach can consistently outperform existing inter-task transferability prediction methods while being conceptually simple and computationally efficient.Our analysis also reveals that the ability of efficiently tuned parameters on transferability prediction is disentangled with their in-task performance.This allows us to use parameters from early checkpoints as task embeddings to further improve efficiency. 1
Wangchunshu Zhou, Canwen Xu, Julian J. McAuley
EMNLP1
2022 VLUE: A Multi-Task Multi-Dimension Benchmark for Evaluating Vision-Language Pre-training
abstract
Recent advances in vision-language pre-training (VLP) have demonstrated impressive performance in a range of vision-language (VL) tasks. However, there exist several challenges for measuring the community’s progress in building general multi-modal intelligence. First, most of the downstream VL datasets are annotated using raw images that are already seen during pre-training, which may result in an overestimation of current VLP models’ generalization ability. Second, recent VLP work mainly focuses on absolute performance but overlooks the efficiency-performance trade-off, which is also an important indicator for measuring progress. To this end, we introduce the Vision-Language Understanding Evaluation (VLUE) benchmark, a multi-task multi-dimension benchmark for evaluating the generalization capabilities and the efficiency-performance trade-off (“Pareto SOTA”) of VLP models. We demonstrate that there is a sizable generalization gap for all VLP models when testing on out-of-distribution test sets annotated on images from a more diverse distribution that spreads across cultures. Moreover, we find that measuring the efficiency-performance trade-off of VLP models leads to complementary insights for several design choices of VLP. We release the VLUE benchmark to promote research on building vision-language models that generalize well to images unseen during pre-training and are practical in terms of efficiency-performance trade-off.
Wangchunshu Zhou, Yan Zeng 0003, Shizhe Diao, Xinsong Zhang
ICML1
2021 Learning from Perturbations: Diverse and Informative Dialogue Generation with Inverse Adversarial Training
abstract
Wangchunshu Zhou, Qifei Li, Chenle Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wangchunshu Zhou, Chenle Li
ACL/IJCNLP (1)1
2021 Beyond Preserved Accuracy: Evaluating Loyalty and Robustness of BERT Compression
abstract
Recent studies on compression of pretrained language models (e.g., BERT) usually use preserved accuracy as the metric for evaluation.In this paper, we propose two new metrics, label loyalty and probability loyalty that measure how closely a compressed model (i.e., student) mimics the original model (i.e., teacher).We also explore the effect of compression with regard to robustness under adversarial attacks.We benchmark quantization, pruning, knowledge distillation and progressive module replacing with loyalty and robustness.By combining multiple compression techniques, we provide a practical strategy to achieve better accuracy, loyalty and robustness. 1
Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Julian J. McAuley, Furu Wei
EMNLP (1)2
2021 Improving Sequence-to-Sequence Pre-training via Sequence Span Rewriting
abstract
In this paper, we propose Sequence Span Rewriting (SSR), a self-supervised task for sequence-to-sequence (Seq2Seq) pre-training.SSR learns to refine the machine-generated imperfect text spans into ground truth text.SSR provides more fine-grained and informative supervision in addition to the original textinfilling objective.Compared to the prevalent text infilling objectives for Seq2Seq pretraining, SSR is naturally more consistent with many downstream generation tasks that require sentence rewriting (e.g., text summarization, question generation, grammatical error correction, and paraphrase generation).We conduct extensive experiments by using SSR to improve the typical Seq2Seq pre-trained model T5 in a continual pre-training setting and show substantial improvements over T5 on various natural language generation tasks. 1
Wangchunshu Zhou, Tao Ge 0001, Canwen Xu, Ke Xu 0001, Furu Wei
EMNLP (1)1
2021 Pre-training Text-to-Text Transformers for Concept-centric Common Sense
Wangchunshu Zhou, Ravi Kiran Selvam, Seyeon Lee, Xiang Ren 0001
ICLR1
2021 Blow the Dog Whistle: A Chinese Dataset for Cant Understanding with Common Sense and World Knowledge
abstract
Canwen Xu, Wangchunshu Zhou, Tao Ge, Ke Xu, Julian McAuley, Furu Wei. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Julian J. McAuley, Furu Wei
NAACL-HLT2
2020 Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation Models
abstract
Automated evaluation of open domain natural language generation (NLG) models remains a challenge and widely used metrics such as BLEU and Perplexity can be misleading in some cases. In our paper, we propose to evaluate natural language generation models by learning to compare a pair of generated sentences by fine-tuning BERT, which has been shown to have good natural language understanding ability. We also propose to evaluate the model-level quality of NLG models with sample-level comparison results with skill rating system. While able to be trained in a fully self-supervised fashion, our model can be further fine-tuned with a little amount of human preference annotation to better imitate human judgment. In addition to evaluating trained models, we propose to apply our model as a performance indicator during training for better hyperparameter tuning and early-stopping. We evaluate our approach on both story generation and chit-chat dialogue response generation. Experimental results show that our model correlates better with human preference compared with previous automated evaluation approaches. Training with the proposed metric yields better performance in human evaluation, which further demonstrates the effectiveness of the proposed model.
Wangchunshu Zhou, Ke Xu 0001
AAAI1
2020 Connecting the Dots Between Fact Verification and Fake News Detection
abstract
Fact verification models have enjoyed a fast advancement in the last two years with the development of pre-trained language models like BERT and the release of large scale datasets such as FEVER.However, the challenging problem of fake news detection has not benefited from the improvement of fact verification models, which is closely related to fake news detection.In this paper, we propose a simple yet effective approach to connect the dots between fact verification and fake news detection.Our approach first employs a text summarization model pre-trained on news corpora to summarize the long news article into a short claim.Then we use a fact verification model pre-trained on the FEVER dataset to detect whether the input news article is real or fake.Our approach makes use of the recent success of fact verification models and enables zero-shot fake news detection, alleviating the need of large scale training data to train fake news detection models.Experimental results on FakenewsNet, a benchmark dataset for fake news detection, demonstrate the effectiveness of our proposed approach.
Wangchunshu Zhou
COLING2
2020 BERT-of-Theseus: Compressing BERT by Progressive Module Replacing
abstract
In this paper, we propose a novel model compression approach to effectively compress BERT by progressive module replacing.Our approach first divides the original BERT into several modules and builds their compact substitutes.Then, we randomly replace the original modules with their substitutes to train the compact modules to mimic the behavior of the original modules.We progressively increase the probability of replacement through the training.In this way, our approach brings a deeper level of interaction between the original and compact models.Compared to the previous knowledge distillation approaches for BERT compression, our approach does not introduce any additional loss function.Our approach outperforms existing knowledge distillation approaches on GLUE benchmark, showing a new perspective of model compression.1
Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Furu Wei, Ming Zhou 0001
EMNLP (1)2
2020 Self-Adversarial Learning with Comparative Discrimination for Text Generation
Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Furu Wei, Ming Zhou 0001
ICLR1
2020 Towards Interpretable Natural Language Understanding with Explanations as Latent Variables
abstract
Recently generating natural language explanations has shown very promising results in not only offering interpretable explanations but also providing additional information and supervision for prediction. However, existing approaches usually require a large set of human annotated explanations for training while collecting a large set of explanations is not only time consuming but also expensive. In this paper, we develop a general framework for interpretable natural language understanding that requires only a small set of human annotated explanations for training. Our framework treats natural language explanations as latent variables that model the underlying reasoning process of a neural model. We develop a variational EM framework for optimization where an explanation generation module and an explanation-augmented prediction module are alternatively optimized and mutually enhance each other. Moreover, we further propose an explanation-based self-training method under this framework for semi-supervised learning. It alternates between assigning pseudo-labels to unlabeled data and generating new explanations to iteratively improve each other. Experiments on two natural language understanding tasks demonstrate that our framework can not only make effective predictions in both supervised and semi-supervised settings, but is also able to generate good natural language explanations.
Wangchunshu Zhou, Jinyi Hu, Hanlin Zhang 0002, Xiaodan Liang, Maosong Sun 0001, Chenyan Xiong, Jian Tang 0005
NeurIPS1
2020 BERT Loses Patience: Fast and Robust Inference with Early Exit
abstract
In this paper, we propose Patience-based Early Exit, a straightforward yet effective inference method that can be used as a plug-and-play technique to simultaneously improve the efficiency and robustness of a pretrained language model (PLM). To achieve this, our approach couples an internal-classifier with each layer of a PLM and dynamically stops inference when the intermediate predictions of the internal classifiers do not change for a pre-defined number of steps. Our approach improves inference efficiency as it allows the model to make a prediction with fewer layers. Meanwhile, experimental results with an ALBERT model show that our method can improve the accuracy and robustness of the model by preventing it from overthinking and exploiting multiple classifiers for prediction, yielding a better accuracy-speed trade-off compared to existing early exit methods.
Wangchunshu Zhou, Canwen Xu, Tao Ge 0001, Julian J. McAuley, Ke Xu 0001, Furu Wei
NeurIPS1
2019 BERT-based Lexical Substitution
abstract
Previous studies on lexical substitution tend to obtain substitute candidates by finding the target word's synonyms from lexical resources (e.g., WordNet) and then rank the candidates based on its contexts.These approaches have two limitations: (1) They are likely to overlook good substitute candidates that are not the synonyms of the target words in the lexical resources;(2) They fail to take into account the substitution's influence on the global context of the sentence.To address these issues, we propose an end-toend BERT-based lexical substitution approach which can propose and validate substitute candidates without using any annotated data or manually curated resources.Our approach first applies dropout to the target word's embedding for partially masking the word, allowing BERT to take balanced consideration of the target word's semantics and contexts for proposing substitute candidates, and then validates the candidates based on their substitution's influence on the global contextualized representation of the sentence.Experiments show our approach performs well in both proposing and ranking substitute candidates, achieving the state-of-the-art results in both LS07 and LS14 benchmarks.
Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Furu Wei, Ming Zhou 0001
ACL (1)1