VLDB 2026 Research / reviewers in the wild / expert
Bowen Zhou 0002
dblp:61/5024-2
· DBLP profile ↗
57ranked-venue papers
1as first author
47since 2021 · last 2026
0000-0003-1062-9526ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 1 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 13 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative ReasoningabstractRecent advancements in Large Language Models (LLMs) have shown that it is promising to utilize Process Reward Models (PRMs) as verifiers to enhance the performance of LLMs. However, current PRMs face three key challenges: (1) limited process supervision and generalization capabilities, (2) dependence on scalar value prediction without leveraging the generative abilities of LLMs, and (3) inability to scale the test-time compute of PRMs. In this work, we introduce GenPRM, a generative process reward model that performs explicit Chain-of-Thought (CoT) reasoning with code verification before providing judgment for each reasoning step. To obtain high-quality process supervision labels and rationale data, we propose Relative Progress Estimation (RPE) and a rationale synthesis framework that incorporates code verification. Experimental results on ProcessBench and several mathematical reasoning tasks show that GenPRM significantly outperforms prior PRMs with only 23K training data from MATH dataset. Through test-time scaling, a 1.5B GenPRM outperforms GPT-4o, and a 7B GenPRM surpasses Qwen2.5-Math-PRM-72B on ProcessBench. Additionally, GenPRM demonstrates strong abilities to serve as a critic model for policy model refinement. This work establishes a new paradigm for process supervision that bridges the gap between PRMs and critic models in LLMs. Jian Zhao 0006, Runze Liu 0002, Zhimu Zhou, Junqi Gao, Dong Li 0016, Jiafei Lyu, Zhouyi Qian, Biqing Qi, Xiu Li 0001, Bowen Zhou 0002 |
AAAI | 11 |
| 2026 | SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language UnderstandingabstractShuang Cheng, Yuhua Jiang, Zineng Zhou, Dawei Liu, Tao Wang, Linfeng Zhang, Biqing Qi, Bowen Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shuang Cheng, Yuhua Jiang, Zineng Zhou, Linfeng Zhang 0001, Biqing Qi, Bowen Zhou 0002 |
ACL (1) | 8 |
| 2026 | Nirvana: A Specialized Generalist Model With Task-Aware Memory MechanismabstractYuhua Jiang, Shuang Cheng, Yihao Liu, Ermo Hua, Che Jiang, Weigao Sun, Yu Cheng, Feifei Gao, Biqing Qi, Bowen Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuhua Jiang, Shuang Cheng, Yihao Liu 0008, Ermo Hua, Che Jiang, Weigao Sun, Yu Cheng 0001, Biqing Qi, Bowen Zhou 0002 |
ACL (1) | 10 |
| 2026 | MARS²: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code GenerationabstractPengfei Li, Shijie Wang, Fangyuan Li, Yikun Fu, Kaifeng Liu, Kaiyan Zhang, Dazhi Zhang, Yuqiang Li, Biqing Qi, Bowen Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Pengfei Li 0011, Yikun Fu, Dazhi Zhang, Biqing Qi, Bowen Zhou 0002 |
ACL (1) | 10 |
| 2026 | I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image EditingabstractJinghan Yu, Junhao Xiao, Chenyu Zhu, Jiaming Li, Jia Li, HanMing Deng, Xirui Wang, Guoli Jia, Jianjun Li, Xiang Bai, Bowen Zhou, Zhiyuan Ma. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jinghan Yu, Chenyu Zhu, HanMing Deng, Xirui Wang, Guoli Jia, Jianjun Li 0010, Xiang Bai, Bowen Zhou 0002, Zhiyuan Ma 0005 |
ACL (1) | 11 |
| 2026 | VC-VTON: Toward Across-View and Multi-Posture-Driven Virtual Try-On via Spatiotemporal-Aware View-Consistency TrainingabstractVirtual try-on (VTON) aims to synthesize specific fashion images dressed in given garments, which possesses great potential in real-world scenarios. Existing methods generally stand on the shoulder of the single-view VTON to train a warping model and then fit the given garments onto the human body under a fixed posture and viewpoint, which often fails to preserve the consistent garment characteristics in across-view and multi-pose guided try-on scenarios due to the lack of both across-view data and effective view consistency training. To alleviate this dilemma, we propose a fresh view consistency-driven VTON task (VC-VTON) and release a multi-view virtual try-on dataset with complete annotation (e.g., viewpoint, text, posture, parsing maps, etc.) to encourage across-view training scenarios. Based on this hard-won dataset, we further propose VC-TwinNet, a Twin-UNet baseline based on spatiotemporal-aware View Consistency training, designed specifically for the challenging task. Specifically, to enable view-aware denoising and sparse-to-continuous view generalization, we introduce RoPE and circle embedding to represent the relative and continuous position relation across viewpoints, serving to distinguish their outfitting appearance and warping states. Afterwards, to implicitly learn the interactions across views under given multiple posture conditions, we further contribute a spatiotemporal-aware view attention module to capture the spatial and temporal details for across-view training. Moreover, we utilize an across-view consistency loss to supervise the model training, to ultimately improve the performance of our VC-VTON. Extensive experiments demonstrate the superiority of our approach and state-of-the-art results on various evaluations without declining single-view performance. And as for practicality and timeliness, our proposed components are essentially plug-and-play and remain effective in the new DiT-centered paradigm. Zhiyuan Ma 0005, Jiabao Wei, Zhihan Cai, Chundi Yang, Ermo Hua, Shulei Xie, Jianjun Li 0010, Bowen Zhou 0002 |
IEEE Trans. Image Process. | 8 |
| 2025 | Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesabstractRetrieval-augmented generation (RAG) has emerged to address the knowledge-intensive visual question answering (VQA) task. Current methods mainly employ separate retrieval and generation modules to acquire external knowledge and generate answers, respectively. We propose ReAuSE, an alternative to the previous RAG model for the knowledge-based VQA task, which seamlessly integrates knowledge retriever into the generative multi-modal large language model, serving as a built-in search engine. Specifically, our model functions both as a generative retriever and an accurate answer generator. It not only helps retrieve documents from the knowledge base by producing identifier for each document, but it also answers visual questions based on the retrieved documents. Furthermore, we also propose a reinforced retrieval calibration module from relevance feedback to improve retrieval performance and align with the preferences for accurate answer generation. Extensive experiments on two representative OKVQA and A-OKVQA datasets demonstrate significant improvements ranging from 2.9% to 9.6% across all evaluation metrics when compared to strong baselines. Xinwei Long, Zhiyuan Ma 0005, Ermo Hua, Biqing Qi, Bowen Zhou 0002 |
AAAI | 6 |
| 2025 | Fusing Highly Specialized Language Models for Comprehensive ExpertiseabstractNing Ding, Yulin Chen, Ganqu Cui, Xingtai Lv, Weilin Zhao, Kaiyan Zhang, Ruobing Xie, Bowen Zhou, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ning Ding 0002, Yulin Chen 0001, Ganqu Cui, Xingtai Lv, Weilin Zhao, Ruobing Xie, Bowen Zhou 0002, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 8 |
| 2025 | Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single ProcessabstractSupervised Fine-Tuning (SFT) and Preference Optimization (PO) are key processes for aligning Language Models (LMs) with human preferences post pre-training.While SFT excels in efficiency and PO in effectiveness, they are often combined sequentially without integrating their optimization objectives.This approach ignores the opportunities to bridge their paradigm gap and take the strengths from both.In this paper, we interpret SFT and PO with two subprocesses -Preference Estimation and Transition Optimization -defined at token level within the Markov Decision Process (MDP).This modeling shows that SFT is only a special case of PO with inferior estimation and optimization.PO estimates the model's preference by its entire generation, while SFT only scores model's subsequent predicted tokens based on prior tokens from ground truth answer.These priors deviates from model's distribution, hindering the preference estimation and transition optimization.Building on this view, we introduce Intuitive Fine-Tuning (IFT) to integrate SFT and PO into a single process.Through a temporal residual connection, IFT brings better estimation and optimization by capturing LMs' intuitive sense of its entire answers.But it solely relies on a single policy and the same volume of non-preference-labeled data as SFT.Our experiments show that IFT performs comparably or even superiorly to SFT and some typical PO methods across several tasks, particularly those requires generation, reasoning, and fact-following abilities.An explainable Frozen Lake game further validates the effectiveness of IFT for getting competitive policy. Ermo Hua, Biqing Qi, Xingtai Lv, Ning Ding 0002, Bowen Zhou 0002 |
ACL (1) | 7 |
| 2025 | Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent SystemabstractHaoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqing Qi, Qi Wu, Hui Li, Wanli Ouyang, Philip Torr, Bowen Zhou, Nanqing Dong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Haoyang Su 0001, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng 0001, Jinzhe Li, Biqing Qi, Hui Li 0037, Wanli Ouyang, Philip Torr 0001, Bowen Zhou 0002, Nanqing Dong |
ACL (1) | 12 |
| 2025 | Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and FeedbackabstractJiakang Yuan, Xiangchao Yan, Bo Zhang, Tao Chen, Botian Shi, Wanli Ouyang, Yu Qiao, Lei Bai, Bowen Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiakang Yuan, Xiangchao Yan, Bo Zhang 0069, Tao Chen 0003, Botian Shi, Wanli Ouyang, Yu Qiao 0001, Lei Bai 0001, Bowen Zhou 0002 |
ACL (1) | 9 |
| 2025 | Less is More: Efficient Model Merging with Binary Task SwitchabstractAs an effective approach to equip models with multitask capabilities without additional training, model merging has garnered significant attention. However, existing merging methods face challenges of redundant parameter conflicts and the excessive storage burden of fine-tuned parameters. In this work, through controlled experiments, we reveal that for fine-tuned task vectors, only those parameters with magnitudes above a certain threshold contribute positively to the task, exhibiting a pulse-like characteristic. We then attempt leveraging this pulse-like characteristic to binarize the task vectors and reduce storage overhead. Further controlled experiments show that the binarized task vectors incur almost no decrease in fine-tuning and merging performance, and even exhibit stronger performance improvements as the proportion of redundant parameters increases. Based on these insights, we propose Task Switch (T-Switch), which decomposes task vectors into three components: 1) an activation switch instantiated by a binarized mask vector, 2) a polarity switch instantiated by a binarized sign vector, and 3) a scaling knob instantiated by a scalar coefficient. By storing task vectors in a binarized form, T-Switch alleviates parameter conflicts while ensuring efficient task parameter storage. Furthermore, to enable automated switch combination in T-Switch, we further introduce Auto-Switch, which enables training-free switch combination via retrieval from a small query set. Experiments indicate that our methods achieve significant performance improvements over existing baselines, requiring only 1-3% of the storage space of full-precision parameters. Biqing Qi, Zhen Wang 0004, Junqi Gao, Dong Li 0016, Peng Ye 0006, Bowen Zhou 0002 |
CVPR | 7 |
| 2025 | ReviewRL: Towards Automated Scientific Review with RLabstractSihang Zeng, Kai Tian, Kaiyan Zhang, Yuru Wang, Junqi Gao, Runze Liu, Sa Yang, Jingxuan Li, Xinwei Long, Jiaheng Ma, Biqing Qi, Bowen Zhou. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Sihang Zeng, Yuru Wang, Junqi Gao, Runze Liu 0002, Sa Yang, Xinwei Long, Jiaheng Ma, Biqing Qi, Bowen Zhou 0002 |
EMNLP | 12 |
| 2025 | AdsQA: Towards Advertisement Video UnderstandingabstractLarge language models (LLMs) have taken a great step towards AGI. Meanwhile, an increasing number of domain-specific problems such as math and programming boost these general-purpose models to continuously evolve via learning deeper expertise. Now is thus the time further to extend the diversity of specialized applications for knowledgeable LLMs, though collecting high quality data with unexpected and informative tasks is challenging. In this paper, we propose to use advertisement (ad) videos as a challenging test-bed to probe the ability of LLMs in perceiving beyond the objective physical content of common visual domain. Our motivation is to take full advantage of the clue-rich and information-dense ad videos' traits, e.g., marketing logic, persuasive strategies, and audience engagement. Our contribution is three-fold: (1) To our knowledge, this is the first attempt to use ad videos with well-designed tasks to evaluate LLMs. We contribute AdsQA, a challenging ad Video QA benchmark derived from 1,544 ad videos with 10,962 clips, totaling 22.7 hours, providing 5 challenging tasks. (2) We propose ReAd-R, a Deepseek-R1 styled RL model that reflects on questions, and generates answers via reward-driven optimization. (3) We benchmark 14 top-tier LLMs on AdsQA, and our \texttt{ReAd-R}~achieves the state-of-the-art outperforming strong competitors equipped with long-chain reasoning capabilities by a clear margin. Xinwei Long, Peng Xu 0005, Guoli Jia, Sa Yang, Yihua Shao, Che Jiang, Jiaheng Ma, Bowen Zhou 0002 |
ICCV | 13 |
| 2025 | Advancing LLM Reasoning Generalists with Preference TreesabstractWe introduce EURUS, a suite of large language models (LLMs) optimized for reasoning. Finetuned from Mistral-7B, Llama-3-8B, and Mixtral-8x22B, EURUS models achieve state-of-the-art results among open-source models on a diverse set of benchmarks covering mathematics, code generation, and logical reasoning problems. Notably, EURUX-8X22B outperforms GPT-3.5 Turbo in reasoning through a comprehensive benchmarking across 12 test sets covering five tasks. The strong performance of EURUS can be primarily attributed to ULTRAINTERACT, our newly-curated large-scale, high-quality training data dataset specifically designed for complex reasoning tasks. ULTRAINTERACT can be used in both supervised fine-tuning, preference learning, and reward modeling. It pairs each instruction with a preference tree consisting of (1) reasoning chains with diverse planning strategies in a unified format, (2) multi-turn interaction trajectories with the environment and the critique, and (3) pairwise positive and negative responses to facilitate preference learning. ULTRAINTERACT allows us to conduct an in-depth exploration of preference learning for reasoning tasks. Our investigation reveals that some well-established preference learning algorithms may be less suitable for reasoning tasks compared to their effectiveness in general conversations. The hypothesis is that in reasoning tasks, the space of correct answers is much smaller than that of incorrect ones, so it is necessary to explicitly increase the reward of chosen data. Therefore, in addition to increasing the reward margin as many preference learning algorithms do, the absolute values of positive responses’ rewards should be positive and may serve as a proxy for performance. Inspired by this, we derive a novel reward modeling objective and empirically that it leads to a stable reward modeling curve and better performance. Together with ULTRAINTERACT, we obtain a strong reward model. Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding 0002, Xingyao Wang 0002, Boji Shan, Zeyuan Liu, Ruobing Xie, Yankai Lin 0001, Zhenghao Liu 0001, Bowen Zhou 0002, Hao Peng 0015, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICLR | 13 |
| 2025 | OpenPRM: Building Open-domain Process-based Reward Models with Preference TreesabstractScaling inference-time computation is increasingly seen as the next frontier in scaling laws for large language models. Previous work in mathematics and coding has demonstrated the remarkable potential for inference-time scaling. During such scaling, fine-grained supervision through process-based reward models (PRMs) is essential for enhancement. However, exploration of inference-time scaling and PRMs in open-domain problems remains limited, where lacking exact answers and obtaining process supervision prove challenging. In this paper, we explore the construction of PRMs for open-domain tasks, specifically for instruction-following tasks. Utilizing existing outcome-based reward models (ORMs), we develop sentence-level preference trees based on the prefix similarity of parallel sampled candidates from datasets like UltraFeedback. This setup allows us to derive weak supervision for processes via back-propagation from outcome-level rewards. Subsequently, we integrate ORMs and PRMs under the same pairwise ranking objectives, resulting in our newly developed reward models, named OpenPRM. This approach significantly enhances the scalability of process-level supervision in open domains at minimal cost. We assess the performance of OpenPRM across various reward benchmarks, demonstrating its competitive edge over traditional ORMs in open domains and PRMs in specialized domains. Additionally, we investigate the scalability of inference-time computation for open-domain instructions. Our results highlight the limitations of ORMs’ scalability, while OpenPRM shows superior performance in scaled settings. Despite these advances, achieving automatic fine-grained supervision for open-domain inference-time scaling remains a substantial challenge. We hope these findings will spur further development of process supervision reward models in open-domain scenarios. Jiayuan Zhang 0001, Haoxin Li, Xuekai Zhu, Ermo Hua, Xingtai Lv, Ning Ding 0002, Biqing Qi, Bowen Zhou 0002 |
ICLR | 9 |
| 2025 | Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length GeneralizationabstractExtending the context length of Language Models (LMs) by improving Rotary Position Embedding (RoPE) has become a trend. While prior works mainly address RoPE’s limitations within attention, this paper uncovers the adverse effects on length generalization from nearly all parts of LMs. Using Discrete Signal Processing theory, we show that RoPE enables periodic attention by implicitly achieving Non-Uniform Discrete Fourier Transform. However, this periodicity is undermined by the spectrum damage caused by: 1) linear layers and activation functions outside of attention; 2) insufficiently trained frequency components brought by time-domain truncation. Building on our observations, we propose Fourier Position Embedding (FoPE), which enhances attention’s frequency-domain properties to improve both its periodic extension and length generalization. FoPE constructs Fourier Series and zero-outs the destructive frequency components, increasing model robustness against the spectrum damage. Experiments across various model scales and benchmarks show that, within varying context windows, FoPE maintains a more stable performance compared to other baselines. Several analyses and ablations bring further support to our method and theoretical modeling. Ermo Hua, Che Jiang, Xingtai Lv, Youbang Sun, Yuchen Fan 0001, Xuekai Zhu, Biqing Qi, Ning Ding 0002, Bowen Zhou 0002 |
ICML | 10 |
| 2025 | Free Process Rewards without Process LabelsabstractDifferent from its counterpart outcome reward models (ORMs), which evaluate the entire responses, a process reward model (PRM) scores a reasoning trajectory step by step, providing denser and more fine-grained rewards. However, training a PRM requires labels annotated at every intermediate step, presenting significant challenges for both manual and automatic data collection. This paper aims to address this challenge. Both theoretically and empirically, we show that an implicit PRM can be obtained at no additional cost, by simply training an ORM on the cheaper response-level labels. The only assumption is to parameterize the outcome reward as the log-likelihood ratios of the policy and reference models r$\phi$(y) = $\beta$ log $\pi$$\phi$(y) $\pi$ref(y) , which can be optimized regardless of the specific choice of loss objectives. In experiments, we instantiate our implicit PRMs with various objectives and evaluate their performance on MATH. We show that our implicit PRM outperforms a strong MCTS-based baseline á la Math-Shepherd (Wang et al., 2023) using less than 1/38 of the training data. Its performance can be further improved with majority voting. We further find that scaling up instructions and responses benefits our implicit PRM, and the latter brings a larger gain. Particularly, we find that our implicit PRM, when instantiated with the cross-entropy (CE) loss, is more data-efficient and can keep improving generation models even when trained with only one response per instruction, the setup that suffers from extreme data scarcity and imbalance. Further, instructions should be relevant to downstream tasks while the diversity of responses does not bring gains. Surprisingly, training on extra Math-Shepherd step labels brings no further improvements to our implicit PRM trained on only outcome data. We hope that our work will encourage a rethinking of PRM training approaches and contribute to making training PRMs more accessible. Lifan Yuan, Wendi Li, Huayu Chen, Ganqu Cui, Ning Ding 0002, Bowen Zhou 0002, Zhiyuan Liu 0001, Hao Peng 0001 |
ICML | 7 |
| 2025 | How to Synthesize Text Data without Model Collapse?abstractModel collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance. With the proliferation of AI models, synthetic data will fundamentally reshape the web data ecosystem. Future GPT-$\{n\}$ models will inevitably be trained on a blend of synthetic and human-produced data. In this paper, we focus on two questions: what is the impact of synthetic data on language model training, and how to synthesize data without model collapse? We first pre-train language models across different proportions of synthetic data, revealing a negative correlation between the proportion of synthetic data and model performance. We further conduct statistical analysis on synthetic data to uncover distributional shift phenomenon and over-concentration of n-gram features. Inspired by the above findings, we propose token editing on human-produced data to obtain semi-synthetic data. As a proof of concept, we theoretically demonstrate that token-level editing can prevent model collapse, as the test error is constrained by a finite upper bound. We conduct extensive experiments on pre-training from scratch, continual pre-training, and supervised fine-tuning. The results validate our theoretical proof that token-level editing improves data quality and enhances model performance. Xuekai Zhu, Daixuan Cheng, Hengli Li, Ermo Hua, Xingtai Lv, Ning Ding 0002, Zhouhan Lin, Zilong Zheng, Bowen Zhou 0002 |
ICML | 10 |
| 2025 | MedXpertQA: Benchmarking Expert-Level Medical Reasoning and UnderstandingabstractWe introduce MedXpertQA, a highly challenging and comprehensive benchmark to evaluate expert-level medical knowledge and advanced reasoning. MedXpertQA includes 4,460 questions spanning 17 specialties and 11 body systems. It includes two subsets, Text for text evaluation and MM for multimodal evaluation. Notably, MM introduces expert-level exam questions with diverse images and rich clinical information, including patient records and examination results, setting it apart from traditional medical multimodal benchmarks with simple QA pairs generated from image captions. MedXpertQA applies rigorous filtering and augmentation to address the insufficient difficulty of existing benchmarks like MedQA, and incorporates specialty board questions to improve clinical relevance and comprehensiveness. We perform data synthesis to mitigate data leakage risk and conduct multiple rounds of expert reviews to ensure accuracy and reliability. We evaluate 18 leading models on MedXpertQA. Moreover, medicine is deeply connected to real-world decision-making, providing a rich and representative setting for assessing reasoning abilities beyond mathematics and code. To this end, we develop a reasoning-oriented subset to facilitate the assessment of o1-like models. Yuxin Zuo, Shang Qu, Zhang-Ren Chen, Xuekai Zhu, Ermo Hua, Ning Ding 0002, Bowen Zhou 0002 |
ICML | 9 |
| 2025 | Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language ModelsabstractLarge Language Models (LLMs) have shown strong abilities in general language tasks, yet adapting them to specific domains remains a challenge.
Current method like Domain Adaptive Pretraining (DAPT) requires costly full-parameter training and suffers from catastrophic forgetting.
Meanwhile, Retrieval-Augmented Generation (RAG) introduces substantial inference latency due to expensive nearest-neighbor searches and longer context.
This paper introduces \textit{Memory Decoder}, a plug-and-play pretrained memory that enables efficient domain adaptation without changing the original model's parameters.
Memory Decoder employs a small transformer decoder that learns to imitate the behavior of an external non-parametric retriever.
Once trained, Memory Decoder can be seamlessly integrated with any pretrained language model that shares the same tokenizer, requiring no model-specific modifications.
Experimental results demonstrate that Memory Decoder enables effective adaptation of various Qwen and Llama models to three distinct specialized domains: biomedicine, finance, and law, reducing perplexity by an average of 6.17 points.
Overall, Memory Decoder introduces a novel paradigm centered on a specially pretrained memory component designed for domain-specific adaptation. This memory architecture can be integrated in a plug-and-play manner, consistently enhancing performance across multiple models within the target domain. Jiaqi Cao 0002, Rubin Wei, Qipeng Guo, Kai Chen 0026, Bowen Zhou 0002, Zhouhan Lin |
NeurIPS | 6 |
| 2025 | DePass: Unified Feature Attributing by Simple Decomposed Forward PassabstractAttributing the behavior of Transformer models to internal computations is a central challenge in mechanistic interpretability. We introduce DePass, a unified framework for feature attribution based on a single decomposed forward pass. DePass decomposes hidden states into customized additive components, then propagates them with attention scores and MLP's activations fixed. It achieves faithful, fine-grained attribution without requiring auxiliary training. We validate DePass across token-level, model component-level, and subspace-level attribution tasks, demonstrating its effectiveness and fidelity. Our experiments highlight its potential to attribute information flow between arbitrary components of a Transformer model. We hope DePass serves as a foundational tool for broader applications in interpretability. Xiangyu Hong, Che Jiang, Biqing Qi, Youbang Sun, Ning Ding 0002, Bowen Zhou 0002 |
NeurIPS | 7 |
| 2025 | TTRL: Test-Time Reinforcement LearningabstractThis paper investigates Reinforcement Learning (RL) on data without explicit labels for reasoning tasks in Large Language Models (LLMs). The core challenge of the problem is reward estimation during inference while not having access to ground-truth information. While this setting appears elusive, we find that common practices in Test-Time Scaling (TTS), such as majority voting, yield surprisingly effective rewards suitable for driving RL training. In this work, we introduce Test-Time Reinforcement Learning (TTRL), a novel method for training LLMs using RL on unlabeled data. TTRL enables self-evolution of LLMs by utilizing the priors in the pre-trained models. Our experiments demonstrate that TTRL consistently improves performance across a variety of tasks and models. Notably, TTRL boosts the pass@1 performance of Qwen-2.5-Math-7B by approximately 211% on the AIME 2024 with only unlabeled test data. Furthermore, although TTRL is only supervised by the Maj@N metric, TTRL has demonstrated performance to consistently surpass the upper limit of the initial model, and approach the performance of models trained directly on test data with ground-truth labels. Our experimental findings validate the general effectiveness of TTRL across various tasks and highlight TTRL's potential for broader tasks and domains. Yuxin Zuo, Shang Qu, Ganqu Cui, Xuekai Zhu, Haozhan Li, Xinwei Long, Ermo Hua, Biqing Qi, Youbang Sun, Zhiyuan Ma 0005, Lifan Yuan, Ning Ding 0002, Bowen Zhou 0002 |
NeurIPS | 16 |
| 2025 | Efficient Diffusion Models: A Comprehensive Survey From Principles to PracticesabstractAs one of the most popular and sought-after generative models in recent years, diffusion models have sparked the interests of many researchers and steadily shown excellent advantage in various generative tasks such as image synthesis, video generation, bioinformatics engineering, 3D scene rendering and multimodal generation, relying on their dense theoretical principles and reliable application practices. The remarkable success of these recent efforts on diffusion models comes largely from progressive design principles and efficient architecture, training, inference, and deployment methodologies. However, there has not been a comprehensive and in-depth review to summarize these principles and practices to help the rapid understanding and application of diffusion models. In this survey, we provide a new efficiency-oriented perspective on these existing efforts, which mainly focuses on the profound principles and efficient practices in architecture designs, model training, fast inference and reliable deployment, to guide further theoretical research, algorithm migration and model application for new scenarios in a reader-friendly way. Zhiyuan Ma 0005, Yuzhu Zhang, Guoli Jia, Yichao Ma, Gaofeng Liu, Ning Ding 0002, Jianjun Li 0010, Bowen Zhou 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 11 |
| 2025 | Guest Editorial: Introduction to the Special Section on Large-Scale Multimodal Learning: Universality, Robustness, Efficiency, and Beyond
Peng Xu 0005, Song Bai 0001, Bowen Zhou 0002, David A. Clifton, Andrea Vedaldi, Mihaela van der Schaar, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Contrastive Augmented Graph2Graph Memory Interaction for Few Shot Continual LearningabstractFew-Shot Class-Incremental Learning (FSCIL) has gained considerable attention in recent years for its pivotal role in addressing continuously arriving classes. However, it encounters additional challenges. The scarcity of samples in new sessions intensifies overfitting, causing incompatibility between the output features of new and old classes, thereby escalating catastrophic forgetting. A prevalent strategy involves mitigating catastrophic forgetting through the Explicit Memory (EM), which comprise of class prototypes. However, current EM-based methods retrieves memory globally by performing Vector-to-Vector (V2V) interaction between features corresponding to the input and prototypes stored in EM, neglecting the geometric structure of local features. This hinders the accurate modeling of their positional relationships. To incorporate information of local geometric structure, we extend the V2V interaction to Graph-to-Graph (G2G) interaction. For enhancing local structures for better G2G alignment and the prevention of local feature collapse, we propose the Local Graph Preservation (LGP) mechanism. Additionally, to address sample scarcity in classes from new sessions, the Contrast-Augmented G2G (CAG2G) is introduced to promote the aggregation of same class features thus helps few-shot learning. Extensive comparisons on CIFAR100, CUB200, and the challenging ImageNet-R dataset demonstrate the superiority of our method over existing methods. Biqing Qi, Junqi Gao, Dong Li 0016, Jianxing Liu, Ligang Wu 0001, Bowen Zhou 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Generative Multi-Modal Knowledge Retrieval with Large Language ModelsabstractKnowledge retrieval with multi-modal queries plays a crucial role in supporting knowledge-intensive multi-modal applications. However, existing methods face challenges in terms of their effectiveness and training efficiency, especially when it comes to training and integrating multiple retrievers to handle multi-modal queries. In this paper, we propose an innovative end-to-end generative framework for multi-modal knowledge retrieval. Our framework takes advantage of the fact that large language models (LLMs) can effectively serve as virtual knowledge bases, even when trained with limited data. We retrieve knowledge via a two-step process: 1) generating knowledge clues related to the queries, and 2) obtaining the relevant document by searching databases using the knowledge clue. In particular, we first introduce an object-aware prefix-tuning technique to guide multi-grained visual learning. Then, we align multi-grained visual features into the textual feature space of the LLM, employing the LLM to capture cross-modal interactions. Subsequently, we construct instruction data with a unified format for model training. Finally, we propose the knowledge-guided generation strategy to impose prior constraints in the decoding steps, thereby promoting the generation of distinctive knowledge clues. Through experiments conducted on three benchmarks, we demonstrate significant improvements ranging from 3.0% to 14.6% across all evaluation metrics when compared to strong baselines. Xinwei Long, Jiali Zeng, Fandong Meng, Zhiyuan Ma 0005, Bowen Zhou 0002, Jie Zhou 0016 |
AAAI | 6 |
| 2024 | AdapEdit: Spatio-Temporal Guided Adaptive Editing Algorithm for Text-Based Continuity-Sensitive Image EditingabstractWith the great success of text-conditioned diffusion models in creative text-to-image generation, various text-driven image editing approaches have attracted the attentions of many researchers. However, previous works mainly focus on discreteness-sensitive instructions such as adding, removing or replacing specific objects, background elements or global styles (i.e., “hard editing”), while generally ignoring subject-binding but semantically fine-changing continuity-sensitive instructions such as actions, poses or adjectives, and so on (i.e., “soft editing”), which hampers generative AI from generating user-customized visual contents. To mitigate this predicament, we propose a spatio-temporal guided adaptive editing algorithm AdapEdit, which realizes adaptive image editing by introducing a soft-attention strategy to dynamically vary the guiding degree from the editing conditions to visual pixels from both temporal and spatial perspectives. Note our approach has a significant advantage in preserving model priors and does not require model training, fine-tuning, extra data, or optimization. We present our results over a wide variety of raw images and editing instructions, demonstrating competitive performance and showing it significantly outperforms the previous approaches. Code is available: https://github.com/AnonymousPony/adap-edit. Zhiyuan Ma 0005, Guoli Jia, Bowen Zhou 0002 |
AAAI | 3 |
| 2024 | LMD: Faster Image Reconstruction with Latent Masking DiffusionabstractAs a class of fruitful approaches, diffusion probabilistic models (DPMs) have shown excellent advantages in high-resolution image reconstruction. On the other hand, masked autoencoders (MAEs), as popular self-supervised vision learners, have demonstrated simpler and more effective image reconstruction and transfer capabilities on downstream tasks. However, they all require extremely high training costs, either due to inherent high temporal-dependence (i.e., excessively long diffusion steps) or due to artificially low spatial-dependence (i.e., human-formulated high mask ratio, such as 0.75). To the end, this paper presents LMD, a faster image reconstruction framework with Latent Masking Diffusion. First, we propose to project and reconstruct images in latent space through a pre-trained variational autoencoder, which is theoretically more efficient than in the pixel-based space. Then, we combine the advantages of MAEs and DPMs to design a progressive masking diffusion model, which gradually increases the masking proportion by three different schedulers and reconstructs the latent features from simple to difficult, without sequentially performing denoising diffusion as in DPMs or using fixed high masking ratio as in MAEs, so as to alleviate the high training time-consumption predicament. Our approach allows for learning high-capacity models and accelerate their training (by 3x or more) and barely reduces the original accuracy. Inference speed in downstream tasks also significantly outperforms the previous approaches. Zhiyuan Ma 0005, Zhihuan Yu, Jianjun Li 0010, Bowen Zhou 0002 |
AAAI | 4 |
| 2024 | CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction FollowingabstractWith the advancement of language models (LMs), their exposure to private data is increasingly inevitable, and their deployment (especially for smaller ones) on personal devices, such as PCs and smartphones, has become a prevailing trend.In contexts laden with user information, enabling models to both safeguard user privacy and execute commands efficiently emerges as an essential research imperative.In this paper, we propose CoGenesis, a collaborative generation framework integrating large (hosted on cloud infrastructure) and small models (deployed on local devices) to address privacy concerns logically.Initially, we design a pipeline to create personalized writing instruction datasets enriched with extensive context details as the testbed of this research issue.Subsequently, we introduce two variants of CoGenesis based on sketch and logits respectively.Our experimental findings, based on our synthesized dataset and two additional open-source datasets, indicate that: 1) Large-scale models perform well when provided with user context but struggle in the absence of such context.2) While specialized smaller models fine-tuned on the synthetic dataset show promise, they still lag behind their larger counterparts.3) Our CoGenesis framework, utilizing mixed-scale models, showcases competitive performance, providing a feasible solution to privacy issues.* Corresponding author 1 This paper defines large LMs (LLMs) as both closed and open-source models, designed for universal application and advanced performance, and intended for cloud deployment.Conversely, small LMs (SLMs) refer to models tailored for specific tasks and deployed on local devices. Jianyu Wang 0012, Ermo Hua, Biqing Qi, Ning Ding 0002, Bowen Zhou 0002 |
ACL (1) | 6 |
| 2024 | Empowering Private Tutoring by Chaining Large Language ModelsabstractArtificial intelligence has been applied in various aspects of online education to facilitate teaching and learning. However, few approaches have been made towards a complete AI-powered tutoring system. In this work, we explore the development of a full-fledged intelligent tutoring system based on large language models (LLMs). The proposed system ChatTutor, powered by state-of-the-art LLMs, is equipped with automatic course planning and adjusting, informative instruction, and adaptive quiz offering and evaluation. ChatTutor is decomposed into three inter-connected core processes: interaction, reflection, and reaction. Each process is implemented by chaining LLM-powered tools along with dynamically updated memory modules. To demonstrate the mechanism of each working module and the benefits of structured memory control and adaptive reflection, we conduct a wide range of analysis based on statistical results and user study. The analysis shows the designed processes boost system consistency and stability under long-term interaction and intentional disruptions, with up to 5% and 20% increase in performance respectively. Meanwhile, we also compare the system with scripts from real-world online learning platform and discuss the potential issues unique to LLM-based systems. Yulin Chen 0001, Ning Ding 0002, Hai-Tao Zheng 0002, Zhiyuan Liu 0001, Maosong Sun 0001, Bowen Zhou 0002 |
CIKM | 6 |
| 2024 | Interactive Continual Learning: Fast and Slow ThinkingabstractAdvanced life forms, sustained by the synergistic interaction of neural cognitive mechanisms, continually acquire and transfer knowledge throughout their lifespan. In contrast, contemporary machine learning paradigms exhibit limitations in emulating the facets of continual learning (CL). Nonetheless, the emergence of large language models (LLMs) presents promising avenues for realizing CL via interactions with these models. Drawing on Complementary Learning System theory, this paper presents a novel Interactive Continual Learning (ICL) framework, enabled by collaborative interactions among models of various sizes. Specifically, we assign the ViT model as System1 and multimodal LLM as System2. To enable the memory module to deduce tasks from class information and enhance Set2Set retrieval, we propose the Class-Knowledge-Task Multi-Head Attention (CKT-MHA). Additionally, to improve memory retrieval in System1 through enhanced geometric representation, we introduce the CL-vMF mechanism, based on the von Mises-Fisher (vMF) distribution. Mean-while, we introduce the von Mises-Fisher Outlier Detection and Interaction (vMF-ODI) strategy to identify hard examples, thus enhancing collaboration between System1 and System2 for complex reasoning realization. Comprehensive evaluation of our proposed ICL demonstrates significant resistance to forgetting and superior performance relative to existing methods. Code is available at github.com/ICL. Biqing Qi, Junqi Gao, Dong Li 0016, Jianxing Liu, Ligang Wu 0001, Bowen Zhou 0002 |
CVPR | 7 |
| 2024 | LAKE-RED: Camouflaged Images Generation by Latent Background Knowledge Retrieval-Augmented DiffusionabstractCamouflaged vision perception is an important vision task with numerous practical applications. Due to the expensive collection and labeling costs, this community struggles with a major bottleneck that the species category of its datasets is limited to a small number of object species. However, the existing camouflaged generation methods require specifying the background manually, thus failing to extend the camouflaged sample diversity in a low-cost manner. In this paper, we propose a Latent Background Knowledge Retrieval-Augmented Diffusion (LAKE-RED) for camouflaged image generation. To our knowledge, our contributions mainly include: (1) For the first time, we propose a camouflaged generation paradigm that does not need to re-eive any background inputs. (2) Our LAKE-RED is the first knowledge retrieval-augmented method with interpretability for camouflaged generation, in which we propose an idea that knowledge retrieval and reasoning enhancement are separated explicitly, to alleviate the task-specific chal-lenges. Moreover, our method is not restricted to specific foreground targets or backgrounds, offering a potential for extending camouflaged vision perception to more diverse domains. (3) Experimental results demonstrate that our method outperforms the existing approaches, generating more realistic camouflage images. Our source code is released on https://github.com/PanchengZhaoILAKE-RED. Pancheng Zhao, Peng Xu 0005, Pengda Qin, Deng-Ping Fan, Guoli Jia, Bowen Zhou 0002, Jufeng Yang |
CVPR | 7 |
| 2024 | MSI-Agent: Incorporating Multi-Scale Insight into Embodied Agents for Superior Planning and Decision-MakingabstractLong-term memory is significant for agents, in which insights play a crucial role.However, the emergence of irrelevant insight and the lack of general insight can greatly undermine the effectiveness of insight.To solve this problem, in this paper, we introduce Multi-Scale Insight Agent (MSI-Agent), an embodied agent designed to improve LLMs' planning and decision-making ability by summarizing and utilizing insight effectively across different scales.MSI achieves this through the experience selector, insight generator, and insight selector.Leveraging a three-part pipeline, MSI can generate task-specific and high-level insight, store it in a database, and then use relevant insight from it to aid in decisionmaking.Our experiments show that MSI outperforms another insight strategy when planning by GPT3.5.Moreover, We delve into the strategies for selecting seed experience and insight, aiming to provide LLM with more useful and relevant insight for better decision-making.Our observations also indicate that MSI exhibits better robustness when facing domainshifting scenarios. Dayuan Fu, Biqing Qi, Yihuai Gao, Che Jiang, Guanting Dong 0001, Bowen Zhou 0002 |
EMNLP | 6 |
| 2024 | Scalable Efficient Training of Large Language Models with Low-dimensional Projected AttentionabstractImproving the effectiveness and efficiency of large language models (LLMs) simultaneously is a critical yet challenging research goal.In this paper, we find that low-rank pre-training, normally considered as efficient methods that will compromise performance, can be scalably effective when reduced parameters are precisely targeted.Specifically, applying the low-dimensional module only to the attention layer -resolves this issue and enhances both effectiveness and efficiency.We refer to this structure as Low-dimensional Projected Attention (LPA) and provide an explanatory analysis.Through extensive experimentation at parameter scales of 130M, 370M, and scaling up to 3B, we have validated the effectiveness and scalability of LPA.Our results show that LPA model can save up to 12.4% in time while achieving an approximate 5% improvement in test perplexity (ppl) and on downstream tasks compared with the vanilla Transformer. Xingtai Lv, Ning Ding 0002, Ermo Hua, Ganqu Cui, Bowen Zhou 0002 |
EMNLP | 6 |
| 2024 | Safe-SD: Safe and Traceable Stable Diffusion with Text Prompt Trigger for Invisible Generative WatermarkingabstractRecently, stable diffusion (SD) models have typically flourished in the field of image synthesis and personalized editing, with a range of photorealistic and unprecedented images being successfully generated. As a result, widespread interest has been ignited to develop and use various SD-based tools for visual content creation. However, the exposure of AI-created content on public platforms could raise both legal and ethical risks. In this regard, the traditional methods of adding watermarks to the already generated images (i.e. post-processing) may face a dilemma (e.g., being erased or modified) in terms of copyright protection and content monitoring, since the powerful image inversion and text-to-image editing techniques have been widely explored in SD-based methods. In this work, we propose a Safe and high-traceable Stable Diffusion framework (namely Safe-SD) to adaptively implant the graphical watermarks (e.g., QR code) into the imperceptible structure-related pixels during the generative diffusion process for supporting text-driven invisible watermarking and detection. Different from the previous high-cost injection-then-detection training framework, we design a simple and unified architecture, which makes it possible to simultaneously train watermark injection and detection in a single network, greatly improving the efficiency and convenience of use. Moreover, to further support text-driven generative watermarking and deeply explore its robustness and high-traceability, we elaborately design a λ-sampling and λ-encryption algorithm to fine-tune a latent diffuser wrapped by a VAE for balancing high-fidelity image synthesis and high-traceable watermark detection. We present our quantitative and qualitative results on two representative datasets LSUN, COCO and FFHQ, demonstrating state-of-the-art performance of Safe-SD and showing it significantly outperforms the previous approaches. Zhiyuan Ma 0005, Guoli Jia, Biqing Qi, Bowen Zhou 0002 |
ACM Multimedia | 4 |
| 2024 | On Large Language Models' Hallucination with Regard to Known FactsabstractChe Jiang, Biqing Qi, Xiangyu Hong, Dayuan Fu, Yang Cheng, Fandong Meng, Mo Yu, Bowen Zhou, Jie Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Che Jiang, Biqing Qi, Xiangyu Hong, Dayuan Fu, Fandong Meng, Mo Yu, Bowen Zhou 0002, Jie Zhou 0016 |
NAACL-HLT | 8 |
| 2024 | PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuningabstractXuekai Zhu, Biqing Qi, Kaiyan Zhang, Xinwei Long, Zhouhan Lin, Bowen Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xuekai Zhu, Biqing Qi, Xinwei Long, Zhouhan Lin, Bowen Zhou 0002 |
NAACL-HLT | 6 |
| 2024 | Neural Residual Diffusion Models for Deep Scalable Vision GenerationabstractThe most advanced diffusion models have recently adopted increasingly deep stacked networks (e.g., U-Net or Transformer) to promote the generative emergence capabilities of vision generation models similar to large language models (LLMs). However, progressively deeper stacked networks will intuitively cause numerical propagation errors and reduce noisy prediction capabilities on generative data, which hinders massively deep scalable training of vision generation models. In this paper, we first uncover the nature that neural networks being able to effectively perform generative denoising lies in the fact that the intrinsic residual unit has consistent dynamic property with the input signal's reverse diffusion process, thus supporting excellent generative abilities.
Afterwards, we stand on the shoulders of two common types of deep stacked networks to propose a unified and massively scalable Neural Residual Diffusion Models framework (Neural-RDM for short), which is a simple yet meaningful change to the common architecture of deep generative networks by introducing a series of learnable gated residual parameters that conform to the generative dynamics. Experimental results on various generative tasks show that the proposed neural residual models obtain state-of-the-art scores on image's and video's generative benchmarks. Rigorous theoretical proofs and extensive experiments also demonstrate the advantages of this simple gated residual mechanism consistent with dynamic modeling in improving the fidelity and consistency of generated content and supporting large-scale scalable training. Zhiyuan Ma 0005, Biqing Qi, Bowen Zhou 0002 |
NeurIPS | 4 |
| 2024 | Exploring Adversarial Robustness of Deep State Space ModelsabstractDeep State Space Models (SSMs) have proven effective in numerous task scenarios but face significant security challenges due to Adversarial Perturbations (APs) in real-world deployments. Adversarial Training (AT) is a mainstream approach to enhancing Adversarial Robustness (AR) and has been validated on various traditional DNN architectures. However, its effectiveness in improving the AR of SSMs remains unclear.
While many enhancements in SSM components, such as integrating Attention mechanisms and expanding to data-dependent SSM parameterizations, have brought significant gains in Standard Training (ST) settings, their potential benefits in AT remain unexplored. To investigate this, we evaluate existing structural variants of SSMs with AT to assess their AR performance. We observe that pure SSM structures struggle to benefit from AT, whereas incorporating Attention yields a markedly better trade-off between robustness and generalization for SSMs in AT compared to other components. Nonetheless, the integration of Attention also leads to Robust Overfitting (RO) issues.
To understand these phenomena, we empirically and theoretically analyze the output error of SSMs under AP. We find that fixed-parameterized SSMs have output error bounds strictly related to their parameters, limiting their AT benefits, while input-dependent SSMs may face the problem of error explosion. Furthermore, we show that the Attention component effectively scales the output error of SSMs during training, enabling them to benefit more from AT, but at the cost of introducing RO due to its high model complexity.
Inspired by this, we propose a simple and effective Adaptive Scaling (AdS) mechanism that brings AT performance close to Attention-integrated SSMs without introducing the issue of RO. Biqing Qi, Yiang Luo, Junqi Gao, Pengfei Li 0011, Zhiyuan Ma 0005, Bowen Zhou 0002 |
NeurIPS | 7 |
| 2024 | UltraMedical: Building Specialized Generalists in BiomedicineabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities across various domains and are moving towards more specialized areas. Recent advanced proprietary models such as GPT-4 and Gemini have achieved significant advancements in biomedicine, which have also raised privacy and security challenges. The construction of specialized generalists hinges largely on high-quality datasets, enhanced by techniques like supervised fine-tuning and reinforcement learning from human or AI feedback, and direct preference optimization. However, these leading technologies (e.g., preference learning) are still significantly limited in the open source community due to the scarcity of specialized data. In this paper, we present the UltraMedical collections, which consist of high-quality manual and synthetic datasets in the biomedicine domain, featuring preference annotations across multiple advanced LLMs. By utilizing these datasets, we fine-tune a suite of specialized medical models based on Llama-3 series, demonstrating breathtaking capabilities across various medical benchmarks. Moreover, we develop powerful reward models skilled in biomedical and general reward benchmark, enhancing further online preference learning within the biomedical LLM community. Sihang Zeng, Ermo Hua, Ning Ding 0002, Zhang-Ren Chen, Zhiyuan Ma 0005, Haoxin Li, Ganqu Cui, Biqing Qi, Xuekai Zhu, Xingtai Lv, Jinfang Hu, Zhiyuan Liu 0001, Bowen Zhou 0002 |
NeurIPS | 14 |
| 2024 | Enhancing Adversarial Transferability via Information Bottleneck ConstraintsabstractFrom the perspective of information bottleneck (IB) theory, we propose a novel framework for performing black-box transferable adversarial attacks named IBTA, which leverages advancements in invariant features. Intuitively, diminishing the reliance of adversarial perturbations on the original data, under equivalent attack performance constraints, encourages a greater reliance on invariant features that contributes most to classification, thereby enhancing the transferability of adversarial attacks. Building on this motivation, we redefine the optimization of transferable attacks using a novel theoretical framework that centers around IB. Specifically, to overcome the challenge of unoptimizable mutual information, we propose a simple and efficient mutual information lower bound (MILB) for approximating computation. Moreover, to quantitatively evaluate mutual information, we utilize the Mutual Information Neural Estimator (MINE) to perform a thorough analysis. Our experiments on the ImageNet dataset well demonstrate the efficiency and scalability of IBTA and derived MILB. Our code is available at github.com/IBTA. Biqing Qi, Junqi Gao, Jianxing Liu, Ligang Wu 0001, Bowen Zhou 0002 |
IEEE Signal Process. Lett. | 5 |
| 2024 | Improving Robustness of Intent Detection Under Adversarial Attacks: A Geometric Constraint PerspectiveabstractDeep neural networks (DNNs)-based natural language processing (NLP) systems are vulnerable to being fooled by adversarial examples presented in recent studies. Intent detection tasks in dialog systems are no exception, however, relatively few works have been attempted on the defense side. The combination of linear classifier and softmax is widely used in most defense methods for other NLP tasks. Unfortunately, it does not encourage the model to learn well-separated feature representations. Thus, it is easy to induce adversarial examples. In this article, we propose a simple, yet efficient defense method from the geometric constraint perspective. Specifically, we first propose an M-similarity metric to shrink variances of intraclass features. Intuitively, better geometric conditions of feature space can bring lower misclassification probability (MP). Therefore, we derive the optimal geometric constraints of anchors within each category from the overall MP (OMP) with theoretical guarantees. Due to the nonconvex characteristic of the optimal geometric condition, it is hard to satisfy the traditional optimization process. To this end, we regard such geometric constraints as manifold optimization processes in the Stiefel manifold, thus naturally avoiding the above challenges. Experimental results demonstrate that our method can significantly improve robustness compared with baselines, while retaining the excellent performance on normal examples. Biqing Qi, Bowen Zhou 0002, Weinan Zhang 0003, Jianxing Liu, Ligang Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Enhancing Chat Language Models by Scaling High-quality Instructional ConversationsabstractFine-tuning on instruction data has been widely validated as an effective practice for implementing chat language models like ChatGPT.Scaling the diversity and quality of such data, although straightforward, stands a great chance of leading to improved performance.This paper aims to push the upper bound of opensource models further.We first provide a systematically designed, diverse, informative, large-scale dataset of instructional conversations, UltraChat, which does not involve human queries.Our objective is to capture the breadth of interactions between a human user and an AI assistant and employs a comprehensive framework to generate multi-turn conversation iteratively.UltraChat contains 1.5 million high-quality multi-turn dialogues and covers a wide range of topics and instructions.Our statistical analysis of UltraChat reveals its superiority in various key metrics, including scale, average length, diversity, coherence, etc., solidifying its position as a leading opensource dataset.Building upon UltraChat, we fine-tune a LLaMA model to create a powerful conversational model, UltraLM.Our evaluations indicate that UltraLM consistently outperforms other open-source models, including WizardLM and Vicuna, the previously recognized state-of-the-art open-source models. Ning Ding 0002, Yulin Chen 0001, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu 0001, Maosong Sun 0001, Bowen Zhou 0002 |
EMNLP | 8 |
| 2023 | Sparse Low-rank Adaptation of Pre-trained Language ModelsabstractFine-tuning pre-trained large language models in a parameter-efficient manner is widely studied for its effectiveness and efficiency.The popular method of low-rank adaptation (LoRA) offers a notable approach, hypothesizing that the adaptation process is intrinsically low-dimensional.Although LoRA has demonstrated commendable performance, it is implemented with a fixed and unalterable intrinsic rank that might not always be the ideal choice.Recognizing the need for more flexible adaptation, we extend the methodology of LoRA to an innovative approach we call sparse low-rank adaptation (SoRA) that enables dynamic adjustments to the intrinsic rank during the adaptation process.We achieve this through the incorporation of a gate unit optimized with proximal gradient method in the training stage, controlling the cardinality of rank under the sparsity of the gate.In the subsequent inference stage, we eliminate the parameter blocks corresponding to the zeroed-out ranks, to reduce each SoRA module back to a concise yet rankoptimal LoRA.Our approach strengthens the representation power of LoRA by initializing it with a higher rank, while efficiently taming a temporarily increased number of parameters via updating in a sparse way.We further introduce a sparsifying scheduler for SoRA, aiming to examine the impact of the number of nonzero parameters on the model's memorization and generalization.Our experimental results demonstrate that SoRA can outperform other baselines even with 70% retained parameters and 70% training time. Ning Ding 0002, Xingtai Lv, Qiaosen Wang, Yulin Chen 0001, Bowen Zhou 0002, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP | 5 |
| 2023 | CRaSh: Clustering, Removing, and Sharing Enhance Fine-tuning without Full Large Language ModelabstractInstruction tuning has recently been recognized as an effective way of aligning Large Language Models (LLMs) to enhance their generalization ability across various tasks.However, when tuning publicly accessible, centralized LLMs with private instruction data, privacy concerns are inevitable.While direct transfer of parameterized modules between models is a plausible approach to address this, its implications and effectiveness need further exploration.This paper focuses on Offsite-Tuning (OFT), a representative technique that transfers transformer blocks between centralized LLMs and downstream emulators.Given the limited understanding of the underlying mechanism of OFT, we perform an empirical analysis on LLMs from the perspectives of representation and functional similarity.Interestingly, our findings reveal a unique modular structure within the layers of LLMs that appears to emerge as the model size expands.Simultaneously, we note subtle but potentially significant changes in representation and intermediate predictions across the layers.Inspired by these observations, we propose CRaSh, involving Clustering, Removing, and Sharing, a training-free strategy to derive improved emulators from LLMs.CRaSh significantly boosts performance of OFT with billions of parameters.Furthermore, we investigate the optimal solutions yielded by fine-tuning with and without full model through the lens of loss landscape.Our findings demonstrate a linear connectivity among these optima falling over the same basin, thereby highlighting the effectiveness of CRaSh and OFT.The source code is publicly available at https://github.com/TsinghuaC3I/CRaSh. Ning Ding 0002, Biqing Qi, Xuekai Zhu, Xinwei Long, Bowen Zhou 0002 |
EMNLP | 6 |
| 2023 | SAM struggles in concealed scenes - empirical study on "Segment Anything"
Ge-Peng Ji, Deng-Ping Fan, Peng Xu 0005, Bowen Zhou 0002, Ming-Ming Cheng, Luc Van Gool |
Sci. China Inf. Sci. | 4 |
| 2018 | R3: Reinforced Ranker-Reader for Open-Domain Question AnsweringabstractIn recent years researchers have achieved considerable success applying neural network methods to question answering (QA). These approaches have achieved state of the art results in simplified closed-domain settings such as the SQuAD (Rajpurkar et al. 2016) dataset, which provides a pre-selected passage, from which the answer to a given question may be extracted. More recently, researchers have begun to tackle open-domain QA, in which the model is given a question and access to a large corpus (e.g., wikipedia) instead of a pre-selected passage (Chen et al. 2017a). This setting is more complex as it requires large-scale search for relevant passages by an information retrieval component, combined with a reading comprehension model that “reads” the passages to generate an answer to the question. Performance in this setting lags well behind closed-domain performance. In this paper, we present a novel open-domain QA system called Reinforced Ranker-Reader (R3), based on two algorithmic innovations. First, we propose a new pipeline for open-domain QA with a Ranker component, which learns to rank retrieved passages in terms of likelihood of extracting the ground-truth answer to a given question. Second, we propose a novel method that jointly trains the Ranker along with an answer-extraction Reader model, based on reinforcement learning. We report extensive experimental results showing that our method significantly improves on the state of the art for multiple open-domain QA datasets. Shuohang Wang, Mo Yu, Tim Klinger, Wei Zhang 0057, Shiyu Chang, Gerald Tesauro, Bowen Zhou 0002, Jing Jiang 0001 |
AAAI | 9 |
| 2017 | Multiresolution Recurrent Neural Networks: An Application to Dialogue Response GenerationabstractWe introduce a new class of models called multiresolution recurrent neural networks, which explicitly model natural language generation at multiple levels of abstraction. The models extend the sequence-to-sequence framework to generate two parallel stochastic processes: a sequence of high-level coarse tokens, and a sequence of natural language words (e.g. sentences). The coarse sequences follow a latent stochastic process with a factorial representation, which helps the models generalize to new examples. The coarse sequences can also incorporate task-specific knowledge, when available. In our experiments, the coarse sequences are extracted using automatic procedures, which are designed to capture compositional structure and semantics. These procedures enable training the multiresolution recurrent neural networks by maximizing the exact joint log-likelihood over both sequences. We apply the models to dialogue response generation in the technical support domain and compare them with several competing models. The multiresolution recurrent neural networks outperform competing models by a substantial margin, achieving state-of-the-art results according to both a human evaluation study and automatic evaluation metrics. Furthermore, experiments show the proposed models generate more fluent, relevant and goal-oriented responses. Iulian Serban, Tim Klinger, Gerald Tesauro, Kartik Talamadupula, Bowen Zhou 0002, Yoshua Bengio, Aaron C. Courville |
AAAI | 5 |
| 2017 | Improved Neural Relation Detection for Knowledge Base Question AnsweringabstractRelation detection is a core component of many NLP applications including Knowledge Base Question Answering (KBQA).In this paper, we propose a hierarchical recurrent neural network enhanced by residual learning which detects KB relations given an input question.Our method uses deep residual bidirectional LSTMs to compare questions and relation names via different levels of abstraction.Additionally, we propose a simple KBQA system that integrates entity linking and our proposed relation detector to make the two components enhance each other.Our experimental results show that our approach not only achieves outstanding relation detection performance, but more importantly, it helps our KBQA system achieve state-of-the-art accuracy for both single-relation (SimpleQuestions) and multi-relation (WebQSP) QA benchmarks. Mo Yu, Wenpeng Yin 0001, Kazi Saidul Hasan, Cícero Nogueira dos Santos, Bing Xiang, Bowen Zhou 0002 |
ACL (1) | 6 |
| 2017 | GaDei: On Scale-Up Training as a Service for Deep LearningabstractDeep learning (DL) training-as-a-service (TaaS) is an important emerging industrial workload. TaaS must satisfy a wide range of customers who have no experience and/or resources to tune DL hyper-parameters (e.g., mini-batch size and learning rate), and meticulous tuning for each user's dataset is prohibitively expensive. Therefore, TaaS hyper-parameters must be fixed with values that are applicable to all users. Unfortunately, few research papers have studied how to design a system for TaaS workloads. By evaluating the IBM Watson Natural Language Classfier (NLC) workloads, the most popular IBM cognitive service used by thousands of enterprise-level clients globally, we provide empirical evidence that only the conservative hyper-parameter setup (e.g., small mini-batch size) can guarantee acceptable model accuracy for a wide range of customers. Unfortunately, smaller mini-batch size requires higher communication bandwidth in a parameter-server based DL training system. In this paper, we characterize the exceedingly high communication bandwidth requirement of TaaS using representative industrial deep learning workloads. We then present GaDei, a highly optimized shared-memory based scale-up parameter server design. We evaluate GaDei using both commercial benchmarks and public benchmarks and demonstrate that GaDei significantly outperforms the state-of-the-art parameter-server based implementation while maintaining the required accuracy. GaDei achieves near-best-possible runtime performance, constrained only by the hardware limitation. Furthermore, to the best of our knowledge, GaDei is the only scale-up DL system that provides fault-tolerance. Wei Zhang 0057, Minwei Feng, Yunhui Zheng, Yufei Ren, Yandong Wang 0001, Peng Liu 0010, Bing Xiang, Li Zhang 0002, Bowen Zhou 0002, Fei Wang 0001 |
ICDM | 10 |
| 2017 | A Structured Self-Attentive Sentence Embedding
Zhouhan Lin, Minwei Feng, Cícero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou 0002, Yoshua Bengio |
ICLR (Poster) | 6 |
| 2016 | Pointing the Unknown WordsabstractThe problem of rare and unknown words is an important issue that can potentially effect the performance of many NLP systems, including traditional count-based and deep learning models.We propose a novel way to deal with the rare and unseen words for the neural network models using attention.Our model uses two softmax layers in order to predict the next word in conditional language models: one predicts the location of a word in the source sentence, and the other predicts a word in the shortlist vocabulary.At each timestep, the decision of which softmax layer to use is adaptively made by an MLP which is conditioned on the context.We motivate this work from a psychological evidence that humans naturally have a tendency to point towards objects in the context or the environment when the name of an object is not known.Using our proposed model, we observe improvements on two tasks, neural machine translation on the Europarl English to French parallel corpora and text summarization on the Gigaword dataset. Caglar Gulcehre, Sungjin Ahn, Ramesh Nallapati, Bowen Zhou 0002, Yoshua Bengio |
ACL (1) | 4 |
| 2016 | Simple Question Answering by Attentive Convolutional Neural NetworkabstractThis work focuses on answering single-relation factoid questions over Freebase. Each question can acquire the answer from a single fact of form (subject, predicate, object) in Freebase. This task, simple question answering (SimpleQA), can be addressed via a two-step pipeline: entity linking and fact selection. In fact selection, we match the subject entity in a fact candidate with the entity mention in the question by a character-level convolutional neural network (char-CNN), and match the predicate in that fact with the question by a word-level CNN (word-CNN). This work makes two main contributions. (i) A simple and effective entity linker over Freebase is proposed. Our entity linker outperforms the state-of-the-art entity linker over SimpleQA task. (ii) A novel attentive maxpooling is stacked over word-CNN, so that the predicate representation can be matched with the predicate-focused question representation more effectively. Experiments show that our system sets new state-of-the-art in this task. Wenpeng Yin 0001, Mo Yu, Bing Xiang, Bowen Zhou 0002, Hinrich Schütze |
COLING | 4 |
| 2016 | Abstractive Text Summarization using Sequence-to-sequence RNNs and BeyondabstractIn this work, we model abstractive text summarization using Attentional Encoder-Decoder Recurrent Neural Networks, and show that they achieve state-of-the-art performance on two different corpora.We propose several novel models that address critical problems in summarization that are not adequately modeled by the basic architecture, such as modeling key-words, capturing the hierarchy of sentence-toword structure, and emitting words that are rare or unseen at training time.Our work shows that many of our proposed models contribute to further improvement in performance.We also propose a new dataset consisting of multi-sentence summaries, and establish performance benchmarks for further research. Ramesh Nallapati, Bowen Zhou 0002, Cícero Nogueira dos Santos, Caglar Gulcehre, Bing Xiang |
CoNLL | 2 |
| 2016 | ABCNN: Attention-Based Convolutional Neural Network for Modeling Sentence PairsabstractHow to model a pair of sentences is a critical issue in many NLP tasks such as answer selection (AS), paraphrase identification (PI) and textual entailment (TE). Most prior work (i) deals with one individual task by fine-tuning a specific system; (ii) models each sentence’s representation separately, rarely considering the impact of the other sentence; or (iii) relies fully on manually designed, task-specific linguistic features. This work presents a general Attention Based Convolutional Neural Network (ABCNN) for modeling a pair of sentences. We make three contributions. (i) The ABCNN can be applied to a wide variety of tasks that require modeling of sentence pairs. (ii) We propose three attention schemes that integrate mutual influence between sentences into CNNs; thus, the representation of each sentence takes into consideration its counterpart. These interdependent sentence pair representations are more powerful than isolated sentence representations. (iii) ABCNNs achieve state-of-the-art performance on AS, PI and TE tasks. We release code at: https://github.com/yinwenpeng/Answer_Selection . Wenpeng Yin 0001, Hinrich Schütze, Bing Xiang, Bowen Zhou 0002 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2013 | The IBM speech-to-speech translation system for smartphone: Improvements for resource-constrained tasks
Bowen Zhou 0002, Songfang Huang, Martin Cmejrek, Wei Zhang 0057, Jia Cui, Bing Xiang, Gregg Daggett, Upendra V. Chaudhari, Sameer Maskey, Etienne Marcheret |
Comput. Speech Lang. | 1 |