EDBT 2026 Demo / reviewers in the wild / expert
Weipeng Chen
dblp:213/1678
· DBLP profile ↗
30ranked-venue papers
1as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 1 first-author · 20 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | G-SAN: Rethinking Audio Navigation Beyond Time in PodcastsabstractPodcasts have become a primary medium for information access and informal learning. However, mainstream audio players still rely on time-based navigation, treating listening as moving through time. This mismatch forces listeners to translate comprehension regulation intentions into trial-and-error seeking on the timeline. While generative AI enables rich semantic analysis of long-form audio, existing tools often demand visual attention and fine-grained interaction, which is a poor fit for mobile, multitasking listening. We present Generative Semantic Audio Navigation (G-SAN), a prototype that brings support for comprehension regulation from the information layer to the control layer by remapping familiar playback controls into meaning-oriented actions over semantic units for local clarification, main-thread grasping, and global orientation. A user study shows that participants were able to understand and adopt these controls during real-time listening, and generally perceived them as helpful for regulating understanding, while also revealing boundary conditions. We conclude with design implications for meaning-oriented audio navigation. Liuziyu Zhao, Weipeng Chen, Pan Hui 0001 |
DIS | 3 |
| 2026 | Hidden Labor behind the Hype: Understanding AI Side Hustles through Platform Narratives and Worker PracticesabstractAI side hustles are increasingly promoted on social media as accessible, empowering, and profitable opportunities. This paper examines the gap between such platform narratives and workers’ lived experiences through a mixed-method study of 7,938 RedNote posts and 16 semi-structured interviews. Our analysis identifies monetization typologies and rhetorical strategies that portray AI work as simple and rewarding, while interview data reveal hidden labor, unstable income, and the devaluation of human contributions. By juxtaposing platform narratives with lived experiences, we show how these narratives structurally foreground ease and reward while downplaying the precarity embedded in actual AI work. This study contributes a critical account of how AI side hustles are framed and experienced, and offers design implications for HCI: platforms should moderate promotional content and provide clearer risk communication, while designers of human–AI collaboration tools should highlight and value human input rather than allowing it to remain invisible. Weipeng Chen, Corey Kewei Xu, Pan Hui 0001 |
CHI | 3 |
| 2026 | AITQE: An Adaptive Image-Text Quality Enhancer for Scalable MLLM PretrainingabstractMultimodal large language models (MLLMs) have made significant strides by integrating visual and textual modalities. A critical factor in training MLLMs is the quality of image-text pairs within multimodal pretraining datasets. However, in the process of high-quality data curation, filter-based paradigms often discard a substantial portion of high-quality images due to inadequate semantic alignment between images and texts, leading to inefficiency in data utilization and scalability. In this paper, we propose the Adaptive Image-Text Quality Enhancer (AITQE), a model that dynamically assesses and enhances the quality of image-text pairs. AITQE employs a text rewriting mechanism for low-quality pairs and incorporates a negative sample learning strategy to improve evaluative capabilities by integrating deliberately generated low-quality samples during training. Unlike prior approaches that significantly alter text distributions, our method minimally adjusts text to preserve data volume while enhancing quality. Experimental results demonstrate that AITQE surpasses existing methods on various benchmarks, effectively leveraging raw data and scaling with increasing data volumes. Codes and model are available at https://github.com/hanhuang22/AITQE. Yuqi Huo, Zijia Zhao, Haoyu Lu, Bingning Wang, Qiang Liu 0006, Weipeng Chen, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration DistillationabstractZican Dong, Junyi Li, Jinhao Jiang, Mingyu Xu, Xin Zhao, Bingning Wang, Weipeng Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zican Dong, Jinhao Jiang, Bingning Wang, Weipeng Chen |
ACL (1) | 7 |
| 2025 | CFBench: A Comprehensive Constraints-Following Benchmark for LLMsabstractTao Zhang, ChengLIn Zhu, Yanjun Shen, Wenjing Luo, Yan Zhang, Hao Liang, Tao Zhang, Fan Yang, Mingan Lin, Yujing Qiao, Weipeng Chen, Bin Cui, Wentao Zhang, Zenan Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tao Zhang 0194, Chenglin Zhu, Wenjing Luo, Yan Zhang 0109, Hao Liang 0017, Fan Yang 0132, Yujing Qiao, Weipeng Chen, Bin Cui 0001, Wentao Zhang 0001, Zenan Zhou |
ACL (1) | 10 |
| 2025 | RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented GenerationabstractRetrieval-augmented generation (RAG) effectively addresses issues of static knowledge and hallucination in large language models. Existing studies mostly focus on question scenarios with clear user intents and concise answers. However, it is prevalent that users issue broad, open-ended queries with diverse sub-intents, for which they desire rich and long-form answers covering multiple relevant aspects. To tackle this important yet underexplored problem, we propose a novel RAG framework, namely RichRAG. It includes a sub-aspect explorer to identify potential sub-aspects of input questions, a multi-faceted retriever to build a candidate pool of diverse external documents related to these sub-aspects, and a generative list-wise ranker, which is a key module to provide the top-k most valuable documents for the final generator. These ranked documents sufficiently cover various query aspects and are aware of the generator’s preferences, hence incentivizing it to produce rich and comprehensive responses for users. The training of our ranker involves a supervised fine-tuning stage to ensure the basic coverage of documents, and a reinforcement learning stage to align downstream LLM’s preferences to the ranking of documents. Experimental results on two publicly available datasets prove that our framework effectively and efficiently provides comprehensive and satisfying responses to users. Shuting Wang 0002, Weipeng Chen, Yutao Zhu 0001, Zhicheng Dou |
COLING | 4 |
| 2025 | Improving Accuracy and Calibration via Differentiated Deep Mutual LearningabstractDeep Neural Networks (DNNs) have achieved remarkable success in a variety of tasks, particularly in terms of prediction accuracy. However, in real-world scenarios, especially in safety-critical applications, accuracy alone is insufficient; reliable uncertainty estimates are essential. Modern DNNs, often trained with cross-entropy loss, tend to exhibit overconfidence, especially on ambiguous samples. Many techniques aim to improve uncertainty calibration, yet they often come at the cost of reduced accuracy or increased computational demands. To address this challenge, we propose Differentiated Deep Mutual Learning (Diff-DML), an efficient ensemble approach that simultaneously enhances accuracy and uncertainty calibration. Diff-DML draws inspiration from Deep Mutual Learning (DML) while introducing two strategies to maintain prediction diversity: (1) Differentiated Training Strategy (DTS) and (2) Diversity-Preserving Learning Objective (DPLO). Our theoretical analysis shows that Diff-DML’s diversified learning framework not only leverages ensemble benefits but also avoids the loss of prediction diversity observed in traditional DML setups, which is crucial for improved calibration. Extensive evaluations on various benchmarks confirm the effectiveness of Diff-DML. For instance, on the CIFAR-100 dataset, Diff-DML on ResNet34/50 models achieved substantial improvements over the previous state-of-the-art method, MDCA, with absolute accuracy gains of 1.3%/3.1%, relative ECE reductions of 49.6%/43.8%, and relative classwise-ECE reductions of 7.7%/13.0%. Peng Cui 0007, Bingning Wang, Weipeng Chen, Jun Zhu 0001, Xiaolin Hu 0001 |
CVPR | 4 |
| 2025 | Efficient Motion-Aware Video MLLMabstractMost current video MLLMs rely on uniform frame sampling and image-level encoders, resulting in inefficient data processing and limited motion awareness. To address these challenges, we introduce EMA, an Efficient Motion-Aware video MLLM that utilizes compressed video structures as inputs. We propose a motion-aware GOP (Group of Pictures) encoder that fuses spatial and motion information within a GOP unit in the compressed video stream, generating compact, informative visual tokens. By integrating fewer but denser RGB frames with more but sparser motion vectors in this native slow-fast input architecture, our approach reduces redundancy and enhances motion representation. Additionally, we introduce MotionBench, a benchmark for evaluating motion understanding across four motion types: linear, curved, rotational, and contact-based. Experimental results show that EMA achieves state-of-the-art performance on both MotionBench and popular video question answering benchmarks, while reducing inference costs. Moreover, EMA demonstrates strong scalability, as evidenced by its competitive performance on long video understanding benchmarks. Zijia Zhao, Yuqi Huo, Tongtian Yue, Longteng Guo, Haoyu Lu, Bingning Wang, Weipeng Chen, Jing Liu 0001 |
CVPR | 7 |
| 2025 | Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual KnowledgeabstractDoes seeing always mean knowing? Large Vision-Language Models (LVLMs) integrate separately pre-trained vision and language components, often using CLIP-ViT as vision backbone. However, these models frequently encounter a core issue of "cognitive misalignment" between the vision encoder (VE) and the large language model (LLM). Specifically, the VE’s representation of visual information may not fully align with LLM’s cognitive framework, leading to a mismatch where visual features exceed the language model’s interpretive range. To address this, we investigate how variations in VE representations influence LVLM comprehension, especially when the LLM faces VE-Unknown data—images whose ambiguous visual representations challenge the VE’s interpretive precision. Accordingly, we construct a multi-granularity landmark dataset and systematically examine the impact of VE-Known and VE-Unknown data on interpretive abilities. Our results show that VE-Unknown data limits LVLM’s capacity for accurate understanding, while VE-Known data, rich in distinctive features, helps reduce cognitive misalignment. Building on these insights, we propose Entity-Enhanced Cognitive Alignment (EECA), a method that employs multi-granularity supervision to generate visually enriched, well-aligned tokens that not only integrate within the embedding space but also align with the LLM’s cognitive framework. This alignment markedly enhances LVLM performance in landmark recognition. Our findings underscore the challenges posed by VE-Unknown data and highlight the essential role of cognitive alignment in advancing multimodal systems. Yuanyang Yin, Victor Shea-Jay Huang, Weipeng Chen, Baoqun Yin, Zenan Zhou |
CVPR | 7 |
| 2025 | Extracting and Combining Abilities For Building Multi-lingual Ability-enhanced Large Language ModelsabstractMulti-lingual ability transfer has become increasingly important for the broad application of large language models (LLMs).Existing work highly relies on training with the multi-lingual ability-related data, which may not be available for low-resource languages.To solve it, we propose a Multilingual Abilities Extraction and Combination approach (MAEC), which decomposes and extracts language-agnostic ability-related weights from LLMs, and combines them across different languages by simple addition and subtraction operations without training.Specifically, our MAEC consists of the extraction and combination stages.In the extraction stage, we firstly locate key neurons that are highly related to specific abilities, and then employ them to extract the transferable ability-related weights.In the combination stage, we further select the ability-related tensors that mitigate the linguistic effects, and design a combining strategy based on them and the languagespecific weights, to build the multi-lingual ability-enhanced LLM.To assess the effectiveness of our approach, we conduct extensive experiments on LLaMA-3 8B on mathematical and scientific tasks in both high-resource and low-resource lingual scenarios.Empirical results have shown that MAEC can effectively and efficiently extract and combine the advanced abilities, achieving comparable performance with PaLM.Resources are available at https://github.com/RUCAIBox/MAET. Zhipeng Chen 0001, Kun Zhou 0002, Wayne Xin Zhao, Bingning Wang, Weipeng Chen, Ji-Rong Wen |
EMNLP | 6 |
| 2025 | FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human FeedbackabstractHuman feedback is crucial in the interactions between humans and Large Language Models (LLMs). However, existing research primarily focuses on benchmarking LLMs in single-turn dialogues. Even in benchmarks designed for multi-turn dialogues, the user utterances are often independent, neglecting the nuanced and complex nature of human feedback within real-world usage scenarios. To fill this research gap, we introduce FB-Bench, a fine-grained, multi-task benchmark designed to evaluate LLMs’ responsiveness to human feedback under real-world usage scenarios in Chinese. Drawing from the two main interaction scenarios, FB-Bench comprises 591 meticulously curated samples, encompassing eight task types, five deficiency types of response, and nine feedback types. We extensively evaluate a broad array of popular LLMs, revealing significant variations in their performance across different interaction scenarios. Further analysis indicates that task, human feedback, and deficiencies of previous responses can also significantly impact LLMs’ responsiveness. Our findings underscore both the strengths and limitations of current models, providing valuable insights and directions for future research. Youquan Li, Miao Zheng, Fan Yang 0132, Guosheng Dong, Bin Cui 0001, Weipeng Chen, Zenan Zhou, Wentao Zhang 0001 |
EMNLP | 6 |
| 2025 | DataSculpt: A Holistic Data Management Framework for Long-Context LLMs TrainingabstractIn recent years, foundation models, particularly large language models (LLMs), have demonstrated significant improvements across a variety of tasks. One of their most important features is long-context capability, which enables them to generate extended text with high semantic coherence, retrieving relevant information, and handling tasks with substantial amounts of text efficiently. The key to improving long-context performance lies in effective data organization and management strategies that integrate data from multiple domains and optimize the context window during training. Through extensive experimental analysis, we identified three key challenges in designing effective data management strategies that enable the model to achieve long-context capability without sacrificing performance in other tasks: (1) a shortage of long documents across multiple domains, (2) effective construction of context windows, and (3) efficient organization of large-scale datasets. To address these challenges, we introduce DataSculpt, a novel data management framework designed for long-context training. We first formulate the organization of training data as a multi-objective optimization problem, focusing on attributes including the relevance among documents within the same training sequence, the quantity of concatenated instances, individual document integrity, and computational cost. Specifically, our approach utilizes a coarse-to-fine method to optimize training data organization effectively. We begin by clustering the data based on semantic similarity (coarse), followed by a multi-objective greedy search within each cluster to score and concatenate documents into various context windows (fine). We have deployed DataSculpt as the data management backend for long-context training in Baichuan Inc. Extensive experiments with diverse downstream tasks show that DataSculpt enhances the model's long-context performance by an average of 15.73%, while maintaining the general capabilities with a 4.63% improvement. Keer Lu, Xiaonan Nie, Da Pan 0003, Shusen Zhang, Keshi Zhao, Weipeng Chen, Zenan Zhou, Guosheng Dong, Bin Cui 0001, Wentao Zhang 0001 |
ICDE | 7 |
| 2025 | Exploring the Design Space of Visual Context Representation in Video MLLMsabstractVideo Multimodal Large Language Models~(MLLMs) have shown remarkable capability of understanding the video semantics on various downstream tasks. Despite the advancements, there is still a lack of systematic research on visual context representation, which refers to the scheme to select frames from a video and further select the tokens from a frame. In this paper, we explore the design space for visual context representation, and aim to improve the performance of video MLLMs by finding more effective representation schemes. Firstly, we formulate the task of visual context representation as a constrained optimization problem, and model the language modeling loss as a function of the number of frames and the number of embeddings (or tokens) per frame, given the maximum visual context window size. Then, we explore the scaling effects in frame selection and token selection respectively, and fit the corresponding function curve by conducting extensive empirical experiments. We examine the effectiveness of typical selection strategies and present empirical findings to determine the two factors. Furthermore, we study the joint effect of frame selection and token selection, and derive the optimal formula for determining the two factors. We demonstrate that the derived optimal settings show alignment with the best-performed results of empirical experiments. The data and code are available at: https://github.com/RUCAIBox/Opt-Visor. Yifan Du 0002, Yuqi Huo, Kun Zhou 0002, Zijia Zhao, Haoyu Lu, Wayne Xin Zhao, Bingning Wang, Weipeng Chen, Ji-Rong Wen |
ICLR | 9 |
| 2025 | Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction TuningabstractLarge Language Models (LLMs) have exhibited significant potential in performing diverse tasks, including the ability to call functions or use external tools to enhance their performance. While current research on function calling by LLMs primarily focuses on single-turn interactions, this paper addresses the overlooked necessity for LLMs to engage in multi-turn function calling—critical for handling compositional, real-world queries that require planning with functions but not only use functions. To facilitate this, we introduce an approach, BUTTON, which generates synthetic compositional instruction tuning data via bottom-up instruction construction and top-down trajectory generation. In the bottom-up phase, we generate simple atomic tasks based on real-world scenarios and build compositional tasks using heuristic strategies based on atomic tasks. Corresponding function definitions are then synthesized for these compositional tasks. The top-down phase features a multi-agent environment where interactions among simulated humans, assistants, and tools are utilized to gather multi-turn function calling trajectories. This approach ensures task compositionality and allows for effective function and trajectory generation by examining atomic tasks within compositional tasks. We produce a dataset BUTTONInstruct comprising 8k data points and demonstrate its effectiveness through extensive experiments across various LLMs. Mingyang Chen 0002, Haoze Sun, Tianpeng Li, Fan Yang 0132, Hao Liang 0017, Keer Lu, Bin Cui 0001, Wentao Zhang 0001, Zenan Zhou, Weipeng Chen |
ICLR | 10 |
| 2025 | SysBench: Can LLMs Follow System Message?abstractLarge Language Models (LLMs) have become instrumental across various applications, with the customization of these models to specific scenarios becoming increasingly critical. System message, a fundamental component of LLMs, is consist of carefully crafted instructions that guide the behavior of model to meet intended goals. Despite the recognized potential of system messages to optimize AI-driven solutions, there is a notable absence of a comprehensive benchmark for evaluating how well LLMs follow system messages. To fill this gap, we introduce SysBench, a benchmark that systematically analyzes system message following ability in terms of three limitations of existing LLMs: constraint violation, instruction misjudgement and multi-turn instability. Specifically, we manually construct evaluation dataset based on six prevalent types of constraints, including 500 tailor-designed system messages and multi-turn user conversations covering various interaction relationships. Additionally, we develop a comprehensive evaluation protocol to measure model performance. Finally, we conduct extensive evaluation across various existing LLMs, measuring their ability to follow specified constraints given in system messages. The results highlight both the strengths and weaknesses of existing models, offering key insights and directions for future research. Yanzhao Qin, Tao Zhang 0194, Wenjing Luo, Haoze Sun, Yan Zhang 0109, Yujing Qiao, Weipeng Chen, Zenan Zhou, Wentao Zhang 0001, Bin Cui 0001 |
ICLR | 9 |
| 2025 | Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMsabstractVideo understanding is a crucial next step for multimodal large language models (MLLMs).
Various benchmarks are introduced for better evaluating the MLLMs.
Nevertheless, current video benchmarks are still inefficient for evaluating video models during iterative development due to the high cost of constructing datasets and the difficulty in isolating specific skills.
In this paper, we propose VideoNIAH (Video Needle in A Haystack), a benchmark construction framework through synthetic video generation.
VideoNIAH decouples video content from their query-responses by inserting unrelated visual 'needles' into original videos.
The framework automates the generation of query-response pairs using predefined rules, minimizing manual labor. The queries focus on specific aspects of video understanding, enabling more skill-specific evaluations. The separation between video content and the queries also allow for increased video variety and evaluations across different lengths.
Utilizing VideoNIAH, we compile a video benchmark, VNBench, which includes tasks such as retrieval, ordering, and counting to evaluate three key aspects of video understanding: temporal perception, chronological ordering, and spatio-temporal coherence. We conduct a comprehensive evaluation of both proprietary and open-source models, uncovering significant differences in their video understanding capabilities across various tasks. Additionally, we perform an in-depth analysis of the test results and model configurations. Based on these findings, we provide some advice for improving video MLLM training, offering valuable insights to guide future research and model development. Zijia Zhao, Haoyu Lu, Yuqi Huo, Yifan Du 0002, Tongtian Yue, Longteng Guo, Bingning Wang, Weipeng Chen, Jing Liu 0001 |
ICLR | 8 |
| 2025 | Maximizing Intermediate Checkpoint Value in LLM Pretraining with Bayesian OptimizationabstractThe rapid proliferation of large language models (LLMs), such as GPT-4 and Gemini, underscores the intense demand for resources during their training processes, posing significant challenges due to substantial computational and environmental costs. In this paper, we introduce a novel checkpoint merging strategy aimed at making efficient use of intermediate checkpoints during LLM pretraining. This method utilizes intermediate checkpoints with shared training trajectories, and is rooted in an extensive search space exploration for the best merging weight via Bayesian optimization. Through various experiments, we demonstrate that: (1) Our proposed methodology exhibits the capacity to augment pretraining, presenting an opportunity akin to obtaining substantial benefits at minimal cost; (2) Our proposed methodology, despite requiring a given held-out dataset, still demonstrates robust generalization capabilities across diverse domains, a pivotal aspect in pretraining. Deyuan Liu, Zecheng Wang, Bingning Wang, Weipeng Chen, Chunshan Li, Zhiying Tu, Dianbo Sui |
ICML | 4 |
| 2025 | KV Shifting Attention Enhances Language ModelingabstractCurrent large language models (LLMs) predominantly rely on decode-only transformer architectures, which exhibit exceptional in-context learning (ICL) capabilities. It is widely acknowledged that the cornerstone of their ICL ability lies in the induction heads mechanism, which necessitates at least two layers of attention. To more effectively harness the model's induction capabilities, we revisit the induction heads mechanism and provide theoretical proof that KV shifting attention reduces the model's dependency on the depth and width of the induction heads mechanism. Our experimental results confirm that KV shifting attention enhances the learning of induction heads and improves language modeling performance. This leads to superior performance or accelerated convergence, spanning from toy models to pre-trained models with over 10 billion parameters. Bingning Wang, Weipeng Chen |
ICML | 3 |
| 2025 | ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningabstractLarge Language Models (LLMs) have shown remarkable capabilities in reasoning, exemplified by the success of OpenAI-o1 and DeepSeek-R1. However, integrating reasoning with external search processes remains challenging, especially for complex multi-hop questions requiring multiple retrieval steps. We propose ReSearch, a novel framework that trains LLMs to Reason with Search via reinforcement learning without using any supervised data on reasoning steps. Our approach treats search operations as integral components of the reasoning chain, where when and how to perform searches is guided by text-based thinking, and search results subsequently influence further reasoning. We train ReSearch on Qwen2.5-7B(-Instruct) and Qwen2.5-32B(-Instruct) models and conduct extensive experiments. Despite being trained on only one dataset, our models demonstrate strong generalizability across various benchmarks. Analysis reveals that ReSearch naturally elicits advanced reasoning capabilities such as reflection and self-correction during the reinforcement learning process. Mingyang Chen 0002, Linzhuang Sun, Tianpeng Li, Haoze Sun, Chenzheng Zhu, Haofen Wang, Jeff Z. Pan, Wen Zhang 0015, Huajun Chen, Fan Yang 0132, Zenan Zhou, Weipeng Chen |
NeurIPS | 13 |
| 2025 | HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG SystemsabstractRetrieval-Augmented Generation (RAG) has been shown to improve knowledge capabilities and alleviate the hallucination problem of LLMs. The Web is a major source of external knowledge used in RAG systems, and many commercial RAG systems have used Web search engines as their major retrieval systems. Typically, such RAG systems retrieve search results, download HTML sources of the results, and then extract plain texts from the HTML sources. Plain text documents or chunks are fed into the LLMs to augment the generation. However, much of the structural and semantic information inherent in HTML, such as headings and table structures, is lost during this plain-text-based RAG process. To alleviate this problem, we propose HtmlRAG, which uses HTML instead of plain text as the format of retrieved knowledge in RAG. We believe HTML is better than plain text in modeling knowledge in external documents, and most LLMs possess robust capacities to understand HTML. However, utilizing HTML presents new challenges. HTML contains additional content such as tags, JavaScript, and CSS specifications, which bring extra input tokens and noise to the RAG system. To address this issue, we propose HTML cleaning, compression, and a two-step block-tree-based pruning strategy, to shorten the HTML while minimizing the loss of information. Experiments on six QA datasets confirm the superiority of using HTML in RAG systems. Our code and datasets are available at https://github.com/plageon/HtmlRAG. Jiejun Tan, Zhicheng Dou, Wen Wang 0016, Weipeng Chen, Ji-Rong Wen |
WWW | 5 |
| 2025 | PQCache: Product Quantization-based KVCache for Long Context LLM InferenceabstractAs the field of Large Language Models (LLMs) continues to evolve, the context length in inference is steadily growing. Key-Value Cache (KVCache), the intermediate representations of tokens within LLM inference, has now become the primary memory bottleneck due to limited GPU memory. Current methods selectively determine suitable keys and values for self-attention computation in LLMs to address the issue. However, they either fall short in maintaining model quality or result in high serving latency. Drawing inspiration from advanced embedding retrieval techniques prevalent in the data management community, we consider the storage and retrieval of KVCache as a typical embedding retrieval problem. We propose PQCache , which employs Product Quantization (PQ) to manage KVCache, maintaining model quality while ensuring low serving latency. During the prefilling phase, we apply PQ to tokens' keys for each LLM layer and head. During the autoregressive decoding phase, we use PQ codes and centroids to approximately identify important preceding tokens, then fetch the corresponding key-value pairs for self-attention computation. Through meticulous design of overlapping and caching, we minimize any additional computation and communication overhead during both phases. Extensive experiments demonstrate that PQCache achieves both effectiveness and efficiency, with 4.60% score improvement over existing methods on InfiniteBench and low system latency in both prefilling and decoding. Hailin Zhang 0004, Fangcheng Fu, Xupeng Miao, Xiaonan Nie, Weipeng Chen, Bin Cui 0001 |
Proc. ACM Manag. Data | 7 |
| 2024 | MetaGPT: Merging Large Language Models Using Model Exclusive Task ArithmeticabstractThe advent of large language models (LLMs) like GPT-4 has catalyzed the exploration of multi-task learning (MTL), in which a single model demonstrates proficiency across diverse tasks.Task arithmetic has emerged as a costeffective approach for MTL.It enables performance enhancement across multiple tasks by adding their corresponding task vectors to a pre-trained model.However, the current lack of a method that can achieve optimal performance with low computational cost and protecting the data privacy, which limits their application to LLMs.In this paper, we propose Model Exclusive Task Arithmetic for merging GPT-scale models (MetaGPT), which formalizes the objective of model merging into a multi-task learning framework, aiming to minimize the average loss difference between the merged model and each individual task model.Since data privacy limits the use of multi-task training data, we leverage LLMs' local linearity and task vectors' orthogonality to separate the data term and scaling coefficients term and derive a model-exclusive task arithmetic method.Our proposed MetaGPT is dataagnostic and bypasses the heavy search process, making it cost-effective and easy to implement for LLMs.Extensive experiments demonstrate that MetaGPT leads to improvements in task arithmetic and achieves state-of-the-art performance on multiple tasks. Yuyan Zhou, Bingning Wang, Weipeng Chen |
EMNLP | 4 |
| 2024 | HGSVerb: Improving Zero-shot Text Classification via Hierarchical Generative Semantic-Aware VerbalizerabstractPrompt-based methods with Pre-trained Language Models (PLMs) have demonstrated remarkable zero-shot performance in text classification tasks. Among these methods, verbalizers are employed to convert model-predicted vocabulary logits into task-specific labels. However, most existing zero-shot approaches face several challenges: (1) a reliance on unlabeled data or additional knowledge bases, which limits their effectiveness in varied data scenarios; (2) a disregard for semantic ambiguity issues in verbalizer construction. This paper presents a novel fully zero-shot approach named Hierarchical Generative Semantic-Aware Generative Verbalizer (HGSVerb). (1) We propose a prompt-based module, Hierarchical Semantic Verbalizer Generation, that explores the full utilization of PLMs and the semantics of label names. As a result, without requiring any extra data or knowledge, our model iteratively generates a verbalizer with hierarchical semantics and structure. (2) We propose a Semantic-Aware Refinement method based on semantic distance between label words and categories to reduce the semantic ambiguity issue. (3) Additionally, HGSVerb introduces a Hierarchical Weight Aggregation to optimize the utilization of our verbalizer by using a decay coefficient to distinguish the importance of label words generated by different iterations. We evaluate HGSVerb on four topic and sentiment classification datasets, and our average accuracy even outperforms those methods utilizing extra resources. We also examine its transferability on datasets with diverse numbers of classes and topics. HGSVerb achieves the best results compared with existing fully zero-shot methods. Zifeng Liu, Weipeng Chen |
IJCNN | 3 |
| 2024 | Exploring Context Window of Large Language Models via Decomposed Positional VectorsabstractTransformer-based large language models (LLMs) typically have a limited context window, resulting in significant performance degradation when processing text beyond the length of the context window. Extensive studies have been proposed to extend the context window and achieve length extrapolation of LLMs, but there is still a lack of in-depth interpretation of these approaches. In this study, we explore the positional information within and beyond the context window for deciphering the underlying mechanism of LLMs. By using a mean-based decomposition method, we disentangle positional vectors from hidden states of LLMs and analyze their formation and effect on attention. Furthermore, when texts exceed the context window, we analyze the change of positional vectors in two settings, i.e., direct extrapolation and context window extension. Based on our findings, we design two training-free context window extension methods, positional vector replacement and attention window extension. Experimental results show that our methods can effectively extend the context window length. Zican Dong, Junyi Li 0001, Xin Men, Wayne Xin Zhao, Bingning Wang, Zhen Tian 0001, Weipeng Chen, Ji-Rong Wen |
NeurIPS | 7 |
| 2024 | Base of RoPE Bounds Context LengthabstractPosition embedding is a core component of current Large Language Models (LLMs). Rotary position embedding (RoPE), a technique that encodes the position information with a rotation matrix, has been the de facto choice for position embedding in many LLMs, such as the Llama series. RoPE has been further utilized to extend long context capability, which is roughly based on adjusting the \textit{base} parameter of RoPE to mitigate out-of-distribution (OOD) problems in position embedding. However, in this paper, we find that LLMs may obtain a superficial long-context ability based on the OOD theory. We revisit the role of RoPE in LLMs and propose a novel property of long-term decay, we derive that the \textit{base of RoPE bounds context length}: there is an absolute lower bound for the base value to obtain certain context length capability. Our work reveals the relationship between context length and RoPE base both theoretically and empirically, which may shed light on future long context training. Xin Men, Bingning Wang, Xianpei Han, Weipeng Chen |
NeurIPS | 7 |
| 2022 | VMEKNet: Visual Memory and External Knowledge Based Network for Medical Report Generation
Weipeng Chen, Haiwei Pan, Kejia Zhang 0001, Qianna Cui |
PRICAI (1) | 1 |
| 2021 | ComQA: Compositional Question Answering via Hierarchical Graph Neural NetworksabstractWith the development of deep learning techniques and large scale datasets, the question answering (QA) systems have been quickly improved, providing more accurate and satisfying answers. However, current QA systems either focus on the sentence-level answer, i.e., answer selection, or phrase-level answer, i.e., machine reading comprehension. How to produce compositional answers has not been throughout investigated. In compositional question answering, the systems should assemble several supporting evidence from the document to generate the final answer, which is more difficult than sentence-level or phrase-level QA. In this paper, we present a large-scale compositional question answering dataset containing more than 120k human-labeled questions. The answer in this dataset is composed of discontiguous sentences in the corresponding document. To tackle the ComQA problem, we proposed a hierarchical graph neural networks, which represent the document from the low-level word to the high-level sentence. We also devise a question selection and node selection task for pre-training. Our proposed model achieves a significant improvement over previous machine reading comprehension methods and pre-training methods. Codes, dataset can be found at https://github.com/benywon/ComQA. Bingning Wang, Weipeng Chen, Jingfang Xu |
WWW | 3 |
| 2020 | Incorporating Knowledge and Content Information to Boost News Recommendation
Zhen Wang 0040, Weizhi Ma, Min Zhang 0006, Weipeng Chen, Jingfang Xu, Yiqun Liu 0001, Shaoping Ma |
NLPCC (1) | 4 |
| 2018 | Utilizing soft constraints to enhance medical relation extraction from the history of present illness in electronic medical records
Yuanju Li, Weipeng Chen, Xinglong Liu, Zhonghua Yu |
J. Biomed. Informatics | 3 |
| 2017 | Automatic Generation and Recommendation for API MashupsabstractUntil today, finding the most suitable APIs to use in an application was burdensome, requiring manual and time-consuming searches across a diverse set of websites, in particular regarding how multiple APIs could be combined and worked together (i.e. API mashups). In this paper, we propose a new method to automatically generate API mashups through real-world data collection, text mining and natural language processing (NLP ) techniques. The generated API mashups are further ranked and recommended to developers based on a quantitative indicator of whether the given API mashup is plausible. To evaluate the overall accuracy of the proposed method, we use the generated API mashups to train several machine learning and deep learning models, and then use an independent mashup dataset collected from Github projects for testing. The experimental results show that our proposed method is feasible and accurate for automatic API mashup generation and recommendation. Qinghan Xue, Weipeng Chen, Mooi Choo Chuah |
ICMLA | 3 |