Monica Xiao Cheng

dblp:352/4083 · also Monica Cheng · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0002-1140-687XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 10 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021
YearPublicationVenuePosition
2026 Prune as You Generate: Online Rollout Pruning for Faster and Better RLVR
abstract
Haobo Xu, Sirui Chen, Ruizhong Qiu, Yuchen Yan, Chen Luo, Monica Xiao Cheng, Jingrui He, Hanghang Tong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ruizhong Qiu, Chen Luo 0003, Monica Xiao Cheng, Jingrui He, Hanghang Tong
ACL (1)6
2026 Attn-GS: Attention-Guided Context Compression for Efficient Personalized LLMs
abstract
Shenglai Zeng, Tianqi Zheng, Chuan Tian, Dante Everaert, Yau-Shian Wang, Yupin Huang, Michael J. Morais, Rohit Patki, Jinjin Tian, Xinnan Dai, Kai Guo, Monica Xiao Cheng, Hui Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shenglai Zeng, Chuan Tian, Dante Everaert, Yau-Shian Wang, Yupin Huang, Michael J. Morais, Rohit Patki, Jinjin Tian, Xinnan Dai, Kai Guo 0003, Monica Xiao Cheng, Hui Liu 0031
ACL (1)12
2026 Building a Production Shopping Agent at Scale
abstract
Deploying a conversational shopping agent at production scale remains challenging despite rapid advances in large language models. Unlike research prototypes, production systems must satisfy strict requirements on accuracy, latency, reliability, and cost while serving millions of customers. We share our year-long journey of building a production shopping agent at scale and show that combining agentic reasoning with production-aware retrieval, tool orchestration, and system optimization enables a single shopping agent to support product discovery, shopping question answering, and agentic actions under real-world traffic. Our experience shows that LLM-based agents can be reliably operated at Amazon scale and provides practical design principles for industrial search and recommendation systems.
Chen Luo 0003, Jason Choi, Ziwei Dong, Rahul Dua, Xuejing Lei, Xin Zhang 0163, Josef Valvoda, Gaurang Sinkar, Binit Jha, Yi Liu 0033, Monica Xiao Cheng
SIGIR13
2025 Unlocking Efficient, Scalable, and Continual Knowledge Editing with Basis-Level Representation Fine-Tuning
abstract
Large language models (LLMs) have achieved remarkable performance on vari- ous natural language tasks. However, they are trained on static corpora and their knowledge can become outdated quickly in the fast-changing world. This moti- vates the development of knowledge editing methods designed to update certain knowledge in LLMs without changing unrelated others. To make selective edits, previous efforts often sought to update a small amount of parameters in some spe- cific layer(s) of a LLM. Nonetheless, in challenging scenarios, they still fall short in making successful edits while preserving knowledge irrelevant to the updates simultaneously, resulting in a notable editing-locality trade-off. In this work, we question if the trade-offs are caused by the fact that parameter-based updates have a global effect, i.e., edited parameters affect all inputs indiscriminately. In light of this, we explore the feasibility of representation fine-tuning, which applied some linear update to a few representations in a learned subspace, for knowledge edit- ing. While being effective to enhance an LLM’s general ability as demonstrated in the previous work, we theoretically show that this linear update imposes a tension in editing-locality trade-off. Subsequently, BaFT is proposed to break the linear- ity. BaFT computes a weight for each basis that spans a dimension of the subspace based on the input representation. This input-dependent weighting mechanism al- lows BaFT to manage different types of knowledge in an adaptive way, thereby achieving a better editing-locality trade-off. Experiments on three LLMs with five editing benchmarks in diverse scenarios show the superiority of our method.
Tianci Liu 0003, Ruirui Li 0002, Yunzhe Qi, Hui Liu 0033, Xianfeng Tang, Qingyu Yin, Monica Xiao Cheng, Jun Huan, Haoyu Wang 0004, Jing Gao 0004
ICLR8
2025 Towards Knowledge Checking in Retrieval-augmented Generation: A Representation Perspective
abstract
Shenglai Zeng, Jiankun Zhang, Bingheng Li, Yuping Lin, Tianqi Zheng, Dante Everaert, Hanqing Lu, Hui Liu, Hui Liu, Yue Xing, Monica Xiao Cheng, Jiliang Tang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Shenglai Zeng, Jiankun Zhang 0001, Bingheng Li, Yuping Lin, Dante Everaert, Hanqing Lu, Hui Liu 0033, Hui Liu 0031, Yue Xing 0002, Monica Xiao Cheng, Jiliang Tang
NAACL (Long Papers)11
2024 RoseLoRA: Row and Column-wise Sparse Low-rank Adaptation of Pre-trained Language Model for Knowledge Editing and Fine-tuning
abstract
Pre-trained language models, trained on largescale corpora, demonstrate strong generalizability across various NLP tasks.Finetuning these models for specific tasks typically involves updating all parameters, which is resource-intensive.Parameter-efficient finetuning (PEFT) methods, such as the popular LoRA family, introduce low-rank matrices to learn only a few parameters efficiently.However, during inference, the product of these matrices updates all pre-trained parameters, complicating tasks like knowledge editing that require selective updates.We propose a novel PEFT method, which conducts row and column-wise sparse low-rank adaptation (RoseLoRA), to address this challenge.RoseLoRA identifies and updates only the most important parameters for a specific task, maintaining efficiency while preserving other model knowledge.By adding a sparsity constraint on the product of low-rank matrices and converting it to row and column-wise sparsity, we ensure efficient and precise model updates.Our theoretical analysis guarantees the lower bound of the sparsity with respective to the matrix product.Extensive experiments on five benchmarks across twenty datasets demonstrate that RoseLoRA outperforms baselines in both general fine-tuning and knowledge editing tasks.
Haoyu Wang 0004, Tianci Liu 0003, Ruirui Li 0002, Monica Xiao Cheng, Tuo Zhao, Jing Gao 0004
EMNLP4
2024 BlendFilter: Advancing Retrieval-Augmented Large Language Models via Query Generation Blending and Knowledge Filtering
abstract
Haoyu Wang, Ruirui Li, Haoming Jiang, Jinjin Tian, Zhengyang Wang, Chen Luo, Xianfeng Tang, Monica Xiao Cheng, Tuo Zhao, Jing Gao. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Haoyu Wang 0004, Ruirui Li 0002, Haoming Jiang, Jinjin Tian, Chen Luo 0003, Xianfeng Tang, Monica Xiao Cheng, Tuo Zhao, Jing Gao 0004
EMNLP8
2024 LightLT: A Lightweight Representation Quantization Framework for Long-Tail Data
abstract
Search tasks require finding items similar to a given query, making it a crucial aspect of various applications. However, storing and computing similarity for millions or billions of item representations can be computationally expensive. To address this, quantization-based hash methods present memory and inference-efficient solutions by converting continuous representations into non-negative integer codes. Despite their advantages, these methods often encounter difficulties in handling long-tail datasets due to imbalanced class distributions. To address this, we propose LightLT, a lightweight representation quantization framework tailored for long-tail datasets. LightLT produces compact codebooks and discrete IDs, enabling efficient inference by computing distances between query and codewords. Our framework includes innovative designs: 1) Quantization Step: We select the most similar codeword for continuous inputs using the differentiable argmax operation. 2) Double Skip Quantization Connection Module: This module promotes codebook diversity and stability during training. 3) Training Loss: Our comprehensive loss includes class-weighted cross-entropy, center loss, and ranking loss. 4) Model Ensemble: We incorporate a model ensemble step to improve generalization. Theoretical analysis confirms LightLT's low space and inference complexity. Experimental results demonstrate superior performance compared to state-of-the-art baselines in terms of search accuracy, efficiency, and memory usage.
Haoyu Wang 0004, Ruirui Li 0002, Xianfeng Tang, Danqing Zhang, Monica Xiao Cheng, Jasha Droppo, Suhang Wang, Jing Gao 0004
ICDE6
2023 Exploiting Intent Evolution in E-commercial Query Recommendation
abstract
Aiming at a better understanding of the search goals in the user search sessions, recent query recommender systems explicitly model the reformulations of queries, which hopes to estimate the intents behind these reformulations and thus benefit the next-query recommendation. However, in real-world e-commercial search scenarios, user intents are much more complicated and may evolve dynamically. Existing methods merely consider trivial reformulation intents from semantic aspects and fail to model dynamic reformulation intent flows in search sessions, leading to sub-optimal capacities to recommend desired queries. To deal with these limitations, we first explicitly define six types of query reformulation intents according to the desired products of two consecutive queries. We then apply two self-attentive encoders on top of two pre-trained large language models to learn the transition dynamics from semantic query and intent reformulation sequences, respectively. We develop an intent-aware query decoder to utilize the predicted intents for suggesting the next queries. We instantiate such a framework with an Intent-aware Variational AutoEncoder (IVAE) under deployment at Amazon. We conduct comprehensive experiments on two real-world e-commercial datasets from Amazon and one public dataset from BestBuy. Specifically, IVAE improves the Recall@15 by 25.44% and 60.47% on two Amazon datasets and 13.91% on BestBuy, respectively.
Yu Wang 0158, Qingyu Yin, Xianfeng Tang, Yinghan Wang, Danqing Zhang, Limeng Cui, Monica Xiao Cheng, Suhang Wang, Philip S. Yu
KDD9
2023 A Unified Framework of Graph Information Bottleneck for Robustness and Membership Privacy
abstract
Graph Neural Networks (GNNs) have achieved great success in modeling graph-structured data. However, recent works show that GNNs are vulnerable to adversarial attacks which can fool the GNN model to make desired predictions of the attacker. In addition, training data of GNNs can be leaked under membership inference attacks. This largely hinders the adoption of GNNs in high-stake domains such as e-commerce, finance and bioinformatics. Though investigations have been made in conducting robust predictions and protecting membership privacy, they generally fail to simultaneously consider the robustness and membership privacy. Therefore, in this work, we study a novel problem of developing robust and membership privacy-preserving GNNs. Our analysis shows that Information Bottleneck (IB) can help filter out noisy information and regularize the predictions on labeled samples, which can benefit robustness and membership privacy. However, structural noises and lack of labels in node classification challenge the deployment of IB on graph-structured data. To mitigate these issues, we propose a novel graph information bottleneck framework that can alleviate structural noises with neighbor bottleneck. Pseudo labels are also incorporated in the optimization to minimize the gap between the predictions on the labeled set and unlabeled set for membership privacy. Extensive experiments on real-world datasets demonstrate that our method can give robust predictions and simultaneously preserve membership privacy.
Enyan Dai, Limeng Cui, Xianfeng Tang, Yinghan Wang, Monica Xiao Cheng, Suhang Wang
KDD6
2023 LightToken: A Task and Model-agnostic Lightweight Token Embedding Framework for Pre-trained Language Models
Haoyu Wang 0004, Ruirui Li 0002, Haoming Jiang, Xianfeng Tang, Bin Bi, Monica Xiao Cheng, Yaqing Wang 0001, Tuo Zhao, Jing Gao 0004
KDD7
2023 Amazon-M2: A Multilingual Multi-locale Shopping Session Dataset for Recommendation and Text Generation
abstract
Modeling customer shopping intentions is a crucial task for e-commerce, as it directly impacts user experience and engagement. Thus, accurately understanding customer preferences is essential for providing personalized recommendations. Session-based recommendation, which utilizes customer session data to predict their next interaction, has become increasingly popular. However, existing session datasets have limitations in terms of item attributes, user diversity, and dataset scale. As a result, they cannot comprehensively capture the spectrum of user behaviors and preferences.To bridge this gap, we present the Amazon Multilingual Multi-locale Shopping Session Dataset, namely Amazon-M2. It is the first multilingual dataset consisting of millions of user sessions from six different locales, where the major languages of products are English, German, Japanese, French, Italian, and Spanish.Remarkably, the dataset can help us enhance personalization and understanding of user preferences, which can benefit various existing tasks as well as enable new tasks. To test the potential of the dataset, we introduce three tasks in this work:(1) next-product recommendation, (2) next-product recommendation with domain shifts, and (3) next-product title generation.With the above tasks, we benchmark a range of algorithms on our proposed dataset, drawing new insights for further research and practice. In addition, based on the proposed dataset and tasks, we hosted a competition in the KDD CUP 2023 https://www.aicrowd.com/challenges/amazon-kdd-cup-23-multilingual-recommendation-challenge and have attracted thousands of users and submissions. The winning solutions and the associated workshop can be accessed at our website~https://kddcup23.github.io/.
Wei Jin 0009, Haitao Mao, Zheng Li 0018, Haoming Jiang, Chen Luo 0003, Hongzhi Wen, Haoyu Han 0001, Hanqing Lu, Ruirui Li 0002, Monica Xiao Cheng, Rahul Goutam, Karthik Subbian, Suhang Wang, Yizhou Sun, Jiliang Tang, Xianfeng Tang
NeurIPS12