EDBT 2026 Demo / reviewers in the wild / expert
Qian Liu 0033
dblp:33/85-33
· DBLP profile ↗
54ranked-venue papers
8as first author
44since 2021 · last 2026
0000-0003-0511-0658ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 8 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How Can Synthetic Data Improve Multilingual Language Model Pretraining? A Data Quality PerspectiveabstractTongyao Zhu, Qian Liu, Chang Ma, Jinghan Zhang, Longxu Dou, Junxian He, Shiqi Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tongyao Zhu, Qian Liu 0033, Jinghan Zhang 0006, Longxu Dou, Junxian He, Shiqi Chen 0002 |
ACL (1) | 2 |
| 2026 | History Doesn't Repeat Itself but Rollouts Rhyme: Accelerating Reinforcement Learning with RhymeRLabstractWith the rapid advancement of large language models (LLMs), reinforcement learning (RL) has emerged as a pivotal methodology for enhancing the reasoning capabilities of LLMs. Unlike traditional pre-training approaches, RL encompasses multiple stages: rollout, reward, and training, which necessitates collaboration among various worker types. However, current RL systems continue to grapple with substantial GPU underutilization, due to two primary factors: (1) The rollout stage dominates the overall RL process due to test-time scaling; (2) Imbalances in rollout lengths (within the same batch) result in GPU bubbles. While prior solutions like asynchronous execution and truncation offer partial relief, they may compromise training accuracy for efficiency. Our key insight stems from a previously overlooked observation: rollout responses exhibit remarkable similarity across adjacent training epochs. Based on the insight, we introduce RhymeRL, an LLM RL system designed to accelerate RL training with two key innovations. First, to enhance rollout generation, we present HistoSpec, a speculative decoding inference engine that utilizes the similarity of historical rollout token sequences to obtain accurate drafts. Second, to tackle rollout bubbles, we introduce HistoPipe, a two-tier scheduling strategy that leverages the similarity of historical rollout distributions to balance workload among rollout workers. Experimental results demonstrate that RhymeRL achieves up to a 2.6x performance improvement over existing methods, without compromising accuracy or modifying the RL paradigm. Jingkai He, Tianjian Li, Erhu Feng, Dong Du 0003, Qian Liu 0033, Yubin Xia, Haibo Chen 0001 |
ASPLOS (2) | 5 |
| 2025 | OpenCoder: The Open Cookbook for Top-Tier Code Large Language ModelsabstractSiming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Yang Xu, Jian Yang, Jiaheng Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, Qiufeng Wang, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang, Yang Gao, Jie Fu, Qian Liu, Houyi Li, Ge Zhang, Yuan Qi, Xu Yinghui, Wei Chu, Zili Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Siming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Jian Yang 0030, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang 0001, Yang Gao 0021, Jie Fu 0001, Qian Liu 0033, Houyi Li, Ge Zhang 0009, Yuan Qi 0001 |
ACL (1) | 20 |
| 2025 | Bootstrapping Language Models with DPO Implicit RewardsabstractHuman alignment in large language models (LLMs) is an active area of research. A recent groundbreaking work, direct preference optimization (DPO), has greatly simplified the process from past work in reinforcement learning from human feedback (RLHF) by bypassing the reward learning stage in RLHF. DPO, after training, provides an implicit reward model. In this work, we make a novel observation that this implicit reward model can by itself be used in a bootstrapping fashion to further align the LLM. Our approach is to use the rewards from a current LLM to construct a preference dataset, which is then used in subsequent DPO rounds. We incorporate two refinements to further improve our approach: 1) length-regularized reward shaping to make the preference dataset length-unbiased; 2) experience replay to enhance the quality of the preference dataset. Our approach, named self-alignment with DPO ImpliCit rEwards (DICE), shows great improvements in alignment. It achieves an increase of more than 8$\\%$ in lengthcontrolled win rate on AlpacaEval 2 for all the different base models that we tried, without relying on external feedback. Our code is available at https://github.com/sail-sg/dice. Changyu Chen, Tianyu Pang, Qian Liu 0033, Arunesh Sinha, Pradeep Varakantham |
ICLR | 5 |
| 2025 | When Attention Sink Emerges in Language Models: An Empirical ViewabstractAuto-regressive language Models (LMs) assign significant attention to the first token, even if it is not semantically important, which is known as **attention sink**. This phenomenon has been widely adopted in applications such as streaming/long context generation, KV cache optimization, inference acceleration, model quantization, and others. Despite its widespread use, a deep understanding of attention sink in LMs is still lacking. In this work, we first demonstrate that attention sinks exist universally in auto-regressive LMs with various inputs, even in small models. Furthermore, attention sink is observed to emerge during the LM pre-training, motivating us to investigate how *optimization*, *data distribution*, *loss function*, and *model architecture* in LM pre-training influence its emergence. We highlight that attention sink emerges after effective optimization on sufficient training data. The sink position is highly correlated with the loss function and data distribution. Most importantly, we find that attention sink acts more like key biases, *storing extra attention scores*, which could be non-informative and not contribute to the value computation. We also observe that this phenomenon (at least partially) stems from tokens' inner dependence on attention scores as a result of softmax normalization. After relaxing such dependence by replacing softmax attention with other attention operations, such as sigmoid attention without normalization, attention sinks do not emerge in LMs up to 1B parameters. The code is available at https://github.com/sail-sg/Attention-Sink. Xiangming Gu, Tianyu Pang, Qian Liu 0033, Fengzhuo Zhang, Cunxiao Du, Ye Wang 0007 |
ICLR | 4 |
| 2025 | Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL WorkflowsabstractReal-world enterprise text-to-SQL workflows often involve complex cloud or local data across various database systems, multiple SQL queries in various dialects, and diverse operations from data transformation to analytics.
We introduce Spider 2.0, an evaluation framework comprising $632$ real-world text-to-SQL workflow problems derived from enterprise-level database use cases.
The databases in Spider 2.0 are sourced from real data applications, often containing over 1,000 columns and stored in local or cloud database systems such as BigQuery and Snowflake.
We show that solving problems in Spider 2.0 frequently requires understanding and searching through database metadata, dialect documentation, and even project-level codebases.
This challenge calls for models to interact with complex SQL workflow environments, process extremely long contexts, perform intricate reasoning, and generate multiple SQL queries with diverse operations, often exceeding $100$ lines, which goes far beyond traditional text-to-SQL challenges.
Our evaluations indicate that based on o1-preview, our code agent framework successfully solves only 21.3\% of the tasks, compared with 91.2\% on Spider 1.0 and 73.0\% on BIRD.
Our results on Spider 2.0 show that while language models have demonstrated remarkable performance in code generation --- especially in prior text-to-SQL benchmarks --- they require significant improvement in order to achieve adequate performance for real-world enterprise usage.
Progress on Spider 2.0 represents crucial steps towards developing intelligent, autonomous, code agents for real-world enterprise settings.
Our code, baseline models, and data are available at [spider2-sql.github.io](spider2-sql.github.io) . Fangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao, Dongchan Shin, Hongjin Su, Zhaoqing Suo, Hongcheng Gao, Wenjing Hu, Victor Zhong, Caiming Xiong, Ruoxi Sun 0002, Qian Liu 0033, Sida I. Wang, Tao Yu 0009 |
ICLR | 14 |
| 2025 | RegMix: Data Mixture as Regression for Language Model Pre-trainingabstractThe data mixture for large language model pre-training significantly impacts performance, yet how to determine an effective mixture remains unclear. We propose RegMix to automatically identify a high-performing data mixture by formulating it as a regression task. RegMix trains many small models on diverse data mixtures, uses regression to predict performance of unseen mixtures, and applies the best predicted mixture to train a large-scale model with orders of magnitude more compute. To empirically validate RegMix, we train 512 models with 1M parameters for 1B tokens to fit the regression model and predict the best data mixture. Using this mixture we train a 1B parameter model for 25B tokens (i.e. 1000× larger and 25× longer) which we find performs best among 64 candidate 1B parameter models with other mixtures. Furthermore, RegMix consistently outperforms human selection in experiments involving models up to 7B models trained on 100B tokens, while matching or exceeding DoReMi using just 10% of the computational resources. Our experiments also show that (1) Data mixtures significantly impact performance; (2) Web corpora rather than data perceived as high-quality like Wikipedia have the strongest positive correlation with downstream performance; (3) Domains interact in complex ways often contradicting common sense, thus automatic approaches like RegMix are needed; (4) Data mixture effects transcend scaling laws. Our code is available at https://github.com/sail-sg/regmix. Qian Liu 0033, Xiaosen Zheng, Niklas Muennighoff, Guangtao Zeng, Longxu Dou, Tianyu Pang, Jing Jiang 0001 |
ICLR | 1 |
| 2025 | Scaling up Masked Diffusion Models on TextabstractMasked diffusion models (MDMs) have shown promise in language modeling, yet their scalability and effectiveness in core language tasks, such as text generation and language understanding, remain underexplored. This paper establishes the first scaling law for MDMs, demonstrating a scaling rate comparable to autoregressive models (ARMs) and a relatively small compute gap. Motivated by their scalability, we train a family of MDMs with up to 1.1 billion (B) parameters to systematically evaluate their performance against ARMs of comparable or larger sizes. Fully leveraging the probabilistic formulation of MDMs, we propose a simple yet effective *unsupervised classifier-free guidance* that effectively exploits large-scale unpaired data, boosting performance for conditional inference. In language understanding, the 1.1B MDM outperforms the 1.1B TinyLlama model trained on the same data across four of eight zero-shot benchmarks. Notably, it achieves competitive math reasoning ability with the 7B Llama-2 model on the GSM8K dataset.
In text generation, MDMs with 16 times more pre-training time offer a flexible trade-off against ARMs with the accelerated sampling technique KV-Cache: MDMs match ARMs in performance while being 1.4 times faster during sampling.
Moreover, MDMs address challenging tasks for ARMs by effectively handling bidirectional reasoning and adapting to temporal shifts in data. Notably, a 1.1B MDM breaks the *reverse curse* encountered by much larger ARMs with significantly more data and computation, such as 13B Llama-2 and 175B GPT-3. Our code is available at https://github.com/ML-GSAI/SMDM. Shen Nie, Fengqi Zhu, Tianyu Pang, Qian Liu 0033, Guangtao Zeng, Chongxuan Li |
ICLR | 5 |
| 2025 | Cheating Automatic LLM Benchmarks: Null Models Achieve High Win RatesabstractAutomatic LLM benchmarks, such as AlpacaEval 2.0, Arena-Hard-Auto, and MT-Bench, have become popular for evaluating language models due to their cost-effectiveness and scalability compared to human evaluation. Achieving high win rates on these benchmarks can significantly boost the promotional impact of newly released language models. This promotional benefit may motivate tricks, such as manipulating model output length or style to game win rates, even though several mechanisms have been developed to control length and disentangle style to reduce gameability. Nonetheless, we show that even a **"null model"** that always outputs a **constant** response (*irrelevant to input instructions*) can cheat automatic benchmarks and achieve top-ranked win rates: an $86.5\\%$ LC win rate on AlpacaEval 2.0; an $83.0$ score on Arena-Hard-Auto; and a $9.55$ score on MT-Bench. Moreover, the crafted cheating outputs are **transferable** because we assume that the instructions of these benchmarks (e.g., $805$ samples of AlpacaEval 2.0) are *private* and cannot be accessed. While our experiments are primarily proof-of-concept, an adversary could use LLMs to generate more imperceptible cheating responses, unethically benefiting from high win rates and promotional impact. Our findings call for the development of anti-cheating mechanisms for reliable automatic benchmarks. The code is available at https://github.com/sail-sg/Cheating-LLM-Benchmarks. Xiaosen Zheng, Tianyu Pang, Qian Liu 0033, Jing Jiang 0001 |
ICLR | 4 |
| 2025 | Unnatural Languages Are Not Bugs but Features for LLMsabstractLarge Language Models (LLMs) have been observed to process non-human-readable text sequences, such as jailbreak prompts, often viewed as a bug for aligned LLMs. In this work, we present a systematic investigation challenging this perception, demonstrating that unnatural languages - strings that appear incomprehensible to humans but maintain semantic meanings for LLMs - contain latent features usable by models. Notably, unnatural languages possess latent features that can be generalized across different models and tasks during inference. Furthermore, models fine-tuned on unnatural versions of instruction datasets perform on-par with those trained on natural language, achieving $49.71$ win rates in Length-controlled AlpacaEval 2.0 in average across various base models. In addition, through comprehensive analysis, we demonstrate that LLMs process unnatural languages by filtering noise and inferring contextual meaning from filtered words. Our code is publicly available at https://github.com/John-AI-Lab/Unnatural_Language. Keyu Duan, Yiran Zhao 0006, Zhili Feng, Jinjie Ni, Tianyu Pang, Qian Liu 0033, Tianle Cai, Longxu Dou, Kenji Kawaguchi, Anirudh Goyal, J. Zico Kolter, Michael Shieh |
ICML | 6 |
| 2025 | Improving Your Model Ranking on Chatbot Arena by Vote RiggingabstractChatbot Arena is an open platform for evaluating LLMs by pairwise battles, in which users vote for their preferred response from two randomly sampled anonymous models. While Chatbot Arena is widely regarded as a reliable LLM ranking leaderboard, we show that crowdsourced voting can be rigged to improve (or decrease) the ranking of a target model $m_{t}$. We first introduce a straightforward target-only rigging strategy that focuses on new battles involving $m_{t}$, identifying it via watermarking or a binary classifier, and exclusively voting for $m_{t}$ wins. However, this strategy is practically inefficient because there are over $190$ models on Chatbot Arena and on average only about 1% of new battles will involve $m_{t}$. To overcome this, we propose an omnipresent rigging strategy, exploiting the Elo rating mechanism of Chatbot Arena that any new vote on a battle can influence the ranking of the target model $m_{t}$, even if $m_{t}$ is not directly involved in the battle. We conduct experiments on around 1.7 million historical votes from the Chatbot Arena Notebook, showing that omnipresent rigging strategy can improve model rankings by rigging only hundreds of new votes. While we have evaluated several defense mechanisms, our findings highlight the importance of continued efforts to prevent vote rigging. Code is publicly available to reproduce all experiments. Tianyu Pang, Qian Liu 0033, Minhao Cheng |
ICML | 4 |
| 2025 | Predictive Data Selection: The Data That Predicts Is the Data That TeachesabstractLanguage model pretraining involves training on extensive corpora, where data quality plays a pivotal role. In this work, we aim to directly estimate the contribution of data during pretraining and select pretraining data in an efficient manner. Specifically, we draw inspiration from recent findings showing that compression efficiency (i.e., normalized loss) of diverse models on certain text correlates strongly with their downstream performance, when the text domain aligns with the downstream benchmarks (Huang et al., 2024). Building on this observation, we hypothesize that data on which model losses are predictive of downstream abilities also contribute effectively to learning, which shares similar intuition with Thrush et al. (2024). To leverage this insight, we introduce predictive data selection (PreSelect), a lightweight and efficient data selection method that requires training and deploying only a fastText-based scorer. Through comprehensive experiments with 1B and 3B parameter models, we demonstrate that models trained on 30B tokens selected with PreSelect surpass the performance of the vanilla baseline trained on 300B tokens, achieving a 10x reduction in compute requirements. Furthermore, PreSelect significantly outperforms other competitive data selection baselines, such as DCLM and FineWeb-Edu on a scale of 3B models trained on 100B tokens. We open-source our trained data selection scorer along with the curated datasets at https://github.com/hkust-nlp/PreSelect. KaShun Shum, Yuzhen Huang 0002, Hongjian Zou, Yixuan Liao, Xiaoxin Chen 0001, Qian Liu 0033, Junxian He |
ICML | 7 |
| 2025 | Programming Every Example: Lifting Pre-training Data Quality Like Experts at ScaleabstractLarge language model pre-training has traditionally relied on human experts to craft heuristics for improving the corpora quality, resulting in numerous rules developed to date. However, these fixed rules lack the flexibility to address the unique characteristics of individual examples, yet crafting sample-wise rules is impractical for human experts. In this paper, we show that even small language models, with only 0.3B parameters, can exhibit substantial data refining capabilities. We propose Programming Every Example (ProX), a novel framework that treats data refinement as a programming task, and enables the model to refine corpora by generating and executing fine-grained operations, such as string normalization, for each individual example at scale. Experiments show that models trained on ProX-refined data consistently outperform other baselines across 10 benchmarks, demonstrating effectiveness across model sizes (up to 1.7B) and pre-training corpora (C4, RedPajama-V2, FineWeb, FineWeb-Edu, and DCLM). ProX also shows great potential in continual pre-training: on math domain, ProX boosts 7B models by up to 20% within 10B tokens—results typically achieved with much larger scale training (e.g., 200B tokens). We believe ProX offers a way to curate high-quality pre-training data, and finally contributes to efficient LLM development. Zengzhi Wang, Qian Liu 0033, Pengfei Liu 0003 |
ICML | 3 |
| 2025 | ZeCO: Zero-Communication Overhead Sequence Parallelism for Linear AttentionabstractLinear attention mechanisms deliver significant advantages for Large Language Models (LLMs) by providing linear computational complexity, enabling efficient processing of ultra-long sequences (e.g., 1M context). However, existing Sequence Parallelism (SP) methods, essential for distributing these workloads across devices, become the primary performance bottleneck due to substantial communication overhead. In this paper, we introduce ZeCO (Zero Communication Overhead) sequence parallelism for linear attention models, a new SP method designed to overcome these limitations and achieve practically end-to-end near-linear scalability for long sequence training. For example, training a model with a 1M sequence length across 64 devices using ZeCO takes roughly the same time as training with an 16k sequence on a single device. At the heart of ZeCO lies All-Scan, a novel collective communication primitive. All-Scan provides each SP rank with precisely the initial operator state it requires while maintaining a minimal communication footprint, effectively eliminating communication overhead. Theoretically, we prove the optimaity of ZeCO, showing that it introduces only negligible time and space overhead. Empirically, we compare the communication costs of different sequence parallelism strategies and demonstrate that All-Scan achieves the fastest communication in SP scenarios. Specifically, on 256 GPUs with an 8M sequence length, ZeCO achieves a 60\% speedup compared to the current state-of-the-art (SOTA) SP method. We believe ZeCO establishes a clear path toward efficiently training next-generation LLMs on previously intractable sequence lengths. Yuhong Chou, Rui-Jie Zhu 0003, Tianjian Li, Congying Chu, Qian Liu 0033, Jibin Wu, Zejun Ma 0001 |
NeurIPS | 7 |
| 2025 | Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency OptimizationabstractLarge Language Models (LLMs) generate functionally correct solutions but often fall short in code efficiency, a critical bottleneck for real-world deployment. In this paper, we introduce a novel test-time iterative optimization framework to address this, employing a closed-loop system where LLMs iteratively refine code based on empirical performance feedback from an execution sandbox. We explore three training strategies: Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Group Relative Policy Optimization~(GRPO). Experiments on our Venus dataset and the APPS benchmark show that SFT and DPO rapidly saturate in efficiency gains. In contrast, GRPO, using reinforcement learning (RL) with execution feedback, continuously optimizes code performance, significantly boosting both pass@1 (from 47% to 62%) and the likelihood of outperforming human submissions in efficiency (from 31% to 45%). Our work demonstrates effective test-time code efficiency improvement and critically reveals the power of RL in teaching LLMs to truly self-improve code efficiency. We released our code and data at https://github.com/Elfsong/Afterburner. Mingzhe Du, Anh Tuan Luu, Yue Liu 0008, Yuhao Qing, Dong Huang 0005, Qian Liu 0033, Zejun Ma 0001, See-Kiong Ng |
NeurIPS | 7 |
| 2025 | General-Reasoner: Advancing LLM Reasoning Across All DomainsabstractReinforcement learning (RL) has recently demonstrated strong potential in enhancing the reasoning capabilities of large language models (LLMs).
Particularly, the "Zero" reinforcement learning introduced by Deepseek-R1-Zero, enables direct RL training of base LLMs without relying on an intermediate supervised fine-tuning stage.
Despite these advancements, current works for LLM reasoning mainly focus on mathematical and coding domains, largely due to data abundance and the ease of answer verification.
This limits the applicability and generalization of such models to broader domains, where questions often have diverse answer representations, and data is more scarce.
In this paper, we propose General-Reasoner, a novel training framework designed to enhance LLM reasoning capabilities across diverse domains. Our key contributions include: (1) constructing a large-scale, high-quality dataset of questions with verifiable answers curated by web crawling, covering a wide range of disciplines; and (2) developing a generative model-based answer verifier, which replaces traditional rule-based verification with the capability of chain-of-thought and context-awareness. We train a series of models and evaluate them on a wide range of datasets covering wide domains like physics, chemistry, finance, electronics etc. Our comprehensive evaluation across these 12 benchmarks (e.g. MMLU-Pro, GPQA, SuperGPQA, TheoremQA, BBEH and MATH AMC) demonstrates that General-Reasoner outperforms existing baseline methods, achieving robust and generalizable reasoning performance while maintaining superior effectiveness in mathematical reasoning tasks. Xueguang Ma, Qian Liu 0033, Dongfu Jiang, Ge Zhang 0009, Zejun Ma 0001, Wenhu Chen |
NeurIPS | 2 |
| 2025 | SkyLadder: Better and Faster Pretraining via Context Window SchedulingabstractRecent advancements in LLM pretraining have featured ever-expanding context windows to process longer sequences. However, our controlled study reveals that models pretrained with shorter context windows consistently outperform their long-context counterparts under a fixed token budget. This finding motivates us to explore an optimal context window scheduling strategy to better balance long-context capability with pretraining efficiency. To this end, we propose SkyLadder, a simple yet effective approach that implements a short-to-long context window transition. SkyLadder preserves strong standard benchmark performance, while matching or exceeding baseline results on long-context tasks. Through extensive experiments, we pretrain 1B-parameter models (up to 32K context) and 3B-parameter models (8K context) on 100B tokens, demonstrating that SkyLadder yields consistent gains of up to 3.7% on common benchmarks, while achieving up to 22% faster training speeds compared to baselines. Tongyao Zhu, Qian Liu 0033, Shiqi Chen 0002, Xiangming Gu, Tianyu Pang, Min-Yen Kan |
NeurIPS | 2 |
| 2025 | MPFToD: a modularized pre-training framework for consistency identification in task-oriented dialogue
Libo Qin 0001, Shijue Huang, Qiguang Chen, Qian Liu 0033, Wanxiang Che, Ruifeng Xu 0001 |
Frontiers Comput. Sci. | 4 |
| 2024 | Self-Distillation Bridges Distribution Gap in Language Model Fine-TuningabstractZhaorui Yang, Tianyu Pang, Haozhe Feng, Han Wang, Wei Chen, Minfeng Zhu, Qian Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhaorui Yang 0001, Tianyu Pang, Haozhe Feng, Wei Chen 0001, Minfeng Zhu 0001, Qian Liu 0033 |
ACL (1) | 7 |
| 2024 | Beyond Memorization: The Challenge of Random Memory Access in Language ModelsabstractRecent developments in Language Models (LMs) have shown their effectiveness in NLP tasks, particularly in knowledge-intensive tasks.However, the mechanisms underlying knowledge storage and memory access within their parameters remain elusive.In this paper, we investigate whether a generative LM (e.g., GPT-2) is able to access its memory sequentially or randomly.Through carefully-designed synthetic tasks, covering the scenarios of full recitation, selective recitation and grounded question answering, we reveal that LMs manage to sequentially access their memory while encountering challenges in randomly accessing memorized content.We find that techniques including recitation and permutation improve the random memory access capability of LMs.Furthermore, by applying this intervention to realistic scenarios of open-domain question answering, we validate that enhancing random access by recitation leads to notable improvements in question answering.The code to reproduce our experiments can be found at https://github. com/sail-sg/lm-random-memory-access. Tongyao Zhu, Qian Liu 0033, Liang Pang 0001, Zhengbao Jiang, Min-Yen Kan |
ACL (1) | 2 |
| 2024 | DataTales: A Benchmark for Real-World Intelligent Data NarrationabstractWe introduce DATATALES, a novel benchmark designed to assess the proficiency of language models in data narration, a task crucial for transforming complex tabular data into accessible narratives.Existing benchmarks often fall short in capturing the requisite analytical complexity for practical applications.DATATALES addresses this gap by offering 4.9k financial reports paired with corresponding market data, showcasing the demand for models to create clear narratives and analyze large datasets while understanding specialized terminology in the field.Our findings highlight the significant challenge that language models face in achieving the necessary precision and analytical depth for proficient data narration, suggesting promising avenues for future model development and evaluation methodologies.The data and code are available at https: //github.com/yajingyang/DataTales/. Yajing Yang, Qian Liu 0033, Min-Yen Kan |
EMNLP | 2 |
| 2024 | OctoPack: Instruction Tuning Code Large Language ModelsabstractFinetuning large language models (LLMs) on instructions leads to vast performance improvements on natural language tasks. We apply instruction tuning using code, leveraging the natural structure of Git commits, which pair code changes with human instructions. We compile CommitPack: 4 terabytes of Git commits across 350 programming languages. We benchmark CommitPack against other natural and synthetic code instructions (xP3x, Self-Instruct, OASST) on the 16B parameter StarCoder model, and achieve state-of-the-art performance among models not trained on OpenAI outputs, on the HumanEval Python benchmark (46.2% pass@1). We further introduce HumanEvalPack, expanding the HumanEval benchmark to a total of 3 coding tasks (Code Repair, Code Explanation, Code Synthesis) across 6 languages (Python, JavaScript, Java, Go, C++, Rust). Our models, OctoCoder and OctoGeeX, achieve the best performance across HumanEvalPack among all permissive models, demonstrating CommitPack's benefits in generalizing to a wider set of languages and natural coding tasks. Code, models and data are freely available at https://github.com/bigcode-project/octopack. Niklas Muennighoff, Qian Liu 0033, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro von Werra, Shayne Longpre |
ICLR | 2 |
| 2024 | Lemur: Harmonizing Natural Language and Code for Language AgentsabstractWe introduce Lemur and Lemur-Chat, openly accessible language models optimized
for both natural language and coding capabilities to serve as the backbone
of versatile language agents. The evolution from language chat models to
functional language agents demands that models not only master human interaction,
reasoning, and planning but also ensure grounding in the relevant environments.
This calls for a harmonious blend of language and coding capabilities
in the models. Lemur and Lemur-Chat are proposed to address this necessity,
demonstrating balanced proficiencies in both domains, unlike existing
open-source models that tend to specialize in either. Through meticulous pretraining
using a code-intensive corpus and instruction fine-tuning on text and code
data, our models achieve state-of-the-art averaged performance across diverse
text and coding benchmarks. Comprehensive experiments demonstrate Lemur’s
superiority over existing open-source models and its proficiency across various
agent tasks involving human communication, tool usage, and interaction under
fully- and partially- observable environments. The harmonization between natural
and programming languages enables Lemur-Chat to significantly narrow the
gap with proprietary models on agent abilities, providing key insights into developing
advanced open-source agents adept at reasoning, planning, and operating
seamlessly across environments. Our model and code have been open-sourced at
https://github.com/OpenLemur/Lemur. Yiheng Xu, Hongjin Su, Chen Xing, Boyu Mi, Qian Liu 0033, Binyuan Hui, Yitao Liu, Tianbao Xie, Zhoujun Cheng, Siheng Zhao, Lingpeng Kong, Bailin Wang, Caiming Xiong, Tao Yu 0009 |
ICLR | 5 |
| 2024 | Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially FastabstractA multimodal large language model (MLLM) agent can receive instructions, capture images, retrieve histories from memory, and decide which tools to use. Nonetheless, red-teaming efforts have revealed that adversarial images/prompts can jailbreak an MLLM and cause unaligned behaviors. In this work, we report an even more severe safety issue in multi-agent environments, referred to as infectious jailbreak. It entails the adversary simply jailbreaking a single agent, and without any further intervention from the adversary, (almost) all agents will become infected exponentially fast and exhibit harmful behaviors. To validate the feasibility of infectious jailbreak, we simulate multi-agent environments containing up to one million LLaVA-1.5 agents, and employ randomized pair-wise chat as a proof-of-concept instantiation for multi-agent interaction. Our results show that feeding an (infectious) adversarial image into the memory of any randomly chosen agent is sufficient to achieve infectious jailbreak. Finally, we derive a simple principle for determining whether a defense mechanism can provably restrain the spread of infectious jailbreak, but how to design a practical defense that meets this principle remains an open question to investigate. Xiangming Gu, Xiaosen Zheng, Tianyu Pang, Qian Liu 0033, Ye Wang 0007, Jing Jiang 0001 |
ICML | 5 |
| 2024 | S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language ModelabstractFangyu Lei, Qian Liu, Yiming Huang, Shizhu He, Jun Zhao, Kang Liu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Fangyu Lei, Qian Liu 0033, Shizhu He, Jun Zhao 0001, Kang Liu 0001 |
NAACL-HLT | 2 |
| 2024 | Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?abstractData science and engineering workflows often span multiple stages, from warehousing to orchestration, using tools like BigQuery, dbt, and Airbyte. As vision language models (VLMs) advance in multimodal understanding and code generation, VLM-based agents could potentially automate these workflows by generating SQL queries, Python code, and GUI operations. This automation can improve the productivity of experts while democratizing access to large-scale data analysis. In this paper, we introduce Spider2-V, the first multimodal agent benchmark focusing on professional data science and engineering workflows, featuring 494 real-world tasks in authentic computer environments and incorporating 20 enterprise-level professional applications. These tasks, derived from real-world use cases, evaluate the ability of a multimodal agent to perform data-related tasks by writing code and managing the GUI in enterprise data software systems. To balance realistic simulation with evaluation simplicity, we devote significant effort to developing automatic configurations for task setup and carefully crafting evaluation metrics for each task. Furthermore, we supplement multimodal agents with comprehensive documents of these enterprise data software systems. Our empirical evaluation reveals that existing state-of-the-art LLM/VLM-based agents do not reliably automate full data workflows (14.0% success). Even with step-by-step guidance, these agents still underperform in tasks that require fine-grained, knowledge-intensive GUI actions (16.2%) and involve remote cloud-hosted workspaces (10.6%). We hope that Spider2-V paves the way for autonomous multimodal agents to transform the automation of data science and engineering workflow. Our code and data are available at https://spider2-v.github.io. Ruisheng Cao, Fangyu Lei, Haoyuan Wu, Jixuan Chen, Yeqiao Fu, Hongcheng Gao, Xinzhuang Xiong, Hanchong Zhang, Wenjing Hu, Tianbao Xie, Hongshen Xu, Sida I. Wang, Ruoxi Sun 0002, Caiming Xiong, Ansong Ni, Qian Liu 0033, Victor Zhong, Lu Chen 0002, Kai Yu 0004, Tao Yu 0009 |
NeurIPS | 19 |
| 2024 | Mercury: A Code Efficiency Benchmark for Code Large Language ModelsabstractAmidst the recent strides in evaluating Large Language Models for Code (Code LLMs), existing benchmarks have mainly focused on the functional correctness of generated code, neglecting the importance of their computational efficiency. To fill the gap, we present Mercury, the first code efficiency benchmark for Code LLMs. It comprises 1,889 Python tasks, each accompanied by adequate solutions that serve as real-world efficiency baselines, enabling a comprehensive analysis of the runtime distribution. Based on the distribution, we introduce a new metric Beyond, which computes a runtime-percentile-weighted Pass score to reflect functional correctness and code efficiency simultaneously. On Mercury, leading Code LLMs can achieve 65% on Pass, while less than 50% on Beyond. Given that an ideal Beyond score would be aligned with the Pass score, it indicates that while Code LLMs exhibit impressive capabilities in generating functionally correct code, there remains a notable gap in their efficiency. Finally, our empirical experiments reveal that Direct Preference Optimization (DPO) serves as a robust baseline for enhancing code efficiency compared with Supervised Fine Tuning (SFT), which paves a promising avenue for future exploration of efficient code generation. Our code and data are available on GitHub: https://github.com/Elfsong/Mercury. Mingzhe Du, Anh Tuan Luu, Bin Ji 0002, Qian Liu 0033, See-Kiong Ng |
NeurIPS | 4 |
| 2024 | Scaling Laws with Vocabulary: Larger Models Deserve Larger VocabulariesabstractResearch on scaling large language models (LLMs) has primarily focused on model parameters and training data size, overlooking the role of vocabulary size. We investigate how vocabulary size impacts LLM scaling laws by training models ranging from 33M to 3B parameters on up to 500B characters with various vocabulary configurations. We propose three complementary approaches for predicting the compute-optimal vocabulary size: IsoFLOPs analysis, derivative estimation, and parametric fit of the loss function. Our approaches converge on the conclusion that the optimal vocabulary size depends on the compute budget, with larger models requiring larger vocabularies. Most LLMs, however, use insufficient vocabulary sizes. For example, we predict that the optimal vocabulary size of Llama2-70B should have been at least 216K, 7 times larger than its vocabulary of 32K. We validate our predictions empirically by training models with 3B parameters across different FLOPs budgets. Adopting our predicted optimal vocabulary size consistently improves downstream performance over commonly used vocabulary sizes. By increasing the vocabulary size from the conventional 32K to 43K, we improve performance on ARC-Challenge from 29.1 to 32.0 with the same 2.3e21 FLOPs. Our work highlights the importance of jointly considering tokenization and model scaling for efficient pre-training. The code and demo are available at https://github.com/sail-sg/scaling-with-vocab and https://hf.co/spaces/sail/scaling-with-vocab-demo. Chaofan Tao, Qian Liu 0033, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo 0002, Ngai Wong 0001 |
NeurIPS | 2 |
| 2024 | Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMsabstractThe recent development of chain-of-thought (CoT) decoding has enabled large language models (LLMs) to generate explicit logical reasoning paths for complex problem-solving. However, research indicates that these paths are not always deliberate and optimal. The tree-of-thought (ToT) method employs tree-searching to extensively explore the reasoning space and find better reasoning paths that CoT decoding might overlook. This deliberation, however, comes at the cost of significantly increased inference complexity. In this work, we demonstrate that fine-tuning LLMs leveraging the search tree constructed by ToT allows CoT to achieve similar or better performance, thereby avoiding the substantial inference burden. This is achieved through \emph{Chain of Preference Optimization} (CPO), where LLMs are fine-tuned to align each step of the CoT reasoning paths with those of ToT using the inherent preference information in the tree-search process. Extensive experimental results show that CPO significantly improves LLM performance in solving a variety of complex problems, including question answering, fact verification, and arithmetic reasoning, demonstrating its effectiveness.
Our code is available at [https://github.com/sail-sg/CPO](https://github.com/sail-sg/CPO). Tianyu Pang, Qian Liu 0033 |
NeurIPS | 4 |
| 2024 | Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their DefensesabstractRecently, Anil et al. (2024) show that many-shot (up to hundreds of) demonstrations can jailbreak state-of-the-art LLMs by exploiting their long-context capability. Nevertheless, is it possible to use few-shot demonstrations to efficiently jailbreak LLMs within limited context sizes? While the vanilla few-shot jailbreaking may be inefficient, we propose improved techniques such as injecting special system tokens like [/INST] and employing demo-level random search from a collected demo pool. These simple techniques result in surprisingly effective jailbreaking against aligned LLMs (even with advanced defenses). For example, our method achieves >80% (mostly >95%) ASRs on Llama-2-7B and Llama-3-8B without multiple restarts, even if the models are enhanced by strong defenses such as perplexity detection and/or SmoothLLM, which is challenging for suffix-based jailbreaking. In addition, we conduct comprehensive and elaborate (e.g., making sure to use correct system prompts) evaluations against other aligned LLMs and advanced defenses, where our method consistently achieves nearly 100% ASRs. Our code is available at https://github.com/sail-sg/I-FSJ. Xiaosen Zheng, Tianyu Pang, Qian Liu 0033, Jing Jiang 0001 |
NeurIPS | 4 |
| 2023 | On Grounded Planning for Embodied Tasks with Language ModelsabstractLanguage models (LMs) have demonstrated their capability in possessing commonsense knowledge of the physical world, a crucial aspect of performing tasks in everyday life. However, it remains unclear whether they have the capacity to generate grounded, executable plans for embodied tasks. This is a challenging task as LMs lack the ability to perceive the environment through vision and feedback from the physical environment. In this paper, we address this important research question and present the first investigation into the topic. Our novel problem formulation, named G-PlanET, inputs a high-level goal and a data table about objects in a specific environment, and then outputs a step-by-step actionable plan for a robotic agent to follow. To facilitate the study, we establish an evaluation protocol and design a dedicated metric, KAS, to assess the quality of the plans. Our experiments demonstrate that the use of tables for encoding the environment and an iterative decoding strategy can significantly enhance the LMs' ability in grounded planning. Our analysis also reveals interesting and non-trivial findings. Bill Y. Lin, Chengsong Huang, Qian Liu 0033, Wenda Gu, Sam Sommerer, Xiang Ren 0001 |
AAAI | 3 |
| 2023 | Active Retrieval Augmented GenerationabstractZhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, Graham Neubig. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu 0033, Jane Dwivedi-Yu, Yiming Yang 0002, Jamie Callan, Graham Neubig |
EMNLP | 5 |
| 2023 | SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific TablesabstractCurrent scientific fact-checking benchmarks exhibit several shortcomings, such as biases arising from crowd-sourced claims and an overreliance on text-based evidence.We present SCITAB, a challenging evaluation dataset consisting of 1.2K expert-verified scientific claims that 1) originate from authentic scientific publications and 2) require compositional reasoning for verification.The claims are paired with evidence-containing scientific tables annotated with labels.Through extensive evaluations, we demonstrate that SCITAB poses a significant challenge to state-of-the-art models, including table-based pretraining models and large language models.All models except GPT-4 achieved performance barely above random guessing.Popular prompting techniques, such as Chain-of-Thought, do not achieve much performance gains on SCITAB.Our analysis uncovers several unique challenges posed by SCITAB, including table grounding, claim ambiguity, and compositional reasoning. Xinyuan Lu, Liangming Pan, Qian Liu 0033, Preslav Nakov, Min-Yen Kan |
EMNLP | 3 |
| 2023 | Generative Table Pre-training Empowers Models for Tabular PredictionabstractRecently, the topic of table pre-training has attracted considerable research interest.However, how to employ table pre-training to boost the performance of tabular prediction remains an open challenge.In this paper, we propose TAPTAP, the first attempt that leverages table pre-training to empower models for tabular prediction.After pre-training on a large corpus of real-world tabular data, TAPTAP can generate high-quality synthetic tables to support various applications on tabular data, including privacy protection, low resource regime, missing value imputation, and imbalanced classification.Extensive experiments on 12 datasets demonstrate that TAPTAP outperforms a total of 16 baselines in different scenarios.Meanwhile, it can be easily combined with various backbone models, including LightGBM, Multilayer Perceptron (MLP) and Transformer.Moreover, with the aid of table pre-training, models trained using synthetic data generated by TAPTAP can even compete with models using the original dataset on half of the experimental datasets, marking a milestone in the development of synthetic tabular data generation.The code and datasets are available at https: //github.com/ZhangTP1996/TapTap. Tianping Zhang, Shaowen Wang 0002, Shuicheng Yan, Li Jian, Qian Liu 0033 |
EMNLP | 5 |
| 2023 | Learning on Large-scale Text-attributed Graphs via Variational Inference
Jianan Zhao 0002, Meng Qu, Chaozhuo Li, Hao Yan 0004, Qian Liu 0033, Rui Li 0086, Xing Xie 0001, Jian Tang 0005 |
ICLR | 5 |
| 2023 | Bag of Tricks for Training Data Extraction from Language ModelsabstractWith the advance of language models, privacy protection is receiving more attention. Training data extraction is therefore of great importance, as it can serve as a potential tool to assess privacy leakage. However, due to the difficulty of this task, most of the existing methods are proof-of-concept and still not effective enough. In this paper, we investigate and benchmark tricks for improving training data extraction using a publicly available dataset. Because most existing extraction methods use a pipeline of generating-then-ranking, i.e., generating text candidates as potential training data and then ranking them based on specific criteria, our research focuses on the tricks for both text generation (e.g., sampling strategy) and text ranking (e.g., token-level criteria). The experimental results show that several previously overlooked tricks can be crucial to the success of training data extraction. Based on the GPT-Neo 1.3B evaluation results, our proposed tricks outperform the baseline by a large margin in most cases, providing a much stronger baseline for future research. The code is available at https://github.com/weichen-yu/LM-Extraction. Weichen Yu, Tianyu Pang, Qian Liu 0033, Bingyi Kang, Shuicheng Yan |
ICML | 3 |
| 2023 | OpenFE: Automated Feature Generation with Expert-level PerformanceabstractThe goal of automated feature generation is to liberate machine learning experts from the laborious task of manual feature generation, which is crucial for improving the learning performance of tabular data. The major challenge in automated feature generation is to efficiently and accurately identify effective features from a vast pool of candidate features. In this paper, we present OpenFE, an automated feature generation tool that provides competitive results against machine learning experts. OpenFE achieves high efficiency and accuracy with two components: 1) a novel feature boosting method for accurately evaluating the incremental performance of candidate features and 2) a two-stage pruning algorithm that performs feature pruning in a coarse-to-fine manner. Extensive experiments on ten benchmark datasets show that OpenFE outperforms existing baseline methods by a large margin. We further evaluate OpenFE in two Kaggle competitions with thousands of data science teams participating. In the two competitions, features generated by OpenFE with a simple baseline model can beat 99.3% and 99.6% data science teams respectively. In addition to the empirical results, we provide a theoretical perspective to show that feature generation can be beneficial in a simple yet representative setting. Tianping Zhang, Zheyu Aqa Zhang, Zhiyuan Fan, Haoyan Luo, Qian Liu 0033, Li Jian |
ICML | 6 |
| 2022 | CGIM: A Cycle Guided Interactive Learning Model for Consistency Identification in Task-oriented DialogueabstractConsistency identification in task-oriented dialog (CI-ToD) usually consists of three subtasks, aiming to identify inconsistency between current system response and current user response, dialog history and the corresponding knowledge base. This work aims to solve CI-ToD task by introducing an explicit interaction paradigm, Cycle Guided Interactive learning Model (CGIM), which achieves to make information exchange explicitly from all the three tasks. Specifically, CGIM relies on two core insights, referred to as guided multi-head attention module and cycle interactive mechanism, that collaborate from each other. On the one hand, each two tasks are linked with the guided multi-head attention module, aiming to explicitly model the interaction across two related tasks. On the other hand, we further introduce cycle interactive mechanism that focuses on facilitating model to exchange information among the three correlated sub-tasks via a cycle interaction manner. Experimental results on CI-ToD benchmark show that our model achieves the state-of-the-art performance, pushing the overall score to 56.3% (5.0% point absolute improvement). In addition, we find that CGIM is robust to the initial task flow order. Libo Qin 0001, Qiguang Chen, Tianbao Xie, Qian Liu 0033, Shijue Huang, Wanxiang Che, Zhou Yu 0005 |
COLING | 4 |
| 2022 | Exploring the Secrets Behind the Learning Difficulty of Meaning Representations for Semantic ParsingabstractPrevious research has shown that the design of Meaning Representation (MR) greatly influences model performance of a neural semantic parser.Therefore, designing a good MR is a long-term goal for semantic parsing.However, it is still an art as there is no quantitative indicator that can tell us which MR among a set of candidates may have the best final model performance.In practice, in order to select an MR, researchers often have to go through the whole training-testing process for all MR candidates, and the process often costs a lot.In this paper, we propose a data-aware metric called ISS (denoting incremental structural stability) of MRs, and demonstrate that ISS is highly correlated with model performance.The finding shows that ISS can be used as an indicator for designing MRs to avoid the costly training-testing process. Zhenwen Li, Qian Liu 0033, Jian-Guang Lou, Tao Xie 0001 |
EMNLP | 3 |
| 2022 | Reasoning Like Program ExecutorsabstractReasoning over natural language is a longstanding goal for the research community.However, studies have shown that existing language models are inadequate in reasoning.To address the issue, we present POET, a novel reasoning pre-training paradigm.Through pretraining language models with programs and their execution results, POET empowers language models to harvest the reasoning knowledge possessed by program executors via a data-driven approach.POET is conceptually simple and can be instantiated by different kinds of program executors.In this paper, we showcase two simple instances POET-Math and POET-Logic, in addition to a complex instance, POET-SQL.Experimental results on six benchmarks demonstrate that POET can significantly boost model performance in natural language reasoning, such as numerical reasoning, logical reasoning, and multi-hop reasoning.POET opens a new gate on reasoningenhancement pre-training, and we hope our analysis would shed light on the future research of reasoning like program executors. Xinyu Pi, Qian Liu 0033, Bei Chen 0008, Morteza Ziyadi, Zeqi Lin, Qiang Fu 0015, Yan Gao 0002, Jian-Guang Lou, Weizhu Chen |
EMNLP | 2 |
| 2022 | TAPEX: Table Pre-training via Learning a Neural SQL Executor
Qian Liu 0033, Bei Chen 0008, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, Jian-Guang Lou |
ICLR | 1 |
| 2022 | Reasoning over Hybrid Chain for Table-and-Text Open Domain Question AnsweringabstractTabular and textual question answering requires systems to perform reasoning over heterogeneous information, considering table structure, and the connections among table and text. In this paper, we propose a ChAin-centric Reasoning and Pre-training framework (CARP). CARP utilizes hybrid chain to model the explicit intermediate reasoning process across table and text for question answering. We also propose a novel chain-centric pre-training method, to enhance the pre-trained model in identifying the cross-modality reasoning process and alleviating the data sparsity problem. This method constructs the large-scale reasoning corpus by synthesizing pseudo heterogeneous reasoning paths from Wikipedia and generating corresponding questions. We evaluate our system on OTT-QA, a large-scale table-and-text open-domain question answering benchmark, and our system achieves the state-of-the-art performance. Further analyses illustrate that the explicit hybrid chain offers substantial performance improvement and interpretablity of the intermediate reasoning process, and the chain-centric pre-training boosts the performance on the chain extraction. Wanjun Zhong, Junjie Huang 0008, Qian Liu 0033, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001, Nan Duan 0001 |
IJCAI | 3 |
| 2021 | Chase: A Large-Scale and Pragmatic Chinese Dataset for Cross-Database Context-Dependent Text-to-SQLabstractJiaqi Guo, Ziliang Si, Yu Wang, Qian Liu, Ming Fan, Jian-Guang Lou, Zijiang Yang, Ting Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ziliang Si, Yu Wang 0093, Qian Liu 0033, Ming Fan 0002, Jian-Guang Lou, Zijiang Yang 0006, Ting Liu 0002 |
ACL/IJCNLP (1) | 4 |
| 2021 | Keep the Structure: A Latent Shift-Reduce Parser for Semantic ParsingabstractTraditional end-to-end semantic parsing models treat a natural language utterance as a holonomic structure. However, hierarchical structures exist in natural languages, which also align with the hierarchical structures of logical forms. In this paper, we propose a latent shift-reduce parser, called LASP, which decomposes both natural language queries and logical form expressions according to their hierarchical structures and finds local alignment between them to enhance semantic parsing. LASP consists of a base parser and a shift-reduce splitter. The splitter dynamically separates an NL query into several spans. The base parser converts the relevant simple spans into logical forms, which are further combined to obtain the final logical form. We conducted empirical studies on two datasets across different domains and different types of logical forms. The results demonstrate that the proposed method significantly improves the performance of semantic parsing, especially on unseen scenarios. Bei Chen 0008, Qian Liu 0033, Yan Gao 0002, Jian-Guang Lou, Yan Zhang 0117, Dongmei Zhang 0001 |
IJCAI | 3 |
| 2020 | You Impress Me: Dialogue Generation via Mutual Persona PerceptionabstractDespite the continuing efforts to improve the engagingness and consistency of chit-chat dialogue systems, the majority of current work simply focus on mimicking human-like responses, leaving understudied the aspects of modeling understanding between interlocutors.The research in cognitive science, instead, suggests that understanding is an essential signal for a high-quality chit-chat conversation.Motivated by this, we propose P 2 BOT, a transmitter-receiver based framework with the aim of explicitly modeling understanding.Specifically, P 2 BOT incorporates mutual persona perception to enhance the quality of personalized dialogue generation.Experiments on a large public dataset, PERSONA-CHAT, demonstrate the effectiveness of our approach, with a considerable boost over the state-of-theart baselines across both automatic metrics and human evaluations. Qian Liu 0033, Bei Chen 0008, Jian-Guang Lou, Dongmei Zhang 0001 |
ACL | 1 |
| 2020 | Benchmarking Meaning Representations in Neural Semantic ParsingabstractMeaning representation is an important component of semantic parsing.Although researchers have designed a lot of meaning representations, recent work focuses on only a few of them.Thus, the impact of meaning representation on semantic parsing is less understood.Furthermore, existing work's performance is often not comprehensively evaluated due to the lack of readily-available execution engines.Upon identifying these gaps, we propose UNIMER, a new unified benchmark on meaning representations, by integrating existing semantic parsing datasets, completing the missing logical forms, and implementing the missing execution engines.The resulting unified benchmark contains the complete enumeration of logical forms and execution engines over three datasets × four meaning representations.A thorough experimental study on UNIMER reveals that neural semantic parsing approaches exhibit notably different performance when they are trained to generate different meaning representations.Also, program alias and grammar rules heavily impact the performance of different meaning representations.Our benchmark, execution engines and implementation can be found on: https Qian Liu 0033, Jian-Guang Lou, Zhenwen Li, Xueqing Liu 0001, Tao Xie 0001, Ting Liu 0002 |
EMNLP (1) | 2 |
| 2020 | "What Do You Mean by That?" A Parser-Independent Interactive Approach for Enhancing Text-to-SQLabstractIn Natural Language Interfaces to Databases systems, the text-to-SQL technique allows users to query databases by using natural language questions. Though significant progress in this area has been made recently, most parsers may fall short when they are deployed in real systems. One main reason stems from the difficulty of fully understanding the users' natural language questions. In this paper, we include human in the loop and present a novel parser-independent interactive approach (PIIA) that interacts with users using multi-choice questions and can easily work with arbitrary parsers. Experiments were conducted on two cross-domain datasets, the WikiSQL and the more complex Spider, with five state-of-the-art parsers. These demonstrated that PIIA is capable of enhancing the text-to-SQL performance with limited interaction turns by using both simulation and human evaluation. Bei Chen 0008, Qian Liu 0033, Yan Gao 0002, Jian-Guang Lou, Yan Zhang 0117, Dongmei Zhang 0001 |
EMNLP (1) | 3 |
| 2020 | Incomplete Utterance Rewriting as Semantic SegmentationabstractRecent years the task of incomplete utterance rewriting has raised a large attention.Previous works usually shape it as a machine translation task and employ sequence to sequence based architecture with copy mechanism.In this paper, we present a novel and extensive approach, which formulates it as a semantic segmentation task.Instead of generating from scratch, such a formulation introduces edit operations and shapes the problem as prediction of a word-level edit matrix.Benefiting from being able to capture both local and global information, our approach achieves state-ofthe-art performance on several public datasets.Furthermore, our approach is four times faster than the standard approach in inference. Qian Liu 0033, Bei Chen 0008, Jian-Guang Lou, Dongmei Zhang 0001 |
EMNLP (1) | 1 |
| 2020 | How Far are We from Effective Context Modeling? An Exploratory Study on Semantic Parsing in ContextabstractRecently semantic parsing in context has received a considerable attention, which is challenging since there are complex contextual phenomena. Previous works verified their proposed methods in limited scenarios, which motivates us to conduct an exploratory study on context modeling methods under real-world semantic parsing in context. We present a grammar-based decoding semantic parser and adapt typical context modeling methods on top of it. We evaluate 13 context modeling methods on two large complex cross-domain datasets, and our best model achieves state-of-the-art performances on both datasets with significant improvements. Furthermore, we summarize the most frequent contextual phenomena, with a fine-grained analysis on representative models, which may shed light on potential research directions. Our code is available at https://github.com/microsoft/ContextualSP. Qian Liu 0033, Bei Chen 0008, Jian-Guang Lou, Dongmei Zhang 0001 |
IJCAI | 1 |
| 2020 | RECPARSER: A Recursive Semantic Parsing Framework for Text-to-SQL TaskabstractNeural semantic parsers usually fail to parse long and complicated utterances into nested SQL queries, due to the large search space. In this paper, we propose a novel recursive semantic parsing framework called RECPARSER to generate the nested SQL query layer-by-layer. It decomposes the complicated nested SQL query generation problem into several progressive non-nested SQL query generation problems. Furthermore, we propose a novel Question Decomposer module to explicitly encourage RECPARSER to focus on different components of an utterance when predicting SQL queries of different layers. Experiments on the Spider dataset show that our approach is more effective compared to the previous works at predicting the nested SQL queries. In addition, we achieve an overall accuracy that is comparable with state-of-the-art approaches. Yan Gao 0002, Bei Chen 0008, Qian Liu 0033, Jian-Guang Lou, Fei Teng 0001, Dongmei Zhang 0001 |
IJCAI | 5 |
| 2020 | Compositional Generalization by Learning Analytical ExpressionsabstractCompositional generalization is a basic and essential intellective capability of human beings, which allows us to recombine known parts readily. However, existing neural network based models have been proven to be extremely deficient in such a capability. Inspired by work in cognition which argues compositionality can be captured by variable slots with symbolic functions, we present a refreshing view that connects a memory-augmented neural model with analytical expressions, to achieve compositional generalization. Our model consists of two cooperative neural modules, Composer and Solver, fitting well with the cognitive argument while being able to be trained in an end-to-end manner via a hierarchical reinforcement learning algorithm. Experiments on the well-known benchmark SCAN demonstrate that our model seizes a great ability of compositional generalization, solving all challenges addressed by previous works with 100% accuracies. Qian Liu 0033, Shengnan An, Jian-Guang Lou, Bei Chen 0008, Zeqi Lin, Yan Gao 0002, Nanning Zheng 0001, Dongmei Zhang 0001 |
NeurIPS | 1 |
| 2019 | FANDA: A Novel Approach to Perform Follow-Up Query AnalysisabstractRecent work on Natural Language Interfaces to Databases (NLIDB) has attracted considerable attention. NLIDB allow users to search databases using natural language instead of SQL-like query languages. While saving the users from having to learn query languages, multi-turn interaction with NLIDB usually involves multiple queries where contextual information is vital to understand the users’ query intents. In this paper, we address a typical contextual understanding problem, termed as follow-up query analysis. In spite of its ubiquity, follow-up query analysis has not been well studied due to two primary obstacles: the multifarious nature of follow-up query scenarios and the lack of high-quality datasets. Our work summarizes typical follow-up query scenarios and provides a new FollowUp dataset with 1000 query triples on 120 tables. Moreover, we propose a novel approach FANDA, which takes into account the structures of queries and employs a ranking model with weakly supervised max-margin learning. The experimental results on FollowUp demonstrate the superiority of FANDA over multiple baselines across multiple metrics. Qian Liu 0033, Bei Chen 0008, Jian-Guang Lou, Dongmei Zhang 0001 |
AAAI | 1 |
| 2019 | A Split-and-Recombine Approach for Follow-up Query AnalysisabstractQian Liu, Bei Chen, Haoyan Liu, Jian-Guang Lou, Lei Fang, Bin Zhou, Dongmei Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Qian Liu 0033, Bei Chen 0008, Haoyan Liu 0001, Jian-Guang Lou, Lei Fang 0004, Dongmei Zhang 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Leveraging Adjective-Noun Phrasing Knowledge for Comparison Relation Prediction in Text-to-SQLabstractHaoyan Liu, Lei Fang, Qian Liu, Bei Chen, Jian-Guang Lou, Zhoujun Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Haoyan Liu 0001, Lei Fang 0004, Qian Liu 0033, Bei Chen 0008, Jian-Guang Lou, Zhoujun Li 0001 |
EMNLP/IJCNLP (1) | 3 |