VLDB 2026 Research / reviewers in the wild / expert
Ziqi Wang 0003
dblp:38/8097-3
· DBLP profile ↗
13ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0001-7772-0338ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Entailment-Preserving First-order Logic Representations in Natural Language EntailmentabstractFirst-order logic (FOL) is often used to represent logical entailment, but determining natural language (NL) entailment using FOL remains a challenge. To address this, we propose the Entailment-Preserving FOL representations (EPF) task and introduce reference-free evaluation metrics for EPF (Entailment-Preserving Rate (EPR) family). In EPF, one should generate FOL representations from multi-premise NL entailment data (e.g., EntailmentBank) so that the automatic prover’s result preserves the entailment labels. Furthermore, we propose a training method specialized for the task, iterative learning-to-rank, which trains an NL-to-FOL translator by using the natural language entailment labels as verifiable rewards. Our method achieves a 1.8–2.7% improvement in EPR and a 17.4–20.6% increase in EPR@16 compared to diverse baselines in three datasets. Further analyses reveal that iterative learning-to-rank effectively suppresses the arbitrariness of FOL representation by reducing the diversity of predicate signatures, and maintains strong performance across diverse inference types and out-of-domain data. Jinu Lee 0001, Runzhi Ma, Vincent Han, Ziqi Wang 0003, Heng Ji 0001, Julia Hockenmaier |
ACL (1) | 5 |
| 2025 | Model Extrapolation Expedites AlignmentabstractGiven the high computational cost of preference alignment training of large language models (LLMs), exploring efficient methods to reduce the training overhead remains an important and compelling research problem.Motivated by the observation that alignment training typically involves only small parameter changes without injecting new knowledge into models, we propose a straightforward method called EXPO (model extrapolation) to expedite LLMs' alignment with human preferences.Given a partially-trained model and its initial SFT checkpoint, EXPO improves the implicit optimization objective of alignment training by simply amplifying the parameter change based on a first-order approximation, without any additional training overhead.Through controlled experiments, we demonstrate that EXPO boosts a DPO model trained with only 20% steps to outperform the fullytrained one.Moreover, we show that EXPO notably improves existing open-source LLMs (ranging from 1.8B to 70B parameters) on the leading AlpacaEval 2.0 and MT-Bench benchmarks, which highlights EXPO's broader utility in efficiently enhancing LLM alignment. Chujie Zheng, Ziqi Wang 0003, Heng Ji 0001, Minlie Huang, Nanyun Peng 0001 |
ACL (1) | 2 |
| 2025 | Eliminating Position Bias of Language Models: A Mechanistic ApproachabstractPosition bias has proven to be a prevalent issue of modern language models (LMs), where the models prioritize content based on its position within the given context. This bias often leads to unexpected model failures and hurts performance, robustness, and reliability across various applications. A simple mechanistic analysis attributes the position bias to two components employed in nearly all state-of-the-art LMs: causal attention and position embedding. Based on the analyses, we propose to **eliminate** position bias (e.g., different retrieved documents' orders in QA affect performance) with a **training-free zero-shot** approach. Our method changes the causal attention to bidirectional attention between documents and utilizes model attention values to decide the relative orders of documents instead of using the order provided in input prompts, therefore enabling Position-INvariant inferencE (PINE) at the document level. By eliminating position bias, models achieve better performance and reliability in downstream tasks, including LM-as-a-judge, retrieval-augmented QA, molecule generation, and math reasoning. Notably, PINE is especially useful when adapting LMs for evaluating reasoning pairs: it consistently provides $8$ to $10$ percentage points performance gains, making Llama-3-70B-Instruct perform even better than GPT-4-0125-preview and GPT-4o-2024-08-06 on the RewardBench reasoning set. Ziqi Wang 0003, Hanlin Zhang 0002, Xiner Li, Kuan-Hao Huang, Chi Han, Shuiwang Ji, Sham M. Kakade, Hao Peng 0009, Heng Ji 0001 |
ICLR | 1 |
| 2025 | LLM Alignment as Retriever Optimization: An Information Retrieval PerspectiveabstractLarge Language Models (LLMs) have revolutionized artificial intelligence with capabilities in reasoning, coding, and communication, driving innovation across industries. Their true potential depends on effective alignment to ensure correct, trustworthy and ethical behavior, addressing challenges like misinformation, hallucinations, bias and misuse. While existing Reinforcement Learning (RL)-based alignment methods are notoriously complex, direct optimization approaches offer a simpler alternative. In this work, we introduce a novel direct optimization approach for LLM alignment by drawing on established Information Retrieval (IR) principles. We present a systematic framework that bridges LLM alignment and IR methodologies, mapping LLM generation and reward models to IR’s retriever-reranker paradigm. Building on this foundation, we propose LLM Alignment as Retriever Preference Optimization (LarPO), a new alignment method that enhances overall alignment quality. Extensive experiments validate LarPO’s effectiveness with 38.9 % and 13.7 % averaged improvement on AlpacaEval2 and MixEval-Hard respectively. Our work opens new avenues for advancing LLM alignment by integrating IR foundations, offering a promising direction for future research. Bowen Jin, Jinsung Yoon, Zhen Qin 0001, Ziqi Wang 0003, Wei Xiong 0015, Yu Meng 0001, Jiawei Han 0001, Sercan Ö. Arik |
ICML | 4 |
| 2024 | Enabling Lanuguage Models to Implicitly Learn Self-ImprovementabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities in open-ended text generation tasks. However, the inherent open-ended nature of these tasks implies that there is always room for improvement in the quality of model responses. To address this challenge, various approaches have been proposed to enhance the performance of LLMs. There has been a growing focus on enabling LLMs to self-improve their response quality, thereby reducing the reliance on extensive human annotation efforts for collecting diverse and high-quality training data. Recently, prompting-based methods have been widely explored among self-improvement methods owing to their effectiveness, efficiency, and convenience. However, those methods usually require explicitly and thoroughly written rubrics as inputs to LLMs. It is expensive and challenging to manually derive and provide all necessary rubrics with a real-world complex goal for improvement (e.g., being more helpfulness and less harmful). To this end, we propose an imPlicit self-ImprovemenT (PIT) framework that implicitly learns the improvement goal from human preference data. PIT only requires preference data that are used to train reward models with no extra human efforts. Specifically, we reformulate the training objective of reinforcement learning from human feedback (RLHF) -- instead of maximizing response quality for a given input, we maximize the quality gap of the response conditioned on a reference response. In this way, PIT is implicitly trained with the improvement goal of better aligning with human preferences. Experiments on two real-world datasets and one synthetic dataset show that our method significantly outperforms prompting-based methods. Ziqi Wang 0003, Le Hou, Tianjian Lu, Yuexin Wu, Yunxuan Li, Hongkun Yu 0001, Heng Ji 0001 |
ICLR | 1 |
| 2024 | Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintabstractThis paper studies the theoretical framework of the alignment process of generative models with Reinforcement Learning from Human Feedback (RLHF). We consider a standard mathematical formulation, the reverse-KL regularized contextual bandit for RLHF. Despite its widespread practical application, a rigorous theoretical analysis of this formulation remains open. We investigate its behavior in three distinct settings—offline, online, and hybrid—and propose efficient algorithms with finite-sample theoretical guarantees. Moving towards practical applications, our framework, with a robust approximation of the information-theoretical policy improvement oracle, naturally gives rise to several novel RLHF algorithms. This includes an iterative version of the Direct Preference Optimization (DPO) algorithm for online settings, and a multi-step rejection sampling strategy for offline scenarios. Our empirical evaluations on real-world alignment experiment of large language model demonstrate that these proposed methods significantly surpass existing strong baselines, such as DPO and Rejection Sampling Optimization (RSO), showcasing the connections between solid theoretical foundations and their potent practical implementations. Wei Xiong 0015, Hanze Dong, Chenlu Ye, Ziqi Wang 0003, Han Zhong 0001, Heng Ji 0001, Nan Jiang 0008, Tong Zhang 0001 |
ICML | 4 |
| 2023 | Augmentation with Projection: Towards an Effective and Efficient Data Augmentation Paradigm for Distillation
Ziqi Wang 0003, Yuexin Wu, Frederick Liu, Daogao Liu, Le Hou, Hongkun Yu 0001, Jing Li 0049, Heng Ji 0001 |
ICLR | 1 |
| 2023 | Recognizing Object by Components With Human Prior Knowledge Enhances Adversarial Robustness of Deep Neural NetworksabstractAdversarial attacks can easily fool object recognition systems based on deep neural networks (DNNs). Although many defense methods have been proposed in recent years, most of them can still be adaptively evaded. One reason for the weak adversarial robustness may be that DNNs are only supervised by category labels and do not have part-based inductive bias like the recognition process of humans. Inspired by a well-known theory in cognitive psychology - recognition-by-components, we propose a novel object recognition model ROCK (Recognizing Object by Components with human prior Knowledge). It first segments parts of objects from images, then scores part segmentation results with predefined human prior knowledge, and finally outputs prediction based on the scores. The first stage of ROCK corresponds to the process of decomposing objects into parts in human vision. The second stage corresponds to the decision process of the human brain. ROCK shows better robustness than classical recognition models across various attack settings. These results encourage researchers to rethink the rationality of currently widely-used DNN-based object recognition models and explore the potential of part-based models, once important but recently ignored, for improving robustness. Xiao Li 0028, Ziqi Wang 0003, Bo Zhang 0010, Fuchun Sun 0001, Xiaolin Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | CLEVE: Contrastive Pre-training for Event ExtractionabstractZiqi Wang, Xiaozhi Wang, Xu Han, Yankai Lin, Lei Hou, Zhiyuan Liu, Peng Li, Juanzi Li, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ziqi Wang 0003, Xiaozhi Wang, Xu Han 0007, Yankai Lin 0001, Lei Hou 0001, Zhiyuan Liu 0001, Peng Li 0030, Juan-Zi Li, Jie Zhou 0016 |
ACL/IJCNLP (1) | 1 |
| 2020 | MAVEN: A Massive General Domain Event Detection DatasetabstractXiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang, Rong Han, Zhiyuan Liu, Juanzi Li, Peng Li, Yankai Lin, Jie Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Xiaozhi Wang, Ziqi Wang 0003, Xu Han 0007, Wangyi Jiang, Zhiyuan Liu 0001, Juan-Zi Li, Peng Li 0030, Yankai Lin 0001, Jie Zhou 0016 |
EMNLP (1) | 2 |
| 2020 | Learning from Explanations with Neural Execution Tree
Ziqi Wang 0003, Yujia Qin, Wenxuan Zhou 0002, Jun Yan 0012, Qinyuan Ye, Leonardo Neves, Zhiyuan Liu 0001, Xiang Ren 0001 |
ICLR | 1 |
| 2020 | NERO: A Neural Rule Grounding Framework for Label-Efficient Relation ExtractionabstractDeep neural models for relation extraction tend to be less reliable when perfectly labeled data is limited, despite their success in label-sufficient scenarios. Instead of seeking more instance-level labels from human annotators, here we propose to annotate frequent surface patterns to form labeling rules. These rules can be automatically mined from large text corpora and generalized via a soft rule matching mechanism. Prior works use labeling rules in an exact matching fashion, which inherently limits the coverage of sentence matching and results in the low-recall issue. In this paper, we present a neural approach to ground rules for RE, named Nero, which jointly learns a relation extraction module and a soft matching module. One can employ any neural relation extraction models as the instantiation for the RE module. The soft matching module learns to match rules with semantically similar sentences such that raw corpora can be automatically labeled and leveraged by the RE module (in a much better coverage) as augmented supervision, in addition to the exactly matched sentences. Extensive experiments and analysis on two public and widely-used datasets demonstrate the effectiveness of the proposed Nero framework, comparing with both rule-based and semi-supervised methods. Through user studies, we find that the time efficiency for a human to annotate rules and sentences are similar (0.30 vs. 0.35 min per label). In particular, Nero’s performance using 270 rules is comparable to the models trained using 3,000 labeled sentences, yielding a 9.5x speedup. Moreover, Nero can predict for unseen relations at test time and provide interpretable predictions. We release our code1 to the community for future research. Wenxuan Zhou 0002, Bill Y. Lin, Ziqi Wang 0003, Junyi Du, Leonardo Neves, Xiang Ren 0001 |
WWW | 4 |
| 2019 | HMEAE: Hierarchical Modular Event Argument ExtractionabstractXiaozhi Wang, Ziqi Wang, Xu Han, Zhiyuan Liu, Juanzi Li, Peng Li, Maosong Sun, Jie Zhou, Xiang Ren. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiaozhi Wang, Ziqi Wang 0003, Xu Han 0007, Zhiyuan Liu 0001, Juan-Zi Li, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016, Xiang Ren 0001 |
EMNLP/IJCNLP (1) | 2 |