EDBT 2026 Demo / reviewers in the wild / expert
Zhepei Wei
dblp:247/2560
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Language models and text generation · 36% Reinforcement learning · 33% Efficient and distributed learning · 9% | |
| Theoretical computer science
2 papers |
Algorithmic game theory and mechanism design · 100% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
retrieval-augmented generation |
1.7 | 2 | 2025 | InstructRAG: Instructing Retrieval-Augmented Generation via Self-Synthesized Rationales · ICLR 2025 Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation · EMNLP 2025 |
Machine learning › Reinforcement learning › bandit
federated bandit |
1.4 | 2 | 2024 | Incentivized Truthful Communication for Federated Bandits · ICLR 2024 Incentivized Communication for Federated Bandits · NeurIPS 2023 |
Algorithmic game theory and mechanism design
incentive mechanism |
1.4 | 2 | 2024 | Incentivized Truthful Communication for Federated Bandits · ICLR 2024 Incentivized Communication for Federated Bandits · NeurIPS 2023 |
Natural language and speech › Language models and text generation
alignment |
1.0 | 1 | 2026 | Aligning Large Language Models via Fully Self-Synthetic Data · ACL (1) 2026 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback › preference-based reinforcement learning
reinforcement learning from AI feedback |
1.0 | 1 | 2026 | Aligning Large Language Models via Fully Self-Synthetic Data · ACL (1) 2026 |
Natural language and speech › Language models and text generation › alignment
self-alignment |
1.0 | 1 | 2026 | Aligning Large Language Models via Fully Self-Synthetic Data · ACL (1) 2026 |
Natural language and speech › Language models and text generation
decoding |
0.9 | 1 | 2025 | AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism · ICML 2025 |
Machine learning › Efficient and distributed learning
inference acceleration |
0.9 | 1 | 2025 | AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism · ICML 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism · ICML 2025 |
Machine learning › Reinforcement learning
multi-turn reinforcement learning |
0.9 | 1 | 2025 | WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning · EMNLP 2025 |
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards |
0.9 | 1 | 2025 | The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning · NeurIPS 2025 |
Natural language and speech › Language models and text generation › LLM agents
web agents |
0.9 | 1 | 2025 | WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning · EMNLP 2025 |
Machine learning › Reinforcement learning › bandit
bandit learning |
0.8 | 1 | 2024 | Incentivized Truthful Communication for Federated Bandits · ICLR 2024 |
Algorithmic game theory and mechanism design › mechanism design
incentive compatibility |
0.8 | 1 | 2024 | Incentivized Truthful Communication for Federated Bandits · ICLR 2024 |
Machine learning › Reinforcement learning
bandit |
0.7 | 1 | 2023 | Incentivized Communication for Federated Bandits · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
discrete latent variable model |
0.6 | 1 | 2022 | Learning Semantic Textual Similarity via Topic-informed Discrete Latent Variables · EMNLP 2022 |
Machine learning › Trustworthy machine learning
interpretability |
0.6 | 1 | 2022 | AttExplainer: Explain Transformer via Attention by Reinforcement Learning · IJCAI 2022 |
Natural language and speech › Information extraction and text analysis › text similarity › semantic similarity
semantic textual similarity |
0.6 | 1 | 2022 | Learning Semantic Textual Similarity via Topic-informed Discrete Latent Variables · EMNLP 2022 |
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
transformer interpretability |
0.6 | 1 | 2022 | AttExplainer: Explain Transformer via Attention by Reinforcement Learning · IJCAI 2022 |
Natural language and speech › Information extraction and text analysis › relation extraction
joint entity and relation extraction |
0.4 | 1 | 2020 | A Novel Cascade Binary Tagging Framework for Relational Triple Extraction · ACL 2020 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.4 | 1 | 2020 | A Novel Cascade Binary Tagging Framework for Relational Triple Extraction · ACL 2020 |
Machine learning › Generative modeling
synthetic data generation |
0.3 | 1 | 2026 | Aligning Large Language Models via Fully Self-Synthetic Data · ACL (1) 2026 |
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks |
0.3 | 1 | 2025 | InstructRAG: Instructing Retrieval-Augmented Generation via Self-Synthesized Rationales · ICLR 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.3 | 1 | 2025 | WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning · EMNLP 2025 |
Information retrieval
retrieval models |
0.3 | 1 | 2025 | Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis
topic model |
0.2 | 1 | 2022 | Learning Semantic Textual Similarity via Topic-informed Discrete Latent Variables · EMNLP 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph
knowledge graph construction |
0.1 | 1 | 2020 | A Novel Cascade Binary Tagging Framework for Relational Triple Extraction · ACL 2020 |
Methods — techniques the papers use, named apart from their topics
regret analysis · 2.8large language model · 1.7critic model · 1.7clustering · 1.7self-evaluation · 1.0preference optimization · 1.0persona role-play · 1.0self-synthesized rationales · 0.9in-context learning · 0.9end-to-end reinforcement learning · 0.9incentive mechanism design · 0.8contextual linear bandits · 0.7UCB · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Aligning Large Language Models via Fully Self-Synthetic DataabstractTraditional reinforcement learning from human feedback (RLHF) for large language models (LLMs) relies on expensive humanannotated datasets, while Reinforcement Learning from AI Feedback (RLAIF) also incurs significant costs, requiring the collection of diverse prompts and corresponding responses, often necessitating external reward models or proprietary models like GPT-4 to annotate preference pairs.In this work, we introduce Self-Alignment Optimization (SAO), a fully self-synthetic framework for LLM alignment, where all training data, including prompts (i.e., user queries), responses, and preferences, are generated by the model itself.Specifically, SAO first instructs the LLM to engage in persona role-play and generate diverse prompts and responses, which are then selfevaluated for preference optimization.Extensive experiments demonstrate that SAO effectively enhances the model's chat capabilities on standard benchmarks like AlpacaEval 2.0, while maintaining strong performance on downstream objective tasks (e.g., questionanswering, math reasoning).Our work provides a practical solution for self-improvement in aligning LLMs, and the code for reproducing our results is available at: https://github. com/SJY8460/SAO. Shangjian Yin, Zhepei Wei, Yu Meng 0001 |
ACL (1) | 2 |
| 2025 | Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented GenerationabstractRetrieval-augmented generation (RAG) enhances large language models (LLMs) by integrating external knowledge sources to address their limitations in accessing up-to-date or specialized information.A natural strategy to increase the likelihood of retrieving relevant information is to expand the number of retrieved documents.However, involving more documents could introduce significant noise, as many documents may be irrelevant or misleading, thereby reducing the overall accuracy of the generated responses.To overcome the challenge associated with handling a larger number of documents, we propose WinnowRAG, a novel RAG framework designed to systematically filter out noisy documents while preserving valuable content -a process we refer to as winnowing.WinnowRAG operates in two stages: In Stage I, we perform queryaware clustering to group similar documents and form distinct topic clusters.Each cluster is assigned to an LLM agent for generating a unique answer.In Stage II, we perform winnowing, wherein a critic LLM evaluates the outputs of multiple agents and iteratively separates useful documents from noisy ones.To retain useful documents when discarding agents, we propose two strategic merging techniques to ensure that only relevant knowledge is used for generating the final response.Crucially, WinnowRAG is model-agnostic and does not require any model fine-tuning, making it easily adaptable to various tasks.Extensive experiments on various realistic datasets demonstrate the effectiveness of WinnowRAG over state-ofthe-art baselines. Song Wang 0013, Zihan Chen 0002, Peng Wang 0105, Zhepei Wei, Zhen Tan 0001, Yu Meng 0001, Cong Shen 0001, Jundong Li |
EMNLP | 4 |
| 2025 | WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement LearningabstractZhepei Wei, Wenlin Yao, Yao Liu, Weizhi Zhang, Qin Lu, Liang Qiu, Changlong Yu, Puyang Xu, Chao Zhang, Bing Yin, Hyokun Yun, Lihong Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zhepei Wei, Wenlin Yao, Changlong Yu, Puyang Xu, Chao Zhang 0014, Hyokun Yun, Lihong Li 0001 |
EMNLP | 1 |
| 2025 | InstructRAG: Instructing Retrieval-Augmented Generation via Self-Synthesized RationalesabstractRetrieval-augmented generation (RAG) has shown promising potential to enhance the accuracy and factuality of language models (LMs). However, imperfect retrievers or noisy corpora can introduce misleading or even erroneous information to the retrieved contents, posing a significant challenge to the generation quality. Existing RAG methods typically address this challenge by directly predicting final answers despite potentially noisy inputs, resulting in an implicit denoising process that is difficult to interpret and verify. On the other hand, the acquisition of explicit denoising supervision is often costly, involving significant human efforts. In this work, we propose InstructRAG, where LMs explicitly learn the denoising process through self-synthesized rationales --- First, we instruct the LM to explain how the ground-truth answer is derived from retrieved documents. Then, these rationales can be used either as demonstrations for in-context learning of explicit denoising or as supervised fine-tuning data to train the model. Compared to standard RAG approaches, InstructRAG requires no additional supervision, allows for easier verification of the predicted answers, and effectively improves generation accuracy. Experiments show InstructRAG consistently outperforms existing RAG methods in both training-free and trainable scenarios, achieving a relative improvement of 8.3% over the best baseline method on average across five knowledge-intensive benchmarks. Extensive analysis indicates that InstructRAG scales well with increased numbers of retrieved documents and consistently exhibits robust denoising ability even in out-of-domain datasets, demonstrating strong generalizability. Zhepei Wei, Yu Meng 0001 |
ICLR | 1 |
| 2025 | AdaDecode: Accelerating LLM Decoding with Adaptive Layer ParallelismabstractLarge language models (LLMs) are increasingly used for long-content generation (e.g., long Chain-of-Thought reasoning) where decoding efficiency becomes a critical bottleneck: Autoregressive decoding is inherently limited by its sequential token generation process, where each token must be generated before the next can be processed. This sequential dependency restricts the ability to fully leverage modern hardware’s parallel processing capabilities. Existing methods like speculative decoding and layer skipping offer potential speedups but have notable drawbacks: speculative decoding relies on an auxiliary “drafter” model, which can be challenging to acquire and increases memory overhead, while layer skipping may introduce discrepancies in the generated outputs due to the missing key-value cache at skipped layers. In this work, we propose AdaDecode, which accelerates LLM decoding without requiring auxiliary models or changes to the original model parameters, while ensuring output consistency. AdaDecode leverages the insight that many tokens—particularly simple or highly-predictable ones—can accurately be generated at intermediate layers, as further layers often do not significantly alter predictions once the model reaches a certain confidence. By adaptively generating tokens at intermediate layers when confidence is high, AdaDecode enables the next token’s computation to begin immediately. The remaining layer computations for early-predicted tokens are deferred and executed in parallel with subsequent tokens when needed, maximizing hardware utilization and reducing decoding latency. A final verification step ensures that early predictions match the results of standard autoregressive decoding, preserving output parity. Experiments across diverse generation tasks shows that AdaDecode consistently achieves superior decoding throughput compared to baselines with up to 1.73$\times$ speedup, while guaranteeing output parity with standard autoregressive decoding. Zhepei Wei, Yu Meng 0001 |
ICML | 1 |
| 2025 | The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningabstractReinforcement learning with verifiable rewards (RLVR) is a promising approach for training language models (LMs) on reasoning tasks that elicit emergent long chains of thought (CoTs). Unlike supervised learning, it updates the model using both correct and incorrect samples via policy gradients. To better understand its mechanism, we decompose the learning signal into reinforcing correct responses and penalizing incorrect ones, referred to as **P**ositive and **N**egative **S**ample **R**einforcement (**PSR** and **NSR**), respectively. We train `Qwen2.5-Math-7B`, `Qwen3-4B` and `Llama-3.1-8B-Instruct` on a mathematical reasoning dataset and uncover a surprising result: training with only negative samples — without reinforcing correct responses — can be highly effective: it consistently improves performance over the base model across the entire Pass@$k$ spectrum $k$ up to 256), often matching or surpassing PPO and GRPO. In contrast, reinforcing only correct responses improves Pass@1 but degrades performance at higher $k$, due to reduced diversity. These inference-scaling trends highlight that solely penalizing incorrect responses may contribute more to performance than previously recognized. Through gradient analysis, we show that NSR works by suppressing incorrect generations and redistributing probability mass toward other plausible candidates, guided by the model's prior beliefs. It refines the model's existing knowledge rather than introducing entirely new behaviors. Building on this insight, we propose a simple variant of the RL objective that upweights NSR, and show that it consistently improves overall Pass@$k$ performance on MATH, AIME 2025, and AMC23. Our code is available at [`https://github.com/TianHongZXY/RLVR-Decomposed`](https://github.com/TianHongZXY/RLVR-Decomposed). Mengzhou Xia, Zhepei Wei, Danqi Chen 0001, Yu Meng 0001 |
NeurIPS | 3 |
| 2024 | Incentivized Truthful Communication for Federated BanditsabstractTo enhance the efficiency and practicality of federated bandit learning, recent advances have introduced incentives to motivate communication among clients, where a client participates only when the incentive offered by the server outweighs its participation cost. However, existing incentive mechanisms naively assume the clients are truthful: they all report their true cost and thus the higher cost one participating client claims, the more the server has to pay. Therefore, such mechanisms are vulnerable to strategic clients aiming to optimize their own utility by misreporting. To address this issue, we propose an incentive compatible (i.e., truthful) communication protocol, named Truth-FedBan, where the incentive for each participant is independent of its self-reported cost, and reporting the true cost is the only way to achieve the best utility. More importantly, Truth-FedBan still guarantees the sub-linear regret and communication cost without any overhead. In other words, the core conceptual contribution of this paper is, for the first time, demonstrating the possibility of simultaneously achieving incentive compatibility and nearly optimal regret in federated bandit learning. Extensive numerical studies further validate the effectiveness of our proposed solution. Zhepei Wei, Chuanhao Li 0002, Tianze Ren, Hongning Wang |
ICLR | 1 |
| 2023 | Incentivized Communication for Federated BanditsabstractMost existing works on federated bandits take it for granted that all clients are altruistic about sharing their data with the server for the collective good whenever needed. Despite their compelling theoretical guarantee on performance and communication efficiency, this assumption is overly idealistic and oftentimes violated in practice, especially when the algorithm is operated over self-interested clients, who are reluctant to share data without explicit benefits. Negligence of such self-interested behaviors can significantly affect the learning efficiency and even the practical operability of federated bandit learning. In light of this, we aim to spark new insights into this under-explored research area by formally introducing an incentivized communication problem for federated bandits, where the server shall motivate clients to share data by providing incentives. Without loss of generality, we instantiate this bandit problem with the contextual linear setting and propose the first incentivized communication protocol, namely, Inc-FedUCB, that achieves near-optimal regret with provable communication and incentive cost guarantees. Extensive empirical experiments on both synthetic and real-world datasets further validate the effectiveness of the proposed method across various environments. Zhepei Wei, Chuanhao Li 0002, Hongning Wang |
NeurIPS | 1 |
| 2022 | Learning Semantic Textual Similarity via Topic-informed Discrete Latent VariablesabstractRecently, discrete latent variable models have received a surge of interest in both Natural Language Processing (NLP) and Computer Vision (CV), attributed to their comparable performance to the continuous counterparts in representation learning, while being more interpretable in their predictions.In this paper, we develop a topic-informed discrete latent variable model for semantic textual similarity, which learns a shared latent space for sentence-pair representation via vector quantization.Compared with previous models limited to local semantic contexts, our model can explore richer semantic information via topic modeling.We further boost the performance of semantic similarity by injecting the quantized representation into a transformer-based language model with a well-designed semanticdriven attention mechanism.We demonstrate, through extensive experiments across various English language datasets, that our model is able to surpass several strong neural baselines in semantic textual similarity tasks. Erxin Yu, Lan Du 0002, Zhepei Wei, Yi Chang 0001 |
EMNLP | 4 |
| 2022 | AttExplainer: Explain Transformer via Attention by Reinforcement LearningabstractTransformer and its variants, built based on attention mechanisms, have recently achieved remarkable performance in many NLP tasks. Most existing works on Transformer explanation tend to reveal and utilize the attention matrix with human subjective intuitions in a qualitative manner. However, the huge size of dimensions directly challenges these methods to quantitatively analyze the attention matrix. Therefore, in this paper, we propose a novel reinforcement learning (RL) based framework for Transformer explanation via attention matrix, namely AttExplainer. The RL agent learns to perform step-by-step masking operations by observing the change in attention matrices. We have adapted our method to two scenarios, perturbation-based model explanation and text adversarial attack. Experiments on three widely used text classification benchmarks validate the effectiveness of the proposed method compared to state-of-the-art baselines. Additional studies show that our method is highly transferable and consistent with human intuition. The code of this paper is available at https://github.com/niuzaisheng/AttExplainer . Runliang Niu, Zhepei Wei |
IJCAI | 2 |
| 2020 | A Novel Cascade Binary Tagging Framework for Relational Triple ExtractionabstractExtracting relational triples from unstructured text is crucial for large-scale knowledge graph construction.However, few existing works excel in solving the overlapping triple problem where multiple relational triples in the same sentence share the same entities.In this work, we introduce a fresh perspective to revisit the relational triple extraction task and propose a novel cascade binary tagging framework (CASREL) derived from a principled problem formulation.Instead of treating relations as discrete labels as in previous works, our new framework models relations as functions that map subjects to objects in a sentence, which naturally handles the overlapping problem.Experiments show that the CAS-REL framework already outperforms state-ofthe-art methods even when its encoder module uses a randomly initialized BERT encoder, showing the power of the new tagging framework.It enjoys further performance boost when employing a pre-trained BERT encoder, outperforming the strongest baseline by 17.5 and 30.2 absolute gain in F1-score on two public datasets NYT and WebNLG, respectively.In-depth analysis on different scenarios of overlapping triples shows that the method delivers consistent performance gain across all these scenarios.The source code and data are released online 1 . Zhepei Wei, Jianlin Su, Yue Wang 0035, Yuan Tian 0016, Yi Chang 0001 |
ACL | 1 |