EDBT 2026 Demo / reviewers in the wild / expert
Qizhi Pei
dblp:322/9716
· DBLP profile ↗
16ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-7242-422XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMsabstractYu Li, Xiaoran Shang, Qizhi Pei, Yun Zhu, Xin Gao, Honglin Lin, Zhanping Zhong, Zhuoshi Pan, Zheng Liu, Xiaoyang Wang, Conghui He, Dahua Lin, Feng Zhao, Lijun Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yu Li 0006, Xiaoran Shang, Qizhi Pei, Yun Zhu 0007, Xin Gao 0001, Honglin Lin, Zhanping Zhong, Zhuoshi Pan, Xiaoyang Wang 0007, Conghui He, Dahua Lin, Feng Zhao 0004, Lijun Wu 0003 |
ACL (1) | 3 |
| 2026 | ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from ScratchabstractZheng Liu, Honglin Lin, Xiaoyang Wang, Xin Gao, Yu Li, Mengzhang Cai, Yun Zhu, Zhanping Zhong, Qizhi Pei, Zhuoshi Pan, Xiaoran Shang, Conghui He, Bin Cui, Wentao Zhang, Lijun Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Honglin Lin, Xiaoyang Wang 0007, Xin Gao 0001, Yu Li 0006, Mengzhang Cai, Yun Zhu 0007, Zhanping Zhong, Qizhi Pei, Zhuoshi Pan, Xiaoran Shang, Conghui He, Bin Cui 0001, Wentao Zhang 0001, Lijun Wu 0003 |
ACL (1) | 9 |
| 2026 | REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at OnceabstractZhuoshi Pan, Qizhi Pei, Yu Li, Zinan Tang, QiYao Sun, H. Vicky Zhao, Conghui He, Lijun Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhuoshi Pan, Qizhi Pei, Yu Li 0006, Zinan Tang 0001, Qiyao Sun, H. Vicky Zhao, Conghui He, Lijun Wu 0003 |
ACL (1) | 2 |
| 2026 | R³: End-to-End Reasoning-based Planning for Multi-step Retrosynthesis via Reinforcement LearningabstractYiFei Wang, Qizhi Pei, Jiangtao Feng, Yuntian Shi, Yi Duan, Lihao Wang, Lei Bai, Lijun Wu, Wei-Ying Ma, Hao Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. YiFei Wang, Qizhi Pei, Jiangtao Feng, Yuntian Shi, Yi Duan, Lei Bai 0001, Lijun Wu 0003, Wei-Ying Ma, Hao Zhou 0012 |
ACL (1) | 2 |
| 2026 | Data Pollination: An Emergent Ecological Process Driving AI Population EvolutionabstractAI development is often framed as the outcome of isolated research and engineering efforts, yet evidence from deployed systems suggests that language models interact through a shared data ecosystem.While the optimization of individual models is extensively studied, the emergent properties of this interconnected population remain largely unexplored, limiting our ability to predict long-term ecosystem trajectories.We term this process data pollination, the unintentional circulation of synthetic model outputs through shared online platforms and web-scale training corpora, and formalize it as a population-based evolutionary framework to investigate stability dynamics under synthetic data training.Our theoretical analysis and controlled experiments involving 320 language models demonstrate that population dynamics can mitigate the model collapse observed in single-lineage recursive training, yielding stable or improving performance across diverse benchmarks.Crucially, we find that ecological diversity functions as a fundamental resilience mechanism that safeguards the ecosystem against collapse, highlighting the critical importance of maintaining model diversity for sustainable AI development. Shufang Xie 0003, Qizhi Pei, Ang Lv, Jingyang Hu, Lijun Wu 0003, Rui Yan 0001 |
ACL (1) | 2 |
| 2026 | Tokenizing 3D Molecule Structure with Quantized Spherical CoordinatesabstractWhile language models (LMs) have demonstrated remarkable general-purpose capabilities across domains, including molecule generation using line notations such as SMILES and SELFIES, their direct application to 3D structure design remains constrained by two interdependent challenges. First, the difficulty in designing a 3D line notation that ensures SE(3)-invariant atomic coordinates and supports autoregressive generation. Second, the incompatibility between continuous spatial coordinates and the discrete token inputs required by LMs. To address this, we propose Mol-StrucTok, a unified framework for tokenizing 3D molecular structures. Our approach comprises two key innovations: (1) a 3D line notation—Spherical Coordinate Notation—that encodes local atomic environments in spherical coordinates, agnostic to 2D notations and inherently SE(3)-invariant; and (2) a structure-aware Vector Quantized Variational Autoencoder (VQ-VAE) for discretizing these coordinates into chemically valid tokens suitable for language model processing. Leveraging this tokenization framework, we train a GPT-2 style model for end-to-end 3D molecular generation. Empirical results demonstrate strong, task-dependent performance: in unconditional generation, Mol-StrucTok achieves diffusion-level stability with ~28× faster inference; in conditional generation, it reduces property-matching mean absolute error (MAE) by 5–8× compared to diffusion-based methods, highlighting the advantage of autoregressive contextual modeling for precise control of molecular attributes. Our code is available at https://github.com/KyGao/Mol-StrucTok. Kaiyuan Gao, Haoxiang Guan, Zun Wang 0006, Qizhi Pei, John E. Hopcroft, Kun He 0001, Lijun Wu 0003 |
KDD (1) | 5 |
| 2025 | A Strategic Coordination Framework of Small LMs Matches Large LMs in Data SynthesisabstractXin Gao, Qizhi Pei, Zinan Tang, Yu Li, Honglin Lin, Jiang Wu, Lijun Wu, Conghui He. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xin Gao 0001, Qizhi Pei, Zinan Tang 0001, Yu Li 0006, Honglin Lin, Jiang Wu 0003, Lijun Wu 0003, Conghui He |
ACL (1) | 2 |
| 2025 | MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction FusionabstractQizhi Pei, Lijun Wu, Zhuoshi Pan, Yu Li, Honglin Lin, Chenlin Ming, Xin Gao, Conghui He, Rui Yan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Qizhi Pei, Lijun Wu 0003, Zhuoshi Pan, Yu Li 0006, Honglin Lin, Chenlin Ming, Xin Gao 0001, Conghui He, Rui Yan 0001 |
ACL (1) | 1 |
| 2025 | Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop LearningabstractSupervised Fine-Tuning (SFT) Large Language Models (LLMs) fundamentally rely on highquality training data.While data selection and data synthesis are two common strategies to improve data quality, existing approaches often face limitations in static dataset curation that fail to adapt to evolving model capabilities.In this paper, we introduce Middo, a self-evolving Model-informed dynamic data optimization framework that uses model-aware data selection and context-preserving data refinement.Unlike conventional one-off filtering/synthesis methods, our framework establishes a closed-loop optimization system: (1) A self-referential diagnostic module proactively identifies suboptimal samples through tri-axial model signals-loss patterns (complexity), embedding cluster dynamics (diversity), and selfalignment scores (quality); (2) An adaptive optimization engine then transforms suboptimal samples into pedagogically valuable training points while preserving semantic integrity; (3) This optimization process continuously evolves with the model's capability through dynamic learning principles.Experiments on multiple benchmarks demonstrate that our Middo consistently enhances the quality of seed data and boosts LLMs' performance, improving accuracy by 7.15% on average while maintaining the original dataset scale.This work establishes a new paradigm for sustainable LLM training through dynamic human-AI co-evolution of data and models. Zinan Tang 0001, Xin Gao 0001, Qizhi Pei, Zhuoshi Pan, Mengzhang Cai, Jiang Wu 0003, Conghui He, Lijun Wu 0003 |
EMNLP | 3 |
| 2025 | 3D-MolT5: Leveraging Discrete Structural Information for Molecule-Text ModelingabstractThe integration of molecular and natural language representations has emerged as a focal point in molecular science, with recent advancements in Language Models (LMs) demonstrating significant potential for comprehensive modeling of both domains. However, existing approaches face notable limitations, particularly in their neglect of three-dimensional (3D) information, which is crucial for understanding molecular structures and functions. While some efforts have been made to incorporate 3D molecular information into LMs using external structure encoding modules, significant difficulties remain, such as insufficient interaction across modalities in pre-training and challenges in modality alignment. To address the limitations, we propose \textbf{3D-MolT5}, a unified framework designed to model molecule in both sequence and 3D structure spaces. The key innovation of our approach lies in mapping fine-grained 3D substructure representations into a specialized 3D token vocabulary. This methodology facilitates the seamless integration of sequence and structure representations in a tokenized format, enabling 3D-MolT5 to encode molecular sequences, molecular structures, and text sequences within a unified architecture. Leveraging this tokenized input strategy, we build a foundation model that unifies the sequence and structure data formats. We then conduct joint pre-training with multi-task objectives to enhance the model's comprehension of these diverse modalities within a shared representation space. Thus, our approach significantly improves cross-modal interaction and alignment, addressing key challenges in previous work. Further instruction tuning demonstrated that our 3D-MolT5 has strong generalization ability and surpasses existing methods with superior performance in multiple downstream tasks, such as nearly 70\% improvement on the molecular property prediction task compared to state-of-the-art methods. Our code is available at \url{https://github.com/QizhiPei/3D-MolT5}. Qizhi Pei, Rui Yan 0001, Kaiyuan Gao, Jinhua Zhu 0001, Lijun Wu 0003 |
ICLR | 1 |
| 2025 | FABind+: Enhancing Molecular Docking through Improved Pocket Prediction and Pose GenerationabstractMolecular docking is a pivotal process in drug discovery. While traditional techniques rely on extensive sampling and simulation governed by physical principles, deep learning has emerged as a promising alternative, offering improvements in both accuracy and efficiency. Building upon the foundational work of FABind, a model focused on speed and accuracy, we introduce FABind+, an enhanced iteration that significantly elevates the performance of its predecessor. We identify pocket prediction as a critical bottleneck in molecular docking and introduce an enhanced approach. In addition to the pocket prediction module, the docking module has also been upgraded with permutation loss and a more refined model design. These designs enable the regression-based FABind+ to surpass most of the generative models. In contrast, while sampling-based models often struggle with inefficiency, they excel in capturing a wide range of potential docking poses, leading to better overall performance. To bridge the gap between sampling and regression docking models, we incorporate a simple yet effective sampling technique coupled with a lightweight confidence model, transforming the regression-based FABind+ into a sampling version without requiring additional training. This involves the introduction of pocket clustering to capture multiple binding sites and dropout sampling for various conformations. The combination of a classification loss and a ranking loss enables the lightweight confidence model to select the most accurate prediction. Experimental results and analysis demonstrate that FABind+ (both the regression and sampling versions) not only significantly outperforms the original FABind, but also achieves competitive state-of-the-art performance. Our code is available at https://github.com/QizhiPei/FABind. Kaiyuan Gao, Qizhi Pei, Jinhua Zhu 0001, Kun He 0001, Lijun Wu 0003 |
KDD (1) | 2 |
| 2025 | Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model ReasoningabstractReasoning capability is pivotal for Large Language Models (LLMs) to solve complex tasks, yet achieving reliable and scalable reasoning remains challenging. While Chain-of-Thought (CoT) prompting has become a mainstream approach, existing methods often suffer from uncontrolled generation, insufficient quality, and limited diversity in reasoning paths.
Recent efforts leverage code to enhance CoT by grounding reasoning in executable steps, but such methods are typically constrained to predefined mathematical problems, hindering scalability and generalizability.
In this work, we propose \texttt{Caco} (Code-Assisted Chain-of-ThOught), a novel framework that automates the synthesis of high-quality, verifiable, and diverse instruction-CoT reasoning data through code-driven augmentation. Unlike prior work, \texttt{Caco} first fine-tunes a code-based CoT generator on existing math and programming solutions in a unified code format, then scales the data generation to a large amount of diverse reasoning traces. Crucially, we introduce automated validation via code execution and rule-based filtering to ensure logical correctness and structural diversity, followed by reverse-engineering filtered outputs into natural language instructions and language CoTs to enrich task adaptability. This closed-loop process enables fully automated, scalable synthesis of reasoning data with guaranteed executability.
Experiments on our created \texttt{Caco}-1.3M dataset demonstrate that \texttt{Caco}-trained models achieve strong competitive performance on mathematical reasoning benchmarks, outperforming existing strong baselines. Further analysis reveals that \texttt{Caco}’s code-anchored verification and instruction diversity contribute to superior generalization across unseen tasks. Our work establishes a paradigm for building self-sustaining, trustworthy reasoning systems without human intervention. Honglin Lin, Qizhi Pei, Zhuoshi Pan, Yu Li 0006, Xin Gao 0001, Juntao Li 0005, Conghui He, Lijun Wu 0003 |
NeurIPS | 2 |
| 2024 | Exploiting Pre-trained Models for Drug Target Affinity Prediction with Nearest NeighborsabstractDrug-Target binding Affinity (DTA) prediction is essential for drug discovery. Despite the application of deep learning methods to DTA prediction, the achieved accuracy remain suboptimal. In this work, inspired by the recent success of retrieval methods, we propose kNN-DTA, a non-parametric embedding-based retrieval method adopted on a pre-trained DTA prediction model, which can extend the power of the DTA model with no or negligible cost. Different from existing methods, we introduce two neighbor aggregation ways from both embedding space and label space that are integrated into a unified framework. Specifically, we propose a label aggregation with pair-wise retrieval and a representation aggregation with point-wise retrieval of the nearest neighbors. This method executes in the inference phase and can efficiently boost the DTA prediction performance with no training cost. In addition, we propose an extension, Ada-kNN-DTA, an instance-wise and adaptive aggregation with lightweight learning. Results on four benchmark datasets show that kNN-DTA brings significant improvements, outperforming previous state-of-the-art (SOTA) results, e.g, on BindingDB IC50 and Ki testbeds, kNN-DTA obtains new records of RMSE 0.684 and 0.750 . The extended Ada-kNN-DTA further improves the performance to be 0.675 and 0.735 RMSE. These results strongly prove the effectiveness of our method. Results in other settings and comprehensive studies/analyses also show the great potential of our kNN-DTA approach. Qizhi Pei, Lijun Wu 0003, Zhenyu He 0012, Jinhua Zhu 0001, Yingce Xia, Shufang Xie 0003, Rui Yan 0001 |
CIKM | 1 |
| 2023 | BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language AssociationsabstractRecent advancements in biological research leverage the integration of molecules, proteins, and natural language to enhance drug discovery.However, current models exhibit several limitations, such as the generation of invalid molecular SMILES, underutilization of contextual information, and equal treatment of structured and unstructured knowledge.To address these issues, we propose BioT5, a comprehensive pretraining framework that enriches cross-modal integration in biology with chemical knowledge and natural language associations.BioT5 utilizes SELFIES for 100% robust molecular representations and extracts knowledge from the surrounding context of bio-entities in unstructured biological literature.Furthermore, BioT5 distinguishes between structured and unstructured knowledge, leading to more effective utilization of information.After fine-tuning, BioT5 shows superior performance across a wide range of tasks, demonstrating its strong capability of capturing underlying relations and properties of bio-entities.Our code is available at https://github.com/QizhiPei/BioT5. Qizhi Pei, Wei Zhang 0399, Jinhua Zhu 0001, Kehan Wu, Kaiyuan Gao, Lijun Wu 0003, Yingce Xia, Rui Yan 0001 |
EMNLP | 1 |
| 2023 | FABind: Fast and Accurate Protein-Ligand BindingabstractModeling the interaction between proteins and ligands and accurately predicting their binding structures is a critical yet challenging task in drug discovery. Recent advancements in deep learning have shown promise in addressing this challenge, with sampling-based and regression-based methods emerging as two prominent approaches. However, these methods have notable limitations. Sampling-based methods often suffer from low efficiency due to the need for generating multiple candidate structures for selection. On the other hand, regression-based methods offer fast predictions but may experience decreased accuracy. Additionally, the variation in protein sizes often requires external modules for selecting suitable binding pockets, further impacting efficiency. In this work, we propose FABind, an end-to-end model that combines pocket prediction and docking to achieve accurate and fast protein-ligand binding. FABind incorporates a unique ligand-informed pocket prediction module, which is also leveraged for docking pose estimation. The model further enhances the docking process by incrementally integrating the predicted pocket to optimize protein-ligand binding, reducing discrepancies between training and inference. Through extensive experiments on benchmark datasets, our proposed FABind demonstrates strong advantages in terms of effectiveness and efficiency compared to existing methods. Our code is available at https://github.com/QizhiPei/FABind. Qizhi Pei, Kaiyuan Gao, Lijun Wu 0003, Jinhua Zhu 0001, Yingce Xia, Shufang Xie 0003, Tao Qin 0001, Kun He 0001, Tie-Yan Liu, Rui Yan 0001 |
NeurIPS | 1 |
| 2023 | Breaking the barriers of data scarcity in drug-target affinity predictionabstractAccurate prediction of drug-target affinity (DTA) is of vital importance in early-stage drug discovery, facilitating the identification of drugs that can effectively interact with specific targets and regulate their activities. While wet experiments remain the most reliable method, they are time-consuming and resource-intensive, resulting in limited data availability that poses challenges for deep learning approaches. Existing methods have primarily focused on developing techniques based on the available DTA data, without adequately addressing the data scarcity issue. To overcome this challenge, we present the Semi-Supervised Multi-task training (SSM) framework for DTA prediction, which incorporates three simple yet highly effective strategies: (1) A multi-task training approach that combines DTA prediction with masked language modeling using paired drug-target data. (2) A semi-supervised training method that leverages large-scale unpaired molecules and proteins to enhance drug and target representations. This approach differs from previous methods that only employed molecules or proteins in pre-training. (3) The integration of a lightweight cross-attention module to improve the interaction between drugs and targets, further enhancing prediction accuracy. Through extensive experiments on benchmark datasets such as BindingDB, DAVIS and KIBA, we demonstrate the superior performance of our framework. Additionally, we conduct case studies on specific drug-target binding activities, virtual screening experiments, drug feature visualizations and real-world applications, all of which showcase the significant potential of our work. In conclusion, our proposed SSM-DTA framework addresses the data limitation challenge in DTA prediction and yields promising results, paving the way for more efficient and accurate drug discovery processes. Qizhi Pei, Lijun Wu 0003, Jinhua Zhu 0001, Yingce Xia, Shufang Xie 0003, Tao Qin 0001, Haiguang Liu, Tie-Yan Liu, Rui Yan 0001 |
Briefings Bioinform. | 1 |