EDBT 2026 Demo / reviewers in the wild / expert
Chi Zhang 0102
dblp:91/195-102
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0000-0323-0715ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spectral Characterization and Mitigation of Sequential Knowledge Editing CollapseabstractChi Zhang, Mengqi Zhang, Xiaotian Ye, Runxi Cheng, Zisheng Zhou, Ying Zhou, Pengjie Ren, Zhumin Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chi Zhang 0102, Mengqi Zhang 0002, Xiaotian Ye, Runxi Cheng, Zisheng Zhou, Pengjie Ren, Zhumin Chen |
ACL (1) | 1 |
| 2026 | Mixture-of-RAG: Integrating Text and Tables with Large Language ModelsabstractLarge language models (LLMs) achieve optimal utility when their responses are grounded in external knowledge sources. However, real-world documents, such as annual reports, scientific papers, and clinical guidelines, frequently combine extensive narrative content with complex, hierarchically structured tables. While existing retrieval-augmented generation (RAG) systems effectively integrate LLMs' generative capabilities with external retrieval-based information, their performance significantly deteriorates especially processing such heterogeneous text-table hierarchies. To address this limitation, we formalize the task of Heterogeneous Document RAG, which requires joint retrieval and reasoning across textual and hierarchical tabular data. We propose MixRAG, a novel three-stage framework: (i) hierarchy row-and-column-level (H-RCL) representation that preserves hierarchical structure and heterogeneous relationship, (ii) an ensemble retriever with LLM-based reranking for evidence alignment, and (iii) multi-step reasoning decomposition via a RECAP prompt strategy. To bridge the gap in available data for this domain, we release the dataset DocRAGLib, a 2k-document corpus paired with automatically aligned text-table summaries and gold document annotations. The comprehensive experiment results demonstrate that MixRAG boosts top-1 retrieval by 46% over strong text-only, table-only, and naive-mixture baselines, establishing new state-of-the-art performance for mixed-modality document grounding. Chi Zhang 0102, Mengqi Zhang 0002 |
KDD (1) | 1 |
| 2025 | PFCA: Efficient Path Filtering with Causal Analysis for Healthcare Risk PredictionabstractElectronic health records (EHRs) store patient medical history in the structured data format, which facilitates automatic healthcare risk prediction, thereby improving personalized healthcare management and treatment. There are two main categories of methods for automatic healthcare risk prediction. The first models time-series information or relationships between visits for enhanced patient representations. However, given the high dimensionality nature of the EHR data, it often obtains compromise results due to the lack of training data. The second exploits external knowledge, e.g., knowledge graphs (KGs), to augment the training data, but less attention has been paid to distinguishing the importance of features and filtering out irrelevant external knowledge, leading to overwhelming noise and inefficiency. Additionally, the joint relationships between patient features were not emphasized, which are highlighted in clinical practice. In this paper, we propose an efficient Path Filtering with Causal Analysis (PFCA) approach for enhanced healthcare risk prediction to address these challenges. PFCA first extracts personalized knowledge graphs (PKGs) consisting of paths linking the patient's features to targets and then devises a fine-grained filtering method based on path messages to remove irrelevant paths for better efficiency. Then we develop an effective similarity-based method to model different features' joint interactions with targets to learn augmented representations for each feature. Furthermore, we design a causal analysis method that includes a novel causal intervention mechanism to mine and prioritize causal features for improved predictive performance. Finally, by exploiting the attention weights of paths in the PKGs, PFCA provides target-oriented interpretations, showing how patients' features lead to targets through significant paths. Experimental results on three public real-world datasets and four healthcare risk prediction tasks confirm PFCA's effectiveness in improving predictive performance compared to ten state-of-the-art baselines, demonstrate its efficiency of path filtering and interpretability. Jiyun Shi, Haochen Xu, Chi Zhang 0102, Zhaojing Luo, Meihui Zhang 0001 |
ICDE | 5 |
| 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language ModelsabstractData selection is of great significance in pretraining large language models, given the variation in quality within the large-scale available training corpora.
To achieve this, researchers are currently investigating the use of data influence to measure the importance of data instances, $i.e.,$ a high influence score indicates that incorporating this instance to the training set is likely to enhance the model performance. Consequently, they select the top-$k$ instances with the highest scores. However, this approach has several limitations.
(1) Calculating the accurate influence of all available data is time-consuming.
(2) The selected data instances are not diverse enough, which may hinder the pretrained model's ability to generalize effectively to various downstream tasks.
In this paper, we introduce $\texttt{Quad}$, a data selection approach that considers both quality and diversity by using data influence to achieve state-of-the-art pretraining results.
To compute the influence ($i.e.,$ the quality) more accurately and efficiently, we incorporate the attention layers to capture more semantic details, which can be accelerated through the Kronecker product.
For the diversity, $\texttt{Quad}$ clusters the dataset into similar data instances within each cluster and diverse instances across different clusters. For each cluster, if we opt to select data from it, we take some samples to evaluate the influence to prevent processing all instances. Overall, we favor clusters with highly influential instances (ensuring high quality) or clusters that have been selected less frequently (ensuring diversity), thereby well balancing between quality and diversity. Experiments on Slimpajama and FineWeb over 7B large language models demonstrate that $\texttt{Quad}$ significantly outperforms other data selection methods with a low FLOPs consumption. Further analysis also validates the effectiveness of our influence calculation. Chi Zhang 0102, Huaping Zhong, Chengliang Chai, Rui Wang 0119, Xinlin Zhuang, Tianyi Bai, Jiantao Qiu, Lei Cao 0004, Ju Fan, Ye Yuan 0001, Guoren Wang, Conghui He |
ICLR | 1 |
| 2025 | Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic OptimizationabstractRecent studies indicate that deep neural networks degrade in generalization performance under noisy supervision. Existing methods focus on isolating clean subsets or correcting noisy labels, facing limitations such as high computational costs, heavy hyperparameter tuning process, and coarse-grained optimization. To address these challenges, we propose a novel two-stage noisy learning framework that enables instance-level optimization through a dynamically weighted loss function, avoiding hyperparameter tuning. To obtain stable and accurate information about noise modeling, we introduce a simple yet effective metric, termed $\textit{wrong event}$, which dynamically models the cleanliness and difficulty of individual samples while maintaining computational costs. Our framework first collects $\textit{wrong event}$ information and builds a strong base model. Then we perform noise-robust training on the base model, using a probabilistic model to handle the $\textit{wrong event}$ information of samples. Experiments on six synthetic and real-world LNL benchmarks demonstrate our method surpasses state-of-the-art methods in performance, achieves a nearly 75\% reduction in storage and computational time, strongly improving model scalability. Our code is available at https://github.com/iTheresaApocalypse/IDO. Chengliang Chai, Jingzhe Xu, Chi Zhang 0102, Ye Yuan 0001, Guoren Wang, Lei Cao 0004 |
NeurIPS | 4 |
| 2025 | AixelAsk: A Stepwise-Guided Retrieval and Reasoning Framework for Large Table QAabstractIn the big data era, Table Question Answering (Table QA) has emerged as a crucial tool for extracting insights from structured data, especially in large table scenarios. There are two main categories of methods for Table QA: Executable Code-driven methods and Language Model based (LM-based) methods. Code-driven methods, e.g. Text-to-SQL based solutions, often struggle with incomplete or mismatching schema information. LM-based methods, include Pre-trained Language Models (PLMs) and Large Language Models (LLMs), also face challenges as PLMs have limited generalization, while LLMs suffer from performance degradation and increased token cost when applied to large tables. To address these challenges, we propose AixelAsk, a novel LLM-based framework designed for Large Table QA. Specifically, AixelAsk incorporates a three-module architecture consisting of Decomposition module, Retrieval module and Reasoning module. The Decomposition module constructs a directed acyclic graph (DAG)-based solution plan by decomposing the question into execution nodes with explicit dependencies, making a clear reasoning path to guide the LLM through a logical process. Inspired by the Retrieval-Augmented Generation, the Retrieval Module extracts key rows and columns from the large table, reducing input token size and focusing on critical information. The Reasoning Module performs step-by-step inferences over the retrieved sub-tables, guided by each execution node in the solution plan, to generate final answer. By tackling the challenges of LLM performance degradation with large inputs and complex questions, AixelAsk achieves superior performance in Large Table QA. Extensive experiments on various baselines across three datasets demonstrate the effectiveness and efficiency of our proposed AixelAsk framework. AixelAsk outperforms the state-of-the-art baseline by 4% - 8% in the exact match score, and at the same time reduces token usage by 86.4%, achieving both high accuracy and cost efficiency in the Large Table QA task. Chi Zhang 0102, Meihui Zhang 0001, Yuxin Yang 0013, Zhaojing Luo |
Proc. ACM Manag. Data | 1 |
| 2024 | Cost-Effective Framework with Optimized Task Decomposition and Batch Prompting for Medical Dialogue SummaryabstractThe generation of medical dialogue notes is essential in healthcare, providing a structured recapitalization of patient-provider interactions. Medical notes are rigorously organized into various sections, including Chief Complaint, History of Present Illness and more. Each section serves a specific purpose to record detailed medical content. Traditionally, this task is labor-intensive, requiring physicians to manually create notes, a process prone to errors. With advancements in AI, it is now feasible to automate the generation of medical notes. There are mainly two categories of methods for automatic medical note generation. Pre-trained language models (PLMs) struggle with unstructured outputs, limited datasets, and inadequate medical terminology. In-context learning (ICL) methods improve accuracy and reduce data requirements but still produce unstructured notes and require high time and cost. To tackle the above challenges, we propose a three-module framework, called CE-DEPT, for accurate, efficient and cost-effective medical note generation. Specifically, the Task Decomposition Module breaks down complete medical dialogues into section-specific dialogues to ensure relevance and accuracy. The Batch Combination Module groups these sections into batches based on disease similarity to reduce costs and improve efficiency. The Note Generation Module employs batch prompting with ICL to generate each section note, followed by combining them into a structured, comprehensive medical note. Experiments on benchmark datasets demonstrated the effectiveness of Task Decomposition and Batch Prompting. Our method, CE-DEPT outperforms the best method by 5% on the ROUGE-1 score, 3% on the Bertscore-F1, a cost-effectiveness improvement of 15%, and a reduction in time consumption of 25% at peak accuracy. Chi Zhang 0102, Jiehao Chen, Jiyun Shi, Zhaojing Luo, Meihui Zhang 0001 |
CIKM | 1 |
| 2024 | KEIM: Knowledge Graph Empowered Interpretable Model for Diagnosis Prediction
Zhaojing Luo, Chi Zhang 0102, Jiyun Shi, Meihui Zhang 0001 |
DASFAA (4) | 2 |
| 2024 | DMRNet: Effective Network for Accurate Discharge Medication RecommendationabstractElectronic Health Records, which contain abundant structured data information of the patients, can help clinicians and data scientists address complex medical issues, particularly medication recommendation. The recommendation of medications is crucial for accurate and timely prescriptions. It is a nuanced task that entails analyzing various sources of healthcare data. Traditional medication recommendation is performed manually, which is labor-intensive and error-prone. The development of Electronic Health Records enables automatic medication recommendation. There are mainly two categories of methods for automatic medication recommendation. The first category uses the patients' current visit information and the drug-drug interactions (DDI). For these methods, both the comprehensive patient's medical history and the significant medication-diagnosis knowledge are not exploited appropriately. The second category utilizes longitudinal patient data, but different history visits are incorporated indiscriminately. Furthermore, in clinical practice, the associations between historical medications and future prescriptions are highlighted. However, they are less emphasized in current methods. Nevertheless, this is less emphasized by current automatic medication recommendation methods. To tackle the above challenges, we propose a three-module Discharge Medication Recommendation Network, called DMRNet, for accurate discharge medication recommendations. Specifically, the Information Integration Module combines information from the current visit and significant external knowledge e.g., the Diagnosis-Medication Co-occurrence (DMC) relationship. The Medication Retention Module is specially designed to capture the associations between the historical medications and the recommended medications. The History Retrieval Module differentiates the significance of different historical visits and incorporates them based on different significance values. Experimental evaluations on benchmark datasets, i.e., MIMIC-III and MIMIC-IV, confirm DMRNet's superiority over state-of-the-art baseline methods in terms of Jaccard Similarity, F1-score, Precision and Recall. Jiyun Shi, Yuqiao Wang, Chi Zhang 0102, Zhaojing Luo, Chengliang Chai, Meihui Zhang 0001 |
ICDE | 3 |