EDBT 2026 Demo / reviewers in the wild / expert
Wei Chen 0088
dblp:181/2832-88
· DBLP profile ↗
23ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0001-9431-9247ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoLink: Autonomous Schema Exploration and Expansion for Scalable Schema Linking in Text-to-SQL at ScaleabstractFor industrial-scale text-to-SQL, supplying the entire database schema to Large Language Models (LLMs) is impractical due to context window limits and irrelevant noise. Schema linking, which filters the schema to a relevant subset, is therefore critical. However, existing methods incur prohibitive costs, struggle to trade off recall and noise, and scale poorly to large databases. We present AutoLink, an autonomous agent framework that reformulates schema linking as an iterative, agent-driven process. Guided by an LLM, AutoLink dynamically explores and expands the linked schema subset, progressively identifying necessary schema components without inputting the full database schema. Our experiments demonstrate AutoLink's superior performance, achieving state-of-the-art strict schema linking recall of 97.4% on Bird-Dev and 91.2% on Spider 2.0-Lite, with competitive execution accuracy, i.e., 68.7% EX on Bird-Dev (better than CHESS) and 34.9% EX on Spider 2.0-Lite (ranking 2nd on the official leaderboard). Crucially, AutoLink exhibits exceptional scalability, maintaining high recall, efficient token consumption, and robust execution accuracy on large schemas (e.g., over 3,000 columns) where existing methods severely degrade—making it a highly scalable, high-recall schema-linking solution for industrial text-to-SQL systems. Yuanlei Zheng, Zhenbiao Cao, Xiaojin Zhang 0002, Zhongyu Wei, Pei Fu, Zhenbo Luo, Wei Chen 0088, Xiang Bai |
AAAI | 8 |
| 2026 | SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical ConsultationabstractMedical consultations are intrinsically speechcentric.However, most prior works focus on long-text-based interactions, which are cumbersome and patient-unfriendly.Recent advances in speech language models (SpeechLMs) have enabled more natural speech-based interaction, yet the scarcity of medical speech data and the inefficiency of directly fine-tuning on speech data jointly hinder the adoption of SpeechLMs in medical consultation.In this paper, we propose SpeechMedAssist, a SpeechLM natively capable of conducting speech-based multi-turn interactions with patients.By exploiting the architectural properties of SpeechLMs, we decouple the conventional one-stage training into a two-stage paradigm consisting of (1) Knowledge & Capability Injection via Text and (2) Modality Re-alignment with Limited Speech Data, thereby reducing the requirement for medical speech data to only 10k synthesized samples.To evaluate SpeechLMs for medical consultation scenarios, we design a benchmark comprising both single-turn question answering and multi-turn simulated interactions.Experimental results show that our model outperforms all baselines in both effectiveness and robustness in most evaluation settings. Sirry Chen, Jieyi Wang, Wei Chen 0088, Zhongyu Wei |
ACL (1) | 3 |
| 2026 | Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic EnvironmentsabstractZheng Jia, Shengbin Yue, Wei Chen, Siyuan Wang, Yidong Liu, Zejun Li, Yun Song, Zhongyu Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zheng Jia, Shengbin Yue, Wei Chen 0088, Siyuan Wang 0025, Yidong Liu, Yun Song, Zhongyu Wei |
ACL (1) | 3 |
| 2026 | Doc-V^*: Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQAabstractYuanlei Zheng, Pei Fu, Hang Li, Ziyang Wang, Yuyi Zhang, Wenyu Ruan, Xiaojin Zhang, Zhongyu Wei, Zhenbo Luo, Jian Luan, Wei Chen, Xiang Bai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuanlei Zheng, Pei Fu, Hang Li 0001, Wenyu Ruan, Xiaojin Zhang 0002, Zhongyu Wei, Zhenbo Luo, Jian Luan 0001, Wei Chen 0088, Xiang Bai |
ACL (1) | 11 |
| 2026 | PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
Xudong Xie, Minghui Liao, Wei Chen 0088, Xiang Bai |
Int. J. Comput. Vis. | 8 |
| 2026 | InsQABench: Benchmarking Chinese insurance domain question answering with large language modelsabstractWe present InsQABench-the first comprehensive benchmark for evaluating LLMs’ capabilities in Chinese insurance QA. InsQABench comprises 95K carefully curated QA pairs derived from real-world insurance documents, covering 3 distinct tasks, 44 question types, and 55 specialized insurance topics. Our experiments evaluated and reported the performance of mainstream LLMs under both fine-tuned and zero-shot settings, demonstrating that fine-tuning on InsQABench can significantly improve model performance. We also introduced two frameworks that further enhanced task-specific performance, achieving 4.91% and 5.11% enhancement in accuracy over the next best-performing model. Binbin Lin 0002, Jiarui Cai, Xiaojin Zhang 0002, Zhongyu Wei, Wei Chen 0088 |
Inf. Process. Manag. | 10 |
| 2025 | FedAA: A Reinforcement Learning Perspective on Adaptive Aggregation for Fair and Robust Federated LearningabstractFederated Learning (FL) has emerged as a promising approach for privacy-preserving model training across decentralized devices. However, it faces challenges such as statistical heterogeneity and susceptibility to adversarial attacks, which can impact model robustness and fairness. Personalized FL attempts to provide some relief by customizing models for individual clients. However, it falls short in addressing server-side aggregation vulnerabilities. We introduce a novel method called FedAA, which optimizes client contributions via Adaptive Aggregation to enhance model robustness against malicious clients and ensure fairness across participants in non-identically distributed settings. To achieve this goal, we propose an approach involving a Deep Deterministic Policy Gradient-based algorithm for continuous control of aggregation weights, an innovative client selection method based on model parameter distances, and a reward mechanism guided by validation set performance. Empirically, extensive experiments demonstrate that, in terms of robustness, FedAA outperforms the state-of-the-art methods, while maintaining comparable levels of fairness, offering a promising solution to build resilient and fair federated systems. Jialuo He, Wei Chen 0088, Xiaojin Zhang 0002 |
AAAI | 2 |
| 2025 | Synergistic Multi-Agent Framework with Trajectory Learning for Knowledge-Intensive TasksabstractRecent advancements in Large Language Models (LLMs) have led to significant breakthroughs in various natural language processing tasks. However, generating factually consistent responses in knowledge-intensive scenarios remains a challenge due to issues such as hallucination, difficulty in acquiring long-tailed knowledge, and limited memory expansion. This paper introduces SMART, a novel multi-agent framework that leverages external knowledge to enhance the interpretability and factual consistency of LLM-generated responses. SMART comprises four specialized agents, each performing a specific sub-trajectory action to navigate complex knowledge-intensive tasks. We propose a multi-agent co-training paradigm, Long-Short Trajectory Learning, which ensures synergistic collaboration among agents while maintaining fine-grained execution by each agent. Extensive experiments on five knowledge-intensive tasks demonstrate SMART's superior performance compared to widely adopted knowledge internalization and knowledge enhancement methods. Our framework can extend beyond knowledge-intensive tasks to more complex scenarios. Shengbin Yue, Siyuan Wang 0025, Wei Chen 0088, Xuanjing Huang 0001, Zhongyu Wei |
AAAI | 3 |
| 2025 | AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction SimulatorabstractArtificial intelligence has significantly revolutionized healthcare, particularly through large language models (LLMs) that demonstrate superior performance in static medical question answering benchmarks. However, evaluating the potential of LLMs for real-world clinical applications remains challenging due to the intricate nature of doctor-patient interactions. To address this, we introduce AI Hospital, a multi-agent framework emulating dynamic medical interactions between Doctor as player and NPCs including Patient and Examiner. This setup allows for more practical assessments of LLMs in simulated clinical scenarios. We develop the Multi-View Medical Evaluation (MVME) benchmark, utilizing high-quality Chinese medical records and multiple evaluation strategies to quantify the performance of LLM-driven Doctor agents on symptom collection, examination recommendations, and diagnoses. Additionally, a dispute resolution collaborative mechanism is proposed to enhance medical interaction capabilities through iterative discussions. Despite improvements, current LLMs (including GPT-4) still exhibit significant performance gaps in multi-turn interactive scenarios compared to non-interactive scenarios. Our findings highlight the need for further research to bridge these gaps and improve LLMs’ clinical decision-making capabilities. Our data, code, and experimental results are all open-sourced at https://github.com/LibertFan/AI_Hospital. Zhihao Fan, Jialong Tang, Wei Chen 0088, Siyuan Wang 0025, Zhongyu Wei, Fei Huang 0002 |
COLING | 4 |
| 2025 | Do Current Video LLMs Have Strong OCR Abilities? A Preliminary StudyabstractWith the rise of multi-modal large language models, accurately extracting and understanding textual information from video content—referred to as video-based optical character recognition (Video OCR)—has become a crucial capability. This paper introduces a novel benchmark designed to evaluate the video OCR performance of multi-modal models in videos. Comprising 1,028 videos and 2,961 question-answer pairs, this benchmark proposes several key challenges through 6 distinct sub-tasks: (1) Recognition of text content itself and its basic visual attributes, (2) Semantic and Spatial Comprehension of OCR objects in videos (3) Dynamic Motion detection and Temporal Localization. We developed this benchmark using a semi-automated approach that integrates the OCR ability of image LLMs with manual refinement, balancing efficiency, cost, and data quality. Our resource aims to help advance research in video LLMs and underscores the need for improving OCR ability for video LLMs. The benchmark will be released on https://github.com/YuHuiGao/FG-Bench.git. Yulin Fei, Yuhui Gao, Xingyuan Xian, Xiaojin Zhang 0002, Wei Chen 0088 |
COLING | 6 |
| 2025 | OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and ReasoningabstractScoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their abilities in certain challenging tasks, such as text localization, handwritten content extraction, and logical reasoning, remain underexplored. To bridge this gap, we introduce OCRBench v2, a large-scale bilingual text-centric benchmark with currently the most comprehensive set of tasks ($4\times$ more tasks than the previous multi-scene benchmark OCRBench), the widest coverage of scenarios ($31$ diverse scenarios), and thorough evaluation metrics, with $10,000$ human-verified question-answering pairs and a high proportion of difficult samples. Moreover, we construct a private test set with $1,500$ manually annotated images. The consistent evaluation trends observed across both public and private test sets validate the OCRBench v2's reliability. After carefully benchmarking state-of-the-art LMMs, we find that most LMMs score below $50$ ($100$ in total) and suffer from five-type limitations, including less frequently encountered text recognition, fine-grained perception, layout perception, complex element parsing, and logical reasoning. The benchmark and evaluation scripts are available at https://github.com/Yuliang-Liu/MultimodalOCR. Zhebin Kuang, Jiajun Song, Mingxin Huang, Linghao Zhu, Qidi Luo, Xinyu Wang 0010, Hao Lu 0003, Guozhi Tang, Bin Shan, Chunhui Lin, Binghong Wu, Hao Feng 0009, Hao Liu 0003, Can Huang 0002, Jingqun Tang, Wei Chen 0088, Xiang Bai |
NeurIPS | 21 |
| 2025 | FinTeam: A Multi-agent Collaborative Intelligence System for Comprehensive Financial Scenarios
Yingqian Wu, Zefei Long, Rong Ye, Zhongtian Lu, Xianyin Zhang, Wei Chen 0088, Zhongyu Wei |
NLPCC (2) | 8 |
| 2025 | No free lunch theorem for privacy-preserving LLM inferenceabstractIndividuals and businesses have been significantly benefited by Large Language Models (LLMs) including PaLM, Gemini and ChatGPT in various ways. For example, LLMs enhance productivity, reduce costs, and enable us to focus on more valuable tasks. Furthermore, LLMs possess the capacity to sift through extensive datasets, uncover underlying patterns, and furnish critical insights that propel the frontiers of technology and science. However, LLMs also pose privacy concerns. Users' interactions with LLMs may expose their sensitive personal or company information. A lack of robust privacy safeguards and legal frameworks could permit the unwarranted intrusion or improper handling of individual data, thereby risking infringements of privacy and the theft of personal identities. To ensure privacy, it is essential to minimize the dependency between shared prompts and private information. Various randomization approaches have been proposed to protect prompts' privacy, but they may incur utility loss compared to unprotected LLMs prompting. Therefore, it is essential to evaluate the balance between the risk of privacy leakage and loss of utility when conducting effective protection mechanisms. The current study develops a framework for inferring privacy-protected Large Language Models (LLMs) and lays down a solid theoretical basis for examining the interplay between privacy preservation and utility. The core insight is encapsulated within a theorem that is called as the NFL (abbreviation of the word No-Free-Lunch) Theorem. Xiaojin Zhang 0002, Yahao Pang, Yan Kang 0001, Wei Chen 0088, Lixin Fan, Hai Jin 0001, Qiang Yang 0001 |
Artif. Intell. | 4 |
| 2025 | Zero-shot text-to-parameter realtime translation for game character auto-creation and identity consistency editing
Weijun Cao, Wei Chen 0088, Haidi Fan, Gen Dong |
Neurocomputing | 5 |
| 2024 | LawLLM: Intelligent Legal System with Legal Reasoning and Verifiable Retrieval
Shengbin Yue, Shujun Liu, Chenchen Shen, Siyuan Wang 0025, Yun Song, Wei Chen 0088, Xuanjing Huang 0001, Zhongyu Wei |
DASFAA (5) | 10 |
| 2023 | A benchmark for automatic medical consultation system: frameworks, tasks and datasetsabstractMOTIVATION: In recent years, interest has arisen in using machine learning to improve the efficiency of automatic medical consultation and enhance patient experience. In this article, we propose two frameworks to support automatic medical consultation, namely doctor-patient dialogue understanding and task-oriented interaction. We create a new large medical dialogue dataset with multi-level fine-grained annotations and establish five independent tasks, including named entity recognition, dialogue act classification, symptom label inference, medical report generation and diagnosis-oriented dialogue policy. RESULTS: We report a set of benchmark results for each task, which shows the usability of the dataset and sets a baseline for future studies. AVAILABILITY AND IMPLEMENTATION: Both code and data are available from https://github.com/lemuria-wchen/imcs21. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wei Chen 0088, Hongyi Fang, Qianyuan Yao, Jianye Hao, Qi Zhang 0001, Xuanjing Huang 0001, Jiajie Peng, Zhongyu Wei |
Bioinform. | 1 |
| 2023 | DxFormer: a decoupled automatic diagnostic system based on decoder-encoder transformer with dense symptom representationsabstractMOTIVATION: Symptom-based automatic diagnostic system queries the patient's potential symptoms through continuous interaction with the patient and makes predictions about possible diseases. A few studies use reinforcement learning (RL) to learn the optimal policy from the joint action space of symptoms and diseases. However, existing RL (or Non-RL) methods focus on disease diagnosis while ignoring the importance of symptom inquiry. Although these systems have achieved considerable diagnostic accuracy, they are still far below its performance upper bound due to few turns of interaction with patients and insufficient performance of symptom inquiry. To address this problem, we propose a new automatic diagnostic framework called DxFormer, which decouples symptom inquiry and disease diagnosis, so that these two modules can be independently optimized. The transition from symptom inquiry to disease diagnosis is parametrically determined by the stopping criteria. In DxFormer, we treat each symptom as a token, and formalize the symptom inquiry and disease diagnosis to a language generation model and a sequence classification model, respectively. We use the inverted version of Transformer, i.e. the decoder-encoder structure, to learn the representation of symptoms by jointly optimizing the reinforce reward and cross-entropy loss. RESULTS: We conduct experiments on three real-world medical dialogue datasets, and the experimental results verify the feasibility of increasing diagnostic accuracy by improving symptom recall. Our model overcomes the shortcomings of previous RL-based methods. By decoupling symptom query from the process of diagnosis, DxFormer greatly improves the symptom recall and achieves the state-of-the-art diagnostic accuracy. AVAILABILITY AND IMPLEMENTATION: Both code and data are available at https://github.com/lemuria-wchen/DxFormer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wei Chen 0088, Jiajie Peng, Zhongyu Wei |
Bioinform. | 1 |
| 2022 | DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response GenerationabstractWei Chen, Yeyun Gong, Song Wang, Bolun Yao, Weizhen Qi, Zhongyu Wei, Xiaowu Hu, Bartuer Zhou, Yi Mao, Weizhu Chen, Biao Cheng, Nan Duan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Wei Chen 0088, Yeyun Gong, Song Wang 0012, Bolun Yao, Weizhen Qi, Zhongyu Wei, Xiaowu Hu, Bartuer Zhou, Weizhu Chen, Biao Cheng, Nan Duan 0001 |
ACL (1) | 1 |
| 2022 | Contextual Fine-to-Coarse Distillation for Coarse-grained Response Selection in Open-Domain ConversationsabstractWei Chen, Yeyun Gong, Can Xu, Huang Hu, Bolun Yao, Zhongyu Wei, Zhihao Fan, Xiaowu Hu, Bartuer Zhou, Biao Cheng, Daxin Jiang, Nan Duan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Wei Chen 0088, Yeyun Gong, Can Xu 0002, Huang Hu, Bolun Yao, Zhongyu Wei, Zhihao Fan, Xiaowu Hu, Bartuer Zhou, Biao Cheng, Daxin Jiang, Nan Duan 0001 |
ACL (1) | 1 |
| 2022 | A Structure-Aware Argument Encoder for Literature Discourse AnalysisabstractExisting research for argument representation learning mainly treats tokens in the sentence equally and ignores the implied structure information of argumentative context. In this paper, we propose to separate tokens into two groups, namely framing tokens and topic ones, to capture structural information of arguments. In addition, we consider high-level structure by incorporating paragraph-level position information. A novel structure-aware argument encoder is proposed for literature discourse analysis. Experimental results on both a self-constructed corpus and a public corpus show the effectiveness of our model. Resources are available at https://github.com/lemuria-wchen/SAE. Yinzi Li, Wei Chen 0088, Zhongyu Wei, Yujun Huang, Chujun Wang, Siyuan Wang 0025, Qi Zhang 0001, Xuanjing Huang 0001, Libo Wu |
COLING | 2 |
| 2022 | Hierarchical reinforcement learning for automatic disease diagnosisabstractMOTIVATION: Disease diagnosis-oriented dialog system models the interactive consultation procedure as the Markov decision process, and reinforcement learning algorithms are used to solve the problem. Existing approaches usually employ a flat policy structure that treat all symptoms and diseases equally for action making. This strategy works well in a simple scenario when the action space is small; however, its efficiency will be challenged in the real environment. Inspired by the offline consultation process, we propose to integrate a hierarchical policy structure of two levels into the dialog system for policy learning. The high-level policy consists of a master model that is responsible for triggering a low-level model, the low-level policy consists of several symptom checkers and a disease classifier. The proposed policy structure is capable to deal with diagnosis problem including large number of diseases and symptoms. RESULTS: Experimental results on three real-world datasets and a synthetic dataset demonstrate that our hierarchical framework achieves higher accuracy and symptom recall in disease diagnosis compared with existing systems. We construct a benchmark including datasets and implementation of existing algorithms to encourage follow-up researches. AVAILABILITY AND IMPLEMENTATION: The code and data are available from https://github.com/FudanDISC/DISCOpen-MedBox-DialoDiagnosis. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kangenbei Liao, Wei Chen 0088, Qianlong Liu, Baolin Peng, Xuanjing Huang 0001, Jiajie Peng, Zhongyu Wei |
Bioinform. | 3 |
| 2021 | Question Generation from Code Snippets and Programming Error Messages
Bolun Yao, Wei Chen 0088, Yeyun Gong, Bartuer Zhou, Zhongyu Wei, Biao Cheng, Nan Duan 0001 |
NLPCC (1) | 2 |
| 2020 | Automatic Generation of Electromyogram Diagnosis ReportabstractElectrophysiological tests, especially, electromyogram (EMG) and nerve conduction velocity (NCV) test are commonly used in clinical practice for diagnosis of muscle and nerve diseases. Report-writing of these tests can be problematic for under-experienced physicians and time-consuming for experienced physicians. In this paper, we apply several neural based natural language generation (NLG) methods to automatically generate diagnosis reports, a first attempt in this domain. Specifically, we use tabular diagnostic records of electrophysiological tests to generate Findings & Impression, which together constitute the diagnostic report. We further use gram-based metrics to evaluate our models and conduct a case study for the result. Qizheng Gu, Cong Nie, Ruixiang Zou, Wei Chen 0088, Chaojun Zheng, Dongqing Zhu, Xiaojun Mao, Zhongyu Wei, Dong Tian |
BIBM | 4 |