EDBT 2026 Demo / reviewers in the wild / expert
Yichao Du
dblp:271/6727
· DBLP profile ↗
20ranked-venue papers
5as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language ModelsabstractZhiwen Ruan, Yichao Du, Jianjie Zheng, Longyue Wang, Yun Chen, Peng Li, Jinsong Su, Yang Liu, Guanhua Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhiwen Ruan, Yichao Du, Jianjie Zheng, Longyue Wang, Yun Chen 0007, Peng Li 0030, Jinsong Su, Yang Liu 0005, Guanhua Chen 0001 |
ACL (1) | 2 |
| 2026 | DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text DetectionabstractThe effective detection and governance of Large Language Model (LLM) generated content has become increasingly critical due to the growing risk of misuse. Despite the impressive performance of existing detectors, their reliability and potential in multilingual, real-world scenarios remain largely underexplored.In this study, we introduce DetectRL-X, a comprehensive multilingual benchmark designed to evaluate advanced detectors across 8 dimensions. The benchmark encompasses 8 languages commonly used in commercial contexts and collects human-written texts from 6 domains highly susceptible to LLM misuse. To better aligned with real-world applications, We create LLM-generated texts using 4 popular commercial LLMs, and include typical AI-assisted writing operations such as polishing, expanding, and condensing to capture authentic usage patterns. Furthermore, we develop a multilingual framework for paraphrasing and perturbation attacks to simulate diverse human modifications and writing noise, enabling stress testing of detectors across languages.Experimental results on DetectRL-X reveal the strengths and limitations of current state-of-the-art detectors when applied to diverse linguistic resources. We further analyze how domains, generators, attack strategies, text length, and refinement operations influence performance in different languages, underscoring DetectRL-X as an effective benchmark for strengthening multilingual and language-specific detectors. Junchao Wu, Yefeng Liu, Chenyu Zhu, Tianqi Shi, Yichao Du, Longyue Wang, Weihua Luo, Jinsong Su, Derek F. Wong |
ACL (1) | 7 |
| 2026 | Precedent Retrieval Based In-Context Learning for Legal Rationale Generation
Linan Yue, Yichao Du, Weibo Gao |
DASFAA (6) | 2 |
| 2025 | MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?abstractWhile text-to-image models like GPT-4o-Image and FLUX are rapidly proliferating, they often encounter challenges such as hallucination, bias, and the production of unsafe, low-quality output. To effectively address these issues, it is crucial to align these models with desired behaviors based on feedback from a multimodal judge. Despite their significance, current multimodal judges frequently undergo inadequate evaluation of their capabilities and limitations, potentially leading to misalignment and unsafe fine-tuning outcomes. To address this issue, we introduce MJ-Bench, a novel benchmark which incorporates a comprehensive preference dataset to evaluate multimodal judges in providing feedback for image generation models across six key perspectives: alignment, safety, image quality, bias, composition, and visualization. Specifically, we evaluate a large variety of multimodal judges including smaller-sized CLIP-based scoring models, open-source VLMs, and close-source VLMs on each decomposed subcategory of our preference dataset. Experiments reveal that close-source VLMs generally provide better feedback, with GPT-4o outperforming other judges in average. Compared with open-source VLMs, smaller-sized scoring models can provide better feedback regarding text-image alignment and image quality, while VLMs provide more accurate feedback regarding safety and generation bias due to their stronger reasoning capabilities. Further studies in feedback scale reveal that VLM judges can generally provide more accurate and stable feedback in natural language than numerical scales. Notably, human evaluations on end-to-end and fine-tuned models using separate feedback from these multimodal judges provide similar conclusions, further confirming the effectiveness of MJ-Bench. Zhaorun Chen, Zichen Wen, Yichao Du, Yiyang Zhou, Chenhang Cui, Siwei Han, Jen Weng, Chaoqi Wang, Zhengwei Tong, Leria Huang, Canyu Chen, Haoqin Tu, Qinghao Ye, Zhihong Zhu 0001, Zhuokai Zhao, Rafael Rafailov, Chelsea Finn, Huaxiu Yao |
NeurIPS | 3 |
| 2025 | Towards Reliable and Faithful Explanations: A Disentanglement-Augmented Approach for Selective RationalizationabstractThe pursuit of model explainability has prompted the selective rationalization (aka, rationale extraction) which can identify important features (i.e., rationales) from the original input to support prediction results. Existing methods typically involve a cascaded approach with a selector responsible for extracting rationales from the input, followed by a predictor that makes predictions based on the selected rationales. However, these approaches often neglect the information contained in the non-rationales, underutilizing the input. Therefore, in our prior work, we introduce the Disentanglement-Augmented Rationale Extraction (DARE) method, which disentangles the input into rationale and non-rationale components, and enhances rationale representations by minimizing the mutual information between them. While DARE demonstrates strong performance in rationalization, it may still rely on shortcuts in the training distribution, leading to unfaithful rationales. To this end, in this paper, we propose Faith-DARE, an extension of DARE that aims to extract more reliable rationales by mitigating shortcut dependencies. Specifically, we treat the non-rationale features identified by DARE as environments that are decorrelated from the predictions. By shuffling and recombining these environments with rationales, we generate counterfactual samples and identify invariant rationales that remain predictive across shifted distributions. Extensive experiments on graph and textual datasets validate the effectiveness of Faith-DARE. Linan Yue, Qi Liu 0003, Yichao Du, Li Wang 0014, Yanqing An, Enhong Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | FedJudge: Federated Legal Large Language Model
Linan Yue, Qi Liu 0003, Yichao Du, Weibo Gao, Ye Liu 0011, Fangzhou Yao |
DASFAA (5) | 3 |
| 2024 | Communication-Efficient Personalized Federated Learning for Speech-to-Text TasksabstractTo protect privacy and meet legal regulations, federated learning (FL) has gained significant attention for training speech-to-text (S2T) systems, including automatic speech recognition (ASR) and speech translation (ST). However, the commonly used FL approach (i.e., FEDAVG) in S2T tasks typically suffers from extensive communication overhead due to multi-round interactions based on the whole model and performance degradation caused by data heterogeneity among clients. To address these issues, we propose a personalized federated S2T framework that introduces FEDLORA, a lightweight LoRA module for client-side tuning and interaction with the server to minimize communication overhead, and FEDMEM, a global model equipped with a k-nearest-neighbor (kNN) classifier that captures client-specific distributional shifts to achieve personalization and overcome data heterogeneity. Extensive experiments based on Conformer and Whisper backbone models on CoVoST and GigaSpeech benchmarks show that our approach significantly reduces the communication overhead on all S2T tasks and effectively personalizes the global model to overcome data heterogeneity. Yichao Du, Zhirui Zhang, Linan Yue, Xu Huang 0008, Tong Xu 0001, Linli Xu 0002, Enhong Chen |
ICASSP | 1 |
| 2024 | Towards Faithful Explanations: Boosting Rationalization with Shortcuts DiscoveryabstractThe remarkable success in neural networks provokes the selective rationalization. It explains the prediction results by identifying a small subset of the inputs sufficient to support them. Since existing methods still suffer from adopting the shortcuts in data to compose rationales and limited large-scale annotated rationales by human, in this paper, we propose a Shortcuts-fused Selective Rationalization (SSR) method, which boosts the rationalization by discovering and exploiting potential shortcuts. Specifically, SSR first designs a shortcuts discovery approach to detect several potential shortcuts. Then, by introducing the identified shortcuts, we propose two strategies to mitigate the problem of utilizing shortcuts to compose rationales. Finally, we develop two data augmentations methods to close the gap in the number of annotated rationales. Extensive experimental results on real-world datasets clearly validate the effectiveness of our proposed method. Linan Yue, Qi Liu 0003, Yichao Du, Li Wang 0014, Weibo Gao, Yanqing An |
ICLR | 3 |
| 2024 | Federated Self-Explaining GNNs with Anti-shortcut AugmentationsabstractGraph Neural Networks (GNNs) have demonstrated remarkable performance in graph classification tasks. However, ensuring the explainability of their predictions remains a challenge. To address this, graph rationalization methods have been introduced to generate concise subsets of the original graph, known as rationales, which serve to explain the predictions made by GNNs. Existing rationalizations often rely on shortcuts in data for prediction and rationale composition. In response, de-shortcut rationalization methods have been proposed, which commonly leverage counterfactual augmentation to enhance data diversity for mitigating the shortcut problem. Nevertheless, these methods have predominantly focused on centralized datasets and have not been extensively explored in the Federated Learning (FL) scenarios. To this end, in this paper, we propose a Federated Graph Rationalization (FedGR) with anti-shortcut augmentations to achieve self-explaining GNNs, which involves two data augmenters. These augmenters are employed to produce client-specific shortcut conflicted samples at each client, which contributes to mitigating the shortcut problem under the FL scenarios. Experiments on real-world benchmarks and synthetic datasets validate the effectiveness of FedGR under the FL scenarios. Linan Yue, Qi Liu 0003, Weibo Gao, Ye Liu 0011, Kai Zhang 0038, Yichao Du, Li Wang 0014, Fangzhou Yao |
ICML | 6 |
| 2024 | Datastore Distillation for Nearest Neighbor Machine TranslationabstractNearest neighbor machine translation (i.e.,$k$NN-MT) is a promising approach to enhance translation quality by equipping pre-trained neural machine translation (NMT) models with the nearest neighbor retrieval. Despite its great success,$k$NN-MT typically requires ample space to store its token-level datastore, causing$k$NN-MT to be less practical in edge devices or online scenarios. In this paper, inspired by the concept of knowledge distillation, we provide a new perspective to ease the storage overhead by datastore distillation, which is formalized as a constrained optimization problem. We further design a novel model-agnostic iterative nearest neighbor merging method for the datastore distillation problem to obtain an effective and efficient solution. Experiments on three benchmark datasets indicate that our approach not only reduces the volume of the datastore by up to 50% without significant performance degradation, but also outperforms other baselines by a large margin at the same compression rate. Another experiment conducted on WikiText-103 further demonstrates the effectiveness of our method in the language model task. Yuhan Dai, Zhirui Zhang, Yichao Du, Shengcai Liu, Lemao Liu, Tong Xu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Nearest Neighbor Machine Translation is Meta-Optimizer on Output Projection LayerabstractNearest Neighbor Machine Translation (kNN-MT) has achieved great success in domain adaptation tasks by integrating pre-trained Neural Machine Translation (NMT) models with domain-specific token-level retrieval.However, the reasons underlying its success have not been thoroughly investigated.In this paper, we comprehensively analyze kNN-MT through theoretical and empirical studies.Initially, we provide new insights into the working mechanism of kNN-MT as an efficient technique to implicitly execute gradient descent on the output projection layer of NMT, indicating that it is a specific case of model fine-tuning.Subsequently, we conduct multi-domain experiments and word-level analysis to examine the differences in performance between kNN-MT and entire-model fine-tuning.Our findings suggest that: (i) Incorporating kNN-MT with adapters yields comparable translation performance to fine-tuning on in-domain test sets, while achieving better performance on out-of-domain test sets; (ii) Fine-tuning significantly outperforms kNN-MT on the recall of in-domain low-frequency words, but this gap could be bridged by optimizing the context representations with additional adapter layers. 1 Zhirui Zhang, Yichao Du, Lemao Liu, Rui Wang 0015 |
EMNLP | 3 |
| 2023 | IMTLab: An Open-Source Platform for Building, Evaluating, and Diagnosing Interactive Machine Translation SystemsabstractXu Huang, Zhirui Zhang, Ruize Gao, Yichao Du, Lemao Liu, Guoping Huang, Shuming Shi, Jiajun Chen, Shujian Huang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Zhirui Zhang, Yichao Du, Lemao Liu, Guoping Huang, Shuming Shi 0001, Jiajun Chen 0001, Shujian Huang |
EMNLP | 4 |
| 2023 | Interventional RationalizationabstractSelective rationalizations improve the explainability of neural networks by selecting a subsequence of the input (i.e., rationales) to explain the prediction results.Although existing methods have achieved promising results, they still suffer from adopting the spurious correlations in data (aka., shortcuts) to compose rationales and make predictions.Inspired by the causal theory, in this paper, we develop an interventional rationalization (Inter-RAT) to discover the causal rationales.Specifically, we first analyse the causalities among the input, rationales and results with a causal graph.Then, we discover spurious correlations between the input and rationales, and between rationales and results, respectively, by identifying the confounder in the causalities.Next, based on the backdoor adjustment, we propose a causal intervention method to remove the spurious correlations between input and rationales.Further, we discuss reasons why spurious correlations between the selected rationales and results exist by analysing the limitations of the sparsity constraint in the rationalization, and employ the causal intervention method to remove these correlations.Extensive experimental results on three realworld datasets clearly validate the effectiveness of our proposed method.The source code of Inter-RAT is available at https://github. com/yuelinan/Codes-of-Inter-RAT. Linan Yue, Qi Liu 0003, Li Wang 0014, Yanqing An, Yichao Du, Zhenya Huang |
EMNLP | 5 |
| 2023 | Simple and Scalable Nearest Neighbor Machine Translation
Yuhan Dai, Zhirui Zhang, Qiuzhi Liu, Qu Cui, Yichao Du, Tong Xu 0001 |
ICLR | 6 |
| 2023 | Federated Nearest Neighbor Machine Translation
Yichao Du, Zhirui Zhang, Bingzhe Wu, Lemao Liu, Tong Xu 0001, Enhong Chen |
ICLR | 1 |
| 2022 | Regularizing End-to-End Speech Translation with Triangular Decomposition AgreementabstractEnd-to-end speech-to-text translation (E2E-ST) is becoming increasingly popular due to the potential of its less error propagation, lower latency, and fewer parameters. Given the triplet training corpus〈speech, transcription, translation〉, the conventional high-quality E2E-ST system leverages the〈speech, transcription〉pair to pre-train the model and then utilizes the〈speech, translation〉pair to optimize it further. However, this process only involves two-tuple data at each stage, and this loose coupling fails to fully exploit the association between triplet data. In this paper, we attempt to model the joint probability of transcription and translation based on the speech input to directly leverage such triplet data. Based on that, we propose a novel regularization method for model training to improve the agreement of dual-path decomposition within triplet data, which should be equal in theory. To achieve this goal, we introduce two Kullback-Leibler divergence regularization terms into the model training objective to reduce the mismatch between output probabilities of dual-path. Then the well-trained model can be naturally transformed as the E2E-ST models by a pre-defined early stop tag. Experiments on the MuST-C benchmark demonstrate that our proposed approach significantly outperforms state-of-the-art E2E-ST baselines on all 8 language pairs while achieving better performance in the automatic speech recognition task. Yichao Du, Zhirui Zhang, Weizhi Wang, Boxing Chen, Tong Xu 0001 |
AAAI | 1 |
| 2022 | Non-Parametric Domain Adaptation for End-to-End Speech TranslationabstractThe end-to-end speech translation (E2E-ST) has received increasing attention due to the potential of its less error propagation, lower latency and fewer parameters.However, the effectiveness of neural-based approaches to this task is severely limited by the available training corpus, especially for domain adaptation where in-domain triplet data is scarce or nonexistent.In this paper, we propose a novel nonparametric method that leverages in-domain text translation corpus to achieve domain adaptation for E2E-ST systems.To this end, we first incorporate an additional encoder into the pre-trained E2E-ST model to realize text translation modeling, based on which the decoder's output representations for text and speech translation tasks are unified by reducing the correspondent representation mismatch in available triplet training data.During domain adaptation, a k-nearest-neighbor (kNN) classifier is introduced to produce the final translation distribution using the external datastore built by the domain-specific text translation corpus, while the universal output representation is adopted to perform a similarity search.Experiments on the Europarl-ST benchmark demonstrate that when in-domain text translation data is involved only, our proposed approach significantly improves baseline by 12.82 BLEU on average in all translation directions, even outperforming the strong in-domain fine-tuning strategy. Yichao Du, Weizhi Wang, Zhirui Zhang, Boxing Chen, Tong Xu 0001, Enhong Chen |
EMNLP | 1 |
| 2022 | DARE: Disentanglement-Augmented Rationale ExtractionabstractRationale extraction can be considered as a straightforward method of improving the model explainability, where rationales are a subsequence of the original inputs, and can be extracted to support the prediction results. Existing methods are mainly cascaded with the selector which extracts the rationale tokens, and the predictor which makes the prediction based on selected tokens. Since previous works fail to fully exploit the original input, where the information of non-selected tokens is ignored, in this paper, we propose a Disentanglement-Augmented Rationale Extraction (DARE) method, which encapsulates more information from the input to extract rationales. Specifically, it first disentangles the input into the rationale representations and the non-rationale ones, and then learns more comprehensive rationale representations for extracting by minimizing the mutual information (MI) between the two disentangled representations. Besides, to improve the performance of MI minimization, we develop a new MI estimator by exploring existing MI estimation methods. Extensive experimental results on three real-world datasets and simulation studies clearly validate the effectiveness of our proposed method. Code is released at https://github.com/yuelinan/DARE. Linan Yue, Qi Liu 0003, Yichao Du, Yanqing An, Li Wang 0014, Enhong Chen |
NeurIPS | 3 |
| 2021 | Inheritance-Guided Hierarchical Assignment for Clinical Automatic Diagnosis
Yichao Du, Pengfei Luo, Tong Xu 0001, Yi Zheng 0007, Enhong Chen |
DASFAA (3) | 1 |
| 2020 | LBBESA: An efficient software-defined networking load-balancing scheme based on elevator scheduling algorithmabstractSummary Elevator scheduling algorithms generally denote methods used to calculate how to use the elevator. These algorithms can distribute elevators to various floors of a building, thereby achieving efficient transportation. From the perspective of the elevator scheduling problem, we address the load‐balancing problem for software‐defined networking (SDN) architecture and propose a load‐balancing method based on the elevator scheduling algorithm, LBBESA. We take advantage of the flexibility of the SDN architecture, obtain the real‐time load of the server through real‐time statistical analyses of the SDN switch port traffic by the controller, and combine this with the idea of regional elevator allocation to coordinate the connection of the client's requests and realize the load balancing of each server in the cluster. Simulation experiments show that, compared with the round‐robin algorithm, LBBESA is more effective in the load balancing of the server pool and can improve the throughput of the server pool to a certain extent. In addition, our scheme is easy to implement and has high scalability. Qiliang Li, Jie Cui 0004, Hong Zhong 0001, Yichao Du, Yonglong Luo, Lu Liu 0001 |
Concurr. Comput. Pract. Exp. | 4 |