EDBT 2026 Demo / reviewers in the wild / expert
Leilei Gan
dblp:213/9040
· DBLP profile ↗
17ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0001-5859-2588ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image GenerationabstractZiwei Huang, Ying Shu, Fanghao, Quanyu Long, Wenya Wang, Qiushi Guo, Tiezheng Ge, Leilei Gan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ziwei Huang 0005, Yin Shu, Quanyu Long, Wenya Wang 0001, Qiushi Guo, Tiezheng Ge, Leilei Gan |
ACL (1) | 8 |
| 2026 | REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment EvaluationabstractEvaluating the alignment between textual prompts and generated images is critical for ensuring the reliability and usability of textto-image (T2I) models.However, most existing evaluation methods rely on coarsegrained metrics or static Question Answering (QA) pipelines, which lack fine-grained interpretability and struggle to reflect human preferences.To address this, we propose REVEALER, a reinforcement-guided visual reasoning framework for element-level textto-image alignment evaluation.Adopting a structured "grounding-reasoning-conclusion" paradigm, our method enables Multimodal Large Language Models (MLLMs) to explicitly localize semantic elements and derive interpretable alignment judgments.We optimize the model via Group Relative Policy Optimization (GRPO) using a multi-dimensional reward function that targets format compliance, localization precision, and alignment accuracy.Extensive experiments confirm that RE-VEALER achieves state-of-the-art results across four benchmarks.Notably, on EvalMuse-40K, it surpasses the strong proprietary Gemini 3 Pro and Training-based baselines with absolute accuracy gains of +4.2% and +13.3%, respectively.Ablation studies further demonstrate the efficacy of our method, contributing a cumulative 19.6% improvement over the base model. Fulin Shi, Wenyi Xiao, Leilei Gan |
ACL (1) | 5 |
| 2026 | From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document UnderstandingabstractYandi Wang, Libin Zhan, Ziwei Huang, Tiancheng Luo, Yuxuan Jiang, Wang Dong, Leilei Gan, Jun Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yandi Wang, Libin Zhan, Ziwei Huang 0005, Tiancheng Luo, Wang Dong, Leilei Gan |
ACL (1) | 7 |
| 2026 | VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models ReasoningabstractLarge Vision Language Models (LVLMs) achieve strong multimodal reasoning but frequently exhibit hallucinations and incorrect responses with high certainty, which hinders their usage in high-stakes domains. Existing verbalized confidence calibration methods, largely developed for text-only LLMs, typically optimize a single holistic confidence score using binary answer-level correctness. This design is mismatched to LVLMs: an incorrect prediction may arise from perceptual failures or from reasoning errors given correct perception, and a single confidence conflates these sources while visual uncertainty is often dominated by language priors. To address these issues, we propose VL-Calibration, a reinforcement learning framework that explicitly decouples confidence into visual and reasoning confidence. To supervise visual confidence without ground-truth perception labels, we introduce an intrinsic visual certainty estimation that combines (i) visual grounding measured by KL-divergence under image perturbations and (ii) internal certainty measured by token entropy. We further propose token-level advantage reweighting to focus optimization on tokens based on visual certainty, suppressing ungrounded hallucinations while preserving valid perception. Experiments on thirteen benchmarks show that VL-Calibration effectively improves calibration while boosting visual reasoning accuracy, and it generalizes to out-of-distribution benchmarks across model scales and architectures. Wenyi Xiao, Xinchi Xu, Leilei Gan |
ACL (1) | 3 |
| 2025 | MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image SynthesisabstractAuto-regressive models have made significant progress in the realm of text-to-image synthesis, yet devising an appropriate model architecture and training strategy to achieve a satisfactory level remains an important avenue of exploration. In this work, we introduce MARS, a novel framework for T2I generation that incorporates a specially designed Semantic Vision-Language Integration Expert (SemVIE). This innovative component integrates pre-trained LLMs by independently processing linguistic and visual information—freezing the textual component while fine-tuning the visual component. This methodology preserves the NLP capabilities of LLMs while imbuing them with exceptional visual understanding. Building upon the powerful base of the pre-trained Qwen-7B, MARS stands out with its bilingual generative capabilities corresponding to both English and Chinese language prompts and the capacity for joint image and text generation. The flexibility of this framework lends itself to migration towards any-to-any task adaptability. Furthermore, MARS employs a multi-stage training strategy that first establishes robust image-text alignment through complementary bidirectional tasks and subsequently concentrates on refining the T2I generation process, significantly augmenting text-image synchrony and the granularity of image details. Notably, MARS requires only 9% of the GPU days needed by SD1.5, yet it achieves remarkable results across a variety of benchmarks, illustrating the training efficiency and the potential for swift deployment in various applications. Wanggui He, Siming Fu, Mushui Liu, Xierui Wang, Wenyi Xiao, Fangxun Shu, Yi Wang 0068, Lei Zhang 0006, Zhelun Yu, Haoyuan Li 0002, Ziwei Huang 0005, Leilei Gan, Hao Jiang 0014 |
AAAI | 12 |
| 2025 | Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI FeedbackabstractThe rapidly developing Large Vision Language Models (LVLMs) still face the hallucination phenomena where the generated responses do not align with the given contexts, significantly restricting the usages of LVLMs. Most previous work detects and mitigates hallucination at the coarse-grained level or requires expensive annotation (e.g., labeling by human experts or proprietary models). To address these issues, we propose detecting and mitigating hallucinations in LVLMs via fine-grained AI feedback. The basic idea is that we generate a small-size sentence-level hallucination annotation dataset by proprietary models, whereby we train a detection model which can perform sentence-level hallucination detection. Then, we propose a detect-then-rewrite pipeline to automatically construct preference dataset for hallucination mitigation training. Furthermore, we propose differentiating the severity of hallucinations, and introducing a Hallucination Severity-Aware Direct Preference Optimization (HSA-DPO) which prioritizes the mitigation of critical hallucination in LVLMs by incorporating the severity of hallucinations into preference learning. Extensive experiments on hallucination detection and mitigation benchmarks demonstrate that our method sets a new state-of-the-art in hallucination detection on MHaluBench, surpassing GPT-4V and Gemini, and reduces the hallucination rate by 36.1% on AMBER and 76.3% on Object HalBench compared to the base model. Wenyi Xiao, Ziwei Huang 0005, Leilei Gan, Wanggui He, Haoyuan Li 0002, Zhelun Yu, Fangxun Shu, Hao Jiang 0014, Linchao Zhu |
AAAI | 3 |
| 2025 | T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive ConceptsabstractZiwei Huang, Wanggui He, Quanyu Long, Yandi Wang, Haoyuan Li, Zhelun Yu, Fangxun Shu, Weilong Dai, Hao Jiang, Fei Wu, Leilei Gan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ziwei Huang 0005, Wanggui He, Quanyu Long, Yandi Wang, Haoyuan Li 0002, Zhelun Yu, Fangxun Shu, Weilong Dai, Hao Jiang 0014, Fei Wu 0001, Leilei Gan |
ACL (1) | 11 |
| 2025 | Fine-tuning Large Language Models for Improving Factuality in Legal Question AnsweringabstractHallucination, or the generation of incorrect or fabricated information, remains a critical challenge in large language models (LLMs), particularly in high-stake domains such as legal question answering (QA). In order to mitigate the hallucination rate in legal QA, we first introduce a benchmark called LegalHalBench and three automatic metrics to evaluate the common hallucinations when LLMs answer legal questions. We then propose a hallucination mitigation method that integrates behavior cloning and a novel Hard Sample-aware Iterative Direct Preference Optimization (HIPO). We conduct extensive real-data experiments to validate the effectiveness of our approach. Our results demonstrate remarkable improvements in various metrics, including the newly proposed Non-Hallucinated Statute Rate, Statute Relevance Rate, Legal Claim Truthfulness, as well as traditional metrics such as METEOR, BERTScore, ROUGE-L, and win rates. Yinghao Hu 0001, Leilei Gan, Wenyi Xiao, Kun Kuang 0001, Fei Wu 0001 |
COLING | 2 |
| 2025 | Fast-Slow Thinking GRPO for Large Vision-Language Model ReasoningabstractWhen applying reinforcement learning—typically through GRPO—to large vision-language model reasoning struggles to effectively scale reasoning length or generates verbose outputs across all tasks with only marginal gains in accuracy.
To address this issue, we present FAST-GRPO, a variant of GRPO that dynamically adapts reasoning depth based on question characteristics.
Through empirical analysis, we establish the feasibility of fast-slow thinking in LVLMs by investigating how response length and data distribution affect performance.
Inspired by these observations, we introduce two complementary metrics to estimate the difficulty of the questions, guiding the model to determine when fast or slow thinking is more appropriate.
Next, we incorporate adaptive length-based rewards and difficulty-aware KL divergence into the GRPO algorithm.
Experiments across seven reasoning benchmarks demonstrate that FAST achieves state-of-the-art accuracy with over 10% relative improvement compared to the base model, while reducing token usage by 32.7-67.3% compared to previous slow-thinking approaches, effectively balancing reasoning length and accuracy. Wenyi Xiao, Leilei Gan |
NeurIPS | 2 |
| 2024 | Deconfounded hierarchical multi-granularity classification
Ziyu Zhao 0001, Leilei Gan, Tao Shen 0002, Kun Kuang 0001, Fei Wu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2022 | Dependency Parsing as MRC-based Span-Span PredictionabstractHigher-order methods for dependency parsing can partially but not fully address the issue that edges in dependency trees should be constructed at the text span/subtree level rather than word level.In this paper, we propose a new method for dependency parsing to address this issue.The proposed method constructs dependency trees by directly modeling span-span (in other words, subtree-subtree) relations.It consists of two modules: the text span proposal module which proposes candidate text spans, each of which represents a subtree in the dependency tree denoted by (root, start, end); and the span linking module, which constructs links between proposed spans.We use the machine reading comprehension (MRC) framework as the backbone to formalize the span linking module, where one span is used as query to extract the text span/subtree it should be linked to.The proposed method has the following merits: (1) it addresses the fundamental problem that edges in a dependency tree should be constructed between subtrees;(2) the MRC framework allows the method to retrieve missing spans in the span proposal stage, which leads to higher recall for eligible spans.Extensive experiments on the PTB, CTB and Universal Dependencies (UD) benchmarks demonstrate the effectiveness of the proposed method. 1 2 Leilei Gan, Yuxian Meng, Kun Kuang 0001, Xiaofei Sun 0001, Chun Fan 0001, Fei Wu 0001, Jiwei Li 0001 |
ACL (1) | 1 |
| 2022 | Investigating the Robustness of Natural Language Generation from Logical Forms via Counterfactual SamplesabstractThe aim of Logic2Text is to generate controllable and faithful texts conditioned on tables and logical forms, which not only requires a deep understanding of the tables and logical forms, but also warrants symbolic reasoning over the tables.State-of-the-art methods based on pre-trained models have achieved remarkable performance on the standard test dataset.However, we question whether these methods really learn how to perform logical reasoning, rather than just relying on the spurious correlations between the headers of the tables and operators of the logical form.To verify this hypothesis, we manually construct a set of counterfactual samples, which modify the original logical forms to generate counterfactual logical forms with rarely co-occurred table headers and logical operators.SOTA methods give much worse results on these counterfactual samples compared with the results on the original test dataset, which verifies our hypothesis.To deal with this problem, we firstly analyze this bias from a causal perspective, based on which we propose two approaches to reduce the model's reliance on the shortcut.The first one incorporates the hierarchical structure of the logical forms into the model.The second one exploits automatically generated counterfactual data for training.Automatic and manual experimental results on the original test dataset and the counterfactual dataset show that our method is effective to alleviate the spurious correlation.Our work points out the weakness of previous methods and takes a further step toward developing Logic2Text models with real logical reasoning ability. Leilei Gan, Kun Kuang 0001, Fei Wu 0001 |
EMNLP | 2 |
| 2022 | Triggerless Backdoor Attack for NLP Tasks with Clean LabelsabstractLeilei Gan, Jiwei Li, Tianwei Zhang, Xiaoya Li, Yuxian Meng, Fei Wu, Yi Yang, Shangwei Guo, Chun Fan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Leilei Gan, Jiwei Li 0001, Tianwei Zhang 0004, Xiaoya Li 0001, Yuxian Meng, Fei Wu 0001, Yi Yang 0001, Shangwei Guo, Chun Fan 0001 |
NAACL-HLT | 1 |
| 2022 | SemGloVe: Semantic Co-Occurrences for GloVe From BERTabstractGloVe learns word embeddings by leveraging statistical information from word co-occurrence matrices. However, word pairs in the matrices are extracted from a predefined local context window, which might lead to limited word pairs and potentially semantic irrelevant word pairs. In this paper, we proposeSemGloVe, which distillssemantic co-occurrencesfrom BERT into static GloVe word embeddings. Particularly, we propose two models to extract co-occurrence statistics based on either the masked language model or the multi-head attention weights of BERT. Our methods can extract word pairs limited by the local window assumption, and can define the co-occurrence weights by directly considering the semantic distance between word pairs. Experiments on several word similarity datasets and external tasks show that SemGloVe can outperform GloVe. Leilei Gan, Zhiyang Teng, Yue Zhang 0004, Linchao Zhu, Fei Wu 0001, Yi Yang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Judgment Prediction via Injecting Legal Knowledge into Neural NetworksabstractLegal Judgment Prediction (LJP) is a key problem in legal artificial intelligence, which is aimed to predict a law case's judgment based on a given text describing the facts of the law case. Most of the previous work treats LJP as a text classification task and generally adopts deep neural networks (DNNs) based methods to solve it. However, existing DNNs based work is data-hungry and hard to explain which legal knowledge is based on to make such a prediction. Thus, injecting legal knowledge into neural networks to interpret the model and improve performance remains a significant problem. In this paper, we propose to represent declarative legal knowledge as a set of first-order logic rules and integrate these logic rules into a co-attention network-based model explicitly. The use of logic rules enhances neural networks with explicit logical reason capabilities and makes the model more interpretable. We take the civil loan scenario as a case study and demonstrate the effectiveness of the proposed method through comprehensive experiments and analysis conducted on the collected dataset. Leilei Gan, Kun Kuang 0001, Yi Yang 0001, Fei Wu 0001 |
AAAI | 1 |
| 2020 | Investigating Self-Attention Network for Chinese Word SegmentationabstractNeural network has become the dominant method for Chinese word segmentation. Most existing models cast the task as sequence labeling, using BiLSTM-CRF for representing the input, and making output predictions. Recently, attention-based sequence models have emerged as a highly competitive alternative to LSTMs, which allow better running speed by parallelization of computation. We investigate self-attention network (SAN) for Chinese word segmentation, making comparisons between BiLSTM-CRF models. In addition, the influence of contextualized character embeddings is investigated using BERT, and a method is proposed for integrating word information into SAN segmentation. Results show that SAN gives highly competitive results compared with BiLSTMs, with BERT, and word information further improving segmentation for in-domain, and cross-domain segmentation. Our final models give the best results for 6 heterogenous domain benchmarks. Leilei Gan, Yue Zhang 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Hybrid computation offloading for smart home automation in mobile cloud computing
Jie Zhang 0053, Zhili Zhou 0001, Leilei Gan, Xuyun Zhang, Lianyong Qi, Xiaolong Xu 0001, Wan-Chun Dou |
Pers. Ubiquitous Comput. | 4 |