VLDB 2026 Research / reviewers in the wild / expert
Yi Liu 0148
dblp:97/4626-148
· DBLP profile ↗
15ranked-venue papers
2as first author
14since 2021 · last 2026
0009-0001-2795-5478ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FineRef: Fine-Grained Error Reflection and Correction for Long-Form Generation with CitationsabstractGenerating with citations is crucial for trustworthy Large Language Models (LLMs), yet even advanced LLMs often produce mismatched or irrelevant citations. Existing methods over-optimize citation fidelity while overlooking relevance to the user query, which degrades answer quality and robustness in real-world settings with noisy or irrelevant retrieved content. Moreover, the prevailing single-pass paradigm struggles to deliver optimal answers in long-form generation that requiring multiple citations. To address these limitations, we propose FineRef, a framework based on Fine-grained error Reflection, which explicitly teaches the model to self-identify and correct two key citation errors—mismatch and irrelevance—on a per-citation basis. FineRef follows a two-stage training strategy. The first stage instills an “attempt–reflect–correct” behavioral pattern via supervised fine-tuning, using fine-grained and controllable reflection data constructed by specialized lightweight models. An online self-reflective bootstrapping strategy is designed to improve generalization by iteratively enriching training data with verified, self-improving examples. To further enhance the self-reflection and correction capability, the second stage applies process-level reinforcement learning with a multi-dimensional reward scheme that promotes reflection accuracy, answer quality, and correction gain. Experiments on the ALCE benchmark demonstrate that FineRef significantly improves both citation performance and answer accuracy. Our 7B model outperforms GPT-4 by up to 18% in Citation F1 and 4% in EM Recall, while also surpassing the state-of-the-art model across key evaluation metrics. FineRef also exhibits strong generalization and robustness in domain transfer settings and noisy retrieval scenarios. Yixing Peng, Licheng Zhang 0002, Shancheng Fang, Yi Liu 0148, Peijian Gu, Quan Wang 0002 |
AAAI | 4 |
| 2026 | In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-FeedbackabstractTraining Large Language Models (LLMs) for chain-of-thought reasoning presents a significant challenge: supervised fine-tuning on a single "golden" rationale hurts generalization as it penalizes equally valid alternatives, whereas reinforcement learning with verifiable rewards struggles with credit assignment and prohibitive computational cost. To tackle these limitations, we introduce InTRO (In-Token Rationality Optimization), a new framework that enables both token-level exploration and self-feedback for accurate and concise reasoning. Instead of directly optimizing an intractable objective over all valid reasoning paths, InTRO leverages correction factors—token-wise importance weights estimated by the information discrepancy between the generative policy and its answer-conditioned counterpart, for informative next-token selection. This approach allows the model to perform token-level exploration and receive self-generated feedback within a single forward pass, ultimately encouraging accurate and concise rationales. Across six math-reasoning benchmarks, InTRO consistently outperforms other baselines, raising solution accuracy by up to 20% relative to the base model. Its chains of thought are also notably more concise, exhibiting reduced verbosity. Beyond this, InTRO enables cross-domain transfer, successfully adapting to out-of-domain reasoning tasks that extend beyond the realm of mathematics, demonstrating robust generalization. Mingye Zhu, Yi Liu 0148, Zheren Fu, Quan Wang 0002, Yongdong Zhang 0001 |
AAAI | 2 |
| 2025 | Mitigating Biases in Language Models via Bias UnlearningabstractMany studies have shown various biases targeting different demographic groups in language models, amplifying discrimination and harming fairness.Recent parameter modification debiasing approaches significantly degrade core capabilities such as text coherence and task accuracy.And Prompt-based debiasing methods, only effective for predefined trigger words, fail to address deeply embedded stereotypical associations in model parameters.In this paper, we propose BiasUnlearn, a novel model debiasing framework which achieves targeted debiasing via dual-pathway unlearning mechanisms coordinating stereotype forgetting with anti-stereotype retention, while preventing bias polarity reversal through adversarial forget set and dynamic dataset swapping.We conducted extensive experiments with multiple language models across various evaluation benchmarks.The results show that BiasUnlearn outperforms existing methods in mitigating bias in language models while retaining language modeling capabilities.Further experiments reveal that debiasing weights are transferable across model variants, confirming that bias representations become entrenched during pre-training and persist through fine-tuning phases. Dianqing Liu, Yi Liu 0148, Guoqing Jin, Zhendong Mao 0001 |
EMNLP | 2 |
| 2025 | DETCP: Self-Detoxifying Language Models With Contrastive PairsabstractInfluenced by context such as tone, emotion and demographic, pre-trained language models may generate harmful text, which limits their widespread application. While detoxifying language models seeks to reduce the likelihood of generating harmful content. There are two categories of detoxification strategies: fine-tuning language models and constraining outputs during inference. Neither category of methods achieved a proper balance between detoxification efficacy, the amount of annotated data, and inference efficiency. In this paper, we introduce a lightweight detoxification approach aiming at guiding the probability distribution of generated tokens towards the opposite direction of toxification, which relies on the language model itself and the contrastive pairs of contexts in the inference phase, without training. Experiments show that our method has state-of-the-art performance in detoxification effect while it has an edge in both fluency and speed of text generation. Dianqing Liu, Yi Liu 0148, Junbo Guo, Zhendong Mao 0001 |
ICASSP | 2 |
| 2025 | On-the-fly Preference Alignment via Principle-Guided DecodingabstractWith the rapidly expanding landscape of large language models, aligning model generations with human values and preferences is becoming increasingly important. Popular alignment methods, such as Reinforcement Learning from Human Feedback, have shown significant success in guiding models with greater control. However, these methods require considerable computational resources, which is inefficient, and substantial collection of training data to accommodate the diverse and pluralistic nature of human preferences, which is impractical. These limitations significantly constrain the scope and efficacy of both task-specific and general preference alignment methods. In this work, we introduce On-the-fly Preference Alignment via Principle-Guided Decoding (OPAD) to directly align
model outputs with human preferences during inference, eliminating the need for fine-tuning. Our approach involves first curating a surrogate solution to an otherwise infeasible optimization problem and then designing a principle-guided reward function based on this surrogate. The final decoding policy is derived by maximizing this customized reward, which exploits the discrepancy between the
constrained policy and its unconstrained counterpart. OPAD directly modifies the model’s predictions during inference, ensuring principle adherence without incurring the computational overhead of retraining or fine-tuning. Experiments show that OPAD achieves competitive or superior performance in both general and personalized alignment tasks, demonstrating its efficiency and effectiveness compared to state-of-the-art baselines. Mingye Zhu, Yi Liu 0148, Lei Zhang 0119, Junbo Guo, Zhendong Mao 0001 |
ICLR | 2 |
| 2025 | Leveraging Importance Sampling to Detach Alignment Modules from Large Language ModelsabstractThe widespread adoption of large language models (LLMs) across industries has increased the demand for high-quality and customizable outputs. However, traditional alignment methods often require retraining large pretrained models, making it difficult to quickly adapt and optimize LLMs for diverse applications. To address this limitation, we propose a novel \textit{Residual Alignment Model} (\textit{RAM}) that formalizes the alignment process as a type of importance sampling. In this framework, the unaligned upstream model serves as the proposal distribution, while the alignment process is framed as secondary sampling based on an autoregressive alignment module that acts as an estimator of the importance weights. This design enables a natural detachment of the alignment module from the target aligned model, improving flexibility and scalability. Based on this model, we derive an efficient sequence-level training strategy for the alignment module, which operates independently of the proposal module. Additionally, we develop a resampling algorithm with iterative token-level decoding to address the common first-token latency issue in comparable methods. Experimental evaluations on two leading open-source LLMs across diverse tasks, including instruction following, domain adaptation, and preference optimization, demonstrate that our approach consistently outperforms baseline models. Yi Liu 0148, Dianqing Liu, Mingye Zhu, Junbo Guo, Yongdong Zhang 0001, Zhendong Mao 0001 |
NeurIPS | 1 |
| 2025 | Leveraging robust optimization for llm alignment under distribution shiftsabstractPreference alignment methods are increasingly critical for steering large language models (LLMs) to generate outputs consistent with human values. While recent approaches often rely on synthetic data generated by LLMs for scalability and cost-efficiency reasons, this reliance can introduce distributional shifts that undermine the nuanced representation of human preferences needed for desirable outputs. In this paper, we propose a novel distribution-aware optimization framework that improves preference alignment despite such shifts. Our approach first leverages well-learned classifiers to assign a calibration value to each training sample, quantifying its alignment with the target human-preferred distribution. These values are then incorporated into a robust optimization objective that minimizes the worst-case loss over regions of the data space most relevant to human preferences. By explicitly focusing optimization on the target distribution, our approach mitigates the impact of distributional mismatch and improves the generation of responses that better reflect intended values. Mingye Zhu, Yi Liu 0148, Zheren Fu, Yongdong Zhang 0001, Zhendong Mao 0001 |
NeurIPS | 2 |
| 2025 | Exploiting Pre-Trained Language Models for Black-Box Attack against Knowledge Graph EmbeddingsabstractDespite the emerging research on adversarial attacks against knowledge graph embedding (KGE) models, most of them focus on white-box attack settings. However, white-box attacks are difficult to apply in practice compared to black-box attacks since they require access to model parameters that are unlikely to be provided. In this article, we propose a novel black-box attack method that only requires access to knowledge graph data, making it more realistic in real-world attack scenarios. Specifically, we utilize pre-trained language models (PLMs) to encode text features of the knowledge graphs, an aspect neglected by previous research. We then employ these encoded text features to identify the most influential triples for constructing corrupted triples for the attack. To improve the transferability of the attack, we further propose to fine-tune the PLM model by enriching triple embeddings with structure information. Extensive experiments conducted on two knowledge graph datasets illustrate the effectiveness of our proposed method. Guangqian Yang, Lei Zhang 0119, Yi Liu 0148, Hongtao Xie 0001, Zhendong Mao 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | Benchmarking Large Language Models on Controllable Generation under Diversified InstructionsabstractWhile large language models (LLMs) have exhibited impressive instruction-following capabilities, it is still unclear whether and to what extent they can respond to explicit constraints that might be entailed in various instructions. As a significant aspect of LLM alignment, it is thus important to formulate such a specialized set of instructions as well as investigate the resulting behavior of LLMs. To address this vacancy, we propose a new benchmark CoDI-Eval to systematically and comprehensively evaluate LLMs' responses to instructions with various constraints. We construct a large collection of constraints-attributed instructions as a test suite focused on both generalization and coverage. Specifically, we advocate an instruction diversification process to synthesize diverse forms of constraint expression and also deliberate the candidate task taxonomy with even finer-grained sub-categories. Finally, we automate the entire evaluation process to facilitate further developments. Different from existing studies on controllable text generation, CoDI-Eval extends the scope to the prevalent instruction-following paradigm for the first time. We provide extensive evaluations of representative LLMs (e.g., ChatGPT, Vicuna) on CoDI-Eval, revealing their limitations in following instructions with specific constraints and there is still a significant gap between open-source and commercial closed-source LLMs. We believe this benchmark will facilitate research into improving the controllability of LLMs' responses to instructions. Our data and code are available at https://github.com/Xt-cyh/CoDI-Eval. Yihan Chen 0001, Benfeng Xu, Quan Wang 0002, Yi Liu 0148, Zhendong Mao 0001 |
AAAI | 4 |
| 2024 | Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute EditingabstractGAN-based image attribute editing firstly leverages GAN Inversion to project real images into the latent space of GAN and then manipulates corresponding latent codes. Recent inversion methods mainly utilize additional high-bit features to improve image details preservation, as low-bit codes cannot faithfully reconstruct source images, leading to the loss of details. However, during editing, existing works fail to accurately complement the lost details and suffer from poor editability. The main reason is they inject all the lost details indiscriminately at one time, which inherently induces the position and quantity of details to overfit source images, resulting in inconsistent content and artifacts in edited images. This work argues that details should be gradually injected into both the reconstruction and editing process in a multi-stage coarse-to-fine manner for better detail preservation and high editability. Therefore, a novel dual-stream framework is proposed to accurately complement details at each stage. The Reconstruction Stream is employed to embed coarse-to-fine lost details into residual features and then adaptively add them to the GAN generator. In the Editing Stream, residual features are accurately aligned by our Selective Attention mechanism and then injected into the editing process in a multi-stage manner. Extensive experiments have shown the superiority of our framework in both reconstruction accuracy and editing quality compared with existing methods. Hao Li 0189, Mengqi Huang, Lei Zhang 0119, Bo Hu 0036, Yi Liu 0148, Zhendong Mao 0001 |
AAAI | 5 |
| 2024 | Visual-Linguistic Dependency Encoding for Image-Text RetrievalabstractImage-text retrieval is a fundamental task to bridge the semantic gap between natural language and vision. Recent works primarily focus on aligning textual meanings with visual appearance. However, they often overlook the semantic discrepancy caused by syntactic structure in natural language expressions and relationships among visual entities. This oversight would lead to sub-optimal alignment and degraded retrieval performance, since the underlying semantic dependencies and object interactions remain inadequately encoded in both textual and visual embeddings. In this paper, we propose a novel Visual-Linguistic Dependency Encoding (VL-DE) framework, which explicitly models the dependency information among textual words and interaction patterns between image regions, improving the discriminative power of cross-modal representations for more accurate image-text retrieval. Specifically, VL-DE enhances textual representations by considering syntactic relationships and dependency types, and visual representations by attending to its spatially neighboring regions. Cross-attention mechanism is then introduced to aggregate aligned region-word pairs into image-text similarities. Analysis on Winoground, a dataset specially designed to measure vision-linguistic compositional structure reasoning, shows that VL-DE outperforms existing methods, demonstrating its effectiveness at this task. Comprehensive experiments on two benchmarks, Flickr30K and MS-COCO, further validates the competitiveness of our approach. Wenxin Guo, Lei Zhang 0119, Kun Zhang 0040, Yi Liu 0148, Zhendong Mao 0001 |
LREC/COLING | 4 |
| 2024 | FlipGuard: Defending Preference Alignment against Update Regression with Constrained OptimizationabstractRecent breakthroughs in preference alignment have significantly improved Large Language Models' ability to generate texts that align with human preferences and values.However, current alignment metrics typically emphasize the post-hoc overall improvement, while overlooking a critical aspect: regression, which refers to the backsliding on previously correctly-handled data after updates.This potential pitfall may arise from excessive fine-tuning on already well-aligned data, which subsequently leads to over-alignment and degeneration.To address this challenge, we propose FlipGuard, a constrained optimization approach to detect and mitigate update regression with focal attention.Specifically, FlipGuard identifies performance degradation using a customized reward characterization and strategically enforces a constraint to encourage conditional congruence with the pre-aligned model during training.Comprehensive experiments demonstrate that FlipGuard effectively alleviates update regression while demonstrating excellent overall performance, with the added benefit of knowledge preservation while aligning preferences. Mingye Zhu, Yi Liu 0148, Quan Wang 0002, Junbo Guo, Zhendong Mao 0001 |
EMNLP | 2 |
| 2024 | Curriculum Learning Driven Domain Adaptation for Low-Resource Machine Reading ComprehensionabstractAlthough the pre-trained language models have achieved great success on machine reading comprehension task, they often rely on large-scale annotated data, while only a little amount of data is available in the most real-world scenarios. To enhance the PTLMs' capabilities in low-resource scenario, we propose a curriculum learning driven domain adaptation method for low-resource machine reading comprehension, the basic paradigm of which is to train a source model with sufficient data and then adaptive it to our target domain. In the adapting procedure, we introduce the curriculum learning strategy, the core idea of which is arranging training examples from easy to difficult, to bridge the gap between source and target domains and enable the source model adapting to the target domain progressively. Specifically, before fine-tuning the well-trained source model using target data, we firstly calculate the loss of each target example using the source model to evaluating the example difficulty accurately. After that, we sample suitable batches based on an increasing sampling function at each fine-tuning step, allowing the source model to start learning from easy examples in the target domain and gradually transition to difficult ones. Experiments conducted on two public datasets have demonstrated the effectiveness of our method. Licheng Zhang 0002, Quan Wang 0002, Benfeng Xu, Yi Liu 0148, Zhendong Mao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2023 | Random Entity Quantization for Parameter-Efficient Compositional Knowledge Graph RepresentationabstractRepresentation Learning on Knowledge Graphs (KGs) is essential for downstream tasks.The dominant approach, KG Embedding (KGE), represents entities with independent vectors and faces the scalability challenge.Recent studies propose an alternative way for parameter efficiency, which represents entities by composing entity-corresponding codewords matched from predefined small-scale codebooks.We refer to the process of obtaining corresponding codewords of each entity as entity quantization, for which previous works have designed complicated strategies.Surprisingly, this paper shows that simple random entity quantization can achieve similar results to current strategies.We analyze this phenomenon and reveal that entity codes, the quantization outcomes for expressing entities, have higher entropy at the code level and Jaccard distance at the codeword level under random entity quantization.Therefore, different entities become more easily distinguished, facilitating effective KG representation.The above results show that current quantization strategies are not critical for KG representation, and there is still room for improvement in entity distinguishability beyond current strategies.The code to reproduce our results is available here. Jiaang Li 0001, Quan Wang 0002, Yi Liu 0148, Licheng Zhang 0002, Zhendong Mao 0001 |
EMNLP | 3 |
| 2014 | Real-Time Scene Text Detection Based on Stroke ModelabstractIn this paper we bring forth a novel stroke-based method which is simple and effective to detect texts in natural scenes. We first introduce a general mathematical model to describe character strokes from the perspective of the scale space along with difference of Gaussian filters. Then we detail a text line aggregation approach utilizing the inherent text layout. Afterwards, we set up the whole scheme with three main steps, i.e. stroke extraction, text line aggregation and verification. Finally, experiments show the advantage of our method. As strokes are considered to be the fundamental component of characters, compared to edge- or other connected-component-based methods, our method is much more reasonable. Yi Liu 0148, Dongming Zhang 0004, Yongdong Zhang 0001, Shouxun Lin |
ICPR | 1 |