EDBT 2026 Demo / reviewers in the wild / expert
Yezi Liu
dblp:197/5408
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-0454-5238ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Debias Once for All: A Data-Centric Strategy for Fair Machine LearningabstractThe increasing use of deep neural networks (DNNs) in high-stakes domains such as hiring, healthcare, and finance has heightened concerns about algorithmic fairness. Because training data can encode historical and societal biases, learned models may exhibit disparate outcomes for underrepresented groups. Prior work is largely model-centric, improving fairness via specialized loss functions or architectural modifications, which can introduce additional training overhead and hinder deployment in modular or rapidly evolving pipelines. We instead study a data-centric alternative: constructing a fair training dataset that promotes equitable behavior without changing the model architecture. We propose FairData, which synthesizes a fair dataset by optimizing a gradient-matching objective that aligns the training dynamics of a randomly initialized model on the synthetic data with those on the original data, while explicitly regularizing for group fairness. The resulting dataset is model-agnostic, lightweight, and remains in the original input space, enabling straightforward reuse across downstream models. Experiments on four benchmark datasets show that FairData consistently reduces group disparities across diverse architectures while maintaining competitive predictive performance, suggesting fairness-aware dataset optimization as a practical complement to model-specific fairness techniques. Yezi Liu, Hanning Chen, Yanning Shen, Mohsen Imani |
WSDM | 1 |
| 2025 | Towards Trustworthy, Efficient, and Scalable Machine LearningabstractThroughout the development of machine learning, researchers have increasingly focused on the challenges of trustworthiness, efficiency, and scalability. Our research specifically addresses these critical aspects. Yezi Liu |
AAAI | 1 |
| 2025 | Enabling Group Fairness in Machine Unlearning via Distribution Correction
Yezi Liu, Yanning Shen |
CIKM | 1 |
| 2025 | DGExplainer: Explaining Dynamic Graph Neural Networks via Relevance Back-propagationabstractDynamic graph neural networks (dynamic GNNs) have demonstrated remarkable effectiveness in analyzing time-varying graph-structured data. However, their black-box nature often hinders users from understanding their predictions, which can limit their applications. In recent years, there has been a surge in research aimed at explaining GNNs, but most studies have focused on static graphs, leaving the explanation of dynamic GNNs relatively unexplored. Explaining dynamic GNNs presents a unique challenge due to their complex spatial and temporal structures. As a result, existing approaches designed for explaining static graphs are not directly applicable to dynamic graphs because they ignore temporal dependencies among graph snapshots. To address this issue, we propose DGExplainer, which offers a reliable explanation of dynamic GNN predictions. DGExplainer utilizes the relevance back-propagation technique both time-wise and layer-wise. Specifically, it incorporates temporal information by computing the relevance of node representations along the inverse of the time evolution. Additionally, for each time step, it calculates layer-wise relevance from a graph-based module by redistributing the relevance of node representations along the back-propagation path. Quantitative and qualitative experimental results on six real-world datasets demonstrate the effectiveness of DGExplainer in identifying important nodes for link prediction and node regression in dynamic GNNs. Yezi Liu, Jiaxuan Xie, Yanning Shen |
IJCAI | 1 |
| 2025 | LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning SegmentationabstractLarge Vision Language Models (LVLMs) have been widely adopted to guide vision foundation models in performing reasoning segmentation tasks, achieving impressive performance. However, the substantial computational overhead associated with LVLMs presents a new challenge. The primary source of this computational cost arises from processing hundreds of image tokens. Therefore, an effective strategy to mitigate such overhead is to reduce the number of image tokens-a process known as image token pruning. Previous studies on image token pruning for LVLMs have primarily focused on high-level visual understanding tasks, such as visual question answering and image captioning. In contrast, guiding vision foundation models to generate accurate visual masks based on textual queries demands precise semantic and spatial reasoning capabilities. Consequently, pruning methods must carefully control individual image tokens throughout the LVLM reasoning process. Our empirical analysis reveals that existing methods struggle to adequately balance reductions in computational overhead with the necessity to maintain high segmentation accuracy. In this work, we propose LVLM_CSP, a novel training-free visual token pruning method specifically designed for LVLM-based reasoning segmentation tasks. LVLM_CSP consists of three stages: clustering, scattering, and pruning. Initially, the LVLM performs coarse-grained visual reasoning using a subset of selected image tokens. Next, fine-grained reasoning is conducted, and finally, most visual tokens are pruned in the last stage. Extensive experiments demonstrate that LVLM_CSP achieves a 65% reduction in image token inference FLOPs with virtually no accuracy degradation, and a 70% reduction with only a minor 1% drop in accuracy on the 7B LVLM. Hanning Chen, Yang Ni 0001, Wenjun Huang 0001, Hyunwoo Oh, Yezi Liu, Tamoghno Das, Mohsen Imani |
ACM Multimedia | 5 |
| 2025 | VLTP: Vision-Language Guided Token Pruning for Task-Oriented SegmentationabstractVision Transformers (ViTs) have emerged as the backbone of many segmentation models, consistently achieving state-of-the-art (SOTA) performance. However, their success comes at a significant computational cost. Image token pruning is one of the most effective strategies to address this complexity. However, previous approaches fall short when applied to more complex task-oriented segmentation (TOS), where the class of each image patch is not predefined but dependent on the specific input task. This work introduces the Vision Language Guided Token Pruning (VLTP), a novel token pruning mechanism that can accelerate ViT-based segmentation models, particularly for TOS guided by multi-modal large language model (MLLM). We argue that ViT does not need to process every image token through all of its layers—only the tokens related to reasoning tasks are necessary. We design a new pruning decoder to take both image tokens and vision-language guidance as input to predict the relevance of each image token to the task. Only image tokens with high relevance are passed to deeper layers of the ViT. Experiments show that the VLTP framework reduces the computational costs of ViT by approximately 25% without performance degradation and by around 40% with only a 1% performance drop. The code associated with this study can be found at this URL. Hanning Chen, Yang Ni 0001, Wenjun Huang 0001, Yezi Liu, Sungheon Jeong 0001, Fei Wen 0003, Nathaniel D. Bastian, Hugo Latapie, Mohsen Imani |
WACV | 4 |
| 2025 | Recoverable Anonymization for Pose Estimation: A Privacy-Enhancing ApproachabstractHuman pose estimation (HPE) is crucial for various applications. However, deploying HPE algorithms in surveillance contexts raises significant privacy concerns due to the potential leakage of sensitive personal information (SPI) such as facial features, and ethnicity. Existing privacy-enhancing methods often compromise either privacy or performance, or they require costly additional modalities. We propose a novel privacy-enhancing system that generates privacy-enhanced portraits while maintaining high HPE performance. Our key innovations include the reversible recovery of SPI for authorized personnel and the preservation of contextual information. By Jointly optimizing a privacy-enhancing module, a privacy recovery module, and a pose estimator, our system ensures robust privacy protection, efficient SPI recovery, and high-performance HPE. Experimental results demonstrate the system's robust performance in privacy enhancement, SPI recovery, and HPE. The code associated with this study can be found at this URL. Wenjun Huang 0001, Yang Ni 0001, Arghavan Rezvani, Sungheon Jeong 0001, Hanning Chen, Yezi Liu, Fei Wen 0003, Mohsen Imani |
WACV | 6 |
| 2023 | FairGraph: Automated Graph Debiasing with Gradient MatchingabstractAs a prevalence data structure in the real world, graphs have found extensive applications ranging from modeling social networks to molecules. However, the existence of diverse biases within graphs gives rise to unfair representations learned by graph neural networks (GNNs). Addressing this issue has typically been approached from a modeling perspective, which not only compromises the integrity of the model structure but also entails additional effort and cost for retraining model parameters when the architecture changes. In this study, we adopt a data-centric standpoint to tackle the problem of fairness, focusing on graph debiasing for Graph Neural Networks. Our specific objective is to eliminate various biases from the input graph by generating a fair synthetic graph. By training GNNs on this fair graph, we aim to achieve an optimal accuracy-fairness trade-off. To this end, we propose FairGraph, which approaches the graph debiasing problem by mimicking the GNN training trajectory of the input graph through an optimization process involving a gradient-matching loss and fairness constraints. Through extensive experiments conducted on three benchmark datasets, we demonstrate the effectiveness of FairGraph and its ability to automatedly generate fair graphs that are transferable across different GNN architectures. Yezi Liu |
CIKM | 1 |
| 2022 | RES: An Interpretable Replicability Estimation System for Research Publications
Zhuoer Wang, Qizhang Feng, Mohinish Chatterjee, Xing Zhao 0003, Yezi Liu, Yuening Li, Frank M. Shipman III, Xia Ben Hu, James Caverlee |
AAAI | 5 |
| 2022 | Contrastive Knowledge Graph Error DetectionabstractKnowledge Graph (KG) errors introduce non-negligible noise, severely affecting KG-related downstream tasks. Detecting errors in KGs is challenging since the patterns of errors are unknown and diverse, while ground-truth labels are rare or even unavailable. A traditional solution is to construct logical rules to verify triples, but it is not generalizable since different KGs have distinct rules with domain knowledge involved. Recent studies focus on designing tailored detectors or ranking triples based on KG embedding loss. However, they all rely on negative samples for training, which are generated by randomly replacing the head or tail entity of existing triples. Such a negative sampling strategy is not enough for prototyping practical KG errors, e.g., (Bruce_Lee, place_of_birth, China), in which the three elements are often relevant, although mismatched. We desire a more effective unsupervised learning mechanism tailored for KG error detection. To this end, we propose a novel framework - ContrAstive knowledge Graph Error Detection (CAGED). It introduces contrastive learning into KG learning and provides a novel way of modeling KG. Instead of following the traditional setting, i.e., considering entities as nodes and relations as semantic edges, CAGED augments a KG into different hyper-views, by regarding each relational triple as a node. After joint training with KG embedding and contrastive learning loss, CAGED assesses the trustworthiness of each triple based on two learning signals, i.e., the consistency of triple representations across multi-views and the self-consistency within the triple. Extensive experiments on three real-world KGs show that CAGED outperforms state-of-the-art methods in KG error detection. Our codes and datasets are available at https://github.com/Qing145/CAGED.git. Qinggang Zhang, Junnan Dong, Keyu Duan, Xiao Huang 0001, Yezi Liu, Linchuan Xu |
CIKM | 5 |
| 2018 | Semi-supervised Multi-label Dimensionality Reduction via Low Rank Representation
Yezi Liu |
ICONIP (3) | 1 |