Yi Li 0018

dblp:59/871-18 · DBLP profile ↗
← Back
28ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0002-2856-7290ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 GSR-MIL: A Geometric Structure-Regularized Multiple Instance Learning Framework for Whole-Slide Image Classification
Hanfu Hu, Shuoyan Wang, Yunyun Xiao, Zhirui Fang, Yi Li 0018
ICIC (21)7
2026 Multi-Source Unsupervised Graph Domain Adaptation via Concise Propagation-Transformation Pipeline
abstract
Unsupervised graph domain adaptation (UGDA) aims to transfer knowledge from a labeled source graph to an unlabeled target graph, addressing the performance degradation caused by distributional shifts in node attributes and graph structures across domains. Despite recent progress, existing UGDA approaches still face two key challenges: (C1) Data-level: Most methods rely on a single source domain, overlooking the complementary knowledge that could be leveraged from multiple sources. (C2) Model-level: Many UGDA models emphasize complex, handcrafted Graph neural network (GNN) architectures, while simpler yet effective designs with propagation (P) & transformation (T) pipeline remain underexplored. To address these challenges, in this paper, we propose a novel approach, which leverages Concise Propagation–Transformation pipeline for multi-source unsupervised Graph Domain Adaptation, dubbed as CPT-GDA, to better capture complementary knowledge from multiple sources in an efficient manner. Specifically, the proposed CPT-GDA adopts a dual-branch GNN architecture with different depths of propagation but the same P-T patterns, which enables the model to efficiently learn node representations to mitigate domain discrepancy. Meanwhile, to facilitate effective knowledge transfer across graphs, we derive three optimization objectives: (1) the classifier loss to learn discriminative representations; (2) the alignment loss weighted by the graph Wasserstein distance to align the structure and feature distribution; and (3) the pseudo-label loss to refine target node representations. Extensive experiments on real-world datasets confirm that the proposed method outperforms recent state-of-the-art baselines, demonstrating its effectiveness.
Yi Li 0018, Xin Zheng 0008, Junyang Chen 0001, Yanqing Guo, Alan Wee-Chung Liew, Shirui Pan
WWW2
2026 Siameformer: Towards a robust source camera identification method on lossy images
Bo Wang 0024, Jiaqi Chi, Weiming Zheng, Fei Wei, Yi Li 0018
Knowl. Based Syst.5
2025 PAFedMIS: Personalized Asynchronous Federated Learning for Medical Image Segmentation
abstract
As privacy protection gains momentum, federated learning has emerged as a cutting-edge approach in medical image analysis. However, the intricacies of medical image segmentation task have led to a dearth of research in this domain, with existing studies falling short in tackling two pivotal challenges: The traditional model with the uniform global model underperforms for certain clients due to the heterogeneity and non-Independent Identically Distributed(non-IID) data across medical institutions. And the communication between the server and clients often incurs significant time costs. This paper introduces a novel Personalized Asynchronous Federated learning for Medical Image Segmentation model, dubbed PAFedMIS, to mitigate the negative impact of the heterogeneous data and fully utilized the waiting time, in medical image segmentation. Comprehensive experiments on ISIC2018 demonstrate the enhanced model accuracy and training efficiency of PAFedMIS.
Yi Li 0018, Yue Hua, Xin Zheng 0008, Yanqing Guo, Bo Wang 0024
ICASSP1
2025 Long-Tailed Federated Learning with Fixed Classifier
abstract
Federated learning (FL) is a machine learning approach where multiple participants train a model together without sharing their privacy data. The challenge of long-tail data heterogeneity in FL causes imbalanced data distribution among clients and significant disparities in data quantities across different classes. To tackle the long-tail distribution in FL, in this paper, we draw inspiration from the equiangular tight frame to establish a fixed balanced classifier, enhancing the model’s generalization ability for classes with fewer instances. Therefore, feature extractors and a loss aligned with fixed classifiers are designed to address the long-tail distribution in FL. The algorithm also utilizes Sample Standardization and Batch Normalization to standardize the feature space of samples, preventing the impact of the magnitude difference of features across different classes on the prediction probability. In addition, we also propose a logit adjustment loss under a mixed low-temperature setting. Experiments demonstrate that our algorithm outperforms state-of-the-art methods in FL settings with long-tail distribution.
Yi Li 0018, Xin Zheng 0008, Haiyan Fu, Yanqing Guo
ICME1
2025 Efficient Hi-Fi Style Transfer via Statistical Attention and Modulation
abstract
Style transfer is a challenging task in computer vision, aiming to blend the stylistic features of one image with the content of another while preserving the content details. Traditional methods often face challenges in terms of computational efficiency and fine-grained content preservation. In this paper, we propose a novel feature modulation mechanism based on parameterized normalization, where the modulation parameters for content and style features are learned using a dual convolution network (BiConv). These parameters adjust the mean and standard deviation of the features, improving both the stability and quality of the style transfer process. To achieve fast inference, we introduce an efficient acceleration technique by leveraging a row and column weighted attention matrix. In addition, we incorporate a contrastive learning scheme to align the local features of the content and the stylized images, improving the fidelity of the generated output. Experimental results demonstrate that our method significantly improves the inference speed and the quality of style transfer while preserving content details, outperforming existing approaches based on both convolution and diffusion.
Zhirui Fang, Yi Li 0018, Chengyan Li, Yanqing Guo
IJCAI2
2025 Trustworthy forgery detection with causal inference
Junxian Duan, Fan Ji, Yi Li 0018, Ran He 0001
Sci. China Inf. Sci.3
2024 Invisible Intruders: Label-Consistent Backdoor Attack Using Re-Parameterized Noise Trigger
abstract
Aremarkable number of backdoor attack methods have been proposed in the literature on deep neural networks (DNNs). However, it hasn't been sufficiently addressed in the existing methods of achieving true senseless backdoor attacks that are visually invisible and label-consistent. In this paper, we propose a new backdoor attack method where the labels of the backdoor images are perfectly aligned with their content, ensuring label consistency. Additionally, the backdoor trigger is meticulously designed, allowing the attack to evade DNN model checks and human inspection. Our approach employs an auto-encoder (AE) to conduct representation learning of benign images and interferes with salient classification features to increase the dependence of backdoor image classification on backdoor triggers. To ensure visual invisibility, we implement a method inspired by image steganography that embeds trigger patterns into the image using the DNN and enable sample-specific backdoor triggers. We conduct comprehensive experiments on multiple benchmark datasets and network architectures to verify the effectiveness of our proposed method under the metric of attack success rate and invisibility. The results also demonstrate satisfactory performance against a variety of defense methods.
Bo Wang 0024, Fei Wei, Yi Li 0018, Wei Wang 0025
IEEE Trans. Multim.4
2023 A Compact Transformer for Adaptive Style Transfer
abstract
Due to the limitation of spatial receptive field, it is challenging for CNN-based style transfer methods to capture rich and long-range semantic concepts in artworks. Though the transformer provides a fresh solution by considering long-range dependencies, it suffers from the heavy burdens of parameter scale and computation cost especially in vision tasks. In this paper, we design a compact transformer architecture AdaFormer to address the problem. The model scale shrinks about 20% compared to state-of-the-art transformer for style transfer. Furthermore, we explore the adaptive style transfer by letting the content to select the detailed style element automatically and adaptively, which encourages the output to be both appealing and reasonable. We evaluate AdaFormer comprehensively in the experiments and the results have shown the effectiveness and superiority of our approach compared to existing artistic methods. Diverse plausible stylized images are obtained with better content preservation, more convincing style sfumato and lower computation complexity.
Yi Li 0018, Haiyan Fu, Xiangyang Luo 0001, Yanqing Guo
ICME1
2023 Federating Hashing Networks Adaptively for Privacy-Preserving Retrieval
abstract
With the rise of neural networks, many deep hashing networks have been successfully trained on the basis of large-scale data. However, the conventional learning process has received increasing challenges from the data privacy concerns and the decentralized storage status, especially in sensitive scenarios like surveillance retrieval. Further considering the probable different distributions of the decentralized data, in this paper, we present a collaborative hashing paradigm FedA-Hash (Federating Adapted Hashing nets) to produce personalized hashing models for the clients without exchanging their local data. To this end, the bilateral knowledge is blended gradually during the learning process between the aggregated global model and the local hashing model, instead of replacing the local model with the global model directly. Extensive experiments are conducted on representative hashing networks, involving tasks as image retrieval and person re-identification. The results show that FedA-Hash significantly enables the collaborated performance among different clients.
Yi Li 0018, Meihua Yu, Haiyan Fu, Yanqing Guo
ICME1
2022 Multi-Attribute Controlled Text Generation with Contrastive-Generator and External-Discriminator
abstract
Though existing researches have achieved impressive results in controlled text generation, they focus mainly on single-attribute control. However, in applications like automatic comments, the topic and sentiment need to be controlled simultaneously. In this work, we propose a new framework for multi-attribute controlled text generation. To achieve this, we design a contrastive-generator that can effectively generate texts with more attributes. In order to increase the convergence of the text on the desired attributes, we adopt an external-discriminator to distinguish whether the generated text holds the desired attributes. Moreover, we propose top-n weighted decoding to further improve the relevance of texts to attributes. Automated evaluations and human evaluations show that our framework achieves remarkable controllability in multi-attribute generation while keeping the text fluent and diverse. It also yields promising performance on zero-shot generation.
Guisheng Liu, Yi Li 0018, Yanqing Guo, Xiangyang Luo 0001, Bo Wang 0024
COLING2
2022 Artistic Style Discovery with Independent Components
abstract
Style transfer has been well studied in recent years with excellent performance processed. While existing methods usually choose CNNs as the powerful tool to accomplish superb stylization, less attention was paid to the latent style space. Rare exploration of underlying dimensions results in the poor style controllability and the limited practical application. In this work, we rethink the internal meaning of style features, further proposing a novel unsupervised algorithm for style discovery and achieving personalized manip-ulation. In particular, we take a closer look into the mechanism of style transfer and obtain different artistic style components from the latent space consisting of different style features. Then fresh styles can be generated by linear combination according to various style components. Experimental results have shown that our approach is superb in 1) restylizing the original output with the diverse artistic styles discovered from the latent space while keeping the content unchanged, and 2) being generic and compatible for various style transfer methods. Our code is available in this page: https://github.com/Shelsin/ArtIns.
Yi Li 0018, Huaibo Huang, Haiyan Fu, Wanwan Wang, Yanqing Guo
CVPR2
2022 Open-Set source camera identification based on envelope of data clustering optimization (EDCO)
Bo Wang 0024, Yue Wang 0132, Jiayao Hou, Yi Li 0018, Yanqing Guo
Comput. Secur.4
2022 Cross-Spectral Iris Recognition by Learning Device-Specific Band
abstract
Cross-spectral recognition is still an open challenge in iris recognition. In cross-spectral iris recognition, there exist distinct device-specific bands between near-infrared (NIR) and visible (VIS) images, resulting in the distribution gap between samples from different spectra and thus severe degradation in recognition performance. To tackle this problem, we propose a new cross-spectral iris recognition method to learn spectral-invariant features by estimating device-specific bands. In the proposed method,GaborTridentNetwork (GTN) first utilizes the Gabor function’s priors to perceive iris textures under different spectra, and then codes the device-specific band as the residual component to assist the generation of spectral-invariant features. By investigating the device-specific band, GTN effectively reduces the impact of device-specific bands on identity features. Besides, we make three efforts to further reduce the distribution gap. First,SpectralAdversarialNetwork (SAN) adopts a class-level adversarial strategy to align feature distributions. Second,Sample-Anchor (SA) loss upgrades triplet loss by pulling samples to their class center and pushing away from other class centers. Third, we develop a higher-order alignment loss to measures the distribution gap according to space bases and distribution shapes. Extensive experiments on five iris datasets demonstrate the efficacy of our proposed method for cross-spectral iris recognition.
Jianze Wei, Yunlong Wang 0003, Yi Li 0018, Ran He 0001, Zhenan Sun
IEEE Trans. Circuits Syst. Video Technol.3
2021 Coupled adversarial learning for semi-supervised heterogeneous face recognition
Ran He 0001, Yi Li 0018, Xiang Wu 0001, Lingxiao Song, Zhenhua Chai, Xiaolin Wei
Pattern Recognit.2
2021 Adversarial Analysis for Source Camera Identification
abstract
Recent studies highlight the vulnerability of convolutional neural networks (CNNs) to adversarial attacks, which also calls into question the reliability of forensic methods. Existing adversarial attacks generate one-to-one noise, which means these methods have not learned the fingerprint information. Therefore, we introduce two powerful attacks, fingerprint copy-move attack, and joint feature-based auto-learning attack. To validate the performance of attack methods, we move a step ahead and introduce the higher possible defense mechanism relation mismatch. which expands the characterization differences of classifiers in the same classification network. Extensive experiments show that relation mismatch is superior in recognizing adversarial examples and prove that the proposed fingerprint-based attacks are more powerful. Both proposed attacks show excellent attack transferability to unknown samples. The Pytorch® implementations of these methods can download from an open-source GitHub projecthttps://github.com/Dlut-lab-zmn/Source-attack.
Bo Wang 0024, Mengnan Zhao 0001, Wei Wang 0025, Xiaorui Dai, Yi Li 0018, Yanqing Guo
IEEE Trans. Circuits Syst. Video Technol.5
2020 Cross-Spectral Face Hallucination via Disentangling Independent Factors
abstract
The cross-sensor gap is one of the challenges that have aroused much research interests in Heterogeneous Face Recognition (HFR). Although recent methods have attempted to fill the gap with deep generative networks, most of them suffer from the inevitable misalignment between different face modalities. Instead of imaging sensors, the misalignment primarily results from facial geometric variations that are independent of the spectrum. Rather than building a monolithic but complex structure, this paper proposes a Pose Aligned Cross-spectral Hallucination (PACH) approach to disentangle the independent factors and deal with them in individual stages. In the first stage, an Unsupervised Face Alignment (UFA) module is designed to align the facial shapes of the near-infrared (NIR) images with those of the visible (VIS) images in a generative way, where UV maps are effectively utilized as the shape guidance. Thus the task of the second stage becomes spectrum translation with aligned paired data. We develop a Texture Prior Synthesis (TPS) module to achieve complexion control and consequently generate more realistic VIS images than existing methods. Experiments on three challenging NIR-VIS datasets verify the effectiveness of our approach in producing visually appealing images and achieving state-of-the-art performance in HFR.
Boyan Duan, Chaoyou Fu, Yi Li 0018, Xingguang Song, Ran He 0001
CVPR3
2020 Informative Sample Mining Network for Multi-domain Image-to-Image Translation
Jie Cao 0002, Huaibo Huang, Yi Li 0018, Ran He 0001, Zhenan Sun
ECCV (19)3
2020 Let's Play Music: Audio-Driven Performance Video Generation
abstract
We propose a new task named Audio-driven Performance Video Generation (APVG), which aims to synthesize the video of a person playing a certain instrument guided by a given music audio clip. It is a challenging task to generate the high-dimensional temporal consistent videos from low-dimensional audio modality. In this paper, we propose a multi-staged framework to generate realistic and synchronized performance video from given music. Firstly, we provide both global appearance and local spatial information by generating the coarse videos and keypoints of body and hands from a given music respectively. Then, we propose to transform the generated keypoints to heatmap via a differentiable space transformer, since the heatmap provides more spatial information but is harder to generate directly from audio. Finally, we propose a Structured Temporal UNet (STU) to extract both intra-frame structured information and interframe temporal consistency. They are obtained via graph-based structure module, and CNN-GRU based high-level temporal module respectively for final video generation. Comprehensive experiments validate the effectiveness of our proposed framework.
Yi Li 0018, Feixia Zhu, Aihua Zheng, Ran He 0001
ICPR2
2020 Arbitrary Talking Face Generation via Attentional Audio-Visual Coherence Learning
abstract
Talking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image. Most existing methods mainly focus on either disentangling the information in a single image or learning temporal information between frames. However, cross-modality coherence between audio and video information has not been well addressed during synthesis. In this paper, we propose a novel arbitrary talking face generation framework by discovering the audio-visual coherence via the proposed Asymmetric Mutual Information Estimator (AMIE). In addition, we propose a Dynamic Attention (DA) block by selectively focusing the lip area of the input image during the training stage, to further enhance lip synchronization. Experimental results on benchmark LRW dataset and GRID dataset transcend the state-of-the-art methods on prevalent metrics with robust high-resolution synthesizing on gender and pose variations.
Huaibo Huang, Yi Li 0018, Aihua Zheng, Ran He 0001
IJCAI3
2020 Disentangled Representation Learning of Makeup Portraits in the Wild
Yi Li 0018, Huaibo Huang, Jie Cao 0002, Ran He 0001, Tieniu Tan
Int. J. Comput. Vis.1
2020 A Survey of Deep Facial Attribute Analysis
Xin Zheng 0008, Yanqing Guo, Huaibo Huang, Yi Li 0018, Ran He 0001
Int. J. Comput. Vis.4
2019 Pose-preserving Cross Spectral Face Hallucination
abstract
To narrow the inherent sensing gap in heterogeneous face recognition (HFR), recent methods have resorted to generative models and explored the ?recognition via generation? framework. Even though, it remains a very challenging task to synthesize photo-realistic visible faces (VIS) from near-infrared (NIR) images especially when paired training data are unavailable. We present an approach to avert the data misalignment problem and faithfully preserve pose, expression and identity information during cross-spectral face hallucination. At the pixel level, we introduce an unsupervised attention mechanism to warping that is jointly learned with the generator to derive pixel-wise correspondence from unaligned data. At the image level, an auxiliary generator is employed to facilitate the learning of mapping from NIR to VIS domain. At the domain level, we first apply the mutual information constraint to explicitly measure the correlation between domains and thus benefit synthesis. Extensive experiments on three heterogeneous face datasets demonstrate that our approach not only outperforms current state-of-the-art HFR methods but also produce visually appealing results at a high resolution.
Junchi Yu, Jie Cao 0002, Yi Li 0018, Xiaofei Jia, Ran He 0001
IJCAI3
2019 Learning a bi-level adversarial network with global and local perception for makeup-invariant face verification
Yi Li 0018, Lingxiao Song, Xiang Wu 0001, Ran He 0001, Tieniu Tan
Pattern Recognit.1
2019 Joint CRF and Locality-Consistent Dictionary Learning for Semantic Segmentation
abstract
Semantic image segmentation can be accomplished by assigning a proper object category label to each meaningful region of an image. Beyond the original bottom-up models, the use of top-down categorization information has been applied to semantic segmentation to improve performance. An excellent example of such a top-down scheme is to integrate a Conditional Random Field (CRF) model with sparse dictionary learning. However, the existing solutions merely consider the discrimination of dictionaries to obtain better sparse codes, without considering the inherent data locality characteristics. In this paper, we explore such characteristics and propose a novel semantic segmentation framework based on an innovative CRF model with locality-consistent dictionary learning. In particular, we propose two new locality-consistent dictionary learning strategies by capturing the local consistencies in the feature space and the label space. In addition, we develop a joint dictionary and a CRF model parameter learning algorithm to seamlessly integrate the proposed locality-consistent dictionary learning strategies into the CRF model. Extensive experiments are conducted with two popular databases of different traits (i.e., Graz-02 and PASCAL-CONTEXT). The simulation results confirm the efficiency of the proposed scheme, especially when training data are limited.
Yi Li 0018, Yanqing Guo, Jun Guo 0008, Xiangwei Kong 0001, Qian Liu 0001
IEEE Trans. Multim.1
2018 Anti-Makeup: Learning A Bi-Level Adversarial Network for Makeup-Invariant Face Verification
abstract
Makeup is widely used to improve facial attractiveness and is well accepted by the public. However, different makeup styles will result in significant facial appearance changes. It remains a challenging problem to match makeup and non-makeup face images. This paper proposes a learning from generation approach for makeup-invariant face verification by introducing a bi-level adversarial network (BLAN). To alleviate the negative effects from makeup, we first generate non-makeup images from makeup ones, and then use the synthesized non-makeup images for further verification. Two adversarial networks in BLAN are integrated in an end-to-end deep network, with the one on pixel level for reconstructing appealing facial images and the other on feature level for preserving identity information. These two networks jointly reduce the sensing gap between makeup and non-makeup images. Moreover, we make the generator well constrained by incorporating multiple perceptual losses. Experimental results on three benchmark makeup face datasets demonstrate that our method achieves state-of-the-art verification accuracy across makeup status and can produce photo-realistic non-makeup face images.
Yi Li 0018, Lingxiao Song, Xiang Wu 0001, Ran He 0001, Tieniu Tan
AAAI1
2017 Image Piece Learning for Weakly Supervised Semantic Segmentation
abstract
The task of semantic segmentation is to infer a predefined category label for each pixel in the image. For most cases, image segmentation is established as a fully supervised task. These methods all built on the basis of having access to sufficient pixel-wise annotated samples for training. However, obtaining the satisfied ground truth is not only labor intensive but also time-consuming, which severely hinders the generality of these fully supervised methods. Instead of pixel-level ground truth, weakly supervised approaches learn their models from much less prior information, e.g., image-level annotation. In this paper, we propose a novel conditional random field (CRF) based framework for weakly supervised semantic segmentation. Enlightened by jigsaw puzzles, we start the approach with merging superpixels from an image into larger pieces by a newly designed strategy. Then pieces from all the training images are gathered and associated with appropriate semantic labels by CRF. Thus, the piece library is constructed, achieving remarkable universality and flexibility. In the case of testing, we compare the superpixels with image pieces in the library and assign them the labels that minimize the potential energy. In addition, the proposed framework is fit for domain adaption and obtains promising results, which is of great practical value. Extensive experimental results on PASCAL VOC 2007, MSRC-21, and VOC 2012 databases demonstrate that our framework outperforms or is comparable to state-of-the-art segmentation methods.
Yi Li 0018, Yanqing Guo, Yueying Kao, Ran He 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2015 Locality sensitive discriminative dictionary learning
abstract
Discriminative dictionary learning (DDL) has been applied to various pattern classification problems. Despite satisfying experimental results, most existing discriminative dictionary learning methods emphasize too much on the role of l0or l1-norm sparsity, while the underlying local structure of original data is totally ignored. In this paper, we present a novel dictionary learning method, named Locality Sensitive Discriminative Dictionary Learning (LSDDL), which combines basic dictionary learning scheme and locality relationship of original data which is propagated to the coding vectors. The learned discriminative dictionary can map the original data points into a new space in which the nearby points with the same label are close to each other while the nearby points with different labels are far apart. Experiments clearly show that our method has very competitive performance in contrast to previous discriminative dictionary learning methods.
Jun Guo 0008, Yanqing Guo, Yi Li 0018, Bo Wang 0024, Ming Li 0011
ICIP3