EDBT 2026 Demo / reviewers in the wild / expert
Li Sun 0012
dblp:57/2405-12
· DBLP profile ↗
25ranked-venue papers
1as first author
19since 2021 · last 2025
0000-0003-0950-4611ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 14 since 2021Artificial intelligence and machine learning · 15 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RobusTReID: Defending Vision Transformer for Robust Image ReIDabstractVision Transformer (ViT) achieves competitive results in person ReID, not only due to the powerful ability on feature representation, but also its resistance to attacks. However, they are vulnerable to specially designed attacks. To enhance their robustness, this paper proposes RobusTReID to defend the ViT-based model against perturbed images without obvious performance drop on clean data. The basic idea is to incorporate the adversarial co-training into ReID, which first disturbs pixels by minimizing adversarial loss in primary feature branch, then optimizes model by ReID task loss computed in all branches on clean and perturbed data. We separate the representation paths for clean and perturbed images. Particularly, a learnable [ADV] token and low-rank positional embeddings (PE) are incorporated to build the feature for perturbed image. Extensive experiments on several ReID datasets show that our method effectively increases the robustness of ReID model under different types of attacks. Tingting Xiao, Li Sun 0012, Qingli Li |
ICME | 3 |
| 2025 | DragLoRA: Online Optimization of LoRA Adapters for Drag-based Image Editing in Diffusion ModelabstractDrag-based editing within pretrained diffusion model provides a precise and flexible way to manipulate foreground objects. Traditional methods optimize the input feature obtained from DDIM inversion directly, adjusting them iteratively to guide handle points towards target locations. However, these approaches often suffer from limited accuracy due to the low representation ability of the feature in motion supervision, as well as inefficiencies caused by the large search space required for point tracking. To address these limitations, we present DragLoRA, a novel framework that integrates LoRA (Low-Rank Adaptation) adapters into the drag-based editing pipeline. To enhance the training of LoRA adapters, we introduce an additional denoising score distillation loss which regularizes the online model by aligning its output with that of the original model. Additionally, we improve the consistency of motion supervision by adapting the input features using the updated LoRA, giving a more stable and accurate input feature for subsequent operations. Building on this, we design an adaptive optimization scheme that dynamically toggles between two modes, prioritizing efficiency without compromising precision. Extensive experiments demonstrate that DragLoRA significantly enhances the control precision and computational efficiency for drag-based image editing. The Codes of DragLoRA are available at: https://github.com/Sylvie-X/DragLoRA. Siwei Xia, Li Sun 0012, Tiantian Sun, Qingli Li |
ICML | 2 |
| 2024 | OD-DETR: Online Distillation for Stabilizing Training of Detection Transformer
Shengjian Wu, Li Sun 0012, Qingli Li |
IJCAI | 2 |
| 2024 | Spatial-Temporal Traffic Prediction Model Based on Adaptive Graphs Fusion and Dual-Graph Collaborative ConvolutionabstractTraffic flow prediction is crucial for intelligent transportation systems (ITS). Traditional graph convolutional networks (GCNs) have limitations in handling road network data. These GCNs can only handle binary relationships between nodes and cannot effectively capture the dynamic and nonlinear spatial dependencies among multiple nodes in the road network. This paper proposes an innovative spatial-temporal traffic flow prediction model to address this challenge. Firstly, by introducing novel graphs fusion and hypergraph encoding module, a fused graph and its dual hypergraph are constructed to provide richer structural information for the model. Then, the model can learn complex relationships among multiple nodes through the collaborative convolution of GCN and hypergraph convolutional network (HGCN). To comprehensively capture the dynamic nature of traffic flow, we utilize a variant of Transformer combined with time encoding information to capture the periodicity in the data, thereby enhancing the model’s ability to recognize periodic traffic flow patterns. Comprehensive experimental results on two publicly accessible real-world traffic datasets demonstrate the superiority of our proposed model over state-of-the-art traffic prediction models. Song Qiu, Li Sun 0012, Dingding Han, Qingli Li, Mingsong Chen 0001 |
IJCNN | 3 |
| 2024 | DSA-SCGC: A Dual Self-Attention Mechanism based on Space-Channel Grouped Compression for Vehicle Re-IdentificationabstractVehicle re-identification (re-ID) has attracted significant attention within the computer vision community due to its wide-ranging applications in intelligent transportation systems and law enforcement. Nevertheless, this field faces considerable challenges owing to the high inter-class similarity and the large intra-class difference among vehicles. To address these challenges, this paper proposes a novel network incorporating a dual self-attention mechanism based on a space-channel grouped compression operation (DSA-SCGC). This innovative approach combines channel and spatial self-attention mechanisms to selectively enhance pivotal channel features and spatial local details while minimizing attention toward backgrounds and occlusions commonly encountered in real-world scenarios. Moreover, to address the issue of spatial information loss in channel attention, we propose a space-channel grouped compression (SCGC) operation that effectively compresses spatial information into channels, thereby significantly preserving spatial information. Comprehensive experiments conducted on the VeRi-776 and VehicleID datasets validate the superiority of our proposed DSA-SCGC model over the existing state-of-the-art vehicle re-identification methods. Yuejun Jiao, Song Qiu, Li Sun 0012, Dingding Han, Qingli Li, Mingsong Chen 0001 |
IJCNN | 3 |
| 2024 | RefineStyle: Dynamic Convolution Refinement for StyleGAN
Siwei Xia, Xueqi Hu, Li Sun 0012, Qingli Li |
PRCV (9) | 3 |
| 2023 | CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text LabelsabstractPre-trained vision-language models like CLIP have recently shown superior performances on various downstream tasks, including image classification and segmentation. However, in fine-grained image re-identification (ReID), the labels are indexes, lacking concrete text descriptions. Therefore, it remains to be determined how such models could be applied to these tasks. This paper first finds out that simply fine-tuning the visual model initialized by the image encoder in CLIP, has already obtained competitive performances in various ReID tasks. Then we propose a two-stage strategy to facilitate a better visual representation. The key idea is to fully exploit the cross-modal description ability in CLIP through a set of learnable text tokens for each ID and give them to the text encoder to form ambiguous descriptions. In the first training stage, image and text encoders from CLIP keep fixed, and only the text tokens are optimized from scratch by the contrastive loss computed within a batch. In the second stage, the ID-specific text tokens and their encoder become static, providing constraints for fine-tuning the image encoder. With the help of the designed loss in the downstream task, the image encoder is able to represent data as vectors in the feature embedding accurately. The effectiveness of the proposed strategy is validated on several datasets for the person or vehicle ReID tasks. Code is available at https://github.com/Syliz517/CLIP-ReID. Li Sun 0012, Qingli Li |
AAAI | 2 |
| 2023 | RecursiveDet: End-to-End Region-based Recursive Object DetectionabstractEnd-to-end region-based object detectors like Sparse R-CNN usually have multiple cascade bounding box decoding stages, which refine the current predictions according to their previous results. Model parameters within each stage are independent, evolving a huge cost. In this paper, we find the general setting of decoding stages is actually redundant. By simply sharing parameters and making a recursive decoder, the detector already obtains a significant improvement. The recursive decoder can be further enhanced by positional encoding (PE) of the proposal box, which makes it aware of the exact locations and sizes of input bounding boxes, thus becoming adaptive to proposals from different stages during the recursion. Moreover, we also design centerness-based PE to distinguish the RoI feature element and dynamic convolution kernels at different positions within the bounding box. To validate the effectiveness of the proposed method, we conduct intensive ablations and build the full model on three recent mainstream region-based detectors. The RecusiveDet is able to achieve obvious performance boosts with even fewer model parameters and slightly increased computation cost. Codes are available at https://github.com/bravezzzzzz/RecursiveDet. Li Sun 0012, Qingli Li |
ICCV | 2 |
| 2022 | IoU-Enhanced Attention for End-to-End Task Specific Object Detection
Shengjian Wu, Li Sun 0012, Qingli Li |
ACCV (5) | 3 |
| 2022 | Style Transformer for Image Inversion and EditingabstractExisting GAN inversion methods fail to provide latent codes for reliable reconstruction and flexible editing simultaneously. This paper presents a transformer-based image inversion and editing model for pretrained StyleGAN which is not only with less distortions, but also of high quality and flexibility for editing. The proposed model employs a CNN encoder to provide multi-scale image features as keys and values. Meanwhile it regards the style code to be determined for different layers of the generator as queries. It first initializes query tokens as learnable parameters and maps them into W+ space. Then the multi-stage alternate self-and cross-attention are utilized, updating queries with the purpose of inverting the input by the generator. Moreover, based on the inverted code, we investigate the reference-and label-based attribute editing through a pretrained latent classifier, and achieve flexible image-to-image translation with high quality results. Extensive experiments are carried out, showing better performances on both inversion and editing tasks within StyleGAN. Codes are available at https://github.com/sapphire497/style-transformer. Xueqi Hu, Qiusheng Huang, Zhengyi Shi, Changxin Gao, Li Sun 0012, Qingli Li |
CVPR | 6 |
| 2022 | QS-Attn: Query-Selected Attention for Contrastive Learning in I2I TranslationabstractUnpaired image-to-image (I2I) translation often requires to maximize the mutual information between the source and the translated images across different domains, which is critical for the generator to keep the source content and prevent it from unnecessary modifications. The self-supervised contrastive learning has already been successfully applied in the I2I. By constraining features from the same location to be closer than those from different ones, it implicitly ensures the result to take content from the source. However, previous work uses the features from random locations to impose the constraint, which may not be appropriate since some locations contain less information of source domain. Moreover, the feature itself does not reflect the relation with others. This paper deals with these problems by intentionally selecting significant anchor points for contrastive learning. We design a query-selected attention (QS-Attn) module, which compares feature distances in the source domain, giving an attention matrix with a probability distribution in each row. Then we select queries according to their measurement of significance, computed from the distribution. The selected ones are regarded as anchors for contrastive loss. At the same time, the reduced attention matrix is employed to route features in both domains, so that source relations maintain in the synthesis. We validate our proposed method in three different I2I datasets, showing that it increases the image quality with-out adding learnable parameters. Codes are available at https://github.com/sapphire497/query-selected-attention. Xueqi Hu, Xinyue Zhou, Qiusheng Huang, Zhengyi Shi, Li Sun 0012, Qingli Li |
CVPR | 5 |
| 2022 | Cross Attention Based Style Distribution for Controllable Person Image Synthesis
Xinyue Zhou, Mingyu Yin, Li Sun 0012, Changxin Gao, Qingli Li |
ECCV (15) | 4 |
| 2022 | Cross-Stage Class-Specific Attention for Image Semantic Segmentation
Zhengyi Shi, Li Sun 0012, Qingli Li |
PRCV (4) | 2 |
| 2022 | A Vehicle Re-ID Algorithm Based on Channel Correlation Self-attention and Lstm Local Information Loss
Tiantian Qi, Song Qiu, Li Sun 0012, Mingsong Chen 0001, Yue Lyu |
PRICAI (2) | 3 |
| 2021 | ID-Unet: Iterative Soft and Hard Deformation for View SynthesisabstractView synthesis is usually done by an autoencoder, in which the encoder maps a source view image into a latent content code, and the decoder transforms it into a target view image according to the condition. However, the source contents are often not well kept in this setting, which leads to unnecessary changes during the view translation. Al-though adding skipped connections, like Unet, alleviates the problem, but it often causes the failure on the view conformity. This paper proposes a new architecture by performing the source-to-target deformation in an iterative way. Instead of simply incorporating the features from multiple layers of the encoder, we design soft and hard deformation modules, which warp the encoder features to the target view at different resolutions, and give results to the decoder to complement the details. Particularly, the current warping flow is not only used to align the feature of the same resolution, but also as an approximation to coarsely deform the high resolution feature. Then the residual flow is estimated and applied in the high resolution, so that the deformation is built up in the coarse-to-fine fashion. To better constrain the model, we synthesize a rough target view image based on the intermediate flows and their warped features. The extensive ablation studies and the final results on two different data sets show the effectiveness of the proposed model. https://github.com/MingyuY/Iterative-view-synthesis Mingyu Yin, Li Sun 0012, Qingli Li |
CVPR | 2 |
| 2021 | Progressive Multi-Stage Feature Mix for Person Re-IdentificationabstractImage features from a small local region often give strong evidence in person re-identification task. However, CNN suffers from paying too much attention on the most salient local areas, thus ignoring other discriminative clues, e.g., hair, shoes or logos on clothes. In this work, we propose a Progressive Multi-stage feature Mix network (PMM), which enables the model to find out the more precise and diverse features in a progressive manner. Specifically, (i) to enforce the model to look for different clues in the image, we adopt a multi-stage classifier and expect that the model is able to focus on a complementary region in each stage. (ii) we propose an Attentive feature Hard-Mix (A-Hard-Mix) to replace the salient feature blocks by the negative example in the current batch, whose label is different from the current sample. (iii) extensive experiments have been carried out on reID datasets such as the Market-1501, DukeMTMC-reID and CUHK03, showing that the proposed method can boost the re-identification performance significantly. Source code1has been released. Binyu He, Li Sun 0012, Qingli Li |
ICASSP | 3 |
| 2021 | Bridging the Gap between Label- and Reference-based Synthesis in Multi-attribute Image-to-Image TranslationabstractThe image-to-image translation (I2TT) model takes a target label or a reference image as the input, and changes a source into the specified target domain. The two types of synthesis, either label- or reference-based, have substantial differences. Particularly, the label-based synthesis reflects the common characteristics of the target domain, and the reference-based shows the specific style similar to the reference. This paper intends to bridge the gap between them in the task of multi-attribute I2TT. We design the label- and reference-based encoding modules (LEM and REM) to compare the domain differences. They first transfer the source image and target label (or reference) into a common embedding space, by providing the opposite directions through the attribute difference vector. Then the two embeddings are simply fused together to form the latent code Srand(or Sref), reflecting the domain style differences, which is injected into each layer of the generator by SPADE. To link LEM and REM, so that two types of results benefit each other, we encourage the two latent codes to be close, and set up the cycle consistency between the forward and backward translations on them. Moreover, the interpolation between the Srandand Srefis also used to synthesize an extra image. Experiments show that label- and reference-based synthesis are indeed mutually promoted, so that we can have the diverse results from LEM, and high quality results with the similar style of the reference. Code will be available at https://github.com/huangqiusheng/BridgeGAN. Qiusheng Huang, Zhilin Zheng, Xueqi Hu, Li Sun 0012, Qingli Li |
ICCV | 4 |
| 2021 | Attribute-specific Control Units in StyleGAN for Fine-grained Image ManipulationabstractImage manipulation with StyleGAN has been an increasing concern in recent years. Recent works have achieved tremendous success in analyzing several semantic latent spaces to edit the attributes of the generated images. However, due to the limited semantic and spatial manipulation precision in these latent spaces, the existing endeavors are defeated in fine-grained StyleGAN image manipulation, i.e., local attribute translation. To address this issue, we discover attribute-specific control units, which consist of multiple channels of feature maps and modulation styles. Specifically, we collaboratively manipulate the modulation style channels and feature maps in control units rather than individual ones to obtain the semantic and spatial disentangled controls. Furthermore, we propose a simple yet effective method to detect the attribute-specific control units. We move the modulation style along a specific sparse direction vector and replace the filter-wise styles used to compute the feature maps to manipulate these control units. We evaluate our proposed method in various face attribute manipulation tasks. Extensive qualitative and quantitative results demonstrate that our proposed method performs favorably against the state-of-the-art methods. The manipulation results of real images further show the effectiveness of our method. Rui Wang 0099, Gang Yu 0002, Li Sun 0012, Changqian Yu, Changxin Gao, Nong Sang |
ACM Multimedia | 4 |
| 2021 | Identification of Melanoma From Hyperspectral Pathology Image Using 3D Convolutional NetworksabstractSkin biopsy histopathological analysis is one of the primary methods used for pathologists to assess the presence and deterioration of melanoma in clinical. A comprehensive and reliable pathological analysis is the result of correctly segmented melanoma and its interaction with benign tissues, and therefore providing accurate therapy. In this study, we applied the deep convolution network on the hyperspectral pathology images to perform the segmentation of melanoma. To make the best use of spectral properties of three dimensional hyperspectral data, we proposed a 3D fully convolutional network named Hyper-net to segment melanoma from hyperspectral pathology images. In order to enhance the sensitivity of the model, we made a specific modification to the loss function with caution of false negative in diagnosis. The performance of Hyper-net surpassed the 2D model with the accuracy over 92%. The false negative rate decreased by nearly 66% using Hyper-net with the modified loss function. These findings demonstrated the ability of the Hyper-net for assisting pathologists in diagnosis of melanoma based on hyperspectral pathology images. Qian Wang 0046, Li Sun 0012, Yan Wang 0033, Mei Zhou, Menghan Hu, Ying Wen 0003, Qingli Li |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Novel View Synthesis on Unpaired Data by Conditional Deformable Variational Auto-Encoder
Mingyu Yin, Li Sun 0012, Qingli Li |
ECCV (28) | 2 |
| 2020 | Disentangling The Spatial Structure And Style In Conditional VAEabstractThis paper proposes a structure in conditional variation autoencoder (cVAE) to disentangle the latent vector into a spatial structure and a style code, complementary to each other, with the one $( z_{s})$ being label relevant and the other $( z_{u})$ irrelevant. Different from traditional cVAE, our network maps the condition label into its relevant code zsthrough a separated module. Depending on whether the label directly relates to the image spatial structure or not, zsoutput from the condition mapping module is used either as the style code with the two spatial dimension of $1 \times 1$, or as the spatial structure code with a single channel. Based on the input image and its corresponding zs, the encoder provides the posterior distribution close to a common prior regardless of its label, thus zusampled from it becomes label irrelevant. The decoder employs zsand zuby two typical adaptive normalization modules to reconstruct the input image. Results on two datasets with different types of labels show the effectiveness of our method. Li Sun 0012, Zhilin Zheng, Qingli Li |
ICIP | 2 |
| 2018 | Position-Squeeze and Excitation Block for Facial Attribute Analysis
Wanxia Shen, Li Sun 0012, Qingli Li |
BMVC | 3 |
| 2018 | Investigation in Spatial-Temporal Domain for Face Spoof DetectionabstractThis paper focuses on face spoofing detection using video. The purpose is to find out the best scheme for this task in the end-to-end learning manner. We investigate 4 different types of structure to fully exploit the raw data in its spatial-temporal domain, which are the pure CNN, CNN with 3D convolution, CNN+LSTM and CNN+Conv-LSTM. Moreover, another stream built on optical flow is also used, and with a proper fusion method, it can improve the accuracy. In experiments, we compare schemes on the raw data in single stream and fusion methods with optical flow in two streams. The performance are not only given within each dataset, but also measured across different datsets, which is crucial to avoid the overfitting. Zhonglin Sun, Li Sun 0012, Qingli Li |
ICASSP | 2 |
| 2018 | Person Re-id by Incorporating PCA Loss in CNN
Li Sun 0012, Song Qiu, Qingli Li |
MMM (2) | 3 |
| 2017 | Facial age estimation through self-paced learningabstractThis paper proposes an age estimation algorithm in Self-Paced Learning (SPL) framework. Facial samples in the training set inherently include both easy and complex images, which is caused by both the characteristic of age and the variation of pose or expression. Furthermore, by randomly hiding patches in face region, data with different difficulty levels can be gradually used by SPL, in which Convolution Neural Network (CNN) is trained to give the estimation. Alternative Optimization Strategy (AVO), for the weight of CNN and the latent weight in SPL regularizer, is adopted in SPL framework. Experiments show that the proposed algorithm is able to give an accurate results especially under pose and expression variation. Li Sun 0012, Song Qiu, Mei Zhou, Qingli Li |
VCIP | 1 |