Sifan Long 0001

dblp:226/5805-1 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0001-7060-1133ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Semantic consistency learning across temporal scales for weakly supervised video anomaly detection
Sifan Long 0001, Hong-Wei Ge, Enxuan Gu, Zidi Li, Zhaoqin Wang
Knowl. Based Syst.2
2025 Prompt-induced prototype alignment for few-shot unsupervised domain adaptation
Yongguang Li, Sifan Long 0001, Sheng-Sheng Wang 0001, Xin Zhao 0021
Expert Syst. Appl.2
2025 Mutual Prompt Leaning for Vision Language Models
Sifan Long 0001, Zhen Zhao 0001, Junkun Yuan, Zichang Tan, Jiangjiang Liu 0006, Jingyuan Feng, Sheng-Sheng Wang 0001, Jingdong Wang 0001
Int. J. Comput. Vis.1
2025 A Slim Prompt-Averaged Consistency prompt learning for vision-language model
Siyu He, Sheng-Sheng Wang 0001, Sifan Long 0001
Knowl. Based Syst.3
2025 SOEDiff: Efficient Distillation for Small Object Editing
abstract
In this article, we delve into a new task known as Small Object Editing (SOE), which focuses on text-based image inpainting within a constrained, small-sized area. Despite the remarkable success have been achieved by current image inpainting approaches, their application to the SOE task generally results in failure cases such as Object Missing, Text-Image Mismatch, and Distortion . These failures stem from the limited use of small-sized objects in training datasets and the down-sampling operations employed by U-Net models, which hinders accurate generation. To overcome these challenges, we introduce a novel training-based approach, SOEDiff, aimed at enhancing the capability of baseline models like StableDiffusion in editing small-sized objects while minimizing training costs. Specifically, our method involves two key components: SO-LoRA , which efficiently fine-tunes low-rank matrices, and Cross-scale score distillation , which leverages high-resolution predictions from the pre-trained teacher diffusion model. Our method presents significant improvements on the test dataset collected from MSCOCO and OpenImage, validating the effectiveness of our proposed method in SOE. In particular, when comparing SOEDiff with SD-I model on the OpenImage-small-val dataset, we observe a 0.99 improvement in CLIP-Score and a reduction of 2.87 in FID.
Yiming Wu 0005, Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Ronghua Liang
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Towards Small Object Editing: A Benchmark Dataset and A Training-Free Approach
abstract
A plethora of text-guided image editing methods has recently been developed by leveraging the impressive capabilities of large-scale diffusion-based generative models especially Stable Diffusion. Despite the success of diffusion models in producing high-quality images, their application to small object generation has been limited due to difficulties in aligning cross-modal attention maps between text and these objects. Our approach offers a training-free method that significantly mitigates this alignment issue with local and global attention guidance, enhancing the model's ability to accurately render small objects in accordance with textual descriptions. We detail the methodology in our approach, emphasizing its divergence from traditional generation techniques and highlighting its advantages. What's more important is that we also provide SOEBench (Small Object Editing), a standardized benchmark for quantitatively evaluating text-based small object generation collected from MSCOCO[22] and OpenImage[18]. Preliminary results demonstrate the effectiveness of our method, showing marked improvements in the fidelity and accuracy of small object generation compared to existing models. This advancement not only contributes to the field of AI and computer vision but also opens up new possibilities for applications in various industries where precise image generation is critical.We will release our dataset on our project page: https://soebench.github.io/
Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Yiming Wu 0005, Wei Ji 0008, Haoran Liang 0001, Ronghua Liang
ACM Multimedia4
2024 Bidirectional feature enhancement transformer for unsupervised domain adaptation
Sheng-Sheng Wang 0001, Sifan Long 0001, Hao Chai
Vis. Comput.3
2023 Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision Transformers
abstract
Vision transformers have achieved significant improvements on various vision tasks but their quadratic interactions between tokens significantly reduce computational efficiency. Many pruning methods have been proposed to remove redundant tokens for efficient vision transformers recently. However, existing studies mainly focus on the token importance to preserve local attentive tokens but completely ignore the global token diversity. In this paper, we emphasize the cruciality of diverse global semantics and propose an efficient token decoupling and merging method that can jointly consider the token importance and diversity for token pruning. According to the class token attention, we decouple the attentive and inattentive tokens. In addition to preserving the most discriminative local tokens, we merge similar inattentive tokens and match homogeneous attentive tokens to maximize the token diversity. Despite its simplicity, our method obtains a promising trade-off between model complexity and classification accuracy. On DeiT-S, our method reduces the FLOPs by 35% with only a 0.2% accuracy drop. Notably, benefiting from maintaining the token diversity, our method can even improve the accuracy of DeiT-T by 0.1% after reducing its FLOPs by 40%.
Sifan Long 0001, Zhen Zhao 0001, Jimin Pi, Sheng-Sheng Wang 0001, Jingdong Wang 0001
CVPR1
2023 Instance-Specific and Model-Adaptive Supervision for Semi-Supervised Semantic Segmentation
abstract
Recently, semi-supervised semantic segmentation has achieved promising performance with a small fraction of labeled data. However, most existing studies treat all unlabeled data equally and barely consider the differences and training difficulties among unlabeled instances. Differentiating unlabeled instances can promote instance-specific supervision to adapt to the model's evolution dynamically. In this paper, we emphasize the cruciality of instance differences and propose an instance-specific and model-adaptive supervision for semi-supervised semantic segmentation, named iMAS. Relying on the model's performance, iMAS employs a class-weighted symmetric intersection-over-union to evaluate quantitative hardness of each unlabeled instance and supervises the training on unlabeled data in a model-adaptive manner. Specifically, iMAS learns from unlabeled instances progressively by weighing their corresponding consistency losses based on the evaluated hardness. Besides, iMAS dynamically adjusts the augmentation for each instance such that the distortion degree of augmented instances is adapted to the model's generalization capability across the training course. Not integrating additional losses and training procedures, iMAS can obtain remarkable performance gains against current state-of-the-art approaches on segmentation benchmarks under different semi-supervised partition protocols11Code and logs: https://github.com/zhenzhao/iMAS.
Zhen Zhao 0001, Sifan Long 0001, Jimin Pi, Jingdong Wang 0001, Luping Zhou
CVPR2
2023 Augmentation Matters: A Simple-Yet-Effective Approach to Semi-Supervised Semantic Segmentation
abstract
Recent studies on semi-supervised semantic segmentation (SSS) have seen fast progress. Despite their promising performance, current state-of-the-art methods tend to increasingly complex designs at the cost of introducing more network components and additional training procedures. Differently, in this work, we follow a standard teacher-student framework and propose AugSeg, a simple and clean approach that focuses mainly on data perturbations to boost the SSS performance. We argue that various data augmentations should be adjusted to better adapt to the semi-supervised scenarios instead of directly applying these techniques from supervised learning. Specifically, we adopt a simplified intensity-based augmentation that selects a random number of data transformations with uniformly sampling distortion strengths from a continuous space. Based on the estimated confidence of the model on different unlabeled samples, we also randomly inject labelled information to augment the unlabeled samples in an adaptive manner. Without bells and whistles, our simple AugSeg can readily achieve new state-of-the-art performance on SSS benchmarks under different partition protocols11Code and logs: https://github.com/zhenzhao/AugSeg..
Zhen Zhao 0001, Lihe Yang, Sifan Long 0001, Jimin Pi, Luping Zhou, Jingdong Wang 0001
CVPR3
2023 Task-Oriented Multi-Modal Mutual Learning for Vision-Language Models
abstract
Prompt learning has become one of the most efficient paradigms for adapting large pre-trained vision-language models to downstream tasks. Current state-of-the-art methods, like CoOp and ProDA, tend to adopt soft prompts to learn an appropriate prompt for each specific task. Recent CoCoOp further boosts the base-to-new generalization performance via an image-conditional prompt. However, it directly fuses identical image semantics to prompts of different labels and significantly weakens the discrimination among different classes as shown in our experiments. Motivated by this observation, we first propose a class-aware text prompt (CTP) to enrich generated prompts with label-related image information. Unlike CoCoOp, CTP can effectively involve image semantics and avoid introducing extra ambiguities into different prompts. On the other hand, instead of reserving the complete image representations, we propose text-guided feature tuning (TFT) to make the image branch attend to class-related representation. A contrastive loss is employed to align such augmented text and image representations on downstream tasks. In this way, the image-to-text CTP and text-to-image TFT can be mutually promoted to enhance the adaptation of VLMs for downstream tasks. Extensive experiments demonstrate that our method outperforms the existing methods by a significant margin. Especially, compared to CoCoOp, we achieve an average improvement of 4.03% on new classes and 3.19% on harmonic-mean over eleven classification benchmarks.
Sifan Long 0001, Zhen Zhao 0001, Junkun Yuan, Zichang Tan, Jiangjiang Liu 0006, Luping Zhou, Sheng-Sheng Wang 0001, Jingdong Wang 0001
ICCV1
2023 HAP: Structure-Aware Masked Image Modeling for Human-Centric Perception
abstract
Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure priors offer significant potential. Motivated by this insight, we further incorporate an intuitive human structure prior - human parts - into pre-training. Specifically, we employ this prior to guide the mask sampling process. Image patches, corresponding to human part regions, have high priority to be masked out. This encourages the model to concentrate more on body structure information during pre-training, yielding substantial benefits across a range of human-centric perception tasks. To further capture human characteristics, we propose a structure-invariant alignment loss that enforces different masked views, guided by the human part prior, to be closely aligned for the same image. We term the entire method as HAP. HAP simply uses a plain ViT as the encoder yet establishes new state-of-the-art performance on 11 human-centric benchmarks, and on-par result on one dataset. For example, HAP achieves 78.1% mAP on MSMT17 for person re-identification, 86.54% mA on PA-100K for pedestrian attribute recognition, 78.2% AP on MS COCO for 2D pose estimation, and 56.0 PA-MPJPE on 3DPW for 3D pose and shape estimation.
Junkun Yuan, Xinyu Zhang 0015, Hao Zhou 0039, Jian Wang 0066, Zhongwei Qiu, Zhiyin Shao, Shaofeng Zhang, Sifan Long 0001, Kun Kuang 0001, Junyu Han, Errui Ding, Lanfen Lin, Fei Wu 0001, Jingdong Wang 0001
NeurIPS8
2023 Sample separation and domain alignment complementary learning mechanism for open set domain adaptation
Sifan Long 0001, Sheng-Sheng Wang 0001, Xin Zhao 0021, Bilin Wang
Appl. Intell.1
2023 Causal view mechanism for adversarial domain adaptation
Sheng-Sheng Wang 0001, Xin Zhao 0021, Sifan Long 0001, Bilin Wang
Multim. Tools Appl.4
2022 Cross-domain feature enhancement for unsupervised domain adaptation
Sifan Long 0001, Sheng-Sheng Wang 0001, Xin Zhao 0021, Bilin Wang
Appl. Intell.1