EDBT 2026 Demo / reviewers in the wild / expert
Sifan Long 0001
dblp:226/5805-1
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0001-7060-1133ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic consistency learning across temporal scales for weakly supervised video anomaly detection
Sifan Long 0001, Hong-Wei Ge, Enxuan Gu, Zidi Li, Zhaoqin Wang |
Knowl. Based Syst. | 2 |
| 2025 | Prompt-induced prototype alignment for few-shot unsupervised domain adaptation
Yongguang Li, Sifan Long 0001, Sheng-Sheng Wang 0001, Xin Zhao 0021 |
Expert Syst. Appl. | 2 |
| 2025 | Mutual Prompt Leaning for Vision Language Models
Sifan Long 0001, Zhen Zhao 0001, Junkun Yuan, Zichang Tan, Jiangjiang Liu 0006, Jingyuan Feng, Sheng-Sheng Wang 0001, Jingdong Wang 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | A Slim Prompt-Averaged Consistency prompt learning for vision-language model
Siyu He, Sheng-Sheng Wang 0001, Sifan Long 0001 |
Knowl. Based Syst. | 3 |
| 2025 | SOEDiff: Efficient Distillation for Small Object EditingabstractIn this article, we delve into a new task known as Small Object Editing (SOE), which focuses on text-based image inpainting within a constrained, small-sized area. Despite the remarkable success have been achieved by current image inpainting approaches, their application to the SOE task generally results in failure cases such as Object Missing, Text-Image Mismatch, and Distortion . These failures stem from the limited use of small-sized objects in training datasets and the down-sampling operations employed by U-Net models, which hinders accurate generation. To overcome these challenges, we introduce a novel training-based approach, SOEDiff, aimed at enhancing the capability of baseline models like StableDiffusion in editing small-sized objects while minimizing training costs. Specifically, our method involves two key components: SO-LoRA , which efficiently fine-tunes low-rank matrices, and Cross-scale score distillation , which leverages high-resolution predictions from the pre-trained teacher diffusion model. Our method presents significant improvements on the test dataset collected from MSCOCO and OpenImage, validating the effectiveness of our proposed method in SOE. In particular, when comparing SOEDiff with SD-I model on the OpenImage-small-val dataset, we observe a 0.99 improvement in CLIP-Score and a reduction of 2.87 in FID. Yiming Wu 0005, Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Ronghua Liang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Towards Small Object Editing: A Benchmark Dataset and A Training-Free ApproachabstractA plethora of text-guided image editing methods has recently been developed by leveraging the impressive capabilities of large-scale diffusion-based generative models especially Stable Diffusion. Despite the success of diffusion models in producing high-quality images, their application to small object generation has been limited due to difficulties in aligning cross-modal attention maps between text and these objects. Our approach offers a training-free method that significantly mitigates this alignment issue with local and global attention guidance, enhancing the model's ability to accurately render small objects in accordance with textual descriptions. We detail the methodology in our approach, emphasizing its divergence from traditional generation techniques and highlighting its advantages. What's more important is that we also provide SOEBench (Small Object Editing), a standardized benchmark for quantitatively evaluating text-based small object generation collected from MSCOCO[22] and OpenImage[18]. Preliminary results demonstrate the effectiveness of our method, showing marked improvements in the fidelity and accuracy of small object generation compared to existing models. This advancement not only contributes to the field of AI and computer vision but also opens up new possibilities for applications in various industries where precise image generation is critical.We will release our dataset on our project page: https://soebench.github.io/ Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Yiming Wu 0005, Wei Ji 0008, Haoran Liang 0001, Ronghua Liang |
ACM Multimedia | 4 |
| 2024 | Bidirectional feature enhancement transformer for unsupervised domain adaptation
Sheng-Sheng Wang 0001, Sifan Long 0001, Hao Chai |
Vis. Comput. | 3 |
| 2023 | Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision TransformersabstractVision transformers have achieved significant improvements on various vision tasks but their quadratic interactions between tokens significantly reduce computational efficiency. Many pruning methods have been proposed to remove redundant tokens for efficient vision transformers recently. However, existing studies mainly focus on the token importance to preserve local attentive tokens but completely ignore the global token diversity. In this paper, we emphasize the cruciality of diverse global semantics and propose an efficient token decoupling and merging method that can jointly consider the token importance and diversity for token pruning. According to the class token attention, we decouple the attentive and inattentive tokens. In addition to preserving the most discriminative local tokens, we merge similar inattentive tokens and match homogeneous attentive tokens to maximize the token diversity. Despite its simplicity, our method obtains a promising trade-off between model complexity and classification accuracy. On DeiT-S, our method reduces the FLOPs by 35% with only a 0.2% accuracy drop. Notably, benefiting from maintaining the token diversity, our method can even improve the accuracy of DeiT-T by 0.1% after reducing its FLOPs by 40%. Sifan Long 0001, Zhen Zhao 0001, Jimin Pi, Sheng-Sheng Wang 0001, Jingdong Wang 0001 |
CVPR | 1 |
| 2023 | Instance-Specific and Model-Adaptive Supervision for Semi-Supervised Semantic SegmentationabstractRecently, semi-supervised semantic segmentation has achieved promising performance with a small fraction of labeled data. However, most existing studies treat all unlabeled data equally and barely consider the differences and training difficulties among unlabeled instances. Differentiating unlabeled instances can promote instance-specific supervision to adapt to the model's evolution dynamically. In this paper, we emphasize the cruciality of instance differences and propose an instance-specific and model-adaptive supervision for semi-supervised semantic segmentation, named iMAS. Relying on the model's performance, iMAS employs a class-weighted symmetric intersection-over-union to evaluate quantitative hardness of each unlabeled instance and supervises the training on unlabeled data in a model-adaptive manner. Specifically, iMAS learns from unlabeled instances progressively by weighing their corresponding consistency losses based on the evaluated hardness. Besides, iMAS dynamically adjusts the augmentation for each instance such that the distortion degree of augmented instances is adapted to the model's generalization capability across the training course. Not integrating additional losses and training procedures, iMAS can obtain remarkable performance gains against current state-of-the-art approaches on segmentation benchmarks under different semi-supervised partition protocols11Code and logs: https://github.com/zhenzhao/iMAS. Zhen Zhao 0001, Sifan Long 0001, Jimin Pi, Jingdong Wang 0001, Luping Zhou |
CVPR | 2 |
| 2023 | Augmentation Matters: A Simple-Yet-Effective Approach to Semi-Supervised Semantic SegmentationabstractRecent studies on semi-supervised semantic segmentation (SSS) have seen fast progress. Despite their promising performance, current state-of-the-art methods tend to increasingly complex designs at the cost of introducing more network components and additional training procedures. Differently, in this work, we follow a standard teacher-student framework and propose AugSeg, a simple and clean approach that focuses mainly on data perturbations to boost the SSS performance. We argue that various data augmentations should be adjusted to better adapt to the semi-supervised scenarios instead of directly applying these techniques from supervised learning. Specifically, we adopt a simplified intensity-based augmentation that selects a random number of data transformations with uniformly sampling distortion strengths from a continuous space. Based on the estimated confidence of the model on different unlabeled samples, we also randomly inject labelled information to augment the unlabeled samples in an adaptive manner. Without bells and whistles, our simple AugSeg can readily achieve new state-of-the-art performance on SSS benchmarks under different partition protocols11Code and logs: https://github.com/zhenzhao/AugSeg.. Zhen Zhao 0001, Lihe Yang, Sifan Long 0001, Jimin Pi, Luping Zhou, Jingdong Wang 0001 |
CVPR | 3 |
| 2023 | Task-Oriented Multi-Modal Mutual Learning for Vision-Language ModelsabstractPrompt learning has become one of the most efficient paradigms for adapting large pre-trained vision-language models to downstream tasks. Current state-of-the-art methods, like CoOp and ProDA, tend to adopt soft prompts to learn an appropriate prompt for each specific task. Recent CoCoOp further boosts the base-to-new generalization performance via an image-conditional prompt. However, it directly fuses identical image semantics to prompts of different labels and significantly weakens the discrimination among different classes as shown in our experiments. Motivated by this observation, we first propose a class-aware text prompt (CTP) to enrich generated prompts with label-related image information. Unlike CoCoOp, CTP can effectively involve image semantics and avoid introducing extra ambiguities into different prompts. On the other hand, instead of reserving the complete image representations, we propose text-guided feature tuning (TFT) to make the image branch attend to class-related representation. A contrastive loss is employed to align such augmented text and image representations on downstream tasks. In this way, the image-to-text CTP and text-to-image TFT can be mutually promoted to enhance the adaptation of VLMs for downstream tasks. Extensive experiments demonstrate that our method outperforms the existing methods by a significant margin. Especially, compared to CoCoOp, we achieve an average improvement of 4.03% on new classes and 3.19% on harmonic-mean over eleven classification benchmarks. Sifan Long 0001, Zhen Zhao 0001, Junkun Yuan, Zichang Tan, Jiangjiang Liu 0006, Luping Zhou, Sheng-Sheng Wang 0001, Jingdong Wang 0001 |
ICCV | 1 |
| 2023 | HAP: Structure-Aware Masked Image Modeling for Human-Centric PerceptionabstractModel pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure priors offer significant potential. Motivated by this insight, we further incorporate an intuitive human structure prior - human parts - into pre-training. Specifically, we employ this prior to guide the mask sampling process. Image patches, corresponding to human part regions, have high priority to be masked out. This encourages the model to concentrate more on body structure information during pre-training, yielding substantial benefits across a range of human-centric perception tasks. To further capture human characteristics, we propose a structure-invariant alignment loss that enforces different masked views, guided by the human part prior, to be closely aligned for the same image. We term the entire method as HAP. HAP simply uses a plain ViT as the encoder yet establishes new state-of-the-art performance on 11 human-centric benchmarks, and on-par result on one dataset. For example, HAP achieves 78.1% mAP on MSMT17 for person re-identification, 86.54% mA on PA-100K for pedestrian attribute recognition, 78.2% AP on MS COCO for 2D pose estimation, and 56.0 PA-MPJPE on 3DPW for 3D pose and shape estimation. Junkun Yuan, Xinyu Zhang 0015, Hao Zhou 0039, Jian Wang 0066, Zhongwei Qiu, Zhiyin Shao, Shaofeng Zhang, Sifan Long 0001, Kun Kuang 0001, Junyu Han, Errui Ding, Lanfen Lin, Fei Wu 0001, Jingdong Wang 0001 |
NeurIPS | 8 |
| 2023 | Sample separation and domain alignment complementary learning mechanism for open set domain adaptation
Sifan Long 0001, Sheng-Sheng Wang 0001, Xin Zhao 0021, Bilin Wang |
Appl. Intell. | 1 |
| 2023 | Causal view mechanism for adversarial domain adaptation
Sheng-Sheng Wang 0001, Xin Zhao 0021, Sifan Long 0001, Bilin Wang |
Multim. Tools Appl. | 4 |
| 2022 | Cross-domain feature enhancement for unsupervised domain adaptation
Sifan Long 0001, Sheng-Sheng Wang 0001, Xin Zhao 0021, Bilin Wang |
Appl. Intell. | 1 |