Yufei Yin

dblp:237/0568 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KF-GS: Kalman filter-guided Gaussian splatting for real-time high-quality dynamic scene reconstruction
Qingyuan Tang, Yufei Yin, Yanming Zhu 0001, Zhou Yu 0001, Zhenzhong Kuang, Jiajun Ding, Jifa He
J. Vis. Commun. Image Represent.2
2026 Enhancing Object Detection with Active Exploration and Spatiotemporal Aggregation
abstract
Classic object detectors are fundamentally limited by their reliance on a single, static viewpoint, which often suffers from occlusions, challenging scales, and ambiguous perspectives. Video-based object detectors can partially mitigate this problem by aggregating information from multiple views, but they usually work with passively gathered sequences and are in no means guaranteed to provide sufficient information for all objects of interest. On the other hand, active vision methods can seek out better views, yet existing active object detectors typically only try to search for a single optimal perspective and discard valuable information gathered along their path. In this article, we introduce a new paradigm for object detection that unifies active exploration with cumulative spatiotemporal aggregation. We train an embodied agent to intelligently explore its environment, guided by a novel, detection-aware reward function that directly encourages seeking out views that resolve visual ambiguities. To leverage this active exploration, we introduce a robust aggregation pipeline that adeptly harnesses SAM 2’s temporal reasoning capabilities to fuse information from the agent’s entire trajectory into a single, coherent set of detections. Through extensive experiments on AI2-THOR, we demonstrate that our framework provides a consistent and substantial performance uplift when applied to a wide range of state-of-the-art detectors, establishing a strong and versatile new baseline for the next generation of active vision systems.
Peiwei Li, Min Wang 0019, Wengang Zhou 0001, Yufei Yin, Yebo Bao, Guodong Shen, Houqiang Li
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Unified Open-World Segmentation with Multi-Modal Prompts
Yang Liu 0357, Yufei Yin, Chenchen Jing, Muzhi Zhu, Hao Chen 0041, Yuling Xi, Hao Wang 0052, Chunhua Shen
ICCV2
2025 Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing
abstract
Text-driven video editing powered by generative diffusion models holds significant promise for applications spanning film production, advertising, and beyond. However, the limited expressiveness of pre-trained word embeddings often restricts nuanced edits, especially when targeting novel concepts with specific attributes. In this work, we present a novel Concept-Augmented Textual Inversion (CATI) framework that flexibly integrates new object information from user-provided concept videos. By fine-tuning only the V (Value) projection in attention via Low-Rank Adaptation (LoRA), our approach preserves the original attention distribution of the diffusion model while efficiently incorporating external concept knowledge. To further stabilize editing results and mitigate the issue of attention dispersion when prompt keywords are modified, we introduce a Dual Prior Supervision (DPS) mechanism. DPS supervises cross-attention between the source and target prompts, preventing undesired changes to non-target areas and improving the fidelity of novel concepts. Extensive evaluations demonstrate that our plug-and-play solution not only maintains spatial and temporal consistency but also outperforms state-of-the-art methods in generating lifelike and stable edited videos. The source code is publicly available at https://guomc9.github.io/STIVE-PAGE/.
Mingce Guo, Jingxuan He 0001, Yufei Yin, Zhangye Wang, Shengeng Tang, Lechao Cheng
IJCAI3
2025 Self-Classification Enhancement and Correction for Weakly Supervised Object Detection
abstract
In recent years, weakly supervised object detection (WSOD) has attracted much attention due to its low labeling cost. The success of recent WSOD models is often ascribed to the two-stage multi-class classification (MCC) task, i.e., multiple instance learning and online classification refinement. Despite achieving non-trivial progresses, these methods overlook potential classification ambiguities between these two MCC tasks and fail to leverage their unique strengths. In this work, we introduce a novel WSOD framework to ameliorate these two issues. For one thing, we propose a self-classification enhancement module that integrates intra-class binary classification (ICBC) to bridge the gap between the two distinct MCC tasks. The ICBC task enhances the network’s discrimination between positive and mis-located samples in a class-wise manner and forges a mutually reinforcing relationship with the MCC task. For another, we propose a self-classification correction algorithm during inference, which combines the results of both MCC tasks to effectively reduce the mis-classified predictions. Extensive experiments on the prevalent VOC 2007 & 2012 datasets demonstrate the superior performance of our framework.
Yufei Yin, Lechao Cheng, Wengang Zhou 0001, Jiajun Deng, Houqiang Li
IJCAI1
2024 Revisiting Open-Set Panoptic Segmentation
abstract
In this paper, we focus on the open-set panoptic segmentation (OPS) task to circumvent the data explosion problem. Different from the close-set setting, OPS targets to detect both known and unknown categories, where the latter is not annotated during training. Different from existing work that only selects a few common categories as unknown ones, we move forward to the real-world scenario by considering the various tail categories (~1k). To this end, we first build a new dataset with long-tail distribution for the OPS task. Based on this dataset, we additionally add a new class type for unknown classes and re-define the training annotations to make the OPS definition more complete and reasonable. Moreover, we analyze the influence of several significant factors in the OPS task and explore the upper bound of performance on unknown classes with different settings. Furthermore, based on the analyses, we design an effective two-phase framework for the OPS task, including thing-agnostic map generation and unknown segment mining. We further adopt semi-supervised learning to improve the OPS performance. Experimental results on different datasets validate the effectiveness of our method.
Yufei Yin, Hao Chen 0041, Wengang Zhou 0001, Jiajun Deng, Houqiang Li
AAAI1
2024 Masked Collaborative Contrast for Weakly Supervised Semantic Segmentation
abstract
This study introduces an efficacious approach, Masked Collaborative Contrast (MCC), to highlight semantic regions in weakly supervised semantic segmentation. MCC adroitly draws inspiration from masked image modeling and contrastive learning to devise a novel framework that induces keys to contract toward semantic regions. Unlike prevalent techniques that directly eradicate patch regions in the input image when generating masks, we scrutinize the neighborhood relations of patch tokens by exploring masks considering keys on the affinity matrix. Moreover, we generate positive and negative samples in contrastive learning by utilizing the masked local output and contrasting it with the global output. Elaborate experiments on commonly employed datasets evidences that the proposed MCC mechanism effectively aligns global and local perspectives within the image, attaining impressive performance. The source code is available at https://github.com/fwu11/MCC.
Fangwen Wu, Jingxuan He 0001, Yufei Yin, Yanbin Hao, Gang Huang 0004, Lechao Cheng
WACV3
2024 Recurrent Generic Contour-Based Instance Segmentation With Progressive Learning
abstract
Contour-based instance segmentation has been actively studied, thanks to its flexibility and elegance in processing visual objects within complex backgrounds. In this work, we propose a novel deep network architecture,i.e., PolySnake, for generic contour-based instance segmentation. Motivated by the classic Snake algorithm, the proposed PolySnake achieves superior and robust segmentation performance with an iterative and progressive contour refinement strategy. Technically, PolySnake introduces a recurrent update operator to estimate the object contour iteratively. It maintains a single estimate of the contour that is progressively deformed toward the object boundary. At each iteration, PolySnake builds a semantic-rich representation for the current contour and feeds it to the recurrent operator for further contour adjustment. Through the iterative refinements, the contour progressively converges to a stable status that tightly encloses the object instance. Beyond the scope of general instance segmentation, extensive experiments are conducted to validate the effectiveness and generalizability of our PolySnake in two additional specific task scenarios, including scene text detection and lane detection. The results demonstrate that the proposed PolySnake outperforms the existing advanced methods on several multiple prevalent benchmarks across the three tasks. The codes and pre-trained models are available at https://github.com/fh2019ustc/PolySnake.
Hao Feng 0009, Keyi Zhou, Wengang Zhou 0001, Yufei Yin, Jiajun Deng, Qi Sun 0005, Houqiang Li
IEEE Trans. Circuits Syst. Video Technol.4
2023 Cyclic-Bootstrap Labeling for Weakly Supervised Object Detection
abstract
Recent progress in weakly supervised object detection is featured by a combination of multiple instance detection networks (MIDN) and ordinal online refinement. However, with only image-level annotation, MIDN inevitably assigns high scores to some unexpected region proposals when generating pseudo labels. These inaccurate high-scoring region proposals will mislead the training of subsequent refinement modules and thus hamper the detection performance. In this work, we explore how to ameliorate the quality of pseudo-labeling in MIDN. Formally, we devise Cyclic-Bootstrap Labeling (CBL), a novel weakly supervised object detection pipeline, which optimizes MIDN with rank information from a reliable teacher network. Specifically, we obtain this teacher network by introducing a weighted exponential moving average strategy to take advantage of various refinement modules. A novel class-specific ranking distillation algorithm is proposed to leverage the output of weighted ensembled teacher network for distilling MIDN with rank information. As a result, MIDN is guided to assign higher scores to accurate proposals among their neighboring ones, thus benefiting the subsequent pseudo labeling. Extensive experiments on the prevalent PASCAL VOC 2007 & 2012 and COCO datasets demonstrate the superior performance of our CBL framework. Code will be available at https://github.com/Yinyf0804/WSOD-CBL/.
Yufei Yin, Jiajun Deng, Wengang Zhou 0001, Li Li 0040, Houqiang Li
ICCV1
2023 FI-WSOD: Foreground Information Guided Weakly Supervised Object Detection
abstract
Existing solutions for weakly supervised object detection (WSOD) generally follow the multiple instance learning (MIL) paradigm to formulate WSOD as a multi-class classification problem over a set of region proposals. However, without the supervision signal of ground-truth boxes, the training objective of multi-class classification makes the detectors devote main efforts to finding the most common pattern of each class, as the common pattern is always the most discriminative evidence for classification. In addition, although learning from distinguishing multiple foreground classes, the detectors can still ignore to differentiate foreground regions from the background ones, which causes false alarm in prediction. These two points account for the limited localization capability of MIL-based WSOD methods. To this end, we propose foreground information guided WSOD (FI-WSOD), a novel framework that introduces an extra foreground-background binary classification (F-BBC) sub-task to the original MIL-based WSOD paradigm. At the training stage, the involvement of F-BBC task not only improves the feature representation of the network, but also provides extra information from the foreground-background perspective. By leveraging the learnt foreground information, a Foreground Guided Self-Training (FGST) module is further proposed to filter out noisy samples, and to mine representative seeds from the remaining proposals. Moreover, a Multi-Seed Training strategy is performed to reduce the impact of noisy labels when training the self-training networks in FGST. We have conducted extensive experiments on the prevalent Pascal VOC 2007, Pascal VOC 2012 and MSCOCO datasets, and report a series of state-of-the-art records achieved by our proposed framework.
Yufei Yin, Jiajun Deng, Wengang Zhou 0001, Li Li 0040, Houqiang Li
IEEE Trans. Multim.1
2022 Dual Decision Improves Open-Set Panoptic Segmentation
Hao Chen 0041, Lingqiao Liu, Yufei Yin
BMVC4
2021 Instance Mining with Class Feature Banks for Weakly Supervised Object Detection
abstract
Recent progress on weakly supervised object detection (WSOD) is characterized by formulating WSOD as a Multiple Instance Learning (MIL) problem and taking online refinement with the selected region proposals from MIL. However, MIL inclines to select the most discriminative part rather than the entire instance as the top-scoring region proposals, which leads to weak localization capability for weakly supervised object detectors. We attribute this problem to the limited intra-class diversity within a single image. Specifically, due to the lack of annotated bounding boxes, the network tends to focus on the most common parts of each class and neglect the diverse parts of objects. To solve the problem, we introduce a novel Instance Mining with Class Feature Banks (IM-CFB) framework, which includes a Class Feature Banks (CFB) module and a Feature Guided Instance Mining (FGIM) algorithm. Concretely, Class Feature Banks (CFB) consist of sub-banks for each class, which are utilized to collect diversity information from a broader view. At the training stage, the RoI features of reliable region proposals are recorded and updated in the CFB. Then, FGIM leverages the features recorded in the CFB to ameliorate the region proposal selection of the MIL branch. Extensive experiments conducted on two publicly available datasets, Pascal VOC 2007 and 2012, demonstrate the effectiveness of our method. More remarkably, our method achieves 54.3% on mAP and 70.7% on CorLoc on Pascal VOC 2007. When further re-trained by a Fast-RCNN detector, we obtain to-date the best reported mAP and CorLoc of 55.8% and 72.2%, respectively.
Yufei Yin, Jiajun Deng, Wengang Zhou 0001, Houqiang Li
AAAI1