EDBT 2026 Demo / reviewers in the wild / expert
Yunfeng Yan
dblp:203/6952
· DBLP profile ↗
20ranked-venue papers
1as first author
19since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ROSE: Remove Objects with Side Effects in VideosabstractVideo object removal has achieved advanced performance due to the recent success of video generative models. However, when addressing the side effects of objects, \textit{e.g.,} their shadows and reflections, existing works struggle to eliminate these effects for the scarcity of paired video data as supervision. This paper presents \method, termed \textbf{R}emove \textbf{O}bjects with \textbf{S}ide \textbf{E}ffects, a framework that systematically studies the object's effects on environment, which can be categorized into five common cases: shadows, reflections, light, translucency and mirror. Given the challenges of curating paired videos exhibiting the aforementioned effects, we leverage a 3D rendering engine for synthetic data generation. We carefully construct a fully-automatic pipeline for data preparation, which simulates a large-scale paired dataset with diverse scenes, objects, shooting angles, and camera trajectories. ROSE is implemented as an video inpainting model built on diffusion transformer. To localize all object-correlated areas, the entire video is fed into the model for reference-based erasing. Moreover, additional supervision is introduced to explicitly predict the areas affected by side effects, which can be revealed through the differential mask between the paired videos. To fully investigate the model performance on various side effect removal, we presents a new benchmark, dubbed ROSE-Bench, incorporating both common scenarios and the five special side effects for comprehensive evaluation. Experimental results demonstrate that \method achieves superior performance compared to existing video object erasing models and generalizes well to real-world video scenarios. Chenxuan Miao, Yutong Feng, Jianshu Zeng, Zixiang Gao, Hantang Liu, Yunfeng Yan, Donglian Qi, Xi Chen 0119, Hengshuang Zhao |
NeurIPS | 6 |
| 2025 | MDRN: Multi-domain representation network for unsupervised domain generalizationabstractAbstract In deep neural networks, performance can degrade when test data distributions differ from training data. Unsupervised Domain Generalization (UDG) aims to improve generalization across unseen domains by leveraging multiple source domains without supervision. Traditional methods focus on extracting domain‐invariant features, potentially at the expense of feature space integrity and generalization potential. We presents a Multi‐Domain Representation Network (MDRN) for unsupervised multi‐domain learning. MDRN innovates by disentangling and preserving both domain‐invariant and domain‐specific features through an unsupervised cross‐domain reconstruction task. It employs content encoders for domain‐invariant features and multi‐domain style encoders for domain‐specific characteristics. By merging these features based on domain similarity, MDRN constructs a comprehensive feature space that enhances image reconstruction across domains. Additionally, MDRN integrates domain‐specific classifiers, which learn domain classification and provide weighted fusion of domain‐specific features. This design facilitates effective inter‐domain distance measurement and feature integration. Experiments on PACS and DomainNet show MDRN's superior performance over existing state‐of‐the‐art UDG approaches, highlighting its effectiveness in handling distribution shifts between source and target domains. Yangyang Zhong, Yunfeng Yan, Pengxin Luo, Weizhen He, Yiheng Deng, Donglian Qi |
IET Image Process. | 2 |
| 2025 | Adept: Annotation-denoising auxiliary tasks with discrete cosine transform map and keypoint for human-centric pretraining
Weizhen He, Yunfeng Yan, Shixiang Tang, Yiheng Deng, Yangyang Zhong, Pengxin Luo, Donglian Qi |
Neurocomputing | 2 |
| 2025 | Instruct-ReID++: Towards Universal Purpose Instruction-Guided Person Re-IdentificationabstractRecently, person re-identification (ReID) has witnessed fast development due to its broad practical applications and proposed various settings, e.g., traditional ReID, clothes-changing ReID, and visible-infrared ReID. However, current studies primarily focus on single specific tasks, which limits model applicability in real-world scenarios. This paper aims to address this issue by introducing a novel instruct-ReID task that unifies 6 existing ReID tasks in one model and retrieves images based on provided visual or textual instructions. Instruct-ReID is the first exploration of a general ReID setting, where 6 existing ReID tasks can be viewed as special cases by assigning different instructions. To facilitate research in this new instruct-ReID task, we propose a large-scale OmniReID++ benchmark equipped with diverse data and comprehensive evaluation methods, e.g., task-specific and task-free evaluation settings. In the task-specific evaluation setting, gallery sets are categorized according to specific ReID tasks. We propose a novel baseline model, IRM, with an adaptive triplet loss to handle various retrieval tasks within a unified framework. For task-free evaluation setting, where target person images are retrieved from task-agnostic gallery sets, we further propose a new method called IRM++ with novel memory bank-assisted learning. Extensive evaluations of IRM and IRM++ on OmniReID++ benchmark demonstrate the superiority of our proposed methods, achieving state-of-the-art performance on 10 test sets. Weizhen He, Yiheng Deng, Yunfeng Yan, Feng Zhu 0006, Yizhou Wang 0007, Lei Bai 0001, Qingsong Xie, Rui Zhao 0001, Donglian Qi, Wanli Ouyang, Shixiang Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Wavelet-Driven Spatiotemporal Predictive Learning: Bridging Frequency and Time VariationsabstractSpatiotemporal predictive learning is a paradigm that empowers models to learn spatial and temporal patterns by predicting future frames from past frames in an unsupervised manner. This method typically uses recurrent units to capture long-term dependencies, but these units often come with high computational costs and limited performance in real-world scenes. This paper presents an innovative Wavelet-based SpatioTemporal (WaST) framework, which extracts and adaptively controls both low and high-frequency components at image and feature levels via 3D discrete wavelet transform for faster processing while maintaining high-quality predictions. We propose a Time-Frequency Aware Translator uniquely crafted to efficiently learn short- and long-range spatiotemporal information by individually modeling spatial frequency and temporal variations. Meanwhile, we design a wavelet-domain High-Frequency Focal Loss that effectively supervises high-frequency variations. Extensive experiments across various real-world scenarios, such as driving scene prediction, traffic flow prediction, human motion capture, and weather forecasting, demonstrate that our proposed WaST achieves state-of-the-art performance over various spatiotemporal prediction methods. Xuesong Nie, Yunfeng Yan, Siyuan Li 0002, Cheng Tan 0012, Xi Chen 0119, Haoyuan Jin, Zhihang Zhu, Stan Z. Li, Donglian Qi |
AAAI | 2 |
| 2024 | Instruct-ReID: A Multi-Purpose Person Re-Identification Task with InstructionsabstractHuman intelligence can retrieve any person according to both visual and language descriptions. However, the current computer vision community studies specific person re-identification (ReID) tasks in different scenarios separately, which limits the applications in the real world. This paper strives to resolve this problem by proposing a new instruct-ReID task that requires the model to retrieve images according to the given image or language instructions. Our instruct-ReID is a more general ReID setting, where existing 6 ReID tasks can be viewed as special cases by designing different instructions. We propose a large-scale OmniReID benchmark and an adaptive triplet loss as a baseline method to facilitate research in this new setting. Experimental results show that the proposed multi-purpose ReID model, trained on our OmniReID benchmark without finetuning, can improve +0.5%, +0.6%, +7.7% mAP on Market1501, MSMT17, CUHK03 for traditional ReID, +6.4%, +7.1%, +11.2% mAP on PRCC, VC-Clothes, LTCC for clothes-changing ReID, +11.7% mAP on COCAS+ real2 for clothes template based clothes-changing ReID when using only RGB images, +24.9% mAP on COCAS+ real2 for our newly defined language-instructed ReID, +4.3% on LLCM for visible-infrared ReID, +2.6% on CUHK-PEDES for text-to-image ReID. The datasets, the model, and code are available at https://github.com/hwz-zju/Instruct-ReID. Weizhen He, Yiheng Deng, Shixiang Tang, Qihao Chen, Qingsong Xie, Yizhou Wang 0007, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001, Wanli Ouyang, Donglian Qi, Yunfeng Yan |
CVPR | 12 |
| 2024 | PredToken: Predicting Unknown Tokens and Beyond with Coarse-to-Fine Iterative DecodingabstractPredictive learning models, which aim to predict future frames based on past observations, are crucial to constructing world models. These models need to maintain low-level consistency and capture high-level dynamics in unannotated spatiotemporal data. Transitioning from frame-wise to token-wise prediction presents a viable strategy for addressing these needs. How to improve token representation and optimize token decoding presents significant challenges. This paper introduces PredToken, a novel predictive framework that addresses these issues by decoupling space-time tokens into distinct components for iterative cascaded decoding. Concretely, we first design a “decomposition, quantization, and reconstruction” schema based on VQGAN to improve the token representation. This scheme disentangles low- and high-frequency representations and employs a dimension-aware quantization model, allowing more low-level details to be preserved. Building on this, we present a “coarse-to-fine iterative decoding” method. It leverages dynamic soft decoding to refine coarse tokens and static soft decoding for fine tokens, enabling more high-level dynamics to be captured. These designs make Pred-Token produce high-quality predictions. Extensive experiments demonstrate the superiority of our method on various real-world spatiotemporal predictive benchmarks. Furthermore, PredToken can also be extended to other visual generative tasks to yield realistic outcomes. Xuesong Nie, Haoyuan Jin, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi |
CVPR | 3 |
| 2024 | SAMP: Adapting Segment Anything Model for Pose EstimationabstractSegment Anything Model (SAM) exhibits superior performance for segmentation. Many follow-up works explore adapting this powerful model to specific domains. However, those works mainly focus on different sub-tasks of segmentation. The cross-task generalization ability of SAM is still not explored. In this paper, we propose SAMP (SAM for Pose), which makes the first attempt to adapt SAM for pose estimation. We observe that SAM could segment different human parts with specific prompts, proving that it contains the knowledge to understand the human structure. Considering that localizing keypoints requires fine-grained perceptual capabilities, we design a Detail-aware Adapter (DA-Adapter), which complements the features of the SAM encoder with multi-scale feature fusion and multi-level supervision. Experimental results demonstrate that SAMP achieves novel state-of-the-art against previously specifically designed pose estimation methods. Specifically, with ViT-B backbone, SAMP achieves 78.1% AP on the COCO val2017, 77.1% AP on the COCO test-dev2017, and 70.5% AP on the CrowdPose dataset. Zhihang Zhu, Yunfeng Yan, Haoyuan Jin, Xuesong Nie, Donglian Qi, Xi Chen 0119 |
ICME | 2 |
| 2024 | Object-Level Pseudo-3D Lifting for Distance-Aware TrackingabstractMulti-object tracking (MOT) is a pivotal task for media interpretation, where reliable motion and appearance cues are essential for cross-frame identity preservation. However, limited by the inherent perspective properties of 2D space, the crowd density and frequent occlusions in real-world scenes expose the fragility of these cues. We observe the natural advantage of objects being well-separated in high-dimensional space and propose a novel 2D MOT framework, "Detecting-Lifting-Tracking'' (DLT). Initially, a pre-trained detector is employed to capture 2D object information. Secondly, we introduce a Mamba Distance Estimator to obtain the distances of objects to a monocular camera with temporal consistency, achieving object-level pseudo-3D lifting. Finally, we thoroughly explore distance-aware tracking via pseudo-3D information. Specifically, we introduce a Score-Distance Hierarchical Matching and Short-Long Terms Association to enhance accurate and robust association capability. Even without appearance cues, our DLT achieves state-of-the-art performance on MOT17, MOT20, and DanceTrack, demonstrating its potential to address occlusion challenges. Haoyuan Jin, Xuesong Nie, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi |
ACM Multimedia | 3 |
| 2024 | Triplet Attention Transformer for Spatiotemporal Predictive LearningabstractSpatiotemporal predictive learning offers a self-supervised learning paradigm that enables models to learn both spatial and temporal patterns by predicting future sequences based on historical sequences. Mainstream methods are dominated by recurrent units, yet they are limited by their lack of parallelization and often underperform in real-world scenarios. To improve prediction quality while maintaining computational efficiency, we propose an innovative triplet attention transformer designed to capture both inter-frame dynamics and intra-frame static features. Specifically, the model incorporates the Triplet Attention Module (TAM), which replaces traditional recurrent units by exploring self-attention mechanisms in temporal, spatial, and channel dimensions. In this configuration: (i) temporal tokens contain abstract representations of inter-frame, facilitating the capture of inherent temporal dependencies; (ii) spatial and channel attention combine to refine the intra-frame representation by performing fine-grained interactions across spatial and channel dimensions. Alternating temporal, spatial, and channel-level attention allows our approach to learn more complex short-and long-range spatiotemporal dependencies. Extensive experiments demonstrate performance surpassing existing recurrent-based and recurrent-free methods, achieving state-of-the-art under multi-scenario examination including moving object trajectory prediction, traffic flow prediction, driving scene prediction, and human motion capture. Xuesong Nie, Xi Chen 0119, Haoyuan Jin, Zhihang Zhu, Yunfeng Yan, Donglian Qi |
WACV | 5 |
| 2024 | ADIR: Advanced domain-invariant representation via decoupling learning and information bottleneckabstractAbstract The discrepancy in data distribution between training and testing scenarios, as well as the inductive bias of convolutional neural networks towards image styles, reduces the model's generalization ability. Many unsupervised domain generalization methods based on feature decoupling suffer from an initial neglect of explicit decoupling of content and style features, resulting in content features that still contain considerable redundant information, thereby restricting improvements in generalization capability. To tackle this problem, this paper optimizes the learning process of domain‐invariant (content) features into an information compression issue, minimizing redundancy in content features. Furthermore, to enhance decoupled learning, this paper introduces innovative cross‐domain loss functions and image reconstruction modules that explicitly decouple and merge content and style across different domains. Extensive experiments demonstrate the method's significant enhancements over recent cutting‐edge approaches. Yangyang Zhong, Yunfeng Yan, Pengxin Luo, Donglian Qi |
IET Image Process. | 2 |
| 2024 | ScopeViT: Scale-Aware Vision Transformer
Xuesong Nie, Haoyuan Jin, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi |
Pattern Recognit. | 3 |
| 2024 | AHOR: Online Multi-Object Tracking With Authenticity Hierarchizing and Occlusion RecoveryabstractDespite extensive exploration of more powerful multi-object tracking (MOT) frameworks, the impact of frequent occlusion has remained a formidable challenge. In this work, we present a novel MOT framework with Authenticity Hierarchizing and Occlusion Recovery (AHOR), that strikingly handles occlusion and demonstrates superior precision and adaptability. Specifically, through an in-depth analysis of the classical tracking-by-detection (TBD) paradigm, we fully upgrade three aspects. Firstly, we propose an Existence Score that provides a more accurate depiction of detection authenticity under occlusion, enhancing the effectiveness and robustness of the hierarchical association. Secondly, we present an ingeniously devised pre-processing method in conjunction with a Recovery Intersection over Union (RIoU) for location similarity measurement, addressing the adverse effects of occlusion-induced disparity between visible and true object regions. Lastly, we introduce an Occluded Person Re-identification Module (ODReID) that extracts appearance features from the restricted visible region, overcoming the critical dependence on object quality. Results of extensive experiments demonstrate that our AHOR achieves state-of-the-art performance on MOT17, MOT20, DanceTrack, and VisDrone test sets. Haoyuan Jin, Xuesong Nie, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | A Multiscale and Multilayer Feature Extraction Network With Dual Attention for Tropical Cyclone Intensity EstimationabstractA tropical cyclone (TC) is a type of catastrophic weather encountered in the tropical or subtropical ocean, and it is of great significance to accurately estimate its intensity. Many estimation methods based on statistics have been proposed, but these methods have obvious problems, such as poor robustness and low accuracy. Therefore, in this study, a TC intensity estimation method is proposed based on satellite image data using the Xception network as a backbone. The main idea of the proposed method is to estimate the TC maximum wind speed by image feature extraction. First, a Laplacian pyramid image fusion method for the infrared (IR) and water vapor (WV) channels of satellite images is adopted to enhance the total amount of information in the basic input data of the model. Second, an optimization strategy for the depth and width of the Xception network model is proposed with the objective of reducing parameter redundancy and improving the estimation accuracy. Third, a multiscale feature extraction module and a multilayer feature fusion module are designed to realize the fusion of different features. In addition, a dual attention module is introduced to allow the model to focus on the key regions of cyclone images. Finally, the proposed network is evaluated on the HURSAT, FY-2, and Gridsat datasets. The results show that, on average, the maximum cyclone wind speed estimation errors, the mean absolute error (MAE), and root mean square error (RMSE), are 8.0% and 11.4% lower than the state-of-the-art models on the three datasets. Zhaoyang Ma 0001, Yunfeng Yan, Jianmin Lin, Dongfang Ma |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | A lightweight vehicle mounted multi-scale traffic sign detector using attention fusion pyramid
Junfan Wang, Yeting Gu, Yunfeng Yan, Mingyu Gao 0002, Zhekang Dong |
J. Supercomput. | 4 |
| 2023 | Distributed Self-Triggered Control for Frequency Restoration and Active Power Sharing in Islanded MicrogridsabstractDistributed event-triggered secondary control in microgrids has been widely investigated to improve system efficiency. But most of them are based on consecutive triggering condition monitors, which would in turn increase the computation burden of the system. To this end, this article presents distributed self-triggered algorithmic solutions to the frequency restoration control and active power sharing control of islanded microgrids. Different from event-triggered control schemes, in our self-triggered solutions, each distributed generator is equipped with a local algorithm that enables it to pre-compute the next triggering time instant according to the states at the previous one. Our starting point is to design a triggering condition with a novel estimate error. Then, the next triggering time instant is determined by solving a quadratic equation established based on the triggering condition, rather than monitoring the triggering condition consecutively. Theoretical analysis and simulation results show that the proposed distributed self-triggered secondary controllers can highly reduce the communication and computation costs simultaneously. Keng-Weng Lao, Donglian Qi, Hongxun Hui, Yunfeng Yan |
IEEE Trans. Ind. Informatics | 6 |
| 2022 | Learning Domain Adaptive Object Detection with Probabilistic TeacherabstractSelf-training for unsupervised domain adaptive object detection is a challenging task, of which the performance depends heavily on the quality of pseudo boxes. Despite the promising results, prior works have largely overlooked the uncertainty of pseudo boxes during self-training. In this paper, we present a simple yet effective framework, termed as Probabilistic Teacher (PT), which aims to capture the uncertainty of unlabeled target data from a gradually evolving teacher and guides the learning of a student in a mutually beneficial manner. Specifically, we propose to leverage the uncertainty-guided consistency training to promote classification adaptation and localization adaptation, rather than filtering pseudo boxes via an elaborate confidence threshold. In addition, we conduct anchor adaptation in parallel with localization adaptation, since anchor can be regarded as a learnable parameter. Together with this framework, we also present a novel Entropy Focal Loss (EFL) to further facilitate the uncertainty-guided self-training. Equipped with EFL, PT outperforms all previous baselines by a large margin and achieve new state-of-the-arts. Meilin Chen, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Xinchao Wang, Lei Zhang 0038, Yunfeng Yan, Donglian Qi, Yueting Zhuang, Di Xie, Shiliang Pu |
ICML | 7 |
| 2022 | Unsupervised Object Detection Pretraining with Joint Object Priors Generation and Detector LearningabstractUnsupervised pretraining methods for object detection aim to learn object discrimination and localization ability from large amounts of images. Typically, recent works design pretext tasks that supervise the detector to predict the defined object priors. They normally leverage heuristic methods to produce object priors, \emph{e.g.,} selective search, which separates the prior generation and detector learning and leads to sub-optimal solutions. In this work, we propose a novel object detection pretraining framework that could generate object priors and learn detectors jointly by generating accurate object priors from the model itself. Specifically, region priors are extracted by attention maps from the encoder, which highlights foregrounds. Instance priors are the selected high-quality output bounding boxes of the detection decoder. By assuming objects as instances in the foreground, we can generate object priors with both region and instance priors. Moreover, our object priors are jointly refined along with the detector optimization. With better object priors as supervision, the model could achieve better detection capability, which in turn promotes the object priors generation. Our method improves the competitive approaches by \textbf{+1.3 AP}, \textbf{+1.7 AP} in 1\% and 10\% COCO low-data regimes object detection. Yizhou Wang 0007, Meilin Chen, Shixiang Tang, Feng Zhu 0006, Haiyang Yang, Lei Bai 0001, Rui Zhao 0001, Yunfeng Yan, Donglian Qi, Wanli Ouyang |
NeurIPS | 8 |
| 2022 | Feature compensation network based on non-uniform quantization of channels for digital image global manipulation forensics
Yuxue Zhang, Yunfeng Yan, Guorui Feng |
Signal Process. Image Commun. | 2 |
| 2019 | Vision-based crater and rock detection using a cascade decision forestabstractBoth crater and rock detection are components of the autonomous landing and hazard avoidance technology (ALHAT) sensor suite, as craters and rocks represent the majority of landing hazards. Furthermore, places with scientific values are very probable next to craters and rocks. Unsupervised approaches, which potentially use the pattern recognition techniques of ring threshold finding, perform quickly; however, they suffer from handling small craters. The supervised pattern recognition method is more powerful but is time‐consuming. To address these issues, here, a simultaneous multi‐size crater and rock detection algorithm is studied. The authors propose a new supervised machine‐learning framework using a cascade decision forest. Sliding windows are utilised in order to search basic features, and a multi‐grained cascade structure is introduced to enhance the framework's ability to learn the representations of the features. The training time of the proposed algorithm on a PC is comparable to that of deep neural networks, and the efficiency is enhanced for a large‐scale database. The outputs of the simulation verify the effectiveness and validity of the introduced technique. Yunfeng Yan, Donglian Qi, Chaoyong Li |
IET Comput. Vis. | 1 |