EDBT 2026 Demo / reviewers in the wild / expert
Xi Chen 0119
dblp:16/3283-119
· DBLP profile ↗
30ranked-venue papers
9as first author
29since 2021 · last 2026
0009-0008-5008-4720ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 9 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Causal Prompts for Open-Vocabulary Video Instance SegmentationabstractOpen-vocabulary Video Instance Segmentation addresses the challenging task of detecting, segmenting, and tracking objects in videos, including categories not encountered during training. However, existing approaches often overlook rich temporal cues from preceding frames, limiting their ability to leverage causal context for robust open-world generalization. To bridge this gap, we propose CPOVIS, a novel framework that introduces causal prompts-dynamically propagated visual and taxonomy prompts from historical frames-to enhance temporal reasoning and semantic consistency. Built upon a Mask2Former architecture with a CLIP backbone, CPOVIS integrates three core innovations: 1) PromptCLIP, which aligns cross-modal embeddings while preserving open-vocabulary capabilities; 2) a Visual Prompt Injector that propagates object-level features to maintain spatial-temporal coherence; and 3) a Taxonomy Prompt Infuser that leverages hierarchical semantic relationships to stabilize unseen category recognition. Furthermore, we introduce a contrastive learning strategy to disentangle object representations across frames and adapt the Segment Anything Model (SAM2) to boost open-vocabulary segmentation and tracking capacity in open-vocabulary video scenarios. Extensive experiments on seven challenging open- and closed-vocabulary video segmentation benchmarks demonstrate CPOVIS's state-of-the-art performance, outperforming existing methods by significant margins. Our findings highlight the critical role of causal prompt propagation in advancing video understanding in open-world scenarios. Rongkun Zheng, Lu Qi 0001, Xi Chen 0119, Yi Wang 0074, Kun Wang 0056, Yu Qiao 0001, Hengshuang Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | UniReal: Universal Image Generation and Editing via Learning Real-world DynamicsabstractWe introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation models that effectively balance consistency and variation across frames, we propose a unifying approach that treats image-level tasks as discontinuous video generation. Specifically, we treat varying numbers of input and output images as frames, enabling seamless support for tasks such as image generation, editing, customization, composition, etc. Although designed for image-level tasks, we leverage videos as a scalable source for universal supervision. UniReal learns world dynamics from large-scale videos, demonstrating advanced capability in handling shadows, reflections, pose variation, and object interaction, while also exhibiting emergent capability for novel applications. Xi Chen 0119, He Zhang 0004, Yuqian Zhou, Soo Ye Kim, Qing Liu 0017, Yijun Li 0001, Jianming Zhang 0001, Nanxuan Zhao, Yilin Wang 0002, Zhe Lin 0001, Hengshuang Zhao |
CVPR | 1 |
| 2025 | ObjectMover: Generative Object Movement with Video PriorabstractSimple as it seems, moving an object to another location within an image is, in fact, a challenging image-editing task that requires re-harmonizing the lighting, adjusting the pose based on perspective, accurately filling occluded regions, and ensuring coherent synchronization of shadows and reflections while maintaining the object identity. In this paper, we present ObjectMover, a generative model that can perform object movement in highly challenging scenes. Our key insight is that we model this task as a sequence-to-sequence problem and fine-tune a video generation model to leverage its knowledge of consistent object generation across video frames. We show that with this approach, our model is able to adjust to complex real-world scenarios, handling extreme lighting harmonization and object effect movement. As large-scale data for object movement are unavailable, we construct a data generation pipeline using a modern game engine to synthesize high-quality data pairs. We further propose a multi-task learning strategy that enables training on real-world video data to improve the model generalization. Through extensive experiments, we demonstrate that ObjectMover achieves outstanding results and adapts well to real-world scenarios. Xin Yu 0004, Tianyu Wang 0003, Soo Ye Kim, Paul Guerrero 0001, Xi Chen 0119, Qing Liu 0017, Zhe Lin 0001, Xiaojuan Qi 0001 |
CVPR | 5 |
| 2025 | DiffDoctor: Diagnosing Image Diffusion Models Before TreatingabstractIn spite of recent progress, image diffusion models still produce artifacts. A common solution is to leverage the feedback provided by quality assessment systems or human annotators to optimize the model, where images are generally rated in their entirety. In this work, we believe problem-solving starts with identification, yielding the request that the model should be aware of not just the presence of defects in an image, but their specific locations. Motivated by this, we propose DiffDoctor, a two-stage pipeline to assist image diffusion models in generating fewer artifacts. Concretely, the first stage targets developing a robust artifact detector, for which we collect a dataset of over 1M flawed synthesized images and set up an efficient human-in-the-loop annotation process, incorporating a carefully designed class-balance strategy. The learned artifact detector is then involved in the second stage to optimize the diffusion model by providing pixel-level feedback. Extensive experiments on text-to-image diffusion models demonstrate the effectiveness of our artifact detector as well as the soundness of our diagnose-then-treat design. Xi Chen 0119, Xiaogang Xu 0002, Sihui Ji, Yu Liu 0063, Yujun Shen, Hengshuang Zhao |
ICCV | 2 |
| 2025 | iDPA: Instance Decoupled Prompt Attention for Incremental Medical Object DetectionabstractExisting prompt-based approaches have demonstrated impressive performance in continual learning, leveraging pre-trained large-scale models for classification tasks; however, the tight coupling between foreground-background information and the coupled attention between prompts and image-text tokens present significant challenges in incremental medical object detection tasks, due to the conceptual gap between medical and natural domains. To overcome these challenges, we introduce the iDPA framework, which comprises two main components: 1) Instance-level Prompt Generation (IPG), which decouples fine-grained instance-level knowledge from images and generates prompts that focus on dense predictions, and 2) Decoupled Prompt Attention (DPA), which decouples the original prompt attention, enabling a more direct and efficient transfer of prompt information while reducing memory usage and mitigating catastrophic forgetting. We collect 13 clinical, cross-modal, multi-organ, and multi-category datasets, referred to as ODinM-13, and experiments demonstrate that iDPA outperforms existing SOTA methods, with FAP improvements of f 5.44%, 4.83%, 12.88%, and 4.59% in full data, 1-shot, 10-shot, and 50-shot settings, respectively. Huahui Yi, Wei Xu 0046, Ziyuan Qin 0001, Xi Chen 0119, Kang Li 0004, Qicheng Lao |
ICML | 4 |
| 2025 | TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference OptimizationabstractRecent advancements in reinforcement learning from human feedback have shown that utilizing fine-grained token-level reward models can substantially enhance the performance of Proximal Policy Optimization (PPO) in aligning large language models. However, it is challenging to leverage such token-level reward as guidance for Direct Preference Optimization (DPO), since DPO is formulated as a sequence-level bandit problem. To address this challenge, this work decomposes the sequence-level PPO into a sequence of token-level proximal policy optimization problems and then frames the problem of token-level PPO with token-level reward guidance, from which closed-form optimal token-level policy and the corresponding token-level reward can be derived. Using the obtained reward and Bradley-Terry model, this work establishes a framework of computable loss functions with token-level reward guidance for DPO, and proposes a practical reward guidance based on the induced DPO reward. This formulation enables different tokens to exhibit varying degrees of deviation from reference policy based on their respective rewards. Experiment results demonstrate that our method achieves substantial performance improvements over DPO, with win rate gains of up to 7.5 points on MT-Bench, 6.2 points on AlpacaEval 2, and 4.3 points on Arena-Hard. Code is available at https://github.com/dvlab-research/TGDPO. Mingkang Zhu, Xi Chen 0119, Zhongdao Wang, Bei Yu 0001, Hengshuang Zhao, Jiaya Jia |
ICML | 2 |
| 2025 | MiCo: Multi-image Contrast for Reinforcement Visual ReasoningabstractThis work explores enabling Chain-of-Thought (CoT) reasoning to link visual cues across multiple images. A straightforward solution is to adapt rule-based reinforcement learning for Vision-Language Models (VLMs). However, such methods typically rely on manually curated question-answer pairs, which can be particularly challenging when dealing with fine-grained visual details and complex logic across images. Inspired by self-supervised visual representation learning, we observe that images contain inherent constraints that can serve as supervision. Based on this insight, we construct image triplets comprising two augmented views of the same image and a third, similar but distinct image. During training, the model is prompted to generate a reasoning process to compare these images (i.e., determine same or different). Then we optimize the model with rule-based reinforcement learning. Due to the high visual similarity and the presence of augmentations, the model must attend to subtle visual cues and perform logical reasoning to succeed. Experimental results demonstrate that, although trained solely on visual comparison tasks, the learned reasoning ability generalizes effectively to a wide range of questions. Without relying on any human-annotated question-answer pairs, our method achieves significant improvements on multi-image reasoning benchmarks and shows strong performance on general vision tasks. Xi Chen 0119, Mingkang Zhu, Shaoteng Liu, Xiaoyang Wu 0002, Xiaogang Xu 0002, Yu Liu 0063, Xiang Bai, Hengshuang Zhao |
NeurIPS | 1 |
| 2025 | ROSE: Remove Objects with Side Effects in VideosabstractVideo object removal has achieved advanced performance due to the recent success of video generative models. However, when addressing the side effects of objects, \textit{e.g.,} their shadows and reflections, existing works struggle to eliminate these effects for the scarcity of paired video data as supervision. This paper presents \method, termed \textbf{R}emove \textbf{O}bjects with \textbf{S}ide \textbf{E}ffects, a framework that systematically studies the object's effects on environment, which can be categorized into five common cases: shadows, reflections, light, translucency and mirror. Given the challenges of curating paired videos exhibiting the aforementioned effects, we leverage a 3D rendering engine for synthetic data generation. We carefully construct a fully-automatic pipeline for data preparation, which simulates a large-scale paired dataset with diverse scenes, objects, shooting angles, and camera trajectories. ROSE is implemented as an video inpainting model built on diffusion transformer. To localize all object-correlated areas, the entire video is fed into the model for reference-based erasing. Moreover, additional supervision is introduced to explicitly predict the areas affected by side effects, which can be revealed through the differential mask between the paired videos. To fully investigate the model performance on various side effect removal, we presents a new benchmark, dubbed ROSE-Bench, incorporating both common scenarios and the five special side effects for comprehensive evaluation. Experimental results demonstrate that \method achieves superior performance compared to existing video object erasing models and generalizes well to real-world video scenarios. Chenxuan Miao, Yutong Feng, Jianshu Zeng, Zixiang Gao, Hantang Liu, Yunfeng Yan, Donglian Qi, Xi Chen 0119, Hengshuang Zhao |
NeurIPS | 8 |
| 2025 | PlayerOne: Egocentric World SimulatorabstractWe introduce PlayerOne, the first egocentric realistic world simulator, facilitating immersive and unrestricted exploration within vividly dynamic environments. Given an egocentric scene image from the user, PlayerOne can accurately construct the corresponding world and generate egocentric videos that are strictly aligned with the real-scene human motion of the user captured by an exocentric camera. PlayerOne is trained in a coarse-to-fine pipeline that first performs pretraining on large-scale egocentric text-video pairs for coarse-level egocentric understanding, followed by finetuning on synchronous motion-video data extracted from egocentric-exocentric video datasets with our automatic construction pipeline. Besides, considering the varying importance of different components, we design a part-disentangled motion injection scheme, enabling precise control of part-level movements. In addition, we devise a joint reconstruction framework that progressively models both the 4D scene and video frames, ensuring scene consistency in the long-form video generation. Experimental results demonstrate its great generalization ability in precise control of varying human movements and world-consistent modeling of diverse scenarios. It marks the first endeavor into egocentric real-world simulation and can pave the way for the community to delve into fresh frontiers of world modeling and its diverse applications. Yuanpeng Tu, Hao Luo 0004, Xi Chen 0119, Xiang Bai, Fan Wang 0019, Hengshuang Zhao |
NeurIPS | 3 |
| 2025 | DiffCamera: Arbitrary Refocusing on ImagesabstractThe depth-of-field (DoF) effect, which introduces aesthetically pleasing blur, enhances photographic quality but is fixed and difficult to modify once the image has been created. This becomes problematic when the applied blur is undesirable (e.g., the subject is out of focus). To address this, we propose DiffCamera, a model that enables flexible refocusing of a created image conditioned on an arbitrary new focus point and a blur level. Specifically, we design a diffusion transformer framework for refocusing learning. However, the training requires pairs of data with different focus planes and bokeh levels in the same scene, which are hard to acquire. To overcome this limitation, we develop a simulation-based pipeline to generate large-scale image pairs with varying focus planes and bokeh levels. With the simulated data, we find that training with only a vanilla diffusion objective often leads to incorrect DoF behaviors due to the complexity of the task. This requires a stronger constraint during training. Inspired by the photographic principle that photos of different focus planes can be linearly blended into a multi-focus image, we propose a stacking constraint during training to enforce precise DoF manipulation. This constraint enhances model training by imposing physically grounded refocusing behavior that the focusing results should be faithfully aligned with the scene structure and the camera conditions so that they can be combined into the correct multi-focus image. We also construct a benchmark to evaluate the effectiveness of our refocusing model. Extensive experiments demonstrate that DiffCamera supports stable refocusing across a wide range of scenes, providing unprecedented control over DoF adjustments for photography and generative AI applications. Xi Chen 0119, Xiaogang Xu 0002, Yu Liu 0063, Hengshuang Zhao |
SIGGRAPH Asia | 2 |
| 2025 | AnyDoor: Zero-Shot Image Customization With Region-to-Region ReferenceabstractThis work presents AnyDoor, a diffusion-based image generator with the power to teleport target objects to new scenes at user-specified locations with desired shapes. Instead of tuning parameters for each object, our model is trained only once and effortlessly generalizes to diverse object-scene combinations at the inference stage. Such a challenging zero-shot setting requires an adequate characterization of a certain object. To this end, we leverage the powerful self-supervised image encoder (i.e., DINOv2) to extract the discriminative dentity feature of the target object. Besides, we complement the identity feature with detail features, which are carefully designed to maintain appearance details yet allow versatile local variations (e.g., lighting, orientation, posture, etc.), supporting the object in favorably blending with different surroundings. We further propose to borrow knowledge from video datasets, where we can observe various forms (i.e., along the time axis) of a single object, leading to stronger model generalizability and robustness. Starting from the task of object insertion, we further extend the framework of AnyDoor to a general solution with region-to-region image reference. With the different definitions of the source region and target region, the tasks of object insertion, object removal, and image variation could be integrated into one model without introducing extra parameters. In addition, we investigate incorporating other conditions like the mask, pose skeleton, and depth map as additional guidance to achieve more controllable generation. Xi Chen 0119, Lianghua Huang, Yu Liu 0063, Yujun Shen, Deli Zhao, Hengshuang Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | UniDetector: Towards Universal Object Detection With Heterogeneous SupervisionabstractIn this paper, we formally address universal object detection, which aims to detect every category in every scene. The dependence on human annotations, the limited visual information, and the novel categories in open world severely restrict the universality of detectors. We propose UniDetector, a universal object detector that recognizes enormous categories in the open world. The critical points for UniDetector are: 1) it leverages images of multiple sources and heterogeneous label spaces in training through image-text alignment, which guarantees sufficient information for universal representations. 2) it involves heterogeneous supervision training, which alleviates the dependence on the limited fully-labeled images. 3) it generalizes to open world easily while keeping the balance between seen and unseen classes. 4) it further promotes generalizing to novel categories through our proposed decoupling training manner and probability calibration. These contributions allow UniDetector to detect over 7 k categories, the largest measurable size so far, with only about 500 classes participating in training. Our UniDetector behaves the strong zero-shot ability on large-vocabulary datasets - it surpasses supervised baselines by more than 5% without seeing any corresponding images. On 13 detection datasets with various scenes, UniDetector also achieves state-of-the-art performance with only a 3% amount of training data. Zhenyu Wang 0005, Yali Li 0001, Xi Chen 0119, Ser-Nam Lim, Antonio Torralba 0001, Hengshuang Zhao, Shengjin Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Wavelet-Driven Spatiotemporal Predictive Learning: Bridging Frequency and Time VariationsabstractSpatiotemporal predictive learning is a paradigm that empowers models to learn spatial and temporal patterns by predicting future frames from past frames in an unsupervised manner. This method typically uses recurrent units to capture long-term dependencies, but these units often come with high computational costs and limited performance in real-world scenes. This paper presents an innovative Wavelet-based SpatioTemporal (WaST) framework, which extracts and adaptively controls both low and high-frequency components at image and feature levels via 3D discrete wavelet transform for faster processing while maintaining high-quality predictions. We propose a Time-Frequency Aware Translator uniquely crafted to efficiently learn short- and long-range spatiotemporal information by individually modeling spatial frequency and temporal variations. Meanwhile, we design a wavelet-domain High-Frequency Focal Loss that effectively supervises high-frequency variations. Extensive experiments across various real-world scenarios, such as driving scene prediction, traffic flow prediction, human motion capture, and weather forecasting, demonstrate that our proposed WaST achieves state-of-the-art performance over various spatiotemporal prediction methods. Xuesong Nie, Yunfeng Yan, Siyuan Li 0002, Cheng Tan 0012, Xi Chen 0119, Haoyuan Jin, Zhihang Zhu, Stan Z. Li, Donglian Qi |
AAAI | 5 |
| 2024 | AnyDoor: Zero-shot Object-level Image CustomizationabstractThis work presents AnyDoor, a diffusion-based image generator with the power to teleport target objects to new scenes at user-specified locations with desired shapes. Instead of tuning parameters for each object, our model is trained only once and effortlessly generalizes to diverse object-scene combinations at the inference stage. Such a challenging zero-shot setting requires an adequate characterization of a certain object. To this end, we complement the commonly used identity feature with detail features, which are carefully designed to maintain appearance details yet allow versatile local variations (e.g., lighting, orientation, posture, etc.), supporting the object in favorably blending with different surroundings. We further propose to borrow knowledge from video datasets, where we can observe various forms (i.e., along the time axis) of a single object, leading to stronger model generalizability and robustness. Extensive experiments demonstrate the superiority of our approach over existing alternatives as well as its great potential in real-world applications, such as virtual try-on, shape editing, and object swapping. Code is released at github.com/ali-vilab/AnyDoor. Xi Chen 0119, Lianghua Huang, Yu Liu 0063, Yujun Shen, Deli Zhao, Hengshuang Zhao |
CVPR | 1 |
| 2024 | PredToken: Predicting Unknown Tokens and Beyond with Coarse-to-Fine Iterative DecodingabstractPredictive learning models, which aim to predict future frames based on past observations, are crucial to constructing world models. These models need to maintain low-level consistency and capture high-level dynamics in unannotated spatiotemporal data. Transitioning from frame-wise to token-wise prediction presents a viable strategy for addressing these needs. How to improve token representation and optimize token decoding presents significant challenges. This paper introduces PredToken, a novel predictive framework that addresses these issues by decoupling space-time tokens into distinct components for iterative cascaded decoding. Concretely, we first design a “decomposition, quantization, and reconstruction” schema based on VQGAN to improve the token representation. This scheme disentangles low- and high-frequency representations and employs a dimension-aware quantization model, allowing more low-level details to be preserved. Building on this, we present a “coarse-to-fine iterative decoding” method. It leverages dynamic soft decoding to refine coarse tokens and static soft decoding for fine tokens, enabling more high-level dynamics to be captured. These designs make Pred-Token produce high-quality predictions. Extensive experiments demonstrate the superiority of our method on various real-world spatiotemporal predictive benchmarks. Furthermore, PredToken can also be extended to other visual generative tasks to yield realistic outcomes. Xuesong Nie, Haoyuan Jin, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi |
CVPR | 4 |
| 2024 | LivePhoto: Real Image Animation with Text-Guided Motion Control
Xi Chen 0119, Yutong Feng, Yu Liu 0063, Yujun Shen, Hengshuang Zhao |
ECCV (18) | 1 |
| 2024 | OpenIns3D: Snap and Lookup for 3D Open-Vocabulary Instance Segmentation
Zhening Huang, Xiaoyang Wu 0002, Xi Chen 0119, Hengshuang Zhao, Lei Zhu 0003, Joan Lasenby |
ECCV (62) | 3 |
| 2024 | LogoSticker: Inserting Logos Into Diffusion Models for Customized Generation
Mingkang Zhu, Xi Chen 0119, Zhongdao Wang, Hengshuang Zhao, Jiaya Jia |
ECCV (60) | 2 |
| 2024 | SAMP: Adapting Segment Anything Model for Pose EstimationabstractSegment Anything Model (SAM) exhibits superior performance for segmentation. Many follow-up works explore adapting this powerful model to specific domains. However, those works mainly focus on different sub-tasks of segmentation. The cross-task generalization ability of SAM is still not explored. In this paper, we propose SAMP (SAM for Pose), which makes the first attempt to adapt SAM for pose estimation. We observe that SAM could segment different human parts with specific prompts, proving that it contains the knowledge to understand the human structure. Considering that localizing keypoints requires fine-grained perceptual capabilities, we design a Detail-aware Adapter (DA-Adapter), which complements the features of the SAM encoder with multi-scale feature fusion and multi-level supervision. Experimental results demonstrate that SAMP achieves novel state-of-the-art against previously specifically designed pose estimation methods. Specifically, with ViT-B backbone, SAMP achieves 78.1% AP on the COCO val2017, 77.1% AP on the COCO test-dev2017, and 70.5% AP on the CrowdPose dataset. Zhihang Zhu, Yunfeng Yan, Haoyuan Jin, Xuesong Nie, Donglian Qi, Xi Chen 0119 |
ICME | 7 |
| 2024 | Object-Level Pseudo-3D Lifting for Distance-Aware TrackingabstractMulti-object tracking (MOT) is a pivotal task for media interpretation, where reliable motion and appearance cues are essential for cross-frame identity preservation. However, limited by the inherent perspective properties of 2D space, the crowd density and frequent occlusions in real-world scenes expose the fragility of these cues. We observe the natural advantage of objects being well-separated in high-dimensional space and propose a novel 2D MOT framework, "Detecting-Lifting-Tracking'' (DLT). Initially, a pre-trained detector is employed to capture 2D object information. Secondly, we introduce a Mamba Distance Estimator to obtain the distances of objects to a monocular camera with temporal consistency, achieving object-level pseudo-3D lifting. Finally, we thoroughly explore distance-aware tracking via pseudo-3D information. Specifically, we introduce a Score-Distance Hierarchical Matching and Short-Long Terms Association to enhance accurate and robust association capability. Even without appearance cues, our DLT achieves state-of-the-art performance on MOT17, MOT20, and DanceTrack, demonstrating its potential to address occlusion challenges. Haoyuan Jin, Xuesong Nie, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi |
ACM Multimedia | 4 |
| 2024 | Zero-shot Image Editing with Reference ImitationabstractImage editing serves as a practical yet challenging task considering the diverse demands from users, where one of the hardest parts is to precisely describe how the edited image should look like. In this work, we present a new form of editing, termed imitative editing, to help users exercise their creativity more conveniently. Concretely, to edit an image region of interest, users are free to directly draw inspiration from some in-the-wild references (e.g., some relative pictures come across online), without having to cope with the fit between the reference and the source. Such a design requires the system to automatically figure out what to expect from the reference to perform the editing. For this purpose, we propose a generative training framework, dubbed MimicBrush, which randomly selects two frames from a video clip, masks some regions of one frame, and learns to recover the masked regions using the information from the other frame. That way, our model, developed from a diffusion prior, is able to capture the semantic correspondence between separate images in a self-supervised manner. We experimentally show the effectiveness of our method under various test cases as well as its superiority over existing alternatives. We also construct a benchmark to facilitate further research. Xi Chen 0119, Yutong Feng, Yu Liu 0063, Yujun Shen, Hengshuang Zhao |
NeurIPS | 1 |
| 2024 | Triplet Attention Transformer for Spatiotemporal Predictive LearningabstractSpatiotemporal predictive learning offers a self-supervised learning paradigm that enables models to learn both spatial and temporal patterns by predicting future sequences based on historical sequences. Mainstream methods are dominated by recurrent units, yet they are limited by their lack of parallelization and often underperform in real-world scenarios. To improve prediction quality while maintaining computational efficiency, we propose an innovative triplet attention transformer designed to capture both inter-frame dynamics and intra-frame static features. Specifically, the model incorporates the Triplet Attention Module (TAM), which replaces traditional recurrent units by exploring self-attention mechanisms in temporal, spatial, and channel dimensions. In this configuration: (i) temporal tokens contain abstract representations of inter-frame, facilitating the capture of inherent temporal dependencies; (ii) spatial and channel attention combine to refine the intra-frame representation by performing fine-grained interactions across spatial and channel dimensions. Alternating temporal, spatial, and channel-level attention allows our approach to learn more complex short-and long-range spatiotemporal dependencies. Extensive experiments demonstrate performance surpassing existing recurrent-based and recurrent-free methods, achieving state-of-the-art under multi-scenario examination including moving object trajectory prediction, traffic flow prediction, driving scene prediction, and human motion capture. Xuesong Nie, Xi Chen 0119, Haoyuan Jin, Zhihang Zhu, Yunfeng Yan, Donglian Qi |
WACV | 2 |
| 2024 | ScopeViT: Scale-Aware Vision Transformer
Xuesong Nie, Haoyuan Jin, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi |
Pattern Recognit. | 4 |
| 2024 | AHOR: Online Multi-Object Tracking With Authenticity Hierarchizing and Occlusion RecoveryabstractDespite extensive exploration of more powerful multi-object tracking (MOT) frameworks, the impact of frequent occlusion has remained a formidable challenge. In this work, we present a novel MOT framework with Authenticity Hierarchizing and Occlusion Recovery (AHOR), that strikingly handles occlusion and demonstrates superior precision and adaptability. Specifically, through an in-depth analysis of the classical tracking-by-detection (TBD) paradigm, we fully upgrade three aspects. Firstly, we propose an Existence Score that provides a more accurate depiction of detection authenticity under occlusion, enhancing the effectiveness and robustness of the hierarchical association. Secondly, we present an ingeniously devised pre-processing method in conjunction with a Recovery Intersection over Union (RIoU) for location similarity measurement, addressing the adverse effects of occlusion-induced disparity between visible and true object regions. Lastly, we introduce an Occluded Person Re-identification Module (ODReID) that extracts appearance features from the restricted visible region, overcoming the critical dependence on object quality. Results of extensive experiments demonstrate that our AHOR achieves state-of-the-art performance on MOT17, MOT20, DanceTrack, and VisDrone test sets. Haoyuan Jin, Xuesong Nie, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Detecting Everything in the Open World: Towards Universal Object DetectionabstractIn this paper, we formally address universal object detection, which aims to detect every scene and predict every category. The dependence on human annotations, the limited visual information, and the novel categories in the open world severely restrict the universality of traditional detectors. We propose UniDetector, a universal object detector that has the ability to recognize enormous categories in the open world. The critical points for the universality of UniDetector are: 1) it leverages images of multiple sources and heterogeneous label spaces for training through the alignment of image and text spaces, which guarantees sufficient information for universal representations. 2) it generalizes to the open world easily while keeping the balance between seen and unseen classes, thanks to abundant information from both vision and language modalities. 3) it further promotes the generalization ability to novel categories through our proposed decoupling training manner and probability calibration. These contributions allow UniDetector to detect over 7k categories, the largest measurable category size so far, with only about 500 classes participating in training. Our UniDetector behaves the strong zero-shot generalization ability on largevocabulary datasets - it surpasses the traditional supervised baselines by more than 4% on average without seeing any corresponding images. On 13 public detection datasets with various scenes, UniDetector also achieves state-of-the-art performance with only a 3% amount of training data.11Codes are available at https://github.com/zhenyuw16/UniDetector. Zhenyu Wang 0005, Yali Li 0001, Xi Chen 0119, Ser-Nam Lim, Antonio Torralba 0001, Hengshuang Zhao, Shengjin Wang |
CVPR | 3 |
| 2023 | Open-vocabulary Panoptic Segmentation with Embedding ModulationabstractOpen-vocabulary image segmentation is attracting increasing attention due to its critical applications in the real world. Traditional closed-vocabulary segmentation methods are not able to characterize novel objects, whereas several recent open-vocabulary attempts obtain unsatisfactory results, i.e., notable performance reduction on the closedvocabulary and massive demand for extra data. To this end, we propose OPSNet, an omnipotent and data-efficient framework for Open-vocabulary Panoptic Segmentation. Specifically, the exquisitely designed Embedding Modulation module, together with several meticulous components, enables adequate embedding enhancement and information exchange between the segmentation model and the visual-linguistic well-aligned CLIP encoder, resulting in superior segmentation performance under both open- and closed-vocabulary settings with much fewer need of additional data. Extensive experimental evaluations are conducted across multiple datasets (e.g., COCO, ADE20K, Cityscapes, and PascalContext) under various circumstances, where the proposed OPSNet achieves state-of-theart results, which demonstrates the effectiveness and generality of the proposed approach. The project page is https://opsnet-page.github.io. Xi Chen 0119, Shuang Li 0013, Ser-Nam Lim, Antonio Torralba 0001, Hengshuang Zhao |
ICCV | 1 |
| 2023 | Uni3DETR: Unified 3D Detection TransformerabstractExisting point cloud based 3D detectors are designed for the particular scene, either indoor or outdoor ones. Because of the substantial differences in object distribution and point density within point clouds collected from various environments, coupled with the intricate nature of 3D metrics, there is still a lack of a unified network architecture that can accommodate diverse scenes. In this paper, we propose Uni3DETR, a unified 3D detector that addresses indoor and outdoor 3D detection within the same framework. Specifically, we employ the detection transformer with point-voxel interaction for object prediction, which leverages voxel features and points for cross-attention and behaves resistant to the discrepancies from data. We then propose the mixture of query points, which sufficiently exploits global information for dense small-range indoor scenes and local information for large-range sparse outdoor ones. Furthermore, our proposed decoupled IoU provides an easy-to-optimize training target for localization by disentangling the $xy$ and $z$ space. Extensive experiments validate that Uni3DETR exhibits excellent performance consistently on both indoor and outdoor 3D detection. In contrast to previous specialized detectors, which may perform well on some particular datasets but suffer a substantial degradation on different scenes, Uni3DETR demonstrates the strong generalization ability under heterogeneous conditions (Fig. 1). Zhenyu Wang 0005, Yali Li 0001, Xi Chen 0119, Hengshuang Zhao, Shengjin Wang |
NeurIPS | 3 |
| 2022 | FocalClick: Towards Practical Interactive Image SegmentationabstractInteractive segmentation allows users to extract target masks by making positive/negative clicks. Although explored by many previous works, there is still a gap between academic approaches and industrial needs: first, existing models are not efficient enough to work on low-power devices; second, they perform poorly when used to refine preexisting masks as they could not avoid destroying the correct part. FocalClick solves both issues at once by predicting and updating the mask in localized areas. For higher efficiency, we decompose the slow prediction on the entire image into two fast inferences on small crops: a coarse segmentation on the Target Crop, and a local refinement on the Focus Crop. To make the model work with preexisting masks, we formulate a sub-task termed Inter-active Mask Correction, and propose Progressive Merge as the solution. Progressive Merge exploits morphological information to decide where to preserve and where to update, enabling users to refine any preexisting mask effectively. FocalClick achieves competitive results against SOTA methods with significantly smaller FLOPs. It also shows significant superiority when making corrections on preexisting masks. Code and data will be released at github.com/XavierCHEN34/ClickSEG Xi Chen 0119, Zhiyan Zhao, Manni Duan, Donglian Qi, Hengshuang Zhao |
CVPR | 1 |
| 2022 | iNL: Implicit non-local network
Yifeng Han, Xi Chen 0119, Songjie Zhang, Donglian Qi |
Neurocomputing | 2 |
| 2020 | State-Aware Tracker for Real-Time Video Object SegmentationabstractIn this work, we address the task of semi-supervised video object segmentation (VOS) and explore how to make efficient use of video property to tackle the challenge of semi-supervision. We propose a novel pipeline called State-Aware Tracker (SAT), which can produce accurate segmentation results with real-time speed. For higher efficiency, SAT takes advantage of the inter-frame consistency and deals with each target object as a tracklet. For more stable and robust performance over video sequences, SAT gets awareness for each state and makes self-adaptation via two feedback loops. One loop assists SAT in generating more stable tracklets. The other loop helps to construct a more robust and holistic target representation. SAT achieves a promising result of 72.3% J&F mean with 39 FPS on DAVIS 2017-Val dataset, which shows a decent trade-off between efficiency and accuracy. Xi Chen 0119, Zuoxin Li, Ye Yuan 0012, Gang Yu 0002, Jianxin Shen, Donglian Qi |
CVPR | 1 |