EDBT 2026 Demo / reviewers in the wild / expert
Xu Zou 0002
dblp:220/4186-2
· DBLP profile ↗
35ranked-venue papers
2as first author
34since 2021 · last 2026
0000-0002-0251-7404ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 1 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FineEdu: a fine-grained class students behavior understanding dataset with jointly action and attention annotations
Zhijun Zhang 0009, Ziyue Feng, Zhiying Yan, Xu Zou 0002, Sheng Zhong 0001 |
Neural Comput. Appl. | 4 |
| 2026 | Supervisory feedback for high-resolution low-textured large-scale multi-view stereo
Yongjian Liao, Shixiang Huang, Chunxi Li, Jiahuan Zhou, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002 |
Pattern Recognit. | 9 |
| 2025 | DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving ScenesabstractVision-centric autonomous driving systems require diverse data for robust training and evaluation, which can be augmented by manipulating object positions and appearances within existing scene captures. While recent advancements in diffusion models have shown promise in video editing, their application to object manipulation in driving scenarios remains challenging due to imprecise positional control and difficulties in preserving high-fidelity object appearances. To address these challenges in position and appearance control, we introduce DriveEditor, the first diffusion-based framework for object editing in driving videos. DriveEditor offers a unified framework for comprehensive object editing operations, including repositioning, replacement, deletion, and insertion. These diverse manipulations are all achieved through a shared set of varying inputs, processed by identical position control and appearance maintenance modules. The position control module projects the given 3D bounding box while preserving depth information and hierarchically injects it into the diffusion process, enabling precise control over object position and orientation. The appearance maintenance module preserves consistent attributes with a single reference image by employing a three-tiered approach: low-level detail preservation, high-level semantic maintenance, and the integration of 3D priors from a novel view synthesis model. Extensive qualitative and quantitative evaluations on the nuScenes dataset demonstrate DriveEditor's exceptional fidelity and controllability in generating diverse driving scene edits, as well as its remarkable ability to facilitate downstream tasks. Yiyuan Liang 0001, Zhiying Yan, Jiahuan Zhou, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002 |
AAAI | 7 |
| 2025 | Incremental Object Keypoint LearningabstractExisting progress in object keypoint estimation primarily benefits from the conventional supervised learning paradigm based on numerous data labeled with pre-defined keypoints. However, these well-trained models can hardly detect the undefined new keypoints in test time, which largely hinders their feasibility for diverse downstream tasks. To handle this, various solutions are explored but still suffer from either limited generalizability or transferability. Therefore, in this paper, we explore a novel keypoint learning paradigm in that we only annotate new keypoints in the new data and incrementally train the model, without retaining any old data, called Incremental object Keypoint Learning (IKL). A two-stage learning scheme as a novel baseline tailored to IKL is developed. In the first Knowledge Association stage, given the data labeled with only new keypoints, an auxiliary KA-Net is trained to automatically associate the old keypoints to these new ones based on their spatial and intrinsic anatomical relations. In the second Mutual Promotion stage, based on a keypoint-oriented spatial distillation loss, we jointly leverage the auxiliary KA-Net and the old model for knowledge consolidation to mutually promote the estimation of all old and new keypoints. Owing to the investigation of the correlations between new and old keypoints, our proposed method can not just effectively mitigate the catastrophic forgetting of old keypoints, but may even further improve the estimation of the old ones and achieve a positive transfer beyond anti-forgetting. Such an observation has been solidly verified by extensive experiments on different keypoint datasets, where our method exhibits superiority in alleviating the forgetting issue and boosting performance while enjoying labeling efficiency even under the low-shot data regime. Mingfu Liang, Jiahuan Zhou, Xu Zou 0002, Ying Wu 0001 |
CVPR | 3 |
| 2025 | STOP: Integrated Spatial-Temporal Dynamic Prompting for Video UnderstandingabstractPre-trained on tremendous image-text pairs, vision-language models like CLIP have demonstrated promising zero-shot generalization across numerous image-based tasks. However, extending these capabilities to video tasks remains challenging due to limited labeled video data and high training costs. Recent video prompting methods attempt to adapt CLIP for video tasks by introducing learnable prompts, but they typically rely on a single static prompt for all video sequences, overlooking the diverse temporal dynamics and spatial variations that exist across frames. This limitation significantly hinders the model’s ability to capture essential temporal information for effective video understanding. To address this, we propose an integrated Spatial-TempOral dynamic Prompting (STOP) model which consists of two complementary modules, the intra-frame spatial prompting and inter-frame temporal prompting. Our intra-frame spatial prompts are designed to adaptively highlight discriminative regions within each frame by leveraging intra-frame attention and temporal variation, allowing the model to focus on areas with substantial temporal dynamics and capture fine-grained spatial details. Additionally, to highlight the varying importance of frames for video understanding, we further introduce inter-frame temporal prompts, dynamically inserting prompts between frames with high temporal variance as measured by frame similarity. This enables the model to prioritize key frames and enhances its capacity to understand temporal dependencies across sequences. Extensive experiments on various video benchmarks demonstrate that STOP consistently achieves superior performance against state-of-the-art methods. The code is available at https://github.com/zhoujiahuan1991/CVPR2025-STOP. Kunlun Xu, Xu Zou 0002, Yuxin Peng 0001, Jiahuan Zhou |
CVPR | 4 |
| 2025 | Self-Reinforcing Prototype Evolution with Dual-Knowledge Cooperation for Semi-Supervised Lifelong Person Re-IdentificationabstractCurrent lifelong person re-identification (LReID) methods predominantly rely on fully labeled data streams. However, in real-world scenarios where annotation resources are limited, a vast amount of unlabeled data coexists with scarce labeled samples, leading to the Semi-Supervised LReID (Semi-LReID) problem where LReID methods suffer severe performance degradation. Existing LReID methods, even when combined with semi-supervised strategies, suffer from limited long-term adaptation performance due to struggling with the noisy knowledge occurring during unlabeled data utilization. In this paper, we pioneer the investigation of Semi-LReID, introducing a novel Self-Reinforcing Prototype Evolution with Dual-Knowledge Cooperation framework (SPRED). Our key innovation lies in establishing a self-reinforcing cycle between dynamic prototype-guided pseudo-label generation and new-old knowledge collaborative purification to enhance the utilization of unlabeled data. Specifically, learnable identity prototypes are introduced to dynamically capture the identity distributions and generate high-quality pseudo-labels. Then, the dual-knowledge cooperation scheme integrates current model specialization and historical model generalization, refining noisy pseudo-labels. Through this cyclic design, reliable pseudo-labels are progressively mined to improve current-stage learning and ensure positive knowledge propagation over long-term learning. Experiments on the established Semi-LReID benchmarks show that our SPRED achieves state-of-the-art performance. Our source code is available at https://github.com/zhoujiahuan1991/ICCV2025-SPRED Kunlun Xu, Fan Zhuo, Jiangmeng Li, Xu Zou 0002, Jiahuan Zhou |
ICCV | 4 |
| 2025 | High-dimension Prototype is a Better Incremental Object Detection LearnerabstractIncremental object detection (IOD), surpassing simple classification, requires the simultaneous overcoming of catastrophic forgetting in both recognition and localization tasks, primarily due to the significantly higher feature space complexity. Integrating Knowledge Distillation (KD) would mitigate the occurrence of catastrophic forgetting. However, the challenge of knowledge shift caused by invisible previous task data hampers existing KD-based methods, leading to limited improvements in IOD performance. This paper aims to alleviate knowledge shift by enhancing the accuracy and granularity in describing complex high-dimensional feature spaces. To this end, we put forth a novel higher-dimension-prototype learning approach for KD-based IOD, enabling a more flexible, accurate, and fine-grained representation of feature distributions without the need to retain any previous task data. Existing prototype learning methods calculate feature centroids or statistical Gaussian distributions as prototypes, disregarding actual irregular distribution information or leading to inter-class feature overlap, which is not directly applicable to the more difficult task of IOD with complex feature space. To address the above issue, we propose a Gaussian Mixture Distribution-based Prototype (GMDP), which explicitly models the distribution relationships of different classes by directly measuring the likelihood of embedding from new and old models into class distribution prototypes in a higher dimension manner. Specifically, GMDP dynamically adapts the component weights and corresponding means/variances of class distribution prototypes to represent both intra-class and inter-class variability more accurately. Progressing into a new task, GMDP constrains the distance between the distribution of new and previous task classes, minimizing overlap with existing classes and thus striking a balance between stability and adaptability. GMDP can be readily integrated into existing IOD methods to enhance performance further. Extensive experiments on the PASCAL VOC and MS-COCO show that our method consistently exceeds four baselines by a large margin and significantly outperforms other SOTA results under various settings. Tianming Zhao 0003, Tao Zhang 0147, Guodong Wang 0001, Luxin Yan, Sheng Zhong 0001, Jiahuan Zhou, Xu Zou 0002 |
ICLR | 9 |
| 2025 | Divide-And-Conquer: Dual-Hierarchical Optimization for Semantic 4D Gaussian SpattingabstractSemantic 4D Gaussians can be used for reconstructing and understanding dynamic scenes, with temporal variations than static scenes. Directly applying static methods to understand dynamic scenes will fail to capture the temporal features. Few works focus on dynamic scene understanding based on Gaussian Splatting, since once the same update strategy is employed for both dynamic and static parts, regardless of the distinction and interaction between Gaussians, significant artifacts and noise appear. We propose Dual-Hierarchical Optimization (DHO), which consists of Hierarchical Gaussian Flow and Hierarchical Gaussian Guidance in a divide-and-conquer manner. The former implements effective division of static and dynamic rendering and features. The latter helps to mitigate the issue of dynamic foreground rendering distortion in textured complex scenes. Extensive experiments show that our method consistently outperforms the baselines on both synthetic and real-world datasets, and supports various downstream tasks. Project Page: https://sweety-yan.github.io/DHO/ Zhiying Yan, Yiyuan Liang 0001, Shilv Cai, Tao Zhang 0147, Sheng Zhong 0001, Luxin Yan, Xu Zou 0002 |
ICME | 7 |
| 2025 | GAPrompt: Geometry-Aware Point Cloud Prompt for 3D Vision ModelabstractPre-trained 3D vision models have gained significant attention for their promising performance on point cloud data. However, fully fine-tuning these models for downstream tasks is computationally expensive and storage-intensive. Existing parameter-efficient fine-tuning (PEFT) approaches, which focus primarily on input token prompting, struggle to achieve competitive performance due to their limited ability to capture the geometric information inherent in point clouds. To address this challenge, we propose a novel Geometry-Aware Point Cloud Prompt (GAPrompt) that leverages geometric cues to enhance the adaptability of 3D vision models. First, we introduce a Point Prompt that serves as an auxiliary input alongside the original point cloud, explicitly guiding the model to capture fine-grained geometric details. Additionally, we present a Point Shift Prompter designed to extract global shape information from the point cloud, enabling instance-specific geometric adjustments at the input level. Moreover, our proposed Prompt Propagation mechanism incorporates the shape information into the model's feature extraction process, further strengthening its ability to capture essential geometric characteristics. Extensive experiments demonstrate that GAPrompt significantly outperforms state-of-the-art PEFT methods and achieves competitive results compared to full fine-tuning on various benchmarks, while utilizing only 2.19\% of trainable parameters. Zixiang Ai, Yuanhang Lei, Zhenyu Cui, Xu Zou 0002, Jiahuan Zhou |
ICML | 5 |
| 2025 | Token Coordinated Prompt Attention is Needed for Visual PromptingabstractVisual prompting techniques are widely used to efficiently fine-tune pretrained Vision Transformers (ViT) by learning a small set of shared prompts for all tokens. However, existing methods overlook the unique roles of different tokens in conveying discriminative information and interact with all tokens using the same prompts, thereby limiting the representational capacity of ViT. This often leads to indistinguishable and biased prompt-extracted features, hindering performance. To address this issue, we propose a plug-and-play Token Coordinated Prompt Attention (TCPA) module, which assigns specific coordinated prompts to different tokens for attention-based interactions. Firstly, recognizing the distinct functions of CLS and image tokens-global information aggregation and local feature extraction, we disentangle the prompts into CLS Prompts and Image Prompts, which interact exclusively with CLS tokens and image tokens through attention mechanisms. This enhances their respective discriminative abilities. Furthermore, as different image tokens correspond to distinct image patches and contain diverse information, we employ a matching function to automatically assign coordinated prompts to individual tokens. This enables more precise attention interactions, improving the diversity and representational capacity of the extracted features. Extensive experiments across various benchmarks demonstrate that TCPA significantly enhances the diversity and discriminative power of the extracted features. Xu Zou 0002, Gang Hua 0001, Jiahuan Zhou |
ICML | 2 |
| 2025 | Componential Prompt-Knowledge Alignment for Domain Incremental LearningabstractDomain Incremental Learning (DIL) aims to learn from non-stationary data streams across domains while retaining and utilizing past knowledge. Although prompt-based methods effectively store multi-domain knowledge in prompt parameters and obtain advanced performance through cross-domain prompt fusion, we reveal an intrinsic limitation: component-wise misalignment between domain-specific prompts leads to conflicting knowledge integration and degraded predictions. This arises from the random positioning of knowledge components within prompts, where irrelevant component fusion introduces interference. To address this, we propose Componential Prompt-Knowledge Alignment (KA-Prompt), a novel prompt-based DIL method that introduces component-aware prompt-knowledge alignment during training, significantly improving both the learning and inference capacity of the model. KA-Prompt operates in two phases: (1) Initial Componential Structure Configuring, where a set of old prompts containing knowledge relevant to the new domain are mined via greedy search, which is then exploited to initialize new prompts to achieve reusable knowledge transfer and establish intrinsic alignment between new and old prompts. (2) Online Alignment Preservation, which dynamically identifies the target old prompts and applies adaptive componential consistency constraints as new prompts evolve. Extensive experiments on DIL benchmarks demonstrate the effectiveness of our KA-Prompt. Our source code is available at https://github.com/zhoujiahuan1991/ICML2025-KA-Prompt. Kunlun Xu, Xu Zou 0002, Gang Hua 0001, Jiahuan Zhou |
ICML | 2 |
| 2025 | State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video UnderstandingabstractRecently, pre-trained state space models have shown great potential for video classification, which sequentially compresses visual tokens in videos with linear complexity, thereby improving the processing efficiency of video data while maintaining high performance. To apply powerful pre-trained models to downstream tasks, prompt learning is proposed to achieve efficient downstream task adaptation with only a small number of fine-tuned parameters. However, the sequentially compressed visual prompt tokens fail to capture the spatial and temporal contextual information in the video, thus limiting the effective propagation of spatial information within a video frame and temporal information between frames in the state compression model and the extraction of discriminative information. To tackle the above issue, we proposed a State Space Prompting (SSP) method for video understanding, which combines intra-frame and inter-frame prompts to aggregate and propagate key spatiotemporal information in the video. Specifically, an Intra-Frame Gathering (IFG) module is designed to aggregate spatial key information within each frame. Besides, an Inter-Frame Spreading (IFS) module is designed to spread discriminative spatio-temporal information across different frames. By adaptively balancing and compressing key spatio-temporal information within and between frames, our SSP effectively propagates discriminative information in videos in a complementary manner. Extensive experiments on four video benchmark datasets verify that our SSP significantly outperforms existing SOTA methods by 2.76\% on average while reducing the overhead of fine-tuning parameters. Jiahuan Zhou, Zhenyu Cui, Xu Zou 0002, Gang Hua 0001 |
NeurIPS | 5 |
| 2025 | Class-aware Domain Knowledge Fusion and Fission for Continual Test-Time AdaptationabstractContinual Test-Time Adaptation (CTTA) aims to quickly fine-tune the model during the test phase so that it can adapt to multiple unknown downstream domain distributions without pre-acquiring downstream domain data.
To this end, existing advanced CTTA methods mainly reduce the catastrophic forgetting of historical knowledge caused by irregular switching of downstream domain data by restoring the initial model or reusing historical models. However, these methods are usually accompanied by serious insufficient learning of new knowledge and interference from potentially harmful historical knowledge, resulting in severe performance degradation. To this end, we propose a class-aware domain Knowledge Fusion and Fission method for continual test-time adaptation, called KFF, which adaptively expands and merges class-aware domain knowledge in old and new domains according to the test-time data from different domains, where discriminative historical knowledge can be dynamically accumulated. Specifically, considering the huge domain gap within streaming data, a domain Knowledge FIssion (KFI) module is designed to adaptively separate new domain knowledge from a paired class-aware domain prompt pool, alleviating the impact of negative knowledge brought by old domains that are distinct from the current domain. Besides, to avoid the cumulative computation and storage overheads from continuously fissioning new knowledge, a domain Knowledge FUsion (KFU) module is further designed to merge the fissioned new knowledge into the existing knowledge pool with minimal cost, where a greedy knowledge dynamic merging strategy is designed to improve the compatibility of new and old knowledge while keeping the computational efficiency. Jiahuan Zhou, Zhenyu Cui, Xu Zou 0002, Gang Hua 0001 |
NeurIPS | 5 |
| 2025 | Long Short-Term Knowledge Decomposition and Consolidation for Lifelong Person Re-IdentificationabstractLifelong person re-identification (LReID) aims to learn from streaming data sources step by step, which suffers from the catastrophic forgetting problem. In this paper, we investigate the exemplar-free LReID setting where no previous exemplar is available during the new step training. Existing exemplar-free LReID methods primarily adopt knowledge distillation to transfer knowledge from an old model to a new one without selection, inevitably introducing erroneous and detrimental information that hinders new knowledge learning. Furthermore, not all critical knowledge can be transferred due to the absence of old data, leading to the permanent loss of undistilled knowledge. To address these limitations, we propose a novel exemplar-free LReID method named Long Short-Term Knowledge Decomposition and Consolidation (LSTKC++). Specifically, an old knowledge rectification mechanism is developed to rectify the old model predictions based on new data annotations, ensuring correct knowledge transfer. Besides, a long-term knowledge consolidation strategy is designed, which first estimates the degree of old knowledge forgetting by leveraging the output difference between the old and new models. Then, a knowledge-guided parameter fusion strategy is developed to balance new and old knowledge, improving long-term knowledge retention. Upon these designs, considering LReID models tend to be biased on the latest seen domains, the fusion weights generated by this process often lead to sub-optimal knowledge balancing. To settle this, we further propose to decompose a single old model into two parts: a long-term old model containing multi-domain knowledge and a short-term model focusing on the latest short-term old knowledge. Then, the incoming new data are explored as an unbiased reference to adjust the old models' fusion weight to achieve backward optimization. Furthermore, an extended complementary knowledge rectification mechanism is developed to mine and retain the correct knowledge in the decomposed models. Extensive experimental results demonstrate that LSTKC++ significantly outperforms state-of-the-art methods by large margins. Kunlun Xu, Xu Zou 0002, Yuxin Peng 0001, Jiahuan Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Distribution-Aware Knowledge Aligning and Prototyping for Non-Exemplar Lifelong Person Re-IdentificationabstractLifelong person re-identification (LReID) suffers from the catastrophic forgetting problem when learning from non-stationary data streams. Existing exemplar-based and knowledge distillation-based LReID methods encounter data privacy and limited acquisition capacity, respectively. In this paper, we introduce the prototype, which is under-investigated in LReID, to better balance knowledge retention and acquisition. Previous prototype-based works primarily focused on the classification task, where prototypes were modeled as discrete points or statistical distributions. However, they either discarded the distribution information or omitted instance-level diversity, which are crucial fine-grained clues for LReID. Furthermore, the domain shifts between data sources result in a feature gap between the new and old data, which restricts the utilization of the fine-grained information in prototypes. To address these challenges, we propose Distribution-aware Knowledge Aligning and Prototyping (DKP++), a novel framework for modeling and leveraging prototypes in LReID. First, an Instance-level Distribution Modeling network is introduced to capture the local diversity of each instance. Next, a Distribution-oriented Prototype Generation algorithm transforms the instance-level diversity into identity-level distributions which are stored as prototypes. Then, a Prototype-based Knowledge Transfer module distills the knowledge within the prototypes to the new model. To mitigate the impact of domain shifts during knowledge transfer, we introduce a privacy-friendly Distribution Aligning module that transforms new input data to fit the historical distribution, which is incorporated with feature-level alignment constraints to enhance the coherence between new and old knowledge, effectively improving historical prototype utilization. Extensive experiments demonstrate that our method achieves a superior balance between plasticity and stability, outperforming state-of-the-art LReID methods by a large margin. Jiahuan Zhou, Kunlun Xu, Fan Zhuo, Xu Zou 0002, Yuxin Peng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Consistent Learning of Sparse Background Features for Infrared Small-Target LabelingabstractRecent years have witnessed many remarkable achievements in infrared small target detection based on deep learning. To achieve high performance in real-world applications, deep learning methods require a substantial number of accurate labels. However, infrared small target annotating is labor intensive as they are very small. Determining and annotating edge pixels demands considerable time and effort, which slows down data expansion and further research in infrared small target detection. To mitigate the issue, we propose a pseudo-label generation method named Consistent Learning of Sparse Background Feature (CLSBF). This approach models the generation of infrared small target pseudo-labels as a domain transformation from local target maps to background ones. It can relax the supervision requirements of deep learning methods, transitioning from absolute pixel-level supervision to point supervision. This method employs an unsupervised approach primarily, supplemented by semi-simulated supervision, to achieve mutual conversion of local images from different domains and obtain the final target pseudo-labels through the differences between target images and transformed background ones. Experiments show that equipped with our method, models trained with the input of a coarsely accurate center label can achieve performance up to 99.94% compared to models trained with official accurate labels. Furthermore, when the official labels are not accurate enough, models trained with pseudo-labels generated by CLSBF consistently show a performance improvement of 1.11% to 4.09% when they are evaluated based on re-labeled bounding boxes. Extensive experiments demonstrate the reliability of CLSBF in generating pseudo-labels and its potential to alleviate the labor-intensive process of manual labeling significantly. We have released an infrared small target labeling tool with CLSBF as an assistant at https://github.com/SeaHifly/CLSBF_software.git. Sheng Zhong 0001, Luxin Yan, Xu Zou 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Robust Point Cloud Registration via Patch MatchingabstractWe study the problem of exacting accurate correspondence pairs for point cloud registration. The existing correspondence methods focus on constructing point descriptors and then extracting correspondence point pairs. However, this process encounters two main issues: 1) point features are unstable and susceptible to noise, leading to a low inlier ratio (IR) for correspondence pairs and 2) the positional deviation of correspondence point pair results in accuracy errors when computing rigid transformations. To address these issues, we propose a robust point cloud registration framework based on patch matching, achieving high positional accuracy and high inlier-rate prediction of correspondence pairs. Specifically, we design a dual-branch point cloud registration network, with one branch dedicated to patch matching and the other branch to predicting the patch anchor, i.e., the coordinates used for patch matching. For patch matching, we integrate the topology of patches into the attention mechanism and adopt a multilevel patch-matching strategy to enhance the matching success rate. For coordinate prediction, we introduce graph convolutional network (GCN) and cross-attention mechanisms to explore local similar points through information interaction and feature correlation of patch pairs. Thanks to the stability of patch descriptors, our method demonstrates higher robustness compared to existing correspondence methods. Extensive experiments conducted on indoor, outdoor, synthetic, and deformable benchmarks validate the superiority of our method. Additionally, our method achieves certain effectiveness in cross-source point clouds. Tianming Zhao 0003, Tian Tian 0006, Xu Zou 0002, Luxin Yan, Sheng Zhong 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Hunt Camouflaged Objects via Revealing Mutation RegionsabstractDue to the high similarity between hidden objects and the surrounding background, camouflaged object detection (COD) remains a challenge. While many recently proposed methods have shown remarkable performance, most of them begin object perception by indiscriminately considering every pixel of the image. However, these early-stage region-insensitive perception methods still struggle to resist background interference, potentially missing subtle pixel changes by not prioritizing potential camouflaged areas initially. Fortunately, we reveal that the availability of an accurate mutation map can significantly enhance camouflaged discrimination ability. To this end, we propose MRNet (Mutation Region Network). MRNet initially generates a mutation map that identifies potential mutation regions exhibiting subtle pixel changes. The generation method involves amplifying and differing pixel changes based on the position and corresponding values of pixels. Subsequently, the selective expansion search operation utilizes the mutation map to extract the mapped graph, effectively reducing interference from background pixels that are distant from the mutation regions. Finally, decoding the mapped graph generates precise masks. Furthermore, we have created the largest test dataset with known categories to advance community research. Extensive experiments conducted on three widely used datasets and our proposed dataset show that MRNet surpasses other methods with superior performance. Source code is publicly available athttps://github.com/XinyueZhangHust/MRNet Xinyue Zhang 0009, Jiahuan Zhou, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Make Lossy Compression Meaningful for Low-Light ImagesabstractLow-light images frequently occur due to unavoidable environmental influences or technical limitations, such as insufficient lighting or limited exposure time. To achieve better visibility for visual perception, low-light image enhancement is usually adopted. Besides, lossy image compression is vital for meeting the requirements of storage and transmission in computer vision applications. To touch the above two practical demands, current solutions can be categorized into two sequential manners: ``Compress before Enhance (CbE)'' or ``Enhance before Compress (EbC)''. However, both of them are not suitable since: (1) Error accumulation in the individual models plagues sequential solutions. Especially, once low-light images are compressed by existing general lossy image compression approaches, useful information (e.g., texture details) would be lost resulting in a dramatic performance decrease in low-light image enhancement. (2) Due to the intermediate process, the sequential solution introduces an additional burden resulting in low efficiency. We propose a novel joint solution to simultaneously achieve a high compression rate and good enhancement performance for low-light images with much lower computational cost and fewer model parameters. We design an end-to-end trainable architecture, which includes the main enhancement branch and the signal-to-noise ratio (SNR) aware branch. Experimental results show that our proposed joint solution achieves a significant improvement over different combinations of existing state-of-the-art sequential ``Compress before Enhance'' or ``Enhance before Compress'' solutions for low-light images, which would make lossy low-light image compression more meaningful. The project is publicly available at: https://github.com/CaiShilv/Joint-IC-LL. Shilv Cai, Sheng Zhong 0001, Luxin Yan, Jiahuan Zhou, Xu Zou 0002 |
AAAI | 6 |
| 2024 | LSTKC: Long Short-Term Knowledge Consolidation for Lifelong Person Re-identificationabstractLifelong person re-identification (LReID) aims to train a unified model from diverse data sources step by step. The severe domain gaps between different training steps result in catastrophic forgetting in LReID, and existing methods mainly rely on data replay and knowledge distillation techniques to handle this issue. However, the former solution needs to store historical exemplars which inevitably impedes data privacy. The existing knowledge distillation-based models usually retain all the knowledge of the learned old models without any selections, which will inevitably include erroneous and detrimental knowledge that severely impacts the learning performance of the new model. To address these issues, we propose an exemplar-free LReID method named LongShort Term Knowledge Consolidation (LSTKC) that contains a Rectification-based Short-Term Knowledge Transfer module (R-STKT) and an Estimation-based Long-Term Knowledge Consolidation module (E-LTKC). For each learning iteration within one training step, R-STKT aims to filter and rectify the erroneous knowledge contained in the old model and transfer the rectified knowledge to facilitate the short-term learning of the new model. Meanwhile, once one training step is finished, E-LTKC proposes to further consolidate the learned long-term knowledge via adaptively fusing the parameters of models from different steps. Consequently, experimental results show that our LSTKC exceeds the state-of-the-art methods by 6.3%/9.4% and 7.9%/4.5%, 6.4%/8.0% and 9.0%/5.5% average mAP/R@1 on seen and unseen domains under two different training orders of the challenging LReID benchmark respectively. Kunlun Xu, Xu Zou 0002, Jiahuan Zhou |
AAAI | 2 |
| 2024 | SNIDA: Unlocking Few-Shot Object Detection with Non-Linear Semantic Decoupling AugmentationabstractOnce only a few-shot annotated samples are available, the performance of learning-based object detection would be heavily dropped. Many few-shot object detection (FSOD) methods have been proposed to tackle this issue by adopting image-level augmentations in linear manners. Nevertheless, those handcrafted enhancements often suf-fer from limited diversity and lack of semantic awareness, resulting in unsatisfactory performance. To this end, we propose a Semantic-guided Nonlinear Instance-level Data Augmentation method (SNIDA) for FSOD by decoupling the foreground and background to increase their diversities respectively. We design a semantic awareness enhancement strategy to separate objects from backgrounds. Concretely, masks of instances are extracted by an unsupervised semantic segmentation module. Then the diversity of samples would be improved by fusing instances into different backgrounds. Considering the shortcomings of augmenting images in a limited transformation space of existing traditional data augmentation methods, we introduce an object reconstruction enhancement module. The aim of this module is to generate sufficient diversity and nonlinear training data at the instance level through a semantic-guided masked autoencoder. In this way, the potential of data can be fully exploited in various object detection scenarios. Extensive experiments on PASCAL VOC and MS-COCO demonstrate that the proposed method outperforms base-lines by a large margin and achieves new state-of-the-art results under different shot settings. Xu Zou 0002, Luxin Yan, Sheng Zhong 0001, Jiahuan Zhou |
CVPR | 2 |
| 2024 | Distribution-Aware Knowledge Prototyping for Non-Exemplar Lifelong Person Re-IdentificationabstractLifelong person re-identification (LReID) suffers from the catastrophic forgetting problem when learning from non-stationary data. Existing exemplar-based and knowl-edge distillation-based LReID methods encounter data pri-vacy and limited acquisition capacity respectively. In this paper, we instead introduce the prototype, which is under-investigated in LReID, to better balance knowledge for-getting and acquisition. Existing prototype-based works primarily focus on the classification task, where the pro-totypes are set as discrete points or statistical distributions. However, they either discard the distribution in-formation or omit instance-level diversity which are cru-cial fine-grained clues for LReID. To address the above problems, we propose Distribution-aware Knowledge Pro-totyping (DKP) where the instance-level diversity of each sample is modeled to transfer comprehensive fine-grained knowledge for prototyping and facilitating LReID learning. Specifically, an Instance-level Distribution Mod-eling network is proposed to capture the local diver-sity of each instance. Then, the Distribution-oriented Prototype Generation algorithm transforms the instance-level diversity into identity-level distributions as proto-types, which is further explored by the designed Prototype-based Knowledge Transfer module to enhance the knowl-edge anti-forgetting and acquisition capacity of the LReID model. Extensive experiments verify that our method achieves superior plasticity and stability balancing and outperforms existing LReID methods by 8.1%19.1% average mAPIR@1 improvement. The code is available at https://github.com/zhoujiahuan1991/CVPR2024-DKP Kunlun Xu, Xu Zou 0002, Yuxin Peng 0001, Jiahuan Zhou |
CVPR | 2 |
| 2024 | Powerful Lossy Compression for Noisy ImagesabstractImage compression and denoising represent fundamental challenges in image processing with many real-world applications. To address practical demands, current solutions can be categorized into two main strategies: 1) sequential method; and 2) joint method. However, sequential methods have the disadvantage of error accumulation as there is information loss between multiple individual models. Recently, the academic community began to make some attempts to tackle this problem through end-to-end joint methods. Most of them ignore that different regions of noisy images have different characteristics. To solve these problems, in this paper, our proposed signal-to-noise ratio (SNR) aware joint solution exploits local and non-local features for image compression and denoising simultaneously. We design an end-to-end trainable network, which includes the main encoder branch, the guidance branch, and the signal-to-noise ratio (SNR) aware branch. We conducted extensive experiments on both synthetic and real-world datasets, demonstrating that our joint solution outperforms existing state-of-the-art methods. Shilv Cai, Xiaoguo Liang, Shuning Cao, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002 |
ICME | 7 |
| 2024 | Perceptual-Distortion Balanced Image Super-Resolution is a Multi-Objective Optimization ProblemabstractTraining Single-Image Super-Resolution (SISR) models using pixel-based regression losses can achieve high distortion metrics scores (e.g., PSNR and SSIM), but often results in blurry images due to insufficient recovery of high-frequency details. Conversely, using GAN or perceptual losses can produce sharp images with high perceptual metric scores (e.g., LPIPS), but may introduce artifacts and incorrect textures. Balancing these two types of losses can help achieve a trade-off between distortion and perception, but the challenge lies in tuning the loss function weights. To address this issue, we propose a novel method that incorporates Multi-Objective Optimization (MOO) into the training process of SISR models to balance perceptual quality and distortion. We conceptualize the relationship between loss weights and image quality assessment (IQA) metrics as black-box objective functions to be optimized within our Multi-Objective Bayesian Optimization Super-Resolution (MOBOSR) framework. This approach automates the hyperparameter tuning process, reduces overall computational cost, and enables the use of numerous loss functions simultaneously. Extensive experiments demonstrate that MOBOSR outperforms state-of-the-art methods in terms of both perceptual quality and distortion, significantly advancing the perception-distortion Pareto frontier. Our work points towards a new direction for future research on balancing perceptual quality and fidelity in nearly all image restoration tasks. The source code and pretrained models are available at: https://github.com/ZhuKeven/MOBOSR. Qiwen Zhu, Shilv Cai, Jiahuan Zhou, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002 |
ACM Multimedia | 8 |
| 2024 | I2C: Invertible Continuous Codec for High-Fidelity Variable-Rate Image CompressionabstractLossy image compression is a fundamental technology in media transmission and storage. Variable-rate approaches have recently gained much attention to avoid the usage of a set of different models for compressing images at different rates. During the media sharing, multiple re-encodings with different rates would be inevitably executed. However, existing Variational Autoencoder (VAE)-based approaches would be readily corrupted in such circumstances, resulting in the occurrence of strong artifacts and the destruction of image fidelity. Based on the theoretical findings of preserving image fidelity via invertible transformation, we aim to tackle the issue of high-fidelity fine variable-rate image compression and thus propose the Invertible Continuous Codec (I2C). We implement the I2C in a mathematical invertible manner with the core Invertible Activation Transformation (IAT) module. I2C is constructed upon a single-rate Invertible Neural Network (INN) based model and the quality level (QLevel) would be fed into the IAT to generate scaling and bias tensors. Extensive experiments demonstrate that the proposed I2C method outperforms state-of-the-art variable-rate image compression methods by a large margin, especially after multiple continuous re-encodings with different rates, while having the ability to obtain a very fine variable-rate control without any performance compromise. Shilv Cai, Zhijun Zhang 0009, Xiangyun Zhao, Jiahuan Zhou, Yuxin Peng 0001, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2024 | Learning Oriented Object Detection via Naive Geometric ComputingabstractDetecting oriented objects along with estimating their rotation information is one crucial step for image analysis, especially for remote sensing images. Despite that many methods proposed recently have achieved remarkable performance, most of them directly learn to predict object directions under the supervision of only one (e.g., the rotation angle) or a few (e.g., several coordinates) groundtruth (GT) values individually. Oriented object detection would be more accurate and robust if extra constraints, with respect to proposal and rotation information regression, are adopted for joint supervision during training. To this end, we propose a mechanism that simultaneously learns the regression of horizontal proposals, oriented proposals, and rotation angles of objects in a consistent manner, via naive geometric computing, as one additional steady constraint. An oriented center prior guided label assignment strategy is proposed for further enhancing the quality of proposals, yielding better performance. Extensive experiments on six datasets demonstrate the model equipped with our idea significantly outperforms the baseline by a large margin and several new state-of-the-art results are achieved without any extra computational burden during inference. Our proposed idea is simple and intuitive that can be readily implemented. Source codes are publicly available at: https://github.com/wangWilson/CGCDet.git. Zhijun Zhang 0009, Guodong Wang 0001, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | Multiscale Multilevel Residual Feature Fusion for Real-Time Infrared Small Target DetectionabstractDetecting infrared dim and small targets is one crucial step for many tasks such as early warning. It remains a continuing challenge since characteristics of infrared small targets, usually represented by only a few pixels, are generally not salient. Despite that many traditional methods have significantly advanced the community, their robustness or efficiency is still lacking. Most recently, CNN-based object detection has achieved remarkable performance and some researchers focus on it. However, these methods are not computationally efficient when implemented on some CPU-only machines and few datasets are available publicly. To promote the detection of infrared small targets in complex backgrounds, we propose a new lightweight CNN-based architecture. The network contains three modules: the feature extraction module is designed for representing multi-scale and multi-level features, the grid resample operation module is proposed to fuse features from all scales, and a decoupled head to distinguish infrared small targets from backgrounds. Moreover, we collect a brand-new infrared small target detection dedicated dataset which consists of 68311 practical captured images with complex backgrounds for alleviating the data dilemma. To validate the proposed model, 54758 images are used for training and 13553 images are used for testing respectively. Extensive experimental results demonstrate that the proposed method outperforms all traditional methods by a large margin and runs much faster than other CNN methods with high precision. The proposed model can be implemented on the Intel i7-10850H CPU (2.3GHz) platform and Jetson Nano for real-time infrared small target detection at 44 FPS and 27 FPS, respectively. It can be even deployed on an Atom x5-Z8500 (1.44GHz) machine at about 25 FPS with 128×128 local images. The source codes and the dataset have been made publicly available at https://github.com/SeaHifly/Infrared-Small-Target. Sheng Zhong 0001, Tianxu Zhang, Xu Zou 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Category-Aware Transformer Network for Better Human-Object Interaction DetectionabstractHuman-Object Interactions (HOI) detection, which aims to localize a human and a relevant object while recognizing their interaction, is crucial for understanding a still image. Recently, tranformer-based models have significantly advanced the progress of HOI detection. However, the capability of these models has not been fully explored since the Object Query of the model is always simply initialized as just zeros, which would affect the performance. In this paper, we try to study the issue of promoting transformer-based HOI detectors by initializing the Object Query with category-aware semantic information. To this end, we innovatively propose the Category-Aware Transformer Network (CATN). Specifically, the Object Query would be initialized via category priors represented by an external object detection model to yield a better performance. Moreover, such category priors can be further used for enhancing the representation ability of features via the attention mechanism. We have firstly verified our idea via the Oracle experiment by initializing the Object Query with the groundtruth category information. And then extensive experiments have been conducted to show that a HOI detection model equipped with our idea outperforms the baseline by a large margin to achieve a new state-of-the-art result. Leizhen Dong, Kunlun Xu, Zhijun Zhang 0009, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002 |
CVPR | 7 |
| 2022 | High-Fidelity Variable-Rate Image Compression via Invertible Activation TransformationabstractLearning-based methods have effectively promoted the community of image compression. Meanwhile, variational autoencoder(VAE) based variable-rate approaches have recently gained much attention to avoid the usage of a set of different networks for various compression rates. Despite the remarkable performance that has been achieved, these approaches would be readily corrupted once multiple compression/decompression operations are executed, resulting in the fact that image quality would be tremendously dropped and strong artifacts would appear. Thus, we try to tackle the issue of high-fidelity fine variable-rate image compression and propose the Invertible Activation Transformation(IAT) module. We implement the IAT in a mathematical invertible manner on a single rate Invertible Neural Network(INN) based model and the quality level(QLevel) would be fed into the IAT to generate scaling and bias tensors. IAT and QLevel together give the image compression model the ability of fine variable-rate control while better maintaining the image fidelity. Extensive experiments demonstrate that the single rate image compression model equipped with our IAT module has the ability to achieve variable-rate control without any compromise. And our IAT-embedded model obtains comparable rate-distortion performance with recent learning-based image compression methods. Furthermore, our method outperforms the state-of-the-art variable-rate image compression method by a large margin, especially after multiple re-encodings. Shilv Cai, Zhijun Zhang 0009, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002 |
ACM Multimedia | 6 |
| 2022 | Effective actor-centric human-object interaction detection
Kunlun Xu, Zhijun Zhang 0009, Leizhen Dong, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002 |
Image Vis. Comput. | 8 |
| 2022 | Learning dynamic background for weakly supervised moving object detection
Zhijun Zhang 0009, Yi Chang 0002, Sheng Zhong 0001, Luxin Yan, Xu Zou 0002 |
Image Vis. Comput. | 5 |
| 2021 | Morphable Detector for Object Detection on DemandabstractMany emerging applications of intelligent robots need to explore and understand new environments, where it is desirable to detect objects of novel classes on the fly with minimum online efforts. This is an object detection on demand (ODOD) task. It is challenging, because it is impossible to annotate a large number of data on the fly, and the embedded systems are usually unable to perform back-propagation which is essential for training. Most existing few-shot detection methods are confronted here as they need extra training. We propose a novel morphable detector (MD), that simply "morphs" some of its changeable parameters online estimated from the few samples, so as to detect novel classes without any extra training. The MD has two sets of parameters, one for the feature embedding and the other for class representation (called "prototypes"). Each class is associated with a hidden prototype to be learned by integrating the visual and semantic embeddings. The learning of the MD is based on the alternate learning of the feature embedding and the prototypes in an EM-like approach which allows the recovery of an unknown prototype from a few samples of a novel class. Once an MD is learned, it is able to use a few samples of a novel class to directly compute its prototype to fulfill the online morphing process. We have shown the superiority of the MD in Pascal [12], COCO [27] and FSOD [13] datasets. Xiangyun Zhao, Xu Zou 0002, Ying Wu 0001 |
ICCV | 2 |
| 2021 | Category-Aware Aircraft Landmark DetectionabstractAircraft landmark detection (ALD) aims at detecting the keypoints of aircraft, which can serve as an important role for subsequent applications such as fine-grained aircraft recognition. In ALD, the physical size discrepancy between different kinds of aircraft may lead to inconsistent landmark structure, which significantly harms landmark detection results. In this letter, we take advantage of the category prior to alleviate the size discrepancy in ALD. The proposed category-aware landmark detection network (CALDN) possesses two streams: a classification stream for size categorization and a localization stream for landmark detection. Instance-level size category information captured by classification stream serves as the guidance in the localization stream for robust landmark detection. Moreover, a category attention module (CAM) is proposed for better-utilizing category information to guide ALD. Benefitting from the adaptive attention mechanism, CAM can automatically highlight category-specific features for ulteriorly reducing the influence of size discrepancy. Furthermore, to advance ALD research, we contribute the first perspective-variant aircraft landmark dataset. Solid experiments demonstrate the superiority of our method. Yi Li 0033, Yi Chang 0002, Yuntong Ye, Xu Zou 0002, Sheng Zhong 0001, Luxin Yan |
IEEE Signal Process. Lett. | 4 |
| 2021 | Towards Unconstrained Facial Landmark Detection Robust to Diverse Cropping MannersabstractFacial landmark detection is one crucial step for face-based image/video analysis. Despite the fact that recently many facial landmark detection models have achieved remarkable performance, most state-of-the-art heatmap regression-based methods heavily rely on initialization of the face detector. However, there inevitably exists semantic gaps among different annotators or face detectors. An improper facial bounding box will tremendously drop off the performance of the facial landmark detection model. Facial landmark detection would be more practical if robust to face images cropped by diverse manners (see Figure 1, the col.1 shows face images cropped by a proper bounding box, col.2 and col.3 show face images cropped by an oversize and a small bounding boxes respectively). To this end, we present a “Unconstrained Facial Landmark Detection(UFLD)” mechanism, that aims at enhancing the robustness of facial landmark detection, to deal with the inconsistent cropping manner issue. UFLD consists of two aspects: a Transformation-Invariant Landmark Detector(TILD) and an Availability-Guided Solver(AGS). TILD gives the ability to detect consistent landmarks for face images cropped by diverse manners. And AGS can alleviate the by-effect of “landmarks outside the image” caused by improper cropping results or TILD, and further promote the performance. The proposed mechanism achieved above 6.5% improvement in standard normalized landmarks mean error reduction on face images cropped by diverse manners compared to baselines. Xu Zou 0002, Luxin Yan, Sheng Zhong 0001, Ying Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Learning Robust Facial Landmark Detection via Hierarchical Structured EnsembleabstractHeatmap regression-based models have significantly advanced the progress of facial landmark detection. However, the lack of structural constraints always generates inaccurate heatmaps resulting in poor landmark detection performance. While hierarchical structure modeling methods have been proposed to tackle this issue, they all heavily rely on manually designed tree structures. The designed hierarchical structure is likely to be completely corrupted due to the missing or inaccurate prediction of landmarks. To the best of our knowledge, in the context of deep learning, no work before has investigated how to automatically model proper structures for facial landmarks, by discovering their inherent relations. In this paper, we propose a novel Hierarchical Structured Landmark Ensemble (HSLE) model for learning robust facial landmark detection, by using it as the structural constraints. Different from existing approaches of manually designing structures, our proposed HSLE model is constructed automatically via discovering the most robust patterns so HSLE has the ability to robustly depict both local and holistic landmark structures simultaneously. Our proposed HSLE can be readily plugged into any existing facial landmark detection baselines for further performance improvement. Extensive experimental results demonstrate our approach significantly outperforms the baseline by a large margin to achieve a state-of-the-art performance. Xu Zou 0002, Sheng Zhong 0001, Luxin Yan, Xiangyun Zhao, Jiahuan Zhou, Ying Wu 0001 |
ICCV | 1 |