VLDB 2026 Research / reviewers in the wild / expert
Zhizhong Zhang 0001
dblp:20/1541-1
· DBLP profile ↗
88ranked-venue papers
7as first author
83since 2021 · last 2026
0000-0001-6905-4478ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 64 · 4 first-author · 60 since 2021Artificial intelligence and machine learning · 51 · 2 first-author · 50 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Human Motion Synthesis in 3D Scenes via Unified Scene Semantic OccupancyabstractHuman motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this paper, we propose a human motion synthesis framework that take an unified Scene Semantic Occupancy (SSO) for scene representation, termed SSOMotion. We design a bi-directional tri-plane decomposition to derive a compact version of the SSO, and scene semantics are mapped to an unified feature space via CLIP encoding and shared linear dimensionality reduction. Such strategy can derive the fine-grained scene semantic structures while significantly reduce redundant computations. We further take these scene hints and movement direction derived from instructions for motion control via frame-wise scene query. Extensive experiments and ablation studies conducted on cluttered scenes using ShapeNet furniture, as well as scanned scenes from PROX and Replica datasets, demonstrate its cutting-edge performance while validating its effectiveness and generalization ability. Jingyu Gong, Kunkun Tong, Zhuoran Chen, Chuanhan Yuan, Mingang Chen, Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006 |
AAAI | 6 |
| 2026 | Diffusion Implicit Policy for Unpaired Scene-aware Motion SynthesisabstractScene-aware motion synthesis has been widely researched recently due to its numerous applications. Prevailing methods rely heavily on paired motion-scene data, while it is difficult to generalize to diverse scenes when trained only on a few specific ones. Thus, we propose a unified framework, termed Diffusion Implicit Policy (DIP), for scene-aware motion synthesis, where paired motion-scene data are no longer necessary. In this paper, we disentangle human-scene interaction from motion synthesis during training, and then introduce an interaction-based implicit policy into motion diffusion during inference. Synthesized motion can be derived through iterative diffusion denoising and implicit policy optimization, thus motion naturalness and interaction plausibility can be maintained simultaneously. For long-term motion synthesis, we introduce motion blending in joint rotation power space. The proposed method is evaluated on synthesized scenes with ShapeNet furniture, and real scenes from PROX and Replica. Results show that our framework presents better motion naturalness and interaction plausibility than cutting-edge methods. This also indicates the feasibility of utilizing the DIP for motion synthesis in more general tasks and versatile scenes. Jingyu Gong, Fengqi Liu, Qianyu Zhou 0001, Xin Tan 0002, Zhizhong Zhang 0001, Yuan Xie 0006 |
AAAI | 7 |
| 2026 | Multi-Step Deformable Gaussian Splatting for Dynamic Scene RenderingabstractReconstructing dynamic scenes has long been a challenging task in 3D vision. Previous mainstream methods based on 3D Gaussian Splatting typically employ a single deformation field to directly model spatiotemporal changes. However, such one-step deformation struggles to capture diverse and complex motion patterns. To address this limitation, we propose decomposing the one-step deformation into a multi-step process, where each step is represented by a deformation layer. Additionally, we introduce a weight prediction mechanism for each layer to control the extent of deformation at every step. We provide two types of deformation layers based on implicit and explicit approaches. Moreover, while the deformation layer is time-conditioned, the Gaussians' behavior may still be influenced by their time-invariant properties. Therefore, we propose a fully time-agnostic scale modulation block to modulate the scaling changes of Gaussians. Extensive experiments on D-NeRF, Neu3D, and HyperNeRF demonstrate that our method achieves state-of-the-art performance. Jiaheng Hu, Zhizhong Zhang 0001, Jingyu Gong, Lizhuang Ma, Xin Tan 0002, Yuan Xie 0006 |
AAAI | 2 |
| 2026 | Zero-Shot Robotic Manipulation via 3D Gaussian Splatting-Enhanced Multimodal Retrieval-Augmented GenerationabstractExisting end-to-end approaches of robotic manipulation often lack generalization to unseen objects or tasks due to limited data and poor interpretability. While recent Multimodal Large Language Models (MLLMs) demonstrate strong commonsense reasoning, they struggle with geometric and spatial understanding required for pose prediction. In this paper, we propose RobMRAG, a 3D Gaussian Splatting-Enhanced Multimodal Retrieval-Augmented Generation (MRAG) framework for zero-shot robotic manipulation. Specifically, We construct a multi-source manipulation knowledge base containing object contact frames, task completion frames, and pose parameters. During inference, a Hierarchical Multimodal Retrieval module first employs hybrid semantic search to find task-relevant object prototypes, then selects the geometrically closest reference example based on pixel-level similarity and Instance Matching Distance (IMD). We further introduce a 3D-Aware Pose Refinement module based on 3D Gaussian Splatting into the MRAG framework, which aligns the pose of the reference object to the target object in 3D space. The aligned results are reprojected onto the image plane and used as input to the MLLM to enhance the generation of the final pose parameters. Extensive experiments show that on a test set containing 30 categories of household objects, our method improves the success rate by 7.76% compared to the best-performing zero-shot baseline under the same setting, and by 6.54% compared to the state-of-the-art supervised baseline. Our results validate that RobMRAG effectively bridges the gap between high-level semantic reasoning and low-level geometric execution, enabling robotic systems that generalize to unseen objects while remaining inherently interpretable. Zilong Xie, Jingyu Gong, Xin Tan 0002, Zhizhong Zhang 0001, Yuan Xie 0006 |
AAAI | 4 |
| 2026 | NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation TasksabstractZhihao Luo, Wentao Yan, Jingyu Gong, Min Wang, Zhizhong Zhang, Xuhong Wang, Yuan Xie, Xin Tan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wentao Yan, Jingyu Gong, Min Wang 0024, Zhizhong Zhang 0001, Xuhong Wang, Yuan Xie 0006, Xin Tan 0002 |
ACL (1) | 5 |
| 2026 | Dynamic expansion orthogonal network for class-incremental learning
Mingda Dong, Zhizhong Zhang 0001, Xin Tan 0002, Jiling Qiu, Yuan Xie 0006 |
Knowl. Based Syst. | 2 |
| 2026 | From sparse semantics to rich instances: Empowering label-efficient LiDAR panoptic segmentation via geometric priors
Wei Zhang 0217, Zhizhong Zhang 0001, Xin Tan 0002, Lizhuang Ma, Yuan Xie 0006 |
Neural Networks | 3 |
| 2026 | Two-stage knowledge distillation for visible-infrared person re-identification
Jiangming Shi, Xiangbo Yin, Demao Zhang, Zhizhong Zhang 0001, Yuan Xie 0001, Yanyun Qu |
Pattern Recognit. | 4 |
| 2026 | A Task-Aware Parameter Decoupling Framework for Continual Anomaly DetectionabstractReal-world industrial scenarios have become increasingly dynamic, with new product types, defect patterns, and operational modes emerging rapidly. In such a context, the one-for-more paradigm enables the use of a single model to economically and continually adapt to evolving distributions or patterns, positioning it as a key component in modern Industrial AI systems. This article proposes a novel one-for-more anomaly detection framework designed to identify anomalies across expanding product lines. The framework incorporates two model-agnostic techniques: instance-aware prompt tuning (IPT) and gradient-aware parameter decoupling (GPD). Our approach is built upon a reconstruction-based vision transformer (ViT) encoder–decoder architecture. IPT addresses the domain gap between pretrained models and industrial data by leveraging an instance-level prompt and a shared memory mechanism, which helps the pretrained model retain previously learned patterns. GPD selectively updates network parameters based on the gradient’s impact on prior tasks, employing orthogonal gradient projection to further minimize interference. In addition, we introduce a new dataset to simulate the one-for-more industrial scenario. Extensive experiments on MVTec and our proposed dataset demonstrate that our framework achieves the state-of-the-art performance across various continual learning settings, significantly outperforming existing methods, particularly in multistep incremental scenarios. Zhizhong Zhang 0001, Guchu Zou, Chengwei Chen, Zhenyi Qi, Jingwen Qi, Yongke Yao, Xiaofan Li 0008, Yuan Xie 0006, Xin Tan 0002 |
IEEE Trans. Ind. Informatics | 1 |
| 2026 | Transporting the Cross-Modal Prototypes for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible infrared person re-identification (USVI-ReID) is a challenging retrieval task that retrieves cross-modality pedestrian images without using any label information. In this task, the large cross-modality variance makes it difficult to generate reliable cross-modality labels, and the lack of annotations also provides additional difficulties for learning modality-invariant features. To facilitate this unsupervised cross-modal learning, we begin by leveraging the information contained in the cross-modality input and its predicted label. Aiming to minimize information loss, we optimize the model by incorporating entropy minimization, uniform label distribution, and cross-modality matching. In our approach, we design a loop iterative training strategy alternating between model training and cross-modality matching, where a uniform prior guided optimal transport assignment is proposed to select matched visible and infrared prototypes. This matching information is then utilized to minimize the intra- and cross-modality entropy. As a result, our model can gradually self-learn useful information, enabling it to generate discriminative representations for unlabeled cross-modal data. Extensive experimental results on benchmarks demonstrate the effectiveness of our method, e.g., 69.4% and 89.4% of Rank-1 accuracy on SYSU-MM01 and RegDB without any annotations. The code will be released soon. Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006 |
IEEE Trans. Image Process. | 1 |
| 2026 | Decoupling 3-D Point Cloud Attributes for Semantic Segmentation via Real-World Prior ExploitationabstractPoint cloud semantic segmentation, which involves assigning a category for each point, is a crucial task in autonomous driving and intelligent transportation systems. Due to the inherently unordered and irregular nature of point clouds, learning robust features that accurately capture real-world distributions from point coordinates and other attributes remains challenging. Following the pioneering work of PointNet, current 3D deep neural networks process point coordinates alongside other attributes without fully exploiting the implicit class prior information embedded in spatial information. In this work, we first conduct a pilot study to evaluate how current 3D networks utilize point coordinates and validate the presence of implicit class priors within them. Subsequently, we design a robust Position-to-Physics (P2P) fusion strategy that learns adaptive weights to dynamically incorporate implicit class priors present in point coordinates into point features. Moreover, we design a dual-branch network architecture and propose a triplet loss to further enhance the adaptive fusion process. Extensive experiments demonstrate that decoupling position attributes from physics attributes facilitates the extraction and utilization of implicit class priors. Our proposed modules consistently improve segmentation performance across various networks and datasets, demonstrating their generalizability and effectiveness. Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Lizhuang Ma, Yuan Xie 0006 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | FastLGS: Speeding Up Language Embedded Gaussians with Feature Grid MappingabstractThe semantically interactive radiance field has always been an appealing task for its potential to facilitate user-friendly and automated real-world 3D scene understanding applications. However, it is a challenging task to achieve high quality, efficiency and zero-shot ability at the same time with semantics in radiance fields. In this work, we present FastLGS, an approach that supports real-time open-vocabulary query within 3D Gaussian Splatting (3DGS) under high resolution. We propose the semantic feature grid to save multi-view CLIP features which are extracted based on Segment Anything Model (SAM) masks, and map the grids to low dimensional features for semantic field training through 3DGS. Once trained, we can restore pixel-aligned CLIP embeddings through feature grids from rendered features for open-vocabulary queries. Comparisons with other state-of-the-art methods prove that FastLGS can achieve the first place performance concerning both speed and accuracy, where FastLGS is 98 times faster than LERF, 4 times faster than LangSplat and 2.5 times faster than LEGaussians. Meanwhile, experiments show that FastLGS is adaptive and compatible with many downstream tasks, such as 3D segmentation and 3D object inpainting, which can be easily applied to other 3D manipulation systems. Yuzhou Ji, Junshu Tang, Wuyi Liu, Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006 |
AAAI | 5 |
| 2025 | DepthFisheye: Efficient Fine-Tuning of Depth Estimation Models for Fisheye Cameras
Zhiwei Zhang 0005, Xin Tan 0002, Zhizhong Zhang 0001, Lizhuang Ma |
CVM (3) | 4 |
| 2025 | One-for-More: Continual Diffusion Model for Anomaly DetectionabstractWith the rise of generative models, there is a growing interest in unifying all tasks within a generative framework. Anomaly detection methods also fall into this scope and utilize diffusion models to generate or reconstruct normal samples when given arbitrary anomaly images. However, our study found that the diffusion model suffers from severe "faithfulness hallucination" and "catastrophic forgetting", which can’t meet the unpredictable pattern increments. To mitigate the above problems, we propose a continual diffusion model that uses gradient projection to achieve stable continual learning. Gradient projection deploys a regularization on the model updating by modifying the gradient towards the direction protecting the learned knowledge. But as a double-edged sword, it also requires huge memory costs brought by the Markov process. Hence, we propose an iterative singular value decomposition method based on the transitive property of linear representation, which consumes tiny memory and incurs almost no performance loss. Finally, considering the risk of "over-fitting" to normal images of the diffusion model, we propose an anomaly-masked network to enhance the condition mechanism of the diffusion model. For continual anomaly detection, ours achieves first place in 17/18 settings on MVTec and VisA. Code is available at https://github.com/FuNz-0/One-for-More Xiaofan Li 0008, Xin Tan 0002, Zhizhong Zhang 0001, Rizen Guo, Guannan Jiang, Yanyun Qu, Lizhuang Ma, Yuan Xie 0006 |
CVPR | 4 |
| 2025 | Efficient Prototypical Classifier for Class-Incremental LearningabstractThe nearest prototypical classifier faces challenges of semantic drift and prototype interference. Previous methods address these issues using data rehearsal and contrastive learning, but these approaches incur high memory costs and slow convergence. In this paper, we propose a novel prototypical minimum distance loss, along with a two-stage training pipeline, to mitigate prototype interference with low memory overhead and fast convergence. Leveraging task-specific prompts and a key-query mechanism, we significantly reduce semantic drift. Additionally, we introduce a continual exponential moving average to enhance model stability and minimize forgetting. Notably, our method is rehearsal-free and avoids generation processes, simplifying training and further reducing memory usage. We validate our approach on four challenging class-incremental learning datasets, achieving significant improvements over state-of-the-art methods. Wei Zhang 0217, Jingyang Qiao, Yuan Xie 0006, Zhizhong Zhang 0001, Xin Tan 0002 |
ICASSP | 4 |
| 2025 | Prototype Alignment with LoRA Fusion for Class-Incremental LearningabstractRecent advancements in pre-trained models have enhanced performance on downstream tasks due to their strong generalizability. Despite this, models fine-tuned continually often face challenges such as catastrophic forgetting and loss of generalization. To address these issues, we propose a novel approach that utilizes distinct Low-Rank Adaptation (LoRA) modules for each task. These modules parameter-efficient, and integrated across tasks to ensure the model maintains strong performance on both old and new classes. Additionally, we investigate semantic relationships between class prototypes to effectively reconstruct old prototypes in the context of new tasks. Our experiments demonstrate that this method significantly outperforms baseline approaches across various class-incremental learning benchmarks, offering an efficient and effective solution for mitigating forgetting and preserving model performance. Wei Zhang 0217, Yuan Xie 0006, Zhizhong Zhang 0001, Xin Tan 0002 |
ICASSP | 3 |
| 2025 | Stylized-Face: A Million-Level Stylized Face Dataset for Face Recognition
Zhengyuan Peng, Jianqing Xu, Yuge Huang, Jinkun Hao, Shouhong Ding, Zhizhong Zhang 0001, Xin Tan 0002, Lizhuang Ma |
ICCV | 6 |
| 2025 | Multi-Schema Proximity Network for Composed Image Retrieval
Jiangming Shi, Xiangbo Yin, Yeyun Chen, Yachao Zhang 0001, Zhizhong Zhang 0001, Yanyun Qu |
ICCV | 5 |
| 2025 | From Enhancement to Understanding: Build a Generalized Bridge for Low-Light Vision via Semantically Consistent Unsupervised Fine-Tuning
Shao Zeng, Tianjun Gu, Zhizhong Zhang 0001, Shouhong Ding, Jun Wang 0006, Xin Tan 0002, Yuan Xie 0006, Lizhuang Ma |
ICCV | 4 |
| 2025 | LFNet: Cross-Modal LiDAR-Fisheye Fusion Network for 3D Semantic SegmentationabstractCross-modal fusion, which leverages images to enhance 3D semantic segmentation, has demonstrated significant effectiveness due to the complementary nature of heterogeneous data. However, existing approaches are limited to pinhole images, leaving fisheye images largely unexplored. In this paper, we introduce the LiDAR-Fisheye Fusion Network (LFNet), a dual-transformer architecture designed for cross-modal fusion (CMF) across hierarchical multi-scale layers. The 3D Transformer extracts point-level features from LiDAR data, while the pre-trained 2D Transformer extracts patch-level features from fisheye images.The CMF module comprises two key components: Local Fusion (LoF) and Global Fusion (GoF). The LoF module interpolates patch-level features to pixel-level for accurate feature alignment and computes precise point-to-pixel mappings for gated fusion. Meanwhile, the GoF module enables points to capture a holistic understanding of the scene via a cross-modal attention mechanism. Experimental results highlight the potential of fisheye images as a promising modality to complement LiDAR data in 3D semantic segmentation. The code will be available at https://github.com/wjzhang642/LFNet. Zhiwei Zhang 0005, Tianfang Sun, Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006 |
ICME | 4 |
| 2025 | Large Continual Instruction AssistantabstractContinual Instruction Tuning (CIT) is adopted to continually instruct Large Models to follow human intent data by data. It is observed that existing gradient update would heavily destroy the performance on previous datasets during CIT process. Instead, Exponential Moving Average (EMA), owns the ability to trace previous parameters, which can aid in decreasing forgetting. Nonetheless, its stable balance weight fails to deal with the ever-changing datasets, leading to the out-of-balance between plasticity and stability. In this paper, we propose a general continual instruction tuning framework to address the challenge. Starting from the trade-off prerequisite and EMA update, we propose the plasticity and stability ideal condition. Based on Taylor expansion in the loss function, we find the optimal balance weight can be automatically determined by the gradients and learned parameters. Therefore, we propose a stable-plasticity balanced coefficient to avoid knowledge interference. Based on the semantic similarity of the instructions, we can determine whether to retrain or expand the training parameters and allocate the most suitable parameters for the testing instances. Extensive experiments across multiple continual instruction tuning benchmarks demonstrate that our approach not only enhances anti-forgetting capabilities but also significantly improves overall continual tuning performance. Our code is available at https://github.com/JingyangQiao/CoIN. Jingyang Qiao, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Shouhong Ding, Yuan Xie 0006 |
ICML | 2 |
| 2025 | PFDepth: Heterogeneous Pinhole-Fisheye Joint Depth Estimation via Distortion-aware Gaussian-Splatted Volumetric FusionabstractIn this paper, we present the first pinhole-fisheye framework for heterogeneous multi-view depth estimation, PFDepth. Our key insight is to exploit the complementary characteristics of pinhole and fisheye imagery (undistorted vs. distorted, small vs. large FOV, far vs. near field) for joint optimization. PFDepth employs a unified architecture capable of processing arbitrary combinations of pinhole and fisheye cameras with varied intrinsics and extrinsics. Within PFDepth, we first explicitly lift 2D features from each heterogeneous view into a canonical 3D volumetric space. Then, a core module termed Heterogeneous Spatial Fusion is designed to process and fuse distortion-aware volumetric features across overlapping and non-overlapping regions. Additionally, we subtly reformulate the conventional voxel fusion into a novel 3D Gaussian representation, in which learnable latent Gaussian spheres dynamically adapt to local image textures for finer 3D aggregation. Finally, fused volume features are rendered into multi-view depth maps. Through extensive experiments, we demonstrate that PFDepth sets a state-of-the-art performance on KITTI-360 and RealHet datasets over current mainstream depth networks. To the best of our knowledge, this is the first systematic study of heterogeneous pinhole-fisheye depth estimation, offering both technical novelty and valuable empirical insights. Zhiwei Zhang 0005, Ruikai Xu, Zhizhong Zhang 0001, Xin Tan 0002, Jingyu Gong, Yuan Xie 0006, Lizhuang Ma |
ACM Multimedia | 4 |
| 2025 | Wandering and feeling the Scenes: Body-Aware Diffusion for 3D Human Motion GenerationabstractAs demand for virtual digital characters grows in fields such as virtual reality, gaming, and animation, generating highly controllable human motion within scenes has become a key research focus. Existing methods for scene-aware motion generation typically rely on global alignment or latent space matching, which provides limited control over the fine-grained movements of individual body parts. This limitation often leads to rigid and unrealistic motions when interacting with complex environments. Therefore, we propose the Body-Aware Interaction Diffusion Model (BA-IDM), which enables fine-grained control of human motion within a scene by leveraging multimodal information. Text descriptions, motion scenes, and movement trajectories can all serve as inputs, allowing for precise control of each body part and facilitating the generation of a wide range of complex actions. Moreover, our approach is designed to operate on de-identified motion data, effectively protecting user privacy throughout the process, which is essential for practical and user-centric applications. Jingyu Gong, Shaohui Lin, Yang Li 0041, Zhizhong Zhang 0001 |
MMAsia | 5 |
| 2025 | Switchable Token-Specific Codebook Quantization For Face Image CompressionabstractWith the ever-increasing volume of visual data, the efficient and lossless transmission, along with its subsequent interpretation and understanding, has become a critical bottleneck in modern information systems. The emerged codebook-based solution
utilize a globally shared codebook to quantize and dequantize each token, controlling the bpp by adjusting the number of tokens or the codebook size.
However, for facial images—which are rich in attributes—such global codebook strategies overlook both the category-specific correlations within images and the semantic differences among tokens, resulting in suboptimal performance, especially at low bpp. Motivated by these observations, we propose a Switchable Token-Specific Codebook Quantization for face image compression, which learns distinct codebook groups for different image categories and assigns an independent codebook to each token.
By recording the codebook group to which each token belongs with a small number of bits, our method can reduce the loss incurred when decreasing the size of each codebook group. This enables a larger total number of codebooks under a lower overall bpp, thereby enhancing the expressive capability and improving reconstruction performance. Owing to its generalizable design, our method can be integrated into any existing codebook-based representation learning approach and has demonstrated its effectiveness on face recognition datasets, achieving an average accuracy of 93.51\% for reconstructed images at 0.05 bpp. Guodong Mu, Jun Wang 0006, Yuan Xie 0001, Zhizhong Zhang 0001, Shouhong Ding |
NeurIPS | 9 |
| 2025 | Point Mask Transformer for Outdoor Point Cloud Semantic SegmentationabstractCurrent outdoor point-cloud segmentation methods typically formulate semantic segmentation as a per-point/voxel-classification task. Although this strategy is straightforward because it classifies each point directly, it ignores the overall relationship of the category. As an alternative paradigm, mask classification decouples category classification from region localization, allowing the model to better capture overall category relationships. In this paper, we propose a novel approach called the point mask transformer (PMFormer), which transforms the semantic segmentation of point clouds from per-point classification to mask classification using a transformer architecture. The proposed model comprises a 3D backbone, transformer decoder, and segmentation head that predicts a series of binary masks, each associated with a global class label. Furthermore, to accommodate the unique characteristics of large and sparse outdoor point-cloud scenes, we propose three enhancements for the integration of point-cloud data with the transformer: MaskMix, 3D position encoding, and attention weights. We evaluate our model using the SemanticKITTI and nuScenes datasets. Our experimental results show that the proposed method outperforms state-of-the-art semantic segmentation approaches. Xin Tan 0002, Zhizhong Zhang 0001, Yuan Xie 0006, Lizhuang Ma |
Comput. Vis. Media | 3 |
| 2025 | Optimal Transport with Arbitrary Prior for Dynamic Resolution Network
Zhizhong Zhang 0001, Chenyang Zhang 0003, Lizhuang Ma, Xin Tan 0002, Yuan Xie 0006 |
Int. J. Comput. Vis. | 1 |
| 2025 | Gradient Projection for Continual Parameter-Efficient TuningabstractParameter-efficient tunings (PETs) have demonstrated impressive performance and promising perspectives in training large models, while they are still confronted with a common problem: the trade-off between learning new content and protecting old knowledge, leading to zero-shot generalization collapse, and cross-modal hallucination. In this paper, we reformulate Adapter, LoRA, Prefix-tuning, and Prompt-tuning from the perspective of gradient projection, and first propose a unified framework called Parameter Efficient Gradient Projection (PEGP). We introduce orthogonal gradient projection into different PET paradigms and theoretically demonstrate that the orthogonal condition for the gradient can effectively resist forgetting even for large-scale models. It therefore modifies the gradient towards the direction that has less impact on the old feature space, with less extra memory space and training time. We extensively evaluate our method with different backbones, including ViT and CLIP, on diverse datasets, and experiments comprehensively demonstrate its efficiency in reducing forgetting in class, online class, domain, task, and multi-modality continual settings. Jingyang Qiao, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Wensheng Zhang 0002, Zhi Han, Yuan Xie 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Identity-aware infrared person image generation and re-identification via controllable diffusion model
Xizhuo Yu, Chaojie Fan, Zhizhong Zhang 0001, Tianjian Yu, Yong Peng 0002 |
Pattern Recognit. | 3 |
| 2025 | Adaptive Pseudo-Label Purification and Debiasing for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised Visible-Infrared Person Re-Identification (USVI-ReID) aims to match visible and infrared person images without relying on prior annotations. Recently, unsupervised contrastive learning methods have become the mainstream approach for USVI-ReID, leveraging clustering algorithms to generate pseudo-labels. However, these methods often suffer from inherent noisy pseudo-labels, which significantly hinders their performance. To address this challenge, we propose a Adaptive Pseudo-label Purification and Debiasing (APPD) framework for USVI-ReID, which is designed to calibrate noisy pseudo-labels and dynamically detects clean pseudo-labels, thereby enhancing the model’s performance and reliability. Specifically, we propose an Adaptive Pseudo-label Calibration and Division (APCD) module, which calibrates noisy pseudo-labels by assessing their reliability and divides pseudo-labels into clean and noisy subsets, ensuring a more focused and accurate learning process. Based on the calibrated pseudo-labels, we develop an Optimal Transport Prototype Matching (OTPM) module to establish robust cross-modality correspondences. For clean pseudo-labels, we propose a Debiased Memory Hybrid Learning (DMHL) module, which jointly captures modality-specific and modality-invariant information while addressing sampling bias to enhance feature representation. To effectively utilize noisy pseudo-labels, we introduce a Neighbor Relation Learning (NRL) module that mitigates intra-class variations by exploring neighbor relationships in the feature space. Comprehensive experiments conducted on two widely recognized USVI-ReID benchmarks demonstrate that APPD achieves state-of-the-art performance, significantly outperforming existing methods. The source code will be made available at https://github.com/XiangboYin/RPNR. Xiangbo Yin, Jiangming Shi, Zhizhong Zhang 0001, Yuan Xie 0006, Yanyun Qu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | GEOcc: Geometrically Enhanced 3D Occupancy Network With Implicit-Explicit Depth Fusion and Contextual Self-Supervisionabstract3D occupancy perception holds a pivotal role in recent vision-centric autonomous driving systems by converting surround-view images into integrated geometric and semantic representations within dense 3D grids. Nevertheless, current models still encounter two main challenges: modeling depth accurately in the 2D-3D view transformation stage, and overcoming the lack of generalizability issues due to sparse LiDAR supervision. To address these issues, this paper presents GEOcc, a Geometric-Enhanced Occupancy network tailored for vision-only surround-view perception. Our approach is three-fold: 1) Integration of explicit lift-based depth prediction and implicit projection-based transformers for depth modeling, enhancing the density and robustness of view transformation. 2) Utilization of mask-based encoder-decoder architecture for fine-grained semantic predictions; 3) Adoption of context-aware self-training loss functions in the pertaining stage to complement LiDAR supervision, involving the re-rendering of 2D depth maps from 3D occupancy features and leveraging image reconstruction loss to obtain denser depth supervision besides sparse LiDAR ground-truths. Our approach achieves State-of-the-Art performance on the Occ3D-nuScenes dataset with the least image resolution needed and the most weightless image backbone compared with current models, marking an improvement of 3.3% due to our proposed contributions. Comprehensive experimentation also demonstrates the consistent superiority of our method over baselines and alternative approaches. Our code is available athttps://github.com/world-executed/GEOcc.git Xin Tan 0002, Zhiwei Zhang 0005, Chaojie Fan, Yong Peng 0002, Zhizhong Zhang 0001, Yuan Xie 0006, Lizhuang Ma |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Cross-Modal Recipe Retrieval With Fine-Grained Prompting Alignment and Evidential Semantic ConsistencyabstractAlignment between the food images and the corresponding recipes is an emerging cross-modal representation learning task. In this task, the recipes are composed of three components, i.e., food title, ingredient lists, and cooking instructions, which require a fine-grained alignment between the features of the two modalities. Existing methods usually aggregate the recipes into global embeddings and then align them with the global image embeddings. Meanwhile, semantic classification is frequently used in these methods to regularize the embeddings of the two modalities. While these methods are efficient, there remain two problems: (1) Forcing the alignment between the global images and recipes embeddings may result in losing the component-specific information. (2) The high diversity of food appearance leads to high uncertainty in the semantic classification of food images and recipes. To solve these problems, we propose a Fine-grained Prompting and Alignment (FPA) model to enhance the feature extraction and bring more component-specific information for fine-grained alignment. Furthermore, to regularize the semantic information contained in the cross-modal features, we design an Evidential Semantic Consistency (ESC) loss to keep the cross-modal semantic consistency. We have conducted comprehensive experiments on the benchmark dataset Recipe1M and the state-of-the-art results on the cross-modal recipe retrieval task demonstrate the effectiveness of our method. Jin Liu 0016, Zhizhong Zhang 0001, Yuan Xie 0006, Yongqiang Tang, Wensheng Zhang 0002, Xiaohui Cui |
IEEE Trans. Multim. | 3 |
| 2025 | Bias to Balance: New-Knowledge-Preferred Few-Shot Class-Incremental Learning via Transition CalibrationabstractHumans can quickly learn new concepts with limited experience, while not forgetting learned knowledge. Such ability in machine learning is referred to as few-shot class-incremental learning (FSCIL). Although some methods try to solve this problem by putting similar efforts to prevent forgetting and promote learning, we find existing techniques do not give enough importance to the new category as new training samples are rather rare. In this article, we propose a new biased-to-unbiased rectification method, which introduces a trainable transition matrix to mitigate the prediction discrepancy between the old classes and the new classes. This transition matrix is to be diagonally dominated, normalized, and differentiable with new-knowledge-preferred prior, to solving the strong bias between heavy old knowledge and limited new knowledge. Hence, we can achieve a balanced solution between learning new concepts and preventing catastrophic forgetting by giving new classes more chances. Extensive experiments on miniImagenet, CIFAR100, and CUB200 demonstrate that our method outperforms the latest state-of-the-art methods by 1.1%, 1.44%, and 2.08%, respectively. Hongquan Zhang, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Yuan Xie 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Beyond the Label Itself: Latent Labels Enhance Semi-supervised Point Cloud Panoptic SegmentationabstractAs the exorbitant expense of labeling autopilot datasets and the growing trend of utilizing unlabeled data, semi-supervised segmentation on point clouds becomes increasingly imperative. Intuitively, finding out more ``unspoken words'' (i.e., latent instance information) beyond the label itself should be helpful to improve performance. In this paper, we discover two types of latent labels behind the displayed label embedded in LiDAR and image data. First, in the LiDAR Branch, we propose a novel augmentation, Cylinder-Mix, which is able to augment more yet reliable samples for training. Second, in the Image Branch, we propose the Instance Position-scale Learning (IPSL) Module to learn and fuse the information of instance position and scale, which is from a 2D pre-trained detector and a type of latent label obtained from 3D to 2D projection. Finally, the two latent labels are embedded into the multi-modal panoptic segmentation network. The ablation of the IPSL module demonstrates its robust adaptability, and the experiments evaluated on SemanticKITTI and nuScenes demonstrate that our model outperforms the state-of-the-art method, LaserMix. Yujun Chen, Xin Tan 0002, Zhizhong Zhang 0001, Yanyun Qu, Yuan Xie 0006 |
AAAI | 3 |
| 2024 | Learning Task-Aware Language-Image Representation for Class-Incremental Object DetectionabstractClass-incremental object detection (CIOD) is a real-world desired capability, requiring an object detector to continuously adapt to new tasks without forgetting learned ones, with the main challenge being catastrophic forgetting. Many methods based on distillation and replay have been proposed to alleviate this problem. However, they typically learn on a pure visual backbone, neglecting the powerful representation capabilities of textual cues, which to some extent limits their performance. In this paper, we propose task-aware language-image representation to mitigate catastrophic forgetting, introducing a new paradigm for language-image-based CIOD. First of all, we demonstrate the significant advantage of language-image detectors in mitigating catastrophic forgetting. Secondly, we propose a learning task-aware language-image representation method that overcomes the existing drawback of directly utilizing the language-image detector for CIOD. More specifically, we learn the language-image representation of different tasks through an insulating approach in the training stage, while using the alignment scores produced by task-specific language-image representation in the inference stage. Through our proposed method, language-image detectors can be more practical for CIOD. We conduct extensive experiments on COCO 2017 and Pascal VOC 2007 and demonstrate that the proposed method achieves state-of-the-art results under the various CIOD settings. Hongquan Zhang, Bin-Bin Gao, Yi Zeng 0006, Xin Tan 0002, Zhizhong Zhang 0001, Yanyun Qu, Jun Liu 0116, Yuan Xie 0006 |
AAAI | 6 |
| 2024 | Explore and Enhance the Generalization of Anomaly DeepFake Detection
Shen Chen 0004, Taiping Yao, Lizhuang Ma, Zhizhong Zhang 0001, Xin Tan 0002 |
CVM (2) | 5 |
| 2024 | Isolation and Integration: A Strong Pre-trained Model-Based Paradigm for Class-Incremental Learning
Wei Zhang 0217, Yuan Xie 0006, Zhizhong Zhang 0001, Xin Tan 0002 |
CVM (2) | 3 |
| 2024 | Building a Strong Pre-Training Baseline for Universal 3D Large-Scale PerceptionabstractAn effective pre-training framework with universal 3D representations is extremely desired in perceiving large- scale dynamic scenes. However, establishing such an ideal framework that is both task-generic and label-efficient poses a challenge in unifying the representation of the same primitive across diverse scenes. The current contrastive 3D pre-training methods typically follow a frame-level consistency, which focuses on the 2D-3D relationships in each detached image. Such inconsiderate consistency greatly hampers the promising path of reaching an universal pre-training framework: (1) The cross-scene semantic self-conflict, i.e., the intense collision between primitive segments of the same semantics from different scenes; (2) Lacking a globally unified bond that pushes the cross-scene semantic consistency into 3D representation learning. To address above challenges, we propose a CSC framework that puts a scene-level semantic consistency in the heart, bridging the connection of the similar semantic segments across various scenes. To achieve this goal, we combine the coherent semantic cues provided by the vision foundation model and the knowledge-rich cross-scene prototypes derived from the complementary multi-modality information. These allow us to train a universal 3D pre-training model that facilitates various downstream tasks with less fine-tuning efforts. Empirically, we achieve consistent improvements over SOTA pre-training approaches in semantic segmentation (+1.4% mIoU), object detection (+ 1.0% mAP), and panoptic segmentation (+3.0% PQ) using their task-specific 3D network on nuScenes. Code is released at https://github.com/chenhaomingbob/CSC, hoping to inspire future research. Haoming Chen, Zhizhong Zhang 0001, Yanyun Qu, Xin Tan 0002, Yuan Xie 0006 |
CVPR | 2 |
| 2024 | PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly DetectionabstractThe vision-language model has brought great improvement to few-shot industrial anomaly detection, which usually needs to design of hundreds of prompts through prompt engineering. For automated scenarios, we first use conventional prompt learning with many-class paradigm as the baseline to automatically learn prompts but found that it can not work well in one-class anomaly detection. To address the above problem, this paper proposes a one-class prompt learning method for few-shot anomaly detection, termed PromptAD. First, we propose semantic concatenation which can transpose normal prompts into anomaly prompts by concatenating normal prompts with anomaly suffixes, thus constructing a large number of negative samples used to guide prompt learning in one-class setting. Furthermore, to mitigate the training challenge caused by the absence of anomaly images, we introduce the concept of explicit anomaly margin, which is used to explicitly control the margin between normal prompt features and anomaly prompt features through a hyper-parameter. For image-level/pixel-level anomaly detection, PromptAD achieves first place in 11/12 few-shot settings on MVTec and VisA. Code is available at https://github.com/FuNz-0/PromptAD.git Xiaofan Li 0008, Zhizhong Zhang 0001, Xin Tan 0002, Chengwei Chen, Yanyun Qu, Yuan Xie 0006, Lizhuang Ma |
CVPR | 2 |
| 2024 | COTR: Compact Occupancy TRansformer for Vision-Based 3D Occupancy PredictionabstractThe autonomous driving community has shown significant interest in 3D occupancy prediction, driven by its exceptional geometric perception and general object recognition capabilities. To achieve this, current works try to construct a Tri-Perspective View (TPV) or Occupancy (OCC) representation extending from the Bird-Eye-View perception. However, compressed views like TPV representation lose 3D geometry information while raw and sparse OCC representation requires heavy but redundant computational costs. To address the above limitations, we propose Compact Occupancy TRansformer (COTR), with a geometry-aware occupancy encoder and a semantic-aware group decoder to reconstruct a compact 3D OCC representation. The occupancy encoder first generates a compact geometrical OCC feature through efficient explicit-implicit view transformation. Then, the occupancy decoder further enhances the semantic discriminability of the compact OCC representation by a coarse-to-fine semantic grouping strategy. Empirical experiments show that there are evident performance gains across multiple baselines, e.g., COTR outperforms baselines with a relative improvement of 8%-15%, demonstrating the superiority of our method. The code is available at https://github.com/NotACracker/COTR. Qihang Ma, Xin Tan 0002, Yanyun Qu, Lizhuang Ma, Zhizhong Zhang 0001, Yuan Xie 0006 |
CVPR | 5 |
| 2024 | Multi-modal In-Context Learning Makes an Ego-evolving Scene Text RecognizerabstractScene text recognition (STR) in the wild frequently en-counters challenges when coping with domain variations, font diversity, shape deformations, etc. A straightforward solution is performing model fine-tuning tailored to a spe-cific scenario, but it is computationally intensive and re-quires multiple model copies for various scenarios. Re-cent studies indicate that large language models (LLMs) can learn from afew demonstration examples in a training-free manner, termed “In-Context Learning” (ICL). Never-theless, applying LLMs as a text recognizer is unacceptably resource-consuming. Moreover, our pilot experiments on LLMs show that ICL fails in STR, mainly attributed to the insufficient incorporation of contextual information from di-verse samples in the training stage. To this end, we intro-duce E2 STR, a STR model trained with context-rich scene text sequences, where the sequences are generated via our proposed in-context training strategy. E2 STR demonstrates that a regular-sized model is sufficient to achieve effective ICL capabilities in STR. Extensive experiments show that E2 STR exhibits remarkable training-free adaptation in var-ious scenarios and outperforms even the fine-tuned state-of-the-art approaches on public benchmarks. The code is released at https://github.com/bytedanceIE2STR. Jingqun Tang, Chunhui Lin, Binghong Wu, Can Huang 0002, Hao Liu 0003, Xin Tan 0002, Zhizhong Zhang 0001, Yuan Xie 0006 |
CVPR | 8 |
| 2024 | Multi-memory Matching for Unsupervised Visible-Infrared Person Re-identification
Jiangming Shi, Xiangbo Yin, Yeyun Chen, Yachao Zhang 0001, Zhizhong Zhang 0001, Yuan Xie 0006, Yanyun Qu |
ECCV (18) | 5 |
| 2024 | Prompt Gradient Projection for Continual LearningabstractPrompt-tuning has demonstrated impressive performance in continual learning by querying relevant prompts for each input instance, which can avoid the introduction of task identifier. Its forgetting is therefore reduced as this instance-wise query mechanism enables us to select and update only relevant prompts. In this paper, we further integrate prompt-tuning with gradient projection approach. Our observation is: prompt-tuning releases the necessity of task identifier for gradient projection method; and gradient projection provides theoretical guarantees against forgetting for prompt-tuning. This inspires a new prompt gradient projection approach (PGP) for continual learning. In PGP, we deduce that reaching the orthogonal condition for prompt gradient can effectively prevent forgetting via the self-attention mechanism in vision-transformer. The condition equations are then realized by conducting Singular Value Decomposition (SVD) on an element-wise sum space between input space and prompt space. We validate our method on diverse datasets and experiments demonstrate the efficiency of reducing forgetting both in class incremental, online class incremental, and task incremental settings. The code is available at https://github.com/JingyangQiao/prompt-gradient-projection. Jingyang Qiao, Zhizhong Zhang 0001, Xin Tan 0002, Chengwei Chen, Yanyun Qu, Yong Peng 0002, Yuan Xie 0006 |
ICLR | 2 |
| 2024 | Mutual Positive and Negative Learning for Weakly-supervised Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation heavily relies on the large-scale point-level annotated dataset, which encourages the weakly-supervised method to prevail gradually. Previous weakly-supervised self-training methods only adopted positive labels, which would be under-performed due to too much noise and lack of supervision. We are the first to present negative labels into the 3D segmentation area, providing extra supervision for hard samples to mitigate the drawbacks induced by noisy labels. Together with the positive labels, we formulate a Mutual Positive-Negative Bi-branch Learning framework to generate positive and negative labels iteratively. Based on that positive branch and negative branch learn complementary knowledge, we build a Mutual Positive-Negative Knowledge Distillation within the bi-branch to further encourage the two branches to learn from each other. Finally, we propose a novel dynamic fusion strategy to fuse predictions from the positive and negative branches, generating more robust predictions. Results on three large-scale datasets show that our method outperforms state-of-the-art weakly-supervised methods by a large margin. Zhizhong Zhang 0001, Yuan Xie 0006, Guchu Zou, Zhenyi Qi, Xin Tan 0002 |
ICME | 3 |
| 2024 | Robust Pseudo-label Learning with Neighbor Relation for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised Visible-Infrared Person Re-identification (USVI-ReID) presents a formidable challenge, which aims to match pedestrian images across visible and infrared modalities without any annotations. Recently, clustered pseudo-label methods have become predominant in USVI-ReID, although the inherent noise in pseudo-labels presents a significant obstacle. Most existing works primarily focus on shielding the model from the harmful effects of noise, neglecting to calibrate noisy pseudo-labels usually associated with hard samples, which will compromise the robustness of the model. To address this issue, we design a Robust Pseudo-label Learning with Neighbor Relation (RPNR) framework for USVI-ReID. To be specific, we first introduce a straightforward yet potent Noisy Pseudo-label Calibration module to correct noisy pseudo-labels. Due to the high intra-class variations, noisy pseudo-labels are difficult to calibrate completely. Therefore, we introduce a Neighbor Relation Learning module to reduce high intra-class variations by modeling potential interactions between all samples. Subsequently, we devise an Optimal Transport Prototype Matching module to establish reliable cross-modality correspondences. On that basis, we design a Memory Hybrid Learning module to jointly learn modality-specific and modality-invariant information. Comprehensive experiments conducted on two widely recognized benchmarks, SYSU-MM01 and RegDB, demonstrate that RPNR outperforms the current state-of-the-art GUR with an average Rank-1 improvement of 10.3%. The code is available at https://github.com/XiangboYin/RPNR. Xiangbo Yin, Jiangming Shi, Yachao Zhang 0001, Yang Lu 0009, Zhizhong Zhang 0001, Yuan Xie 0006, Yanyun Qu |
ACM Multimedia | 5 |
| 2024 | Learning Commonality, Divergence and Variety for Unsupervised Visible-Infrared Person Re-identificationabstractUnsupervised visible-infrared person re-identification (USVI-ReID) aims to match specified persons in infrared images to visible images without annotations, and vice versa. USVI-ReID is a challenging yet underexplored task. Most existing methods address the USVI-ReID through cluster-based contrastive learning, which simply employs the cluster center to represent an individual. However, the cluster center primarily focuses on commonality, overlooking divergence and variety. To address the problem, we propose a Progressive Contrastive Learning with Hard and Dynamic Prototypes for USVI-ReID. In brief, we generate the hard prototype by selecting the sample with the maximum distance from the cluster center. We reveal that the inclusion of the hard prototype in contrastive loss helps to emphasize divergence. Additionally, instead of rigidly aligning query images to a specific prototype, we generate the dynamic prototype by randomly picking samples within a cluster. The dynamic prototype is used to encourage variety. Finally, we introduce a progressive learning strategy to gradually shift the model's attention towards divergence and variety, avoiding cluster deterioration. Extensive experiments conducted on the publicly available SYSU-MM01 and RegDB datasets validate the effectiveness of the proposed method. Jiangming Shi, Xiangbo Yin, Yachao Zhang 0001, Zhizhong Zhang 0001, Yuan Xie 0001, Yanyun Qu |
NeurIPS | 4 |
| 2024 | Harmonizing Visual Text Comprehension and GenerationabstractIn this work, we present TextHarmony, a unified and versatile multimodal generative model proficient in comprehending and generating visual text. Simultaneously generating images and texts typically results in performance degradation due to the inherent inconsistency between vision and language modalities. To overcome this challenge, existing approaches resort to modality-specific data for supervised fine-tuning, necessitating distinct model instances. We propose Slide-LoRA, which dynamically aggregates modality-specific and modality-agnostic LoRA experts, partially decoupling the multimodal generation space. Slide-LoRA harmonizes the generation of vision and language within a singular model instance, thereby facilitating a more unified generative process. Additionally, we develop a high-quality image caption dataset, DetailedTextCaps-100K, synthesized with a sophisticated closed-source MLLM to enhance visual text generation capabilities further. Comprehensive experiments across various benchmarks demonstrate the effectiveness of the proposed approach. Empowered by Slide-LoRA, TextHarmony achieves comparable performance to modality-specific fine-tuning results with only a 2% increase in parameters and shows an average improvement of 2.5% in visual text comprehension tasks and 4.0% in visual text generation tasks. Our work delineates the viability of an integrated approach to multimodal generation within the visual text domain, setting a foundation for subsequent inquiries. Code is available at https://github.com/bytedance/TextHarmony. Jingqun Tang, Binghong Wu, Chunhui Lin, Shu Wei, Hao Liu 0003, Xin Tan 0002, Zhizhong Zhang 0001, Can Huang 0002, Yuan Xie 0006 |
NeurIPS | 8 |
| 2024 | Uni-to-Multi Modal Knowledge Distillation for Bidirectional LiDAR-Camera Semantic SegmentationabstractCombining LiDAR points and images for robust semantic segmentation has shown great potential. However, the heterogeneity between the two modalities (e.g. the density, the field of view) poses challenges in establishing a bijective mapping between each point and pixel. This modality alignment problem introduces new challenges in network design and data processing for cross-modal methods. Specifically, 1) points that are projected outside the image planes; 2) the complexity of maintaining geometric consistency limits the deployment of many data augmentation techniques. To address these challenges, we propose a cross-modal knowledge imputation and transition approach. First, we introduce a bidirectional feature fusion strategy that imputes missing image features and performs cross-modal fusion simultaneously. This allows us to generate reliable predictions even when images are missing. Second, we propose a Uni-to-Multi modal Knowledge Distillation (U2MKD) framework, leveraging the transfer of informative features from a single-modality teacher to a cross-modality student. This overcomes the issues of augmentation misalignment and enables us to train the student effectively. Extensive experiments on the nuScenes, Waymo, and SemanticKITTI datasets demonstrate the effectiveness of our approach. Notably, our method achieves an 8.3 mIoU gain over the LiDAR-only baseline on the nuScenes validation set and achieves state-of-the-art performance on the three datasets. Tianfang Sun, Zhizhong Zhang 0001, Xin Tan 0002, Yong Peng 0002, Yanyun Qu, Yuan Xie 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Variational Distillation for Multi-View LearningabstractInformation Bottleneck (IB) provides an information-theoretic principle for multi-view learning by revealing the various components contained in each viewpoint. This highlights the necessity to capture their distinct roles to achieve view-invariance and predictive representations but remains under-explored due to the technical intractability of modeling and organizing innumerable mutual information (MI) terms. Recent studies show that sufficiency and consistency play such key roles in multi-view representation learning, and could be preserved via a variational distillation framework. But when it generalizes to arbitrary viewpoints, such strategy fails as the mutual information terms of consistency become complicated. This paper presents Multi-View Variational Distillation (MV$^{2}$D), tackling the above limitations for generalized multi-view learning. Uniquely, MV$^{2}$D can recognize useful consistent information and prioritize diverse components by their generalization ability. This guides an analytical and scalable solution to achieving both sufficiency and consistency. Additionally, by rigorously reformulating the IB objective, MV$^{2}$D tackles the difficulties in MI optimization and fully realizes the theoretical advantages of the information bottleneck principle. We extensively evaluate our model on diverse tasks to verify its effectiveness, where the considerable gains provide key insights into achieving generalized multi-view representations under a rigorous information-theoretic principle. Zhizhong Zhang 0001, Cong Wang 0039, Wensheng Zhang 0002, Yanyun Qu, Lizhuang Ma, Zongze Wu 0001, Yuan Xie 0006, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Dynamic image super-resolution via progressive contrastive self-distillation
Zhizhong Zhang 0001, Yuan Xie 0006, Yanbo Wang 0003, Yanyun Qu, Shaohui Lin, Lizhuang Ma, Qi Tian 0001 |
Pattern Recognit. | 1 |
| 2024 | Glass Makes Blurs: Learning the Visual Blurriness for Glass Surface DetectionabstractGlass surface detection is challenging as glass normally borrows similar visual appearances from the arbitrary objects/scenes behind it. Although some methods have been proposed to address this problem, they may fail if the reference objects are nonexistent or the additional annotations are missing. This article aims to address the glass surface detection problem by utilizing the intrinsic glass properties without reference objects and additional annotations. We observe glass makes blurs naturally. Based on the investigation of this intrinsic visual blurriness cue, we propose a novel visual blurriness aggregation module to model visual blurriness as a learnable residual in order to extract and aggregate multiscale valuable visual blurriness features used for guiding the backbone features to detect glass precisely. Besides, we note the ratio of the blurred area assists in utilizing the visual blurriness cue caused by glass and propose a visual blurriness driven refinement module to refine glass maps with this ratio to better leverage the visual blurriness information. Extensive experiments show that the proposed method achieves state-of-the-art performance on popular glass surface datasets. Fulin Qi, Xin Tan 0002, Zhizhong Zhang 0001, Mingang Chen, Yuan Xie 0006, Lizhuang Ma |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Image Understands Point Cloud: Weakly Supervised 3D Semantic Segmentation via Association LearningabstractWeakly supervised point cloud semantic segmentation methods that require 1% or fewer labels with the aim of realizing almost the same performance as fully supervised approaches have recently attracted extensive research attention. A typical solution in this framework is to use self-training or pseudo-labeling to mine the supervision from the point cloud itself while ignoring the critical information from images. In fact, cameras widely exist in LiDAR scenarios, and this complementary information seems to be highly important for 3D applications. In this paper, we propose a novel cross-modality weakly supervised method for 3D segmentation that incorporates complementary information from unlabeled images. We design a dual-branch network equipped with an active labeling strategy to maximize the power of tiny parts of labels and to directly realize 2D-to-3D knowledge transfer. Afterward, we establish a cross-modal self-training framework, which iterates between parameter updating and pseudolabel estimation. In the training phase, we propose cross-modal association learning to mine complementary supervision from images by reinforcing the cycle consistency between 3D points and 2D superpixels. In the pseudolabel estimation phase, a pseudolabel self-rectification mechanism is derived to filter noisy labels, thus providing more accurate labels for the networks to be fully trained. The extensive experimental results demonstrate that our method even outperforms the state-of-the-art fully supervised competitors with less than 1% actively selected annotations. Tianfang Sun, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Yuan Xie 0006 |
IEEE Trans. Image Process. | 2 |
| 2023 | High-Resolution GAN Inversion for Degraded Images in Large Diverse DatasetsabstractThe last decades are marked by massive and diverse image data, which shows increasingly high resolution and quality. However, some images we obtained may be corrupted, affecting the perception and the application of downstream tasks. A generic method for generating a high-quality image from the degraded one is in demand. In this paper, we present a novel GAN inversion framework that utilizes the powerful generative ability of StyleGAN-XL for this problem. To ease the inversion challenge with StyleGAN-XL, Clustering \& Regularize Inversion (CRI) is proposed. Specifically, the latent space is firstly divided into finer-grained sub-spaces by clustering. Instead of initializing the inversion with the average latent vector, we approximate a centroid latent vector from the clusters, which generates an image close to the input image. Then, an offset with a regularization term is introduced to keep the inverted latent vector within a certain range. We validate our CRI scheme on multiple restoration tasks (i.e., inpainting, colorization, and super-resolution) of complex natural images, and show preferable quantitative and qualitative results. We further demonstrate our technique is robust in terms of data and different GAN models. To our best knowledge, we are the first to adopt StyleGAN-XL for generating high-quality natural images from diverse degraded inputs. Code is available at https://github.com/Booooooooooo/CRI. Yanbo Wang 0003, Chuming Lin, Donghao Luo 0001, Ying Tai, Zhizhong Zhang 0001, Yuan Xie 0006 |
AAAI | 5 |
| 2023 | Self-supervised Contrastive Feature Refinement for Few-Shot Class-Incremental Learning
Shengjin Ma, Wang Yuan, Xin Tan 0002, Zhizhong Zhang 0001, Lizhuang Ma |
CAD/Graphics | 5 |
| 2023 | Multi-Centroid Task Descriptor for Dynamic Class Incremental InferenceabstractIncremental learning could be roughly divided into two categories, i.e., class- and task-incremental learning. The main difference is whether the task ID is given during evaluation. In this paper, we show this task information is indeed a strong prior knowledge, which will bring significant improvement over class-incremental learning baseline, e.g., DER [39]. Based on this observation, we propose a gate network to predict the task ID for class incremental inference. This is challenging as there is no explicit semantic relationship between categories in the concept of task. Therefore, we propose a multi-centroid task descriptor by assuming the data within a task can form multiple clusters. The cluster centers are optimized by pulling relevant sample-centroid pairs while pushing others away, which ensures that there is at least one centroid close to a given sample. To select relevant pairs, we use class prototypes as proxies and solve a bipartite matching problem, making the task descriptor representative yet not degenerate to uni-modal. As a result, our dynamic inference network is trained independently of baseline and provides a flexible, efficient solution to distinguish between tasks. Extensive experiments show our approach achieves state-of-the-art results, e.g., we achieve 72.41% average accuracy on CIFAR100-BOS50, outperforming DER by 3.40%. Tenghao Cai, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Guannan Jiang, Chengjie Wang 0001, Yuan Xie 0006 |
CVPR | 2 |
| 2023 | Rethinking Gradient Projection Continual Learning: Stability/Plasticity Feature Space DecouplingabstractContinual learning aims to incrementally learn novel classes over time, while not forgetting the learned knowledge. Recent studies have found that learning would not forget if the updated gradient is orthogonal to the feature space. However, previous approaches require the gradient to be fully orthogonal to the whole feature space, leading to poor plasticity, as the feasible gradient direction becomes narrow when the tasks continually come, i.e., feature space is unlimitedly expanded. In this paper, we propose a space decoupling (SD) algorithm to decouple the feature space into a pair of complementary subspaces, i.e., the stability space$\mathcal{I}$and the plasticity space$\mathcal{R}. \mathcal{I}$is established by conducting space intersection between the historic and current feature space, and thus$\mathcal{I}$contains more task-shared bases.$\mathcal{R}$is constructed by seeking the orthogonal complementary subspace of$T$and thus$\mathcal{R}$mainly contains task-specific bases. By putting distinguishing constraints on$\mathcal{R}$and$\mathcal{I}$, our method achieves a better balance between stability and plasticity. Extensive experiments are conducted by applying SD to gradient projection baselines, and show SD is model-agnostic and achieves SOTA results on publicly available datasets. Zhizhong Zhang 0001, Xin Tan 0002, Jun Liu 0116, Yanyun Qu, Yuan Xie 0006, Lizhuang Ma |
CVPR | 2 |
| 2023 | Dual Pseudo-Labels Interactive Self-Training for Semi-Supervised Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) aims to match a specific person from a gallery of images captured from non-overlapping visible and infrared cameras. Most works focus on fully supervised VI-ReID, which requires substantial cross-modality annotation that is more expensive than the annotation in single-modality. To reduce the extensive cost of annotation, we explore two practical semi-supervised settings: uni-semi-supervised (annotating only visible images) and bi-semi-supervised (annotating partially in both modalities). These two semi-supervised settings face two challenges due to the large cross-modality discrepancies and the lack of correspondence supervision between visible and infrared images. Thus, it is diffi-cult to generate reliable pseudo-labels and learn modality-invariant features from noise pseudo-labels. In this paper, we propose a dual pseudo-label interactive self-training (DPIS) for these two semi-supervised VI-ReID. Our DPIS integrates two pseudo-labels generated by distinct models into a hybrid pseudo-label for unlabeled data. However, the hybrid pseudo-label still inevitably contains noise. To eliminate the negative effect of noise pseudo-labels, we introduce three modules: noise label penalty (NLP), noise correspondence calibration (NCC), and unreliable anchor learning (UAL). Specifically, NLP penalizes noise labels, NCC calibrates noisy correspondences, and UAL mines the hard-to-discriminate features. Extensive experimental results on SYSU-MM01 and RegDB demonstrate that our DPIS achieves impressive performance under these two semi-supervised settings. Jiangming Shi, Yachao Zhang 0001, Xiangbo Yin, Yuan Xie 0006, Zhizhong Zhang 0001, Jianping Fan 0007, Zhongchao Shi, Yanyun Qu |
ICCV | 5 |
| 2023 | Instance and Category Supervision are Alternate Learners for Continual LearningabstractContinual Learning (CL) is the constant development of complex behaviors by building upon previously acquired skills. Yet, current CL algorithms tend to incur class-level forgetting as the label information is often quickly overwritten by new knowledge. This motivates attempts to mine instance-level discrimination by resorting to recent self-supervised learning (SSL) techniques. However, previous works have pointed out that the self-supervised learning objective is essentially a trade-off between invariance to distortion and preserving sample information, which seriously hinders the unleashing of instance-level discrimination.In this work, we reformulate SSL from the information-theoretic perspective by disentangling the goal of instance-level discrimination, and tackle the trade-off to promote compact representations with maximally preserved invariance to distortion. On this basis, we develop a novel alternate learning paradigm to enjoy the complementary merits of instance-level and category-level supervision, which yields improved robustness against forgetting and better adaptation to each task. To verify the proposed method, we conduct extensive experiments on four different benchmarks using both class-incremental and task-incremental settings, where the leap in performance and thorough ablation studies demonstrate the efficacy and efficiency of our modeling strategy. Zhizhong Zhang 0001, Xin Tan 0002, Jun Liu 0116, Chengjie Wang 0001, Yanyun Qu, Guannan Jiang, Yuan Xie 0006 |
ICCV | 2 |
| 2023 | LiDAR-Camera Panoptic Segmentation via Geometry-Consistent and Semantic-Aware Alignmentabstract3D panoptic segmentation is a challenging perception task that requires both semantic segmentation and instance segmentation. In this task, we notice that images could provide rich texture, color, and discriminative information, which can complement LiDAR data for evident performance improvement, but their fusion remains a challenging problem. To this end, we propose LCPS, the first LiDAR-Camera Panoptic Segmentation network. In our approach, we conduct LiDAR-Camera fusion in three stages: 1) an Asynchronous Compensation Pixel Alignment (ACPA) module that calibrates the coordinate misalignment caused by asynchronous problems between sensors; 2) a Semantic-Aware Region Alignment (SARA) module that extends the one-to-one point-pixel mapping to one-to-many semantic relations; 3) a Point-to-Voxel feature Propagation (PVP) module that integrates both geometric and semantic fusion information for the entire point cloud. Our fusion strategy improves about 6.9% PQ performance over the LiDAR-only baseline on NuScenes dataset. Extensive quantitative and qualitative experiments further demonstrate the effectiveness of our novel framework. The code will be released at https://github.com/zhangzw12319/lcps.git. Zhiwei Zhang 0005, Zhizhong Zhang 0001, Ran Yi 0002, Yuan Xie 0006, Lizhuang Ma |
ICCV | 2 |
| 2023 | CVTE-Poly: A New Benchmark for Chinese Polyphone Disambiguation
Siheng Zhang, Xingjun Tan, Yanqiang Lei, Xianxiang Wang, Zhizhong Zhang 0001, Yuan Xie 0006 |
INTERSPEECH | 5 |
| 2023 | Unveiling the Power of CLIP in Unsupervised Visible-Infrared Person Re-IdentificationabstractLarge-scale Vision-Language Pre-training (VLP) model, e.g., CLIP, has demonstrated its natural advantage in generating textual descriptions for images. These textual descriptions afford us greater semantic monitoring insights while not requiring any domain knowledge. In this paper, we propose a new prompt learning paradigm for unsupervised visible-infrared person re-identification (USL-VI-ReID) by taking full advantage of the visual-text representation ability from CLIP. In our framework, we establish a learnable cluster-aware prompt for person images and obtain textual descriptions allowing for subsequent unsupervised training. This description complements the rigid pseudo-labels and provides an important semantic supervised signal. On that basis, we propose a new memory-swapping contrastive learning, where we first find the correlated cross-modal prototypes by the Hungarian matching method and then swap the prototype pairs in the memory. Thus typical contrastive learning without any change could easily associate the cross-modal information. Extensive experiments on the benchmark datasets demonstrate the effectiveness of our method. For example, on SYSU-MM01 we arrive at 54.0% in terms of Rank-1 accuracy, over 9% improvement against state-of-the-art approaches. Code is available at https://github.com/CzAngus/CCLNet. Zhong Chen 0007, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Yuan Xie 0006 |
ACM Multimedia | 2 |
| 2023 | Improving Cross-Modal Recipe Retrieval with Component-Aware Prompted CLIP EmbeddingabstractCross-modal recipe retrieval is an emerging visual-textual retrieval task, which aims at matching food images with the corresponding recipes. Although large-scale Vision-Language Pre-training (VLP) models have achieved impressive performance on a wide range of downstream tasks, they still perform unsatisfactorily on this cross-modal retrieval task due to the following two problems: (1) Features from food images and recipes need to be aligned, simply fine-tuning the pre-trained VLP model's image encoder does not explicitly help with this goal. (2) The text content in the recipe is more structured than the text caption in the VLP model's pre-training corpus, which prevents the VLP model from adapting to the recipe retrieval task. In this paper, we propose a Component-aware Instance-specific Prompt learning (CIP) model that fully exploits the ability of large-scale VLP models. CIP enables us to learn the structured recipe information and therefore allows for aligning visual-textual representations without fine-tuning. Furthermore, we construct a recipe encoder termed Adaptive Recipe Merger (ARM) based on hierarchical Transformers, encouraging the model to learn more effective recipe representations. Extensive experiments on the public Recipe1M dataset demonstrate the superiority of our proposed method by outperforming the state-of-the-art methods on cross-modal recipe retrieval task. Jin Liu 0016, Zhizhong Zhang 0001, Yuan Xie 0006 |
ACM Multimedia | 3 |
| 2023 | Incremental Learning Based on Dual-Branch Network
Mingda Dong, Zhizhong Zhang 0001, Yuan Xie 0006 |
PRCV (3) | 2 |
| 2023 | Cross-stream contrastive learning for self-supervised skeleton-based action recognition
Ding Li 0006, Yongqiang Tang, Zhizhong Zhang 0001, Wensheng Zhang 0002 |
Image Vis. Comput. | 3 |
| 2023 | Positive-Negative Receptive Field Reasoning for Omni-Supervised 3D SegmentationabstractHidden features in the neural networks usually fail to learn informative representation for 3D segmentation as supervisions are only given on output prediction, while this can be solved by omni-scale supervision on intermediate layers. In this paper, we bring the first omni-scale supervision method to 3D segmentation via the proposed gradual Receptive Field Component Reasoning (RFCR), where target Receptive Field Component Codes (RFCCs) is designed to record categories within receptive fields for hidden units in the encoder. Then, target RFCCs will supervise the decoder to gradually infer the RFCCs in a coarse-to-fine categories reasoning manner, and finally obtain the semantic labels. To purchase more supervisions, we also propose an RFCR-NL model with complementary negative codes (i.e., Negative RFCCs, NRFCCs) with negative learning. Because many hidden features are inactive with tiny magnitudes and make minor contributions to RFCC prediction, we propose Feature Densification with a centrifugal potential to obtain more unambiguous features, and it is in effect equivalent to entropy regularization over features. More active features can unleash the potential of omni-supervision method. We embed our method into three prevailing backbones, which are significantly improved in all three datasets on both fully and weakly supervised segmentation tasks and achieve competitive performances. Xin Tan 0002, Qihang Ma, Jingyu Gong, Zhizhong Zhang 0001, Yanyun Qu, Yuan Xie 0006, Lizhuang Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Boosting Night-Time Scene Parsing With Learnable FrequencyabstractNight-Time Scene Parsing (NTSP) is essential to many vision applications, especially for autonomous driving. Most of the existing methods are proposed for day-time scene parsing. They rely on modeling pixel intensity-based spatial contextual cues under even illumination. Hence, these methods do not perform well in night-time scenes as such spatial contextual cues are buried in the over-/under-exposed regions in night-time scenes. In this paper, we first conduct an image frequency-based statistical experiment to interpret the day-time and night-time scene discrepancies. We find that image frequency distributions differ significantly between day-time and night-time scenes, and understanding such frequency distributions is critical to NTSP problem. Based on this, we propose to exploit the image frequency distributions for night-time scene parsing. First, we propose a Learnable Frequency Encoder (LFE) to model the relationship between different frequency coefficients to measure all frequency components dynamically. Second, we propose a Spatial Frequency Fusion module (SFF) that fuses both spatial and frequency information to guide the extraction of spatial context features. Extensive experiments show that our method performs favorably against the state-of-the-art methods on the NightCity, NightCity+ and BDD100K-night datasets. In addition, we demonstrate that our method can be applied to existing day-time scene parsing methods and boost their performance on night-time scenes. The code is available at https://github.com/wangsen99/FDLNet. Ke Xu 0010, Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006, Lizhuang Ma |
IEEE Trans. Image Process. | 4 |
| 2022 | Task-Level Self-Supervision for Cross-Domain Few-Shot LearningabstractLearning with limited labeled data is a long-standing problem. Among various solutions, episodic training progres-sively classifies a series of few-shot tasks and thereby is as-sumed to be beneficial for improving the model’s generalization ability. However, recent studies show that it is eveninferior to the baseline model when facing domain shift between base and novel classes. To tackle this problem, we pro-pose a domain-independent task-level self-supervised (TL-SS) method for cross-domain few-shot learning.TL-SS strategy promotes the general idea of label-based instance-levelsupervision to task-level self-supervision by augmenting mul-tiple views of tasks. Two regularizations on task consistencyand correlation metric are introduced to remarkably stabi-lize the training process and endow the generalization ability into the prediction model. We also propose a high-order associated encoder (HAE) being adaptive to various tasks.By utilizing 3D convolution module, HAE is able to generate proper parameters and enables the encoder to flexibly toany unseen tasks. Two modules complement each other andshow great promotion against state-of-the-art methods experimentally. Finally, we design a generalized task-agnostic test,where our intriguing findings highlight the need to re-think the generalization ability of existing few-shot approaches. Wang Yuan, Zhizhong Zhang 0001, Cong Wang 0039, Yuan Xie 0006, Lizhuang Ma |
AAAI | 2 |
| 2022 | Optimization over Disentangled Encoding: Unsupervised Cross-Domain Point Cloud Completion via Occlusion Factor Manipulation
Jingyu Gong, Fengqi Liu, Min Wang 0024, Xin Tan 0002, Zhizhong Zhang 0001, Ran Yi 0002, Yuan Xie 0006, Lizhuang Ma |
ECCV (2) | 6 |
| 2022 | Mutually Reinforcing Structure with Proposal Contrastive Consistency for Few-Shot Object Detection
TianXue Ma, Mingwei Bi, Jian Zhang 0079, Wang Yuan, Zhizhong Zhang 0001, Yuan Xie 0006, Shouhong Ding, Lizhuang Ma |
ECCV (20) | 5 |
| 2022 | Optimal Transport for Label-Efficient Visible-Infrared Person Re-Identification
Jiangming Wang, Zhizhong Zhang 0001, Mingang Chen, Cong Wang 0039, Bin Sheng 0001, Yanyun Qu, Yuan Xie 0006 |
ECCV (24) | 2 |
| 2022 | Self-Mimic Mutual-Distillation for Cross-Modality Person Re-IdentificationabstractCross-modality person re-identification is a newly rising and challenging problem, as there is a significant gap between the visible and infrared images. Though recent methods rapidly narrow the gap, the intra-modality variance is often ignored before inter-modality alignment. In this paper, we study this problem in the knowledge distillation perspective and design a self-mimic mutual-distillation method to reduce the discrepancy of each person from intra-modality feature alignment to cross-modality feature alignment. For intra-modality feature alignment, the self-mimic mechanism is implemented to simultaneously learn globally viewed, stable, and distinguish prototypes for each ID and minimize the intra-modality discrepancy. For inter-modality feature alignment, the mutual distillation is conducted to minimize the cross-modality distribution discrepancy of each person. Extensive experimental results on SYSU-MM01 and RegDB demonstrate that the proposed method achieves the best performance, outperforming state-of-the-art methods by a large margin without adding extra network parameters to the baseline. Especially, on the SYSU-MM01 dataset, our method achieves 64.8% Rank-1 and 60.2% mAP with significant gains over the latest related method. Demao Zhang, Ming Hong, Zheng Wang 0007, Zhizhong Zhang 0001, Xiaotong Luo, Yuan Xie 0006, Yanyun Qu |
ICME | 5 |
| 2022 | Cross-Domain and Cross-Modal Knowledge Distillation in Domain Adaptation for 3D Semantic SegmentationabstractWith the emergence of multi-modal datasets where LiDAR and camera are synchronized and calibrated, cross-modal Unsupervised Domain Adaptation (UDA) has attracted increasing attention because it reduces the laborious annotation of target domain samples. To alleviate the distribution gap between source and target domains, existing methods conduct feature alignment by using adversarial learning. However, it is well-known to be highly sensitive to hyperparameters and difficult to train. In this paper, we propose a novel model (Dual-Cross) that integrates Cross-Domain Knowledge Distillation (CDKD) and Cross-Modal Knowledge Distillation (CMKD) to mitigate domain shift. Specifically, we design the multi-modal style transfer to convert source image and point cloud to target style. With these synthetic samples as input, we introduce a target-aware teacher network to learn knowledge of the target domain. Then we present dual-cross knowledge distillation when the student is learning on source domain. CDKD constrains teacher and student predictions under same modality to be consistent. It can transfer target-aware knowledge from the teacher to the student, making the student more adaptive to the target domain. CMKD generates hybrid-modal prediction from the teacher predictions and constrains it to be consistent with both 2D and 3D student predictions. It promotes the information interaction between two modalities to make them complement each other. From the evaluation results on various domain adaptation settings, Dual-Cross significantly outperforms both uni-modal and cross-modal state-of-the-art methods. Miaoyu Li, Yachao Zhang 0001, Yuan Xie 0006, Zuodong Gao, Cuihua Li, Zhizhong Zhang 0001, Yanyun Qu |
ACM Multimedia | 6 |
| 2022 | Not All Pixels Are Matched: Dense Contrastive Learning for Cross-Modality Person Re-IdentificationabstractVisible-Infrared Person Re-Identification (VI-ReID) has become an emerging task for night-time surveillance systems. In order to reduce the cross-modality discrepancy, previous works either align the features via metric learning or generate synthesized cross-modality images by Generative Adversary Network. However, feature-level alignment ignores the heterogeneous data itself while generative framework suffers from the low generation quality, limiting their applications. In this paper, we propose a dense contrastive learning framework (DCLNet), which performs pixel-to-pixel dense alignment acting on the intermediate representations, rather than the final deep feature. It is a new loss function that brings views of positive pixels with same semantic information closer in shallow representation space, whilst pushing views of negative pixels apart. It naturally provides additional dense supervision and captures fine-grained pixel correspondence, reducing the modality gap from a new perspective. To implement it, a Part Aware Parsing (PAP) module and a Semantic Rectification Module (SRM) are introduced to learn and refine a semantic-guided mask, allowing us to efficiently find positive pairs only requiring instance-level supervision. Extensive experiments on the public SYSU-MM01 and RegDB datasets demonstrate the superiority of our pipeline over state-of-the-arts. Code is available at https://github.com/sunhz0117/DCLNet. Hanzhe Sun, Jun Liu 0116, Zhizhong Zhang 0001, Chengjie Wang 0001, Yanyun Qu, Yuan Xie 0006, Lizhuang Ma |
ACM Multimedia | 3 |
| 2022 | Hierarchical Walking Transformer for Object Re-IdentificationabstractRecently, transformer purely based on attention mechanism has been applied to a wide range of tasks and achieved impressive performance. Though extensive efforts have been made, there are still drawbacks to the transformer architecture which hinder its further applications: (i) the quadratic complexity brought by attention mechanism; (ii) barely incorporated inductive bias. Jun Liu 0116, Zhizhong Zhang 0001, Chengjie Wang 0001, Yanyun Qu, Yuan Xie 0006, Lizhuang Ma |
ACM Multimedia | 3 |
| 2022 | Self-supervised Exclusive Learning for 3D Segmentation with Cross-Modal Unsupervised Domain Adaptationabstract2D-3D unsupervised domain adaptation (UDA) tackles the lack of annotations in a new domain by capitalizing the relationship between 2D and 3D data. Existing methods achieve considerable improvements by performing cross-modality alignment in a modality-agnostic way, failing to exploit modality-specific characteristic for modeling complementarity. In this paper, we present self-supervised exclusive learning for cross-modal semantic segmentation under the UDA scenario, which avoids the prohibitive annotation. Specifically, two self-supervised tasks are designed, named "plane-to-spatial'' and "discrete-to-textured''. The former helps the 2D network branch improve the perception of spatial metrics, and the latter supplements structured texture information for the 3D network branch. In this way, modality-specific exclusive information can be effectively learned, and the complementarity of multi-modality is strengthened, resulting in a robust network to different domains. With the help of the self-supervised tasks supervision, we introduce a mixed domain to enhance the perception of the target domain by mixing the patches of the source and target domain samples. Besides, we propose a domain-category adversarial learning with category-wise discriminators by constructing the category prototypes for learning domain-invariant features. We evaluate our method on various multi-modality domain adaptation settings, where our results significantly outperform both uni-modality and multi-modality state-of-the-art competitors. Yachao Zhang 0001, Miaoyu Li, Yuan Xie 0006, Cuihua Li, Cong Wang 0039, Zhizhong Zhang 0001, Yanyun Qu |
ACM Multimedia | 6 |
| 2022 | Dual Mutual Learning for Cross-Modality Person Re-IdentificationabstractCross-modality person re-identification (Re-ID) is more challenging than traditional visible Re-ID due to the huge cross-modality gap from heterogeneous images. To alleviate this problem, existing methods often utilize a dual path learning framework equipped with metric loss to learn discriminative features. Despite effectiveness, the inevitable degeneration of intra-modality discrimination by taking cross-modality discrimination into consideration is unsolvable. Such degeneration substantially hinders the model’s capability of further improving feature representations. To mitigate this degeneration, we propose a Dual Mutual Learning (DML) method for cross-modality Re-ID which conducts mutual learning between the cross-modality and each of two single modalities. We design a triple-branch deep model containing the RGB and IR branches and the cross-modality branch. The cross-modality branch is designed to learn modality-invariant feature subspace for appearance similarity measurement. Both the RGB branch and IR branch provide attention supervision information to the cross-modality branch for attention feature alignment so as to enhance the intra-modality discrimination. Experimental results on two standard benchmarks demonstrate DML is superior to state-of-the-art methods. Demao Zhang, Zhizhong Zhang 0001, Ying Ju 0002, Cong Wang 0039, Yuan Xie 0006, Yanyun Qu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | One-Step Multiview Subspace Segmentation via Joint Skinny Tensor Learning and Latent ClusteringabstractMultiview subspace clustering (MSC) has attracted growing attention due to the extensive value in various applications, such as natural language processing, face recognition, and time-series analysis. In this article, we are devoted to address two crucial issues in MSC: 1) high computational cost and 2) cumbersome multistage clustering. Existing MSC approaches, including tensor singular value decomposition (t-SVD)-MSC that has achieved promising performance, generally utilize the dataset itself as the dictionary and regard representation learning and clustering process as two separate parts, thus leading to the high computational overhead and unsatisfactory clustering performance. To remedy these two issues, we propose a novel MSC model called joint skinny tensor learning and latent clustering (JSTC), which can learn high-order skinny tensor representations and corresponding latent clustering assignments simultaneously. Through such a joint optimization strategy, the multiview complementary information and latent clustering structure can be exploited thoroughly to improve the clustering performance. An alternating direction minimization algorithm, which owns low computational complexity and can be run in parallel when solving several key subproblems, is carefully designed to optimize the JSTC model. Such a nice property makes our JSTC an appealing solution for large-scale MSC problems. We conduct extensive experiments on ten popular datasets and compare our JSTC with 12 competitors. Five commonly used metrics, including four external measures (NMI, ACC, F-score, and RI) and one internal metric (SI), are adopted to evaluate the clustering quality. The experimental results with the Wilcoxon statistical test demonstrate the superiority of the proposed method in both clustering performance and operational efficiency. Yongqiang Tang, Yuan Xie 0006, Changqing Zhang 0002, Zhizhong Zhang 0001, Wensheng Zhang 0002 |
IEEE Trans. Cybern. | 4 |
| 2021 | Farewell to Mutual Information: Variational Distillation for Cross-Modal Person Re-IdentificationabstractThe Information Bottleneck (IB) provides an information theoretic principle for representation learning, by retaining all information relevant for predicting label while minimizing the redundancy. Though IB principle has been applied to a wide range of applications, its optimization remains a challenging problem which heavily relies on the accurate estimation of mutual information. In this paper, we present a new strategy, Variational Self-Distillation (VSD), which provides a scalable, flexible and analytic solution to essentially fitting the mutual information but without explicitly estimating it. Under rigorously theoretical guarantee, VSD enables the IB to grasp the intrinsic correlation between representation and label for supervised training. Further-more, by extending VSD to multi-view learning, we introduce two other strategies, Variational Cross-Distillation (VCD) and Variational Mutual-Learning (VML), which significantly improve the robustness of representation to view-changes by eliminating view-specific and task-irrelevant in-formation. To verify our theoretically grounded strategies, we apply our approaches to cross-modal person Re-ID, and conduct extensive experiments, where the superior performance against state-of-the-art methods are demonstrated. Our intriguing findings highlight the need to rethink the way to estimate mutual information. Zhizhong Zhang 0001, Shaohui Lin, Yanyun Qu, Yuan Xie 0006, Lizhuang Ma |
CVPR | 2 |
| 2021 | Contrastive Learning for Compact Single Image DehazingabstractSingle image dehazing is a challenging ill-posed problem due to the severe information degeneration. However, existing deep learning based dehazing methods only adopt clear images as positive samples to guide the training of dehazing network while negative information is unexploited. Moreover, most of them focus on strengthening the dehazing network with an increase of depth and width, leading to a significant requirement of computation and memory. In this paper, we propose a novel contrastive regularization (CR) built upon contrastive learning to exploit both the information of hazy images and clear images as negative and positive samples, respectively. CR ensures that the restored image is pulled to closer to the clear image and pushed to far away from the hazy image in the representation space.Furthermore, considering trade-off between performance and memory storage, we develop a compact dehazing network based on autoencoder-like (AE) framework. It involves an adaptive mixup operation and a dynamic feature enhancement module, which can benefit from preserving information flow adaptively and expanding the receptive field to improve the network’s transformation capability, respectively. We term our dehazing network with autoencoder and contrastive regularization as AECR-Net. The extensive experiments on synthetic and real-world datasets demonstrate that our AECR-Net surpass the state-of-the-art approaches. The code is released in https://github.com/GlassyWu/AECR-Net. Haiyan Wu, Yanyun Qu, Shaohui Lin, Ruizhi Qiao, Zhizhong Zhang 0001, Yuan Xie 0006, Lizhuang Ma |
CVPR | 6 |
| 2021 | Non-Adversarial Novelty Detection with Generative Latent Nearest NeighborsabstractNovelty detection is the task of identifying whether a new data point is considered to be an inlier or an outlier. Generative Adversarial Networks (GAN)-based methods suffer from mode dropping and unstable training issue, which poses the greatest threat to learn the target class distribution. To solve mode dropping issues, the nearest neighbor generator is designed to ensure that for every training image there exists a candidate generated image that is near to it at optimality. The generator considers the entire distribution of training data without mode dropping. To avoid the instability training issue, we consider capturing the distribution of the target class by non-adversarial strategy. In addition, to provide great image priors and fully diversity candidate samples for the generator, we also design a two-step mapping process. Finally, Experiments show that our model has clear superiority over cutting-edge novelty detectors and achieves state-of-the-art results on the datasets. Chengwei Chen, Zhizhong Zhang 0001, Yuan Xie 0006, Lizhuang Ma |
ICME | 2 |
| 2021 | Both Comparison and Induction are Indispensable for Cross-Domain Few-Shot LearningabstractFew-shot learning (FSL), aiming to extract new knowledge from very small amount of labeled samples, has attracted noticeable attentions recently. However, most of existing methods often fail when facing huge domain shift between seen and unseen classes. We think this should be attributed to the episode strategy which ignore utilizing support samples to induct the test classes. So in this paper, for the first time, we propose a bilevel episode strategy (BL-ES) to train a inductive graph network (IGN) that learn to both comparison and induction. Specifically, first, outer episodes in BL-ES simulate the cross-domain few-shot tasks constantly, while inner episodes learn to drive IGN to induct the common features of test classes. Then, the propsoed IGN captures the correlation among all samples to update meta points of each category in induction module. Finally, we introduce a geometrical constraint term utilizing meta points into the training loss, to update the nodes and edges in feature space. This way improves the robustness of training process. Extensive experiments show that our framework outperforms the state-of-the-art FSL alternatives, and are more suitable for real-world applications. Wang Yuan, TianXue Ma, Yuan Xie 0006, Zhizhong Zhang 0001, Lizhuang Ma |
ICME | 5 |
| 2021 | Learn from Concepts: Towards the Purified Memory for Few-shot LearningabstractHuman beings have a great generalization ability to recognize a novel category by only seeing a few number of samples. This is because humans possess the ability to learn from the concepts that already exist in our minds. However, many existing few-shot approaches fail in addressing such a fundamental problem, {\it i.e.,} how to utilize the knowledge learned in the past to improve the prediction for the new task. In this paper, we present a novel purified memory mechanism that simulates the recognition process of human beings. This new memory updating scheme enables the model to purify the information from semantic labels and progressively learn consistent, stable, and expressive concepts when episodes are trained one by one. On its basis, a Graph Augmentation Module (GAM) is introduced to aggregate these concepts and knowledge learned from new tasks via a graph neural network, making the prediction more accurate. Generally, our approach is model-agnostic and computing efficient with negligible memory cost. Extensive experiments performed on several benchmarks demonstrate the proposed method can consistently outperform a vast number of state-of-the-art few-shot learning methods. Xuncheng Liu, Shaohui Lin, Yanyun Qu, Lizhuang Ma, Wang Yuan, Zhizhong Zhang 0001, Yuan Xie 0006 |
IJCAI | 7 |
| 2021 | Towards Compact Single Image Super-Resolution via Contrastive Self-distillationabstractConvolutional neural networks (CNNs) are highly successful for super-resolution (SR) but often require sophisticated architectures with heavy memory cost and computational overhead significantly restricts their practical deployments on resource-limited devices. In this paper, we proposed a novel contrastive self-distillation (CSD) framework to simultaneously compress and accelerate various off-the-shelf SR models. In particular, a channel-splitting super-resolution network can first be constructed from a target teacher network as a compact student network. Then, we propose a novel contrastive loss to improve the quality of SR images and PSNR/SSIM via explicit knowledge transfer. Extensive experiments demonstrate that the proposed CSD scheme effectively compresses and accelerates several standard SR models such as EDSR, RCAN and CARN. Code is available at https://github.com/Booooooooooo/CSD. Yanbo Wang 0003, Shaohui Lin, Yanyun Qu, Haiyan Wu, Zhizhong Zhang 0001, Yuan Xie 0006, Angela Yao |
IJCAI | 5 |
| 2021 | Improving Domain-Adaptive Person Re-Identification by Dual-Alignment Learning With Camera-Aware Image GenerationabstractDomain adaptation in person re-identification (re-ID) has always been challenging, especially for the lack of supervision information on the target domain. Existing methods generally introduced extra supervision by adversarial learning techniques, then added all the augmented data in the training process to optimize the re-ID model. However, the direct utilization of all the generated data not only increases additional computational cost but also ignores the potential correlation between the origin and generated data. In this article, we propose a novel dual-alignment learning framework (DAL) with camera-aware image generation to efficiently and effectively tackle this issue. Specifically, we propose a camera transfer matching module to generate additional training images with different camera styles, and construct the matching pairs with each containing a origin image and one corresponding camera transferred image. To strengthen the correlation of images for each matching pair, we align the pseudo-labels via clustering algorithm to reduce the pseudo-labels distribution discrepancy between the origin and generated images. Besides, to avoid model degeneration affected by some inaccurate pseudo-labels on unlabelled data, we maximize the mutual information to align the image feature representations of matching pair. The DAL allows us to decrease the camera variance and enhance the discrimination ability of re-ID model. Extensive experiments on three large-scale benchmarks demonstrate the superiority of DAL over state-of-the-art methods. Chenyang Zhang 0003, Yongqiang Tang, Zhizhong Zhang 0001, Ding Li 0006, Xuebing Yang, Wensheng Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Multi-scale 2D Representation Learning for weakly-supervised moment retrieval
Ding Li 0006, Yongqiang Tang, Zhizhong Zhang 0001, Wensheng Zhang 0002 |
ICPR | 4 |
| 2020 | A Deep Nonnegative Matrix Factorization Approach via Autoencoder for Nonlinear Fault DetectionabstractIn the era of big data, data-driven fault detection is vital for modern industrial systems. This article considers the potential complexity of fault detection and proposes a novel nonlinear method based on nonnegative matrix factorization (NMF). Motivated by an autoencoder, in this article we first utilize the input data to learn an appropriate nonlinear mapping function, which transforms the original space into a high-dimensional feature space. Then, according to the decomposition rule of NMF, we divide the learned feature space into two subspaces, and two statistics in these subspaces are designed appropriately for nonlinear fault detection. The established method, i.e., deep nonnegative matrix factorization (DNMF), is implemented by three parts: an encoder module, an NMF module, and a decoder module. Unlike conventional NMF-based nonlinear methods using implicit and predetermined kernels, DNMF provides a new nonlinear scheme applied to NMF via a deep autoencoder framework and realizes nonlinear mapping for input data automatically. Our proposed nonlinear framework can be further generalized to other linear methods. Besides, DNMF greatly expands the NMF application scope by breaking through the limitation of nonnegative input. The Tennessee Eastman process as an industrial benchmark is employed to verify the effectiveness of the proposed method. Zelin Ren, Wensheng Zhang 0002, Zhizhong Zhang 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Learning to Align via Wasserstein for Person Re-IdentificationabstractExisting successful person re-identification (Re-ID) models often employ the part-level representation to extract the fine-grained information, but commonly use the loss that is particularly designed for global features, ignoring the relationship between semantic parts. In this paper, we present a novel triplet loss that emphasizes the salient parts and also takes the consideration of alignment. This loss is based on the crossing-bing matching metric that also known as Wasserstein Distance. It measures how much effort it would take to move the embeddings of local features to align two distributions, such that it is able to find an optimal transport matrix to re-weight the distance of different local parts. The distributions in support of local parts is produced via a new attention mechanism, which is calculated by the inner product between high-level global feature and local features, representing the importance of different semantic parts w.r.t. identification. We show that the obtained optimal transport matrix can not only distinguish the relevant and misleading parts, and hence assign different weights to them, but also rectify the original distance according to the learned distributions, resulting in an elegant solution for the mis-alignment issue. Besides, the proposed method is easily implemented in most Re-ID learning system with end-to-end training style, and can obviously improve their performance. Extensive experiments and comparisons with recent Re-ID methods manifest the competitive performance of our method. Zhizhong Zhang 0001, Yuan Xie 0006, Ding Li 0006, Wensheng Zhang 0002, Qi Tian 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Tensor Multi-Task Learning for Person Re-IdentificationabstractThis paper presents a tensor multi-task model for person re-identification (Re-ID). Due to discrepancy among cameras, our approach regards Re-ID from multiple cameras as different but related classification tasks, each task corresponding to a specific camera. In each task, we distinguish the person identity as a one-vs-all linear classification problem, where one classifier is associated with a specific person. By constructing all classifiers into a task-specific projection matrix, the proposed method could utilize all the matrices to form a tensor structure, and jointly train all the tasks in a uniform tensor space. In this space, by assuming the features of the same person under different cameras are generated from a latent subspace, and different identities under the same perspective share similar patterns, the high-order correlations, not only across different tasks but also within a certain task, can be captured by utilizing a new type of low-rank tensor constraint. Therefore, the learned classifiers transform the original feature vector into the latent space, where feature distributions across cameras can be well-aligned. Moreover, this model can be incorporated into multiple visual features to boost the performance, and easily extended to the unsupervised setting. Extensive experiments and comparisons with recent Re-ID methods manifest the competitive performance of our method. Zhizhong Zhang 0001, Yuan Xie 0006, Wensheng Zhang 0002, Yongqiang Tang, Qi Tian 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Effective Image Retrieval via Multilinear Multi-Index FusionabstractMulti-index fusion has demonstrated impressive performances in the retrieval task by integrating different visual representations in a unified framework. However, previous works mainly consider propagating similarities via a neighbor structure, ignoring the high-order information among different visual representations. In this paper, we propose a new multi-index fusion scheme for image retrieval. By formulating this procedure as a multilinear-based optimization problem, the complementary information hidden in different indexes can be explored more thoroughly. Specifically, we first build our multiple indexes from various visual representations. Then, a so-called index-specific functional matrix, which aims to propagate similarities, is introduced to update the original index. The functional matrices are then optimized in a unified tensor space to achieve a refinement, such that the relevant images can be pushed closer. The optimization problem can be efficiently solved by the augmented Lagrangian method with a theoretical convergence guarantee. Unlike the traditional multi-index fusion scheme, our approach embeds the multi-index subspace structure into the new indexes with sparse constraint and, thus, it has little additional memory consumption in the online query stage. Experimental evaluation on three benchmark datasets reveals that the proposed approach achieves state-of-the-art performance, that is, N-score 3.94 on UKBench, mAP 94.1% on Holiday, and 62.39% on Market-1501. Zhizhong Zhang 0001, Yuan Xie 0006, Wensheng Zhang 0002, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |