Xin Tan 0002

dblp:89/6413-2 · DBLP profile ↗
← Back
110ranked-venue papers
7as first author
98since 2021 · last 2026
0000-0001-9346-1196ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 76 · 4 first-author · 69 since 2021Artificial intelligence and machine learning · 61 · 2 first-author · 54 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
abstract
Human motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this paper, we propose a human motion synthesis framework that take an unified Scene Semantic Occupancy (SSO) for scene representation, termed SSOMotion. We design a bi-directional tri-plane decomposition to derive a compact version of the SSO, and scene semantics are mapped to an unified feature space via CLIP encoding and shared linear dimensionality reduction. Such strategy can derive the fine-grained scene semantic structures while significantly reduce redundant computations. We further take these scene hints and movement direction derived from instructions for motion control via frame-wise scene query. Extensive experiments and ablation studies conducted on cluttered scenes using ShapeNet furniture, as well as scanned scenes from PROX and Replica datasets, demonstrate its cutting-edge performance while validating its effectiveness and generalization ability.
Jingyu Gong, Kunkun Tong, Zhuoran Chen, Chuanhan Yuan, Mingang Chen, Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006
AAAI7
2026 Diffusion Implicit Policy for Unpaired Scene-aware Motion Synthesis
abstract
Scene-aware motion synthesis has been widely researched recently due to its numerous applications. Prevailing methods rely heavily on paired motion-scene data, while it is difficult to generalize to diverse scenes when trained only on a few specific ones. Thus, we propose a unified framework, termed Diffusion Implicit Policy (DIP), for scene-aware motion synthesis, where paired motion-scene data are no longer necessary. In this paper, we disentangle human-scene interaction from motion synthesis during training, and then introduce an interaction-based implicit policy into motion diffusion during inference. Synthesized motion can be derived through iterative diffusion denoising and implicit policy optimization, thus motion naturalness and interaction plausibility can be maintained simultaneously. For long-term motion synthesis, we introduce motion blending in joint rotation power space. The proposed method is evaluated on synthesized scenes with ShapeNet furniture, and real scenes from PROX and Replica. Results show that our framework presents better motion naturalness and interaction plausibility than cutting-edge methods. This also indicates the feasibility of utilizing the DIP for motion synthesis in more general tasks and versatile scenes.
Jingyu Gong, Fengqi Liu, Qianyu Zhou 0001, Xin Tan 0002, Zhizhong Zhang 0001, Yuan Xie 0006
AAAI6
2026 Multi-Step Deformable Gaussian Splatting for Dynamic Scene Rendering
abstract
Reconstructing dynamic scenes has long been a challenging task in 3D vision. Previous mainstream methods based on 3D Gaussian Splatting typically employ a single deformation field to directly model spatiotemporal changes. However, such one-step deformation struggles to capture diverse and complex motion patterns. To address this limitation, we propose decomposing the one-step deformation into a multi-step process, where each step is represented by a deformation layer. Additionally, we introduce a weight prediction mechanism for each layer to control the extent of deformation at every step. We provide two types of deformation layers based on implicit and explicit approaches. Moreover, while the deformation layer is time-conditioned, the Gaussians' behavior may still be influenced by their time-invariant properties. Therefore, we propose a fully time-agnostic scale modulation block to modulate the scaling changes of Gaussians. Extensive experiments on D-NeRF, Neu3D, and HyperNeRF demonstrate that our method achieves state-of-the-art performance.
Jiaheng Hu, Zhizhong Zhang 0001, Jingyu Gong, Lizhuang Ma, Xin Tan 0002, Yuan Xie 0006
AAAI5
2026 LidarPainter: One-Step Away from Any Lidar View to Novel Guidance
abstract
Dynamic driving scene reconstruction is of great importance in fields like digital twin system and autonomous driving simulation. However, unacceptable degradation occurs when the view deviates from the input trajectory, leading to corrupted background and vehicle models. To improve reconstruction quality on novel trajectory, existing methods are subject to various limitations including inconsistency, deformation, and time consumption. This paper proposes LidarPainter, a one-step diffusion model that recovers consistent driving views from sparse LiDAR condition and artifact-corrupted renderings in real-time, enabling high-fidelity lane shifts in driving scene reconstruction. Extensive experiments show that LidarPainter outperforms state-of-the-art methods in speed, quality and resource efficiency, specifically 7 × faster than StreetCrafter with only one fifth of GPU memory required. LidarPainter also supports stylized generation using text prompts such as “foggy” and “night”, allowing for a diverse expansion of the existing asset library.
Yuzhou Ji, Anchun Zhang, Lizhuang Ma, Xin Tan 0002
AAAI6
2026 Zero-Shot Robotic Manipulation via 3D Gaussian Splatting-Enhanced Multimodal Retrieval-Augmented Generation
abstract
Existing end-to-end approaches of robotic manipulation often lack generalization to unseen objects or tasks due to limited data and poor interpretability. While recent Multimodal Large Language Models (MLLMs) demonstrate strong commonsense reasoning, they struggle with geometric and spatial understanding required for pose prediction. In this paper, we propose RobMRAG, a 3D Gaussian Splatting-Enhanced Multimodal Retrieval-Augmented Generation (MRAG) framework for zero-shot robotic manipulation. Specifically, We construct a multi-source manipulation knowledge base containing object contact frames, task completion frames, and pose parameters. During inference, a Hierarchical Multimodal Retrieval module first employs hybrid semantic search to find task-relevant object prototypes, then selects the geometrically closest reference example based on pixel-level similarity and Instance Matching Distance (IMD). We further introduce a 3D-Aware Pose Refinement module based on 3D Gaussian Splatting into the MRAG framework, which aligns the pose of the reference object to the target object in 3D space. The aligned results are reprojected onto the image plane and used as input to the MLLM to enhance the generation of the final pose parameters. Extensive experiments show that on a test set containing 30 categories of household objects, our method improves the success rate by 7.76% compared to the best-performing zero-shot baseline under the same setting, and by 6.54% compared to the state-of-the-art supervised baseline. Our results validate that RobMRAG effectively bridges the gap between high-level semantic reasoning and low-level geometric execution, enabling robotic systems that generalize to unseen objects while remaining inherently interpretable.
Zilong Xie, Jingyu Gong, Xin Tan 0002, Zhizhong Zhang 0001, Yuan Xie 0006
AAAI3
2026 NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
abstract
Zhihao Luo, Wentao Yan, Jingyu Gong, Min Wang, Zhizhong Zhang, Xuhong Wang, Yuan Xie, Xin Tan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wentao Yan, Jingyu Gong, Min Wang 0024, Zhizhong Zhang 0001, Xuhong Wang, Yuan Xie 0006, Xin Tan 0002
ACL (1)8
2026 Dual-driven synergy of blockchain and federated learning for trustworthy medical data sharing in internet of medical things
Chenquan Gan, Xin Tan 0002, Qingyi Zhu, Akanksha Saini, Deepak Kumar Jain 0001, Abebe Abeshu Diro
J. Inf. Secur. Appl.2
2026 Dynamic expansion orthogonal network for class-incremental learning
Mingda Dong, Zhizhong Zhang 0001, Xin Tan 0002, Jiling Qiu, Yuan Xie 0006
Knowl. Based Syst.3
2026 From sparse semantics to rich instances: Empowering label-efficient LiDAR panoptic segmentation via geometric priors
Wei Zhang 0217, Zhizhong Zhang 0001, Xin Tan 0002, Lizhuang Ma, Yuan Xie 0006
Neural Networks4
2026 F3-SD: Focal feature fusion with self-distillation on large vision-language models for cross-modal retrieval
Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Xiaocui Li 0001, Maobin Tang, Meie Fang, Wensheng Zhang 0002
Pattern Recognit.4
2026 Cross-domain distillation for unsupervised domain adaptation with large vision-language models
Xingwei Deng, Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Maobin Tang, Meie Fang, Wensheng Zhang 0002
Pattern Recognit.4
2026 Unlocking the candidates: Beam-aware reasoning for audio-visual speech recognition
Shao Zeng, Tianjun Gu, Shouhong Ding, Xin Tan 0002, Yang Gao 0001
Pattern Recognit.10
2026 From static to adaptive multi-view: Nuanced expert prompt tuning for Fine-Grained Image Retrieval
Ke-Yue Zhang, Jingyu Gong, Yang Gao 0001, Xin Tan 0002, Lizhuang Ma
Pattern Recognit.6
2026 DANIM: Domain adaptation network with intermediate domain masking for night-time scene parsing
Qijian Tian, Ran Yi 0002, Zufeng Zhang, Bin Sheng 0001, Xin Tan 0002, Lizhuang Ma
Pattern Recognit.6
2026 MMoFusion: Multi-modal co-speech motion generation with diffusion model
Jiangning Zhang, Xin Tan 0002, Chengjie Wang 0001, Lizhuang Ma
Pattern Recognit.3
2026 SHTOcc: Effective 3D occupancy prediction with sparse head and tail voxels
Qiucheng Yu, Yuan Xie 0006, Xin Tan 0002
Pattern Recognit.3
2026 A Task-Aware Parameter Decoupling Framework for Continual Anomaly Detection
abstract
Real-world industrial scenarios have become increasingly dynamic, with new product types, defect patterns, and operational modes emerging rapidly. In such a context, the one-for-more paradigm enables the use of a single model to economically and continually adapt to evolving distributions or patterns, positioning it as a key component in modern Industrial AI systems. This article proposes a novel one-for-more anomaly detection framework designed to identify anomalies across expanding product lines. The framework incorporates two model-agnostic techniques: instance-aware prompt tuning (IPT) and gradient-aware parameter decoupling (GPD). Our approach is built upon a reconstruction-based vision transformer (ViT) encoder–decoder architecture. IPT addresses the domain gap between pretrained models and industrial data by leveraging an instance-level prompt and a shared memory mechanism, which helps the pretrained model retain previously learned patterns. GPD selectively updates network parameters based on the gradient’s impact on prior tasks, employing orthogonal gradient projection to further minimize interference. In addition, we introduce a new dataset to simulate the one-for-more industrial scenario. Extensive experiments on MVTec and our proposed dataset demonstrate that our framework achieves the state-of-the-art performance across various continual learning settings, significantly outperforming existing methods, particularly in multistep incremental scenarios.
Zhizhong Zhang 0001, Guchu Zou, Chengwei Chen, Zhenyi Qi, Jingwen Qi, Yongke Yao, Xiaofan Li 0008, Yuan Xie 0006, Xin Tan 0002
IEEE Trans. Ind. Informatics10
2026 Transporting the Cross-Modal Prototypes for Unsupervised Visible-Infrared Person Re-Identification
abstract
Unsupervised visible infrared person re-identification (USVI-ReID) is a challenging retrieval task that retrieves cross-modality pedestrian images without using any label information. In this task, the large cross-modality variance makes it difficult to generate reliable cross-modality labels, and the lack of annotations also provides additional difficulties for learning modality-invariant features. To facilitate this unsupervised cross-modal learning, we begin by leveraging the information contained in the cross-modality input and its predicted label. Aiming to minimize information loss, we optimize the model by incorporating entropy minimization, uniform label distribution, and cross-modality matching. In our approach, we design a loop iterative training strategy alternating between model training and cross-modality matching, where a uniform prior guided optimal transport assignment is proposed to select matched visible and infrared prototypes. This matching information is then utilized to minimize the intra- and cross-modality entropy. As a result, our model can gradually self-learn useful information, enabling it to generate discriminative representations for unlabeled cross-modal data. Extensive experimental results on benchmarks demonstrate the effectiveness of our method, e.g., 69.4% and 89.4% of Rank-1 accuracy on SYSU-MM01 and RegDB without any annotations. The code will be released soon.
Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006
IEEE Trans. Image Process.3
2026 Decoupling 3-D Point Cloud Attributes for Semantic Segmentation via Real-World Prior Exploitation
abstract
Point cloud semantic segmentation, which involves assigning a category for each point, is a crucial task in autonomous driving and intelligent transportation systems. Due to the inherently unordered and irregular nature of point clouds, learning robust features that accurately capture real-world distributions from point coordinates and other attributes remains challenging. Following the pioneering work of PointNet, current 3D deep neural networks process point coordinates alongside other attributes without fully exploiting the implicit class prior information embedded in spatial information. In this work, we first conduct a pilot study to evaluate how current 3D networks utilize point coordinates and validate the presence of implicit class priors within them. Subsequently, we design a robust Position-to-Physics (P2P) fusion strategy that learns adaptive weights to dynamically incorporate implicit class priors present in point coordinates into point features. Moreover, we design a dual-branch network architecture and propose a triplet loss to further enhance the adaptive fusion process. Extensive experiments demonstrate that decoupling position attributes from physics attributes facilitates the extraction and utilization of implicit class priors. Our proposed modules consistently improve segmentation performance across various networks and datasets, demonstrating their generalizability and effectiveness.
Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Lizhuang Ma, Yuan Xie 0006
IEEE Trans. Intell. Transp. Syst.3
2025 FastLGS: Speeding Up Language Embedded Gaussians with Feature Grid Mapping
abstract
The semantically interactive radiance field has always been an appealing task for its potential to facilitate user-friendly and automated real-world 3D scene understanding applications. However, it is a challenging task to achieve high quality, efficiency and zero-shot ability at the same time with semantics in radiance fields. In this work, we present FastLGS, an approach that supports real-time open-vocabulary query within 3D Gaussian Splatting (3DGS) under high resolution. We propose the semantic feature grid to save multi-view CLIP features which are extracted based on Segment Anything Model (SAM) masks, and map the grids to low dimensional features for semantic field training through 3DGS. Once trained, we can restore pixel-aligned CLIP embeddings through feature grids from rendered features for open-vocabulary queries. Comparisons with other state-of-the-art methods prove that FastLGS can achieve the first place performance concerning both speed and accuracy, where FastLGS is 98 times faster than LERF, 4 times faster than LangSplat and 2.5 times faster than LEGaussians. Meanwhile, experiments show that FastLGS is adaptive and compatible with many downstream tasks, such as 3D segmentation and 3D object inpainting, which can be easily applied to other 3D manipulation systems.
Yuzhou Ji, Junshu Tang, Wuyi Liu, Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006
AAAI6
2025 DrivingForward: Feed-forward 3D Gaussian Splatting for Driving Scene Reconstruction from Flexible Surround-view Input
abstract
We propose DrivingForward, a feed-forward Gaussian Splatting model that reconstructs driving scenes from flexible surround-view input. Driving scene images from vehicle-mounted cameras are typically sparse, with limited overlap, and the movement of the vehicle further complicates the acquisition of camera extrinsics. To tackle these challenges and achieve real-time reconstruction, we jointly train a pose network, a depth network, and a Gaussian network to predict the Gaussian primitives that represent the driving scenes. The pose network and depth network determine the position of the Gaussian primitives in a self-supervised manner, without using depth ground truth and camera extrinsics during training. The Gaussian network independently predicts primitive parameters from each input image, including covariance, opacity, and spherical harmonics coefficients. At the inference stage, our model can achieve feed-forward reconstruction from flexible multi-frame surround-view input. Experiments on the nuScenes dataset show that our model outperforms existing state-of-the-art feed-forward and scene-optimized reconstruction methods in terms of reconstruction.
Qijian Tian, Xin Tan 0002, Yuan Xie 0006, Lizhuang Ma
AAAI2
2025 DepthFisheye: Efficient Fine-Tuning of Depth Estimation Models for Fisheye Cameras
Zhiwei Zhang 0005, Xin Tan 0002, Zhizhong Zhang 0001, Lizhuang Ma
CVM (3)3
2025 TAD: A Plug-and-Play Task Arithmetic Approach for Augmenting Diffusion Models
Qingyi Zhu, Ruochen Jin, Zhiwei Zhang 0005, Yishen Xue, Xin Tan 0002, Lizhuang Ma
CVM (2)5
2025 One-for-More: Continual Diffusion Model for Anomaly Detection
abstract
With the rise of generative models, there is a growing interest in unifying all tasks within a generative framework. Anomaly detection methods also fall into this scope and utilize diffusion models to generate or reconstruct normal samples when given arbitrary anomaly images. However, our study found that the diffusion model suffers from severe "faithfulness hallucination" and "catastrophic forgetting", which can’t meet the unpredictable pattern increments. To mitigate the above problems, we propose a continual diffusion model that uses gradient projection to achieve stable continual learning. Gradient projection deploys a regularization on the model updating by modifying the gradient towards the direction protecting the learned knowledge. But as a double-edged sword, it also requires huge memory costs brought by the Markov process. Hence, we propose an iterative singular value decomposition method based on the transitive property of linear representation, which consumes tiny memory and incurs almost no performance loss. Finally, considering the risk of "over-fitting" to normal images of the diffusion model, we propose an anomaly-masked network to enhance the condition mechanism of the diffusion model. For continual anomaly detection, ours achieves first place in 17/18 settings on MVTec and VisA. Code is available at https://github.com/FuNz-0/One-for-More
Xiaofan Li 0008, Xin Tan 0002, Zhizhong Zhang 0001, Rizen Guo, Guannan Jiang, Yanyun Qu, Lizhuang Ma, Yuan Xie 0006
CVPR2
2025 MOS: Modeling Object-Scene Associations in Generalized Category Discovery
abstract
Generalized Category Discovery (GCD) is a classification task that aims to classify both base and novel classes in un-labeled images, using knowledge from a labeled dataset. In GCD, previous research overlooks scene information or treats it as noise, reducing its impact during model training. However, in this paper, we argue that scene information should be viewed as a strong prior for inferring novel classes. We attribute the misinterpretation of scene information to a key factor: the Ambiguity Challenge inherent in GCD. Specifically, novel objects in base scenes might be wrongly classified into base categories, while base objects in novel scenes might be mistakenly recognized as novel categories. Once the ambiguity challenge is addressed, scene information can reach its full potential, significantly enhancing the performance of GCD models. To more effectively leverage scene information, we propose the Modeling Object-Scene Associations (MOS) framework, which utilizes a simple MLP-based scene-awareness module to enhance GCD performance. It achieves an exceptional average accuracy improvement of 4% on the challenging fine-grained datasets compared to state-of-the-art methods, emphasizing its superior performance in fine-grained GCD. The code is publicly available at https://github.com/JethroPeng/MOS.
Zhengyuan Peng, Jinpeng Ma, Zhimin Sun, Ran Yi 0002, Xin Tan 0002, Lizhuang Ma
CVPR6
2025 Efficient Prototypical Classifier for Class-Incremental Learning
abstract
The nearest prototypical classifier faces challenges of semantic drift and prototype interference. Previous methods address these issues using data rehearsal and contrastive learning, but these approaches incur high memory costs and slow convergence. In this paper, we propose a novel prototypical minimum distance loss, along with a two-stage training pipeline, to mitigate prototype interference with low memory overhead and fast convergence. Leveraging task-specific prompts and a key-query mechanism, we significantly reduce semantic drift. Additionally, we introduce a continual exponential moving average to enhance model stability and minimize forgetting. Notably, our method is rehearsal-free and avoids generation processes, simplifying training and further reducing memory usage. We validate our approach on four challenging class-incremental learning datasets, achieving significant improvements over state-of-the-art methods.
Wei Zhang 0217, Jingyang Qiao, Yuan Xie 0006, Zhizhong Zhang 0001, Xin Tan 0002
ICASSP5
2025 Prototype Alignment with LoRA Fusion for Class-Incremental Learning
abstract
Recent advancements in pre-trained models have enhanced performance on downstream tasks due to their strong generalizability. Despite this, models fine-tuned continually often face challenges such as catastrophic forgetting and loss of generalization. To address these issues, we propose a novel approach that utilizes distinct Low-Rank Adaptation (LoRA) modules for each task. These modules parameter-efficient, and integrated across tasks to ensure the model maintains strong performance on both old and new classes. Additionally, we investigate semantic relationships between class prototypes to effectively reconstruct old prototypes in the context of new tasks. Our experiments demonstrate that this method significantly outperforms baseline approaches across various class-incremental learning benchmarks, offering an efficient and effective solution for mitigating forgetting and preserving model performance.
Wei Zhang 0217, Yuan Xie 0006, Zhizhong Zhang 0001, Xin Tan 0002
ICASSP4
2025 Stylized-Face: A Million-Level Stylized Face Dataset for Face Recognition
Zhengyuan Peng, Jianqing Xu, Yuge Huang, Jinkun Hao, Shouhong Ding, Zhizhong Zhang 0001, Xin Tan 0002, Lizhuang Ma
ICCV7
2025 From Enhancement to Understanding: Build a Generalized Bridge for Low-Light Vision via Semantically Consistent Unsupervised Fine-Tuning
Shao Zeng, Tianjun Gu, Zhizhong Zhang 0001, Shouhong Ding, Jun Wang 0006, Xin Tan 0002, Yuan Xie 0006, Lizhuang Ma
ICCV9
2025 LFNet: Cross-Modal LiDAR-Fisheye Fusion Network for 3D Semantic Segmentation
abstract
Cross-modal fusion, which leverages images to enhance 3D semantic segmentation, has demonstrated significant effectiveness due to the complementary nature of heterogeneous data. However, existing approaches are limited to pinhole images, leaving fisheye images largely unexplored. In this paper, we introduce the LiDAR-Fisheye Fusion Network (LFNet), a dual-transformer architecture designed for cross-modal fusion (CMF) across hierarchical multi-scale layers. The 3D Transformer extracts point-level features from LiDAR data, while the pre-trained 2D Transformer extracts patch-level features from fisheye images.The CMF module comprises two key components: Local Fusion (LoF) and Global Fusion (GoF). The LoF module interpolates patch-level features to pixel-level for accurate feature alignment and computes precise point-to-pixel mappings for gated fusion. Meanwhile, the GoF module enables points to capture a holistic understanding of the scene via a cross-modal attention mechanism. Experimental results highlight the potential of fisheye images as a promising modality to complement LiDAR data in 3D semantic segmentation. The code will be available at https://github.com/wjzhang642/LFNet.
Zhiwei Zhang 0005, Tianfang Sun, Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006
ICME5
2025 Large Continual Instruction Assistant
abstract
Continual Instruction Tuning (CIT) is adopted to continually instruct Large Models to follow human intent data by data. It is observed that existing gradient update would heavily destroy the performance on previous datasets during CIT process. Instead, Exponential Moving Average (EMA), owns the ability to trace previous parameters, which can aid in decreasing forgetting. Nonetheless, its stable balance weight fails to deal with the ever-changing datasets, leading to the out-of-balance between plasticity and stability. In this paper, we propose a general continual instruction tuning framework to address the challenge. Starting from the trade-off prerequisite and EMA update, we propose the plasticity and stability ideal condition. Based on Taylor expansion in the loss function, we find the optimal balance weight can be automatically determined by the gradients and learned parameters. Therefore, we propose a stable-plasticity balanced coefficient to avoid knowledge interference. Based on the semantic similarity of the instructions, we can determine whether to retrain or expand the training parameters and allocate the most suitable parameters for the testing instances. Extensive experiments across multiple continual instruction tuning benchmarks demonstrate that our approach not only enhances anti-forgetting capabilities but also significantly improves overall continual tuning performance. Our code is available at https://github.com/JingyangQiao/CoIN.
Jingyang Qiao, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Shouhong Ding, Yuan Xie 0006
ICML3
2025 EyeSeg: An Uncertainty-Aware Eye Segmentation Framework for AR/VR
abstract
Human-machine interaction through augmented reality (AR) and virtual reality (VR) is increasingly prevalent, requiring accurate and efficient gaze estimation which hinges on the accuracy of eye segmentation to enable smooth user experiences. We introduce EyeSeg, a novel eye segmentation framework designed to overcome key challenges that existing approaches struggle with: motion blur, eyelid occlusion, and train-test domain gaps. In these situations, existing models struggle to extract robust features, leading to suboptimal performance. Noting that these challenges can be generally quantified by uncertainty, we design EyeSeg as an uncertainty-aware eye segmentation framework for AR/VR wherein we explicitly model the uncertainties by performing Bayesian uncertainty learning of a posterior under the closed set prior. Theoretically, we prove that a statistic of the learned posterior indicates segmentation uncertainty levels and empirically outperforms existing methods in downstream tasks, such as gaze estimation. EyeSeg outputs an uncertainty score and the segmentation result, weighting and fusing multiple gaze estimates for robustness, which proves to be effective especially under motion blur, eyelid occlusion and cross-domain challenges. Moreover, empirical results suggest that EyeSeg achieves segmentation improvements of MIoU, E1, F1, and ACC surpassing previous approaches.
Zhengyuan Peng, Jianqing Xu, Shen Li 0004, Jiazhen Ji, Yuge Huang, Jinmin Li, Shouhong Ding, Rizen Guo, Xin Tan 0002, Lizhuang Ma
IJCAI10
2025 PFDepth: Heterogeneous Pinhole-Fisheye Joint Depth Estimation via Distortion-aware Gaussian-Splatted Volumetric Fusion
abstract
In this paper, we present the first pinhole-fisheye framework for heterogeneous multi-view depth estimation, PFDepth. Our key insight is to exploit the complementary characteristics of pinhole and fisheye imagery (undistorted vs. distorted, small vs. large FOV, far vs. near field) for joint optimization. PFDepth employs a unified architecture capable of processing arbitrary combinations of pinhole and fisheye cameras with varied intrinsics and extrinsics. Within PFDepth, we first explicitly lift 2D features from each heterogeneous view into a canonical 3D volumetric space. Then, a core module termed Heterogeneous Spatial Fusion is designed to process and fuse distortion-aware volumetric features across overlapping and non-overlapping regions. Additionally, we subtly reformulate the conventional voxel fusion into a novel 3D Gaussian representation, in which learnable latent Gaussian spheres dynamically adapt to local image textures for finer 3D aggregation. Finally, fused volume features are rendered into multi-view depth maps. Through extensive experiments, we demonstrate that PFDepth sets a state-of-the-art performance on KITTI-360 and RealHet datasets over current mainstream depth networks. To the best of our knowledge, this is the first systematic study of heterogeneous pinhole-fisheye depth estimation, offering both technical novelty and valuable empirical insights.
Zhiwei Zhang 0005, Ruikai Xu, Zhizhong Zhang 0001, Xin Tan 0002, Jingyu Gong, Yuan Xie 0006, Lizhuang Ma
ACM Multimedia5
2025 Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models
abstract
Large vision-language models (LVLMs) are vulnerable to harmful input compared to their language-only backbones. We investigated this vulnerability by exploring LVLMs internal dynamics, framing their inherent safety understanding in terms of three key capabilities. Specifically, we define these capabilities as safety perception, semantic understanding, and alignment for linguistic expression, and experimentally pinpointed their primary locations within the model architecture. The results indicate that safety perception often emerges before comprehensive semantic understanding, leading to the reduction in safety. Motivated by these findings, we propose Self-Aware Safety Augmentation (SASA), a technique that projects informative semantic representations from intermediate layers onto earlier safety-oriented layers. This approach leverages the model's inherent semantic understanding to enhance safety recognition without fine-tuning. Then, we employ linear probing to articulate the model's internal semantic comprehension to detect the risk before the generation process. Extensive experiments on various datasets and tasks demonstrate that SASA significantly improves the safety of LVLMs, with minimal impact on the utility.
Wanying Wang, Zeyu Ma 0003, Xin Tan 0002, Mingang Chen
ACM Multimedia4
2025 Domain-Incremental Learning Paradigm for scene understanding via Pseudo-Replay Generation
abstract
Scene understanding is a computer vision task that involves grasping the pixel-level distribution of objects. Unlike most research focuses on single-scene models, we consider a more versatile proposal: domain-incremental learning for scene understanding. This allows us to adapt well-studied single-scene models into multi-scene models, reducing data requirements and ensuring model flexibility. However, domain-incremental learning that leverages correlations between scene domains has yet to be explored. To address this challenge, we propose a Domain-Incremental Learning Paradigm (D-ILP) for scene understanding, along with a new strategy of Pseudo-Replay Generation (PRG) that does not require manual labeling. Specifically, D-ILP leverages pre-trained single-scene models and incremental images for supervised training to acquire new knowledge from other scenes. As a pre-trained generation model, PRG can controllably generate pseudo-replays resembling source images from incremental images and text prompts. These pseudo-replays are utilized to minimize catastrophic forgetting in the original scene. We perform experiments with three publicly accessible models: Mask2Former, Segformer, and DeepLabv3+. With successfully transforming these single-scene models into multi-scene models, we achieve high-quality parsing results for original and new scenes simultaneously. Meanwhile, the validity and rationality of our method are proved by the analysis of D-ILP.
Qile He, Mengtian Li 0002, Xin Tan 0002
Graph. Model.5
2025 Point Mask Transformer for Outdoor Point Cloud Semantic Segmentation
abstract
Current outdoor point-cloud segmentation methods typically formulate semantic segmentation as a per-point/voxel-classification task. Although this strategy is straightforward because it classifies each point directly, it ignores the overall relationship of the category. As an alternative paradigm, mask classification decouples category classification from region localization, allowing the model to better capture overall category relationships. In this paper, we propose a novel approach called the point mask transformer (PMFormer), which transforms the semantic segmentation of point clouds from per-point classification to mask classification using a transformer architecture. The proposed model comprises a 3D backbone, transformer decoder, and segmentation head that predicts a series of binary masks, each associated with a global class label. Furthermore, to accommodate the unique characteristics of large and sparse outdoor point-cloud scenes, we propose three enhancements for the integration of point-cloud data with the transformer: MaskMix, 3D position encoding, and attention weights. We evaluate our model using the SemanticKITTI and nuScenes datasets. Our experimental results show that the proposed method outperforms state-of-the-art semantic segmentation approaches.
Xin Tan 0002, Zhizhong Zhang 0001, Yuan Xie 0006, Lizhuang Ma
Comput. Vis. Media2
2025 Optimal Transport with Arbitrary Prior for Dynamic Resolution Network
Zhizhong Zhang 0001, Chenyang Zhang 0003, Lizhuang Ma, Xin Tan 0002, Yuan Xie 0006
Int. J. Comput. Vis.5
2025 Gradient Projection for Continual Parameter-Efficient Tuning
abstract
Parameter-efficient tunings (PETs) have demonstrated impressive performance and promising perspectives in training large models, while they are still confronted with a common problem: the trade-off between learning new content and protecting old knowledge, leading to zero-shot generalization collapse, and cross-modal hallucination. In this paper, we reformulate Adapter, LoRA, Prefix-tuning, and Prompt-tuning from the perspective of gradient projection, and first propose a unified framework called Parameter Efficient Gradient Projection (PEGP). We introduce orthogonal gradient projection into different PET paradigms and theoretically demonstrate that the orthogonal condition for the gradient can effectively resist forgetting even for large-scale models. It therefore modifies the gradient towards the direction that has less impact on the old feature space, with less extra memory space and training time. We extensively evaluate our method with different backbones, including ViT and CLIP, on diverse datasets, and experiments comprehensively demonstrate its efficiency in reducing forgetting in class, online class, domain, task, and multi-modality continual settings.
Jingyang Qiao, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Wensheng Zhang 0002, Zhi Han, Yuan Xie 0006
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 GEOcc: Geometrically Enhanced 3D Occupancy Network With Implicit-Explicit Depth Fusion and Contextual Self-Supervision
abstract
3D occupancy perception holds a pivotal role in recent vision-centric autonomous driving systems by converting surround-view images into integrated geometric and semantic representations within dense 3D grids. Nevertheless, current models still encounter two main challenges: modeling depth accurately in the 2D-3D view transformation stage, and overcoming the lack of generalizability issues due to sparse LiDAR supervision. To address these issues, this paper presents GEOcc, a Geometric-Enhanced Occupancy network tailored for vision-only surround-view perception. Our approach is three-fold: 1) Integration of explicit lift-based depth prediction and implicit projection-based transformers for depth modeling, enhancing the density and robustness of view transformation. 2) Utilization of mask-based encoder-decoder architecture for fine-grained semantic predictions; 3) Adoption of context-aware self-training loss functions in the pertaining stage to complement LiDAR supervision, involving the re-rendering of 2D depth maps from 3D occupancy features and leveraging image reconstruction loss to obtain denser depth supervision besides sparse LiDAR ground-truths. Our approach achieves State-of-the-Art performance on the Occ3D-nuScenes dataset with the least image resolution needed and the most weightless image backbone compared with current models, marking an improvement of 3.3% due to our proposed contributions. Comprehensive experimentation also demonstrates the consistent superiority of our method over baselines and alternative approaches. Our code is available athttps://github.com/world-executed/GEOcc.git
Xin Tan 0002, Zhiwei Zhang 0005, Chaojie Fan, Yong Peng 0002, Zhizhong Zhang 0001, Yuan Xie 0006, Lizhuang Ma
IEEE Trans. Intell. Transp. Syst.1
2025 WV-LUT: Wide Vision Lookup Tables for Real-Time Low-Light Image Enhancement
abstract
In recent years, the lookup tables (LUTs) with deep learning for image enhancement have achieved remarkable results with extremely high inference efficiency. However, when dealing with severely degraded low-light images, lookup-table-based methods tend to exhibit poor enhancement results due to the lack of contextual and global information. To address the limitations of current lookup-table-based methods in the low-light image enhancement task, we propose the novel Wide Vision Lookup Tables (WV-LUT) by introducing Complementary-Hierarchical 4D-LUTs into 3D-LUT, which allows 3D-LUT to have a wider range of vision. Specifically, the 4D-LUTs are used to expand the receptive field and process local information on a single channel, while a 3D-LUT is used for sRGB channel post-processing. Additionally, we propose a lightweight Global Adjustment Module that further enhances the performance and generalization of WV-LUT by obtaining global adjustment parameters for gamma and color correction matrix to adaptively process images. Experimental results demonstrate that our method outperforms other state-of-the-art methods in low-light image enhancement with the highest average ranking and superior inference efficiency. Furthermore, deployment experiments on mobile devices demonstrate that our WV-LUT achieves superior results and inference efficiency, showcasing promising application prospects for edge devices.
Canlin Li, Haowen Su, Xin Tan 0002, Xiangfei Zhang, Lizhuang Ma
IEEE Trans. Multim.3
2025 Bias to Balance: New-Knowledge-Preferred Few-Shot Class-Incremental Learning via Transition Calibration
abstract
Humans can quickly learn new concepts with limited experience, while not forgetting learned knowledge. Such ability in machine learning is referred to as few-shot class-incremental learning (FSCIL). Although some methods try to solve this problem by putting similar efforts to prevent forgetting and promote learning, we find existing techniques do not give enough importance to the new category as new training samples are rather rare. In this article, we propose a new biased-to-unbiased rectification method, which introduces a trainable transition matrix to mitigate the prediction discrepancy between the old classes and the new classes. This transition matrix is to be diagonally dominated, normalized, and differentiable with new-knowledge-preferred prior, to solving the strong bias between heavy old knowledge and limited new knowledge. Hence, we can achieve a balanced solution between learning new concepts and preventing catastrophic forgetting by giving new classes more chances. Extensive experiments on miniImagenet, CIFAR100, and CUB200 demonstrate that our method outperforms the latest state-of-the-art methods by 1.1%, 1.44%, and 2.08%, respectively.
Hongquan Zhang, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Yuan Xie 0006
IEEE Trans. Neural Networks Learn. Syst.3
2025 AttentionPainter: An Efficient and Adaptive Stroke Predictor for Scene Painting
abstract
Stroke-based Rendering (SBR) aims to decompose an input image into a sequence of parameterized strokes, which can be rendered into a painting that resembles the input image. Recently, Neural Painting methods that utilize deep learning and reinforcement learning models to predict the stroke sequences have been developed, but suffer from longer inference time or unstable training. To address these issues, we propose AttentionPainter, an efficient and adaptive model for single-step neural painting. First, we propose a novel scalable stroke predictor, which predicts a large number of stroke parameters within a single forward process, instead of the iterative prediction of previous Reinforcement Learning or auto-regressive methods, which makes AttentionPainter faster than previous neural painting methods. To further increase the training efficiency, we propose a Fast Stroke Stacking algorithm, which brings 13 times acceleration for training. Moreover, we propose Stroke-density Loss, which encourages the model to use small strokes for detailed information, to help improve the reconstruction quality. Finally, we design a Stroke Diffusion Model as an application of AttentionPainter, which conducts the denoising process in the stroke parameter space and facilitates stroke-based inpainting and editing applications helpful for human artists' design. Extensive experiments show that AttentionPainter outperforms the state-of-the-art neural painting methods.
Yizhe Tang, Yue Wang 0020, Ran Yi 0002, Xin Tan 0002, Lizhuang Ma, Yukun Lai, Paul L. Rosin
IEEE Trans. Vis. Comput. Graph.5
2025 D2U-Net: a dual-path hybrid UNet architecture for precise medical image segmentation
Noor Ahmed 0002, Xin Tan 0002, Lizhuang Ma
Vis. Comput.2
2025 Learnable scene prior for point cloud semantic segmentation
Yuanhao Chai, Jingyu Gong, Xin Tan 0002, Yuan Xie 0006, Lizhuang Ma
Vis. Comput.3
2025 Learning Pulse Image with Deep Dynamic Frequency Network for Cardiovascular Diseases Diagnosis
Ji Cui, Litai Pang, Shiju Zhao, Zhengyuan Peng, Lingzhi Zeng, Tao Jiang 0032, Mengchen Liang, Jinlian Huang, Wang Yuan, Xin Tan 0002, Lizhuang Ma, Jiatuo Xu
Vis. Comput.11
2025 Innovative collaborative multi-lookup table for real-time enhancement of low-light images
Canlin Li, Haowen Su, Xin Tan 0002, Lihua Bi, Xiangfei Zhang, Lizhuang Ma
Vis. Comput.3
2025 OSH-Splat: optimizable semantic hyperplanes for enhanced 3D language feature Gaussian splatting
Yuzhou Ji, Xin Tan 0002, Lizhuang Ma
Vis. Comput.3
2024 Beyond the Label Itself: Latent Labels Enhance Semi-supervised Point Cloud Panoptic Segmentation
abstract
As the exorbitant expense of labeling autopilot datasets and the growing trend of utilizing unlabeled data, semi-supervised segmentation on point clouds becomes increasingly imperative. Intuitively, finding out more ``unspoken words'' (i.e., latent instance information) beyond the label itself should be helpful to improve performance. In this paper, we discover two types of latent labels behind the displayed label embedded in LiDAR and image data. First, in the LiDAR Branch, we propose a novel augmentation, Cylinder-Mix, which is able to augment more yet reliable samples for training. Second, in the Image Branch, we propose the Instance Position-scale Learning (IPSL) Module to learn and fuse the information of instance position and scale, which is from a 2D pre-trained detector and a type of latent label obtained from 3D to 2D projection. Finally, the two latent labels are embedded into the multi-modal panoptic segmentation network. The ablation of the IPSL module demonstrates its robust adaptability, and the experiments evaluated on SemanticKITTI and nuScenes demonstrate that our model outperforms the state-of-the-art method, LaserMix.
Yujun Chen, Xin Tan 0002, Zhizhong Zhang 0001, Yanyun Qu, Yuan Xie 0006
AAAI2
2024 Domain-Hallucinated Updating for Multi-Domain Face Anti-spoofing
abstract
Multi-Domain Face Anti-Spoofing (MD-FAS) is a practical setting that aims to update models on new domains using only novel data while ensuring that the knowledge acquired from previous domains is not forgotten. Prior methods utilize the responses from models to represent the previous domain knowledge or map the different domains into separated feature spaces to prevent forgetting. However, due to domain gaps, the responses of new data are not as accurate as those of previous data. Also, without the supervision of previous data, separated feature spaces might be destroyed by new domains while updating, leading to catastrophic forgetting. Inspired by the challenges posed by the lack of previous data, we solve this issue from a new standpoint that generates hallucinated previous data for updating FAS model. To this end, we propose a novel Domain-Hallucinated Updating (DHU) framework to facilitate the hallucination of data. Specifically, Domain Information Explorer learns representative domain information of the previous domains. Then, Domain Information Hallucination module transfers the new domain data to pseudo-previous domain ones. Moreover, Hallucinated Features Joint Learning module is proposed to asymmetrically align the new and pseudo-previous data for real samples via dual levels to learn more generalized features, promoting the results on all domains. Our experimental results and visualizations demonstrate that the proposed method outperforms state-of-the-art competitors in terms of effectiveness.
Chengyang Hu, Ke-Yue Zhang, Taiping Yao, Shice Liu, Shouhong Ding, Xin Tan 0002, Lizhuang Ma
AAAI6
2024 Continuous Piecewise-Affine Based Motion Model for Image Animation
abstract
Image animation aims to bring static images to life according to driving videos and create engaging visual content that can be used for various purposes such as animation, entertainment, and education. Recent unsupervised methods utilize affine and thin-plate spline transformations based on keypoints to transfer the motion in driving frames to the source image. However, limited by the expressive power of the transformations used, these methods always produce poor results when the gap between the motion in the driving frame and the source image is large. To address this issue, we propose to model motion from the source image to the driving frame in highly-expressive diffeomorphism spaces. Firstly, we introduce Continuous Piecewise-Affine based (CPAB) transformation to model the motion and present a well-designed inference algorithm to generate CPAB transformation from control keypoints. Secondly, we propose a SAM-guided keypoint semantic loss to further constrain the keypoint extraction process and improve the semantic consistency between the corresponding keypoints on the source and driving images. Finally, we design a structure alignment loss to align the structure-related features extracted from driving and generated images, thus helping the generator generate results that are more consistent with the driving action. Extensive experiments on four datasets demonstrate the effectiveness of our method against state-of-the-art competitors quantitatively and qualitatively. Code will be publicly available at: https://github.com/DevilPG/AAAI2024-CPABMM.
Fengqi Liu, Qianyu Zhou 0001, Ran Yi 0002, Xin Tan 0002, Lizhuang Ma
AAAI5
2024 Learning Task-Aware Language-Image Representation for Class-Incremental Object Detection
abstract
Class-incremental object detection (CIOD) is a real-world desired capability, requiring an object detector to continuously adapt to new tasks without forgetting learned ones, with the main challenge being catastrophic forgetting. Many methods based on distillation and replay have been proposed to alleviate this problem. However, they typically learn on a pure visual backbone, neglecting the powerful representation capabilities of textual cues, which to some extent limits their performance. In this paper, we propose task-aware language-image representation to mitigate catastrophic forgetting, introducing a new paradigm for language-image-based CIOD. First of all, we demonstrate the significant advantage of language-image detectors in mitigating catastrophic forgetting. Secondly, we propose a learning task-aware language-image representation method that overcomes the existing drawback of directly utilizing the language-image detector for CIOD. More specifically, we learn the language-image representation of different tasks through an insulating approach in the training stage, while using the alignment scores produced by task-specific language-image representation in the inference stage. Through our proposed method, language-image detectors can be more practical for CIOD. We conduct extensive experiments on COCO 2017 and Pascal VOC 2007 and demonstrate that the proposed method achieves state-of-the-art results under the various CIOD settings.
Hongquan Zhang, Bin-Bin Gao, Yi Zeng 0006, Xin Tan 0002, Zhizhong Zhang 0001, Yanyun Qu, Jun Liu 0116, Yuan Xie 0006
AAAI5
2024 Domain Alignment with Large Vision-language Models for Cross-domain Remote Sensing Image Retrieval
abstract
Cross-domain remote sensing image retrieval has been a hotspot in the past few years. Most of the existing methods focus on combining semantic learning with domain adaptation on well-labeled source domain and unlabeled target domain. However, they face two serious challenges. (1) They cannot deal with practical scenarios where the source domain lacks sufficient label supervision. (2) They suffer from severe performance degradation when the data distribution between the source domain and target domain becomes highly inconsistent. To address these challenges, we propose D omain A lignment with L arge V ision-language models for cross-domain remote sensing image retrieval (termed as DALV). First, we design a dual-modality prototype guided pseudo-labeling mechanism, which leverages the pre-trained large vision-language model (i.e., CLIP) to assign pseudo-labels for all unlabeled source domain images and target domain images. Second, we compute the confidence scores for these pseudo-labels to distinguish their reliability. Next, we devise a loss reweighting strategy, which incorporates the confidence scores as weight values into the contrastive loss to mitigate the impact of noisy pseudo-labels. Finally, the low-rank adaptation fine-tuning means is adapted to update our model and achieve domain alignment to obtain class discriminative features. Extensive experiments on 12 cross-domain remote sensing image retrieval tasks show that our proposed DALV outperforms the state-of-the-art approaches. The source code is available at https://github.com/ptyy01/DALV.
Guocan Cai, Fufang Li, Yangtao Wang, Xin Tan 0002, Xiaocui Li 0001
CIKM5
2024 Image-text Retrieval with Main Semantics Consistency
abstract
Image-text retrieval (ITR) has been one of the primary tasks in cross-modal retrieval, serving as a crucial bridge between computer vision and natural language processing. Significant progress has been made to achieve global alignment and local alignment between images and texts by mapping images and texts into a common space to establish correspondences between these two modalities. However, the rich semantic content contained in each image may bring false matches, resulting in the matched text ignoring the main semantics but focusing on the secondary or other semantics of this image. To address this issue, this paper proposes a semantically optimized approach with a novel Main Semantics Consistency (MSC) loss function, which aims to rank the semantically most similar images (or texts) corresponding to the given query at the top position during the retrieval process. First, in each batch of image-text pairs, we separately compute (i) the image-image similarity, i.e., the similarity between every two images, (ii) the text-text similarity, i.e., the similarity between a group of texts (that belong to a certain image) and another group of texts (that belong to another image), and (iii) the image-text similarity, i.e., the similarity between each image and each text. Afterward, our proposed MSC effectively aligns the above image-image, image-text, and text-text similarity, since the main semantics of every two images will be highly close if their text descriptions remain highly semantically consistent. By this means, we can capture the main semantics of each image to be matched with its corresponding texts, prioritizing the semantically most related retrieval results. Extensive experiments on MSCOCO and FLICKR30K verify the superior performance of MSC compared with the SOTA image-text retrieval methods. The source code of this project is released at GitHub: https://github.com/xyi007/MSC.
Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Jingjing Li 0001, Xiaocui Li 0001, Weilong Peng, Maobin Tang, Meie Fang
CIKM4
2024 Leveraging Panoptic Prior for 3D Zero-Shot Semantic Understanding Within Language Embedded Radiance Fields
Yuzhou Ji, Xin Tan 0002, Wuyi Liu, Yuan Xie 0006, Lizhuang Ma
CVM (1)2
2024 Explore and Enhance the Generalization of Anomaly DeepFake Detection
Shen Chen 0004, Taiping Yao, Lizhuang Ma, Zhizhong Zhang 0001, Xin Tan 0002
CVM (2)6
2024 Isolation and Integration: A Strong Pre-trained Model-Based Paradigm for Class-Incremental Learning
Wei Zhang 0217, Yuan Xie 0006, Zhizhong Zhang 0001, Xin Tan 0002
CVM (2)4
2024 Building a Strong Pre-Training Baseline for Universal 3D Large-Scale Perception
abstract
An effective pre-training framework with universal 3D representations is extremely desired in perceiving large- scale dynamic scenes. However, establishing such an ideal framework that is both task-generic and label-efficient poses a challenge in unifying the representation of the same primitive across diverse scenes. The current contrastive 3D pre-training methods typically follow a frame-level consistency, which focuses on the 2D-3D relationships in each detached image. Such inconsiderate consistency greatly hampers the promising path of reaching an universal pre-training framework: (1) The cross-scene semantic self-conflict, i.e., the intense collision between primitive segments of the same semantics from different scenes; (2) Lacking a globally unified bond that pushes the cross-scene semantic consistency into 3D representation learning. To address above challenges, we propose a CSC framework that puts a scene-level semantic consistency in the heart, bridging the connection of the similar semantic segments across various scenes. To achieve this goal, we combine the coherent semantic cues provided by the vision foundation model and the knowledge-rich cross-scene prototypes derived from the complementary multi-modality information. These allow us to train a universal 3D pre-training model that facilitates various downstream tasks with less fine-tuning efforts. Empirically, we achieve consistent improvements over SOTA pre-training approaches in semantic segmentation (+1.4% mIoU), object detection (+ 1.0% mAP), and panoptic segmentation (+3.0% PQ) using their task-specific 3D network on nuScenes. Code is released at https://github.com/chenhaomingbob/CSC, hoping to inspire future research.
Haoming Chen, Zhizhong Zhang 0001, Yanyun Qu, Xin Tan 0002, Yuan Xie 0006
CVPR5
2024 PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly Detection
abstract
The vision-language model has brought great improvement to few-shot industrial anomaly detection, which usually needs to design of hundreds of prompts through prompt engineering. For automated scenarios, we first use conventional prompt learning with many-class paradigm as the baseline to automatically learn prompts but found that it can not work well in one-class anomaly detection. To address the above problem, this paper proposes a one-class prompt learning method for few-shot anomaly detection, termed PromptAD. First, we propose semantic concatenation which can transpose normal prompts into anomaly prompts by concatenating normal prompts with anomaly suffixes, thus constructing a large number of negative samples used to guide prompt learning in one-class setting. Furthermore, to mitigate the training challenge caused by the absence of anomaly images, we introduce the concept of explicit anomaly margin, which is used to explicitly control the margin between normal prompt features and anomaly prompt features through a hyper-parameter. For image-level/pixel-level anomaly detection, PromptAD achieves first place in 11/12 few-shot settings on MVTec and VisA. Code is available at https://github.com/FuNz-0/PromptAD.git
Xiaofan Li 0008, Zhizhong Zhang 0001, Xin Tan 0002, Chengwei Chen, Yanyun Qu, Yuan Xie 0006, Lizhuang Ma
CVPR3
2024 COTR: Compact Occupancy TRansformer for Vision-Based 3D Occupancy Prediction
abstract
The autonomous driving community has shown significant interest in 3D occupancy prediction, driven by its exceptional geometric perception and general object recognition capabilities. To achieve this, current works try to construct a Tri-Perspective View (TPV) or Occupancy (OCC) representation extending from the Bird-Eye-View perception. However, compressed views like TPV representation lose 3D geometry information while raw and sparse OCC representation requires heavy but redundant computational costs. To address the above limitations, we propose Compact Occupancy TRansformer (COTR), with a geometry-aware occupancy encoder and a semantic-aware group decoder to reconstruct a compact 3D OCC representation. The occupancy encoder first generates a compact geometrical OCC feature through efficient explicit-implicit view transformation. Then, the occupancy decoder further enhances the semantic discriminability of the compact OCC representation by a coarse-to-fine semantic grouping strategy. Empirical experiments show that there are evident performance gains across multiple baselines, e.g., COTR outperforms baselines with a relative improvement of 8%-15%, demonstrating the superiority of our method. The code is available at https://github.com/NotACracker/COTR.
Qihang Ma, Xin Tan 0002, Yanyun Qu, Lizhuang Ma, Zhizhong Zhang 0001, Yuan Xie 0006
CVPR2
2024 Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer
abstract
Scene text recognition (STR) in the wild frequently en-counters challenges when coping with domain variations, font diversity, shape deformations, etc. A straightforward solution is performing model fine-tuning tailored to a spe-cific scenario, but it is computationally intensive and re-quires multiple model copies for various scenarios. Re-cent studies indicate that large language models (LLMs) can learn from afew demonstration examples in a training-free manner, termed “In-Context Learning” (ICL). Never-theless, applying LLMs as a text recognizer is unacceptably resource-consuming. Moreover, our pilot experiments on LLMs show that ICL fails in STR, mainly attributed to the insufficient incorporation of contextual information from di-verse samples in the training stage. To this end, we intro-duce E2 STR, a STR model trained with context-rich scene text sequences, where the sequences are generated via our proposed in-context training strategy. E2 STR demonstrates that a regular-sized model is sufficient to achieve effective ICL capabilities in STR. Extensive experiments show that E2 STR exhibits remarkable training-free adaptation in var-ious scenarios and outperforms even the fine-tuned state-of-the-art approaches on public benchmarks. The code is released at https://github.com/bytedanceIE2STR.
Jingqun Tang, Chunhui Lin, Binghong Wu, Can Huang 0002, Hao Liu 0003, Xin Tan 0002, Zhizhong Zhang 0001, Yuan Xie 0006
CVPR7
2024 Prompt Gradient Projection for Continual Learning
abstract
Prompt-tuning has demonstrated impressive performance in continual learning by querying relevant prompts for each input instance, which can avoid the introduction of task identifier. Its forgetting is therefore reduced as this instance-wise query mechanism enables us to select and update only relevant prompts. In this paper, we further integrate prompt-tuning with gradient projection approach. Our observation is: prompt-tuning releases the necessity of task identifier for gradient projection method; and gradient projection provides theoretical guarantees against forgetting for prompt-tuning. This inspires a new prompt gradient projection approach (PGP) for continual learning. In PGP, we deduce that reaching the orthogonal condition for prompt gradient can effectively prevent forgetting via the self-attention mechanism in vision-transformer. The condition equations are then realized by conducting Singular Value Decomposition (SVD) on an element-wise sum space between input space and prompt space. We validate our method on diverse datasets and experiments demonstrate the efficiency of reducing forgetting both in class incremental, online class incremental, and task incremental settings. The code is available at https://github.com/JingyangQiao/prompt-gradient-projection.
Jingyang Qiao, Zhizhong Zhang 0001, Xin Tan 0002, Chengwei Chen, Yanyun Qu, Yong Peng 0002, Yuan Xie 0006
ICLR3
2024 Mutual Positive and Negative Learning for Weakly-supervised Point Cloud Semantic Segmentation
abstract
Point cloud semantic segmentation heavily relies on the large-scale point-level annotated dataset, which encourages the weakly-supervised method to prevail gradually. Previous weakly-supervised self-training methods only adopted positive labels, which would be under-performed due to too much noise and lack of supervision. We are the first to present negative labels into the 3D segmentation area, providing extra supervision for hard samples to mitigate the drawbacks induced by noisy labels. Together with the positive labels, we formulate a Mutual Positive-Negative Bi-branch Learning framework to generate positive and negative labels iteratively. Based on that positive branch and negative branch learn complementary knowledge, we build a Mutual Positive-Negative Knowledge Distillation within the bi-branch to further encourage the two branches to learn from each other. Finally, we propose a novel dynamic fusion strategy to fuse predictions from the positive and negative branches, generating more robust predictions. Results on three large-scale datasets show that our method outperforms state-of-the-art weakly-supervised methods by a large margin.
Zhizhong Zhang 0001, Yuan Xie 0006, Guchu Zou, Zhenyi Qi, Xin Tan 0002
ICME7
2024 LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
abstract
Visual Spatial Description (VSD) aims to generate texts that describe the spatial relationships between objects within images. Traditional visual spatial relationship classification (VSRC) methods typically output the spatial relationship between two objects in an image, often neglecting world knowledge and lacking general language capabilities. In this paper, we propose a Large Language-and-Vision Assistant for Visual Spatial Description, named LLaVA-VSD, which is designed for the classification, description, and open-ended description of visual spatial relationships. Specifically, the model first constructs a visual spatial instruction-following dataset using given figure-caption pairs for the three tasks. It then employs LoRA to fine-tune a Large Language and Vision Assistant for VSD, which has 13 billion parameters and supports high-resolution images. Finally, a large language model is used to refine the generated sentences, enhancing their diversity and accuracy. LLaVA-VSD demonstrates excellent multimodal conversational capabilities and can follow open-ended instructions to assist with inquiries about object relationships in images.
Yizhang Jin, Jian Li 0062, Jiangning Zhang, Jianlong Hu, Zhenye Gan, Xin Tan 0002, Yong Liu 0032, Yabiao Wang, Chengjie Wang 0001, Lizhuang Ma
ACM Multimedia6
2024 Harmonizing Visual Text Comprehension and Generation
abstract
In this work, we present TextHarmony, a unified and versatile multimodal generative model proficient in comprehending and generating visual text. Simultaneously generating images and texts typically results in performance degradation due to the inherent inconsistency between vision and language modalities. To overcome this challenge, existing approaches resort to modality-specific data for supervised fine-tuning, necessitating distinct model instances. We propose Slide-LoRA, which dynamically aggregates modality-specific and modality-agnostic LoRA experts, partially decoupling the multimodal generation space. Slide-LoRA harmonizes the generation of vision and language within a singular model instance, thereby facilitating a more unified generative process. Additionally, we develop a high-quality image caption dataset, DetailedTextCaps-100K, synthesized with a sophisticated closed-source MLLM to enhance visual text generation capabilities further. Comprehensive experiments across various benchmarks demonstrate the effectiveness of the proposed approach. Empowered by Slide-LoRA, TextHarmony achieves comparable performance to modality-specific fine-tuning results with only a 2% increase in parameters and shows an average improvement of 2.5% in visual text comprehension tasks and 4.0% in visual text generation tasks. Our work delineates the viability of an integrated approach to multimodal generation within the visual text domain, setting a foundation for subsequent inquiries. Code is available at https://github.com/bytedance/TextHarmony.
Jingqun Tang, Binghong Wu, Chunhui Lin, Shu Wei, Hao Liu 0003, Xin Tan 0002, Zhizhong Zhang 0001, Can Huang 0002, Yuan Xie 0006
NeurIPS7
2024 MSPAN: Multi-scale pyramid attention network for efficient skin cancer lesion segmentation
abstract
Abstract Skin cancer is common and deadly, needs to be detected and treated properly. Deep learning algorithms like UNet have shown potential results in medical imaging. Such approaches still struggle to capture fine‐grained details and scale differences in skin lesions‐based occlusions' appearance, size etc. This research proposes a redesign UNet, the Multi‐Scale Pyramid Attention Network (MSPAN), to improve skin cancer lesion segmentation. The input data is processed at numerous scales with varied receptive fields. This enhances the network's ability to identify lesion locations by capturing local and global context. Attention approaches also help the network to suppress noise by focusing on informative features. We have evaluated MSPAN model on the publicly available ISIC2018 benchmark dataset for skin lesion segmentation. The method surpasses traditional UNet and other current methods in accuracy and effectiveness. The model also has a post‐processing to estimate lesion area for fast inference, making it suitable for extensive screening. Redesigned UNet with the Multi‐Scale Pyramid Attention Network improves skin cancer lesion segmentation. The model's ability to collect fine‐grained information and handle occlusions allows for more accurate skin cancer diagnosis and treatment. The MSPAN design can improve computer‐aided diagnosis systems and help dermatologists make precise clinical decisions.
Noor Ahmed 0002, Xin Tan 0002, Lizhuang Ma
IET Image Process.2
2024 Uni-to-Multi Modal Knowledge Distillation for Bidirectional LiDAR-Camera Semantic Segmentation
abstract
Combining LiDAR points and images for robust semantic segmentation has shown great potential. However, the heterogeneity between the two modalities (e.g. the density, the field of view) poses challenges in establishing a bijective mapping between each point and pixel. This modality alignment problem introduces new challenges in network design and data processing for cross-modal methods. Specifically, 1) points that are projected outside the image planes; 2) the complexity of maintaining geometric consistency limits the deployment of many data augmentation techniques. To address these challenges, we propose a cross-modal knowledge imputation and transition approach. First, we introduce a bidirectional feature fusion strategy that imputes missing image features and performs cross-modal fusion simultaneously. This allows us to generate reliable predictions even when images are missing. Second, we propose a Uni-to-Multi modal Knowledge Distillation (U2MKD) framework, leveraging the transfer of informative features from a single-modality teacher to a cross-modality student. This overcomes the issues of augmentation misalignment and enables us to train the student effectively. Extensive experiments on the nuScenes, Waymo, and SemanticKITTI datasets demonstrate the effectiveness of our approach. Notably, our method achieves an 8.3 mIoU gain over the LiDAR-only baseline on the nuScenes validation set and achieves state-of-the-art performance on the three datasets.
Tianfang Sun, Zhizhong Zhang 0001, Xin Tan 0002, Yong Peng 0002, Yanyun Qu, Yuan Xie 0006
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Glass Makes Blurs: Learning the Visual Blurriness for Glass Surface Detection
abstract
Glass surface detection is challenging as glass normally borrows similar visual appearances from the arbitrary objects/scenes behind it. Although some methods have been proposed to address this problem, they may fail if the reference objects are nonexistent or the additional annotations are missing. This article aims to address the glass surface detection problem by utilizing the intrinsic glass properties without reference objects and additional annotations. We observe glass makes blurs naturally. Based on the investigation of this intrinsic visual blurriness cue, we propose a novel visual blurriness aggregation module to model visual blurriness as a learnable residual in order to extract and aggregate multiscale valuable visual blurriness features used for guiding the backbone features to detect glass precisely. Besides, we note the ratio of the blurred area assists in utilizing the visual blurriness cue caused by glass and propose a visual blurriness driven refinement module to refine glass maps with this ratio to better leverage the visual blurriness information. Extensive experiments show that the proposed method achieves state-of-the-art performance on popular glass surface datasets.
Fulin Qi, Xin Tan 0002, Zhizhong Zhang 0001, Mingang Chen, Yuan Xie 0006, Lizhuang Ma
IEEE Trans. Ind. Informatics2
2024 Image Understands Point Cloud: Weakly Supervised 3D Semantic Segmentation via Association Learning
abstract
Weakly supervised point cloud semantic segmentation methods that require 1% or fewer labels with the aim of realizing almost the same performance as fully supervised approaches have recently attracted extensive research attention. A typical solution in this framework is to use self-training or pseudo-labeling to mine the supervision from the point cloud itself while ignoring the critical information from images. In fact, cameras widely exist in LiDAR scenarios, and this complementary information seems to be highly important for 3D applications. In this paper, we propose a novel cross-modality weakly supervised method for 3D segmentation that incorporates complementary information from unlabeled images. We design a dual-branch network equipped with an active labeling strategy to maximize the power of tiny parts of labels and to directly realize 2D-to-3D knowledge transfer. Afterward, we establish a cross-modal self-training framework, which iterates between parameter updating and pseudolabel estimation. In the training phase, we propose cross-modal association learning to mine complementary supervision from images by reinforcing the cycle consistency between 3D points and 2D superpixels. In the pseudolabel estimation phase, a pseudolabel self-rectification mechanism is derived to filter noisy labels, thus providing more accurate labels for the networks to be fully trained. The extensive experimental results demonstrate that our method even outperforms the state-of-the-art fully supervised competitors with less than 1% actively selected annotations.
Tianfang Sun, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Yuan Xie 0006
IEEE Trans. Image Process.3
2024 PIG: Prompt Images Guidance for Night-Time Scene Parsing
abstract
Night-time scene parsing aims to extract pixel-level semantic information in night images, aiding downstream tasks in understanding scene object distribution. Due to limited labeled night image datasets, unsupervised domain adaptation (UDA) has become the predominant method for studying night scenes. UDA typically relies on paired day-night image pairs to guide adaptation, but this approach hampers dataset construction and restricts generalization across night scenes in different datasets. Moreover, UDA, focusing on network architecture and training strategies, faces difficulties in handling classes with few domain similarities. In this paper, we leverage Prompt Images Guidance (PIG) to enhance UDA with supplementary night knowledge. We propose a Night-Focused Network (NFNet) to learn night-specific features from both target domain images and prompt images. To generate high-quality pseudo-labels, we propose Pseudo-label Fusion via Domain Similarity Guidance (FDSG). Classes with fewer domain similarities are predicted by NFNet, which excels in parsing night features, while classes with more domain similarities are predicted by UDA, which has rich labeled semantics. Additionally, we propose two data augmentation strategies: the Prompt Mixture Strategy (PMS) and the Alternate Mask Strategy (AMS), aimed at mitigating the overfitting of the NFNet to a few prompt images. We conduct extensive experiments on four night-time datasets: NightCity, NightCity+, Dark Zurich, and ACDC. The results indicate that utilizing PIG can enhance the parsing accuracy of UDA. The code is available at https://github.com/qiurui4shu/PIG.
Xin Tan 0002, Yuan Xie 0006, Lizhuang Ma
IEEE Trans. Image Process.4
2024 CSFwinformer: Cross-Space-Frequency Window Transformer for Mirror Detection
abstract
Mirror detection is a challenging task since mirrors do not possess a consistent visual appearance. Even the Segment Anything Model (SAM), which boasts superior zero-shot performance, cannot accurately detect the position of mirrors. Existing methods determine the position of the mirror under hypothetical conditions, such as the correspondence between objects inside and outside the mirror, and the semantic association between the mirror and surrounding objects. However, these assumptions do not apply to all scenarios. For instance, there may be no corresponding real objects to the reflected objects in the scene, or it may be challenging to extract meaningful semantic associations in complex scenes. On the other hand, humans can easily recognize mirrors through the specular texture caused by materials. To mine mirror features in more general scenes, we propose a Cross-Space-Frequency Window Transformer (CSFwinformer) to extract spatial and frequency features for texture analysis. Specifically, we design a Spatial-Frequency Window Alignment module (SFWA) to calculate spatial-frequency feature affinities and learn the difference between mirror and non-mirror textures. We then propose a Dilated Window Attention (DWA) to extract global features to complement the limitation of window alignment. Besides, we propose a Cross-Modality Context Contrast module (CMCC) to fuse cross-modality features and global features, which enables information flow between different windows to take full advantage of cross-modality information. Extensive experiments show that our method performs favorably against state-of-the-art methods on three mirror detection benchmarks and significantly improved SAM performance on mirror detection. The code is available at https://github.com/wangsen99/CSFwinformer.
Qiucheng Yu, Xin Tan 0002, Yuan Xie 0006
IEEE Trans. Image Process.4
2023 Self-supervised Contrastive Feature Refinement for Few-Shot Class-Incremental Learning
Shengjin Ma, Wang Yuan, Xin Tan 0002, Zhizhong Zhang 0001, Lizhuang Ma
CAD/Graphics4
2023 Multi-Centroid Task Descriptor for Dynamic Class Incremental Inference
abstract
Incremental learning could be roughly divided into two categories, i.e., class- and task-incremental learning. The main difference is whether the task ID is given during evaluation. In this paper, we show this task information is indeed a strong prior knowledge, which will bring significant improvement over class-incremental learning baseline, e.g., DER [39]. Based on this observation, we propose a gate network to predict the task ID for class incremental inference. This is challenging as there is no explicit semantic relationship between categories in the concept of task. Therefore, we propose a multi-centroid task descriptor by assuming the data within a task can form multiple clusters. The cluster centers are optimized by pulling relevant sample-centroid pairs while pushing others away, which ensures that there is at least one centroid close to a given sample. To select relevant pairs, we use class prototypes as proxies and solve a bipartite matching problem, making the task descriptor representative yet not degenerate to uni-modal. As a result, our dynamic inference network is trained independently of baseline and provides a flexible, efficient solution to distinguish between tasks. Extensive experiments show our approach achieves state-of-the-art results, e.g., we achieve 72.41% average accuracy on CIFAR100-BOS50, outperforming DER by 3.40%.
Tenghao Cai, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Guannan Jiang, Chengjie Wang 0001, Yuan Xie 0006
CVPR3
2023 Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data
abstract
Semi-supervised learning (SSL) has attracted enormous attention due to its vast potential of mitigating the dependence on large labeled datasets. The latest methods (e.g., FixMatch) use a combination of consistency regularization and pseudo-labeling to achieve remarkable successes. However, these methods all suffer from the waste of complicated examples since all pseudo-labels have to be selected by a high threshold to filter out noisy ones. Hence, the examples with ambiguous predictions will not contribute to the training phase. For better leveraging all unlabeled examples, we propose two novel techniques: Entropy Meaning Loss (EML) and Adaptive Negative Learning (ANL). EML incorporates the prediction distribution of non-target classes into the optimization objective to avoid competition with target class, and thus generating more high-confidence predictions for selecting pseudo-label. ANL introduces the additional negative pseudo-label for all unlabeled data to leverage low-confidence examples. It adaptively allocates this label by dynamically evaluating the top-k performance of the model. EML and ANL do not introduce any additional parameter and hyperparameter. We integrate these techniques with FixMatch, and develop a simple yet powerful framework called FullMatch. Extensive experiments on several common SSL benchmarks (CIFAR-10/100, SVHN, STL-10 and ImageNet) demonstrate that FullMatch exceeds FixMatch by a large margin. Integrated with FlexMatch (an advanced FixMatch-based framework), we achieve state-of-the-art performance. Source code is available at https://github.com/megvii-research/FullMatch.
Xin Tan 0002, Borui Zhao, Zhaowei Chen, Renjie Song, Jiajun Liang, Xuequan Lu
CVPR2
2023 Learning to Detect Mirrors from Videos via Dual Correspondences
abstract
Detecting mirrors from static images has received significant research interest recently. However, detecting mirrors over dynamic scenes is still under-explored due to the lack of a high-quality dataset and an effective method for video mirror detection (VMD). To the best of our knowledge, this is the first work to address the VMD problem from a deep-learning-based perspective. Our observation is that there are often correspondences between the contents inside (reflected) and outside (real) of a mirror, but such correspondences may not always appear in every frame, e.g., due to the change of camera pose. This inspires us to propose a video mirror detection method, named VMD-Net, that can tolerate spatially missing correspondences by considering the mirror correspondences at both the intra-frame level as well as inter-frame level via a dual correspondence module that looks over multiple frames spatially and temporally for correlating correspondences. We further propose a first large-scale dataset for VMD (named VMD-D), which contains 14,987 image frames from 269 videos with corresponding manually annotated masks. Experimental results show that the proposed method outperforms SOTA methods from relevant fields. To enable real-time VMD, our method efficiently utilizes the backbone features by removing the redundant multi-level module design and gets rid of post-processing of the output maps commonly used in existing methods, making it very efficient and practical for real-time video-based applications. Code, dataset, and models are available at https://jiaying.link/cvpr2023-vmd/
Jiaying Lin 0001, Xin Tan 0002, Rynson W. H. Lau
CVPR2
2023 Rethinking Gradient Projection Continual Learning: Stability/Plasticity Feature Space Decoupling
abstract
Continual learning aims to incrementally learn novel classes over time, while not forgetting the learned knowledge. Recent studies have found that learning would not forget if the updated gradient is orthogonal to the feature space. However, previous approaches require the gradient to be fully orthogonal to the whole feature space, leading to poor plasticity, as the feasible gradient direction becomes narrow when the tasks continually come, i.e., feature space is unlimitedly expanded. In this paper, we propose a space decoupling (SD) algorithm to decouple the feature space into a pair of complementary subspaces, i.e., the stability space$\mathcal{I}$and the plasticity space$\mathcal{R}. \mathcal{I}$is established by conducting space intersection between the historic and current feature space, and thus$\mathcal{I}$contains more task-shared bases.$\mathcal{R}$is constructed by seeking the orthogonal complementary subspace of$T$and thus$\mathcal{R}$mainly contains task-specific bases. By putting distinguishing constraints on$\mathcal{R}$and$\mathcal{I}$, our method achieves a better balance between stability and plasticity. Extensive experiments are conducted by applying SD to gradient projection baselines, and show SD is model-agnostic and achieves SOTA results on publicly available datasets.
Zhizhong Zhang 0001, Xin Tan 0002, Jun Liu 0116, Yanyun Qu, Yuan Xie 0006, Lizhuang Ma
CVPR3
2023 Instance and Category Supervision are Alternate Learners for Continual Learning
abstract
Continual Learning (CL) is the constant development of complex behaviors by building upon previously acquired skills. Yet, current CL algorithms tend to incur class-level forgetting as the label information is often quickly overwritten by new knowledge. This motivates attempts to mine instance-level discrimination by resorting to recent self-supervised learning (SSL) techniques. However, previous works have pointed out that the self-supervised learning objective is essentially a trade-off between invariance to distortion and preserving sample information, which seriously hinders the unleashing of instance-level discrimination.In this work, we reformulate SSL from the information-theoretic perspective by disentangling the goal of instance-level discrimination, and tackle the trade-off to promote compact representations with maximally preserved invariance to distortion. On this basis, we develop a novel alternate learning paradigm to enjoy the complementary merits of instance-level and category-level supervision, which yields improved robustness against forgetting and better adaptation to each task. To verify the proposed method, we conduct extensive experiments on four different benchmarks using both class-incremental and task-incremental settings, where the leap in performance and thorough ablation studies demonstrate the efficacy and efficiency of our modeling strategy.
Zhizhong Zhang 0001, Xin Tan 0002, Jun Liu 0116, Chengjie Wang 0001, Yanyun Qu, Guannan Jiang, Yuan Xie 0006
ICCV3
2023 Unveiling the Power of CLIP in Unsupervised Visible-Infrared Person Re-Identification
abstract
Large-scale Vision-Language Pre-training (VLP) model, e.g., CLIP, has demonstrated its natural advantage in generating textual descriptions for images. These textual descriptions afford us greater semantic monitoring insights while not requiring any domain knowledge. In this paper, we propose a new prompt learning paradigm for unsupervised visible-infrared person re-identification (USL-VI-ReID) by taking full advantage of the visual-text representation ability from CLIP. In our framework, we establish a learnable cluster-aware prompt for person images and obtain textual descriptions allowing for subsequent unsupervised training. This description complements the rigid pseudo-labels and provides an important semantic supervised signal. On that basis, we propose a new memory-swapping contrastive learning, where we first find the correlated cross-modal prototypes by the Hungarian matching method and then swap the prototype pairs in the memory. Thus typical contrastive learning without any change could easily associate the cross-modal information. Extensive experiments on the benchmark datasets demonstrate the effectiveness of our method. For example, on SYSU-MM01 we arrive at 54.0% in terms of Rank-1 accuracy, over 9% improvement against state-of-the-art approaches. Code is available at https://github.com/CzAngus/CCLNet.
Zhong Chen 0007, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Yuan Xie 0006
ACM Multimedia3
2023 Multi-domain mixup for scenario-universal face anti-spoofing
Shitao Lu, Shice Liu, Keyue Zhang, Mingang Chen, Xin Tan 0002, Lizhuang Ma
Comput. Graph.5
2023 LW-CovidNet: Automatic covid-19 lung infection detection from chest X-ray images
abstract
Coronavirus Disease 2019 (Covid-19) overtook the worldwide in early 2020, placing the world's health in threat. Automated lung infection detection using Chest X-ray images has a ton of potential for enhancing the traditional covid-19 treatment strategy. However, there are several challenges to detect infected regions from Chest X-ray images, including significant variance in infected features similar spatial characteristics, multi-scale variations in texture shapes and sizes of infected regions. Moreover, high parameters with transfer learning are also a constraints to deploy deep convolutional neural network(CNN) models in real time environment. A novel covid-19 lightweight CNN(LW-CovidNet) method is proposed to automatically detect covid-19 infected regions from Chest X-ray images to address these challenges. In our proposed hybrid method of integrating Standard and Depth-wise Separable convolutions are used to aggregate the high level features and also compensate the information loss by increasing the Receptive Field of the model. The detection boundaries of disease regions representations are then enhanced via an Edge-Attention method by applying heatmaps for accurate detection of disease regions. Extensive experiments indicate that the proposed LW-CovidNet surpasses most cutting-edge detection methods and also contributes to the advancement of state-of-the-art performance. It is envisaged that with reliable accuracy, this method can be introduced for clinical practices in the future.
Noor Ahmed 0002, Xin Tan 0002, Lizhuang Ma
IET Image Process.2
2023 A new method proposed to Melanoma-skin cancer lesion detection and segmentation based on hybrid convolutional neural network
Noor Ahmed 0002, Xin Tan 0002, Lizhuang Ma
Multim. Tools Appl.2
2023 Mirror Detection With the Visual Chirality Cue
abstract
Mirror detection is challenging because the visual appearances of mirrors change depending on those of their surroundings. As existing mirror detection methods are mainly based on extracting contextual contrast and relational similarity between mirror and non-mirror regions, they may fail to identify a mirror region if these assumptions are violated. Inspired by a recent study of applying a CNN to help distinguish whether an image is flipped or not based on the visual chirality property, in this paper, we rethink this image-level visual chirality property and reformulate it as a learnable pixel level cue for mirror detection. Specifically, we first propose a novel flipping-convolution-flipping (FCF) transformation to model visual chirality as learnable commutative residual. We then propose a novel visual chirality embedding (VCE) module to exploit this commutative residual in multi-scale feature maps, to embed the visual chirality features into our mirror detection model. Besides, we also propose a visual chirality-guided edge detection (CED) module to integrate the visual chirality features with contextual features for detection refinement. Extensive experiments show that the proposed method outperforms state-of-the-art methods on three benchmark datasets.
Xin Tan 0002, Jiaying Lin 0001, Ke Xu 0010, Lizhuang Ma, Rynson W. H. Lau
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Positive-Negative Receptive Field Reasoning for Omni-Supervised 3D Segmentation
abstract
Hidden features in the neural networks usually fail to learn informative representation for 3D segmentation as supervisions are only given on output prediction, while this can be solved by omni-scale supervision on intermediate layers. In this paper, we bring the first omni-scale supervision method to 3D segmentation via the proposed gradual Receptive Field Component Reasoning (RFCR), where target Receptive Field Component Codes (RFCCs) is designed to record categories within receptive fields for hidden units in the encoder. Then, target RFCCs will supervise the decoder to gradually infer the RFCCs in a coarse-to-fine categories reasoning manner, and finally obtain the semantic labels. To purchase more supervisions, we also propose an RFCR-NL model with complementary negative codes (i.e., Negative RFCCs, NRFCCs) with negative learning. Because many hidden features are inactive with tiny magnitudes and make minor contributions to RFCC prediction, we propose Feature Densification with a centrifugal potential to obtain more unambiguous features, and it is in effect equivalent to entropy regularization over features. More active features can unleash the potential of omni-supervision method. We embed our method into three prevailing backbones, which are significantly improved in all three datasets on both fully and weakly supervised segmentation tasks and achieve competitive performances.
Xin Tan 0002, Qihang Ma, Jingyu Gong, Zhizhong Zhang 0001, Yanyun Qu, Yuan Xie 0006, Lizhuang Ma
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Semantic-Aware Dehazing Network With Adaptive Feature Fusion
abstract
Despite that convolutional neural networks (CNNs) have shown high-quality reconstruction for single image dehazing, recovering natural and realistic dehazed results remains a challenging problem due to semantic confusion in the hazy scene. In this article, we show that it is possible to recover textures faithfully by incorporating semantic prior into dehazing network since objects in haze-free images tend to show certain shapes, textures, and colors. We propose a semantic-aware dehazing network (SDNet) in which the semantic prior is taken as a color constraint for dehazing, benefiting the acquisition of a reasonable scene configuration. In addition, we design a densely connected block to capture global and local information for dehazing and semantic prior estimation. To eliminate the unnatural appearance of some objects, we propose to fuse the features from shallow and deep layers adaptively. Experimental results demonstrate that our proposed model performs favorably against the state-of-the-art single image dehazing approaches.
Shengdong Zhang, Wenqi Ren, Xin Tan 0002, Zhi-Jie Wang 0009, Yong Liu 0018, Jingang Zhang, Xiaoqin Zhang 0002, Xiaochun Cao
IEEE Trans. Cybern.3
2023 Boosting Night-Time Scene Parsing With Learnable Frequency
abstract
Night-Time Scene Parsing (NTSP) is essential to many vision applications, especially for autonomous driving. Most of the existing methods are proposed for day-time scene parsing. They rely on modeling pixel intensity-based spatial contextual cues under even illumination. Hence, these methods do not perform well in night-time scenes as such spatial contextual cues are buried in the over-/under-exposed regions in night-time scenes. In this paper, we first conduct an image frequency-based statistical experiment to interpret the day-time and night-time scene discrepancies. We find that image frequency distributions differ significantly between day-time and night-time scenes, and understanding such frequency distributions is critical to NTSP problem. Based on this, we propose to exploit the image frequency distributions for night-time scene parsing. First, we propose a Learnable Frequency Encoder (LFE) to model the relationship between different frequency coefficients to measure all frequency components dynamically. Second, we propose a Spatial Frequency Fusion module (SFF) that fuses both spatial and frequency information to guide the extraction of spatial context features. Extensive experiments show that our method performs favorably against the state-of-the-art methods on the NightCity, NightCity+ and BDD100K-night datasets. In addition, we demonstrate that our method can be applied to existing day-time scene parsing methods and boost their performance on night-time scenes. The code is available at https://github.com/wangsen99/FDLNet.
Ke Xu 0010, Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006, Lizhuang Ma
IEEE Trans. Image Process.5
2023 Frequency-aware Camouflaged Object Detection
abstract
Camouflaged object detection (COD) is important as it has various potential applications. Unlike salient object detection (SOD), which tries to identify visually salient objects, COD tries to detect objects that are visually very similar to the surrounding background. We observe that recent COD methods try to fuse features from different levels using some context aggregation strategies originally developed for SOD. Such an approach, however, may not be appropriate for COD as these existing context aggregation strategies are good at detecting distinctive objects while weakening the features from less discriminative objects. To address this problem, we propose in this article to exploit frequency learning to suppress the confusing high-frequency texture information, to help separate camouflaged objects from their surrounding background, and a frequency-based method, called FBNet, for camouflaged object detection. Specifically, we design a frequency-aware context aggregation (FACA) module to suppress high-frequency information and aggregate multi-scale features from a frequency perspective, an adaptive frequency attention (AFA) module to enhance the features of the learned important frequency components, and a gradient-weighted loss function to guide the proposed method to pay more attention to contour details. Experimental results show that our model outperforms relevant state-of-the-art methods.
Jiaying Lin 0001, Xin Tan 0002, Ke Xu 0010, Lizhuang Ma, Rynson W. H. Lau
ACM Trans. Multim. Comput. Commun. Appl.2
2023 HSNet: hierarchical semantics network for scene parsing
Xin Tan 0002, Ying Cao 0001, Ke Xu 0010, Lizhuang Ma, Rynson W. H. Lau
Vis. Comput.1
2022 Rethinking Efficient Lane Detection via Curve Modeling
abstract
This paper presents a novel parametric curve-based method for lane detection in RGB images. Unlike state-of-the-art segmentation-based and point detection-based methods that typically require heuristics to either decode predictions or formulate a large sum of anchors, the curve-based methods can learn holistic lane representations naturally. To handle the optimization difficulties of existing poly-nomial curve methods, we propose to exploit the parametric Bézier curve due to its ease of computation, stability, and high freedom degrees of transformations. In addition, we propose the deformable convolution-based feature flip fusion, for exploiting the symmetry properties of lanes in driving scenes. The proposed method achieves a new state-of-the-art performance on the popular LLAMAS benchmark. It also achieves favorable accuracy on the TuSimple and CULane datasets, while retaining both low latency (>150 FPS) and small model size (<10M). Our method can serve as a new baseline, to shed the light on the parametric curves modeling for lane detection. Codes of our model and PytorchAutoDrive: a unified framework for self-driving perception, are available at: https://github.com/voldemortX/pytorch-auto-drive.
Zhengyang Feng, Shaohua Guo, Xin Tan 0002, Ke Xu 0010, Min Wang 0024, Lizhuang Ma
CVPR3
2022 Optimization over Disentangled Encoding: Unsupervised Cross-Domain Point Cloud Completion via Occlusion Factor Manipulation
Jingyu Gong, Fengqi Liu, Min Wang 0024, Xin Tan 0002, Zhizhong Zhang 0001, Ran Yi 0002, Yuan Xie 0006, Lizhuang Ma
ECCV (2)5
2022 DMT: Dynamic mutual training for semi-supervised learning
Zhengyang Feng, Qianyu Zhou 0001, Xin Tan 0002, Xuequan Lu, Jianping Shi, Lizhuang Ma
Pattern Recognit.4
2022 Sketch-to-photo face generation based on semantic consistency preserving and similar connected component refinement
Junshu Tang, Zhiwen Shao, Xin Tan 0002, Lizhuang Ma
Vis. Comput.4
2021 Boundary-Aware Geometric Encoding for Semantic Segmentation of Point Clouds
abstract
Boundary information plays a significant role in 2D image segmentation, while usually being ignored in 3D point cloud segmentation where ambiguous features might be generated in feature extraction, leading to misclassification in the transition area between two objects. In this paper, firstly, we propose a Boundary Prediction Module (BPM) to predict boundary points. Based on the predicted boundary, a boundary-aware Geometric Encoding Module (GEM) is designed to encode geometric information and aggregate features with discrimination in a neighborhood, so that the local features belonging to different categories will not be polluted by each other. To provide extra geometric information for boundary-aware GEM, we also propose a light-weight Geometric Convolution Operation (GCO), making the extracted features more distinguishing. Built upon the boundary-aware GEM, we build our network and test it on benchmarks like ScanNet v2, S3DIS. Results show our methods can significantly improve the baseline and achieve state-of-the-art performance.
Jingyu Gong, Xin Tan 0002, Jie Zhou 0029, Yanyun Qu, Yuan Xie 0006, Lizhuang Ma
AAAI3
2021 Omni-Supervised Point Cloud Segmentation via Gradual Receptive Field Component Reasoning
abstract
Hidden features in neural network usually fail to learn informative representation for 3D segmentation as supervisions are only given on output prediction, while this can be solved by omni-scale supervision on intermediate layers. In this paper, we bring the first omni-scale supervision method to point cloud segmentation via the proposed gradual Receptive Field Component Reasoning (RFCR), where target Receptive Field Component Codes (RFCCs) are designed to record categories within receptive fields for hidden units in the encoder. Then, target RFCCs will supervise the decoder to gradually infer the RFCCs in a coarse-to-fine categories reasoning manner, and finally obtain the semantic labels. Because many hidden features are inactive with tiny magnitude and make minor contributions to RFCC prediction, we propose a Feature Densification with a centrifugal potential to obtain more unambiguous features, and it is in effect equivalent to entropy regularization over features. More active features can further unleash the potential of our omni-supervision method. We embed our method into four prevailing backbones and test on three challenging benchmarks. Our method can significantly improve the backbones in all three datasets. Specifically, our method brings new state-of-the-art performances for S3DIS as well as Semantic3D and ranks the 1st in the ScanNet benchmark among all the point-based methods. Code is publicly available at https://github.com/azuki-miho/RFCR.
Jingyu Gong, Xin Tan 0002, Yanyun Qu, Yuan Xie 0006, Lizhuang Ma
CVPR3
2021 Confident Semantic Ranking Loss for Part Parsing
abstract
Part parsing is taken as a dense prediction task, assigning each pixel a semantic part label. Some previous methods tried to model the human-known relationships among different parts (inter-part). However, these methods are hard to be used for multi-object part parsing since the given relationships are highly dependent on human priors which require the special model to learn. In addition, pixels in the same part (intra-part) are always assumed equally important. In fact, even they belong to the same part, some pixels are quite uncertain for their predictions while some are with high confidence, but theoretically they are representing the same feature. In this paper, we study the inequality and uncertainty of intra-part and inter-part pixels and propose the confident-semantic-ranking (CO-Rank) loss function, which maximizes the similarities of different groups of pixels to alleviate the uncertainty and models the pixel relationships among intra-/inter-parts. In addition, previous feature maps lost some of part-level relationships due to simply using the global max/average pooling, hence, we propose a new Global Object Pooling layer (GOP) to encode the abundant global information while preserving the geometry details. The experimental results show that our proposed method achieves new state-of-the-art performance on multi-class part parsing benchmark Pascal-Part dataset.
Xin Tan 0002, Jinkun Hao, Lizhuang Ma
ICME1
2021 Self-supervised Compressed Video Action Recognition via Temporal-Consistent Sampling
Shaohui Lin, Xin Tan 0002, Lizhuang Ma
ICONIP (4)5
2021 Novelty Detection via Contrastive Learning with Negative Data Augmentation
abstract
Novelty detection is the process of determining whether a query example differs from the learned training distribution. Previous generative adversarial networks based methods and self-supervised approaches suffer from instability training, mode dropping, and low discriminative ability. We overcome such problems by introducing a novel decoder-encoder framework. Firstly, a generative network (decoder) learns the representation by mapping the initialized latent vector to an image. In particular, this vector is initialized by considering the entire distribution of training data to avoid the problem of mode-dropping. Secondly, a contrastive network (encoder) aims to ``learn to compare'' through mutual information estimation, which directly helps the generative network to obtain a more discriminative representation by using a negative data augmentation strategy. Extensive experiments show that our model has significant superiority over cutting-edge novelty detectors and achieves new state-of-the-art results on various novelty detection benchmarks, e.g. CIFAR10 and DCASE. Moreover, our model is more stable for training in a non-adversarial manner, compared to other adversarial based novelty detection methods.
Chengwei Chen, Yuan Xie 0006, Shaohui Lin, Ruizhi Qiao, Xin Tan 0002, Lizhuang Ma
IJCAI6
2021 Single Image Deraining via detail-guided Efficient Channel Attention Network
Xiao Lin 0012, Qi Huang 0005, Xin Tan 0002, Meie Fang, Lizhuang Ma
Comput. Graph.4
2021 Weakly-Supervised Saliency Detection via Salient Object Subitizing
abstract
Salient object detection aims at detecting the most visually distinct objects and producing the corresponding masks. As the cost of pixel-level annotations is high, image tags are usually used as weak supervisions. However, an image tag can only be used to annotate one class of objects. In this paper, we introduce saliency subitizing as the weak supervision since it is class-agnostic. This allows the supervision to be aligned with the property of saliency detection, where the salient objects of an image could be from more than one class. To this end, we propose a model with two modules, Saliency Subitizing Module (SSM) and Saliency Updating Module (SUM). While SSM learns to generate the initial saliency masks using the subitizing information, without the need for any unsupervised methods or some random seeds, SUM helps iteratively refine the generated saliency masks. We conduct extensive experiments on five benchmark datasets. The experimental results show that our method outperforms other weakly-supervised methods and even performs comparable to some fully-supervised methods.
Xin Tan 0002, Jie Zhou 0029, Lizhuang Ma, Rynson W. H. Lau
IEEE Trans. Circuits Syst. Video Technol.2
2021 Night-Time Scene Parsing With a Large Real Dataset
abstract
Although huge progress has been made on scene analysis in recent years, most existing works assume the input images to be in day-time with good lighting conditions. In this work, we aim to address the night-time scene parsing (NTSP) problem, which has two main challenges: 1) labeled night-time data are scarce, and 2) over- and under-exposures may co-occur in the input night-time images and are not explicitly modeled in existing pipelines. To tackle the scarcity of night-time data, we collect a novel labeled dataset, named NightCity, of 4,297 real night-time images with ground truth pixel-level semantic annotations. To our knowledge, NightCity is the largest dataset for NTSP. In addition, we also propose an exposure-aware framework to address the NTSP problem through augmenting the segmentation process with explicitly learned exposure features. Extensive experiments show that training on NightCity can significantly improve NTSP performances and that our exposure-aware model outperforms the state-of-the-art methods, yielding top performances on our dataset as well as existing datasets.
Xin Tan 0002, Ke Xu 0010, Ying Cao 0001, Lizhuang Ma, Rynson W. H. Lau
IEEE Trans. Image Process.1
2020 A Shape-Aware Feature Extraction Module for Semantic Segmentation of 3D Point Clouds
Jie Zhou 0029, Xin Tan 0002, Lizhuang Ma
ICONIP (4)3
2020 Learning Object Deformation and Motion Adaption for Semi-supervised Video Object Segmentation
Xin Tan 0002, Jianming Guo, Lizhuang Ma
ICPR2
2020 SceneEncoder: Scene-Aware Semantic Segmentation of Point Clouds with A Learnable Scene Descriptor
abstract
Besides local features, global information plays an essential role in semantic segmentation, while recent works usually fail to explicitly extract the meaningful global information and make full use of it. In this paper, we propose a SceneEncoder module to impose a scene-aware guidance to enhance the effect of global information. The module predicts a scene descriptor, which learns to represent the categories of objects existing in the scene and directly guides the point-level semantic segmentation through filtering out categories not belonging to this scene. Additionally, to alleviate segmentation noise in local region, we design a region similarity loss to propagate distinguishing features to their own neighboring points with the same label, leading to the enhancement of the distinguishing ability of point-wise features. We integrate our methods into several prevailing networks and conduct extensive experiments on benchmark datasets ScanNet and ShapeNet. Results show that our methods greatly improve the performance of baselines and achieve state-of-the-art performance.
Jingyu Gong, Jie Zhou 0029, Xin Tan 0002, Yuan Xie 0006, Lizhuang Ma
IJCAI4
2020 Deep multi-center learning for face alignment
Zhiwen Shao, Hengliang Zhu, Xin Tan 0002, Yangyang Hao, Lizhuang Ma
Neurocomputing3
2019 Object-Level Salience Detection by Progressively Enhanced Network
Wang Yuan, Xin Tan 0002, Chengwei Chen, Shouhong Ding, Lizhuang Ma
ICANN (3)3
2019 Learning the Spiral Sharing Network with Minimum Salient Region Regression for Saliency Detection
abstract
With the development of convolutional neural networks (CNNs), saliency detection methods have made a big progress in recent years. However, the previous methods sometimes mistakenly highlight the non-salient region, especially in complex backgrounds. To solve this problem, a two-stage method for saliency detection is proposed in this paper. In the first stage, a network is used to regress the minimum salient region (RMSR) containing all salient objects. Then in the second stage, in order to fuse the multi-level features, the spiral sharing network (SSN) is proposed for pixel-level detection on the result of RMSR. Experimental results on four public datasets show that our model is effective over the state-of-the-art approaches.
Zukai Chen, Xin Tan 0002, Hengliang Zhu, Shouhong Ding, Lizhuang Ma
ICASSP2
2019 Re-ID Driven Localization Refinement for Person Search
abstract
Person search aims at localizing and identifying a query person from a gallery of uncropped scene images. Different from person re-identification (re-ID), its performance also depends on the localization accuracy of a pedestrian detector. The state-of-the-art methods train the detector individually, and the detected bounding boxes may be sub-optimal for the following re-ID task. To alleviate this issue, we propose a re-ID driven localization refinement framework for providing the refined detection boxes for person search. Specifically, we develop a differentiable ROI transform layer to effectively transform the bounding boxes from the original images. Thus, the box coordinates can be supervised by the re-ID training other than the original detection task. With this supervision, the detector can generate more reliable bounding boxes, and the downstream re-ID model can produce more discriminative embeddings based on the refined person localizations. Extensive experimental results on the widely used benchmarks demonstrate that our proposed method performs favorably against the state-of-the-art person search methods.
Chuchu Han, Jiacheng Ye, Yunshan Zhong, Xin Tan 0002, Chi Zhang 0026, Changxin Gao, Nong Sang
ICCV4
2019 Accurate and Efficient Object Detection with Context Enhancement Block
abstract
Recently feature pyramid composed of multi-level feature maps has been extensively used in region-free detectors to address multi-scale object detection. However, the contradiction between scale and context in the feature pyramid limits the detection performance, extraordinarily on small objects. Most works introduce an extra top-down path to overcome the limitation yet suffering from high computational burden. In this paper, we propose a novel Expansion Receptive Field Block (ERFB) to capture multiple strong contextual features at low computational cost, and then apply the Feature Attention Block (FAB) to eliminate the inconsistency between different features to generate more discriminative features. To be further, we construct an efficient and accurate detector (named CEBNet) mainly consists of Context Enhancement Blocks (CEBs), which are cascaded with ERFB and FAB. The extensive experiments on Pascal VOC and MS COCO demonstrate that CEBNet achieves state-of-the-art detection accuracy at a real-time processing speed.
Min Zhao 0010, Xin Tan 0002, Dihua Sun
ICME3
2019 MCCH: A novel convex hull prior based solution for saliency detection
Xiao Lin 0012, Zhi-Jie Wang 0009, Xin Tan 0002, Meie Fang, Naixue Xiong, Lizhuang Ma
Inf. Sci.3
2018 Saliency Detection by Deep Network with Boundary Refinement and Global Context
abstract
A novel end-to-end fully convolutional neural network for saliency detection is proposed in this paper, aiming at refining the boundary and covering the global context (GBR-Net). Previous CNN based methods for saliency detection are universally accompanied with blurring edge and ambiguous salient object. To tackle this problem, we propose to embed the boundary enhancement block (BEB) into the network to refine edge. It keeps the details by the mutual-coupling con-volutionallayers. Besides, we employ a pooling pyramid that utilizes the multi-level feature informations to search global context, and it also contributes as an auxiliary supervision. The final saliency map is obtained by fusing the edge refinement with global context extraction. Experiments on four benchmark datasets prove that the proposed saliency detection model gains an edge over the state-of-the-art approaches.
Xin Tan 0002, Hengliang Zhu, Zhiwen Shao, Xiao-Nan Hou, Yangyang Hao, Lizhuang Ma
ICME1
2018 Multi-Path Feature Fusion Network for Saliency Detection
abstract
Recent saliency detection methods have made great progress with the fully convolutional network. However, we find that the saliency maps are usually coarse and fuzzy, especially near the boundary of salient object. To deal with this problem, in this paper, we exploit a multi-path feature fusion model for saliency detection. The proposed model is a fully convolutional network with raw images as input and saliency maps as output. In particular, we propose a multi-path fusion strategy for deriving the intrinsic features of salient objects. The structure has the ability of capturing the low-level visual features and generating the boundary-preserving saliency maps. Moreover, a coupled structure module is proposed in our model, which helps to explore the high-level semantic properties of salient objects. Extensive experiments on four public benchmarks indicate that our saliency model is effective and outperforms state-of-the-art methods.
Hengliang Zhu, Xin Tan 0002, Zhiwen Shao, Yangyang Hao, Lizhuang Ma
ICME2
2018 Facial Landmark Detection Under Large Pose
Yangyang Hao, Hengliang Zhu, Zhiwen Shao, Xin Tan 0002, Lizhuang Ma
ICONIP (4)4