Wei Shen 0002

dblp:71/3692-2 · DBLP profile ↗
← Back
126ranked-venue papers
19as first author
79since 2021 · last 2026
0000-0002-1235-598XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 86 · 16 first-author · 57 since 2021Graphics, computer vision, multimedia, augmented reality and games · 83 · 8 first-author · 53 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Dereflection Any Image with Diffusion Priors and Diversified Data
abstract
Reflection removal of a single image remains a highly challenging task due to the complex entanglement between target scenes and unwanted reflections. Despite significant progress, existing methods are hindered by the scarcity of high-quality, diverse data and insufficient restoration priors, resulting in limited generalization across various real-world scenarios. In this paper, we propose Dereflection Any Image, a comprehensive solution with an efficient data preparation pipeline and a generalizable model for robust reflection removal. First, we introduce a dataset named Diverse Reflection Removal (DRR) created by randomly rotating reflective mediums in target scenes, enabling variation of reflection angles and intensities, and setting a new benchmark in scale, quality, and diversity. Second, we propose a diffusion-based framework with one-step diffusion for deterministic outputs and fast inference. To ensure stable learning, we design a three-stage progressive training strategy, including reflection-invariant finetuning to encourage consistent outputs across varying reflection patterns that characterize our dataset. Extensive experiments show that our method achieves SOTA performance on both common benchmarks and challenging in-the-wild images, showing superior generalization across diverse real-world scenes.
Jichen Hu, Chen Yang 0023, Zanwei Zhou, Jiemin Fang, Qi Tian 0001, Wei Shen 0002
AAAI6
2026 WorldGrow: Generating Infinite 3D World
abstract
We tackle the challenge of generating the infinitely extendable 3D world -- large, continuous environments with coherent geometry and realistic appearance. Existing methods face key challenges: 2D-lifting approaches suffer from geometric and appearance inconsistencies across views, 3D implicit representations are hard to scale up, and current 3D foundation models are mostly object-centric, limiting their applicability to scene-level generation. Our key insight is leveraging strong generation priors from pre-trained 3D models for structured scene block generation. To this end, we propose WorldGrow, a hierarchical framework for unbounded 3D scene synthesis. Our method features three core components: (1) a data curation pipeline that extracts high-quality scene blocks for training, making the 3D structured latent representations suitable for scene generation; (2) a 3D block inpainting mechanism that enables context-aware scene extension; and (3) a coarse-to-fine generation strategy that ensures both global layout plausibility and local geometric/textural fidelity. Evaluated on the large-scale 3D-FRONT dataset, WorldGrow achieves SOTA performance in geometry reconstruction, while uniquely supporting infinite scene generation with photorealistic and structurally consistent outputs. These results highlight its capability for constructing large-scale virtual environments and potential for building future world models.
Sikuang Li, Chen Yang 0023, Jiemin Fang, Taoran Yi, Jiazhong Cen, Lingxi Xie, Wei Shen 0002, Qi Tian 0001
AAAI8
2026 CP-CLIP: Customized Parameter Generation for Open-vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation aims to assign pixel-level labels to images based on textual descriptions, even for categories beyond predefined closed sets. While vision-language foundation models like CLIP are widely used for this task, fine-tuning them for pixel-level predictions often compromises their generalization capabilities. To address this, we propose a novel fine-tuning strategy, CP-CLIP, which generates customized parameters for CLIP without sacrificing its generalization. Our method employs a customized parameter generator that produces newly added parameters based on random noise, using local visual features from CLIP's image encoder as conditions, enabling generalization to new images from unseen scenarios. Additionally, we introduce an orthogonal adaptation technique to ensure the update direction is orthogonal to the pre-trained weights, largely preserving the initial generalization ability. Extensive experiments demonstrate that CP-CLIP achieves state-of-the-art performance across multiple benchmarks in open-vocabulary semantic segmentation.
Zelin Peng, Zhengqin Xu, Wei Shen 0002
AAAI4
2026 Efficient Segmentation with Multimodal Large Language Model via Token Routing
abstract
Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in addressing open-world segmentation tasks. However, the substantial computational cost of the LLM components presents a significant challenge, especially in segmentation tasks, where efficiency has long been a central concern. Existing efficient MLLM approaches typically reduce computation cost by pruning visual tokens in the early layers, as they account for the majority of the input sequence. Despite their efficiency, this is incompatible with dense prediction tasks such as segmentation, since removing visual tokens leads to the loss of essential object parts and spatial details. To better understand the roles of visual tokens in segmentation, we analyze the attention weights of both image and mask tokens within LLM. We find that image tokens are important throughout all layers, whereas mask tokens only attend to image tokens at deeper layers. Based on the observation, we build an efficient segmentation framework based on MLLMs by introducing a sophisticated token routing strategy. This strategy dynamically determines when and how different tokens participate in computation: For mask tokens, they are only inserted at deeper layers of the LLM to reduce redundant computation, since they rarely attend to image tokens in early layers; For image tokens, only a small number of them, named proxies, are updated via full feedforward network (FFN) computation, while the update of the remaining tokens is guided by these proxies, i.e., efficiently computed through a lightweight projector applied on the difference of the proxies during their update. Our method achieves a 1.5× acceleration over the original LLM process by reducing its FLOPs to 56%, while maintaining the same segmentation performance.
Changsong Wen, Zelin Peng, Wei Shen 0002
AAAI4
2026 Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
abstract
Flow-based 3D generation models typically require dozens of sampling steps during inference. Though few-step distillation methods, particularly Consistency Models (CMs), have achieved substantial advancements in accelerating 2D diffusion models, they remain under-explored for more complex 3D generation tasks. In this study, we propose a novel framework, MDT-dist, for few-step 3D flow distillation. Our approach is built upon a primary objective: distilling the pretrained model to learn the Marginal-Data Transport. Directly learning this objective needs to integrate the velocity fields, while this integral is intractable to be implemented. Therefore, we propose two optimizable objectives, Velocity Matching (VM) and Velocity Distillation (VD), to equivalently convert the optimization target from the transport level to the velocity and the distribution level respectively. Velocity Matching (VM) learns to stably match the velocity fields between the student and the teacher, but inevitably provides biased gradient estimates. Velocity Distillation (VD) further enhances the optimization process by leveraging the learned velocity fields to perform probability density distillation. When evaluated on the pioneer 3D generation framework TRELLIS, our method reduces sampling steps of each flow transformer from 25 to 1–2, achieving 0.68s (1 step x2) and 0.94s (2 steps x2) latency with 9.0x and 6.5x speedup on A800, while preserving high visual and geometric fidelity. Experiments demonstrate that our method significantly outperforms existing CM distillation methods, and enables TRELLIS to achieve superior performance in few-step 3D generation.
Zanwei Zhou, Taoran Yi, Jiemin Fang, Chen Yang 0023, Lingxi Xie, Xinggang Wang, Wei Shen 0002, Qi Tian 0001
AAAI7
2026 Task-Specific Directions: Definition, Exploration, and Utilization in Parameter Efficient Fine-Tuning
abstract
Large language models demonstrate impressive performance on downstream tasks, yet requiring extensive resource consumption when fully fine-tuning all parameters. To mitigate this, Parameter Efficient Fine-Tuning (PEFT) strategies, such as LoRA, have been developed. In this paper, we delve into the concept of task-specific directions (TSDs)-critical for transitioning large models from pretrained states to task-specific enhancements in PEFT. We propose a framework to clearly define these directions and explore their properties, and practical utilization challenges. We then introduce a novel approach, LoRA-Dash, which aims to maximize the impact of TSDs during the fine-tuning process, thereby enhancing model performance on targeted tasks. Additionally, based on our exploration of TSD, we focus on an important issue in PEFT: the initialization of LoRA. While some works have pointed out the significance of initialization for LoRA's performance and proposed various strategies, these methods are often empirical and not task-specific. To address this issue, we propose LoRA-Init. Starting from TSD, we identify the directions that require the most adjustment during fine-tuning for downstream tasks. By initializing the matrices in LoRA with these directions, LoRA-Init significantly enhances LoRA's performance. Moreover, we can combine LoRA-Dash and LoRA-Init to create the final version of LoRA based on TSDs, which we refer to as LoRA-TSD. Extensive experiments have conclusively demonstrated the effectiveness of these methods, and in-depth analyses further reveal the underlying mechanisms of these methods.
Chongjie Si, Zhiyi Shi, Shifan Zhang, Xiaokang Yang 0001, Hanspeter Pfister, Wei Shen 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 DMformer: Difficulty-Adapted Masked Transformer for Semi-Supervised Medical Image Segmentation
abstract
The shared anatomy among different human bodies can serve as a strong prior for effectively leveraging unlabeled data in semi-supervised medical image segmentation. Inspired by the success of masked image modeling, we notice that this prior can be explicitly realized by incorporating an auxiliary unsupervised gross anatomy reconstruction task into a teacher-student semi-supervised segmentation framework. In this auxiliary task, consistency is maintained between the student's predictions on masked images and the teacher's predictions on the original images. Despite its potential, we observe that the reconstruction difficulties of different organs/tissues can vary significantly and therefore reconstructing them requires tailored learning strategies. To address this issue, we introduce a difficulty-adapted mask mechanism based on the teacher-student framework, wherein the reconstruction difficulty is adapted to facilitate training. Specifically, we control the reconstruction difficulty by modulating two important factors: masked region ratio and masked class ratio. Accordingly, we design two corresponding mask strategies. 1) Region-based masking: randomly masks a fraction of each class according to an automatically computed mask ratio. 2) Class-based masking: masks the entire regions of the specific classes according to the class confidence predicted by the teacher model. During training, a conflict-aware gradient computation strategy is introduced to mitigate potential optimization conflicts arising from modulating the two reconstruction factors simultaneously. By building on vision transformers, we develop an Difficulty-adapted Masked Transformer (DMformer) for semi-supervised medical image segmentation. Extensive experiments demonstrate the superiority of DMformer, which outperforms the previous SOTA by 9.53% and 4.63% in terms of DSC on ACDC dataset with 5% labeled images and Synapse dataset with 30% labeled images, respectively.
Zelin Peng, Guanchun Wang, Zhengqin Xu, Xiaokang Yang 0001, Wei Shen 0002
IEEE J. Biomed. Health Informatics5
2025 Segment Any 3D Gaussians
abstract
This paper presents SAGA (Segment Any 3D GAussians), a highly efficient 3D promptable segmentation method based on 3D Gaussian Splatting (3D-GS). Given 2D visual prompts as input, SAGA can segment the corresponding 3D target represented by 3D Gaussians within 4 ms. This is achieved by attaching a scale-gated affinity feature to each 3D Gaussian to endow it a new property towards multi-granularity segmentation. Specifically, a scale-aware contrastive training strategy is proposed for the scale-gated affinity feature learning. It 1) distills the segmentation capability of the Segment Anything Model (SAM) from 2D masks into the affinity features and 2) employs a soft scale gate mechanism to deal with multi-granularity ambiguity in 3D segmentation through adjusting the magnitude of each feature channel according to a specified 3D physical scale. Evaluations demonstrate that SAGA achieves real-time multi-granularity segmentation with quality comparable to state-of-the-art methods. As one of the first methods addressing promptable segmentation in 3D-GS, the simplicity and effectiveness of SAGA pave the way for future advancements in this field.
Jiazhong Cen, Jiemin Fang, Chen Yang 0023, Lingxi Xie, Xiaopeng Zhang 0008, Wei Shen 0002, Qi Tian 0001
AAAI6
2025 FATE: Feature-Adapted Parameter Tuning for Vision-Language Models
abstract
Following the recent popularity of vision language models, several attempts, e.g., parameter-efficient fine-tuning (PEFT), have been made to extend them to different downstream tasks. Previous PEFT works motivate their methods from the view of introducing new parameters for adaptation but still need to learn this part of weight from scratch, i.e., random initialization. In this paper, we present a novel strategy that incorporates the potential of prompts, e.g., vision features, to facilitate the initial parameter space adapting to new scenarios. We introduce a Feature-Adapted parameTer Efficient tuning paradigm for vision-language models, dubbed as FATE, which injects informative features from the vision encoder into language encoder's parameters space. Specifically, we extract vision features from the last layer of CLIP's vision encoder and, after projection, treat them as parameters for fine-tuning each layer of CLIP's language encoder. By adjusting these feature-adapted parameters, we can directly enable communication between the vision and language branches, facilitating CLIP's adaptation to different scenarios. Experimental results show that FATE exhibits superior generalization performance on 11 datasets with a very small amount of extra parameters and computation.
Zhengqin Xu, Zelin Peng, Xiaokang Yang 0001, Wei Shen 0002
AAAI4
2025 EndoSD-SLAM: Real-time Deformable Endoscopic SLAM via Sparse-Dense Hybrid Representation
abstract
Accurate tissue tracking and dense reconstruction in endoscopic procedures present fundamental challenges, particularly when addressing complex tissue deformations captured by monocular cameras. Most existing approaches compromise between dynamic tracking accuracy and reconstruction density due to the ill-posed nature of monocular deformable scene re-construction. In this paper, we introduce EndoSD-SLAM, a real-time dense monocular SLAM framework designed for deformable endoscopic scenes that simultaneously achieves precise tracking and high-fidelity dense reconstruction. Our approach leverages a novel hybrid representation: (1) a sparse deformable keypoint map with visco-elastic constraints that enables robust camera tracking under tissue motion, and (2) a dense 3D Gaussian splatting map optimized through differentiable rendering for photorealistic reconstruction. To seamlessly integrate these com-plementary representations, we propose a depth alignment mod-ule that establishes constraints between sparse geometric priors and dense rendering, eliminating the dependency on external depth sensors while maintaining geometric consistency and scale accuracy. Extensive experiments on both synthetic (C3VD) and real-world (EndoMapper, StereoMIS) datasets demonstrate that EndoSD-SLAM significantly outperforms existing methods in reconstruction quality while maintaining real-time performance, showcasing its potential for enhanced intraoperative visualization and surgical guidance. The code will be released soon.
Kailing Wang, Chen Yang 0023, Rong Lin, Xiaokang Yang 0001, Wei Shen 0002
BIBM5
2025 Star with Bilinear Mapping
abstract
Contextual modeling is crucial for robust visual representation learning, especially in computer vision. Although Transformers have become a leading architecture for vision tasks due to their attention mechanism, the quadratic complexity of full attention operations presents substantial computational challenges. To address this, we introduce Star with Bilinear Mapping (SBM), a Transformer-Like architecture that achieves global contextual modeling with linear complexity. SBM employs a bilinear mapping module (BM) with low-rank decomposition strategy and star operations (element-wise multiplication) to efficiently capture global contextual information. Our model demonstrates competitive performance on image classification and semantic segmentation tasks, delivering significant computational efficiency gains compared to traditional attention-based models. Code is available at https://github.com/SJTU-DeepVisionLab/SBM.
Zelin Peng, Zhengqin Xu, Xiaokang Yang 0001, Wei Shen 0002
CVPR7
2025 Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation seeks to label each pixel in an image with arbitrary text descriptions. Vision-language foundation models, especially CLIP, have recently emerged as powerful tools for acquiring open-vocabulary capabilities. However, fine-tuning CLIP to equip it with pixel-level prediction ability often suffers three issues: 1) high computational cost, 2) misalignment between the two inherent modalities of CLIP, and 3) degraded generalization ability on unseen categories. To address these issues, we propose H-CLIP, a symmetrical parameter-efficient fine-tuning (PEFT) strategy conducted in hyperspherical space for both of the two CLIP modalities. Specifically, the PEFT strategy is achieved by a series of efficient block-diagonal learnable transformation matrices and a dual cross-relation communication module among all learnable matrices. Since the PEFT strategy is conducted symmetrically to the two CLIP modalities, the misalignment between them is mitigated. Furthermore, we apply an additional constraint to PEFT on the CLIP text encoder according to the hyperspherical energy principle, i.e., minimizing hyperspherical energy during fine-tuning preserves the intrinsic structure of the original parameter space, to prevent the destruction of the generalization ability offered by the CLIP text encoder. Extensive evaluations across various benchmarks show that H-CLIP achieves new SOTA open-vocabulary semantic segmentation results while only requiring updating approximately 4% of the total parameters of CLIP. The code is available at: https://github.com/SJTU-DeepVisionLab/H-CLIP.
Zelin Peng, Zhengqin Xu, Zhilin Zeng, Wei Shen 0002
CVPR6
2025 Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space
abstract
CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing the text encoder preserves its powerful embeddings, recent studies show that fine-tuning both the text and image encoders jointly significantly enhances segmentation performance, especially for classes from open sets. In this work, we explain this phenomenon from the perspective of hierarchical alignment, since during fine-tuning, the hierarchy level of image embeddings shifts from image-level to pixel-level. We achieve this by leveraging hyperbolic space, which naturally encoders hierarchical structures. Our key observation is that, during fine-tuning, the hyperbolic radius of CLIP’s text embeddings decreases, facilitating better alignment with the pixel-level hierarchical structure of visual data. Building on this insight, we propose HyperCLIP, a novel fine-tuning strategy that adjusts the hyperbolic radius of the text embeddings through scaling transformations. By doing so, HyperCLIP equips CLIP with segmentation capability while introducing only a small number of learnable parameters. Our experiments demonstrate that HyperCLIP achieves state-of-the-art performance on open-vocabulary semantic segmentation tasks across three benchmarks, while fine-tuning only approximately 4% of the total parameters of CLIP. More importantly, we observe that after adjustment, CLIP’s text embeddings exhibit a relatively fixed hyperbolic radius across datasets, suggesting that the granularity required for this segmentation task might be quantified using the hyperbolic radius.
Zelin Peng, Zhengqin Xu, Zhilin Zeng, Changsong Wen, Menglin Yang 0001, Wei Shen 0002
CVPR8
2025 Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
abstract
Multi-modal Large Language Models (MLLMs) have introduced a novel dimension to document understanding, i.e., they endow large language models with visual comprehension capabilities; however, how to design a suitable image-text pre-training task for bridging the visual and language modality in document-level MLLMs remains underexplored. In this study, we introduce a novel visuallanguage alignment method that casts the key issue as a Visual Question Answering with Mask generation (VQA-Mask) task, optimizing two tasks simultaneously: VQA-based text parsing and mask generation. The former allows the model to implicitly align images and text at the semantic level. The latter introduces an additional mask generator (discarded during inference) to explicitly ensure alignment between visual texts within images and their corresponding image regions at a spatially-aware level. Together, they can prevent model hallucinations when parsing visual text and effectively promote spatially-aware feature representation learning. To support the proposed VQAMask task, we construct a comprehensive image-mask generation pipeline and provide a large-scale dataset with 6M data (MTMask6M). Subsequently, we demonstrate that introducing the proposed mask generation task yields competitive document-level understanding performance. Leveraging the proposed VQAMask, we introduce Marten, a trainingefficient MLLM tailored for document-level understanding. Extensive experiments show that our Marten consistently achieves significant improvements among 8B-MLLMs in document-centric tasks. Code and datasets are available at https://github.com/PriNing/Marten.
Tongkun Guan, Pei Fu, Chen Duan, Qianyi Jiang, Zhentao Guo, Shan Guo, Junfeng Luo, Wei Shen 0002, Xiaokang Yang 0001
CVPR9
2025 Domain Generalization in CLIP via Learning with Diverse Text Prompts
abstract
Domain generalization (DG) aims to train a model on source domains that can generalize well to unseen domains. Recent advances in Vision-Language Models (VLMs), such as CLIP, exhibit remarkable generalization capabilities across a wide range of data distributions, benefiting tasks like DG. However, CLIP is pre-trained by aligning images with their descriptions, which inevitably captures domain-specific details. Moreover, adapting CLIP to source domains with limited feature diversity introduces bias. These limitations hinder the model’s ability to generalize across domains. In this paper, we propose a new DG approach by learning with diverse text prompts. These text prompts incorporate varied contexts to imitate different domains, enabling DG model to learn domain-invariant features. The text prompts guide DG model learning in three aspects: feature suppression, which uses these prompts to identify domain-sensitive features and suppress them; feature consistency, which ensures the model’s features are robust to domain variations imitated by the diverse prompts; and feature diversification, which diversifies features based on the prompts to mitigate bias. Experimental results show that our approach improves domain generalization performance on five datasets from the DomainBed benchmark, achieving state-of-the-art results.
Changsong Wen, Zelin Peng, Xiaokang Yang 0001, Wei Shen 0002
CVPR5
2025 MDN: Mamba-Driven Dualstream Network For Medical Hyperspectral Image Segmentation
abstract
Medical Hyperspectral Imaging (MHSI) offers potential for computational pathology and precision medicine. However, existing CNN and Transformer struggle to balance segmentation accuracy and speed due to high spatial-spectral dimensionality. In this study, we leverage Mamba’s global context modeling to propose a dual-stream architecture for joint spatial-spectral feature extraction. To address the limitation of Mamba’s unidirectional aggregation, we introduce a recurrent spectral sequence representation to capture low-redundancy global spectral features. Experiments on a public Multi-Dimensional Choledoch dataset and a private Cervical Cancer dataset show that our method outperforms state-of-the-art approaches in segmentation accuracy while minimizing resource usage and achieving the fastest inference speed. Our code will be available at https://github.com/DeepMed-Lab-ECNU/MDN.
Shijie Lin, Boxiang Yun, Wei Shen 0002, Qingli Li, Anqiang Yang, Yan Wang 0033
ICASSP3
2025 A Token-Level Text Image Foundation Model for Document Understanding
Tongkun Guan, Pei Fu, Zhengtao Guo, Wei Shen 0002, Tiezhu Yue, Chen Duan, Qianyi Jiang, Junfeng Luo, Xiaokang Yang 0001
ICCV5
2025 Generalized Tensor-Based Parameter-Efficient Fine-Tuning via Lie Group Transformations
abstract
Adapting pre-trained foundation models for diverse downstream tasks is a core practice in artificial intelligence. However, the wide range of tasks and high computational costs make full fine-tuning impractical. To overcome this, parameter-efficient fine-tuning (PEFT) methods like LoRA have emerged and are becoming a growing research focus. Despite the success of these methods, they are primarily designed for linear layers, focusing on two-dimensional matrices while largely ignoring higher-dimensional parameter spaces like convolutional kernels. Moreover, directly applying these methods to higher-dimensional parameter spaces often disrupts their structural relationships. Given the rapid advancements in matrix-based PEFT methods, rather than designing a specialized strategy, we propose a generalization that extends matrix-based PEFT methods to higher-dimensional parameter spaces without compromising their structural properties. Specifically, we treat parameters as elements of a Lie group, with updates modeled as perturbations in the corresponding Lie algebra. These perturbations are mapped back to the Lie group through the exponential map, ensuring smooth, consistent updates that preserve the inherent structure of the parameter space. Extensive experiments on computer vision and natural language processing validate the effectiveness and versatility of our approach, demonstrating clear improvements over existing methods.
Chongjie Si, Zhiyi Shi, Yichen Xiao, Xiaokang Yang 0001, Wei Shen 0002
ICCV6
2025 Unleashing the Power of Task-Specific Directions in Parameter Efficient Fine-tuning
abstract
Large language models demonstrate impressive performance on downstream tasks, yet requiring extensive resource consumption when fully fine-tuning all parameters. To mitigate this, Parameter Efficient Fine-Tuning (PEFT) strategies, such as LoRA, have been developed. In this paper, we delve into the concept of task-specific directions (TSDs)—critical for transitioning large models from pretrained states to task-specific enhancements in PEFT. We propose a framework to clearly define these directions and explore their properties, and practical utilization challenges. We then introduce a novel approach, LoRA-Dash, which aims to maximize the impact of TSDs during the fine-tuning process, thereby enhancing model performance on targeted tasks. Extensive experiments have conclusively demonstrated the effectiveness of LoRA-Dash, and in-depth analyses further reveal the underlying mechanisms of LoRA-Dash.
Chongjie Si, Zhiyi Shi, Shifan Zhang, Xiaokang Yang 0001, Hanspeter Pfister, Wei Shen 0002
ICLR6
2025 Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuning
abstract
Adapting pre-trained foundation models for various downstream tasks has been prevalent in artificial intelligence. Due to the vast number of tasks and high costs, adjusting all parameters becomes unfeasible. To mitigate this, several fine-tuning techniques have been developed to update the pre-trained model weights in a more resource-efficient manner, such as through low-rank adjustments. Yet, almost all of these methods focus on linear weights, neglecting the intricacies of parameter spaces in higher dimensions like 4D. Alternatively, some methods can be adapted for high-dimensional parameter space by compressing changes in the original space into two dimensions and then employing low-rank matrix adaptations. However, these approaches destructs the structural integrity of the involved high-dimensional spaces. To tackle the diversity of dimensional spaces across different foundation models and provide a more precise representation of the changes within these spaces, this paper introduces a generalized parameter-efficient fine-tuning framework, designed for various dimensional parameter space. Specifically, our method asserts that changes in each dimensional parameter space are based on a low-rank core space which maintains the consistent topological structure with the original space. It then models the changes through this core space alongside corresponding weights to reconstruct alterations in the original space. It effectively preserves the structural integrity of the change of original N-dimensional parameter space, meanwhile models it via low-rank tensor adaptation. Extensive experiments on computer vision, natural language processing and multi-modal tasks validate the effectiveness of our method.
Chongjie Si, Xue Yang 0005, Zhengqin Xu, Qingyun Li, Jifeng Dai, Yu Qiao 0001, Xiaokang Yang 0001, Wei Shen 0002
ICLR9
2025 Tackling View-Dependent Semantics in 3D Language Gaussian Splatting
abstract
Recent advancements in 3D Gaussian Splatting (3D-GS) enable high-quality 3D scene reconstruction from RGB images. Many studies extend this paradigm for language-driven open-vocabulary scene understanding. However, most of them simply project 2D semantic features onto 3D Gaussians and overlook a fundamental gap between 2D and 3D understanding: a 3D object may exhibit various semantics from different viewpoints—a phenomenon we term **view-dependent semantics**. To address this challenge, we propose **LaGa** (**La**nguage **Ga**ussians), which establishes cross-view semantic connections by decomposing the 3D scene into objects. Then, it constructs view-aggregated semantic representations by clustering semantic descriptors and reweighting them based on multi-view semantics. Extensive experiments demonstrate that LaGa effectively captures key information from view-dependent semantics, enabling a more comprehensive understanding of 3D scenes. Notably, under the same settings, LaGa achieves a significant improvement of **+18.7\% mIoU** over the previous SOTA on the LERF-OVS dataset. Our code is available at: https://github.com/https://github.com/SJTU-DeepVisionLab/LaGa.
Jiazhong Cen, Jiemin Fang, Changsong Wen, Lingxi Xie, Xiaopeng Zhang 0008, Wei Shen 0002, Qi Tian 0001
ICML7
2025 EndoDAV: Depth Any Video in Endoscopy with Spatiotemporal Accuracy
Zanwei Zhou, Chen Yang 0023, Piao Yang, Xiaokang Yang 0001, Wei Shen 0002
MICCAI (9)5
2025 HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
abstract
Multi-modal large language models (MLLMs) have emerged as a transformative approach for aligning visual and textual understanding. They typically require extremely high computational resources (e.g., thousands of GPUs) for training to achieve cross-modal alignment at multi-granularity levels. We argue that a key source of this inefficiency lies in the vision encoders they widely equip with, e.g., CLIP and SAM, which lack the alignment with language at multi-granularity levels. To address this issue, in this paper, we leverage hyperbolic space, which inherently models hierarchical levels and thus provides a principled framework for bridging the granularity gap between visual and textual modalities at an arbitrary granularity level. Concretely, we propose an efficient training paradigm for MLLMs, dubbed as \blg, which can optimize visual representations to align with their textual counterparts at an arbitrary granularity level through dynamic hyperbolic radius adjustment in hyperbolic space. \alg employs learnable matrices with M\"{o}bius multiplication operations, implemented via three effective configurations: diagonal scaling matrices, block-diagonal matrices, and banded matrices, providing a flexible yet efficient parametrization strategy. Comprehensive experiments across multiple MLLM benchmarks demonstrate that \alg consistently improves both existing pre-training and fine-tuning MLLMs clearly with less than 1\% additional parameters. Code is available at \url{https://github.com/godlin-sjtu/HyperET}.
Zelin Peng, Zhengqin Xu, Xiaokang Yang 0001, Wei Shen 0002
NeurIPS5
2025 OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance Information
abstract
Open-vocabulary semantic segmentation assigns every pixel a label drawn from an open-ended, text-defined space. Vision–language models such as CLIP excel at zero-shot recognition, yet their image-level pre-training hinders dense prediction. Current approaches either fine-tune CLIP—at high computational cost—or adopt training-free attention refinements that favor local smoothness while overlooking global semantics. In this paper, we present OPMapper, a lightweight, plug-and-play module that injects both local compactness and global connectivity into attention maps of CLIP. It combines Context-aware Attention Injection, which embeds spatial and semantic correlations, and Semantic Attention Alignment, which iteratively aligns the enriched weights with textual prompts. By jointly modeling token dependencies and leveraging textual guidance, OPMapper enhances visual understanding. OPMapper is highly flexible and can be seamlessly integrated into both training-based and training-free paradigms with minimal computational overhead. Extensive experiments demonstrate its effectiveness, yielding significant improvements across 8 open-vocabulary segmentation benchmarks.
Chongjie Si, Xue Yang 0005, Yuzhi Zhao, Wenhai Wang, Xiaokang Yang 0001, Wei Shen 0002
NeurIPS7
2025 Segment Anything in 3D with Radiance Fields
Jiazhong Cen, Jiemin Fang, Zanwei Zhou, Chen Yang 0023, Lingxi Xie, Xiaopeng Zhang 0008, Wei Shen 0002, Qi Tian 0001
Int. J. Comput. Vis.7
2025 CCDPlus: Towards Accurate Character to Character Distillation for Text Recognition
abstract
Existing scene text recognition methods leverage large-scale labeled synthetic data (LSD) to reduce reliance on labor-intensive annotation tasks and improve recognition capability in real-world scenarios. However, the emergence of a synth-to-real domain gap still limits their efficiency and robustness. Consequently, harvesting the meaningful intrinsic qualities of unlabeled real data (URD) is of great importance, given the prevalence of text-laden images. Toward the target, recent efforts have focused on pre-training on URD through sequence-to-sequence self-supervised learning, followed by fine-tuning on LSD via supervised learning. Nevertheless, they encounter three important issues: coarse representation learning units, inflexible data augmentation, and an emerging real-to-synth domain drift. To overcome these challenges, we propose CCDPlus, an accurate character-to-character distillation method for scene text recognition with a joint supervised and self-supervised learning framework. Specifically, tailored for text images, CCDPlus delineates the fine-grained character structures on URD as representation units by transferring knowledge learned from LSD online. Without requiring extra bounding box or pixel-level annotations, this process allows CCDPlus to enable character-to-character distillation flexibly with versatile data augmentation, which effectively extracts general real-world character-level feature representations. Meanwhile, the unified framework combines self-supervised learning on URD with supervised learning on LSD, effectively solving the domain inconsistency and enhancing the recognition performance. Extensive experiments demonstrate that CCDPlus outperforms previous state-of-the-art (SOTA) supervised, semi-supervised, and self-supervised methods by an average of 1.8%, 0.6%, and 1.1% on standard datasets, respectively. Additionally, it achieves a 6.1% improvement on the more challenging Union14M-L dataset.
Tongkun Guan, Wei Shen 0002, Xiaokang Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 SAM-PARSER: Fine-Tuning SAM Efficiently by Parameter Space Reconstruction
abstract
Segment Anything Model (SAM) has received remarkable attention as it offers a powerful and versatile solution for object segmentation in images. However, fine-tuning SAM for downstream segmentation tasks under different scenarios remains a challenge, as the varied characteristics of different scenarios naturally requires diverse model parameter spaces. Most existing fine-tuning methods attempt to bridge the gaps among different scenarios by introducing a set of new parameters to modify SAM's original parameter space. Unlike these works, in this paper, we propose fine-tuning SAM efficiently by parameter space reconstruction (SAM-PARSER), which introduce nearly zero trainable parameters during fine-tuning. In SAM-PARSER, we assume that SAM's original parameter space is relatively complete, so that its bases are able to reconstruct the parameter space of a new scenario. We obtain the bases by matrix decomposition, and fine-tuning the coefficients to reconstruct the parameter space tailored to the new scenario by an optimal linear combination of the bases. Experimental results show that SAM-PARSER exhibits superior segmentation performance across various scenarios, while reducing the number of trainable parameters by approximately 290 times compared with current parameter-efficient fine-tuning methods.
Zelin Peng, Zhengqin Xu, Zhilin Zeng, Xiaokang Yang 0001, Wei Shen 0002
AAAI5
2024 Partial Label Learning with a Partner
abstract
In partial label learning (PLL), each instance is associated with a set of candidate labels among which only one is ground-truth. The majority of the existing works focuses on constructing robust classifiers to estimate the labeling confidence of candidate labels in order to identify the correct one. However, these methods usually struggle to rectify mislabeled samples. To help existing PLL methods identify and rectify mislabeled samples, in this paper, we introduce a novel partner classifier and propose a novel ``mutual supervision'' paradigm. Specifically, we instantiate the partner classifier predicated on the implicit fact that non-candidate labels of a sample should not be assigned to it, which is inherently accurate and has not been fully investigated in PLL. Furthermore, a novel collaborative term is formulated to link the base classifier and the partner one. During each stage of mutual supervision, both classifiers will blur each other's predictions through a blurring mechanism to prevent overconfidence in a specific label. Extensive experiments demonstrate that the performance and disambiguation ability of several well-established stand-alone and deep-learning based PLL approaches can be significantly improved by coupling with this learning paradigm.
Chongjie Si, Zekun Jiang, Yan Wang 0033, Xiaokang Yang 0001, Wei Shen 0002
AAAI6
2024 DeCo-Net: Robust Multimodal Brain Tumor Segmentation via Decoupled Complementary Knowledge Distillation
abstract
Automated brain tumor segmentation with multimodal magnetic resonance imaging (MRI) plays a pivotal rule in clinical application. However, most existing algorithms require complete image modalities as input, which is often impractical to obtain for every patient in real clinical practice. Therefore, a robust multimodal algorithm that is capable of handling various modality-incomplete data is highly desirable. In this paper, we propose DeCo-Net, a Decoupled Complementary knowledge distillation framework for multimodal brain tumor segmentation with incomplete modalities. Specifically, our approach decouples the feature learning of the modality-incomplete data into two branches: one dedicated to extracting the inherent features from the available modalities and the other focused on inferring the complementary missing modal information. We employ a teacher-student co-training framework where the teacher network is collaboratively trained to dynamically transfer the complementary knowledge to the student model based on the specific type of modality-incomplete data fed to student. To this end, we propose a modality-aware contrastive distillation strategy that guides the student model to distill a discriminative and complementary knowledge representation that acts as supplements to the original modality-incomplete representation. Extensive evaluations on the BraTS2018, BraTS2020 and BraTS2023 datasets demonstrate that our method achieves state-of-the-art performance in multimodal brain tumor segmentation with incomplete modalities.
Zhilin Zeng, Zelin Peng, Xiaokang Yang 0001, Wei Shen 0002
BIBM4
2024 Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model
abstract
Parameter-efficient fine-tuning (PEFT) is an effective methodology to unleash the potential of large foundation models in novel scenarios with limited training data. In the computer vision community, PEFT has shown effectiveness in image classification, but little research has studied its ability for image segmentation. Fine-tuning segmentation models usually requires a heavier adjustment of parameters to align the proper projection directions in the parameter space for new scenarios. This raises a challenge to existing PEFT algorithms, as they often inject a limited number of individual parameters into each block, which prevents substantial adjustment of the projection direction of the parameter space due to the limitation of Hidden Markov Chain along blocks. In this paper, we equip PEFT with a cross-block orchestration mechanism to enable the adaptation of the Segment Anything Model (SAM) to various downstream scenarios. We introduce a novel inter-block communication module, which integrates a learnable relation matrix to facilitate communication among different coefficient sets of each PEFT block's parameter space. Moreover, we propose an intra-block enhancement module, which introduces a linear projection head whose weights are generated from a hyper-complex layer, further enhancing the impact of the adjustment of projection directions on the entire parameter space. Extensive experiments on diverse benchmarks demonstrate that our proposed approach consistently improves the segmentation performance significantly on novel scenarios with only around 1K additional parameters.
Zelin Peng, Zhengqin Xu, Zhilin Zeng, Lingxi Xie, Qi Tian 0001, Wei Shen 0002
CVPR6
2024 UniProcessor: A Text-Induced Unified Low-Level Image Processor
Huiyu Duan, Xiongkuo Min, Sijing Wu, Wei Shen 0002, Guangtao Zhai
ECCV (67)4
2024 PosFormer: Recognizing Complex Handwritten Mathematical Expression with Position Forest Transformer
Tongkun Guan, Chengyu Lin 0002, Wei Shen 0002, Xiaokang Yang 0001
ECCV (22)3
2024 Bridging Synthetic and Real Worlds for Pre-Training Scene Text Detectors
Tongkun Guan, Wei Shen 0002, Xue Yang 0005, Xiaokang Yang 0001
ECCV (44)2
2024 Tendency-Driven Mutual Exclusivity for Weakly Supervised Incremental Semantic Segmentation
Chongjie Si, Xiaokang Yang 0001, Wei Shen 0002
ECCV (35)4
2024 Domain-Adaptive Semantic Segmentation Emerges From Vision-Language Supervised Domain-Debiased Self-Training
abstract
Unsupervised domain adaptive semantic segmentation leverages synthetic data to train a segmentation model and transfers it to unlabeled real images. Due to the style difference, the transferred model suffers from the domain gap. Even worse, some classes exhibit the extreme domain gap, where the feature distributions undergo a complete shift between the two domains. To alleviate it, we propose a domain-debiased self-training strategy with CLIP to distill its domain-agnostic knowledge. Specifically, we enforce the consistency between the feature maps from our segmentation model and the image encoder of CLIP. Meanwhile, the text embeddings from the text encoder for each class serve as a domain-agnostic classifier to support a domain-debiased feature learning condition. Experimental results under standard UDA settings demonstrate that our proposed strategy consistently improves the UDA segmentation performance based on different backbones and with different large pre-trained models.
Huayu Wang, Zekun Jiang, Lingxi Xie, Dongsheng Jiang, Wei Shen 0002, Qi Tian 0001
ICASSP5
2024 EndoGSLAM: Real-Time Dense Reconstruction and Tracking in Endoscopic Surgeries Using Gaussian Splatting
Kailing Wang, Chen Yang 0023, Yuehao Wang, Sikuang Li, Yan Wang 0033, Qi Dou 0001, Xiaokang Yang 0001, Wei Shen 0002
MICCAI (6)8
2024 Missing as Masking: Arbitrary Cross-Modal Feature Reconstruction for Incomplete Multimodal Brain Tumor Segmentation
Zhilin Zeng, Zelin Peng, Xiaokang Yang 0001, Wei Shen 0002
MICCAI (8)4
2024 Adaptive feature alignment for adversarial training
abstract
Recent studies reveal that Convolutional Neural Networks (CNNs) are typically vulnerable to adversarial attacks . Many adversarial defense methods have been proposed to improve the robustness against adversarial samples. Moreover, these methods can only defend adversarial samples of a specific strength, reducing their flexibility against attacks of varying strengths. Moreover, these methods often enhance adversarial robustness at the expense of accuracy on clean samples. In this paper, we first observed that features of adversarial images change monotonically and smoothly w.r.t the rising of attacking strength. This intriguing observation suggests that features of adversarial images with various attacking strengths can be approximated by interpolating between the features of adversarial images with the strongest and weakest attacking strengths. Due to the monotonicity property, the interpolation weight can be easily learned by a neural network . Based on the observation, we proposed the adaptive feature alignment (AFA) that automatically align features to defense adversarial attacks of various attacking strengths. During training, our method learns the statistical information of adversarial samples with various attacking strengths using a dual batchnorm architecture. In this architecture, each batchnorm process handles samples of a specific attacking strength. During inference, our method automatically adjusts to varying attacking strengths by linearly interpolating the dual-BN features. Unlike previous methods that need to either retrain the model or manually tune hyper-parameters for a new attacking strength, our method can deal with arbitrary attacking strengths with a single model without introducing any hyper-parameter. Additionally, our method improves the model robustness against adversarial samples without incurring much loss of accuracy on clean images. Experiments on CIFAR-10, SVHN and tiny-ImageNet datasets demonstrate that our method outperforms the state-of-the-art under various attacking strengths and even improve accuracy on clean samples. Code will be made open available upon acceptance.
Kai Zhao 0012, Wei Shen 0002
Pattern Recognit. Lett.4
2024 Consensus Synergizes With Memory: A Simple Approach for Anomaly Segmentation in Urban Scenes
abstract
Anomaly segmentation is a critical task for safety-critical applications, such as autonomous driving in urban environments. Its objective is to detect out-of-distribution (OOD) samples with unseen categories, given a pre-trained segmentation model. The core challenge of this task is how to distinguish hard in-distribution samples from OOD samples, which has not been explicitly discussed in previous research. In this paper, we propose a simple yet effective approach named CosMe (Consensus Synergizes with Memory) to address this challenge. CosMe consists of two key components: 1) building a memory bank comprising seen prototypes extracted from multiple layers of the given segmentation model, and 2) training an auxiliary model that mimics the behavior of the given model and using the consensus of their mid-level features as complementary cues that synergize with the memory bank. The former serves as a baseline that can detect all potential outliers, including both OOD and hard in-distribution samples; the latter assists in distinguishing between these two types of outliers. Experimental results on several urban scene anomaly segmentation datasets demonstrate that CosMe outperforms previous approaches by a significant margin.
Jiazhong Cen, Zekun Jiang, Lingxi Xie, Dongsheng Jiang, Wei Shen 0002, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Un-Gaze: A Unified Transformer for Joint Gaze-Location and Gaze-Object Detection
abstract
This paper proposes an efficient and effective method for joint gaze location detection (GL-D) and gaze object detection (GO-D), i.e., gaze following detection. Current approaches frame GL-D and GO-D as two separate tasks, employing a multi-stage framework where human head crops must first be detected and then be fed into a subsequent GL-D sub-network, which is further followed by an additional object detector for GO-D. In contrast, we reframe the gaze following detection task as detecting human head locations and their gaze followings simultaneously, aiming at jointly detect human gaze location and gaze object in a unified and single-stage pipeline. To this end, we propose GTR, short for Gaze following detection TRansformer, streamlining the gaze following detection pipeline by eliminating all additional components, leading to the first unified paradigm that unites GL-D and GO-D in a fully end-to-end manner. GTR enables an iterative interaction between holistic semantics and human head features through a hierarchical structure, inferring the relations of salient objects and human gaze from the global image context and resulting in an impressive accuracy. Concretely, GTR achieves a 12.1 mAP gain ($\mathbf {25.1}\%$) on GazeFollowing and a 18.2 mAP gain ($\mathbf {43.3\%}$) on VideoAttentionTarget for GL-D, as well as a 19 mAP improvement ($\mathbf {45.2\%}$) on GOO-Real for GO-D. Meanwhile, unlike existing systems detecting gaze following sequentially due to the need for a human head as input, GTR has the flexibility to comprehend any number of people’s gaze followings simultaneously, resulting in high efficiency. Specifically, GTR introduces over a$\times 9$improvement in FPS and the relative gap becomes more pronounced as the human number grows.
Danyang Tu, Wei Shen 0002, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.2
2024 SpecTr: Spectral Transformer for Microscopic Hyperspectral Pathology Image Segmentation
abstract
Hyperspectral imaging (HSI) unlocks the huge potential to a wide variety of applications relying on high-precision pathology image segmentation, such as computational pathology. It can acquire biochemical properties even invisible to naked eyes from histological specimens. Since 1) spectra contain discriminative and continuous patterns for differentiating tissues/cells, and 2) the discriminability of spectra relies on both fine-grained relations in the high-resolution spectrum and coarse relations in the low-resolution spectrum, the key to achieving high-precision hyperspectral pathology image segmentation is to felicitously model the intra- and inter-scale context especially for spectra. In this paper, we propose a spectral transformer (SpecTr) for hyperspectral pathology image segmentation, which first captures global context for intra-scale spectral features, and subsequently extract coarse and fine-grained discriminative spectral information from inter-scale features, respectively. To learn intra-scale spectral context, we propose a Spectral Attentive Module (SAM). Unlike the existing Transformer model that is designed for modalities such as natural images, our proposed SAM is efficient in capturing sparse and pivotal spectral context while avoiding the heterogeneous underlying distributions and noises of different bands. Besides, to reduce the computational complexity of the HSI segmentation model, we further propose a global-local attention module to effectively learn a condensed spectral feature. Experiments show that HSIs can become a more powerful image modality for understanding microscopic pathology images than RGB images, and the proposed SpecTr outperforms other competing methods for hyperspectral pathology image segmentation, with an improvement of 3% compared with the popular 3D-nnUNet and other transformer-based methods. Our code is available at https://github.com/DeepMed-Lab-ECNU/SpecTr.
Boxiang Yun, Bai Ying Lei, Jieneng Chen, Song Qiu, Wei Shen 0002, Qingli Li, Yan Wang 0033
IEEE Trans. Circuits Syst. Video Technol.6
2024 Efficient Deformable Tissue Reconstruction via Orthogonal Neural Plane
abstract
Intraoperative imaging techniques for reconstructing deformable tissues in vivo are pivotal for advanced surgical systems. Existing methods either compromise on rendering quality or are excessively computationally intensive, often demanding dozens of hours to perform, which significantly hinders their practical application. In this paper, we introduce Fast Orthogonal Plane (Forplane), a novel, efficient framework based on neural radiance fields (NeRF) for the reconstruction of deformable tissues. We conceptualize surgical procedures as 4D volumes, and break them down into static and dynamic fields comprised of orthogonal neural planes. This factorization discretizes the four-dimensional space, leading to a decreased memory usage and faster optimization. A spatiotemporal importance sampling scheme is introduced to improve performance in regions with tool occlusion as well as large motions and accelerate training. An efficient ray marching method is applied to skip sampling among empty regions, significantly improving inference speed. Forplane accommodates both binocular and monocular endoscopy videos, demonstrating its extensive applicability and flexibility. Our experiments, carried out on two in vivo datasets, the EndoNeRF and Hamlyn datasets, demonstrate the effectiveness of our framework. In all cases, Forplane substantially accelerates both the optimization process (by over 100 times) and the inference process (by over 15 times) while maintaining or even improving the quality across a variety of non-rigid deformations. This significant performance improvement promises to be a valuable asset for future intraoperative surgical applications. The code of our project is now available at https://github.com/Loping151/ForPlane.
Chen Yang 0023, Kailing Wang, Yuehao Wang, Qi Dou 0001, Xiaokang Yang 0001, Wei Shen 0002
IEEE Trans. Medical Imaging6
2024 GaussianObject: High-Quality 3D Object Reconstruction from Four Views with Gaussian Splatting
abstract
Reconstructing and rendering 3D objects from highly sparse views is of critical importance for promoting applications of 3D vision techniques and improving user experience. However, images from sparse views only contain very limited 3D information, leading to two significant challenges: 1) Difficulty in building multi-view consistency as images for matching are too few; 2) Partially omitted or highly compressed object information as view coverage is insufficient. To tackle these challenges, we propose GaussianObject, a framework to represent and render the 3D object with Gaussian splatting that achieves high rendering quality with only 4 input images. We first introduce techniques of visual hull and floater elimination, which explicitly inject structure priors into the initial optimization process to help build multi-view consistency, yielding a coarse 3D Gaussian representation. Then we construct a Gaussian repair model based on diffusion models to supplement the omitted object information, where Gaussians are further refined. We design a self-generating strategy to obtain image pairs for training the repair model. We further design a COLMAP-free variant, where pre-given accurate camera poses are not required, which achieves competitive quality and facilitates wider applications. GaussianObject is evaluated on several challenging datasets, including MipNeRF360, OmniObject3D, OpenIllumination, and our-collected unposed images, achieving superior performance from only four views and significantly outperforming previous SOTA methods.
Chen Yang 0023, Sikuang Li, Jiemin Fang, Ruofan Liang, Lingxi Xie, Xiaopeng Zhang 0008, Wei Shen 0002, Qi Tian 0001
ACM Trans. Graph.7
2023 Intriguing Findings of Frequency Selection for Image Deblurring
abstract
Blur was naturally analyzed in the frequency domain, by estimating the latent sharp image and the blur kernel given a blurry image. Recent progress on image deblurring always designs end-to-end architectures and aims at learning the difference between blurry and sharp image pairs from pixel-level, which inevitably overlooks the importance of blur kernels. This paper reveals an intriguing phenomenon that simply applying ReLU operation on the frequency domain of a blur image followed by inverse Fourier transform, i.e., frequency selection, provides faithful information about the blur pattern (e.g., the blur direction and blur level, implicitly shows the kernel pattern). Based on this observation, we attempt to leverage kernel-level information for image deblurring networks by inserting Fourier transform, ReLU operation, and inverse Fourier transform to the standard ResBlock. 1 × 1 convolution is further added to let the network modulate flexible thresholds for frequency selection. We term our newly built block as Res FFT-ReLU Block, which takes advantages of both kernel-level and pixel-level features via learning frequency-spatial dual-domain representations. Extensive experiments are conducted to acquire a thorough analysis on the insights of the method. Moreover, after plugging the proposed block into NAFNet, we can achieve 33.85 dB in PSNR on GoPro dataset. Our method noticeably improves backbone architectures without introducing many parameters, while maintaining low computational complexity. Code is available at https://github.com/DeepMed-Lab/DeepRFT-AAAI2023.
Xintian Mao, Fengze Liu, Qingli Li, Wei Shen 0002, Yan Wang 0033
AAAI5
2023 Bidirectional Copy-Paste for Semi-Supervised Medical Image Segmentation
abstract
In semi-supervised medical image segmentation, there exist empirical mismatch problems between labeled and un-labeled data distribution. The knowledge learned from the labeled data may be largely discarded if treating labeled and unlabeled data separately or in an inconsistent manner. We propose a straightforward method for alleviating the problem-copy-pasting labeled and unlabeled data bidirectionally, in a simple Mean Teacher architecture. The method encourages unlabeled data to learn comprehensive common semantics from the labeled data in both inward and outward directions. More importantly, the consistent learning procedure for labeled and unlabeled data can largely reduce the empirical distribution gap. In detail, we copy-paste a random crop from a labeled image (foreground) onto an unlabeled image (background) and an unlabeled image (foreground) onto a labeled image (background), respectively. The two mixed images are fed into a Student network and supervised by the mixed supervisory signals of pseudo-labels and ground-truth. We reveal that the simple mechanism of copy-pasting bidirectionally between labeled and unlabeled data is good enough and the experiments show solid gains (e.g., over 21% Dice improvement on ACDC dataset with 5% labeled data) compared with other state-of-the-arts on various semi-supervised medical image segmentation datasets. Code is avaiable at https://github.com/DeepMed-Lab-ECNU/BCP.
Yunhao Bai, Duowen Chen 0002, Qingli Li, Wei Shen 0002, Yan Wang 0033
CVPR4
2023 MagicNet: Semi-Supervised Multi-Organ Segmentation via Magic-Cube Partition and Recovery
abstract
We propose a novel teacher-student model for semi-supervised multi-organ segmentation. In teacher-student model, data augmentation is usually adopted on unlabeled data to regularize the consistent training between teacher and student. We start from a key perspective that fixed relative locationsand variable sizes of different organs can provide distribution information where a multi-organ CT scan is drawn. Thus, we treat the prior anatomy as a strong tool to guide the data augmentation and reduce the mismatch between labeled and unlabeled images for semi-supervised learning. More specifically, we propose a data augmentation strategy based on partition-and-recovery N3cubes cross-and within-labeled and unlabeled images. Our strategy encourages unlabeled images to learn organ semantics in relative locations from the labeled images (cross-branch) and enhances the learning ability for small organs (within-branch). For within-branch, we further propose to refine the quality of pseudo labels by blending the learned representations from small cubes to incorporate local attributes. Our method is termed as MagicNet, since it treats the CT volume as a magic-cube and N3-cube partition-and-recovery process matches with the rule of playing a magic-cube. Extensive experiments on two public CT multi-organ datasets demonstrate the effectiveness of MagicNet, and noticeably outperforms state-of-the-art semi-supervised medical image segmentation approaches, with + 7% DSC improvement on MACT dataset with 10% labeled images. Code is avaiable at https://github.com/DeepMed-Lab-ECNU/MagicNet.
Duowen Chen 0002, Yunhao Bai, Wei Shen 0002, Qingli Li, Lequan Yu, Yan Wang 0033
CVPR3
2023 Self-Supervised Implicit Glyph Attention for Text Recognition
abstract
The attention mechanism has become the de facto module in scene text recognition (STR) methods, due to its capability of extracting character-level representations. These methods can be summarized into implicit attention based and supervised attention based, depended on how the attention is computed, i.e., implicit attention and supervised attention are learned from sequence-level text annotations and or character-level bounding box annotations, respectively. Implicit attention, as it may extract coarse or even incorrect spatial regions as character attention, is prone to suffering from an alignment-drifted issue. Supervised attention can alleviate the above issue, but it is character category-specific, which requires extra laborious character-level bounding box annotations and would be memory-intensive when handling languages with larger character categories. To address the aforementioned issues, we propose a novel attention mechanism for STR, self-supervised implicit glyph attention (SICA). SICA delineates the glyph structures of text images by jointly self-supervised text seg-mentation and implicit attention alignment, which serve as the supervision to improve attention correctness without extra character-level annotations. Experimental results demonstrate that SIGA performs consistently and significantly better than previous attention-based STR methods, in terms of both attention correctness and final recognition performance on publicly available context benchmarks and our contributed contextless benchmarks.
Tongkun Guan, Chaochen Gu, Jingzheng Tu, Xue Yang 0005, Yudi Zhao, Wei Shen 0002
CVPR7
2023 NeRFVS: Neural Radiance Fields for Free View Synthesis via Geometry Scaffolds
abstract
We present NeRFVS, a novel neural radiance fields (NeRF) based method to enable free navigation in a room. NeRF achieves impressive performance in rendering images for novel views similar to the input views while suffering for novel views that are significantly different from the training views. To address this issue, we utilize the holistic priors, including pseudo depth maps and view coverage information, from neural reconstruction to guide the learning of implicit neural representations of 3D indoor scenes. Concretely, an off-the-shelf neural reconstruction method is leveraged to generate a geometry scaffold. Then, two loss functions based on the holistic priors are proposed to improve the learning of NeRF: 1) A robust depth loss that can tolerate the error of the pseudo depth map to guide the geometry learning of NeRF; 2) A variance loss to regularize the variance of implicit neural representations to reduce the geometry and color ambiguity in the learning procedure. These two loss functions are modulated during NeRF optimization according to the view coverage information to reduce the negative influence brought by the view coverage imbalance. Extensive results demonstrate that our NeRFVS outperforms state-of-the-art view synthesis methods quantitatively and qualitatively on indoor scenes, achieving high-fidelity free navigation results.
Chen Yang 0023, Peihao Li 0003, Zanwei Zhou, Shanxin Yuan, Xiaokang Yang 0001, Weichao Qiu, Wei Shen 0002
CVPR8
2023 Self-supervised Character-to-Character Distillation for Text Recognition
abstract
When handling complicated text images (e.g., irregular structures, low resolution, heavy occlusion, and uneven illumination), existing supervised text recognition methods are data-hungry. Although these methods employ large-scale synthetic text images to reduce the dependence on annotated real images, the domain gap still limits the recognition performance. Therefore, exploring the robust text feature representations on unlabeled real images by self-supervised learning is a good solution. However, existing self-supervised text recognition methods conduct sequence-to-sequence representation learning by roughly splitting the visual features along the horizontal axis, which limits the flexibility of the augmentations, as large geometric-based augmentations may lead to sequence-to-sequence feature inconsistency. Motivated by this, we propose a novel self-supervised Character-to-Character Distillation method, CCD, which enables versatile augmentations to facilitate general text representation learning. Specifically, we delineate the character structures of unlabeled real images by designing a self-supervised character segmentation module. Following this, CCD easily enriches the diversity of local characters while keeping their pairwise alignment under flexible augmentations, using the transformation matrix between two augmented views from images. Experiments demonstrate that CCD achieves state-of-the-art results, with average performance gains of 1.38% in text recognition, 1.7% in text segmentation, 0.24 dB (PSNR) and 0.0321 (SSIM) in text super-resolution. Code is available at https://github.com/TongkunGuan/CCD.
Tongkun Guan, Wei Shen 0002, Xue Yang 0005, Zekun Jiang, Xiaokang Yang 0001
ICCV2
2023 USAGE: A Unified Seed Area Generation Paradigm for Weakly Supervised Semantic Segmentation
abstract
Seed area generation is usually the starting point of weakly supervised semantic segmentation (WSSS). Computing the Class Activation Map (CAM) from a multi-label classification network is the de facto paradigm for seed area generation, but CAMs generated from Convolutional Neural Networks (CNNs) and Transformers are prone to be under- and over-activated, respectively, which makes the strategies to refine CAMs for CNNs usually inappropriate for Transformers, and vice versa. In this paper, we propose a Unified optimization paradigm for Seed Area GEneration (USAGE) for both types of networks, in which the objective function to be optimized consists of two terms: One is a generation loss, which controls the shape of seed areas by a temperature parameter following a deterministic principle for different types of networks; The other is a regularization loss, which ensures the consistency between the seed areas that are generated by self-adaptive network adjustment from different views, to overturn false activation in seed areas. Experimental results show that USAGE consistently improves seed area generation for both CNNs and Transformers by large margins, e.g., outperforming state-of-the-art methods by a mIoU of 4.1% on PASCAL VOC. Moreover, based on the USAGE-generated seed areas on Transformers, we achieve state-of-the-art WSSS results on both PASCAL VOC and MS COCO.
Zelin Peng, Guanchun Wang, Lingxi Xie, Dongsheng Jiang, Wei Shen 0002, Qi Tian 0001
ICCV5
2023 Agglomerative Transformer for Human-Object Interaction Detection
abstract
We propose an agglomerative Transformer (AGER) that enables Transformer-based human-object interaction (HOI) detectors to flexibly exploit extra instance-level cues in a single-stage and end-to-end manner for the first time. AGER acquires instance tokens by dynamically clustering patch tokens and aligning cluster centers to instances with textual guidance, thus enjoying two benefits: 1) Integrality: each instance token is encouraged to contain all discriminative feature regions of an instance, which demonstrates a significant improvement in the extraction of different instance-level cues and subsequently leads to a new state-of-the-art performance of HOI detection with 36.75 mAP on HICO-Det. 2) Efficiency: the dynamical clustering mechanism allows AGER to generate instance tokens jointly with the feature learning of the Transformer encoder, eliminating the need of an additional object detector or instance decoder in prior methods, thus allowing the extraction of desirable extra cues for HOI detection in a single-stage and end-to-end pipeline. Concretely, AGER reduces GFLOPs by 8.5% and improves FPS by 36%, even compared to a vanilla DETR-like pipeline without extra cue extraction. The code will be available at https://github.com/six6607/AGER.git.
Danyang Tu, Wei Sun 0029, Guangtao Zhai, Wei Shen 0002
ICCV4
2023 Neural LerPlane Representations for Fast 4D Reconstruction of Deformable Tissues
Chen Yang 0023, Kailing Wang, Yuehao Wang, Xiaokang Yang 0001, Wei Shen 0002
MICCAI (9)5
2023 Segment Anything in 3D with NeRFs
abstract
Recently, the Segment Anything Model (SAM) emerged as a powerful vision foundation model which is capable to segment anything in 2D images. This paper aims to generalize SAM to segment 3D objects. Rather than replicating the data acquisition and annotation procedure which is costly in 3D, we design an efficient solution, leveraging the Neural Radiance Field (NeRF) as a cheap and off-the-shelf prior that connects multi-view 2D images to the 3D space. We refer to the proposed solution as SA3D, for Segment Anything in 3D. It is only required to provide a manual segmentation prompt (e.g., rough points) for the target object in a single view, which is used to generate its 2D mask in this view with SAM. Next, SA3D alternately performs mask inverse rendering and cross-view self-prompting across various views to iteratively complete the 3D mask of the target object constructed with voxel grids. The former projects the 2D mask obtained by SAM in the current view onto 3D mask with guidance of the density distribution learned by the NeRF; The latter extracts reliable prompts automatically as the input to SAM from the NeRF-rendered 2D mask in another view. We show in experiments that SA3D adapts to various scenes and achieves 3D segmentation within minutes. Our research offers a generic and efficient methodology to lift a 2D vision foundation model to 3D, as long as the 2D model can steadily address promptable segmentation across multiple views.
Jiazhong Cen, Zanwei Zhou, Jiemin Fang, Chen Yang 0023, Wei Shen 0002, Lingxi Xie, Dongsheng Jiang, Xiaopeng Zhang 0008, Qi Tian 0001
NeurIPS5
2023 A Survey on Label-Efficient Deep Image Segmentation: Bridging the Gap Between Weak Supervision and Dense Prediction
abstract
The rapid development of deep learning has made a great progress in image segmentation, one of the fundamental tasks of computer vision. However, the current segmentation algorithms mostly rely on the availability of pixel-level annotations, which are often expensive, tedious, and laborious. To alleviate this burden, the past years have witnessed an increasing attention in building label-efficient, deep-learning-based image segmentation algorithms. This paper offers a comprehensive review on label-efficient image segmentation methods. To this end, we first develop a taxonomy to organize these methods according to the supervision provided by different types of weak labels (including no supervision, inexact supervision, incomplete supervision and inaccurate supervision) and supplemented by the types of segmentation problems (including semantic segmentation, instance segmentation and panoptic segmentation). Next, we summarize the existing label-efficient image segmentation methods from a unified perspective that discusses an important question: how to bridge the gap between weak supervision and dense prediction - the current methods are mostly based on heuristic priors, such as cross-pixel similarity, cross-label constraint, cross-view consistency, and cross-image relation. Finally, we share our opinions about the future research directions for label-efficient deep image segmentation.
Wei Shen 0002, Zelin Peng, Huayu Wang, Jiazhong Cen, Dongsheng Jiang, Lingxi Xie, Xiaokang Yang 0001, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 BNET: Batch Normalization With Enhanced Linear Transformation
abstract
Batch normalization (BN) is a fundamental unit in modern deep neural networks. However, BN and its variants focus on normalization statistics but neglect the recovery step that uses linear transformation to improve the capacity of fitting complex data distributions. In this paper, we demonstrate that the recovery step can be improved by aggregating the neighborhood of each neuron rather than just considering a single neuron. Specifically, we propose a simple yet effective method named batch normalization with enhanced linear transformation (BNET) to embed spatial contextual information and improve representation ability. BNET can be easily implemented using the depth-wise convolution and seamlessly transplanted into existing architectures with BN. To our best knowledge, BNET is the first attempt to enhance the recovery step for BN. Furthermore, BN is interpreted as a special case of BNET from both spatial and spectral views. Experimental results demonstrate that BNET achieves consistent performance gains based on various backbones in a wide range of visual tasks. Moreover, BNET can accelerate the convergence of network training and enhance spatial information by assigning important neurons with large weights accordingly.
Yuhui Xu 0002, Lingxi Xie, Cihang Xie, Wenrui Dai, Jieru Mei, Siyuan Qiao, Wei Shen 0002, Hongkai Xiong, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 Develop Then Rival: A Human Vision-Inspired Framework for Superimposed Image Decomposition
abstract
A single superimposed image containing two image views causes visual confusion for both human vision and computer vision. Human vision needs a “develop-then-rival” process to decompose the superimposed image into two individual images, which effectively suppresses visual confusion. However, separating individual image views from a single superimposed image has been an important but challenging task in computer vision area for a long time. In this paper, we propose a human vision-inspired framework for single superimposed image decomposition. We first propose a network to simulate the development stage, which tries to understand and distinguish the semantic information of the two layers of a single superimposed image. To further simulate the rivalry activation/suppression process in human brains, we carefully design a rivalry stage, which incorporates the original mixed input (superimposed image), the activated visual information (outputs of the development stage) together, and then rivals to get images without ambiguity. Experimental results show that our novel framework effectively separates the superimposed images and significantly improves the performance with better output quality compared with state-of-the-art methods. The proposed method also achieves state-of-the-art results on related applications including single image reflection removal, single image rain removal, single image shadow removal, and illumination correction,etc., which validates the generalization of the framework.
Huiyu Duan, Wei Shen 0002, Xiongkuo Min, Yuan Tian 0017, Jae-Hyun Jung, Xiaokang Yang 0001, Guangtao Zhai
IEEE Trans. Multim.2
2023 ChildPredictor: A Child Face Prediction Framework With Disentangled Learning
abstract
The appearances of children are inherited from their parents, which makes it feasible to predict them. Predicting realistic children's faces may help settle many social problems, such as age-invariant face recognition, kinship verification, and missing child identification. It can be regarded as an image-to-image translation task. Existing approaches usually assume domain information in the image-to-image translation can be interpreted by “style”, i.e., the separation of image content and style. However, such separation is improper for the child face prediction, because the facial contours between children and parents are not the same. To address this issue, we propose a new disentangled learning strategy for children's face prediction. We assume that children's faces are determined by genetic factors (compact family features, e.g., face contour), external factors (facial attributes irrelevant to prediction, such as moustaches and glasses), and variety factors (individual properties for each child). On this basis, we formulate predictions as a mapping from parents’ genetic factors to children's genetic factors, and disentangle them from external and variety factors. In order to obtain accurate genetic factors and perform the mapping, we propose a ChildPredictor framework. It transfers human faces to genetic factors by encoders and back by generators. Then, it learns the relationship between the genetic factors of parents and children through a mapping function. To ensure the generated faces are realistic, we collect a large Family Face Database to train ChildPredictor and evaluate it on the FF-Database validation set. Experimental results demonstrate that ChildPredictor is superior to other well-known image-to-image translation methods in predicting realistic and diverse child faces. Implementation codes can be found athttps://github.com/zhaoyuzhi/ChildPredictor.
Yuzhi Zhao, Lai-Man Po, Qiong Yan, Wei Shen 0002, Yujia Zhang 0002, Wei Liu 0004, Chun Kit Wong, Chiu-Sing Pang, Weifeng Ou, Wing Yin Yu, Buhua Liu
IEEE Trans. Multim.5
2022 End-to-End Human-Gaze-Target Detection with Transformers
abstract
In this paper, we propose an effective and efficient method for Human-Gaze-Target (HGT) detection, i.e., gaze following. Current approaches decouple the HGT detection task into separate branches of salient object detection and human gaze prediction, employing a two-stage framework where human head locations must first be detected and then be fed into the next gaze target prediction sub-network. In contrast, we redefine the HGT detection task as detecting human head locations and their gaze targets, simultaneously. By this way, our method, named Human-Gaze-Target detection TRansformer or HGTTR, streamlines the HGT detection pipeline by eliminating all other additional components. HGTTR reasons about the relations of salient objects and human gaze from the global image context. Moreover, unlike existing two-stage methods that require human head locations as input and can predict only one human's gaze target at a time, HGTTR can directly predict the locations of all people and their gaze targets at one time in an end-to-end manner. The effectiveness and robustness of our proposed method are verified with extensive experiments on the two standard benchmark datasets, GazeFollowing and VideoAttentionTarget. Without bells and whistles, HGTTR outperforms existing state-of-the-art methods by large margins (6.4 mAP gain on GazeFollowing and 10.3 mAP gain on VideoAttentionTarget) with a much simpler architecture.
Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen 0002
CVPR6
2022 ContrastMask: Contrastive Learning to Segment Every Thing
abstract
Partially-supervised instance segmentation is a task which requests segmenting objects from novel categories via learning on limited base categories with annotated masks thus eliminating demands of heavy annotation burden. The key to addressing this task is to build an effective class-agnostic mask segmentation model. Unlike previous methods that learn such models only on base categories, in this paper, we propose a new method, named ContrastMask, which learns a mask segmentation model on both base and novel categories under a unified pixel-level contrastive learning framework. In this framework, annotated masks of base categories and pseudo masks of novel categories serve as a prior for contrastive learning, where features from the mask regions (foreground) are pulled together, and are contrasted against those from the background, and vice versa. Through this framework, feature discrimination between foreground and background is largely improved, facilitating learning of the class-agnostic mask segmentation model. Exhaustive experiments on the COCO dataset demonstrate the superiority of our method, which outperforms previous state-of-the-arts.
Kai Zhao 0012, Shouhong Ding, Yan Wang 0033, Wei Shen 0002
CVPR6
2022 Iwin: Human-Object Interaction Detection via Transformer with Irregular Windows
Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen 0002
ECCV (4)6
2022 CP2: Copy-Paste Contrastive Pretraining for Semantic Segmentation
Feng Wang 0047, Chen Wei 0005, Alan L. Yuille, Wei Shen 0002
ECCV (30)5
2022 BézierPalm: A Free Lunch for Palmprint Recognition
Kai Zhao 0012, Chuhan Zhou, Shouhong Ding, Wei Jia 0001, Wei Shen 0002
ECCV (13)9
2022 A Unified Two-Stage Model for Separating Superimposed Images
abstract
A single superimposed image containing two image views causes visual confusion for both human vision and computer vision. Human vision needs a "develop-then-rival" process to decompose the superimposed image into two individual images, which effectively suppresses visual confusion. In this paper, we propose a human vision-inspired framework for separating superimposed images. We first propose a network to simulate the development stage, which tries to understand and distinguish the semantic information of the two layers of a single superimposed image. To further simulate the rivalry activation/suppression process in human brains, we carefully design a rivalry stage, which incorporates the original mixed input (superimposed image), the activated visual information (outputs of the development stage) together, and then rivals to get images without ambiguity. Experimental results show that our novel framework effectively separates the superimposed images and significantly improves the performance with better output quality compared with state-of-the-art methods.
Huiyu Duan, Xiongkuo Min, Wei Shen 0002, Guangtao Zhai
ICASSP3
2022 Image BERT Pre-training with Online Tokenizer
Jinghao Zhou, Chen Wei 0005, Wei Shen 0002, Cihang Xie, Alan L. Yuille, Tao Kong
ICLR4
2022 Saliency in Augmented Reality
abstract
With the rapid development of multimedia technology, Augmented Reality (AR) has become a promising next-generation mobile platform. The primary theory underlying AR is human visual confusion, which allows users to perceive the real-world scenes and augmented contents (virtual-world scenes) simultaneously by superimposing them together. To achieve good Quality of Experience (QoE), it is important to understand the interaction between two scenarios, and harmoniously display AR contents. However, studies on how this superimposition will influence the human visual attention are lacking. Therefore, in this paper, we mainly analyze the interaction effect between background (BG) scenes and AR contents, and study the saliency prediction problem in AR. Specifically, we first construct a Saliency in AR Dataset (SARD), which contains 450 BG images, 450 AR images, as well as 1350 superimposed images generated by superimposing BG and AR images in pair with three mixing levels. A large-scale eye-tracking experiment among 60 subjects is conducted to collect eye movement data. To better predict the saliency in AR, we propose a vector quantized saliency prediction method and generalize it for AR saliency prediction. For comparison, three benchmark methods are proposed and evaluated together with our proposed method on our SARD. Experimental results demonstrate the superiority of our proposed method on both of the common saliency prediction problem and the AR saliency prediction problem over benchmark methods. Our dataset and code are available at: https://github.com/DuanHuiyu/ARSaliency.
Huiyu Duan, Wei Shen 0002, Xiongkuo Min, Danyang Tu, Jing Li 0026, Guangtao Zhai
ACM Multimedia2
2022 Skeleton2Humanoid: Animating Simulated Characters for Physically-plausible Motion In-betweening
abstract
Human motion synthesis is a long-standing problem with various applications in digital twins and the Metaverse. However, modern deep learning based motion synthesis approaches barely consider the physical plausibility of synthesized motions and consequently they usually produce unrealistic human motions. In order to solve this problem, we propose a system "Skeleton2Humanoid" which performs physics-oriented motion correction at test time by regularizing synthesized skeleton motions in a physics simulator. Concretely, our system consists of three sequential stages: (I) test time motion synthesis network adaptation, (II) skeleton to humanoid matching and (III) motion imitation based on reinforcement learning (RL). Stage I introduces a test time adaptation strategy, which improves the physical plausibility of synthesized human skeleton motions by optimizing skeleton joint locations. Stage II performs an analytical inverse kinematics strategy, which converts the optimized human skeleton motions to humanoid robot motions in a physics simulator, then the converted humanoid robot motions can be served as reference motions for the RL policy to imitate. Stage III introduces a curriculum residual force control policy, which drives the humanoid robot to mimic complex converted reference motions in accordance with the physical law. We verify our system on a typical human motion synthesis task, motion-in-betweening. Experiments on the challenging LaFAN1 dataset show our system can outperform prior methods significantly in terms of both physical plausibility and accuracy. Code will be released for research purposes at: https://github.com/michaelliyunhao/Skeleton2Humanoid.
Zhenbo Yu, Yucheng Zhu, Bingbing Ni, Guangtao Zhai, Wei Shen 0002
ACM Multimedia6
2022 Video-based Human-Object Interaction Detection from Tubelet Tokens
abstract
We present a novel vision Transformer, named TUTOR, which is able to learn tubelet tokens, served as highly-abstracted spatial-temporal representations, for video-based human-object interaction (V-HOI) detection. The tubelet tokens structurize videos by agglomerating and linking semantically-related patch tokens along spatial and temporal domains, which enjoy two benefits: 1) Compactness: each token is learned by a selective attention mechanism to reduce redundant dependencies from others; 2) Expressiveness: each token is enabled to align with a semantic instance, i.e., an object or a human, thanks to agglomeration and linking. The effectiveness and efficiency of TUTOR are verified by extensive experiments. Results show our method outperforms existing works by large margins, with a relative mAP gain of $16.14\%$ on VidHOI and a 2 points gain on CAD-120 as well as a $4 \times$ speedup.
Danyang Tu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Wei Shen 0002
NeurIPS5
2022 Rethinking mask heads for partially supervised instance segmentation
Kai Zhao 0012, Wei Shen 0002
Neurocomputing5
2022 Distribution alignment for cross-device palmprint recognition
Kai Zhao 0012, Wei Shen 0002
Pattern Recognit.5
2022 Poxture: Human Posture Imitation Using Neural Texture
abstract
Human pose imitation, which aims to generate an image with a source character’s appearance, the source character’s shape, and a target character’s posture, has many potential applications in virtual reality, augmented reality, games, movies, etc. It is incredibly challenging due to non-rigid human body motions, significant variations in clothing textures, and self-occluded human bodies in 2D images. In this paper, we propose Poxture, a novel human posture imitation method with neural texture, to address the challenges mentioned above. Concretely, first, we build a dense mapping between a source SMPL human body model (shape and posture) and its corresponding texture (appearance). Then, we apply a neural texture generator to recover the complete texture of the source character. At last, we wrap the source neural texture to the source SMLP model with a target pose to generate the desired image by a GAN model. Poxture does not require any annotations, and our framework can fully disentangle the source character’s appearance, shape, and pose, which enjoys several advantages: 1) It can synthesize high-resolution images with detailed textures, thanks to the learned neural textures containing both visible and invisible parts and high-frequency information; 2) It can imitate complex actions with various appearances and body figures since the complete texture of the source character is acquired. We compare our method with previous methods, showing state-of-the-art results on two challenging benchmarks. Extensive experiments demonstrate that, given any character, our method can manipulate this avatar imitating arbitrary posture.
Chen Yang 0023, Zanwei Zhou, Bin Ji 0004, Guangtao Zhai, Wei Shen 0002
IEEE Trans. Circuits Syst. Video Technol.6
2022 External Attention Assisted Multi-Phase Splenic Vascular Injury Segmentation With Limited Data
abstract
The spleen is one of the most commonly injured solid organs in blunt abdominal trauma. The development of automatic segmentation systems from multi-phase CT for splenic vascular injury can augment severity grading for improving clinical decision support and outcome prediction. However, accurate segmentation of splenic vascular injury is challenging for the following reasons: 1) Splenic vascular injury can be highly variant in shape, texture, size, and overall appearance; and 2) Data acquisition is a complex and expensive procedure that requires intensive efforts from both data scientists and radiologists, which makes large-scale well-annotated datasets hard to acquire in general. In light of these challenges, we hereby design a novel framework for multi-phase splenic vascular injury segmentation, especially with limited data. On the one hand, we propose to leverage external data to mine pseudo splenic masks as the spatial attention, dubbed external attention, for guiding the segmentation of splenic vascular injury. On the other hand, we develop a synthetic phase augmentation module, which builds upon generative adversarial networks, for populating the internal data by fully leveraging the relation between different phases. By jointly enforcing external attention and populating internal data representation during training, our proposed method outperforms other competing methods and substantially improves the popular DeepLab-v3+ baseline by more than 7% in terms of average DSC, which confirms its effectiveness.
Yuyin Zhou, David Dreizin, Yan Wang 0033, Fengze Liu, Wei Shen 0002, Alan L. Yuille
IEEE Trans. Medical Imaging5
2021 Deeply Shape-Guided Cascade for Instance Segmentation
abstract
The key to a successful cascade architecture for precise instance segmentation is to fully leverage the relationship between bounding box detection and mask segmentation across multiple stages. Although modern instance segmentation cascades achieve leading performance, they mainly make use of a unidirectional relationship, i.e., mask segmentation can benefit from iteratively refined bounding box detection. In this paper, we investigate an alternative direction, i.e., how to take the advantage of precise mask segmentation for bounding box detection in a cascade architecture. We propose a Deeply Shape-guided Cascade (DSC) for instance segmentation, which iteratively imposes the shape guidances extracted from mask prediction at previous stage on bounding box detection at current stage. It forms a bi-directional relationship between the two tasks by introducing three key components: (1) Initial shape guidance: A mask-supervised Region Proposal Network (mPRN) with the ability to generate class-agnostic masks; (2) Explicit shape guidance: A mask-guided regionof-interest (RoI) feature extractor, which employs mask segmentation at previous stage to focus feature extraction at current stage within a region aligned well with the shape of the instance-of-interest rather than a rectangular RoI; (3) Implicit shape guidance: A feature fusion operation which feeds intermediate mask features at previous stage to the bounding box head at current stage. Experimental results show that DSC outperforms the state-of-the-art instance segmentation cascade, Hybrid Task Cascade (HTC), by a large margin and achieves 51.8 box AP and 45.5 mask AP on COCO test-dev. The code is released at: https://github.com/hding2455/DSC.
Hao Ding 0021, Siyuan Qiao, Alan L. Yuille, Wei Shen 0002
CVPR4
2021 Dual Attention Guided Gaze Target Detection in the Wild
abstract
Gaze target detection aims to infer where each person in a scene is looking. Existing works focus on 2D gaze and 2D saliency, but fail to exploit 3D contexts. In this work, we propose a three-stage method to simulate the human gaze inference behavior in 3D space. In the first stage, we introduce a coarse-to-fine strategy to robustly estimate a 3D gaze orientation from the head. The predicted gaze is decomposed into a planar gaze on the image plane and a depth-channel gaze. In the second stage, we develop a Dual Attention Module (DAM), which takes the planar gaze to produce the filed of view and masks interfering objects regulated by depth information according to the depth-channel gaze. In the third stage, we use the generated dual attention as guidance to perform two sub-tasks: (1) identifying whether the gaze target is inside or out of the image; (2) locating the target if inside. Extensive experiments demonstrate that our approach performs favorably against state-of-the-art methods on GazeFollow and VideoAttentionTarget datasets.
Yi Fang 0009, Jiapeng Tang, Wang Shen, Wei Shen 0002, Xiao Gu 0001, Li Song 0001, Guangtao Zhai
CVPR4
2021 Looking here or there? Gaze Following in 360-Degree Images
abstract
Gaze following, i.e., detecting the gaze target of a human subject, in 2D images has become an active topic in computer vision. However, it usually suffers from the out of frame issue due to the limited field-of-view (FoV) of 2D images. In this paper, we introduce a novel task, gaze following in 360-degree images which provide an omnidirectional FoV and can alleviate the out of frame issue. We collect the first dataset, "GazeFollow360"1, for this task, containing around 10,000 360-degree images with complex gaze behaviors under various scenes. Existing 2D gaze following methods suffer from performance degradation in 360degree images since they may use the assumption that a gaze target is in the 2D gaze sight line. However, this assumption is no longer true for long-distance gaze behaviors in 360-degree images, due to the distortion brought by sphere-to-plane projection. To address this challenge, we propose a 3D sight line guided dual-pathway framework, to detect the gaze target within a local region (here) and from a distant region (there), parallelly. Specifically, the local region is obtained as a 2D cone-shaped field along the 2D projection of the sight line starting at the human subject’s head position, and the distant region is obtained by searching along the sight line in 3D sphere space. Finally, the location of the gaze target is determined by fusing the estimations from both the local region and the distant region. Experimental results show that our method achieves significant improvements over previous 2D gaze following methods on our GazeFollow360 dataset.
Wei Shen 0002, Zhongpai Gao, Yucheng Zhu, Guangtao Zhai, Guodong Guo
ICCV2
2021 CO2: Consistent Contrast for Unsupervised Visual Representation Learning
Chen Wei 0005, Wei Shen 0002, Alan L. Yuille
ICLR3
2021 Shape-Texture Debiased Neural Network Training
Yingwei Li 0002, Qihang Yu, Mingxing Tan, Jieru Mei, Peng Tang 0005, Wei Shen 0002, Alan L. Yuille, Cihang Xie
ICLR6
2021 Glance-and-Gaze Vision Transformer
abstract
Recently, there emerges a series of vision Transformers, which show superior performance with a more compact model size than conventional convolutional neural networks, thanks to the strong ability of Transformers to model long-range dependencies. However, the advantages of vision Transformers also come with a price: Self-attention, the core part of Transformer, has a quadratic complexity to the input sequence length. This leads to a dramatic increase of computation and memory cost with the increase of sequence length, thus introducing difficulties when applying Transformers to the vision tasks that require dense predictions based on high-resolution feature maps.In this paper, we propose a new vision Transformer, named Glance-and-Gaze Transformer (GG-Transformer), to address the aforementioned issues. It is motivated by the Glance and Gaze behavior of human beings when recognizing objects in natural scenes, with the ability to efficiently model both long-range dependencies and local context. In GG-Transformer, the Glance and Gaze behavior is realized by two parallel branches: The Glance branch is achieved by performing self-attention on the adaptively-dilated partitions of the input, which leads to a linear complexity while still enjoying a global receptive field; The Gaze branch is implemented by a simple depth-wise convolutional layer, which compensates local image context to the features obtained by the Glance mechanism. We empirically demonstrate our method achieves consistently superior performance over previous state-of-the-art Transformers on various vision tasks and benchmarks.
Qihang Yu, Yingda Xia, Yutong Bai, Yongyi Lu, Alan L. Yuille, Wei Shen 0002
NeurIPS6
2021 Deep Differentiable Random Forests for Age Estimation
abstract
Age estimation from facial images is typically cast as a label distribution learning or regression problem, since aging is a gradual progress. Its main challenge is the facial feature space w.r.t. ages is inhomogeneous, due to the large variation in facial appearance across different persons of the same age and the non-stationary property of aging. In this paper, we propose two Deep Differentiable Random Forests methods, Deep Label Distribution Learning Forest (DLDLF) and Deep Regression Forest (DRF), for age estimation. Both of them connect split nodes to the top layer of convolutional neural networks (CNNs) and deal with inhomogeneous data by jointly learning input-dependent data partitions at the split nodes and age distributions at the leaf nodes. This joint learning follows an alternating strategy: (1) Fixing the leaf nodes and optimizing the split nodes and the CNN parameters by Back-propagation; (2) Fixing the split nodes and optimizing the leaf nodes by Variational Bounding. Two Deterministic Annealing processes are introduced into the learning of the split and leaf nodes, respectively, to avoid poor local optima and obtain better estimates of tree parameters free of initial values. Experimental results show that DLDLF and DRF achieve state-of-the-art performance on three age estimation datasets.
Wei Shen 0002, Yilu Guo, Yan Wang 0033, Kai Zhao 0012, Bo Wang 0044, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Learning Inductive Attention Guidance for Partially Supervised Pancreatic Ductal Adenocarcinoma Prediction
abstract
Pancreatic ductal adenocarcinoma (PDAC) is the third most common cause of cancer death in the United States. Predicting tumors like PDACs (including both classification and segmentation) from medical images by deep learning is becoming a growing trend, but usually a large number of annotated data are required for training, which is very labor-intensive and time-consuming. In this paper, we consider a partially supervised setting, where cheap image-level annotations are provided for all the training data, and the costly per-voxel annotations are only available for a subset of them. We propose an Inductive Attention Guidance Network (IAG-Net) to jointly learn a global image-level classifier for normal/PDAC classification and a local voxel-level classifier for semi-supervised PDAC segmentation. We instantiate both the global and the local classifiers by multiple instance learning (MIL), where the attention guidance, indicating roughly where the PDAC regions are, is the key to bridging them: For global MIL based normal/PDAC classification, attention serves as a weight for each instance (voxel) during MIL pooling, which eliminates the distraction from the background; For local MIL based semi-supervised PDAC segmentation, the attention guidance is inductive, which not only provides bag-level pseudo-labels to training data without per-voxel annotations for MIL training, but also acts as a proxy of an instance-level classifier. Experimental results show that our IAG-Net boosts PDAC segmentation accuracy by more than 5% compared with the state-of-the-arts.
Yan Wang 0033, Peng Tang 0005, Yuyin Zhou, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille
IEEE Trans. Medical Imaging4
2020 Deep Distance Transform for Tubular Structure Segmentation in CT Scans
abstract
Tubular structure segmentation in medical images, e.g., segmenting vessels in CT scans, serves as a vital step in the use of computers to aid in screening early stages of related diseases. But automatic tubular structure segmentation in CT scans is a challenging problem, due to issues such as poor contrast, noise and complicated background. A tubular structure usually has a cylinder-like shape which can be well represented by its skeleton and cross-sectional radii (scales). Inspired by this, we propose a geometry-aware tubular structure segmentation method, Deep Distance Transform (DDT), which combines intuitions from the classical distance transform for skeletonization and modern deep segmentation networks. DDT first learns a multi-task network to predict a segmentation mask for a tubular structure and a distance map. Each value in the map represents the distance from each tubular structure voxel to the tubular structure surface. Then the segmentation mask is refined by leveraging the shape prior reconstructed from the distance map. We apply our DDT on six medical image datasets. Results show that (1) DDT can boost tubular structure segmentation performance significantly (e.g., over 13% DSC improvement for pancreatic duct segmentation), and (2) DDT additionally provides a geometrical measurement for a tubular structure, which is important for clinical diagnosis (e.g., the cross-sectional scale of a pancreatic duct can be an indicator for pancreatic cancer).
Yan Wang 0033, Fengze Liu, Jieneng Chen, Yuyin Zhou, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille
CVPR6
2020 Synthesize Then Compare: Detecting Failures and Anomalies for Semantic Segmentation
Yingda Xia, Yi Zhang 0099, Fengze Liu, Wei Shen 0002, Alan L. Yuille
ECCV (1)4
2020 Domain Adaptive Relational Reasoning for 3D Multi-organ Segmentation
Shuhao Fu, Yongyi Lu, Yan Wang 0033, Yuyin Zhou, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille
MICCAI (1)5
2020 Detecting Pancreatic Ductal Adenocarcinoma in Multi-phase CT Scans via Alignment Ensemble
Yingda Xia, Qihang Yu, Wei Shen 0002, Yuyin Zhou, Elliot K. Fishman, Alan L. Yuille
MICCAI (3)3
2020 Robust Face Detection via Learning Small Faces on Hard Images
abstract
Recent anchor-based deep face detectors have achieved promising performance, but they are still struggling to detect hard faces, such as small, blurred and partially occluded faces. One reason is that they treat all images and faces equally, and ignore the imbalance between easy images and hard images; however large amounts of training images only contain easy faces, which are less helpful to learn robust detectors for hard faces. In this paper, we propose that the robustness of a face detector against hard faces can be improved by learning small faces on hard images. Our intuitions are (1) hard images are the images which contain at least one hard face, thus they facilitate training robust face detectors; (2) most hard faces are small faces and other types of hard faces can be easily shrunk to small faces. To this end, we build an anchor-based deep face detector, which only outputs a single high-resolution feature map with small anchors, to specifically learn small faces and train it by a novel hard image mining strategy which automatically adjusts training weights on images according to their difficulties. Extensive experiments have been conducted on WIDER FACE, FDDB, Pascal Faces, and AFW datasets and our method achieves APs of 95.7, 94.9 and 89.7 on easy, medium and hard WIDER FACE val dataset respectively, which verify the effectiveness of our methods, especially on detecting hard faces. Our detector is also lightweight and enjoys a fast inference speed. Code and model are available at https://github.com/bairdzhang/smallhardface.
Zhishuai Zhang, Wei Shen 0002, Siyuan Qiao, Yan Wang 0033, Bo Wang 0044, Alan L. Yuille
WACV2
2020 Resisting Large Data Variations via Introspective Transformation Network
abstract
Training deep networks that generalize to a wide range of variations in test data is essential to building accurate and robust image classifiers. Data variations in this paper include but not limited to unseen affine transformations and warping in the training data. One standard strategy to overcome this problem is to apply data augmentation to synthetically enlarge the training set. However, data augmentation is essentially a brute-force method which generates uniform samples from some pre-defined set of transformations. In this paper, we propose a principled approach named introspective transformation network (ITN) that significantly improves network resistance to large variations between training and testing data. This is achieved by embedding a learnable transformation module into the introspective network, which is a convolutional neural network (CNN) classifier empowered with generative capabilities. Our approach alternates between synthesizing pseudo-negative samples and transformed positive examples based on the current model, and optimizing model predictions on these synthesized samples. Experimental results verify that our approach significantly improves the ability of deep networks to resist large variations between training and testing data and achieves classification accuracy improvements on several benchmark datasets, including MNIST, affNIST, SVHN, CIFAR-10 and miniImageNet.
Yunhan Zhao, Charless C. Fowlkes, Wei Shen 0002, Alan L. Yuille
WACV4
2020 PCL: Proposal Cluster Learning for Weakly Supervised Object Detection
abstract
Weakly Supervised Object Detection (WSOD), using only image-level annotations to train object detectors, is of growing importance in object recognition. In this paper, we propose a novel deep network for WSOD. Unlike previous networks that transfer the object detection problem to an image classification problem using Multiple Instance Learning (MIL), our strategy generates proposal clusters to learn refined instance classifiers by an iterative process. The proposals in the same cluster are spatially adjacent and associated with the same object. This prevents the network from concentrating too much on parts of objects instead of whole objects. We first show that instances can be assigned object or background labels directly based on proposal clusters for instance classifier refinement, and then show that treating each cluster as a small new bag yields fewer ambiguities than the directly assigning label method. The iterative instance classifier refinement is implemented online using multiple streams in convolutional neural networks, where the first is an MIL network and the others are for instance classifier refinement supervised by the preceding one. Experiments are conducted on the PASCAL VOC, ImageNet detection, and MS-COCO benchmarks for WSOD. Results show that our method outperforms the previous state of the art significantly.
Peng Tang 0005, Xinggang Wang, Song Bai 0001, Wei Shen 0002, Xiang Bai, Wenyu Liu 0001, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Semi-Supervised 3D Abdominal Multi-Organ Segmentation Via Deep Multi-Planar Co-Training
abstract
In multi-organ segmentation of abdominal CT scans, most existing fully supervised deep learning algorithms require lots of voxel-wise annotations, which are usually difficult, expensive, and slow to obtain. In comparison, massive unlabeled 3D CT volumes are usually easily accessible. Current mainstream works to address semi-supervised biomedical image segmentation problem are mostly graph-based. By contrast, deep network based semi-supervised learning methods have not drawn much attention in this field. In this work, we propose Deep Multi-Planar Co-Training (DMPCT), whose contributions can be divided into two folds: 1) The deep model is learned in a co-training style which can mine consensus information from multiple planes like the sagittal, coronal, and axial planes; 2) Multi-planar fusion is applied to generate more reliable pseudo-labels, which alleviates the errors occurring in the pseudo-labels and thus can help to train better segmentation networks. Experiments are done on our newly collected large dataset with 100 unlabeled cases as well as 210 labeled cases where 16 anatomical structures are manually annotated by four radiologists and confirmed by a senior expert. The results suggest that DMPCT significantly outperforms the fully supervised method by more than 4% especially when only a small set of annotations is used.
Yuyin Zhou, Yan Wang 0033, Peng Tang 0005, Song Bai 0001, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille
WACV5
2019 Proposal pyramid networks for fast face detection
Dan Zeng 0001, Fan Zhao 0004, Shiming Ge, Wei Shen 0002, Zhijiang Zhang
Inf. Sci.5
2019 Abdominal multi-organ segmentation with organ-attention networks and statistical fusion
Yan Wang 0033, Yuyin Zhou, Wei Shen 0002, Seyoun Park, Elliot K. Fishman, Alan L. Yuille
Medical Image Anal.3
2019 Fast cascade face detection with pyramid network
Dan Zeng 0001, Fan Zhao 0004, Shiming Ge, Wei Shen 0002
Pattern Recognit. Lett.4
2018 A 3D Coarse-to-Fine Framework for Volumetric Medical Image Segmentation
abstract
In this paper, we adopt 3D Convolutional Neural Networks to segment volumetric medical images. Although deep neural networks have been proven to be very effective on many 2D vision tasks, it is still challenging to apply them to 3D tasks due to the limited amount of annotated 3D data and limited computational resources. We propose a novel 3D-based coarse-to-fine framework to effectively and efficiently tackle these challenges. The proposed 3D-based framework outperforms the 2D counterpart to a large margin since it can leverage the rich spatial information along all three axes. We conduct experiments on two datasets which include healthy and pathological pancreases respectively, and achieve the current state-of-the-art in terms of Dice-Sørensen Coefficient (DSC). On the NIH pancreas segmentation dataset, we outperform the previous best by an average of over 2%, and the worst case is improved by 7% to reach almost 70%, which indicates the reliability of our framework in clinical applications.
Zhuotun Zhu, Yingda Xia, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille
3DV3
2018 Deep Regression Forests for Age Estimation
abstract
Age estimation from facial images is typically cast as a nonlinear regression problem. The main challenge of this problem is the facial feature space w.r.t. ages is inhomogeneous, due to the large variation in facial appearance across different persons of the same age and the non-stationary property of aging patterns. In this paper, we propose Deep Regression Forests (DRFs), an end-to-end model, for age estimation. DRFs connect the split nodes to a fully connected layer of a convolutional neural network (CNN) and deal with inhomogeneous data by jointly learning input-dependant data partitions at the split nodes and data abstractions at the leaf nodes. This joint learning follows an alternating strategy: First, by fixing the leaf nodes, the split nodes as well as the CNN parameters are optimized by Back-propagation; Then, by fixing the split nodes, the leaf nodes are optimized by iterating a step-size free update rule derived from Variational Bounding. We verify the proposed DRFs on three standard age estimation benchmarks and achieve state-of-the-art results on all of them.
Wei Shen 0002, Yilu Guo, Yan Wang 0033, Kai Zhao 0012, Bo Wang 0044, Alan L. Yuille
CVPR1
2018 Few-Shot Image Recognition by Predicting Parameters From Activations
abstract
In this paper, we are interested in the few-shot learning problem. In particular, we focus on a challenging scenario where the number of categories is large and the number of examples per novel category is very limited, e.g. 1, 2, or 3. Motivated by the close relationship between the parameters and the activations in a neural network associated with the same category, we propose a novel method that can adapt a pre-trained neural network to novel categories by directly predicting the parameters from the activations. Zero training is required in adaptation to novel categories, and fast inference is realized by a single forward pass. We evaluate our method by doing few-shot image recognition on the ImageNet dataset, which achieves the state-of-the-art classification accuracy on novel categories by a significant margin while keeping comparable performance on the large-scale categories. We also test our method on the MiniImageNet dataset and it strongly outperforms the previous state-of-the-art methods.
Siyuan Qiao, Chenxi Liu 0001, Wei Shen 0002, Alan L. Yuille
CVPR3
2018 Single-Shot Object Detection With Enriched Semantics
abstract
We propose a novel single shot object detection network named Detection with Enriched Semantics (DES). Our motivation is to enrich the semantics of object detection features within a typical deep detector, by a semantic segmentation branch and a global activation module. The segmentation branch is supervised by weak segmentation ground-truth, i.e., no extra annotation is required. In conjunction with that, we employ a global activation module which learns relationship between channels and object classes in a self-supervised manner. Comprehensive experimental results on both PASCAL VOC and MS COCO detection datasets demonstrate the effectiveness of the proposed method. In particular, with a VGG16 based DES, we achieve an mAP of 81.7 on VOC2007 test and an mAP of 32.8 on COCO test-dev with an inference speed of 31.5 milliseconds per image on a Titan Xp GPU. With a lower resolution version, we achieve an mAP of 79.7 on VOC2007 with an inference speed of 13.0 milliseconds per image.
Zhishuai Zhang, Siyuan Qiao, Cihang Xie, Wei Shen 0002, Bo Wang 0044, Alan L. Yuille
CVPR4
2018 Deep Co-Training for Semi-Supervised Image Recognition
Siyuan Qiao, Wei Shen 0002, Zhishuai Zhang, Bo Wang 0044, Alan L. Yuille
ECCV (15)2
2018 Gradually Updated Neural Networks for Large-Scale Image Recognition
abstract
Depth is one of the keys that make neural networks succeed in the task of large-scale image recognition. The state-of-the-art network architectures usually increase the depths by cascading convolutional layers or building blocks. In this paper, we present an alternative method to increase the depth. Our method is by introducing computation orderings to the channels within convolutional layers or blocks, based on which we gradually compute the outputs in a channel-wise manner. The added orderings not only increase the depths and the learning capacities of the networks without any additional computation costs, but also eliminate the overlap singularities so that the networks are able to converge faster and perform better. Experiments show that the networks based on our method achieve the state-of-the-art performances on CIFAR and ImageNet datasets.
Siyuan Qiao, Zhishuai Zhang, Wei Shen 0002, Bo Wang 0044, Alan L. Yuille
ICML3
2018 Hi-Fi: Hierarchical Feature Integration for Skeleton Detection
abstract
In natural images, the scales (thickness) of object skeletons may dramatically vary among objects and object parts. Thus, robust skeleton detection requires powerful multi-scale feature integration ability. To address this issue, we present a new convolutional neural network (CNN) architecture by introducing a novel hierarchical feature integration mechanism, named Hi-Fi, to address the object skeleton detection problem. The proposed CNN-based approach intrinsically captures high-level semantics from deeper layers, as well as low-level details from shallower layers. By hierarchically integrating different CNN feature levels with bidirectional guidance, our approach (1) enables mutual refinement across features of different levels, and (2) possesses the strong ability to capture both rich object context and high-resolution details. Experimental results show that our method significantly outperforms the state-of-the-art methods in terms of effectively fusing features from very different scales, as evidenced by a considerable performance improvement on several benchmarks.
Kai Zhao 0012, Wei Shen 0002, Shanghua Gao, Ming-Ming Cheng
IJCAI2
2018 Training Multi-organ Segmentation Networks with Sample Selection by Relaxed Upper Confident Bound
Yan Wang 0033, Yuyin Zhou, Peng Tang 0005, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille
MICCAI (4)4
2018 Bag of Shape Features with a learned pooling function for shape recognition
Wei Shen 0002, Chenting Du, Yuan Jiang 0002, Dan Zeng 0001, Zhijiang Zhang
Pattern Recognit. Lett.1
2018 Multi-oriented text detection from natural scene images based on a CNN and pruning non-adjacent graph edges
Yuanwang Wei, Wei Shen 0002, Dan Zeng 0001, Lihua Ye, Zhijiang Zhang
Signal Process. Image Commun.2
2017 ScaleNet: Guiding Object Proposal Generation in Supermarkets and Beyond
abstract
Motivated by product detection in supermarkets, this paper studies the problem of object proposal generation in supermarket images and other natural images. We argue that estimation of object scales in images is helpful for generating object proposals, especially for supermarket images where object scales are usually within a small range. Therefore, we propose to estimate object scales of images before generating object proposals. The proposed method for predicting object scales is called ScaleNet. To validate the effectiveness of ScaleNet, we build three supermarket datasets, two of which are real-world datasets used for testing and the other one is a synthetic dataset used for training. In short, we extend the previous state-of-the-art object proposal methods by adding a scale prediction phase. The resulted method outperforms the previous state-of-the-art on the supermarket datasets by a large margin. We also show that the approach works for object proposal on other natural images and it outperforms the previous state-of-the-art object proposal methods on the MS COCO dataset. The supermarket datasets, the virtual supermarkets, and the tools for creating more synthetic datasets will be made public.
Siyuan Qiao, Wei Shen 0002, Weichao Qiu, Chenxi Liu 0001, Alan L. Yuille
ICCV2
2017 Multi-stage Multi-recursive-input Fully Convolutional Networks for Neuronal Boundary Detection
abstract
In the field of connectomics, neuroscientists seek to identify cortical connectivity comprehensively. Neuronal boundary detection from the Electron Microscopy (EM) images is often done to assist the automatic reconstruction of neuronal circuit. But the segmentation of EM images is a challenging problem, as it requires the detector to be able to detect both filament-like thin and blob-like thick membrane, while suppressing the ambiguous intracellular structure. In this paper, we propose multi-stage multi-recursiveinput fully convolutional networks to address this problem. The multiple recursive inputs for one stage, i.e., the multiple side outputs with different receptive field sizes learned from the lower stage, provide multi-scale contextual boundary information for the consecutive learning. This design is biologically-plausible, as it likes a human visual system to compare different possible segmentation solutions to address the ambiguous boundary issue. Our multi-stage networks are trained end-to-end. It achieves promising results on two public available EM segmentation datasets, the mouse piriform cortex dataset and the ISBI 2012 EM dataset.
Wei Shen 0002, Bin Wang 0027, Yuan Jiang 0002, Yan Wang 0033, Alan L. Yuille
ICCV1
2017 Shape recognition by bag of contour fragments with a learned pooling function
abstract
Bag of Contour Fragments (BoCF), derived from the well-known Bag-of-Features (BoF), is an effective framework for shape representation. The feature pooling in this framework is a critical step, while either max pooling or average pooling is not a learnable process. In this paper, we aim at learning a pooling function which is adaptive to the input contour fragment features instead. Towards this end, we formulate our pooling function as a weighted sum of max pooling and average pooling, where the weight is expressed by an activation function of the input contour fragment features. To automatically learn this weight, the output of the pooling function is fed into a SVM classifier and they are trained jointly to minimize a shape classification loss. Experimental results on several standard shape datasets demonstrate the effectiveness of the proposed learned pooling function, which can achieve considerable improvements compared with BoCF.
Wei Shen 0002, Wenjing Gao, Yuan Jiang 0002, Dan Zeng 0001, Zhijiang Zhang
ICIP1
2017 A Fixed-Point Model for Pancreas Segmentation in Abdominal CT Scans
Yuyin Zhou, Lingxi Xie, Wei Shen 0002, Yan Wang 0033, Elliot K. Fishman, Alan L. Yuille
MICCAI (1)3
2017 Label Distribution Learning Forests
abstract
Label distribution learning (LDL) is a general learning framework, which assigns to an instance a distribution over a set of labels rather than a single label or multiple labels. Current LDL methods have either restricted assumptions on the expression form of the label distribution or limitations in representation learning, e.g., to learn deep features in an end-to-end manner. This paper presents label distribution learning forests (LDLFs) - a novel label distribution learning algorithm based on differentiable decision trees, which have several advantages: 1) Decision trees have the potential to model any general form of label distributions by a mixture of leaf node predictions. 2) The learning of differentiable decision trees can be combined with representation learning. We define a distribution-based loss function for a forest, enabling all the trees to be learned jointly, and show that an update function for leaf node predictions, which guarantees a strict decrease of the loss function, can be derived by variational bounding. The effectiveness of the proposed LDLFs is verified on several LDL tasks and a computer vision application, showing significant improvements to the state-of-the-art LDL methods.
Wei Shen 0002, Kai Zhao 0012, Yilu Guo, Alan L. Yuille
NIPS1
2017 Neighborhood geometry based feature matching for geostationary satellite remote sensing image
Dan Zeng 0001, Wei Shen 0002, Qi Tian 0001
Neurocomputing4
2017 Directional Edge Boxes: Exploiting Inner Normal Direction Cues for Effective Object Proposal Generation
Xiang Bai, Zheng Zhang 0022, Wei Shen 0002
J. Comput. Sci. Technol.4
2017 Text detection in scene images based on exhaustive segmentation
Yuanwang Wei, Zhijiang Zhang, Wei Shen 0002, Dan Zeng 0001, Mei Fang, Shifu Zhou
Signal Process. Image Commun.3
2017 DeepSkeleton: Learning Multi-Task Scale-Associated Deep Side Outputs for Object Skeleton Extraction in Natural Images
abstract
Object skeletons are useful for object representation and object detection. They are complementary to the object contour, and provide extra information, such as how object scale (thickness) varies among object parts. But object skeleton extraction from natural images is very challenging, because it requires the extractor to be able to capture both local and non-local image context in order to determine the scale of each skeleton pixel. In this paper, we present a novel fully convolutional network with multiple scale-associated side outputs to address this problem. By observing the relationship between the receptive field sizes of the different layers in the network and the skeleton scales they can capture, we introduce two scale-associated side outputs to each stage of the network. The network is trained by multi-task learning, where one task is skeleton localization to classify whether a pixel is a skeleton pixel or not, and the other is skeleton scale prediction to regress the scale of each skeleton pixel. Supervision is imposed at different stages by guiding the scale-associated side outputs toward the ground-truth skeletons at the appropriate scales. The responses of the multiple scale-associated side outputs are then fused in a scale-specific way to detect skeleton pixels using multiple scales effectively. Our method achieves promising results on two skeleton extraction datasets, and significantly outperforms other competitors. In addition, the usefulness of the obtained skeletons and scales (thickness) are verified on two object detection applications: foreground object segmentation and object proposal detection.
Wei Shen 0002, Kai Zhao 0012, Yuan Jiang 0002, Yan Wang 0033, Xiang Bai, Alan L. Yuille
IEEE Trans. Image Process.1
2016 Object Skeleton Extraction in Natural Images by Fusing Scale-Associated Deep Side Outputs
abstract
Object skeleton is a useful cue for object detection, complementary to the object contour, as it provides a structural representation to describe the relationship among object parts. While object skeleton extraction in natural images is a very challenging problem, as it requires the extractor to be able to capture both local and global image context to determine the intrinsic scale of each skeleton pixel. Existing methods rely on per-pixel based multi-scale feature computation, which results in difficult modeling and high time consumption. In this paper, we present a fully convolutional network with multiple scale-associated side outputs to address this problem. By observing the relationship between the receptive field sizes of the sequential stages in the network and the skeleton scales they can capture, we introduce a scale-associated side output to each stage. We impose supervision to different stages by guiding the scale-associated side outputs toward groundtruth skeletons of different scales. The responses of the multiple scaleassociated side outputs are then fused in a scale-specific way to localize skeleton pixels with multiple scales effectively. Our method achieves promising results on two skeleton extraction datasets, and significantly outperforms other competitors.
Wei Shen 0002, Kai Zhao 0012, Yuan Jiang 0002, Yan Wang 0033, Zhijiang Zhang, Xiang Bai
CVPR1
2016 Multi-oriented Text Detection with Fully Convolutional Networks
abstract
In this paper, we propose a novel approach for text detection in natural images. Both local and global cues are taken into account for localizing text lines in a coarse-to-fine procedure. First, a Fully Convolutional Network (FCN) model is trained to predict the salient map of text regions in a holistic manner. Then, text line hypotheses are estimated by combining the salient map and character components. Finally, another FCN classifier is used to predict the centroid of each character, in order to remove the false hypotheses. The framework is general for handling text in multiple orientations, languages and fonts. The proposed method consistently achieves the state-of-the-art performance on three text detection benchmarks: MSRA-TD500, ICDAR2015 and ICDAR2013.
Zheng Zhang 0022, Chengquan Zhang, Wei Shen 0002, Cong Yao, Wenyu Liu 0001, Xiang Bai
CVPR3
2016 Multiple instance subspace learning via partial random projection tree for local reflection symmetry in natural images
Wei Shen 0002, Xiang Bai, Zihao Hu, Zhijiang Zhang
Pattern Recognit.1
2016 Shape recognition by bag of skeleton-associated contour parts
Wei Shen 0002, Yuan Jiang 0002, Wenjing Gao, Dan Zeng 0001, Xinggang Wang
Pattern Recognit. Lett.1
2016 Spatial-temporal convolutional neural networks for anomaly detection and localization in crowded scenes
Shifu Zhou, Wei Shen 0002, Dan Zeng 0001, Mei Fang, Yuanwang Wei, Zhijiang Zhang
Signal Process. Image Commun.2
2015 DeepContour: A deep convolutional feature learned by positive-sharing loss for contour detection
abstract
Contour detection serves as the basis of a variety of computer vision tasks such as image segmentation and object recognition. The mainstream works to address this problem focus on designing engineered gradient features. In this work, we show that contour detection accuracy can be improved by instead making the use of the deep features learned from convolutional neural networks (CNNs). While rather than using the networks as a blackbox feature extractor, we customize the training strategy by partitioning contour (positive) data into subclasses and fitting each subclass by different model parameters. A new loss function, named positive-sharing loss, in which each subclass shares the loss for the whole positive class, is proposed to learn the parameters. Compared to the sofmax loss function, the proposed one, introduces an extra regularizer to emphasizes the losses for the positive and negative classes, which facilitates to explore more discriminative features. Our experimental results demonstrate that learned deep features can achieve top performance on Berkeley Segmentation Dataset and Benchmark (BSDS500) and obtain competitive cross dataset generalization result on the NYUD dataset.
Wei Shen 0002, Xinggang Wang, Yan Wang 0033, Xiang Bai, Zhijiang Zhang
CVPR1
2015 Symmetry-based text line detection in natural scenes
abstract
Recently, a variety of real-world applications have triggered huge demand for techniques that can extract textual information from natural scenes. Therefore, scene text detection and recognition have become active research topics in computer vision. In this work, we investigate the problem of scene text detection from an alternative perspective and propose a novel algorithm for it. Different from traditional methods, which mainly make use of the properties of single characters or strokes, the proposed algorithm exploits the symmetry property of character groups and allows for direct extraction of text lines from natural images. The experiments on the latest ICDAR benchmarks demonstrate that the proposed algorithm achieves state-of-the-art performance. Moreover, compared to conventional approaches, the proposed algorithm shows stronger adaptability to texts in challenging scenarios.
Zheng Zhang 0022, Wei Shen 0002, Cong Yao, Xiang Bai
CVPR2
2015 Unusual event detection in crowded scenes by trajectory analysis
abstract
Anomaly detection in crowded scenes is a challenge task due to variation of the definitions for both abnormality and normality, the low resolution on the target, ambiguity of appearance, and severe occlusions of inter-object. In this paper, we propose a novel statistical framework to detect abnormal behaviors of the crowded scene by modeling trajectories of pedestrians. First, the trajectories are acquired by Kanade-Lucas-Tomasi Feature Tracker (KLT). Then trajectories are grouped to form representative trajectories, which characterize the underlying motion patterns of the crowd. Finally, trajectories are modeled by Multi-Observation Hidden Markov Model (MOHMM) to determine whether frames are normal or abnormal. The experiments are conducted on a well-known crowded scene dataset. Experimental results show that the proposed method can capture abnormal crowd behaviors successfully and achieves state-of-the-art performances.
Shifu Zhou, Wei Shen 0002, Dan Zeng 0001, Zhijiang Zhang
ICASSP2
2014 Regularity Guaranteed Human Pose Correction
Wei Shen 0002, Rui Lei, Dan Zeng 0001, Zhijiang Zhang
ACCV (2)1
2014 Exemplar-Based Human Action Pose Correction
abstract
The launch of Xbox Kinect has built a very successful computer vision product and made a big impact on the gaming industry. This sheds lights onto a wide variety of potential applications related to action recognition. The accurate estimation of human poses from the depth image is universally a critical step. However, existing pose estimation systems exhibit failures when facing severe occlusion. In this paper, we propose an exemplar-based method to learn to correct the initially estimated poses. We learn an inhomogeneous systematic bias by leveraging the exemplar information within a specific human action domain. Furthermore, as an extension, we learn a conditional model by incorporation of pose tags to further increase the accuracy of pose correction. In the experiments, significant improvements on both joint-based skeleton correction and tag prediction are observed over the contemporary approaches, including what is delivered by the current Kinect system. Our experiments for the facial landmark correction also illustrate that our algorithm can improve the accuracy of other detection/estimation systems.
Wei Shen 0002, Xiang Bai, Tommer Leyvand, Baining Guo, Zhuowen Tu
IEEE Trans. Cybern.1
2013 Skeleton pruning as trade-off between skeleton simplicity and reconstruction error
Wei Shen 0002, Xiang Bai, Xingwei Yang, Longin Jan Latecki
Sci. China Inf. Sci.1
2013 Face identification using reference-based features with message passing model
Wei Shen 0002, Bo Wang 0044, Xiang Bai, Longin Jan Latecki
Neurocomputing1
2013 Shape clustering: Common structure discovery
Wei Shen 0002, Yan Wang 0033, Xiang Bai, Longin Jan Latecki
Pattern Recognit.1
2012 Exemplar-based human action pose correction and tagging
abstract
The launch of Xbox Kinect has built a very successful computer vision product and made a big impact to the gaming industry; this sheds lights onto a wide variety of potential applications related to action recognition. The accurate estimation of human poses from the depth image is universally a critical step. However, existing pose estimation systems exhibit failures when faced severe occlusion. In this paper, we propose an exemplar-based method to learn to correct the initially estimated poses. We learn an inhomogeneous systematic bias by leveraging the exemplar information within specific human action domain. Our algorithm is illustrated on both joint-based skeleton correction and tag prediction. In the experiments, significant improvement is observed over the contemporary approaches, including what is delivered by the current Kinect system.
Wei Shen 0002, Xiang Bai, Tommer Leyvand, Baining Guo, Zhuowen Tu
CVPR1
2011 Skeleton growing and pruning with bending potential ratio
Wei Shen 0002, Xiang Bai, Longin Jan Latecki
Pattern Recognit.1
2010 Shape Classification Using Tree -Unions
abstract
In this paper, we proposed a novel approach to shape classification. A new shape tree based on junction nodes can represent the global structure in a simple way. The statistic distribution of junctions can be learned by merging the shape trees. In the process of learning, context of a junction node is obtained to improve the rate of classification. We illustrate the utility of the proposed method on the problem of 2D shape classification using the new shape tree representation.
Bo Wang 0044, Wei Shen 0002, Wenyu Liu 0001, Xinge You, Xiang Bai
ICPR2
2010 Recursive spatiotemporal subspace learning for gait recognition
Wei Shen 0002
Neurocomputing2