Lingxi Xie

dblp:123/2869 · DBLP profile ↗
← Back
160ranked-venue papers
19as first author
95since 2021 · last 2026
0000-0003-4831-9451ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 117 · 10 first-author · 71 since 2021Graphics, computer vision, multimedia, augmented reality and games · 116 · 15 first-author · 57 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-authorComputer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 WorldGrow: Generating Infinite 3D World
abstract
We tackle the challenge of generating the infinitely extendable 3D world -- large, continuous environments with coherent geometry and realistic appearance. Existing methods face key challenges: 2D-lifting approaches suffer from geometric and appearance inconsistencies across views, 3D implicit representations are hard to scale up, and current 3D foundation models are mostly object-centric, limiting their applicability to scene-level generation. Our key insight is leveraging strong generation priors from pre-trained 3D models for structured scene block generation. To this end, we propose WorldGrow, a hierarchical framework for unbounded 3D scene synthesis. Our method features three core components: (1) a data curation pipeline that extracts high-quality scene blocks for training, making the 3D structured latent representations suitable for scene generation; (2) a 3D block inpainting mechanism that enables context-aware scene extension; and (3) a coarse-to-fine generation strategy that ensures both global layout plausibility and local geometric/textural fidelity. Evaluated on the large-scale 3D-FRONT dataset, WorldGrow achieves SOTA performance in geometry reconstruction, while uniquely supporting infinite scene generation with photorealistic and structurally consistent outputs. These results highlight its capability for constructing large-scale virtual environments and potential for building future world models.
Sikuang Li, Chen Yang 0023, Jiemin Fang, Taoran Yi, Jiazhong Cen, Lingxi Xie, Wei Shen 0002, Qi Tian 0001
AAAI7
2026 Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
abstract
Flow-based 3D generation models typically require dozens of sampling steps during inference. Though few-step distillation methods, particularly Consistency Models (CMs), have achieved substantial advancements in accelerating 2D diffusion models, they remain under-explored for more complex 3D generation tasks. In this study, we propose a novel framework, MDT-dist, for few-step 3D flow distillation. Our approach is built upon a primary objective: distilling the pretrained model to learn the Marginal-Data Transport. Directly learning this objective needs to integrate the velocity fields, while this integral is intractable to be implemented. Therefore, we propose two optimizable objectives, Velocity Matching (VM) and Velocity Distillation (VD), to equivalently convert the optimization target from the transport level to the velocity and the distribution level respectively. Velocity Matching (VM) learns to stably match the velocity fields between the student and the teacher, but inevitably provides biased gradient estimates. Velocity Distillation (VD) further enhances the optimization process by leveraging the learned velocity fields to perform probability density distillation. When evaluated on the pioneer 3D generation framework TRELLIS, our method reduces sampling steps of each flow transformer from 25 to 1–2, achieving 0.68s (1 step x2) and 0.94s (2 steps x2) latency with 9.0x and 6.5x speedup on A800, while preserving high visual and geometric fidelity. Experiments demonstrate that our method significantly outperforms existing CM distillation methods, and enables TRELLIS to achieve superior performance in few-step 3D generation.
Zanwei Zhou, Taoran Yi, Jiemin Fang, Chen Yang 0023, Lingxi Xie, Xinggang Wang, Wei Shen 0002, Qi Tian 0001
AAAI5
2026 DiffCrack: A semantic-structural controllable framework for crack image generation in complex scenes
abstract
Robust pavement crack detection in complex scenes remains a significant challenge. This stems not merely from the scarcity of annotated data, but more critically, from the severe lack of pattern diversity within existing datasets. Key variations in morphology, scale, background texture, and imaging conditions are often underrepresented, which fundamentally impedes the generalization capability of recognition models. While generative approaches (e.g., GANs and diffusion models) offer a potential path to augment this diversity synthetically, they commonly suffer from poor background realism, entangled structural-appearance representations, and a lack of precise control over generated defects. This paper presents the Crack Diffusion Generator (DiffCrack), a diffusion-based framework designed for semantic-structural controllability in crack image synthesis. DiffCrack decouples crack geometry and visual appearance through two conditioning inputs: a binary mask to anchor spatial layout, and a Hierarchical Prompt Attention (HPA) module to independently modulate attributes such as width, depth, color, and texture. This design enables the controllable and targeted generation of diverse, photorealistic crack patterns that are often missing from real-world datasets. Extensive experiments on real datasets demonstrate that training with DiffCrack-generated images enhances the F1-score of segmentation models by up to 23% on average under complex scene conditions. This result validates that DiffCrack is a scalable, pattern-aware data augmentation tool. By addressing the critical bottleneck of data diversity, our framework offers a principled pathway to improving the robustness and generalization of infrastructure inspection models.
Lingxi Xie, Qin Zou 0001, Qi Tian 0001, Qingquan Li 0001
Pattern Recognit.2
2026 EinsPT: Efficient Instance-Aware Pre-Training of Vision Foundation Models
abstract
In this study, we introduce EinsPT, an efficient instance-aware pre-training paradigm designed to reduce the transfer gap between vision foundation models and downstream instance-level tasks. Unlike conventional image-level pre-training that relies solely on unlabeled images, EinsPT leverages both image reconstruction and instance annotations to learn representations that are spatially coherent and instance discriminative. To achieve this efficiently, we propose a proxy-foundation architecture that decouples high-resolution and low-resolution learning: the foundation model processes masked low-resolution images for global semantics, while a lightweight proxy model operates on complete high-resolution images to preserve fine-grained details. The two branches are jointly optimized through reconstruction and instance-level prediction losses on fused features. Extensive experiments demonstrate that EinsPT consistently enhances recognition accuracy across various downstream tasks with substantially reduced computational cost, while qualitative results further reveal improved instance perception and completeness in visual representations. Code is available at github.com/feufhd/EinsPT.
Zhaozhi Wang, Yunjie Tian, Lingxi Xie, Yaowei Wang 0001, Qixiang Ye
IEEE Trans. Image Process.3
2025 Segment Any 3D Gaussians
abstract
This paper presents SAGA (Segment Any 3D GAussians), a highly efficient 3D promptable segmentation method based on 3D Gaussian Splatting (3D-GS). Given 2D visual prompts as input, SAGA can segment the corresponding 3D target represented by 3D Gaussians within 4 ms. This is achieved by attaching a scale-gated affinity feature to each 3D Gaussian to endow it a new property towards multi-granularity segmentation. Specifically, a scale-aware contrastive training strategy is proposed for the scale-gated affinity feature learning. It 1) distills the segmentation capability of the Segment Anything Model (SAM) from 2D masks into the affinity features and 2) employs a soft scale gate mechanism to deal with multi-granularity ambiguity in 3D segmentation through adjusting the magnitude of each feature channel according to a specified 3D physical scale. Evaluations demonstrate that SAGA achieves real-time multi-granularity segmentation with quality comparable to state-of-the-art methods. As one of the first methods addressing promptable segmentation in 3D-GS, the simplicity and effectiveness of SAGA pave the way for future advancements in this field.
Jiazhong Cen, Jiemin Fang, Chen Yang 0023, Lingxi Xie, Xiaopeng Zhang 0008, Wei Shen 0002, Qi Tian 0001
AAAI4
2025 ChatterBox: Multimodal Referring and Grounding with Chain-of-Questions
abstract
In this study, we establish a benchmark and a baseline approach for Multimodal referring and grounding with Chain-of-Questions (MCQ), opening up a promising direction for ‘logical’ multimodal dialogues. The newly collected dataset, named CB-300K, spans challenges including probing dialogues with spatial relationship among multiple objects, consistent reasoning, and complex question chains. The baseline approach, termed ChatterBox, involves a modularized design and a referent feedback mechanism to ensure logical coherence in continuous referring and grounding tasks. This design reduces the risk of referential confusion, simplifies the training process, and presents validity in retaining the language model’s generation ability. Experiments show that ChatterBox demonstrates superiority in MCQ both quantitatively and qualitatively, paving a new path towards multimodal dialogue scenarios with logical interactions.
Yunjie Tian, Tianren Ma, Lingxi Xie, Qixiang Ye
AAAI3
2025 Adaptive Keyframe Sampling for Long Video Understanding
abstract
Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. However, when the visual input changes from a single image to a long video, the above paradigm encounters difficulty because the vast amount of video tokens has significantly exceeded the maximal capacity of MLLMs. Therefore, existing video-based MLLMs are mostly established upon sampling a small portion of tokens from input data, which can cause key information to be lost and thus produce incorrect answers. This paper presents a simple yet effective algorithm named Adaptive Keyframe Sampling (AKS). It inserts a plug-and-play module known as keyframe selection, which aims to maximize the useful information with a fixed number of video tokens. We formulate keyframe selection as an optimization involving (1) the relevance between the keyframes and the prompt, and (2) the coverage of the keyframes over the video, and present an adaptive algorithm to approximate the best solution. Experiments on two long video understanding benchmarks validate that AKS improves video QA accuracy (beyond strong baselines) upon selecting informative keyframes. Our study reveals the importance of information pre-filtering in video-based MLLMs. Our codes are available at https://github.com/ncTimTang/AKS
Jihao Qiu, Lingxi Xie, Yunjie Tian, Jianbin Jiao, Qixiang Ye
CVPR3
2025 CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic Segmentation
Jiannan Ge, Lingxi Xie, Hongtao Xie 0001, Pandeng Li, Sun-Ao Liu, Xiaopeng Zhang 0008, Qi Tian 0001, Yongdong Zhang 0001
ICCV2
2025 SAM-CP: Marrying SAM with Composable Prompts for Versatile Segmentation
abstract
The Segment Anything model (SAM) has shown a generalized ability to group image pixels into patches, but applying it to semantic-aware segmentation still faces major challenges. This paper presents SAM-CP, a simple approach that establishes two types of composable prompts beyond SAM and composes them for versatile segmentation. Specifically, given a set of classes (in texts) and a set of SAM patches, the Type-I prompt judges whether a SAM patch aligns with a text label, and the Type-II prompt judges whether two SAM patches with the same text label also belong to the same instance. To decrease the complexity in dealing with a large number of semantic classes and patches, we establish a unified framework that calculates the affinity between (semantic and instance) queries and SAM patches, and then merges patches with high affinity to the query. Experiments show that SAM-CP achieves semantic, instance, and panoptic segmentation in both open and closed domains. In particular, it achieves state-of-the-art performance in open-vocabulary segmentation. Our research offers a novel and generalized methodology for equipping vision foundation models like SAM with multi-grained semantic perception abilities. Codes are released on https://github.com/ucas-vg/SAM-CP.
Pengfei Chen 0004, Lingxi Xie, Xinyue Huo, Xuehui Yu, Xiaopeng Zhang 0008, Yingfei Sun, Zhenjun Han, Qi Tian 0001
ICLR2
2025 ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
abstract
Aligning vision and language concepts at a finer level remains an essential topic of multimodal large language models (MLLMs), particularly for tasks such as referring and grounding. Existing methods, such as *proxy encoding* and *geometry encoding* genres, incorporate additional syntax to encode spatial information, imposing extra burdens when communicating between language with vision modules. In this study, we propose ClawMachine, offering a new methodology that explicitly notates each entity using **token collectives**—groups of visual tokens that collaboratively represent higher-level semantics. A hybrid perception mechanism is also explored to perceive and understand scenes from both discrete and continuous spaces. Our method unifies the prompt and answer of visual referential tasks without using additional syntax. By leveraging a joint vision-language vocabulary, ClawMachine integrates referring and grounding in an auto-regressive manner, demonstrating great potential with scaled up pre-training data. Experiments show that ClawMachine achieves superior performance on scene-level and referential understanding tasks with higher efficiency. It also exhibits the potential to integrate multi-source information for complex visual reasoning, which is beyond the capability of many MLLMs. Our code is available at https://github.com/martian422/ClawMachine.
Tianren Ma, Lingxi Xie, Yunjie Tian, Boyu Yang 0002, Qixiang Ye
ICLR2
2025 Tackling View-Dependent Semantics in 3D Language Gaussian Splatting
abstract
Recent advancements in 3D Gaussian Splatting (3D-GS) enable high-quality 3D scene reconstruction from RGB images. Many studies extend this paradigm for language-driven open-vocabulary scene understanding. However, most of them simply project 2D semantic features onto 3D Gaussians and overlook a fundamental gap between 2D and 3D understanding: a 3D object may exhibit various semantics from different viewpoints—a phenomenon we term **view-dependent semantics**. To address this challenge, we propose **LaGa** (**La**nguage **Ga**ussians), which establishes cross-view semantic connections by decomposing the 3D scene into objects. Then, it constructs view-aggregated semantic representations by clustering semantic descriptors and reweighting them based on multi-view semantics. Extensive experiments demonstrate that LaGa effectively captures key information from view-dependent semantics, enabling a more comprehensive understanding of 3D scenes. Notably, under the same settings, LaGa achieves a significant improvement of **+18.7\% mIoU** over the previous SOTA on the LERF-OVS dataset. Our code is available at: https://github.com/https://github.com/SJTU-DeepVisionLab/LaGa.
Jiazhong Cen, Jiemin Fang, Changsong Wen, Lingxi Xie, Xiaopeng Zhang 0008, Wei Shen 0002, Qi Tian 0001
ICML5
2025 Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
abstract
Multimodal Large Language Models (MLLMs) are experiencing rapid growth, yielding a plethora of novel works recently. The prevailing trend involves adopting data-driven methodologies, wherein diverse instruction-following datasets were collected. However, these approaches always face the challenge of limited visual perception capabilities, as they solely utilizing CLIP-like encoders to extract visual information from inputs. Though these encoders are pre-trained on billions of image-text pairs, they still grapple with the information loss dilemma, given that textual captions only partially capture the contents depicted in images. To address this limitation, this paper proposes to improve the visual perception ability of MLLMs through a mixture-of-experts knowledge enhancement mechanism. Specifically, this work introduces a novel method that incorporates multi-task encoders and existing visual tools into the MLLMs training and inference pipeline, aiming to provide a more comprehensive summarization of visual inputs. Extensive experiments have evaluated its effectiveness of advancing MLLMs, showcasing improved visual perception capability achieved through the integration of visual experts.
Longhui Wei, Lingxi Xie, Qi Tian 0001
IJCAI3
2025 Segment Anything in 3D with Radiance Fields
Jiazhong Cen, Jiemin Fang, Zanwei Zhou, Chen Yang 0023, Lingxi Xie, Xiaopeng Zhang 0008, Wei Shen 0002, Qi Tian 0001
Int. J. Comput. Vis.5
2025 P2Object: Single Point Supervised Object Detection and Instance Segmentation
Pengfei Chen 0004, Xuehui Yu, Xumeng Han, Kuiran Wang, Guorong Li, Lingxi Xie, Zhenjun Han, Jianbin Jiao
Int. J. Comput. Vis.6
2025 Beyond masking: Demystifying token-based pre-training for vision transformers
Yunjie Tian, Lingxi Xie, Jiemin Fang, Jianbin Jiao, Qi Tian 0001
Pattern Recognit.2
2025 Denoised and Dynamic Alignment Enhancement for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) focuses on recognizing unseen categories by aligning visual features with semantic information. Recent advancements have shown that aligning each attribute with its corresponding visual region significantly improves zero-shot learning performance. However, the crude semantic proxies used in these methods fail to capture the varied appearances of each attribute, and are also easily confused by the presence of semantically redundant backgrounds, leading to suboptimal alignment. To combat these issues, we introduce a novel Alignment-Enhanced Network (AENet), designed to denoise the visual features and dynamically perceive semantic information, thus enhancing visual-semantic alignment. Our approach comprises two key innovations. (1) A visual denoising encoder, employing a class-agnostic mask to filter out semantically redundant visual information, thus producing refined visual features adaptable to unseen classes. (2) A dynamic semantic generator that crafts content-aware semantic proxies adaptively, steered by visual features, enabling AENet to discriminate fine-grained variations in visual contents. Additionally, we integrate a cross-fusion module to ensure comprehensive interaction between the denoised visual features and the generated dynamic semantic proxies, further facilitating visual-semantic alignment. Through extensive experiments across three datasets, the proposed method demonstrates that it narrows down the visual-semantic gap and sets a new benchmark in this setting.
Jiannan Ge, Pandeng Li, Lingxi Xie, Yongdong Zhang 0001, Qi Tian 0001, Hongtao Xie 0001
IEEE Trans. Image Process.4
2025 Exploring Complicated Search Spaces With Interleaving-Free Sampling
abstract
Conventional neural architecture search (NAS) algorithms typically work on search spaces with short-distance node connections. We argue that such designs, though safe and stable, are obstacles to exploring more effective network architectures. In this brief, we explore the search algorithm upon a complicated search space with long-distance connections and show that existing weight-sharing search algorithms fail due to the existence of interleaved connections (ICs). Based on the observation, we present a simple-yet-effective algorithm, termed interleaving-free neural architecture search (IF-NAS). We further design a periodic sampling strategy to construct subnetworks during the search procedure, avoiding the ICs to emerge in any of them. In the proposed search space, IF-NAS outperforms both random sampling and previous weight-sharing search algorithms by significant margins. It can also be well-generalized to the microcell-based spaces. This study emphasizes the importance of macrostructure and we look forward to further efforts in this direction. The code is available at github.com/sunsmarterjie/IFNAS.
Yunjie Tian, Lingxi Xie, Jiemin Fang, Jianbin Jiao, Qixiang Ye, Qi Tian 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model
abstract
Parameter-efficient fine-tuning (PEFT) is an effective methodology to unleash the potential of large foundation models in novel scenarios with limited training data. In the computer vision community, PEFT has shown effectiveness in image classification, but little research has studied its ability for image segmentation. Fine-tuning segmentation models usually requires a heavier adjustment of parameters to align the proper projection directions in the parameter space for new scenarios. This raises a challenge to existing PEFT algorithms, as they often inject a limited number of individual parameters into each block, which prevents substantial adjustment of the projection direction of the parameter space due to the limitation of Hidden Markov Chain along blocks. In this paper, we equip PEFT with a cross-block orchestration mechanism to enable the adaptation of the Segment Anything Model (SAM) to various downstream scenarios. We introduce a novel inter-block communication module, which integrates a learnable relation matrix to facilitate communication among different coefficient sets of each PEFT block's parameter space. Moreover, we propose an intra-block enhancement module, which introduces a linear projection head whose weights are generated from a hyper-complex layer, further enhancing the impact of the adjustment of projection directions on the entire parameter space. Extensive experiments on diverse benchmarks demonstrate that our proposed approach consistently improves the segmentation performance significantly on novel scenarios with only around 1K additional parameters.
Zelin Peng, Zhengqin Xu, Zhilin Zeng, Lingxi Xie, Qi Tian 0001, Wei Shen 0002
CVPR4
2024 GaussianEditor: Editing 3D Gaussians Delicately with Text Instructions
abstract
Recently, impressive results have been achieved in 3D scene editing with text instructions based on a 2D diffusion model. However, current diffusion models primarily generate images by predicting noise in the latent space, and the editing is usually applied to the whole image, which makes it challenging to perform delicate, especially localized, editing for 3D scenes. Inspired by recent 3D Gaussian splatting, we propose a systematic framework, named Gaus-sianEditor, to edit 3D scenes delicately via 3D Gaussians with text instructions. Benefiting from the explicit property of 3D Gaussians, we design a series of techniques to achieve delicate editing. Specifically, we first extract the region of interest (RoI) corresponding to the text instruction, aligning it to 3D Gaussians. The Gaussian RoI is further used to control the editing process. Our framework can achieve more delicate and precise editing of 3D scenes than previous methods while enjoying much faster training speed, i.e. within 20 minutes on a single V100 GPU, more than twice as fast as Instruct-NeRF2NeRF (45 minutes - 2 hours)11The editing time varies in different scenes according to the scene structure complexity.. The project page is at GaussianEditor. github.io.
Junjie Wang 0012, Jiemin Fang, Xiaopeng Zhang 0008, Lingxi Xie, Qi Tian 0001
CVPR4
2024 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
abstract
Representing and rendering dynamic scenes has been an important but challenging task. Especially, to accurately model complex motions, high efficiency is usually hard to guarantee. To achieve real-time dynamic scene rendering while also enjoying high training and storage efficiency, we propose 4D Gaussian Splatting (4D-GS) as a holistic representation for dynamic scenes rather than applying 3D-GS for each individual frame. In 4D-GS, a novel explicit representation containing both 3D Gaussians and 4D neural voxels is proposed. A decomposed neural voxel encoding algorithm inspired by HexPlane is proposed to efficiently build Gaussian features from 4D neural voxels and then a lightweight MLP is applied to predict Gaussian deformations at novel timestamps. Our 4D-GS method achieves real-time rendering under high resolutions, 82 FPS at an 800x800 resolution on an RTX 3090 GPU while maintaining comparable or better quality than previous state- of-the-art methods. More demos and code are available at https://guanjunwu.github.io/4dgs/.
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang 0008, Wei Wei 0002, Wenyu Liu 0001, Qi Tian 0001, Xinggang Wang
CVPR4
2024 GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models
abstract
In recent times, the generation of 3D assets from text prompts has shown impressive results. Both 2D and 3D diffusion models can help generate decent 3D objects based on prompts. 3D diffusion models have good 3D consistency, but their quality and generalization are limited as trainable 3D data is expensive and hard to obtain. 2D diffusion models enjoy strong abilities of generalization and fine generation, but 3D consistency is hard to guarantee. This paper attempts to bridge the power from the two types of diffusion models via the recent explicit and efficient 3D Gaussian splatting representation. A fast 3D object gener-ation framework, named as GaussianDreamer, is proposed, where the 3D diffusion model provides priors for initial-ization and the 2D diffusion model enriches the geometry and appearance. Operations of noisy point growing and color perturbation are introduced to enhance the initialized Gaussians. Our GaussianDreamer can generate a high-quality 3D instance or 3D avatar within 15 minutes on one GPU, much faster than previous methods, while the generated instances can be directly rendered in real time. Demos and code are available at https://taoranyi.com/gaussiandreamer/.
Taoran Yi, Jiemin Fang, Junjie Wang 0012, Guanjun Wu, Lingxi Xie, Xiaopeng Zhang 0008, Wenyu Liu 0001, Qi Tian 0001, Xinggang Wang
CVPR5
2024 Cascade-Zero123: One Image to Highly Consistent 3D with Self-prompted Nearby Views
Yabo Chen, Jiemin Fang, Taoran Yi, Xiaopeng Zhang 0008, Lingxi Xie, Xinggang Wang, Wenrui Dai, Hongkai Xiong, Qi Tian 0001
ECCV (41)6
2024 AlignZeg: Mitigating Objective Misalignment for Zero-Shot Semantic Segmentation
Jiannan Ge, Lingxi Xie, Hongtao Xie 0001, Pandeng Li, Xiaopeng Zhang 0008, Yongdong Zhang 0001, Qi Tian 0001
ECCV (43)2
2024 Domain-Adaptive Semantic Segmentation Emerges From Vision-Language Supervised Domain-Debiased Self-Training
abstract
Unsupervised domain adaptive semantic segmentation leverages synthetic data to train a segmentation model and transfers it to unlabeled real images. Due to the style difference, the transferred model suffers from the domain gap. Even worse, some classes exhibit the extreme domain gap, where the feature distributions undergo a complete shift between the two domains. To alleviate it, we propose a domain-debiased self-training strategy with CLIP to distill its domain-agnostic knowledge. Specifically, we enforce the consistency between the feature maps from our segmentation model and the image encoder of CLIP. Meanwhile, the text embeddings from the text encoder for each class serve as a domain-agnostic classifier to support a domain-debiased feature learning condition. Experimental results under standard UDA settings demonstrate that our proposed strategy consistently improves the UDA segmentation performance based on different backbones and with different large pre-trained models.
Huayu Wang, Zekun Jiang, Lingxi Xie, Dongsheng Jiang, Wei Shen 0002, Qi Tian 0001
ICASSP3
2024 QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models
abstract
Recently years have witnessed a rapid development of large language models (LLMs). Despite the strong ability in many language-understanding tasks, the heavy computational burden largely restricts the application of LLMs especially when one needs to deploy them onto edge devices. In this paper, we propose a quantization-aware low-rank adaptation (QA-LoRA) algorithm. The motivation lies in the imbalanced degrees of freedom of quantization and adaptation, and the solution is to use group-wise operators which increase the degree of freedom of quantization meanwhile decreasing that of adaptation. QA-LoRA is easily implemented with a few lines of code, and it equips the original LoRA with two-fold abilities: (i) during fine-tuning, the LLM's weights are quantized (e.g., into INT4) to reduce time and memory usage; (ii) after fine-tuning, the LLM and auxiliary weights are naturally integrated into a quantized model without loss of accuracy. We apply QA-LoRA to the LLaMA and LLaMA2 model families and validate its effectiveness in different fine-tuning datasets and downstream scenarios. The code is made available at https://github.com/yuhuixu1993/qa-lora.
Yuhui Xu 0002, Lingxi Xie, Xiaotao Gu, Xin Chen 0033, Heng Chang, Hengheng Zhang, Zhengsu Chen, Xiaopeng Zhang 0008, Qi Tian 0001
ICLR2
2024 VMamba: Visual State Space Model
abstract
Designing computationally efficient network architectures remains an ongoing necessity in computer vision. In this paper, we adapt Mamba, a state-space language model, into VMamba, a vision backbone with linear time complexity. At the core of VMamba is a stack of Visual State-Space (VSS) blocks with the 2D Selective Scan (SS2D) module. By traversing along four scanning routes, SS2D bridges the gap between the ordered nature of 1D selective scan and the non-sequential structure of 2D vision data, which facilitates the collection of contextual information from various sources and perspectives. Based on the VSS blocks, we develop a family of VMamba architectures and accelerate them through a succession of architectural and implementation enhancements. Extensive experiments demonstrate VMamba’s promising performance across diverse visual perception tasks, highlighting its superior input scaling efficiency compared to existing benchmark models. Source code is available at https://github.com/MzeroMiko/VMamba
Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang 0001, Qixiang Ye, Jianbin Jiao, Yunfan Liu 0001
NeurIPS5
2024 Artemis: Towards Referential Understanding in Complex Videos
abstract
Videos carry rich visual information including object description, action, interaction, etc., but the existing multimodal large language models (MLLMs) fell short in referential understanding scenarios such as video-based referring. In this paper, we present Artemis, an MLLM that pushes video-based referential understanding to a finer level. Given a video, Artemis receives a natural-language question with a bounding box in any video frame and describes the referred target in the entire video. The key to achieving this goal lies in extracting compact, target-specific video features, where we set a solid baseline by tracking and selecting spatiotemporal features from the video. We train Artemis on the newly established ViderRef45K dataset with 45K video-QA pairs and design a computationally efficient, three-stage training procedure. Results are promising both quantitatively and qualitatively. Additionally, we show that Artemis can be integrated with video grounding and text summarization tools to understand more complex scenarios. Code and data are available at https://github.com/NeurIPS24Artemis/Artemis.
Jihao Qiu, Lingxi Xie, Tianren Ma, Pengyu Yan, David S. Doermann, Qixiang Ye, Yunjie Tian
NeurIPS4
2024 Domain-Agnostic Priors for Semantic Segmentation Under Unsupervised Domain Adaptation and Domain Generalization
Xinyue Huo, Lingxi Xie, Hengtong Hu, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001
Int. J. Comput. Vis.2
2024 CenterNet++ for Object Detection
abstract
There are two mainstream approaches for object detection: top-down and bottom-up. The state-of-the-art approaches are mainly top-down methods. In this paper, we demonstrate that bottom-up approaches show competitive performance compared with top-down approaches and have higher recall rates. Our approach, named CenterNet, detects each object as a triplet of keypoints (top-left and bottom-right corners and the center keypoint). We first group the corners according to some designed cues and confirm the object locations based on the center keypoints. The corner keypoints allow the approach to detect objects of various scales and shapes and the center keypoint reduces the confusion introduced by a large number of false-positive proposals. Our approach is an anchor-free detector because it does not need to define explicit anchor boxes. We adapt our approach to backbones with different structures, including 'hourglass'-like networks and 'pyramid'-like networks, which detect objects in single-resolution and multi-resolution feature maps, respectively. On the MS-COCO dataset, CenterNet with Res2Net-101 and Swin-Transformer achieve average precisions (APs) of 53.7% and 57.1%, respectively, outperforming all existing bottom-up detectors and achieving state-of-the-art performance. We also design a real-time CenterNet model, which achieves a good trade-off between accuracy and speed, with an AP of 43.6% at 30.5 frames per second (FPS).
Kaiwen Duan, Song Bai 0001, Lingxi Xie, Honggang Qi, Qingming Huang, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network With Token Migration
abstract
We propose integrally pre-trained transformer pyramid network (iTPN), towards jointly optimizing the network backbone and the neck, so that transfer gap between representation models and downstream tasks is minimal. iTPN is born with two elaborated designs: 1) The first pre-trained feature pyramid upon vision transformer (ViT). 2) Multi-stage supervision to the feature pyramid using masked feature modeling (MFM). iTPN is updated to Fast-iTPN, reducing computational memory overhead and accelerating inference through two flexible designs. 1) Token migration: dropping redundant tokens of the backbone while replenishing them in the feature pyramid without attention operations. 2) Token gathering: reducing computation cost caused by global attention by introducing few gathering tokens. The base/large-level Fast-iTPN achieve 88.75%/89.5% top-1 accuracy on ImageNet-1 K. With 1× training schedule using DINO, the base/large-level Fast-iTPN achieves 58.4%/58.8% box AP on COCO object detection, and a 57.5%/58.7% mIoU on ADE20 K semantic segmentation using MaskDINO. Fast-iTPN can accelerate the inference procedure by up to 70%, with negligible performance loss, demonstrating the potential to be a powerful backbone for downstream vision tasks.
Yunjie Tian, Lingxi Xie, Jihao Qiu, Jianbin Jiao, Yaowei Wang 0001, Qi Tian 0001, Qixiang Ye
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Consensus Synergizes With Memory: A Simple Approach for Anomaly Segmentation in Urban Scenes
abstract
Anomaly segmentation is a critical task for safety-critical applications, such as autonomous driving in urban environments. Its objective is to detect out-of-distribution (OOD) samples with unseen categories, given a pre-trained segmentation model. The core challenge of this task is how to distinguish hard in-distribution samples from OOD samples, which has not been explicitly discussed in previous research. In this paper, we propose a simple yet effective approach named CosMe (Consensus Synergizes with Memory) to address this challenge. CosMe consists of two key components: 1) building a memory bank comprising seen prototypes extracted from multiple layers of the given segmentation model, and 2) training an auxiliary model that mimics the behavior of the given model and using the consensus of their mid-level features as complementary cues that synergize with the memory bank. The former serves as a baseline that can detect all potential outliers, including both OOD and hard in-distribution samples; the latter assists in distinguishing between these two types of outliers. Experimental results on several urban scene anomaly segmentation datasets demonstrate that CosMe outperforms previous approaches by a significant margin.
Jiazhong Cen, Zekun Jiang, Lingxi Xie, Dongsheng Jiang, Wei Shen 0002, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Multi-Granularity Matching Transformer for Text-Based Person Search
abstract
Text-based person search aims to retrieve the most relevant pedestrian images from an image gallery based on textual descriptions. Most existing methods rely on two separate encoders to extract the image and text features, and then elaborately design various schemes to bridge the gap between image and text modalities. However, the shallow interaction between both modalities in these methods is still insufficient to eliminate the modality gap. To address the above problem, we propose TransTPS, a transformer-based framework that enables deeper interaction between both modalities through the self-attention mechanism in transformer, effectively alleviating the modality gap. In addition, due to the small inter-class variance and large intra-class variance in image modality, we further develop two techniques to overcome these limitations. Specifically, Cross-modal Multi-Granularity Matching (CMGM) is proposed to address the problem caused by small inter-class variance and facilitate distinguishing pedestrians with similar appearance. Besides, Contrastive Loss with Weakly Positive pairs (CLWP) is introduced to mitigate the impact of large intra-class variance and contribute to the retrieval of more target images. Experiments on CUHK-PEDES and RSTPReID datasets demonstrate that our proposed framework achieves state-of-the-art performance compared to previous methods.
Liping Bao, Longhui Wei, Wengang Zhou 0001, Lin Liu 0016, Lingxi Xie, Houqiang Li, Qi Tian 0001
IEEE Trans. Multim.5
2024 Towards Discriminative Feature Generation for Generalized Zero-Shot Learning
abstract
Generalized Zero-Shot Learning (GZSL) aims to recognize both seen and unseen categories by establishing visual and semantic relations. Recently, generation-based methods that focus on synthesizing fictitious visual features from corresponding attributes have gained significant attention. However, these generated features often lack discriminative capabilities due to inadequate training of the generative model. To address this issue, we propose a novel Discriminative Enhanced Network (DENet) to harness the potential of the generative model by adapting the training features and imposing constraints on the generated features. Our approach incorporates three pivotal modules: (1) Before the generative network training, we implement a Pre-Tuning Module (PTM) to eliminate irrelevant background noise in the raw features extracted from a fixed CNN backbone. Therefore, PTM can provide tuned training features without redundant noise for generative model. (2) During the generative network training, we propose an Asymmetry Cross-authenticity Contrastive (AC2) loss to group visual features of the same category while repel features from different categories by optimizing a large number of sample pairs. Additionally, we incorporate intra-class and relation-specific inter-class boundaries within the AC2 loss to enrich sample diversity and preserve valid semantic information. (3) Also within the generative network training, a Dual-semantic Alignment Module (DAM) is designed to align visual features with both attributes and label embeddings, enabling the model to learn attribute-related information and discriminative extended semantics. Experiments on four standard benchmarks demonstrate that our approach learns more discriminative features and surpasses the existing methods.
Jiannan Ge, Hongtao Xie 0001, Pandeng Li, Lingxi Xie, Shaobo Min, Yongdong Zhang 0001
IEEE Trans. Multim.4
2024 GaussianObject: High-Quality 3D Object Reconstruction from Four Views with Gaussian Splatting
abstract
Reconstructing and rendering 3D objects from highly sparse views is of critical importance for promoting applications of 3D vision techniques and improving user experience. However, images from sparse views only contain very limited 3D information, leading to two significant challenges: 1) Difficulty in building multi-view consistency as images for matching are too few; 2) Partially omitted or highly compressed object information as view coverage is insufficient. To tackle these challenges, we propose GaussianObject, a framework to represent and render the 3D object with Gaussian splatting that achieves high rendering quality with only 4 input images. We first introduce techniques of visual hull and floater elimination, which explicitly inject structure priors into the initial optimization process to help build multi-view consistency, yielding a coarse 3D Gaussian representation. Then we construct a Gaussian repair model based on diffusion models to supplement the omitted object information, where Gaussians are further refined. We design a self-generating strategy to obtain image pairs for training the repair model. We further design a COLMAP-free variant, where pre-given accurate camera poses are not required, which achieves competitive quality and facilitates wider applications. GaussianObject is evaluated on several challenging datasets, including MipNeRF360, OmniObject3D, OpenIllumination, and our-collected unposed images, achieving superior performance from only four views and significantly outperforming previous SOTA methods.
Chen Yang 0023, Sikuang Li, Jiemin Fang, Ruofan Liang, Lingxi Xie, Xiaopeng Zhang 0008, Wei Shen 0002, Qi Tian 0001
ACM Trans. Graph.5
2024 One-Bit Supervision for Image Classification: Problem, Solution, and Beyond
abstract
This article presents one-bit supervision, a novel setting of learning with fewer labels, for image classification. Instead of the training model using the accurate label of each sample, our setting requires the model to interact with the system by predicting the class label of each sample and learn from the answer whether the guess is correct, which provides one bit (yes or no) of information. An intriguing property of the setting is that the burden of annotation largely is alleviated in comparison to offering the accurate label. There are two keys to one-bit supervision: (i) improving the guess accuracy and (ii) making good use of the incorrect guesses. To achieve these goals, we propose a multi-stage training paradigm and incorporate negative label suppression into an off-the-shelf semi-supervised learning algorithm. Theoretical analysis shows that one-bit annotation is more efficient than full-bit annotation in most cases and gives the conditions of combining our approach with active learning. Inspired by this, we further integrate the one-bit supervision framework into the self-supervised learning algorithm, which yields an even more efficient training schedule. Different from training from scratch, when self-supervised learning is used for initialization, both hard example mining and class balance are verified to be effective in boosting the learning performance. However, these two frameworks still need full-bit labels in the initial stage. To cast off this burden, we utilize unsupervised domain adaptation to train the initial model and conduct pure one-bit annotations on the target dataset. In multiple benchmarks, the learning efficiency of the proposed approach surpasses that using full-bit, semi-supervised supervision.
Hengtong Hu, Lingxi Xie, Xinyue Huo, Richang Hong, Qi Tian 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Visual Recognition by Request
abstract
Humans have the ability of recognizing visual semantics in an unlimited granularity, but existing visual recognition algorithms cannot achieve this goal. In this paper, we establish a new paradigm named visual recognition by request (ViRReq11We recommend the readers to pronounce ViRReqas/virik/.) to bridge the gap. The key lies in decomposing visual recognition into atomic tasks named requests and leveraging a knowledge base, a hierarchical and text-based dictionary, to assist task definition. ViRReq allows for (i) learning complicated whole-part hierarchies from highly incomplete annotations and (ii) inserting new concepts with minimal efforts. We also establish a solid baseline by integrating language-driven recognition into recent semantic and instance segmentation methods, and demonstrate its flexible recognition ability on CPP and ADE20K, two datasets with hierarchical whole-part annotations.
Chufeng Tang, Lingxi Xie, Xiaopeng Zhang 0008, Xiaolin Hu 0001, Qi Tian 0001
CVPR2
2023 Integrally Pre-Trained Transformer Pyramid Networks
abstract
In this paper, we present an integral pre-training framework based on masked image modeling (MIM). We advocate for pre-training the backbone and neck jointly so that the transfer gap between MIM and downstream recognition tasks is minimal. We make two technical contributions. First, we unify the reconstruction and recognition necks by inserting a feature pyramid into the pre-training stage. Second, we complement mask image modeling (MIM) with masked feature modeling (MFM) that offers multi-stage supervision to the feature pyramid. The pre-trained models, termed integrally pre-trained transformer pyramid networks (iTPNs), serve as powerful foundation models for visual recognition. In particular, the base/large-level iTPN achieves an 86.2%/87.8% top-1 accuracy on ImageNet-1K, a 53.2%/55.6% box AP on COCO object detection with 1× training schedule using Mask-RCNN, and a 54.7%/57.7% mIoU on ADE20K semantic segmentation using UPerHead – all these results set new records. Our work inspires the community to work on unifying upstream pre-training and downstream fine-tuning tasks. Code is available at github.com/sunsmarterjie/iTPN.
Yunjie Tian, Lingxi Xie, Zhaozhi Wang, Longhui Wei, Xiaopeng Zhang 0008, Jianbin Jiao, Yaowei Wang 0001, Qi Tian 0001, Qixiang Ye
CVPR2
2023 Focus on Your Target: A Dual Teacher-Student Framework for Domain-adaptive Semantic Segmentation
abstract
We study unsupervised domain adaptation (UDA) for semantic segmentation. Currently, a popular UDA framework lies in self-training which endows the model with two-fold abilities: (i) learning reliable semantics from the labeled images in the source domain, and (ii) adapting to the target domain via generating pseudo labels on the unlabeled images. We find that, by decreasing/increasing the proportion of training samples from the target domain, the ‘learning ability’ is strengthened/weakened while the ‘adapting ability’ goes in the opposite direction, implying a conflict between these two abilities, especially for a single model. To alleviate the issue, we propose a novel dual teacher-student (DTS) framework and equip it with a bidirectional learning strategy. By increasing the proportion of target-domain data, the second teacher-student model learns to ‘Focus on Your Target’ while the first model is not affected. DTS is easily plugged into existing self-training approaches. In a standard UDA scenario (training on synthetic, labeled data and real, unlabeled data), DTS shows consistent gains over the baselines and sets new state-of-the-art results of 76.5% and 75.1% mIoUs on GTAv→Cityscapes and SYNTHIA→Cityscapes, respectively. The implementation is available at https://github.com/xinyuehuo/DTS.
Xinyue Huo, Lingxi Xie, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001
ICCV2
2023 USAGE: A Unified Seed Area Generation Paradigm for Weakly Supervised Semantic Segmentation
abstract
Seed area generation is usually the starting point of weakly supervised semantic segmentation (WSSS). Computing the Class Activation Map (CAM) from a multi-label classification network is the de facto paradigm for seed area generation, but CAMs generated from Convolutional Neural Networks (CNNs) and Transformers are prone to be under- and over-activated, respectively, which makes the strategies to refine CAMs for CNNs usually inappropriate for Transformers, and vice versa. In this paper, we propose a Unified optimization paradigm for Seed Area GEneration (USAGE) for both types of networks, in which the objective function to be optimized consists of two terms: One is a generation loss, which controls the shape of seed areas by a temperature parameter following a deterministic principle for different types of networks; The other is a regularization loss, which ensures the consistency between the seed areas that are generated by self-adaptive network adjustment from different views, to overturn false activation in seed areas. Experimental results show that USAGE consistently improves seed area generation for both CNNs and Transformers by large margins, e.g., outperforming state-of-the-art methods by a mIoU of 4.1% on PASCAL VOC. Moreover, based on the USAGE-generated seed areas on Transformers, we achieve state-of-the-art WSSS results on both PASCAL VOC and MS COCO.
Zelin Peng, Guanchun Wang, Lingxi Xie, Dongsheng Jiang, Wei Shen 0002, Qi Tian 0001
ICCV3
2023 HiViT: A Simpler and More Efficient Design of Hierarchical Vision Transformer
Xiaosong Zhang 0004, Yunjie Tian, Lingxi Xie, Qi Dai 0001, Qixiang Ye, Qi Tian 0001
ICLR3
2023 Segment Anything in 3D with NeRFs
abstract
Recently, the Segment Anything Model (SAM) emerged as a powerful vision foundation model which is capable to segment anything in 2D images. This paper aims to generalize SAM to segment 3D objects. Rather than replicating the data acquisition and annotation procedure which is costly in 3D, we design an efficient solution, leveraging the Neural Radiance Field (NeRF) as a cheap and off-the-shelf prior that connects multi-view 2D images to the 3D space. We refer to the proposed solution as SA3D, for Segment Anything in 3D. It is only required to provide a manual segmentation prompt (e.g., rough points) for the target object in a single view, which is used to generate its 2D mask in this view with SAM. Next, SA3D alternately performs mask inverse rendering and cross-view self-prompting across various views to iteratively complete the 3D mask of the target object constructed with voxel grids. The former projects the 2D mask obtained by SAM in the current view onto 3D mask with guidance of the density distribution learned by the NeRF; The latter extracts reliable prompts automatically as the input to SAM from the NeRF-rendered 2D mask in another view. We show in experiments that SA3D adapts to various scenes and achieves 3D segmentation within minutes. Our research offers a generic and efficient methodology to lift a 2D vision foundation model to 3D, as long as the 2D model can steadily address promptable segmentation across multiple views.
Jiazhong Cen, Zanwei Zhou, Jiemin Fang, Chen Yang 0023, Wei Shen 0002, Lingxi Xie, Dongsheng Jiang, Xiaopeng Zhang 0008, Qi Tian 0001
NeurIPS6
2023 VoxSeP: semi-positive voxels assist self-supervised 3D medical segmentation
Zijie Yang, Lingxi Xie, Xinyue Huo, Longhui Wei, Qi Tian 0001, Sheng Tang
Multim. Syst.2
2023 Learnable Distribution Calibration for Few-Shot Class-Incremental Learning
abstract
Few-shot class-incremental learning (FSCIL) faces the challenges of memorizing old class distributions and estimating new class distributions given few training samples. In this study, we propose a learnable distribution calibration (LDC) approach, to systematically solve these two challenges using a unified framework. LDC is built upon a parameterized calibration unit (PCU), which initializes biased distributions for all classes based on classifier vectors (memory-free) and a single covariance matrix. The covariance matrix is shared by all classes, so that the memory costs are fixed. During base training, PCU is endowed with the ability to calibrate biased distributions by recurrently updating sampled features under supervision of real distributions. During incremental learning, PCU recovers distributions for old classes to avoid 'forgetting', as well as estimating distributions and augmenting samples for new classes to alleviate 'over-fitting' caused by the biased distributions of few-shot samples. LDC is theoretically plausible by formatting a variational inference procedure. It improves FSCIL's flexibility as the training procedure requires no class similarity priori. Experiments on CUB200, CIFAR100, and mini-ImageNet datasets show that LDC respectively outperforms the state-of-the-arts by 4.64%, 1.98%, and 3.97%. LDC's effectiveness is also validated on few-shot learning scenarios.
Binghao Liu, Boyu Yang 0002, Lingxi Xie, Qi Tian 0001, Qixiang Ye
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 GAIA-Universe: Everything is Super-Netify
abstract
Pre-training on large-scale datasets has played an increasingly significant role in computer vision and natural language processing recently. However, as there exist numerous application scenarios that have distinctive demands such as certain latency constraints and specialized data distributions, it is prohibitively expensive to take advantage of large-scale pre-training for per-task requirements. we focus on two fundamental perception tasks (object detection and semantic segmentation) and present a complete and flexible system named GAIA-Universe(GAIA), which could automatically and efficiently give birth to customized solutions according to heterogeneous downstream needs through data union and super-net training. GAIA is capable of providing powerful pre-trained weights and searching models that conform to downstream demands such as hardware constraints, computation constraints, specified data domains, and telling relevant data for practitioners who have very few datapoints on their tasks. With GAIA, we achieve promising results on COCO, Objects365, Open Images, BDD100 k, and UODB which is a collection of datasets including KITTI, VOC, WiderFace, DOTA, Clipart, Comic, and more. Taking COCO as an example, GAIA is able to efficiently produce models covering a wide range of latency from 16 ms to 53 ms, and yields AP from 38.2 to 46.5 without whistles and bells. GAIA is released at https://github.com/GAIA-vision.
Junran Peng, Xingyuan Bu, Lingxi Xie, Xiaopeng Zhang 0008, Qi Tian 0001, Zhaoxiang Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Conformer: Local Features Coupling Global Representations for Recognition and Detection
abstract
With convolution operations, Convolutional Neural Networks (CNNs) are good at extracting local features but experience difficulty to capture global representations. With cascaded self-attention modules, vision transformers can capture long-distance feature dependencies but unfortunately deteriorate local feature details. In this paper, we propose a hybrid network structure, termed Conformer, to take both advantages of convolution operations and self-attention mechanisms for enhanced representation learning. Conformer roots in feature coupling of CNN local features and transformer global representations under different resolutions in an interactive fashion. Conformer adopts a dual structure so that local details and global dependencies are retained to the maximum extent. We also propose a Conformer-based detector (ConformerDet), which learns to predict and refine object proposals, by performing region-level feature coupling in an augmented cross-attention fashion. Experiments on ImageNet and MS COCO datasets validate Conformer's superiority for visual recognition and object detection, demonstrating its potential to be a general backbone network.
Zhiliang Peng, Zonghao Guo, Yaowei Wang 0001, Lingxi Xie, Jianbin Jiao, Qi Tian 0001, Qixiang Ye
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 A Survey on Label-Efficient Deep Image Segmentation: Bridging the Gap Between Weak Supervision and Dense Prediction
abstract
The rapid development of deep learning has made a great progress in image segmentation, one of the fundamental tasks of computer vision. However, the current segmentation algorithms mostly rely on the availability of pixel-level annotations, which are often expensive, tedious, and laborious. To alleviate this burden, the past years have witnessed an increasing attention in building label-efficient, deep-learning-based image segmentation algorithms. This paper offers a comprehensive review on label-efficient image segmentation methods. To this end, we first develop a taxonomy to organize these methods according to the supervision provided by different types of weak labels (including no supervision, inexact supervision, incomplete supervision and inaccurate supervision) and supplemented by the types of segmentation problems (including semantic segmentation, instance segmentation and panoptic segmentation). Next, we summarize the existing label-efficient image segmentation methods from a unified perspective that discusses an important question: how to bridge the gap between weak supervision and dense prediction - the current methods are mostly based on heuristic priors, such as cross-pixel similarity, cross-label constraint, cross-view consistency, and cross-image relation. Finally, we share our opinions about the future research directions for label-efficient deep image segmentation.
Wei Shen 0002, Zelin Peng, Huayu Wang, Jiazhong Cen, Dongsheng Jiang, Lingxi Xie, Xiaokang Yang 0001, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 BNET: Batch Normalization With Enhanced Linear Transformation
abstract
Batch normalization (BN) is a fundamental unit in modern deep neural networks. However, BN and its variants focus on normalization statistics but neglect the recovery step that uses linear transformation to improve the capacity of fitting complex data distributions. In this paper, we demonstrate that the recovery step can be improved by aggregating the neighborhood of each neuron rather than just considering a single neuron. Specifically, we propose a simple yet effective method named batch normalization with enhanced linear transformation (BNET) to embed spatial contextual information and improve representation ability. BNET can be easily implemented using the depth-wise convolution and seamlessly transplanted into existing architectures with BN. To our best knowledge, BNET is the first attempt to enhance the recovery step for BN. Furthermore, BN is interpreted as a special case of BNET from both spatial and spectral views. Experimental results demonstrate that BNET achieves consistent performance gains based on various backbones in a wide range of visual tasks. Moreover, BNET can accelerate the convergence of network training and enhance spatial information by assigning important neurons with large weights accordingly.
Yuhui Xu 0002, Lingxi Xie, Cihang Xie, Wenrui Dai, Jieru Mei, Siyuan Qiao, Wei Shen 0002, Hongkai Xiong, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Seed the Views: Hierarchical Semantic Alignment for Contrastive Representation Learning
abstract
Self-supervised learning based on instance discrimination has shown remarkable progress. In particular, contrastive learning, which regards each image as well as its augmentations as an individual class and tries to distinguish them from all other images, has been verified effective for representation learning. However, conventional contrastive learning does not model the relation between semantically similar samples explicitly. In this paper, we propose a general module that considers the semantic similarity among images. This is achieved by expanding the views generated by a single image to Cross-Samples and Multi-Levels, and modeling the invariance to semantically similar images in a hierarchical way. Specifically, the cross-samples are generated by a data mixing operation, which is constrained within samples that are semantically similar, while the multi-level samples are expanded at the intermediate layers of a network. In this way, the contrastive loss is extended to allow for multiple positives per anchor, and explicitly pulling semantically similar images together at different layers of the network. Our method, termed as CSML, has the ability to integrate multi-level representations across samples in a robust way. CSML is applicable to current contrastive based methods and consistently improves the performance. Notably, using MoCo v2 as an instantiation, CSML achieves 76.6% top-1 accuracy with linear evaluation using ResNet-50 as backbone, 66.7% and 75.1% top-1 accuracy with only 1% and 10% labels, respectively. All these numbers set the new state-of-the-art. The code is available at https://github.com/haohang96/CSML.
Haohang Xu, Xiaopeng Zhang 0008, Hao Li 0090, Lingxi Xie, Wenrui Dai, Hongkai Xiong, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Progressive Instance-Aware Feature Learning for Compositional Action Recognition
abstract
In order to enable the model to generalize to unseen "action-objects" (compositional action), previous methods encode multiple pieces of information (i.e., the appearance, position, and identity of visual instances) independently and concatenate them for classification. However, these methods ignore the potential supervisory role of instance information (i.e., position and identity) in the process of visual perception. To this end, we present a novel framework, namely Progressive Instance-aware Feature Learning (PIFL), to progressively extract, reason, and predict dynamic cues of moving instances from videos for compositional action recognition. Specifically, this framework extracts features from foreground instances that are likely to be relevant to human actions (Position-aware Appearance Feature Extraction in Section III-B1), performs identity-aware reasoning among instance-centric features with semantic-specific interactions (Identity-aware Feature Interaction in Section III-B2), and finally predicts instances' position from observed states to force the model into perceiving their movement (Semantic-aware Position Prediction in Section III-B3). We evaluate our approach on two compositional action recognition benchmarks, namely, Something-Else and IKEA-Assembly. Our approach achieves consistent accuracy gain beyond off-the-shelf action recognition algorithms in terms of both ground truth and detected position of instances.
Rui Yan 0010, Lingxi Xie, Xiangbo Shu, Liyan Zhang 0001, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 HiGCIN: Hierarchical Graph-Based Cross Inference Network for Group Activity Recognition
abstract
Group activity recognition (GAR) is a challenging task aimed at recognizing the behavior of a group of people. It is a complex inference process in which visual cues collected from individuals are integrated into the final prediction, being aware of the interaction between them. This paper goes one step further beyond the existing approaches by designing a Hierarchical Graph-based Cross Inference Network (HiGCIN), in which three levels of information, i.e., the body-region level, person level, and group-activity level, are constructed, learned, and inferred in an end-to-end manner. Primarily, we present a generic Cross Inference Block (CIB), which is able to concurrently capture the latent spatiotemporal dependencies among body regions and persons. Based on the CIB, two modules are designed to extract and refine features for group activities at each level. Experiments on two popular benchmarks verify the effectiveness of our approach, particularly in the ability to infer with multilevel visual cues. In addition, training our approach does not require individual action labels to be provided, which greatly reduces the amount of labor required in data annotation.
Rui Yan 0010, Lingxi Xie, Jinhui Tang 0001, Xiangbo Shu, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 CIPS-3D++: End-to-End Real-Time High-Resolution 3D-Aware GANs for GAN Inversion and Stylization
abstract
Style-based GANs achieve state-of-the-art results for generating high-quality images, but lack explicit and precise control over camera poses. Recently proposed NeRF-based GANs have made great progress towards 3D-aware image generation. However, the methods either rely on convolution operators which are not rotationally invariant, or utilize complex yet suboptimal training procedures to integrate both NeRF and CNN sub-structures, yielding un-robust, low-quality images with a large computational burden. This article presents an upgraded version called CIPS-3D++, aiming at high-robust, high-resolution and high-efficiency 3D-aware GANs. On the one hand, our basic model CIPS-3D, encapsulated in a style-based architecture, features a shallow NeRF-based 3D shape encoder as well as a deep MLP-based 2D image decoder, achieving robust image generation/editing with rotation-invariance. On the other hand, our proposed CIPS-3D++, inheriting the rotational invariance of CIPS-3D, together with geometric regularization and upsampling operations, encourages high-resolution high-quality image generation/editing with great computational efficiency. Trained on raw single-view images, without any bells and whistles, CIPS-3D++ sets new records for 3D-aware image synthesis, with an impressive FID of 3.2 on FFHQ at the 1024×1024 resolution. In the meantime, CIPS-3D++ runs efficiently and enjoys a low GPU memory footprint so that it can be trained end-to-end on high-resolution images directly, in contrast to previous alternate/progressive methods. Based on the infrastructure of CIPS-3D++, we propose a 3D-aware GAN inversion algorithm named FlipInversion, which can reconstruct the 3D object from a single-view image. We also provide a 3D-aware stylization method for real images based on CIPS-3D++ and FlipInversion. In addition, we analyze the problem of mirror symmetry suffered in training, and solve it by introducing an auxiliary discriminator for the NeRF network. Overall, CIPS-3D++ provides a strong base model that can serve as a testbed for transferring GAN-based image editing methods from 2D to 3D.
Peng Zhou 0010, Lingxi Xie, Bingbing Ni, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Exploring the diversity and invariance in yourself for visual pre-training task
Longhui Wei, Lingxi Xie, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001
Pattern Recognit.2
2023 M²NAS: Joint Neural Architecture Optimization System With Network Transmission
abstract
Differentiable neural architecture search (NAS) methods have achieved comparable results for low search costs and high performance. Existing differentiable methods focus on searching microstructures in micro space, lacking in exploring macrostructures. However, different networks should have different macrostructures rather than all evenly distributed. This article proposes an M2NAS optimization system to optimize macrostructure and microstructure jointly. Specifically, we initialize a network with maximum complexity, where down-samplings occur at the beginning. Then, iteratively optimize the macrostructure and microstructure. For macro search, we establish a macro space and explore this space by generating candidates and initializing weights and architectural parameters for candidate networks, i.e., network transmission. For micro search, we propose a progressive pruning method to eliminate the vast quantization error caused by one-time pruning. With the number of parameters decreasing during the search, M2NAS obtains a series of networks with different complexity through network selection, forming a Pareto-optimal set. Results show that networks obtained by M2NAS have a better tradeoff between accuracy and complexity. A network searched on ImageNet achieves 77.4% and 76.3% top-1 accuracy, with and without data augment during retraining. When taking different pretrained models as backbones and combining them with Faster RCNN on COCO, our model can get 35.6%$(AP)$and 55.7%$(AP_{50})$, higher than typical NAS models, second only to ResNet-50 results.
Lanfei Wang, Lingxi Xie, Kaifeng Bi, Kaili Zhao, Jun Guo 0002, Qi Tian 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Understanding and Mitigating Overfitting in Prompt Tuning for Vision-Language Models
abstract
Pretrained vision-language models (VLMs) such as CLIP have shown impressive generalization capability in downstream vision tasks with appropriate text prompts. Instead of designing prompts manually, Context Optimization (CoOp) has been recently proposed to learn continuous prompts using task-specific training data. Despite the performance improvements on downstream tasks, several studies have reported that CoOp suffers from the overfitting issue in two aspects: (i) the test accuracy on base classes first improves and then worsens during training; (ii) the test accuracy on novel classes keeps decreasing. However, none of the existing studies can understand and mitigate such overfitting problems. In this study, we first explore the cause of overfitting by analyzing the gradient flow. Comparative experiments reveal that CoOp favors generalizable and spurious features in the early and later training stages, respectively, leading to the non-overfitting and overfitting phenomena. Given those observations, we propose Subspace Prompt Tuning (Sub PT) to project the gradients in back-propagation onto the low-rank subspace spanned by the early-stage gradient flow eigenvectors during the entire training process and successfully eliminate the overfitting problem. In addition, we equip CoOp with a Novel Feature Learner (NFL) to enhance the generalization ability of the learned prompts onto novel categories beyond the training set, needless of image training data. Extensive experiments on 11 classification datasets demonstrate that Sub PT+NFL consistently boost the performance of CoOp and outperform the state-of-the-art CoCoOp approach. Experiments on more challenging vision downstream tasks, including open-vocabulary object detection and zero-shot semantic segmentation, also verify the effectiveness of the proposed method. Codes can be found athttps://tinyurl.com/mpe64f89.
Chengcheng Ma, Yang Liu 0356, Jiankang Deng, Lingxi Xie, Weiming Dong, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.4
2023 HRInversion: High-Resolution GAN Inversion for Cross-Domain Image Synthesis
abstract
We investigate GAN inversion problems of using pre-trained GANs to reconstruct real images. Recent methods for such problems typically employ a VGG perceptual loss to measure the difference between images. While the perceptual loss has achieved remarkable success in various computer vision tasks, it may cause unpleasant artifacts and is sensitive to changes in input scale. This paper delivers an important message that algorithm details are crucial for achieving satisfying performance. In particular, we propose two important but undervalued design principles: (i) not down-sampling the input of the perceptual loss to avoid high-frequency artifacts; and (ii) calculating the perceptual loss using convolutional features which are robust to scale. Integrating these designs derives the proposed framework, HRInversion, that achieves superior performance in reconstructing image details. We validate the effectiveness of HRInversion on a cross-domain image synthesis task and propose a post-processing approach named local style optimization (LSO) to synthesize clean and controllable stylized images. For the evaluation of the cross-domain images, we introduce a metric named ID retrieval which captures the similarity of face identities of stylized images to content images. We also test HRInversion on non-square images. Equipped with implicit neural representation, HRInversion applies to ultra-high resolution images with more than 10 million pixels. Furthermore, we show applications of style transfer and 3D-aware GAN inversion, paving the way for extending the application range of HRInversion.
Peng Zhou 0010, Lingxi Xie, Bingbing Ni, Lin Liu 0016, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Can Semantic Labels Assist Self-Supervised Visual Representation Learning?
abstract
Recently, contrastive learning has largely advanced the progress of unsupervised visual representation learning. Pre-trained on ImageNet, some self-supervised algorithms reported higher transfer learning performance compared to fully-supervised methods, seeming to deliver the message that human labels hardly contribute to learning transferrable visual features. In this paper, we defend the usefulness of semantic labels but point out that fully-supervised and self-supervised methods are pursuing different kinds of features. To alleviate this issue, we present a new algorithm named Supervised Contrastive Adjustment in Neighborhood (SCAN) that maximally prevents the semantic guidance from damaging the appearance feature embedding. In a series of downstream tasks, SCAN achieves superior performance compared to previous fully-supervised and self-supervised methods, and sometimes the gain is significant. More importantly, our study reveals that semantic labels are useful in assisting self-supervised methods, opening a new direction for the community.
Longhui Wei, Lingxi Xie, Xiaopeng Zhang 0008, Qi Tian 0001
AAAI2
2022 DATA: Domain-Aware and Task-Aware Self-supervised Learning
abstract
The paradigm of training models on massive data without label through self-supervised learning (SSL) and fine-tuning on many downstream tasks has become a trend recently. However, due to the high training costs and the un-consciousness of downstream usages, most self-supervised learning methods lack the capability to correspond to the diversities of downstream scenarios, as there are various data domains, different vision tasks and latency constraints on models. Neural architecture search (NAS) is one universally acknowledged fashion to conquer the issues above, but applying NAS on SSL seems impossible as there is no label or metric provided for judging model selection. In this paper, we present DATA, a simple yet effective NAS approach specialized for SSL that provides Domain-Aware and Task-Aware pre-training. Specifically, we (i) train a supernet which could be deemed as a set of millions of networks covering a wide range of model scales without any label, (ii) propose a flexible searching mechanism compatible with SSL that enables finding networks of different computation costs, for various downstream vision tasks and data domains without explicit metric provided. Instantiated With MoCo v2, our method achieves promising results across a wide range of computation costs on down-stream tasks, including image classification, object detection and semantic segmentation. DATA is orthogonal to most existing SSL methods and endows them the ability of customization on downstream needs. Extensive experiments on other SSL methods demonstrate the generalizability of the proposed method. Code is released at https://github.com/GAIA-vision/GAIA-ssl.
Junran Peng, Lingxi Xie, Qi Tian 0001, Zhaoxiang Zhang 0001
CVPR3
2022 MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens
abstract
Transformers have offered a new methodology of designing neural networks for visual recognition. Compared to convolutional networks, Transformers enjoy the ability of referring to global features at each stage, yet the attention module brings higher computational overhead that obstructs the application of Transformers to process highresolution visual data. This paper aims to alleviate the conflict between efficiency and flexibility, for which we propose a specialized token for each region that serves as a messenger (MSG). Hence, by manipulating these MSG tokens, one can flexibly exchange visual information across regions and the computational complexity is reduced. We then integrate the MSG token into a multi-scale architecture named MSG-Transformer. In standard image classification and object detection, MSG-Transformer achieves competitive performance and the inference on both GPU and CPU is accelerated. Code is available at https://github.com/hustvl/MSG-Transformer.
Jiemin Fang, Lingxi Xie, Xinggang Wang, Xiaopeng Zhang 0008, Wenyu Liu 0001, Qi Tian 0001
CVPR2
2022 Domain-Agnostic Prior for Transfer Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) is an important topic in the computer vision community. The key difficulty lies in defining a common property between the source and target domains so that the source-domain features can align with the target-domain semantics. In this paper, we present a simple and effective mechanism that regularizes cross-domain representation learning with a domain-agnostic prior (DAP) that constrains the features extracted from source and target domains to align with a domain-agnostic space. In practice, this is easily implemented as an extra loss term that requires a little extra costs. In the standard evaluation protocol of transferring synthesized data to real data, we validate the effectiveness of different types of DAP, especially that borrowed from a text embedding model that shows favorable performance beyond the state-of-the-art UDA approaches in terms of segmentation accuracy. Our research reveals that UDA benefits much from better proxies, possibly from other data modalities.
Xinyue Huo, Lingxi Xie, Hengtong Hu, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001
CVPR2
2022 One-bit Active Query with Contrastive Pairs
abstract
How to achieve better results with fewer labeling costs remains a challenging task. In this paper, we present a new active learning framework, which for the first time incorporates contrastive learning into recently proposed one-bit supervision. Here one-bit supervision denotes a simple Yes or No query about the correctness of the model's prediction, and is more efficient than previous active learning methods requiring assigning accurate labels to the queried samples. We claim that such one-bit information is intrinsically in accordance with the goal of contrastive loss that pulls positive pairs together and pushes negative samples away. Towards this goal, we design an uncertainty metric to actively select samples for query. These samples are then fed into different branches according to the queried results. The Yes query is treated as positive pairs of the queried category for contrastive pulling, while the No query is treated as hard negative pairs for contrastive repelling. Additionally, we design a negative loss that penalizes the negative samples away from the incorrect predicted class, which can be treated as optimizing hard negatives for the corresponding category. Our method, termed as ObCP, produces a more powerful active learning framework, and experiments on several benchmarks demonstrate its superiority.
Yuhang Zhang 0012, Xiaopeng Zhang 0008, Lingxi Xie, Jie Li 0002, Robert C. Qiu, Hengtong Hu, Qi Tian 0001
CVPR3
2022 Vibration-Based Uncertainty Estimation for Learning from Limited Supervision
Hengtong Hu, Lingxi Xie, Xinyue Huo, Richang Hong, Qi Tian 0001
ECCV (30)2
2022 Skeleton-Parted Graph Scattering Networks for 3D Human Motion Prediction
Maosen Li, Siheng Chen, Lingxi Xie, Qi Tian 0001, Ya Zhang 0002
ECCV (6)4
2022 TAPE: Task-Agnostic Prior Embedding for Image Restoration
Lin Liu 0016, Lingxi Xie, Xiaopeng Zhang 0008, Shanxin Yuan, Xiangyu Chen 0006, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001
ECCV (18)2
2022 Active Pointly-Supervised Instance Segmentation
Chufeng Tang, Lingxi Xie, Xiaopeng Zhang 0008, Qi Tian 0001, Xiaolin Hu 0001
ECCV (28)2
2022 Cornerformer: Purifying Instances for Corner-Based Detectors
Xin Chen 0033, Lingxi Xie, Qi Tian 0001
ECCV (10)3
2022 MVP: Multimodality-Guided Visual Pre-training
Longhui Wei, Lingxi Xie, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001
ECCV (30)2
2022 Bag of Instances Aggregation Boosts Self-supervised Distillation
Haohang Xu, Jiemin Fang, Xiaopeng Zhang 0008, Lingxi Xie, Xinggang Wang, Wenrui Dai, Hongkai Xiong, Qi Tian 0001
ICLR4
2022 Finding the Host from the Lesion by Iteratively Mining the Registration Graph
abstract
Voxel-level annotation has always been a burden of training medical image segmentation models. This paper investigates an interesting problem that finds the host organ of a lesion without actually labeling the organ. To remedy the missing annotation, we construct a graph using an off-the-shelf registration algorithm, on which lesion labels over the training set are accumulated to obtain the pseudo organ for each case. These pseudo labels are used to train a deep network, whose predictions determine the affinity of each lesion on the registration graph. We iteratively update the pseudo labels with the affinity until the training convergence. Our method is evaluated on the MSD Liver and KiTS datasets, without seeing any organ annotation, we achieve the test Dice score of 93% for liver and 92% for kidney, and boosts the accuracy of tumor segmentation to a considerable degree, $3%$, which even surpasses the model trained with ground-truth of both organ and tumor.
Zijie Yang, Lingxi Xie, Xinyue Huo, Sheng Tang, Qi Tian 0001, Yongdong Zhang 0001
ACM Multimedia2
2022 Fine-Grained Semantically Aligned Vision-Language Pre-Training
abstract
Large-scale vision-language pre-training has shown impressive advances in a wide range of downstream tasks. Existing methods mainly model the cross-modal alignment by the similarity of the global representations of images and text, or advanced cross-modal attention upon image and text features. However, they fail to explicitly learn the fine-grained semantic alignment between visual regions and textual phrases, as only global image-text alignment information is available. In this paper, we introduce LOUPE, a fine-grained semantically aLigned visiOn-langUage PrE-training framework, which learns fine-grained semantic alignment from the novel perspective of game-theoretic interactions. To efficiently estimate the game-theoretic interactions, we further propose an uncertainty-aware neural Shapley interaction learning module. Experiments show that LOUPE achieves state-of-the-art performance on a variety of vision-language tasks. Without any object-level human annotations and fine-tuning, LOUPE achieves competitive performance on object detection and visual grounding. More importantly, LOUPE opens a new promising direction of learning fine-grained semantics from large-scale raw image-text pairs.
Juncheng Li 0006, Longhui Wei, Linchao Zhu, Lingxi Xie, Yueting Zhuang, Qi Tian 0001, Siliang Tang
NeurIPS6
2022 Fast Dynamic Radiance Fields with Time-Aware Neural Voxels
abstract
Neural radiance fields (NeRF) have shown great success in modeling 3D scenes and synthesizing novel-view images. However, most previous NeRF methods take much time to optimize one single scene. Explicit data structures, e.g. voxel features, show great potential to accelerate the training process. However, voxel features face two big challenges to be applied to dynamic scenes, i.e. modeling temporal information and capturing different scales of point motions. We propose a radiance field framework by representing scenes with time-aware voxel features, named as TiNeuVox. A tiny coordinate deformation network is introduced to model coarse motion trajectories and temporal information is further enhanced in the radiance network. A multi-distance interpolation method is proposed and applied on voxel features to model both small and large motions. Our framework significantly accelerates the optimization of dynamic radiance fields while maintaining high rendering quality. Empirical evaluation is performed on both synthetic and real scenes. Our TiNeuVox completes training with only 8 minutes and 8-MB storage cost while showing similar or even better rendering performance than previous dynamic NeRF methods. Code is available at https://github.com/hustvl/TiNeuVox.
Jiemin Fang, Taoran Yi, Xinggang Wang, Lingxi Xie, Xiaopeng Zhang 0008, Wenyu Liu 0001, Matthias Nießner, Qi Tian 0001
SIGGRAPH Asia4
2022 Network Adjustment: Channel and Block Search Guided by Resource Utilization Ratio
Zhengsu Chen, Lingxi Xie, Jianwei Niu 0002, Xuefeng Liu 0001, Longhui Wei, Qi Tian 0001
Int. J. Comput. Vis.2
2022 Scalable NAS with factorizable architectural parameters
Lanfei Wang, Lingxi Xie, Kaili Zhao, Jun Guo 0002, Qi Tian 0001
Neurocomputing2
2022 Progressive privileged knowledge distillation for online action detection
Peisen Zhao, Lingxi Xie, Ya Zhang 0002, Qi Tian 0001
Pattern Recognit.2
2022 Actionness-Guided Transformer for Anchor-Free Temporal Action Localization
abstract
Temporal action localization, detecting actions in untrimmed videos, is widely studied by anchor-based approaches that first generate excessive action proposals,i.e., temporal windows, then evaluate and classify these proposals. To reduce the number of action proposals, recent studies use an anchor-free approach that leverages each time point rather than a temporal window to represent an action instance. However, this point representation, usually modeled by temporal convolutions, may have the fixed and limited receptive field to detect an entire action. So we propose an Actionness-guided Transformer (Ag-Trans) model to learn representations for each point proposal. Ag-Trans first predicts the actionness,i.e., time sequences of the action starting, continuing, and ending phases, then the corresponding action phase can be embedded to model the point representation. Experimental results show that the Ag-Trans model outperforms the CNN-based model under the same experiment settings, especially for long-duration actions.
Peisen Zhao, Lingxi Xie, Ya Zhang 0002, Qi Tian 0001
IEEE Signal Process. Lett.2
2022 Searching Towards Class-Aware Generators for Conditional Generative Adversarial Networks
abstract
Conditional generative adversarial networks (cGANs) are designed to generate images based on the provided conditions,e.g., class-level distributions, semantic label maps,etc. Existing methods have used the same generator architecture for all classes. This paper presents an idea that adopts neural architecture search (NAS) to find a class-aware architecture for each class. The search space contains regular and class-modulated convolutions, where the latter is designed to introduce class-specific information while avoiding the reduction of training data for each class generator. The search algorithm follows a weight-sharing pipeline with mixed-architecture optimization so that the search cost does not grow with the number of classes. To learn the sampling policy, a Markov decision process is embedded into the search algorithm, and a moving average is applied for better stability. Class-aware generators show advantages over class-agnostic architectures experimentally. Moreover, we discover two intriguing phenomena that are inspirational to craft cGANs by hand.
Peng Zhou 0010, Lingxi Xie, Bingbing Ni, Qi Tian 0001
IEEE Signal Process. Lett.2
2022 Adaptive Spatial Location With Balanced Loss for Video Captioning
abstract
Many pioneering approaches have verified the effectiveness of utilizing the global temporal and local object information for video understanding tasks and have achieved significant progress. However, existing methods utilize object detectors to extract all objects overall video frames. This may bring performance degradation due to the information redundancy both spatially and temporally. To address this problem, we propose an adaptive spatial location module for the video captioning task which dynamically predicts an important position of each video frame in the procedure of generating the description sentence. The proposed adaptive spatial location method not only makes our model focus on local object information, but also reduces time and memory consumption brought by the temporal redundancy in extensive video frames and improves the accuracy of generated description. Besides, we propose a balanced loss function to address the class imbalance problem existing in training data. The proposed balanced loss assigns different weight to each word of ground-truth sentence in the training process which can generate more diversified description sentences. Extensive experimental results on the MSVD and MSR-VTT dataset show that the proposed method achieves competitive performance compared to state-of-the-art methods.
Linghui Li 0001, Yongdong Zhang 0001, Sheng Tang, Lingxi Xie, Xiaoyong Li 0003, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Camera-Based Batch Normalization: An Effective Distribution Alignment Method for Person Re-Identification
abstract
Person re-identification (ReID) aims at matching identities across disjoint cameras. Its fundamental difficulty lies in associating images across individual cameras, where a key clue,i.e., identity appearance, is prone to the environmental factors of cameras and, consequently, subject to distinct image distributions due to the environmental differences between cameras. To associate images from training cameras, ReID methods strongly demand expensive inter-camera annotations for learning the relations between the distribution of these cameras, yet trained models are still not guaranteed to transfer well to unseen cameras. This problem significantly limits the application of ReID. This paper rethinks the working mechanism of conventional ReID approaches and puts forward a new solution. With an effective operator named Camera-based Batch Normalization (CBN), we guarantee an invariant input distribution independent of all cameras. Thus, the training and testing procedures are always conducted under the same input distribution. This alignment brings three benefits. First, ReID models enjoy better abilities to generalize across testing scenarios with unseen cameras and transfer across multiple training sets. Second, it makes better use of intra-camera annotations, which have been undervalued before due to the lack of cross-camera information. Ideally, the cost of inter-camera annotations can be largely reduced. Third, cross-modality tasks can be better defined through aligning visible/infrared cameras’ distributions. Experiments on a wide range of ReID tasks demonstrate the effectiveness of our approach.
Zijie Zhuang, Longhui Wei, Lingxi Xie, Haizhou Ai, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Heterogeneous Contrastive Learning: Encoding Spatial Information for Compact Visual Representations
abstract
Unsupervised pretraining is of great significance for visual representation. Especially, contrastive learning has achieved great success recently, but existing approaches mostly ignored spatial information which is often crucial for visual representation. Strong semantic embedding has an inherent advantage for classification, but dense prediction tasks require more spatial and low-level representation. This paper presentsheterogeneous contrastive learning(HCL), an effective approach that adds spatial information to the encoding stage to alleviate the learning inconsistency between the contrastive objective and strong data augmentation operations. We demonstrate the effectiveness of HCL by showing that (i) it achieves higher accuracy in instance discrimination, (ii) it surpasses existing pre-training methods in a series of downstream tasks (iii) and it shrinks the pre-training costs by half for almost 800 GPU-hours. More importantly, we show that our approach achieves higher efficiency in visual representations, and thus delivers a key message to inspire the future research of self-supervised visual representation learning.
Xinyue Huo, Lingxi Xie, Longhui Wei, Xiaopeng Zhang 0008, Xin Chen 0033, Hao Li 0090, Zijie Yang, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001
IEEE Trans. Multim.2
2021 Fitting the Search Space of Weight-sharing NAS with Graph Convolutional Networks
abstract
Neural architecture search has attracted wide attentions in both academia and industry. To accelerate it, researchers proposed weight-sharing methods which first train a super-network to reuse computation among different operators, from which exponentially many sub-networks can be sampled and efficiently evaluated. These methods enjoy great advantages in terms of computational costs, but the sampled sub-networks are not guaranteed to be estimated precisely unless an individual training process is taken. This paper owes such inaccuracy to the inevitable mismatch between assembled network layers, so that there is a random error term added to each estimation. We alleviate this issue by training a graph convolutional network to fit the performance of sampled sub-networks so that the impact of random errors becomes minimal. With this strategy, we achieve a higher rank correlation coefficient in the selected set of candidates, which consequently leads to better performance of the final architecture. In addition, our approach also enjoys the flexibility of being used under different hardware constraints, since the graph convolutional network has provided an efficient lookup table of the performance of architectures in the entire search space.
Xin Chen 0033, Lingxi Xie, Jun Wu 0006, Longhui Wei, Yuhui Xu 0002, Qi Tian 0001
AAAI2
2021 MagDR: Mask-Guided Detection and Reconstruction for Defending Deepfakes
abstract
1Deepfakes raised serious concerns on the authenticity of visual contents. Prior works revealed the possibility to disrupt deepfakes by adding adversarial perturbations to the source data, but we argue that the threat has not been eliminated yet. This paper presents MagDR, a mask-guided detection and reconstruction pipeline for defending deepfakes from adversarial attacks. MagDR starts with a detection module that defines a few criteria to judge the abnormality of the output of deepfakes, and then uses it to guide a learnable reconstruction procedure. Adaptive masks are extracted to capture the change in local facial regions. In experiments, MagDR defends three main tasks of deepfakes, and the learned reconstruction pipeline transfers across input data, showing promising performance in defending both black-box and white-box attacks.
Lingxi Xie, Shanmin Pang, Bo Zhang 0010
CVPR2
2021 ATSO: Asynchronous Teacher-Student Optimization for Semi-Supervised Image Segmentation
abstract
Semi-supervised learning is a useful tool for image segmentation, mainly due to its ability in extracting knowledge from unlabeled data to assist learning from labeled data. This paper focuses on a popular pipeline known as self-learning, where we point out a weakness named lazy mimicking that refers to the inertia that a model retains the prediction from itself and thus resists updates. To alleviate this issue, we propose the Asynchronous Teacher-Student Optimization (ATSO) algorithm that (i) breaks up continual learning from teacher to student and (ii) partitions the unlabeled training data into two subsets and alternately uses one subset to fine-tune the model which updates the labels on the other. We show the ability of ATSO on medical and natural image segmentation. In both scenarios, our method reports competitive performance, on par with the state-of-the-arts, in either using partial labeled data in the same dataset or transferring the trained model to an unlabeled dataset.
Xinyue Huo, Lingxi Xie, Zijie Yang, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001
CVPR2
2021 UnrealPerson: An Adaptive Pipeline Towards Costless Person Re-Identification
abstract
The main difficulty of person re-identification (ReID) lies in collecting annotated data and transferring the model across different domains. This paper presents UnrealPerson, a novel pipeline that makes full use of unreal image data to decrease the costs in both the training and deployment stages. Its fundamental part is a system that can generate synthesized images of high-quality and from controllable distributions. Instance-level annotation goes with the synthesized data and is almost free. We point out some details in image synthesis that largely impact the data quality. With 3,000 IDs and 120,000 instances, our method achieves a 38.5% rank-1 accuracy when being directly transferred to MSMT17. It almost doubles the former record using synthesized data and even surpasses previous direct transfer records using real data. This offers a good basis for unsupervised domain adaption, where our pre-trained model is easily plugged into the state-of-the-art algorithms towards higher accuracy. In addition, the data distribution can be flexibly adjusted to fit some corner ReID scenarios, which widens the application of our pipeline. We publish our data synthesis toolkit and synthesized data in https://github.com/FlyHighest/UnrealPerson.
Lingxi Xie, Longhui Wei, Zijie Zhuang, Yongfei Zhang, Bo Li 0006, Qi Tian 0001
CVPR2
2021 Visformer: The Vision-friendly Transformer
abstract
The past year has witnessed the rapid development of applying the Transformer module to vision problems. While some researchers have demonstrated that Transformer-based models enjoy a favorable ability of fitting data, there are still growing number of evidences showing that these models suffer over-fitting especially when the training data is limited. This paper offers an empirical study by performing step-by-step operations to gradually transit a Transformer-based model to a convolution-based model. The results we obtain during the transition process deliver useful messages for improving visual recognition. Based on these observations, we propose a new architecture named Visformer, which is abbreviated from the ‘Vision-friendly Transformer’. With the same computational complexity, Visformer outperforms both the Transformer-based and convolution-based models in terms of ImageNet classification accuracy, and the advantage becomes more significant when the model complexity is lower or the training set is smaller. The code is available at https://github.com/danczs/Visformer.
Zhengsu Chen, Lingxi Xie, Jianwei Niu 0002, Xuefeng Liu 0001, Longhui Wei, Qi Tian 0001
ICCV2
2021 Conformer: Local Features Coupling Global Representations for Visual Recognition
abstract
Within Convolutional Neural Network (CNN), the convolution operations are good at extracting local features but experience difficulty to capture global representations. Within visual transformer, the cascaded self-attention modules can capture long-distance feature dependencies but unfortunately deteriorate local feature details. In this paper, we propose a hybrid network structure, termed Conformer, to take advantage of convolutional operations and self-attention mechanisms for enhanced representation learning. Conformer roots in the Feature Coupling Unit (FCU), which fuses local features and global representations under different resolutions in an interactive fashion. Conformer adopts a concurrent structure so that local features and global representations are retained to the maximum extent. Experiments show that Conformer, under the comparable parameter complexity, outperforms the visual transformer (DeiT-B) by 2.3% on ImageNet. On MSCOCO, it outperforms ResNet-101 by 3.7% and 3.6% mAPs for object detection and instance segmentation, respectively, demonstrating the great potential to be a general backbone network. Code is available at github.com/pengzhiliang/Conformer.
Zhiliang Peng, Shanzhi Gu, Lingxi Xie, Yaowei Wang 0001, Jianbin Jiao, Qixiang Ye
ICCV4
2021 Omni-GAN: On the Secrets of cGANs and Beyond
abstract
The conditional generative adversarial network (cGAN) is a powerful tool of generating high-quality images, but existing approaches mostly suffer unsatisfying performance or the risk of mode collapse. This paper presents Omni-GAN, a variant of cGAN that reveals the devil in designing a proper discriminator for training the model. The key is to ensure that the discriminator receives strong supervision to perceive the concepts and moderate regularization to avoid collapse. Omni-GAN is easily implemented and freely integrated with off-the-shelf encoding methods (e.g., implicit neural representation, INR). Experiments validate the superior performance of Omni-GAN and Omni-INR-GAN in a wide range of image generation and restoration tasks. In particular, Omni-INR-GAN sets new records on the ImageNet dataset with impressive Inception scores of 262.85 and 343.22 for the image sizes of 128 and 256, respectively, surpassing the previous records by 100+ points. Moreover, leveraging the generator prior, Omni-INR-GAN can extrapolate low-resolution images to arbitrary resolution, even up to ×60+ higher resolution. Code is available1.
Peng Zhou 0010, Lingxi Xie, Bingbing Ni, Cong Geng, Qi Tian 0001
ICCV2
2021 Towards Multiple Black-boxes Attack via Adversarial Example Generation Network
abstract
The current research on adversarial attacks aims at a single model while the research on attacking multiple models simultaneously is still challenging. In this paper, we propose a novel black-box attack method, referred to as MBbA, which can attack multiple black-boxes at the same time. By encoding input image and its target category into an associated space, each decoder seeks the appropriate attack areas from the image through the designed loss functions, and then generates effective adversarial examples. This process realizes end-to-end adversarial example generation without involving substitute models for the black-box scenario. On the other hand, adopting the adversarial examples generated by MBbA for adversarial training, the robustness of the attacked models are greatly improved. More importantly, those adversarial examples can achieve satisfactory attack performance, even if these black-box models are trained with the adversarial examples generated by other black-box attack methods, which show good transferability. Finally, extensive experiments show that compared with other state-of-the-art methods: (1) MBbA takes the least time to obtain the most effective attack effects in multi-black-box attack scenario. Furthermore, MBbA achieves the highest attack success rates in a single black-box attack scenario; (2) the adversarial examples generated by MBbA can effectively improve the robustness of the attacked models and exhibit good transferability.
Mingxing Duan, Kenli Li 0001, Lingxi Xie, Qi Tian 0001, Bin Xiao 0001
ACM Multimedia3
2021 Rectifying the Shortcut Learning of Background for Few-Shot Learning
abstract
The category gap between training and evaluation has been characterised as one of the main obstacles to the success of Few-Shot Learning (FSL). In this paper, we for the first time empirically identify image background, common in realistic images, as a shortcut knowledge helpful for in-class classification but ungeneralizable beyond training categories in FSL. A novel framework, COSOC, is designed to tackle this problem by extracting foreground objects in images at both training and evaluation without any extra supervision. Extensive experiments carried on inductive FSL tasks demonstrate the effectiveness of our approaches.
Xu Luo 0003, Longhui Wei, Liangjian Wen, Lingxi Xie, Zenglin Xu, Qi Tian 0001
NeurIPS5
2021 Appending Adversarial Frames for Universal Video Attack
abstract
This paper investigates the problem of generating adversarial examples for video classification. We project all videos onto a semantic space and a perception space, and point out that adversarial attack is to find a counterpart which is close to the target in the perception space but far from the target in the semantic space. Based on this formulation, we notice that conventional attacking methods mostly used Euclidean distance to measure the perception space, but we propose to make full use of the property of videos and assume a modified video with a few consecutive frames replaced by dummy contents (e.g., a black frame with texts of `thank you for watching' on it) to be close to the original video in the perception space though they have a large Euclidean gap. This leads to a new attack approach which only adds perturbations on the newly-added frames. We show its high success rates in attacking six state-of-the-art video classification networks, as well as its universality, i.e., transferring well across videos and models.
Lingxi Xie, Shanmin Pang, Qi Tian 0001
WACV2
2021 Progressive DARTS: Bridging the Optimization Gap for NAS in the Wild
Xin Chen 0033, Lingxi Xie, Jun Wu 0006, Qi Tian 0001
Int. J. Comput. Vis.2
2021 Cyclic CNN: Image Classification With Multiscale and Multilocation Contexts
abstract
Improving the capability of models at limited computational cost is an urgent demand in many vision-based Internet-of-Things applications. Recent progress on deep convolutional neural network (CNN) has largely accelerated the development of image classification. Although the hierarchical structure of CNN naturally helps to extract image features in different scales and locations progressively, conventional convolution can only handle contexts of one scale and on a limited area of a single location in a specific layer, limiting the utilization of multiscale and multilocation information. In this work, we present a cyclic CNN framework, which enables sufficient utilization of multiscale and multilocation contexts in a single layer of convolution. The cyclic CNN is an extremely simple but effective improvement upon conventional convolution, which occupies no additional parameter and negligible computation (even less than 0.1%). Moreover, cyclic CNN can be easily plugged into many existing CNN pipelines, e.g., the ResNet family, obtaining extremely low-cost performance gain upon them. Extensive experiments on both small-scale (CIFAR10 and CIFAR100) and large-scale (ILSVRC2012) image classification benchmarks demonstrate that a consistent performance promotion is obtained with the help of cyclic CNN.
Xin Chen 0033, Lingxi Xie, Jun Wu 0006, Qi Tian 0001
IEEE Internet Things J.2
2021 Partially-Connected Neural Architecture Search for Reduced Computational Redundancy
abstract
Differentiable architecture search (DARTS) enables effective neural architecture search (NAS) using gradient descent, but suffers from high memory and computational costs. In this paper, we propose a novel approach, namely Partially-Connected DARTS (PC-DARTS), to achieve efficient and stable neural architecture search by reducing the channel and spatial redundancies of the super-network. In the channel level, partial channel connection is presented to randomly sample a small subset of channels for operation selection to accelerate the search process and suppress the over-fitting of the super-network. Side operation is introduced for bypassing (non-sampled) channels to guarantee the performance of searched architectures under extremely low sampling rates. In the spatial level, input features are down-sampled to eliminate spatial redundancy and enhance the efficiency of the mixed computation for operation selection. Furthermore, edge normalization is developed to maintain the consistency of edge selection based on channel sampling with the architectural parameters for edges. Theoretical analysis shows that partial channel connection and parameterized side operation are equivalent to regularizing the super-network on the weights and architectural parameters during bilevel optimization. Experimental results demonstrate that the proposed approach achieves higher search speed and training stability than DARTS. PC-DARTS obtains a top-1 error rate of 2.55 percent on CIFAR-10 with 0.07 GPU-days for architecture search, and a state-of-the-art top-1 error rate of 24.1 percent on ImageNet (under the mobile setting) within 2.8 GPU-days.
Yuhui Xu 0002, Lingxi Xie, Wenrui Dai, Xiaopeng Zhang 0008, Xin Chen 0033, Guo-Jun Qi, Hongkai Xiong, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Discretization-aware architecture search
Yunjie Tian, Chang Liu 0042, Lingxi Xie, Jianbin Jiao, Qixiang Ye
Pattern Recognit.3
2021 3D-GAT: 3D-Guided adversarial transform network for person re-identification in unseen domains
Hengheng Zhang, Ying Li 0016, Zijie Zhuang, Lingxi Xie, Qi Tian 0001
Pattern Recognit.4
2021 Interaction-Integrated Network for Natural Language Moment Localization
abstract
Natural language moment localization aims at localizing video clips according to a natural language description. The key to this challenging task lies in modeling the relationship between verbal descriptions and visual contents. Existing approaches often sample a number of clips from the video, and individually determine how each of them is related to the query sentence. However, this strategy can fail dramatically, in particular when the query sentence refers to some visual elements that appear outside of, or even are distant from, the target clip. In this paper, we address this issue by designing an Interaction-Integrated Network (I2N), which contains a few Interaction-Integrated Cells (I2Cs). The idea lies in the observation that the query sentence not only provides a description to the video clip, but also contains semantic cues on the structure of the entire video. Based on this, I2Cs go one step beyond modeling short-term contexts in the time domain by encoding long-term video content into every frame feature. By stacking a few I2Cs, the obtained network, I2N, enjoys an improved ability of inference, brought by both (I) multi-level correspondence between vision and language and (II) more accurate cross-modal alignment. When evaluated on a challenging video moment localization dataset named DiDeMo, I2N outperforms the state-of-the-art approach by a clear margin of 1.98%. On other two challenging datasets, Charades-STA and TACoS, I2N also reports competitive performance.
Lingxi Xie, Jianzhuang Liu, Fei Wu 0001, Qi Tian 0001
IEEE Trans. Image Process.2
2021 Universal-to-Specific Framework for Complex Action Recognition
abstract
Video-based action recognition has recently attracted much attention in the field of computer vision. To solve more complex recognition tasks, it has become necessary to distinguish different levels of interclass variations. Inspired by a common flowchart based on the human decision-making process that first narrows down the probable classes and then applies a "rethinking" process for finer-level recognition, we propose an effective universal-to-specific (U2S) framework for complex action recognition. The U2S framework is composed of three subnetworks: a universal network, a category-specific network, and a mask network. The universal network first learns universal feature representations. The mask network then generates attention masks for confusing classes through category regularization based on the output of the universal network. The mask is further used to guide the category-specific network for class-specific feature representations. The entire framework is optimized in an end-to-end manner. Experiments on a variety of benchmark datasets, e.g., the Something-Something, UCF101, and HMDB51 datasets, demonstrate the effectiveness of the U2S framework; i.e., U2S can focus on discriminative spatiotemporal regions for confusing categories. We further visualize the relationship between different classes, showing that U2S indeed improves the discriminability of learned features. Moreover, the proposed U2S model is a general framework and may adopt any base recognition network.
Peisen Zhao, Lingxi Xie, Ya Zhang 0002, Qi Tian 0001
IEEE Trans. Multim.2
2020 Pruning from Scratch
abstract
Network pruning is an important research field aiming at reducing computational costs of neural networks. Conventional approaches follow a fixed paradigm which first trains a large and redundant network, and then determines which units (e.g., channels) are less important and thus can be removed. In this work, we find that pre-training an over-parameterized model is not necessary for obtaining the target pruned structure. In fact, a fully-trained over-parameterized model will reduce the search space for the pruned structure. We empirically show that more diverse pruned structures can be directly pruned from randomly initialized weights, including potential models with better performance. Therefore, we propose a novel network pruning pipeline which allows pruning from scratch with little training overhead. In the experiments for compressing classification models on CIFAR10 and ImageNet datasets, our approach not only greatly reduces the pre-training burden of traditional pruning methods, but also achieves similar or even higher accuracy under the same computation budgets. Our results facilitate the community to rethink the effectiveness of existing techniques used for network pruning.
Lingxi Xie, Jun Zhou 0011, Hang Su 0006, Bo Zhang 0010, Xiaolin Hu 0001
AAAI3
2020 Single Camera Training for Person Re-Identification
abstract
Person re-identification (ReID) aims at finding the same person in different cameras. Training such systems usually requires a large amount of cross-camera pedestrians to be annotated from surveillance videos, which is labor-consuming especially when the number of cameras is large. Differently, this paper investigates ReID in an unexplored single-camera-training (SCT) setting, where each person in the training set appears in only one camera. To the best of our knowledge, this setting was never studied before. SCT enjoys the advantage of low-cost data collection and annotation, and thus eases ReID systems to be trained in a brand new environment. However, it raises major challenges due to the lack of cross-camera person occurrences, which conventional approaches heavily rely on to extract discriminative features. The key to dealing with the challenges in the SCT setting lies in designing an effective mechanism to complement cross-camera annotation. We start with a regular deep network for feature extraction, upon which we propose a novel loss function named multi-camera negative loss (MCNL). This is a metric learning loss motivated by probability, suggesting that in a multi-camera system, one image is more likely to be closer to the most similar negative sample in other cameras than to the most similar negative sample in the same camera. In experiments, MCNL significantly boosts ReID accuracy in the SCT setting, which paves the way of fast deployment of ReID systems with good performance on new target scenes.
Lingxi Xie, Longhui Wei, Yongfei Zhang, Bo Li 0006, Qi Tian 0001
AAAI2
2020 Network Adjustment: Channel Search Guided by FLOPs Utilization Ratio
abstract
Automatic designing computationally efficient neural networks has received much attention in recent years. Existing approaches either utilize network pruning or leverage the network architecture search methods. This paper presents a new framework named network adjustment, which considers network accuracy as a function of FLOPs, so that under each network configuration, one can estimate the FLOPs utilization ratio (FUR) for each layer and use it to determine whether to increase or decrease the number of channels on the layer. Note that FUR, like the gradient of a non-linear function, is accurate only in a small neighborhood of the current network. Hence, we design an iterative mechanism so that the initial network undergoes a number of steps, each of which has a small 'adjusting rate' to control the changes to the network. The computational overhead of the entire search process is reasonable, i.e., comparable to that of re-training the final model from scratch. Experiments on standard image classification datasets and a wide range of base networks demonstrate the effectiveness of our approach, which consistently outperforms the pruning counterpart. The code is available at https://github.com/danczs/NetworkAdjustment.
Zhengsu Chen, Jianwei Niu 0002, Lingxi Xie, Xuefeng Liu 0001, Longhui Wei, Qi Tian 0001
CVPR3
2020 Creating Something From Nothing: Unsupervised Knowledge Distillation for Cross-Modal Hashing
abstract
In recent years, cross-modal hashing (CMH) has attracted increasing attentions, mainly because its potential ability of mapping contents from different modalities, especially in vision and language, into the same space, so that it becomes efficient in cross-modal data retrieval. There are two main frameworks for CMH, differing from each other in whether semantic supervision is required. Compared to the unsupervised methods, the supervised methods often enjoy more accurate results, but require much heavier labors in data annotation. In this paper, we propose a novel approach that enables guiding a supervised method using outputs produced by an unsupervised method. Specifically, we make use of teacher-student optimization for propagating knowledge. Experiments are performed on two popular CMH benchmarks, i.e., the MIRFlickr and NUS-WIDE datasets. Our approach outperforms all existing unsupervised methods by a large margin.
Hengtong Hu, Lingxi Xie, Richang Hong, Qi Tian 0001
CVPR2
2020 Unsupervised Person Re-Identification via Softened Similarity Learning
abstract
Person re-identification (re-ID) is an important topic in computer vision. This paper studies the unsupervised setting of re-ID, which does not require any labeled information and thus is freely deployed to new scenarios. There are very few studies under this setting, and one of the best approach till now used iterative clustering and classification, so that unlabeled images are clustered into pseudo classes for a classifier to get trained, and the updated features are used for clustering and so on. This approach suffers two problems, namely, the difficulty of determining the number of clusters, and the hard quantization loss in clustering. In this paper, we follow the iterative training mechanism but discard clustering, since it incurs loss from hard quantization, yet its only product, image-level similarity, can be easily replaced by pairwise computation and a softened classification task. With these improvements, our approach becomes more elegant and is more robust to hyper-parameter changes. Experiments on two image-based and video-based datasets demonstrate state-of-the-art performance under the unsupervised re-ID setting.
Yutian Lin, Lingxi Xie, Yu Wu 0011, Chenggang Yan 0001, Qi Tian 0001
CVPR2
2020 Corner Proposal Network for Anchor-Free, Two-Stage Object Detection
Kaiwen Duan, Lingxi Xie, Honggang Qi, Song Bai 0001, Qingming Huang, Qi Tian 0001
ECCV (3)2
2020 Reinforced Axial Refinement Network for Monocular 3D Object Detection
Lijie Liu, Chufan Wu, Jiwen Lu, Lingxi Xie, Jie Zhou 0001, Qi Tian 0001
ECCV (17)4
2020 Circumventing Outliers of AutoAugment with Knowledge Distillation
Longhui Wei, An Xiao, Lingxi Xie, Xiaopeng Zhang 0008, Xin Chen 0033, Qi Tian 0001
ECCV (3)3
2020 Social Adaptive Module for Weakly-Supervised Group Activity Recognition
Rui Yan 0010, Lingxi Xie, Jinhui Tang 0001, Xiangbo Shu, Qi Tian 0001
ECCV (8)2
2020 Bottom-Up Temporal Action Localization with Mutual Regularization
Peisen Zhao, Lingxi Xie, Chen Ju, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
ECCV (8)2
2020 Rethinking the Distribution Gap of Person Re-identification with Camera-Based Batch Normalization
Zijie Zhuang, Longhui Wei, Lingxi Xie, Hengheng Zhang, Haozhe Wu, Haizhou Ai, Qi Tian 0001
ECCV (12)3
2020 Cross-VAE: Towards Disentangling Expression from Identity For Human Faces
abstract
Facial expression and identity are two independent yet intertwined components for representing a face. For facial expression recognition, identity can contaminate the training procedure by providing tangled but irrelevant information. In this paper, we propose to learn clearly disentangled and discriminative features that are invariant of identities for expression recognition. However, such disentanglement normally requires annotations of both expression and identity on one large dataset, which is often unavailable. Our solution is to extend conditional VAE to a crossed version named Cross-VAE, which is able to use partially labeled data to disentangle expression from identity. We emphasis the following novel characteristics of our Cross-VAE: (1) It is based on an independent assumption that the two latent representations' distributions are orthogonal. This ensures both encoded representations to be disentangled and expressive. (2) It utilizes a symmetric training procedure where the output of each encoder is fed as the condition of the other. Thus two partially labeled sets can be jointly used. Extensive experiments show that our proposed method is capable of encoding expressive and disentangled features for facial expression. Compared with the baseline methods, our model shows an improvement of 3.56% on average in terms of accuracy.
Haozhe Wu, Jia Jia 0001, Lingxi Xie, Guo-Jun Qi, Yuanchun Shi, Qi Tian 0001
ICASSP3
2020 PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search
Yuhui Xu 0002, Lingxi Xie, Xiaopeng Zhang 0008, Xin Chen 0033, Guo-Jun Qi, Qi Tian 0001, Hongkai Xiong
ICLR2
2020 Polar Relative Positional Encoding for Video-Language Segmentation
abstract
In this paper, we tackle a challenging task named video-language segmentation. Given a video and a sentence in natural language, the goal is to segment the object or actor described by the sentence in video frames. To accurately denote a target object, the given sentence usually refers to multiple attributes, such as nearby objects with spatial relations, etc. In this paper, we propose a novel Polar Relative Positional Encoding (PRPE) mechanism that represents spatial relations in a ``linguistic'' way, i.e., in terms of direction and range. Sentence feature can interact with positional embeddings in a more direct way to extract the implied relative positional relations. We also propose parameterized functions for these positional embeddings to adapt real-value directions and ranges. With PRPE, we design a Polar Attention Module (PAM) as the basic module for vision-language fusion. Our method outperforms previous best method by a large margin of 11.4% absolute improvement in terms of mAP on the challenging A2D Sentences dataset. Our method also achieves competitive performances on the J-HMDB Sentences dataset.
Lingxi Xie, Fei Wu 0001, Qi Tian 0001
IJCAI2
2020 One-bit Supervision for Image Classification
abstract
This paper presents one-bit supervision, a novel setting of learning from incomplete annotations, in the scenario of image classification. Instead of training a model upon the accurate label of each sample, our setting requires the model to query with a predicted label of each sample and learn from the answer whether the guess is correct. This provides one bit (yes or no) of information, and more importantly, annotating each sample becomes much easier than finding the accurate label from many candidate classes. There are two keys to training a model upon one-bit supervision: improving the guess accuracy and making use of incorrect guesses. For these purposes, we propose a multi-stage training paradigm which incorporates negative label suppression into an off-the-shelf semi-supervised learning algorithm. In three popular image classification benchmarks, our approach claims higher efficiency in utilizing the limited amount of annotations.
Hengtong Hu, Lingxi Xie, Zewei Du, Richang Hong, Qi Tian 0001
NeurIPS2
2020 Recurrent Saliency Transformation Network for Tiny Target Segmentation in Abdominal CT Scans
abstract
We aim at segmenting a wide variety of organs, including tiny targets (e.g., adrenal gland), and neoplasms (e.g., pancreatic cyst), from abdominal CT scans. This is a challenging task in two aspects. First, some organs (e.g., the pancreas), are highly variable in both anatomy and geometry, and thus very difficult to depict. Second, the neoplasms often vary a lot in its size, shape, as well as its location within the organ. Third, the targets (organs and neoplasms) can be considerably small compared to the human body, and so standard deep networks for segmentation are often less sensitive to these targets and thus predict less accurately especially around their boundaries. In this paper, we present an end-to-end framework named recurrent saliency transformation network (RSTN) for segmenting tiny and/or variable targets. The RSTN is a coarse-to-fine approach that uses prediction from the first (coarse) stage to shrink the input region for the second (fine) stage. A saliency transformation module is inserted between these two stages so that 1) the coarse-scaled segmentation mask can be transferred as spatial weights and applied to the fine stage and 2) the gradients can be back-propagated from the loss layer to the entire network so that the two stages are optimized in a joint manner. In the testing stage, we perform segmentation iteratively to improve accuracy. In this extended journal paper, we allow a gradual optimization to improve the stability of the RSTN, and introduce a hierarchical version named H-RSTN to segment tiny and variable neoplasms such as pancreatic cysts. Experiments are performed on several CT datasets including a public pancreas segmentation dataset, our own multi-organ dataset, and a cystic pancreas dataset. In all these cases, the RSTN outperforms the baseline (a stage-wise coarse-to-fine approach) significantly. Confirmed by the radiologists in our team, these promising segmentation results can help early diagnosis of pancreatic cancer. The code and pre-trained models of our project were made available at https://github.com/198808xc/OrganSegRSTN.
Lingxi Xie, Qihang Yu, Yuyin Zhou, Yan Wang 0033, Elliot K. Fishman, Alan L. Yuille
IEEE Trans. Medical Imaging1
2019 G2C: A Generator-to-Classifier Framework Integrating Multi-Stained Visual Cues for Pathological Glomerulus Classification
abstract
Pathological glomerulus classification plays a key role in the diagnosis of nephropathy. As the difference between different subcategories is subtle, doctors often refer to slides from different staining methods to make decisions. However, creating correspondence across various stains is labor-intensive, bringing major difficulties in collecting data and training a vision-based algorithm to assist nephropathy diagnosis.This paper provides an alternative solution for integrating multi-stained visual cues for glomerulus classification. Our approach, named generator-to-classifier (G2C), is a twostage framework. Given an input image from a specified stain, several generators are first applied to estimate its appearances in other staining methods, and a classifier follows to combine visual cues from different stains for prediction (whether it is pathological, or which type of pathology it has). We optimize these two stages in a joint manner. To provide a reasonable initialization, we pre-train the generators in an unlabeled reference set under an unpaired image-to-image translation task, and then fine-tune them together with the classifier.We conduct experiments on a glomerulus type classification dataset collected by ourselves (there are no publicly available datasets for this purpose). Although joint optimization slightly harms the authenticity of the generated patches, it boosts classification performance, suggesting more effective visual cues are extracted in an automatic way. We also transfer our model to a public dataset for breast cancer classification, and outperform the state-of-the-arts significantly.
Bingzhe Wu, Shiwan Zhao, Lingxi Xie, Caihong Zeng, Guangyu Sun 0003
AAAI4
2019 Training Deep Neural Networks in Generations: A More Tolerant Teacher Educates Better Students
abstract
We focus on the problem of training a deep neural network in generations. The flowchart is that, in order to optimize the target network (student), another network (teacher) with the same architecture is first trained, and used to provide part of supervision signals in the next stage. While this strategy leads to a higher accuracy, many aspects (e.g., why teacher-student optimization helps) still need further explorations.This paper studies this problem from a perspective of controlling the strictness in training the teacher network. Existing approaches mostly used a hard distribution (e.g., one-hot vectors) in training, leading to a strict teacher which itself has a high accuracy, but we argue that the teacher needs to be more tolerant, although this often implies a lower accuracy. The implementation is very easy, with merely an extra loss term added to the teacher network, facilitating a few secondary classes to emerge and complement to the primary class. Consequently, the teacher provides a milder supervision signal (a less peaked distribution), and makes it possible for the student to learn from inter-class similarity and potentially lower the risk of over-fitting. Experiments are performed on standard image classification tasks (CIFAR100 and ILSVRC2012). Although the teacher network behaves less powerful, the students show a persistent ability growth and eventually achieve higher classification accuracies than other competitors. Model ensemble and transfer feature extraction also verify the effectiveness of our approach.
Lingxi Xie, Siyuan Qiao, Alan L. Yuille
AAAI2
2019 Attention-Guided Unified Network for Panoptic Segmentation
abstract
This paper studies panoptic segmentation, a recently proposed task which segments foreground (FG) objects at the instance level as well as background (BG) contents at the semantic level. Existing methods mostly dealt with these two problems separately, but in this paper, we reveal the underlying relationship between them, in particular, FG objects provide complementary cues to assist BG understanding. Our approach, named the Attention-guided Unified Network (AUNet), is a unified framework with two branches for FG and BG segmentation simultaneously. Two sources of attentions are added to the BG branch, namely, RPN and FG segmentation mask to provide object-level and pixel-level attentions, respectively. Our approach is generalized to different backbones with consistent accuracy gain in both FG and BG segmentation, and also sets new state-of-the-arts both in the MS-COCO (46.5% PQ) and Cityscapes (59.0% PQ) benchmarks.
Xinze Chen, Lingxi Xie, Guan Huang 0003, Dalong Du, Xingang Wang 0003
CVPR4
2019 SIXray: A Large-Scale Security Inspection X-Ray Benchmark for Prohibited Item Discovery in Overlapping Images
abstract
In this paper, we present a large-scale dataset and establish a baseline for prohibited item discovery in Security Inspection X-ray images. Our dataset, named SIXray, consists of 1,059,231 X-ray images, in which 6 classes of 8,929 prohibited items are manually annotated. It raises a brand new challenge of overlapping image data, meanwhile shares the same properties with existing datasets, including complex yet meaningless contexts and class imbalance. We propose an approach named class-balanced hierarchical refinement (CHR) to deal with these difficulties. CHR assumes that each input image is sampled from a mixture distribution, and that deep networks require an iterative process to infer image contents accurately. To accelerate, we insert reversed connections to different network backbones, delivering high-level visual cues to assist mid-level features. In addition, a class-balanced loss function is designed to maximally alleviate the noise introduced by easy negative samples. We evaluate CHR on SIXray with different ratios of positive/negative samples. Compared to the baselines, CHR enjoys a better ability of discriminating objects especially using mid-level features, which offers the possibility of using a weakly-supervised approach towards accurate object localization. In particular, the advantage of CHR is more significant in the scenarios with fewer positive training samples, which demonstrates its potential application in real-world security inspection.
Caijing Miao, Lingxi Xie, Fang Wan 0001, Chi Su, Hongye Liu, Jianbin Jiao, Qixiang Ye
CVPR2
2019 Elastic Boundary Projection for 3D Medical Image Segmentation
abstract
We focus on an important yet challenging problem: using a 2D deep network to deal with 3D segmentation for medical image analysis. Existing approaches either applied multi-view planar (2D) networks or directly used volumetric (3D) networks for this purpose, but both of them are not ideal: 2D networks cannot capture 3D contexts effectively, and 3D networks are both memory-consuming and less stable arguably due to the lack of pre-trained models. In this paper, we bridge the gap between 2D and 3D using a novel approach named Elastic Boundary Projection (EBP). The key observation is that, although the object is a 3D volume, what we really need in segmentation is to find its boundary which is a 2D surface. Therefore, we place a number of pivot points in the 3D space, and for each pivot, we determine its distance to the object boundary along a dense set of directions. This creates an elastic shell around each pivot which is initialized as a perfect sphere. We train a 2D deep network to determine whether each ending point falls within the object, and gradually adjust the shell so that it gradually converges to the actual shape of the boundary and thus achieves the goal of segmentation. EBP allows boundary-based segmentation without cutting a 3D volume into slices or patches, which stands out from conventional 2D and 3D approaches. EBP achieves promising accuracy in abdominal organ segmentation. Our code will be released on https://github.com/twni2016/Elastic-Boundary-Projection .
Tianwei Ni, Lingxi Xie, Huangjie Zheng, Elliot K. Fishman, Alan L. Yuille
CVPR2
2019 Iterative Reorganization With Weak Spatial Constraints: Solving Arbitrary Jigsaw Puzzles for Unsupervised Representation Learning
abstract
Learning visual features from unlabeled image data is an important yet challenging task, which is often achieved by training a model on some annotation-free information. We consider spatial contexts, for which we solve so-called jigsaw puzzles, i.e., each image is cut into grids and then disordered, and the goal is to recover the correct configuration. Existing approaches formulated it as a classification task by defining a fixed mapping from a small subset of configurations to a class set, but these approaches ignore the underlying relationship between different configurations and also limit their applications to more complex scenarios. This paper presents a novel approach which applies to jigsaw puzzles with an arbitrary grid size and dimensionality. We provide a fundamental and generalized principle, that weaker cues are easier to be learned in an unsupervised manner and also transfer better. In the context of puzzle recognition, we use an iterative manner which, instead of solving the puzzle all at once, adjusts the order of the patches in each step until convergence. In each step, we combine both unary and binary features of each patch into a cost function judging the correctness of the current configuration. Our approach, by taking similarity between puzzles into consideration, enjoys a more efficient way of learning visual knowledge. We verify the effectiveness of our approach from two aspects. First, it solves arbitrarily complex puzzles, including high-dimensional puzzles, that prior methods are difficult to handle. Second, it serves as a reliable way of network initialization, which leads to better transfer performance in visual recognition tasks including classification, detection and segmentation.
Chen Wei 0005, Lingxi Xie, Xutong Ren, Yingda Xia, Chi Su, Jiaying Liu 0001, Qi Tian 0001, Alan L. Yuille
CVPR2
2019 Snapshot Distillation: Teacher-Student Optimization in One Generation
abstract
Optimizing a deep neural network is a fundamental task in computer vision, yet direct training methods often suffer from over-fitting. Teacher-student optimization aims at providing complementary cues from a model trained previously, but these approaches are often considerably slow due to the pipeline of training a few generations in sequence, i.e., time complexity is increased by several times. This paper presents snapshot distillation (SD), the first framework which enables teacher-student optimization in one generation. The idea of SD is very simple: instead of borrowing supervision signals from previous generations, we extract such information from earlier epochs in the same generation, meanwhile make sure that the difference between teacher and student is sufficiently large so as to prevent under-fitting. To achieve this goal, we implement SD in a cyclic learning rate policy, in which the last snapshot of each cycle is used as the teacher for all iterations in the next cycle, and the teacher signal is smoothed to provide richer information. In standard image classification benchmarks such as CIFAR100 and ILSVRC2012, SD achieves consistent accuracy gain without heavy computational overheads. We also verify that models pre-trained with SD transfers well to object detection and semantic segmentation in the PascalVOC dataset.
Lingxi Xie, Chi Su, Alan L. Yuille
CVPR2
2019 Adversarial Attacks Beyond the Image Space
abstract
Generating adversarial examples is an intriguing problem and an important way of understanding the working mechanism of deep neural networks. Most existing approaches generated perturbations in the image space, i.e., each pixel can be modified independently. However, in this paper we pay special attention to the subset of adversarial examples that correspond to meaningful changes in 3D physical properties (like rotation and translation, illumination condition, etc.). These adversaries arguably pose a more serious concern, as they demonstrate the possibility of causing neural network failure by easy perturbations of real-world 3D objects and scenes. In the contexts of object classification and visual question answering, we augment state-of-the-art deep neural networks that receive 2D input images with a rendering module (either differentiable or not) in front, so that a 3D scene (in the physical space) is rendered into a 2D image (in the image space), and then mapped to a prediction (in the output space). The adversarial perturbations can now go beyond the image space, and have clear meanings in the 3D physical world. Though image-space adversaries can be interpreted as per-pixel albedo change, we verify that they cannot be well explained along these physically meaningful dimensions, which often have a non-local effect. But it is still possible to successfully attack beyond the image space on the physical space, though this is more difficult than image-space attacks, reflected in lower success rates and heavier perturbations required.
Xiaohui Zeng, Chenxi Liu 0001, Yu-Siang Wang, Weichao Qiu, Lingxi Xie, Yu-Wing Tai, Chi-Keung Tang, Alan L. Yuille
CVPR5
2019 CRAVES: Controlling Robotic Arm With a Vision-Based Economic System
abstract
Training a robotic arm to accomplish real-world tasks has been attracting increasing attention in both academia and industry. This work discusses the role of computer vision algorithms in this field. We focus on low-cost arms on which no sensors are equipped and thus all decisions are made upon visual recognition, e.g., real-time 3D pose estimation. This requires annotating a lot of training data, which is not only time-consuming but also laborious. In this paper, we present an alternative solution, which uses a 3D model to create a large number of synthetic data, trains a vision model in this virtual domain, and applies it to real-world images after domain adaptation. To this end, we design a semi-supervised approach, which fully leverages the geometric constraints among keypoints. We apply an iterative algorithm for optimization. Without any annotations on real images, our algorithm generalizes well and produces satisfying results on 3D pose estimation, which is evaluated on two real-world datasets. We also construct a vision-based control system for task accomplishment, for which we train a reinforcement learning agent in a virtual environment and apply it to the real-world. Moreover, our approach, with merely a 3D model being required, has the potential to generalize to other types of multi-rigid-body dynamic systems.
Yiming Zuo 0001, Weichao Qiu, Lingxi Xie, Fangwei Zhong, Yizhou Wang 0001, Alan L. Yuille
CVPR3
2019 Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image Retrieval
abstract
Sketch-based image retrieval (SBIR) is widely recognized as an important vision problem which implies a wide range of real-world applications. Recently, research interests arise in solving this problem under the more realistic and challenging setting of zero-shot learning. In this paper, we investigate this problem from the viewpoint of domain adaptation which we show is critical in improving feature embedding in the zero-shot scenario. Based on a framework which starts with a pre-trained model on ImageNet and fine-tunes it on the training set of SBIR benchmark, we advocate the importance of preserving previously acquired knowledge, e.g., the rich discriminative features learned from ImageNet, to improve the model's transfer ability. For this purpose, we design an approach named Semantic-Aware Knowledge prEservation (SAKE), which fine-tunes the pre-trained model in an economical way and leverages semantic information, e.g., inter-class relationship, to achieve the goal of knowledge preservation. Zero-shot experiments on two extended SBIR datasets, TU-Berlin and Sketchy, verify the superior performance of our approach. Extensive diagnostic experiments validate that knowledge preserved benefits SBIR in zero-shot settings, as a large fraction of the performance gain is from the more properly structured feature embedding for photo images.
Qing Liu 0017, Lingxi Xie, Alan L. Yuille
ICCV2
2019 Semantic Part Detection via Matching: Learning to Generalize to Novel Viewpoints From Limited Training Data
abstract
Detecting semantic parts of an object is a challenging task, particularly because it is hard to annotate semantic parts and construct large datasets. In this paper, we present an approach which can learn from a small annotated dataset containing a limited range of viewpoints and generalize to detect semantic parts for a much larger range of viewpoints. The approach is based on our matching algorithm, which is used for finding accurate spatial correspondence between two images and transplanting semantic parts annotated on one image to the other. Images in the training set are matched to synthetic images rendered from a 3D CAD model, following which a clustering algorithm is used to automatically annotate semantic parts of the CAD model. During the testing period, this CAD model can synthesize annotated images under every viewpoint. These synthesized images are matched to images in the testing set to detect semantic parts in novel viewpoints. Our algorithm is simple, intuitive, and contains very few parameters. Experiments show our method outperforms standard deep learning approaches and, in particular, performs much better on novel viewpoints. For facilitating the future research, code is available: https://github.com/ytongbai/SemanticPartDetection.
Yutong Bai, Qing Liu 0017, Lingxi Xie, Weichao Qiu, Alan L. Yuille
ICCV3
2019 Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and Evaluation
abstract
Recently, differentiable search methods have made major progress in reducing the computational costs of neural architecture search. However, these approaches often report lower accuracy in evaluating the searched architecture or transferring it to another dataset. This is arguably due to the large gap between the architecture depths in search and evaluation scenarios. In this paper, we present an efficient algorithm which allows the depth of searched architectures to grow gradually during the training procedure. This brings two issues, namely, heavier computational overheads and weaker search stability, which we solve using search space approximation and regularization, respectively. With a significantly reduced search time (~7 hours on a single GPU), our approach achieves state-of-the-art performance on both the proxy dataset (CIFAR10 or CIFAR100) and the target dataset (ImageNet). Code is available at https://github.com/chenxin061/pdarts.
Xin Chen 0033, Lingxi Xie, Jun Wu 0006, Qi Tian 0001
ICCV2
2019 CenterNet: Keypoint Triplets for Object Detection
abstract
In object detection, keypoint-based approaches often experience the drawback of a large number of incorrect object bounding boxes, arguably due to the lack of an additional assessment inside cropped regions. This paper presents an efficient solution that explores the visual patterns within individual cropped regions with minimal costs. We build our framework upon a representative one-stage keypoint-based detector named CornerNet. Our approach, named CenterNet, detects each object as a triplet, rather than a pair, of keypoints, which improves both precision and recall. Accordingly, we design two customized modules, cascade corner pooling, and center pooling, that enrich information collected by both the top-left and bottom-right corners and provide more recognizable information from the central regions. On the MS-COCO dataset, CenterNet achieves an AP of 47.0 %, outperforming all existing one-stage detectors by at least 4.9%. Furthermore, with a faster inference speed than the top-ranked two-stage detectors, CenterNet demonstrates a comparable performance to these detectors. Code is available at https://github.com/Duankaiwen/CenterNet.
Kaiwen Duan, Song Bai 0001, Lingxi Xie, Honggang Qi, Qingming Huang, Qi Tian 0001
ICCV3
2019 Multi-scale Coarse-to-Fine Segmentation for Screening Pancreatic Ductal Adenocarcinoma
Zhuotun Zhu, Yingda Xia, Lingxi Xie, Elliot K. Fishman, Alan L. Yuille
MICCAI (6)3
2019 Fast Non-Local Neural Networks with Spectral Residual Learning
abstract
Effectively modeling long-range spatial correlation is crucial in context-sensitive visual computing tasks, such as human pose estimation and video classification. Enlarging receptive field is popularly adopted in building such non-local deep networks. However, current solutions, including dilation convolution or self-attention based operators, mostly suffer from either low computational efficacy or insufficient receptive field. This paper proposes spectral residual learning (SRL), a novel network architectural design for achieving fully global receptive field. A neural block that implements SRL has three key components: a local-to-global transform that projects some ordinary local features into a spectral domain, compiled operations in the spectral domain, and a global-to-local transform that converts all data back to the original local format. We show its equivalence to conducting residual learning in some spectral domain and carefully re-formulate a variety of neural layers into their spectral forms, such as ReLU or convolutions. The benefits of SRL is three-fold: first, all operations have global receptive field, namely any update affects all image positions. This can extract richer context information in various vision tasks; Secondly, the local-to-global / global-to-local transforms in SRL are defined by bi-linear unitary matrices, which is both computation and parameter economic; Lastly, SRL is a generic formulation, here instantiated by Fourier transform and real orthogonal matrix. We conduct comprehensive evaluations on two challenging tasks, including human pose estimation from images and video classification. All experiments clearly show performance improvement by large margins in comparison with conventional non-local network designs.
Lu Chi, Guiyu Tian, Yadong Mu, Lingxi Xie, Qi Tian 0001
ACM Multimedia4
2019 Det2Seg: A Two-Stage Approach for Road Object Segmentation from 3D Point Clouds
abstract
Object segmentation from 3D point clouds is an important topic in real-world applications. However, due to the sparsity and irregularity of point clouds, instance segmentation often suffers unsatisfying performance in various scenes such as autonomous driving. There are two main difficulties. First, in a wild scene, background noise often occupies a majority of the entire point set. Second, small-scale objects are often difficult to be recognized due to larger uncertainty. In this paper, we propose Det2Seg, a two-stage approach which alleviates the above issues towards more accurate instance segmentation of road-objects. In the first stage, we adopt Pointpillars [1] to detect the regions-of- interest that can localize and classify objects in a coarse level; in the second stage, we extract points from the detected regions into pillars, encode them into a new data format and feed it into a 2D convolutional neural network to perform fine-grained, domain-specific instance segmentation. We evaluate our approach on raw LiDAR (Light Detection And Ranging) data from the KITTI dataset [2]. The experimental results show that our approach largely outperforms the prior researches. In particular, our approach stands out for its significant ability on recognizing and segmenting small-scale objects, i.e., an improvement of over 20% in terms of Intersection over Union (IoU), beyond state-of- the-arts, is obtained for the cyclist class.
Wei Zhang 0196, Lingxi Xie, Qi Tian 0001, Hongkai Xiong
VCIP3
2018 Identity-Enhanced Network for Facial Expression Recognition
Xingang Wang 0003, Shilei Zhang, Lingxi Xie, Hongyuan Yu
ACCV (4)4
2018 SampleAhead: Online Classifier-Sampler Communication for Learning from Synthesized Data
Qi Chen 0014, Weichao Qiu, Yi Zhang 0099, Lingxi Xie, Alan L. Yuille
BMVC4
2018 Recurrent Saliency Transformation Network: Incorporating Multi-Stage Visual Cues for Small Organ Segmentation
abstract
We aim at segmenting small organs (e.g., the pancreas) from abdominal CT scans. As the target often occupies a relatively small region in the input image, deep neural networks can be easily confused by the complex and variable background. To alleviate this, researchers proposed a coarse-to-fine approach [46], which used prediction from the first (coarse) stage to indicate a smaller input region for the second (fine) stage. Despite its effectiveness, this algorithm dealt with two stages individually, which lacked optimizing a global energy function, and limited its ability to incorporate multi-stage visual cues. Missing contextual information led to unsatisfying convergence in iterations, and that the fine stage sometimes produced even lower segmentation accuracy than the coarse stage. This paper presents a Recurrent Saliency Transformation Network. The key innovation is a saliency transformation module, which repeatedly converts the segmentation probability map from the previous iteration as spatial weights and applies these weights to the current iteration. This brings us two-fold benefits. In training, it allows joint optimization over the deep networks dealing with different input scales. In testing, it propagates multi-stage visual information throughout iterations to improve segmentation accuracy. Experiments in the NIH pancreas segmentation dataset demonstrate the state-of-the-art accuracy, which outperforms the previous best by an average of over 2%. Much higher accuracies are also reported on several small organs in a larger dataset collected by ourselves. In addition, our approach enjoys better convergence properties, making it more efficient and reliable in practice.
Qihang Yu, Lingxi Xie, Yan Wang 0033, Yuyin Zhou, Elliot K. Fishman, Alan L. Yuille
CVPR2
2018 DeepVoting: A Robust and Explainable Deep Network for Semantic Part Detection Under Partial Occlusion
abstract
In this paper, we study the task of detecting semantic parts of an object, e.g., a wheel of a car, under partial occlusion. We propose that all models should be trained without seeing occlusions while being able to transfer the learned knowledge to deal with occlusions. This setting alleviates the difficulty in collecting an exponentially large dataset to cover occlusion patterns and is more essential. In this scenario, the proposal-based deep networks, like RCNN-series, often produce unsatisfactory results, because both the proposal extraction and classification stages may be confused by the irrelevant occluders. To address this, [25] proposed a voting mechanism that combines multiple local visual cues to detect semantic parts. The semantic parts can still be detected even though some visual cues are missing due to occlusions. However, this method is manually-designed, thus is hard to be optimized in an end-to-end manner. In this paper, we present DeepVoting, which incorporates the robustness shown by [25] into a deep network, so that the whole pipeline can be jointly optimized. Specifically, it adds two layers after the intermediate features of a deep network, e.g., the pool-4 layer of VGGNet. The first layer extracts the evidence of local visual cues, and the second layer performs a voting mechanism by utilizing the spatial relationship between visual cues and semantic parts. We also propose an improved version DeepVoting+ by learning visual cues from context outside objects. In experiments, DeepVoting achieves significantly better performance than several baseline methods, including Faster-RCNN, for semantic part detection under occlusion. In addition, DeepVoting enjoys explainability as the detection results can be diagnosed via looking up the voting cues.
Zhishuai Zhang, Cihang Xie, Jianyu Wang 0001, Lingxi Xie, Alan L. Yuille
CVPR4
2018 Multi-scale Spatially-Asymmetric Recalibration for Image Classification
Yan Wang 0033, Lingxi Xie, Siyuan Qiao, Ya Zhang 0002, Wenjun Zhang 0001, Alan L. Yuille
ECCV (13)2
2018 Bridging the Gap Between 2D and 3D Organ Segmentation with Volumetric Fusion Net
Yingda Xia, Lingxi Xie, Fengze Liu, Zhuotun Zhu, Elliot K. Fishman, Alan L. Yuille
MICCAI (4)2
2018 Attention-based Pyramid Aggregation Network for Visual Place Recognition
abstract
Visual place recognition is challenging in the urban environment and is usually viewed as a large scale image retrieval task. The intrinsic challenges in place recognition exist that the confusing objects such as cars and trees frequently occur in the complex urban scene, and buildings with repetitive structures may cause over-counting and the burstiness problem degrading the image representations. To address these problems, we present an Attention-based Pyramid Aggregation Network (APANet), which is trained in an end-to-end manner for place recognition. One main component of APANet, the spatial pyramid pooling, can effectively encode the multi-size buildings containing geo-information. The other one, the attention block, is adopted as a region evaluator for suppressing the confusing regional features while highlighting the discriminative ones. When testing, we further propose a simple yet effective PCA power whitening strategy, which significantly improves the widely used PCA whitening by reasonably limiting the impact of over-counting. Experimental evaluations demonstrate that the proposed APANet outperforms the state-of-the-art methods on two place recognition benchmarks, and generalizes well on standard image retrieval datasets.
Yingying Zhu 0001, Lingxi Xie, Liang Zheng 0001
ACM Multimedia3
2017 Detecting Semantic Parts on Partially Occluded Objects
Jianyu Wang 0001, Zhishuai Zhang, Cihang Xie, Jun Zhu 0001, Lingxi Xie, Alan L. Yuille
BMVC5
2017 SORT: Second-Order Response Transform for Visual Recognition
abstract
In this paper, we reveal the importance and benefits of introducing second-order operations into deep neural networks. We propose a novel approach named Second-Order Response Transform (SORT), which appends element-wise product transform to the linear sum of a two-branch network module. A direct advantage of SORT is to facilitate cross-branch response propagation, so that each branch can update its weights based on the current status of the other branch. Moreover, SORT augments the family of transform operations and increases the nonlinearity of the network, making it possible to learn flexible functions to fit the complicated distribution of feature space. SORT can be applied to a wide range of network architectures, including a branched variant of a chain-styled network and a residual network, with very light-weighted modifications. We observe consistent accuracy gain on both small (CIFAR10, CIFAR100 and SVHN) and big (ILSVRC2012) datasets. In addition, SORT is very efficient, as the extra computation overhead is less than 5%.
Yan Wang 0033, Lingxi Xie, Chenxi Liu 0001, Siyuan Qiao, Ya Zhang 0002, Wenjun Zhang 0001, Qi Tian 0001, Alan L. Yuille
ICCV2
2017 Adversarial Examples for Semantic Segmentation and Object Detection
abstract
It has been well demonstrated that adversarial examples, i.e., natural images with visually imperceptible perturbations added, cause deep networks to fail on image classification. In this paper, we extend adversarial examples to semantic segmentation and object detection which are much more difficult. Our observation is that both segmentation and detection are based on classifying multiple targets on an image (e.g., the target is a pixel or a receptive field in segmentation, and an object proposal in detection). This inspires us to optimize a loss function over a set of targets for generating adversarial perturbations. Based on this, we propose a novel algorithm named Dense Adversary Generation (DAG), which applies to the state-of-the-art networks for segmentation and detection. We find that the adversarial perturbations can be transferred across networks with different training data, based on different architectures, and even for different recognition tasks. In particular, the transfer ability across networks with the same architecture is more significant than in other cases. Besides, we show that summing up heterogeneous perturbations often leads to better transfer performance, which provides an effective method of black-box adversarial attack.
Cihang Xie, Jianyu Wang 0001, Zhishuai Zhang, Yuyin Zhou, Lingxi Xie, Alan L. Yuille
ICCV5
2017 Genetic CNN
abstract
The deep convolutional neural network (CNN) is the state-of-the-art solution for large-scale visual recognition. Following some basic principles such as increasing network depth and constructing highway connections, researchers have manually designed a lot of fixed network architectures and verified their effectiveness. In this paper, we discuss the possibility of learning deep network structures automatically. Note that the number of possible network structures increases exponentially with the number of layers in the network, which motivates us to adopt the genetic algorithm to efficiently explore this large search space. The core idea is to propose an encoding method to represent each network structure in a fixed-length binary string. The genetic algorithm is initialized by generating a set of randomized individuals. In each generation, we define standard genetic operations, e.g., selection, mutation and crossover, to generate competitive individuals and eliminate weak ones. The competitiveness of each individual is defined as its recognition accuracy, which is obtained via a standalone training process on a reference dataset. We run the genetic process on CIFAR10, a small-scale dataset, demonstrating its ability to find high-quality structures which are little studied before. The learned powerful structures are also transferrable to the ILSVRC2012 dataset for large-scale visual recognition.
Lingxi Xie, Alan L. Yuille
ICCV1
2017 Object Recognition with and without Objects
abstract
While recent deep neural networks have achieved a promising performance on object recognition, they rely implicitly on the visual contents of the whole image. In this paper, we train deep neural networks on the foreground (object) and background (context) regions of images respectively. Considering human recognition in the same situations, networks trained on the pure background without objects achieves highly reasonable recognition performance that beats humans by a large margin if only given context. However, humans still outperform networks with pure object available, which indicates networks and human beings have different mechanisms in understanding an image. Furthermore, we straightforwardly combine multiple trained networks to explore different visual cues learned by different networks. Experiments show that useful visual hints can be explicitly learned separately and then combined to achieve higher performance, which verifies the advantages of the proposed framework.
Zhuotun Zhu, Lingxi Xie, Alan L. Yuille
IJCAI2
2017 Deep Supervision for Pancreatic Cyst Segmentation in Abdominal CT Scans
Yuyin Zhou, Lingxi Xie, Elliot K. Fishman, Alan L. Yuille
MICCAI (3)2
2017 A Fixed-Point Model for Pancreas Segmentation in Abdominal CT Scans
Yuyin Zhou, Lingxi Xie, Wei Shen 0002, Yan Wang 0033, Elliot K. Fishman, Alan L. Yuille
MICCAI (1)2
2017 Towards Reversal-Invariant Image Representation
Lingxi Xie, Jingdong Wang 0001, Weiyao Lin, Bo Zhang 0010, Qi Tian 0001
Int. J. Comput. Vis.1
2016 DisturbLabel: Regularizing CNN on the Loss Layer
abstract
During a long period of time we are combating overfitting in the CNN training process with model regularization, including weight decay, model averaging, data augmentation, etc. In this paper, we present DisturbLabel, an extremely simple algorithm which randomly replaces a part of labels as incorrect values in each iteration. Although it seems weird to intentionally generate incorrect training labels, we show that DisturbLabel prevents the network training from over-fitting by implicitly averaging over exponentially many networks which are trained with different label sets. To the best of our knowledge, DisturbLabel serves as the first work which adds noises on the loss layer. Meanwhile, DisturbLabel cooperates well with Dropout to provide complementary regularization functions. Experiments demonstrate competitive recognition results on several popular image recognition datasets.
Lingxi Xie, Jingdong Wang 0001, Meng Wang 0001, Qi Tian 0001
CVPR1
2016 InterActive: Inter-Layer Activeness Propagation
abstract
An increasing number of computer vision tasks can be tackled with deep features, which are the intermediate outputs of a pre-trained Convolutional Neural Network. Despite the astonishing performance, deep features extracted from low-level neurons are still below satisfaction, arguably because they cannot access the spatial context contained in the higher layers. In this paper, we present InterActive, a novel algorithm which computes the activeness of neurons and network connections. Activeness is propagated through a neural network in a top-down manner, carrying highlevel context and improving the descriptive power of lowlevel and mid-level neurons. Visualization indicates that neuron activeness can be interpreted as spatial-weighted neuron responses. We achieve state-of-the-art classification performance on a wide range of image datasets.
Lingxi Xie, Liang Zheng 0001, Jingdong Wang 0001, Alan L. Yuille, Qi Tian 0001
CVPR1
2016 Geometric Neural Phrase Pooling: Modeling the Spatial Co-occurrence of Neurons
Lingxi Xie, Qi Tian 0001, John Flynn, Jingdong Wang 0001, Alan L. Yuille
ECCV (1)1
2016 Fast Nearest Neighbor Search in the Hamming Space
Zhansheng Jiang, Lingxi Xie, Xiaotie Deng, Weiwei Xu 0003, Jingdong Wang 0001
MMM (1)2
2016 Incorporating visual adjectives for image classification
Lingxi Xie, Jingdong Wang 0001, Bo Zhang 0010, Qi Tian 0001
Neurocomputing1
2016 Simple Techniques Make Sense: Feature Pooling and Normalization for Image Classification
abstract
Image classification is a fundamental task in computer vision, implying a wide range of challenging problems, such as object recognition, scene understanding, and image tagging. One of the most popular approaches to image classification, the bag-of-features (BoF) model, represents an image with a long feature vector and adopts machine learning algorithms for training and testing. Owing to its simplicity and scalability, the BoF model is widely used in both academic research studies and industrial applications. This paper discusses the feature summarization stage, including pooling and normalization, in the BoF model. We show that these two modules, although devalued sometimes, have important impacts on image classification performance. We present two algorithms, i.e., generalized regular spatial pooling for constructing a better group of spatial bins and hierarchical feature normalization for assigning proper weights for regional feature normalization. Both algorithms are independent of the descriptor extraction and feature encoding stages, and therefore, they could be freely transplanted onto many other classification frameworks based on local feature statistics. We further provide insightful discussions for the nature of designing efficient image classification models. Experiments verify that the proposed algorithm achieves state-of-the-art results on a wide range of image classification data sets.
Lingxi Xie, Qi Tian 0001, Bo Zhang 0010
IEEE Trans. Circuits Syst. Video Technol.1
2015 RIDE: Reversal Invariant Descriptor Enhancement
abstract
In many fine-grained object recognition datasets, image orientation (left/right) might vary from sample to sample. Since handcrafted descriptors such as SIFT are not reversal invariant, the stability of image representation based on them is consequently limited. A popular solution is to augment the datasets by adding a left-right reversed copy for each original image. This strategy improves recognition accuracy to some extent, but also brings the price of almost doubled time and memory consumptions. In this paper, we present RIDE (Reversal Invariant Descriptor Enhancement) for fine-grained object recognition. RIDE is a generalized algorithm which cancels out the impact of image reversal by estimating the orientation of local descriptors, and guarantees to produce the identical representation for an image and its left-right reversed copy. Experimental results reveal the consistent accuracy gain of RIDE with various types of descriptors. We also provide insightful discussions on the working mechanism of RIDE and its generalization to other applications.
Lingxi Xie, Jingdong Wang 0001, Weiyao Lin, Bo Zhang 0010, Qi Tian 0001
ICCV1
2015 Fine-grained visual categorization with fine-tuned segmentation
abstract
Fine-grained visual categorization (FGVC) refers to the task of classifying objects that belong to the same basic-level class (e.g., different bird species). Since the subtle inter-class variation often exists on small parts (e.g., beak, belly, etc.), it is reasonable to localize semantic parts of an object before describing it. However, unsupervised part-segmentation methods often suffer from over-segmentation which harms the quality of image representation. In this paper, we present a fine-tuning approach to tackle this problem. To this end, we perform a greedy algorithm to optimize an intuitive objective function, preserving principal parts meanwhile filtering noises, and further construct mid-level parts beyond the refined parts toward a more descriptive representation. Experiments demonstrate that our approach achieves competitive classification accuracy on the CUB-200-2011 dataset with both Fisher vectors and deep conv-net features.
Yanqing Guo, Lingxi Xie, Xiangwei Kong 0001, Qi Tian 0001
ICIP3
2015 Image Classification and Retrieval are ONE
abstract
In this paper, we demonstrate that the essentials of image classification and retrieval are the same, since both tasks could be tackled by measuring the similarity between images. To this end, we propose ONE (Online Nearest-neighbor Estimation), a unified algorithm for both image classification and retrieval. ONE is surprisingly simple, which only involves manual object definition, regional description and nearest-neighbor search. We take advantage of PCA and PQ approximation and GPU parallelization to scale our algorithm up to large-scale image search. Experimental results verify that ONE achieves state-of-the-art accuracy in a wide range of image classification and retrieval benchmarks.
Lingxi Xie, Richang Hong, Bo Zhang 0010, Qi Tian 0001
ICMR1
2015 Heterogeneous Graph Propagation for Large-Scale Web Image Search
abstract
State-of-the-art web image search frameworks are often based on the bag-of-visual-words (BoVWs) model and the inverted index structure. Despite the simplicity, efficiency, and scalability, they often suffer from low precision and/or recall, due to the limited stability of local features and the considerable information loss on the quantization stage. To refine the quality of retrieved images, various postprocessing methods have been adopted after the initial search process. In this paper, we investigate the online querying process from a graph-based perspective. We introduce a heterogeneous graph model containing both image and feature nodes explicitly, and propose an efficient reranking approach consisting of two successive modules, i.e., incremental query expansion and image-feature voting, to improve the recall and precision, respectively. Compared with the conventional reranking algorithms, our method does not require using geometric information of visual words, therefore enjoys low consumptions of both time and memory. Moreover, our method is independent of the initial search process, and could cooperate with many BoVW-based image search pipelines, or adopted after other postprocessing algorithms. We evaluate our approach on large-scale image search tasks and verify its competitive search performance.
Lingxi Xie, Qi Tian 0001, Wengang Zhou 0001, Bo Zhang 0010
IEEE Trans. Image Process.1
2015 Fine-Grained Image Search
abstract
Large-scale image search has been attracting lots of attention from both academic and commercial fields. The conventional bag-of-visual-words (BoVW) model with inverted index is verified efficient at retrieving near-duplicate images, but it is less capable of discovering fine-grained concepts in the query and returning semantically matched search results. In this paper, we suggest that instance search should return not only near-duplicate images, but also fine-grained results, which is usually the actual intention of a user. We propose a new and interesting problem named fine-grained image search, which means that we prefer those images containing the same fine-grained concept with the query. We formulate the problem by constructing a hierarchical database and defining an evaluation method. We thereafter introduce a baseline system using fine-grained classification scores to represent and co-index images so that the semantic attributes are better incorporated in the online querying stage. Large-scale experiments reveal that promising search results are achieved with reasonable time and memory consumption. We hope this paper will be the foundation for future work on image search. We also expect more follow-up efforts along this research topic and look forward to commercial fine-grained image search engines.
Lingxi Xie, Jingdong Wang 0001, Bo Zhang 0010, Qi Tian 0001
IEEE Trans. Multim.1
2014 Orientational Pyramid Matching for Recognizing Indoor Scenes
abstract
Scene recognition is a basic task towards image understanding. Spatial Pyramid Matching (SPM) has been shown to be an efficient solution for spatial context modeling. In this paper, we introduce an alternative approach, Orientational Pyramid Matching (OPM), for orientational context modeling. Our approach is motivated by the observation that the 3D orientations of objects are a crucial factor to discriminate indoor scenes. The novelty lies in that OPM uses the 3D orientations to form the pyramid and produce the pooling regions, which is unlike SPM that uses the spatial positions to form the pyramid. Experimental results on challenging scene classification tasks show that OPM achieves the performance comparable with SPM and that OPM and SPM make complementary contributions so that their combination gives the state-of-the-art performance.
Lingxi Xie, Jingdong Wang 0001, Baining Guo, Bo Zhang 0010, Qi Tian 0001
CVPR1
2014 Max-SIFT: Flipping invariant descriptors for Web logo search
abstract
Logo search is widely required in many real-world applications. As a special case of near-duplicate images, logo pictures have some particular properties, for instance, suffering from flipping operations, e.g., geometry-inverted and brightness-inverted operations. Such operations completely change the spatial structure of local descriptors, such as SIFT, so that image search algorithms based on Bag-of-Visual-Words (BoVW) often fail to retrieve the flipped logos. We propose a novel descriptor named Max-SIFT, which finds the maximal SIFT value sequence for detecting flipping operations. Compared with previous algorithms, our algorithm is extremely easy to implement yet very efficient to carry out. We evaluate the improved descriptor on a large-scale Web logo search dataset, and demonstrate that our method enjoys good performance and low computational costs.
Lingxi Xie, Qi Tian 0001, Bo Zhang 0010
ICIP1
2014 Fast and accurate near-duplicate image search with affinity propagation on the ImageWeb
Lingxi Xie, Qi Tian 0001, Wengang Zhou 0001, Bo Zhang 0010
Comput. Vis. Image Underst.1
2014 Spatial Pooling of Heterogeneous Features for Image Classification
abstract
In image classification tasks, one of the most successful algorithms is the bag-of-features (BoFs) model. Although the BoF model has many advantages, such as simplicity, generality, and scalability, it still suffers from several drawbacks, including the limited semantic description of local descriptors, lack of robust structures upon single visual words, and missing of efficient spatial weighting. To overcome these shortcomings, various techniques have been proposed, such as extracting multiple descriptors, spatial context modeling, and interest region detection. Though they have been proven to improve the BoF model to some extent, there still lacks a coherent scheme to integrate each individual module together. To address the problems above, we propose a novel framework with spatial pooling of complementary features. Our model expands the traditional BoF model on three aspects. First, we propose a new scheme for combining texture and edge-based local features together at the descriptor extraction level. Next, we build geometric visual phrases to model spatial context upon complementary features for midlevel image representation. Finally, based on a smoothed edgemap, a simple and effective spatial weighting scheme is performed to capture the image saliency. We test the proposed framework on several benchmark data sets for image classification. The extensive results show the superior performance of our algorithm over the state-of-the-art methods.
Lingxi Xie, Qi Tian 0001, Meng Wang 0001, Bo Zhang 0010
IEEE Trans. Image Process.1
2013 Hierarchical Part Matching for Fine-Grained Visual Categorization
abstract
As a special topic in computer vision, fine-grained visual categorization (FGVC) has been attracting growing attention these years. Different with traditional image classification tasks in which objects have large inter-class variation, the visual concepts in the fine-grained datasets, such as hundreds of bird species, often have very similar semantics. Due to the large inter-class similarity, it is very difficult to classify the objects without locating really discriminative features, therefore it becomes more important for the algorithm to make full use of the part information in order to train a robust model. In this paper, we propose a powerful flowchart named Hierarchical Part Matching (HPM) to cope with fine-grained classification tasks. We extend the Bag-of-Features (BoF) model by introducing several novel modules to integrate into image representation, including foreground inference and segmentation, Hierarchical Structure Learning (HSL), and Geometric Phrase Pooling (GPP). We verify in experiments that our algorithm achieves the state-of-the-art classification accuracy in the Caltech-UCSD-Birds-200-2011 dataset by making full use of the ground-truth part annotations.
Lingxi Xie, Qi Tian 0001, Richang Hong, Shuicheng Yan, Bo Zhang 0010
ICCV1
2013 Feature normalization for part-based image classification
abstract
Part-based Bag-of-Features (BoF) models such as Spatial Pyramid Matching (SPM) play an important role in image classification. Before sending the feature vectors into classifiers for training and testing, it is required to normalize them in order to approximately equalize ranges of the attributes and make them have comparable effects in distance computation. Although some works have been focused on general feature normalization, we do not see any discussion on specialized normalization algorithms for part-based BoF models. In this paper, we fill in the blank with extensive experiments and discussions. Based on solid normalization parameters (power and coefficient), we further study two straightforward part-based properties, i.e., the independent assumption and the hierarchical-contribution assumption, to scale the feature super-vectors separately. Finally, we test our algorithm on challenging image sets, i.e., Caltech 101 and CUB-200-2011, for general and fine-grained classification, and show its efficiency, scalability and adaptability in both scenarios.
Lingxi Xie, Qi Tian 0001, Bo Zhang 0010
ICIP1
2012 Spatial pooling of heterogeneous features for image applications
abstract
The Bag-of-Features (BoF) model has played an important role for image representation in many multimedia applications. It has been extensively applied to many tasks including image classification, image retrieval, scene understanding, and so on. Despite the advantages of this model such as simplicity, efficiency and generality, there are also notable drawbacks for this model, including poor power of semantic expression of local descriptors, and lack of robust structures upon single visual words. To overcome these problems, various techniques have been proposed, such as multiple descriptors, spatial context modeling and interest region detection. Though they have been proven to improve the BoF model to some extent, there still lacks a coherent scheme to integrate each individual module.
Lingxi Xie, Qi Tian 0001, Bo Zhang 0010
ACM Multimedia1