EDBT 2026 Demo / reviewers in the wild / expert
Xiaopeng Zhang 0008
dblp:167/1261-8
· DBLP profile ↗
77ranked-venue papers
11as first author
57since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 4 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 11 first-author · 32 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | O-DisCo-Edit: Object Distortion Control for Unified Realistic Video EditingabstractDiffusion models have recently advanced video editing, yet controllable editing remains challenging due to the need for precise manipulation of diverse object properties. Current methods require different control signal for diverse editing tasks, which complicates model design and demands significant training resources. To address this, we propose O-DisCo-Edit, a unified framework that incorporates a novel object distortion control (O-DisCo). This signal, based on random and adaptive noise, flexibly encapsulates a wide range of editing cues within a single representation. Paired with a “copy-form” preservation module for preserving non-edited regions, O-DisCo-Edit enables efficient, high-fidelity editing through an effective training paradigm. Extensive experiments and comprehensive human evaluations consistently demonstrate that O-DisCo-Edit surpasses both specialized and multitask state-of-the-art methods across various video editing tasks. Junjie Wang 0012, Lin Liu 0016, Ruihang Chu, Xiaopeng Zhang 0008, Qi Tian 0001, Yujiu Yang 0001 |
AAAI | 5 |
| 2026 | Vision-Language Efficient Tuning for Mitigating Catastrophic Forgetting in Multi-Modal Learning
Yuchen Liu 0006, Wenrui Dai, Xiaopeng Zhang 0008, Junni Zou, Qi Tian 0001, Hongkai Xiong |
Int. J. Comput. Vis. | 4 |
| 2025 | Segment Any 3D GaussiansabstractThis paper presents SAGA (Segment Any 3D GAussians), a highly efficient 3D promptable segmentation method based on 3D Gaussian Splatting (3D-GS). Given 2D visual prompts as input, SAGA can segment the corresponding 3D target represented by 3D Gaussians within 4 ms. This is achieved by attaching a scale-gated affinity feature to each 3D Gaussian to endow it a new property towards multi-granularity segmentation. Specifically, a scale-aware contrastive training strategy is proposed for the scale-gated affinity feature learning. It 1) distills the segmentation capability of the Segment Anything Model (SAM) from 2D masks into the affinity features and 2) employs a soft scale gate mechanism to deal with multi-granularity ambiguity in 3D segmentation through adjusting the magnitude of each feature channel according to a specified 3D physical scale. Evaluations demonstrate that SAGA achieves real-time multi-granularity segmentation with quality comparable to state-of-the-art methods. As one of the first methods addressing promptable segmentation in 3D-GS, the simplicity and effectiveness of SAGA pave the way for future advancements in this field. Jiazhong Cen, Jiemin Fang, Chen Yang 0023, Lingxi Xie, Xiaopeng Zhang 0008, Wei Shen 0002, Qi Tian 0001 |
AAAI | 5 |
| 2025 | IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot MannerabstractControllability of video generation has been recently concerned in addition to the quality of generated videos. The main challenge to controllable video generation is to synthesize videos based on user-specified instance spatial locations and movement trajectories. However, existing methods suffer from a dilemma between the resource consumption, generation quality, and user controllability. As an efficient alternative to prohibitive training-based video generation, existing zero-shot video generation methods cannot generate high-quality and motion-consistent videos under the control of layouts and movement trajectories. In this paper, we propose a novel zero-shot method named IM-Zero that ameliorates instance-level motion controllable video generation with enhanced control accuracy, motion consistency, and richness of details to address this problem. Specifically, we first present a motion generation stage that extracts motion and textural guidance from keyframe candidates from pre-trained grounded text-to-image model to generate the desired coarse motion video. Subsequently, we develop a video refinement stage that injects the motion priors of pre-trained text-to-video models and detail priors of pre-trained text-to-image models into the latents of coarse motion videos to further enhance video motion consistency and richness of details. To our best knowledge, IM-Zero is the first to simultaneously achieve high-quality video generation and allow to control both layouts and movement trajectories in a zero-shot manner. Extensive experiments demonstrate that IM-Zero outperforms existing methods in terms of video quality, inter-frame consistency, and the alignment of location and trajectory. Furthermore, compared with existing methods, IM-Zero enjoys extra advantages of versatility in video generation, including motion control of subparts within instances, finer control of specifying instance shapes via masks, and more difficult tasks of motion transfer for customizing fine-grained motion patterns through reference videos and high-quality text-to-video generation. Yabo Chen, Xiaopeng Zhang 0008, Wenrui Dai, Junni Zou, Hongkai Xiong, Qi Tian 0001 |
CVPR | 4 |
| 2025 | CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic Segmentation
Jiannan Ge, Lingxi Xie, Hongtao Xie 0001, Pandeng Li, Sun-Ao Liu, Xiaopeng Zhang 0008, Qi Tian 0001, Yongdong Zhang 0001 |
ICCV | 6 |
| 2025 | METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models
Yuchen Liu 0006, Bowen Shi 0003, Xiaopeng Zhang 0008, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
ICCV | 4 |
| 2025 | SAM-CP: Marrying SAM with Composable Prompts for Versatile SegmentationabstractThe Segment Anything model (SAM) has shown a generalized ability to group image pixels into patches, but applying it to semantic-aware segmentation still faces major challenges. This paper presents SAM-CP, a simple approach that establishes two types of composable prompts beyond SAM and composes them for versatile segmentation. Specifically, given a set of classes (in texts) and a set of SAM patches, the Type-I prompt judges whether a SAM patch aligns with a text label, and the Type-II prompt judges whether two SAM patches with the same text label also belong to the same instance. To decrease the complexity in dealing with a large number of semantic classes and patches, we establish a unified framework that calculates the affinity between (semantic and instance) queries and SAM patches, and then merges patches with high affinity to the query. Experiments show that SAM-CP achieves semantic, instance, and panoptic segmentation in both open and closed domains. In particular, it achieves state-of-the-art performance in open-vocabulary segmentation. Our research offers a novel and generalized methodology for equipping vision foundation models like SAM with multi-grained semantic perception abilities. Codes are released on https://github.com/ucas-vg/SAM-CP. Pengfei Chen 0004, Lingxi Xie, Xinyue Huo, Xuehui Yu, Xiaopeng Zhang 0008, Yingfei Sun, Zhenjun Han, Qi Tian 0001 |
ICLR | 5 |
| 2025 | Tackling View-Dependent Semantics in 3D Language Gaussian SplattingabstractRecent advancements in 3D Gaussian Splatting (3D-GS) enable high-quality 3D scene reconstruction from RGB images. Many studies extend this paradigm for language-driven open-vocabulary scene understanding. However, most of them simply project 2D semantic features onto 3D Gaussians and overlook a fundamental gap between 2D and 3D understanding: a 3D object may exhibit various semantics from different viewpoints—a phenomenon we term **view-dependent semantics**. To address this challenge, we propose **LaGa** (**La**nguage **Ga**ussians), which establishes cross-view semantic connections by decomposing the 3D scene into objects. Then, it constructs view-aggregated semantic representations by clustering semantic descriptors and reweighting them based on multi-view semantics. Extensive experiments demonstrate that LaGa effectively captures key information from view-dependent semantics, enabling a more comprehensive understanding of 3D scenes. Notably, under the same settings, LaGa achieves a significant improvement of **+18.7\% mIoU** over the previous SOTA on the LERF-OVS dataset. Our code is available at: https://github.com/https://github.com/SJTU-DeepVisionLab/LaGa. Jiazhong Cen, Jiemin Fang, Changsong Wen, Lingxi Xie, Xiaopeng Zhang 0008, Wei Shen 0002, Qi Tian 0001 |
ICML | 6 |
| 2025 | Diffusion-Driven Progressive Target Manipulation for Source-Free Domain AdaptationabstractSource-free domain adaptation (SFDA) is a challenging task that tackles domain shifts using only a pre-trained source model and unlabeled target data. Existing SFDA methods are restricted by the fundamental limitation of source-target domain discrepancy. Non-generation SFDA methods suffer from unreliable pseudo-labels in challenging scenarios with large domain discrepancies, while generation-based SFDA methods are evidently degraded due to enlarged domain discrepancies in creating pseudo-source data. To address this limitation, we propose a novel generation-based framework named Diffusion-Driven Progressive Target Manipulation (DPTM) that leverages unlabeled target data as references to reliably generate and progressively refine a pseudo-target domain for SFDA. Specifically, we divide the target samples into a trust set and a non-trust set based on the reliability of pseudo-labels to sufficiently and reliably exploit their information. For samples from the non-trust set, we develop a manipulation strategy to semantically transform them into the newly assigned categories, while simultaneously maintaining them in the target distribution via a latent diffusion model. Furthermore, we design a progressive refinement mechanism that progressively reduces the domain discrepancy between the pseudo-target domain and the real target domain via iterative refinement. Experimental results demonstrate that DPTM outperforms existing methods by a large margin and achieves state-of-the-art performance on four prevailing SFDA benchmark datasets with different scales. Remarkably, DPTM can significantly enhance the performance by up to 18.6\% in scenarios with large source-target gaps. Yabo Chen, Junyu Zhou 0001, Wenrui Dai, Xiaopeng Zhang 0008, Junni Zou, Hongkai Xiong, Qi Tian 0001 |
NeurIPS | 5 |
| 2025 | Rethinking visual prompt learning as masked visual token modeling
Ning Liao, Bowen Shi 0003, Xiaopeng Zhang 0008, Min Cao 0005, Junchi Yan, Qi Tian 0001 |
Artif. Intell. | 3 |
| 2025 | Segment Anything in 3D with Radiance Fields
Jiazhong Cen, Jiemin Fang, Zanwei Zhou, Chen Yang 0023, Lingxi Xie, Xiaopeng Zhang 0008, Wei Shen 0002, Qi Tian 0001 |
Int. J. Comput. Vis. | 6 |
| 2025 | Contrastive Learning via Variational Information BottleneckabstractRecent advances in self-supervised learning have witnessed great achievements, especially with the introduction of contrastive learning, where the goal is to maximize the mutual information between different augmentations of the same image, i.e., positive pairs. However, such optimization does not necessarily correspond to optimal representation due to noisy samples, thus inevitably being over-confident in the relevance between views. As a result, the learned model would capture spurious correlation and retain superfluous information that deteriorates representations. In this paper, we facilitate contrastive learning by reducing superfluous relevance between positive views. To this end, we introduce the representation entropy minimization regularization over the objective of vanilla contrastive learning, which forces representations to retain possibly the least information, thus alleviating superfluous relevance from irrelevant views. Then, we derive the analytical expression of the proposed objective by converting it to an information bottleneck problem and solving via variation approximation, which leads to a novel contrastive learning framework, termed as CLIMB, short for Contrastive Learning via variational InforMation Bottleneck. Experiments over multiple benchmarks demonstrate that CLIMB brings consistent improvement. Notably, using DINO as an instantiation, CLIMB achieves 4.5% and 3.5% gain under the k-NN classification metric with EfficientNet-B0 and ResNet-50 as backbones, respectively. Jin Li 0057, Xiaopeng Zhang 0008, Dongsheng Jiang, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | M-Tuning: Prompt Tuning With Mitigated Label Bias in Open-Set ScenariosabstractIn realistic open-set scenarios where labels of a part of testing data are totally unknown, when vision-language (VL) prompt learning methods encounter inputs related to unknown classes (i.e., not seen during training), they always predict them as one of the training classes. The exhibited label bias causes difficulty in open set recognition (OSR), in which an image should be correctly predicted as one of the known classes or the unknown one. To achieve this goal, we propose a vision-language prompt tuning method with mitigated label bias (M-Tuning). It introduces open words from the WordNet to extend the range of words forming the prompt texts from only closed-set label words to more, and thus prompts are tuned in a simulated open-set scenario. Besides, inspired by the observation that classifying directly on large datasets causes a much higher false positive rate than on small datasets, we propose a Combinatorial Tuning and Testing (CTT) strategy for improving performance. CTT decomposes M-Tuning on large datasets as multiple independent group-wise tuning on fewer classes, then makes accurate and comprehensive predictions by selecting the optimal sub-prompt. Finally, given the lack of VL-based OSR baselines in the literature, especially for prompt methods, we contribute new baselines for fair comparisons. Our method achieves the best performance on datasets with various scales, and extensive ablation studies also validate its effectiveness. Ning Liao, Xiaopeng Zhang 0008, Min Cao 0005, Junchi Yan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | MENSA: Multi-Dataset Harmonized Pretraining for Semantic SegmentationabstractExisting pretraining methods for semantic segmentation are hampered by the task gap between global image -level pretraining and local pixel-level finetuning. Joint dense-level pretraining is a promising alternative to exploit off-the-shelf annotations from diverse segmentation datasets but suffers from low-quality class embeddings and inconsistent data and supervision signals across multiple datasets by directly employing CLIP. To overcome these challenges, we propose a novelMulti-datasEt harmoNized pretraining framework forSemantic sEgmentation (MENSA). MENSA incorporates high-quality language embeddings and momentum-updated visual embeddings to effectively model the class relationships in the embedding space and thereby provide reliable supervision information for each category. To further adapt to multiple datasets, we achieve one-to-many pixel-embedding pairing with cross-dataset multi-label mapping through cross-modal information exchange to mitigate inconsistent supervision signals and introduce region-level and pixel-level cross-dataset mixing for varying data distribution. Experimental results demonstrate that MENSA is a powerful foundation segmentation model that consistently outperforms popular supervised or unsupervised ImageNet pretrained models for various benchmarks under standard fine-tuning. Furthermore, MENSA is shown to significantly benefit frozen-backbone fine-tuning and zero-shot learning by endowing pixel-level distinctiveness to learned representations. Bowen Shi 0003, Xiaopeng Zhang 0008, Wenrui Dai, Junni Zou, Hongkai Xiong |
IEEE Trans. Multim. | 2 |
| 2024 | GaussianEditor: Editing 3D Gaussians Delicately with Text InstructionsabstractRecently, impressive results have been achieved in 3D scene editing with text instructions based on a 2D diffusion model. However, current diffusion models primarily generate images by predicting noise in the latent space, and the editing is usually applied to the whole image, which makes it challenging to perform delicate, especially localized, editing for 3D scenes. Inspired by recent 3D Gaussian splatting, we propose a systematic framework, named Gaus-sianEditor, to edit 3D scenes delicately via 3D Gaussians with text instructions. Benefiting from the explicit property of 3D Gaussians, we design a series of techniques to achieve delicate editing. Specifically, we first extract the region of interest (RoI) corresponding to the text instruction, aligning it to 3D Gaussians. The Gaussian RoI is further used to control the editing process. Our framework can achieve more delicate and precise editing of 3D scenes than previous methods while enjoying much faster training speed, i.e. within 20 minutes on a single V100 GPU, more than twice as fast as Instruct-NeRF2NeRF (45 minutes - 2 hours)11The editing time varies in different scenes according to the scene structure complexity.. The project page is at GaussianEditor. github.io. Junjie Wang 0012, Jiemin Fang, Xiaopeng Zhang 0008, Lingxi Xie, Qi Tian 0001 |
CVPR | 3 |
| 2024 | 4D Gaussian Splatting for Real-Time Dynamic Scene RenderingabstractRepresenting and rendering dynamic scenes has been an important but challenging task. Especially, to accurately model complex motions, high efficiency is usually hard to guarantee. To achieve real-time dynamic scene rendering while also enjoying high training and storage efficiency, we propose 4D Gaussian Splatting (4D-GS) as a holistic representation for dynamic scenes rather than applying 3D-GS for each individual frame. In 4D-GS, a novel explicit representation containing both 3D Gaussians and 4D neural voxels is proposed. A decomposed neural voxel encoding algorithm inspired by HexPlane is proposed to efficiently build Gaussian features from 4D neural voxels and then a lightweight MLP is applied to predict Gaussian deformations at novel timestamps. Our 4D-GS method achieves real-time rendering under high resolutions, 82 FPS at an 800x800 resolution on an RTX 3090 GPU while maintaining comparable or better quality than previous state- of-the-art methods. More demos and code are available at https://guanjunwu.github.io/4dgs/. Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang 0008, Wei Wei 0002, Wenyu Liu 0001, Qi Tian 0001, Xinggang Wang |
CVPR | 5 |
| 2024 | GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion ModelsabstractIn recent times, the generation of 3D assets from text prompts has shown impressive results. Both 2D and 3D diffusion models can help generate decent 3D objects based on prompts. 3D diffusion models have good 3D consistency, but their quality and generalization are limited as trainable 3D data is expensive and hard to obtain. 2D diffusion models enjoy strong abilities of generalization and fine generation, but 3D consistency is hard to guarantee. This paper attempts to bridge the power from the two types of diffusion models via the recent explicit and efficient 3D Gaussian splatting representation. A fast 3D object gener-ation framework, named as GaussianDreamer, is proposed, where the 3D diffusion model provides priors for initial-ization and the 2D diffusion model enriches the geometry and appearance. Operations of noisy point growing and color perturbation are introduced to enhance the initialized Gaussians. Our GaussianDreamer can generate a high-quality 3D instance or 3D avatar within 15 minutes on one GPU, much faster than previous methods, while the generated instances can be directly rendered in real time. Demos and code are available at https://taoranyi.com/gaussiandreamer/. Taoran Yi, Jiemin Fang, Junjie Wang 0012, Guanjun Wu, Lingxi Xie, Xiaopeng Zhang 0008, Wenyu Liu 0001, Qi Tian 0001, Xinggang Wang |
CVPR | 6 |
| 2024 | Cascade-Zero123: One Image to Highly Consistent 3D with Self-prompted Nearby Views
Yabo Chen, Jiemin Fang, Taoran Yi, Xiaopeng Zhang 0008, Lingxi Xie, Xinggang Wang, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
ECCV (41) | 5 |
| 2024 | AlignZeg: Mitigating Objective Misalignment for Zero-Shot Semantic Segmentation
Jiannan Ge, Lingxi Xie, Hongtao Xie 0001, Pandeng Li, Xiaopeng Zhang 0008, Yongdong Zhang 0001, Qi Tian 0001 |
ECCV (43) | 5 |
| 2024 | DomainFusion: Generalizing to Unseen Domains with Latent Diffusion Models
Yabo Chen, Yuchen Liu 0006, Xiaopeng Zhang 0008, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
ECCV (41) | 4 |
| 2024 | UMG-CLIP: A Unified Multi-granularity Vision Generalist for Open-World Understanding
Bowen Shi 0003, Peisen Zhao, Yuhang Zhang 0012, Jin Li 0057, Wenrui Dai, Junni Zou, Hongkai Xiong, Qi Tian 0001, Xiaopeng Zhang 0008 |
ECCV (38) | 11 |
| 2024 | Hybrid Distillation: Connecting Masked Autoencoders with Contrastive LearnersabstractAs two prominent strategies for representation learning, Contrastive Learning (CL) and Masked Image Modeling (MIM) have witnessed significant progress. Previous studies have demonstrated the advantages of each approach in specific scenarios. CL, resembling supervised pre-training, excels at capturing longer-range global patterns and enhancing feature discrimination, while MIM is adept at introducing local and diverse attention across transformer layers. Considering the respective strengths, previous studies utilize feature distillation to inherit both discrimination and diversity. In this paper, we thoroughly examine previous feature distillation methods and observe that the increase in diversity mainly stems from asymmetric designs, which may in turn compromise the discrimination ability. To strike a balance between the two properties, we propose a simple yet effective strategy termed Hybrid Distill, which leverages both the CL and MIM teachers to jointly guide the student model. Hybrid Distill emulates the token relations of the MIM teacher at intermediate layers for diversity, while simultaneously distilling the final features of the CL teacher to enhance discrimination. A progressive redundant token masking strategy is employed to reduce the expenses associated with distillation and aid in preventing the model from converging to local optima. Experimental results demonstrate that Hybrid Distill achieves superior performance on various benchmark datasets. Bowen Shi 0003, Xiaopeng Zhang 0008, Jin Li 0057, Wenrui Dai, Junni Zou, Hongkai Xiong, Qi Tian 0001 |
ICLR | 2 |
| 2024 | BarLeRIa: An Efficient Tuning Framework for Referring Image SegmentationabstractPre-training followed by full fine-tuning has gradually been substituted by Parameter-Efficient Tuning (PET) in the field of computer vision. PET has gained popularity, especially in the context of large-scale models, due to its ability to reduce transfer learning costs and conserve hardware resources. However, existing PET approaches primarily focus on recognition tasks and typically support uni-modal optimization, while neglecting dense prediction tasks and vision language interactions. To address this limitation, we propose a novel PET framework called **B**i-direction**a**l Inte**r**twined Vision **L**anguage Effici**e**nt Tuning for **R**eferring **I**mage Segment**a**tion (**BarLeRIa**), which leverages bi-directional intertwined vision language adapters to fully exploit the frozen pre-trained models' potential in cross-modal dense prediction tasks. In BarLeRIa, two different tuning modules are employed for efficient attention, one for global, and the other for local, along with an intertwined vision language tuning module for efficient modal fusion.
Extensive experiments conducted on RIS benchmarks demonstrate the superiority of BarLeRIa over prior PET methods with a significant margin, i.e., achieving an average improvement of 5.6\%. Remarkably, without requiring additional training datasets, BarLeRIa even surpasses SOTA full fine-tuning approaches. The code is available at https://github.com/NastrondAd/BarLeRIa. Jin Li 0057, Xiaopeng Zhang 0008, Bowen Shi 0003, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
ICLR | 3 |
| 2024 | QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language ModelsabstractRecently years have witnessed a rapid development of large language models (LLMs). Despite the strong ability in many language-understanding tasks, the heavy computational burden largely restricts the application of LLMs especially when one needs to deploy them onto edge devices. In this paper, we propose a quantization-aware low-rank adaptation (QA-LoRA) algorithm. The motivation lies in the imbalanced degrees of freedom of quantization and adaptation, and the solution is to use group-wise operators which increase the degree of freedom of quantization meanwhile decreasing that of adaptation. QA-LoRA is easily implemented with a few lines of code, and it equips the original LoRA with two-fold abilities: (i) during fine-tuning, the LLM's weights are quantized (e.g., into INT4) to reduce time and memory usage; (ii) after fine-tuning, the LLM and auxiliary weights are naturally integrated into a quantized model without loss of accuracy. We apply QA-LoRA to the LLaMA and LLaMA2 model families and validate its effectiveness in different fine-tuning datasets and downstream scenarios. The code is made available at https://github.com/yuhuixu1993/qa-lora. Yuhui Xu 0002, Lingxi Xie, Xiaotao Gu, Xin Chen 0033, Heng Chang, Hengheng Zhang, Zhengsu Chen, Xiaopeng Zhang 0008, Qi Tian 0001 |
ICLR | 8 |
| 2024 | ControlVideo: Training-free Controllable Text-to-video GenerationabstractText-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart lags behind due to the excessive training cost.
To avert the training burden, we propose a training-free ControlVideo to produce high-quality videos based on the provided text prompts and motion sequences.
Specifically, ControlVideo adapts a pre-trained text-to-image model (i.e., ControlNet) for controllable text-to-video generation.
To generate continuous videos without flicker effect, we propose an interleaved-frame smoother to smooth the intermediate frames.
In particular, interleaved-frame smoother splits the whole videos with successive three-frame clips, and stabilizes each clip by updating the middle frame with the interpolation among other two frames in latent space.
Furthermore, a fully cross-frame interaction mechanism have been exploited to further enhance the frame consistency, while a hierarchical sampler is employed to produce long videos efficiently.
Extensive experiments demonstrate that our ControlVideo outperforms the state-of-the-arts both quantitatively and qualitatively.
It is worthy noting that, thanks to the efficient designs, ControlVideo could generate both short and long videos within several minutes using one NVIDIA 2080Ti.
Code and videos are available at [this link](https://github.com/YBYBZhang/ControlVideo). Yabo Zhang, Yuxiang Wei 0001, Dongsheng Jiang, Xiaopeng Zhang 0008, Wangmeng Zuo, Qi Tian 0001 |
ICLR | 4 |
| 2024 | Bootstrap AutoEncoders With Contrastive Paradigm for Self-supervised Gaze EstimationabstractExisting self-supervised methods for gaze estimation using the dominant streams of contrastive and generative approaches are restricted to eye images and could fail in general full-face settings. In this paper, we reveal that contrastive methods are ineffective in data augmentation for self-supervised full-face gaze estimation, while generative methods are prone to trivial solutions due to the absence of explicit regularization on semantic representations. To address this challenge, we propose a novel approach called **B**ootstrap auto-**e**ncoders with **C**ontrastive p**a**radigm (**BeCa**), which combines the strengths of both generative and contrastive methods. Specifically, we revisit the Auto-Encoder used in generative approaches and incorporate the contrastive paradigm to introduce explicit regularization on gaze representation. Furthermore, we design the InfoMSE loss as an alternative to the vanilla MSE loss for Auto-Encoder to mitigate the inconsistency between reconstruction and representation learning. Experimental results demonstrate that the proposed approaches outperform state-of-the-art unsupervised gaze approaches on extensive datasets (including wild scenes) under both within-dataset and cross-dataset protocols. Jin Li 0057, Wenrui Dai, Bowen Shi 0003, Xiaopeng Zhang 0008, Hongkai Xiong |
ICML | 5 |
| 2024 | SDPT: Semantic-Aware Dimension-Pooling Transformer for Image SegmentationabstractImage segmentation plays a critical role in autonomous driving by providing vehicles with a detailed and accurate understanding of their surroundings. Transformers have recently shown encouraging results in image segmentation. However, transformer-based models are challenging to strike a better balance between performance and efficiency. The computational complexity of the transformer-based models is quadratic with the number of inputs, which severely hinders their application in dense prediction tasks. In this paper, we present the semantic-aware dimension-pooling transformer (SDPT) to mitigate the conflict between accuracy and efficiency. The proposed model comprises an efficient transformer encoder for generating hierarchical features and a semantic-balanced decoder for predicting semantic masks. In the encoder, a dimension-pooling mechanism is used in the multi-head self-attention (MHSA) to reduce the computational cost, and a parallel depth-wise convolution is used to capture local semantics. Simultaneously, we further apply this dimension-pooling attention (DPA) to the decoder as a refinement module to integrate multi-level features. With such a simple yet powerful encoder-decoder framework, we empirically demonstrate that the proposed SDPT achieves excellent performance and efficiency on various popular benchmarks, including ADE20K, Cityscapes, and COCO-Stuff. For example, our SDPT achieves 48.6$\%$mIOU on the ADE20K dataset, which outperforms the current methods with fewer computational costs. The codes can be found at https://github.com/HuCaoFighting/SDPT. Hu Cao, Guang Chen 0001, Hengshuang Zhao, Dongsheng Jiang, Xiaopeng Zhang 0008, Qi Tian 0001, Alois C. Knoll |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Learnable Central Similarity Quantization for Efficient Image and Video RetrievalabstractData-dependent hashing methods aim to learn hash functions from the pairwise or triplet relationships among the data, which often lead to low efficiency and low collision rate by only capturing the local distribution of the data. To solve the limitation, we propose central similarity, in which the hash codes of similar data pairs are encouraged to approach a common center and those of dissimilar pairs to converge to different centers. As a new global similarity metric, central similarity can improve the efficiency and retrieval accuracy of hash learning. By introducing a new concept, hash centers, we principally formulate the computation of the proposed central similarity metric, in which the hash centers refer to a set of points scattered in the Hamming space with a sufficient mutual distance between each other. To construct well-separated hash centers, we provide two efficient methods: 1) leveraging the Hadamard matrix and Bernoulli distributions to generate data-independent hash centers and 2) learning data-dependent hash centers from data representations. Based on the proposed similarity metric and hash centers, we propose central similarity quantization (CSQ) that optimizes the central similarity between data points with respect to their hash centers instead of optimizing the local similarity to generate a high-quality deep hash function. We also further improve the CSQ with data-dependent hash centers, dubbed as CSQ with learnable center (CSQLC). The proposed CSQ and CSQLC are generic and applicable to image and video hashing scenarios. We conduct extensive experiments on large-scale image and video retrieval tasks, and the proposed CSQ yields noticeably boosted retrieval performance, i.e., 3%-20% in mean average precision (mAP) over the previous state-of-the-art methods, which also demonstrates that our methods can generate cohesive hash codes for similar data pairs and dispersed hash codes for dissimilar pairs. Li Yuan 0007, Tao Wang 0053, Xiaopeng Zhang 0008, Francis E. H. Tay, Zequn Jie, Yonghong Tian 0001, Wei Liu 0005, Jiashi Feng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | GaussianObject: High-Quality 3D Object Reconstruction from Four Views with Gaussian SplattingabstractReconstructing and rendering 3D objects from highly sparse views is of critical importance for promoting applications of 3D vision techniques and improving user experience. However, images from sparse views only contain very limited 3D information, leading to two significant challenges: 1) Difficulty in building multi-view consistency as images for matching are too few; 2) Partially omitted or highly compressed object information as view coverage is insufficient. To tackle these challenges, we propose GaussianObject, a framework to represent and render the 3D object with Gaussian splatting that achieves high rendering quality with only 4 input images. We first introduce techniques of visual hull and floater elimination, which explicitly inject structure priors into the initial optimization process to help build multi-view consistency, yielding a coarse 3D Gaussian representation. Then we construct a Gaussian repair model based on diffusion models to supplement the omitted object information, where Gaussians are further refined. We design a self-generating strategy to obtain image pairs for training the repair model. We further design a COLMAP-free variant, where pre-given accurate camera poses are not required, which achieves competitive quality and facilitates wider applications. GaussianObject is evaluated on several challenging datasets, including MipNeRF360, OmniObject3D, OpenIllumination, and our-collected unposed images, achieving superior performance from only four views and significantly outperforming previous SOTA methods. Chen Yang 0023, Sikuang Li, Jiemin Fang, Ruofan Liang, Lingxi Xie, Xiaopeng Zhang 0008, Wei Shen 0002, Qi Tian 0001 |
ACM Trans. Graph. | 6 |
| 2023 | Visual Recognition by RequestabstractHumans have the ability of recognizing visual semantics in an unlimited granularity, but existing visual recognition algorithms cannot achieve this goal. In this paper, we establish a new paradigm named visual recognition by request (ViRReq11We recommend the readers to pronounce ViRReqas/virik/.) to bridge the gap. The key lies in decomposing visual recognition into atomic tasks named requests and leveraging a knowledge base, a hierarchical and text-based dictionary, to assist task definition. ViRReq allows for (i) learning complicated whole-part hierarchies from highly incomplete annotations and (ii) inserting new concepts with minimal efforts. We also establish a solid baseline by integrating language-driven recognition into recent semantic and instance segmentation methods, and demonstrate its flexible recognition ability on CPP and ADE20K, two datasets with hierarchical whole-part annotations. Chufeng Tang, Lingxi Xie, Xiaopeng Zhang 0008, Xiaolin Hu 0001, Qi Tian 0001 |
CVPR | 3 |
| 2023 | Integrally Pre-Trained Transformer Pyramid NetworksabstractIn this paper, we present an integral pre-training framework based on masked image modeling (MIM). We advocate for pre-training the backbone and neck jointly so that the transfer gap between MIM and downstream recognition tasks is minimal. We make two technical contributions. First, we unify the reconstruction and recognition necks by inserting a feature pyramid into the pre-training stage. Second, we complement mask image modeling (MIM) with masked feature modeling (MFM) that offers multi-stage supervision to the feature pyramid. The pre-trained models, termed integrally pre-trained transformer pyramid networks (iTPNs), serve as powerful foundation models for visual recognition. In particular, the base/large-level iTPN achieves an 86.2%/87.8% top-1 accuracy on ImageNet-1K, a 53.2%/55.6% box AP on COCO object detection with 1× training schedule using Mask-RCNN, and a 54.7%/57.7% mIoU on ADE20K semantic segmentation using UPerHead – all these results set new records. Our work inspires the community to work on unifying upstream pre-training and downstream fine-tuning tasks. Code is available at github.com/sunsmarterjie/iTPN. Yunjie Tian, Lingxi Xie, Zhaozhi Wang, Longhui Wei, Xiaopeng Zhang 0008, Jianbin Jiao, Yaowei Wang 0001, Qi Tian 0001, Qixiang Ye |
CVPR | 5 |
| 2023 | Adapting Shortcut with Normalizing Flow: An Efficient Tuning Framework for Visual RecognitionabstractPretraining followed by fine-tuning has proven to be effective in visual recognition tasks. However, fine-tuning all parameters can be computationally expensive, particularly for large-scale models. To mitigate the computational and storage demands, recent research has explored Parameter-Efficient Fine-Tuning (PEFT), which focuses on tuning a minimal number of parameters for efficient adaptation. Existing methods, however, fail to analyze the impact of the additional parameters on the model, resulting in an unclear and suboptimal tuning process. In this paper, we introduce a novel and effective PEFT paradigm, named SNF (Shortcut adaptation via Normalization Flow), which utilizes normalizing flows to adjust the shortcut layers. We highlight that layers without Lipschitz constraints can lead to error propagation when adapting to downstream datasets. Since modifying the over-parameterized residual connections in these layers is expensive, we focus on adjusting the cheap yet crucial shortcuts. Moreover, learning new information with few parameters in PEFT can be challenging, and information loss can result in label information degradation. To address this issue, we propose an information-preserving normalizing flow. Experimental results demonstrate the effectiveness of SNF. Specifically, with only 0.036M parameters, SNF surpasses previous approaches on both the FGVC and VTAB-1k benchmarks using ViT/B-16 as the backbone. The code is available at https://github.com/Wang-Yaoming/SNF Bowen Shi 0003, Xiaopeng Zhang 0008, Jin Li 0057, Yuchen Liu 0006, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
CVPR | 3 |
| 2023 | Prune Spatio-temporal Tokens by Semantic-aware Temporal AccumulationabstractTransformers have become the primary backbone of the computer vision community due to their impressive performance. However, the unfriendly computation cost impedes their potential in the video recognition domain. To optimize the speed-accuracy trade-off, we propose Semanticaware Temporal Accumulation score (STA) to prune spatiotemporal tokens integrally. STA score considers two critical factors: temporal redundancy and semantic importance. The former depicts a specific region based on whether it is a new occurrence or a seen entity by aggregating token-to-token similarity in consecutive frames while the latter evaluates each token based on its contribution to the overall prediction. As a result, tokens with higher scores of STA carry more temporal redundancy as well as lower semantics thus being pruned. Based on the STA score, we are able to progressively prune the tokens without introducing any additional parameters or requiring further re-training. We directly apply the STA module to off-the-shelf ViT and VideoSwin backbones, and the empirical results on Kinetics-400 and Something-Something V2 achieve over 30% computation reduction with a negligible ~ 0.2% accuracy drop. The code is released at https://github.com/Mark12Ding/STA. Shuangrui Ding, Peisen Zhao, Xiaopeng Zhang 0008, Rui Qian 0001, Hongkai Xiong, Qi Tian 0001 |
ICCV | 3 |
| 2023 | The KFIoU Loss for Rotated Object Detection
Xue Yang 0005, Yue Zhou 0005, Gefan Zhang, Jirui Yang, Wentao Wang 0009, Junchi Yan, Xiaopeng Zhang 0008, Qi Tian 0001 |
ICLR | 7 |
| 2023 | Progressively Compressed Auto-Encoder for Self-supervised Representation Learning
Jin Li 0057, Xiaopeng Zhang 0008, Yabo Chen, Dongsheng Jiang, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
ICLR | 3 |
| 2023 | VioLET: Vision-Language Efficient Tuning with Collaborative Multi-modal GradientsabstractParameter-Efficient Tuning (PET) has emerged as a leading advancement in both Natural Language Processing and Computer Vision, enabling efficient accommodation of downstream tasks without costly fine-tuning. However, most existing PET approaches are limited to uni-modal tuning, even for vision-language models like CLIP. We investigate this limitation and demonstrate that simultaneous tuning of the two modalities in such models leads to multi-modal forgetting and catastrophic performance degradation, particularly when generalizing to new classes. To address this issue, we propose a novel PET approach called VioLET (Vision Language Efficient Tuning) that utilizes collaborative multi-modal gradients to unlock the full potential of both modalities. Specifically, we incorporate an additional visual encoder without learnable parameters and use these two visual encoders to compute the gradients of the context parameters separately. When conflicts arise, we replace the original gradient with an orthogonal gradient. Extensive experiments are conducted on few-shot recognition and unseen class generalization tasks using ResNet-50 or ViT/B-16 as the backbone. VioLET consistently outperforms several state-of-the-art methods on 11 datasets, showcasing its superiority over existing PET approaches. The code is available at https://github.com/Wang-Yaoming/VioLET. Yuchen Liu 0006, Xiaopeng Zhang 0008, Jin Li 0057, Bowen Shi 0003, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
ACM Multimedia | 3 |
| 2023 | Segment Anything in 3D with NeRFsabstractRecently, the Segment Anything Model (SAM) emerged as a powerful vision foundation model which is capable to segment anything in 2D images. This paper aims to generalize SAM to segment 3D objects. Rather than replicating the data acquisition and annotation procedure which is costly in 3D, we design an efficient solution, leveraging the Neural Radiance Field (NeRF) as a cheap and off-the-shelf prior that connects multi-view 2D images to the 3D space. We refer to the proposed solution as SA3D, for Segment Anything in 3D. It is only required to provide a manual segmentation prompt (e.g., rough points) for the target object in a single view, which is used to generate its 2D mask in this view with SAM. Next, SA3D alternately performs mask inverse rendering and cross-view self-prompting across various views to iteratively complete the 3D mask of the target object constructed with voxel grids. The former projects the 2D mask obtained by SAM in the current view onto 3D mask with guidance of the density distribution learned by the NeRF; The latter extracts reliable prompts automatically as the input to SAM from the NeRF-rendered 2D mask in another view. We show in experiments that SA3D adapts to various scenes and achieves 3D segmentation within minutes. Our research offers a generic and efficient methodology to lift a 2D vision foundation model to 3D, as long as the 2D model can steadily address promptable segmentation across multiple views. Jiazhong Cen, Zanwei Zhou, Jiemin Fang, Chen Yang 0023, Wei Shen 0002, Lingxi Xie, Dongsheng Jiang, Xiaopeng Zhang 0008, Qi Tian 0001 |
NeurIPS | 8 |
| 2023 | AiluRus: A Scalable ViT Framework for Dense PredictionabstractVision transformers (ViTs) have emerged as a prevalent architecture for vision tasks owing to their impressive performance. However, their complexity dramatically increases when handling long token sequences, particularly for dense prediction tasks that require high-resolution input. Notably, dense prediction tasks, such as semantic segmentation or object detection, emphasize more on the contours or shapes of objects, while the texture inside objects is less informative. Motivated by this observation, we propose to apply adaptive resolution for different regions in the image according to their importance. Specifically, at the intermediate layer of the ViT, we select anchors from the token sequence using the proposed spatial-aware density-based clustering algorithm. Tokens that are adjacent to anchors are merged to form low-resolution regions, while others are preserved independently as high-resolution. This strategy could significantly reduce the number of tokens, and the following layers only handle the reduced token sequence for acceleration. At the output end, the resolution of the feature map is recovered by unfolding merged tokens for task prediction. Consequently, we can considerably accelerate ViTs for dense prediction tasks. The proposed method is evaluated across three different datasets and demonstrates promising performance. For instance, "Segmenter ViT-L" can be accelerated by 48\% FPS without fine-tuning, while maintaining the performance. Moreover, our method can also be applied to accelerate fine-tuning. Experiments indicate that we can save 52\% training time while accelerating 2.46$\times$ FPS with only a 0.09\% performance drop. Jin Li 0057, Xiaopeng Zhang 0008, Bowen Shi 0003, Dongsheng Jiang, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
NeurIPS | 3 |
| 2023 | GAIA-Universe: Everything is Super-NetifyabstractPre-training on large-scale datasets has played an increasingly significant role in computer vision and natural language processing recently. However, as there exist numerous application scenarios that have distinctive demands such as certain latency constraints and specialized data distributions, it is prohibitively expensive to take advantage of large-scale pre-training for per-task requirements. we focus on two fundamental perception tasks (object detection and semantic segmentation) and present a complete and flexible system named GAIA-Universe(GAIA), which could automatically and efficiently give birth to customized solutions according to heterogeneous downstream needs through data union and super-net training. GAIA is capable of providing powerful pre-trained weights and searching models that conform to downstream demands such as hardware constraints, computation constraints, specified data domains, and telling relevant data for practitioners who have very few datapoints on their tasks. With GAIA, we achieve promising results on COCO, Objects365, Open Images, BDD100 k, and UODB which is a collection of datasets including KITTI, VOC, WiderFace, DOTA, Clipart, Comic, and more. Taking COCO as an example, GAIA is able to efficiently produce models covering a wide range of latency from 16 ms to 53 ms, and yields AP from 38.2 to 46.5 without whistles and bells. GAIA is released at https://github.com/GAIA-vision. Junran Peng, Xingyuan Bu, Lingxi Xie, Xiaopeng Zhang 0008, Qi Tian 0001, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Seed the Views: Hierarchical Semantic Alignment for Contrastive Representation LearningabstractSelf-supervised learning based on instance discrimination has shown remarkable progress. In particular, contrastive learning, which regards each image as well as its augmentations as an individual class and tries to distinguish them from all other images, has been verified effective for representation learning. However, conventional contrastive learning does not model the relation between semantically similar samples explicitly. In this paper, we propose a general module that considers the semantic similarity among images. This is achieved by expanding the views generated by a single image to Cross-Samples and Multi-Levels, and modeling the invariance to semantically similar images in a hierarchical way. Specifically, the cross-samples are generated by a data mixing operation, which is constrained within samples that are semantically similar, while the multi-level samples are expanded at the intermediate layers of a network. In this way, the contrastive loss is extended to allow for multiple positives per anchor, and explicitly pulling semantically similar images together at different layers of the network. Our method, termed as CSML, has the ability to integrate multi-level representations across samples in a robust way. CSML is applicable to current contrastive based methods and consistently improves the performance. Notably, using MoCo v2 as an instantiation, CSML achieves 76.6% top-1 accuracy with linear evaluation using ResNet-50 as backbone, 66.7% and 75.1% top-1 accuracy with only 1% and 10% labels, respectively. All these numbers set the new state-of-the-art. The code is available at https://github.com/haohang96/CSML. Haohang Xu, Xiaopeng Zhang 0008, Hao Li 0090, Lingxi Xie, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Semi-Supervised Contrastive Learning With Similarity Co-CalibrationabstractSemi-supervised learning acts as an effective way to leverage massive unlabeled data. In this paper, we propose a novel training strategy, termed asSemi-supervised Contrastive Learning (SsCL), which combines the well-known contrastive loss in self-supervised learning with the cross entropy loss in semi-supervised learning, and jointly optimizes the two objectives in an end-to-end way. The highlight is that different from self-training based semi-supervised learning that conducts prediction and retraining over the same model weights, SsCL interchanges the predictions over the unlabeled data between the two branches, and thus formulates a co-calibration procedure, which we find is beneficial for better prediction and avoids being trapped in local minimum. Towards this goal, the contrastive loss branch models pairwise similarities among samples, using the pseudo labels generated from the cross entropy branch, and in turn calibrates the prediction distribution of the cross entropy branch with the contrastive similarity. We show that SsCL produces more discriminative representation and is beneficial to semi-supervised learning. Notably, on ImageNet with ResNet50 as the backbone, SsCL achieves$\bm {60.2\%}$and$\bm {72.1\%}$top-1 accuracy with 1% and 10% labeled samples respectively, which significantly outperforms the baseline, and is better than previous semi-supervised and self-supervised methods. Yuhang Zhang 0012, Xiaopeng Zhang 0008, Jie Li 0002, Robert C. Qiu, Haohang Xu, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Can Semantic Labels Assist Self-Supervised Visual Representation Learning?abstractRecently, contrastive learning has largely advanced the progress of unsupervised visual representation learning. Pre-trained on ImageNet, some self-supervised algorithms reported higher transfer learning performance compared to fully-supervised methods, seeming to deliver the message that human labels hardly contribute to learning transferrable visual features. In this paper, we defend the usefulness of semantic labels but point out that fully-supervised and self-supervised methods are pursuing different kinds of features. To alleviate this issue, we present a new algorithm named Supervised Contrastive Adjustment in Neighborhood (SCAN) that maximally prevents the semantic guidance from damaging the appearance feature embedding. In a series of downstream tasks, SCAN achieves superior performance compared to previous fully-supervised and self-supervised methods, and sometimes the gain is significant. More importantly, our study reveals that semantic labels are useful in assisting self-supervised methods, opening a new direction for the community. Longhui Wei, Lingxi Xie, Xiaopeng Zhang 0008, Qi Tian 0001 |
AAAI | 4 |
| 2022 | MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger TokensabstractTransformers have offered a new methodology of designing neural networks for visual recognition. Compared to convolutional networks, Transformers enjoy the ability of referring to global features at each stage, yet the attention module brings higher computational overhead that obstructs the application of Transformers to process highresolution visual data. This paper aims to alleviate the conflict between efficiency and flexibility, for which we propose a specialized token for each region that serves as a messenger (MSG). Hence, by manipulating these MSG tokens, one can flexibly exchange visual information across regions and the computational complexity is reduced. We then integrate the MSG token into a multi-scale architecture named MSG-Transformer. In standard image classification and object detection, MSG-Transformer achieves competitive performance and the inference on both GPU and CPU is accelerated. Code is available at https://github.com/hustvl/MSG-Transformer. Jiemin Fang, Lingxi Xie, Xinggang Wang, Xiaopeng Zhang 0008, Wenyu Liu 0001, Qi Tian 0001 |
CVPR | 4 |
| 2022 | One-bit Active Query with Contrastive PairsabstractHow to achieve better results with fewer labeling costs remains a challenging task. In this paper, we present a new active learning framework, which for the first time incorporates contrastive learning into recently proposed one-bit supervision. Here one-bit supervision denotes a simple Yes or No query about the correctness of the model's prediction, and is more efficient than previous active learning methods requiring assigning accurate labels to the queried samples. We claim that such one-bit information is intrinsically in accordance with the goal of contrastive loss that pulls positive pairs together and pushes negative samples away. Towards this goal, we design an uncertainty metric to actively select samples for query. These samples are then fed into different branches according to the queried results. The Yes query is treated as positive pairs of the queried category for contrastive pulling, while the No query is treated as hard negative pairs for contrastive repelling. Additionally, we design a negative loss that penalizes the negative samples away from the incorrect predicted class, which can be treated as optimizing hard negatives for the corresponding category. Our method, termed as ObCP, produces a more powerful active learning framework, and experiments on several benchmarks demonstrate its superiority. Yuhang Zhang 0012, Xiaopeng Zhang 0008, Lingxi Xie, Jie Li 0002, Robert C. Qiu, Hengtong Hu, Qi Tian 0001 |
CVPR | 2 |
| 2022 | SdAE: Self-distillated Masked Autoencoder
Yabo Chen, Yuchen Liu 0006, Dongsheng Jiang, Xiaopeng Zhang 0008, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
ECCV (30) | 4 |
| 2022 | TAPE: Task-Agnostic Prior Embedding for Image Restoration
Lin Liu 0016, Lingxi Xie, Xiaopeng Zhang 0008, Shanxin Yuan, Xiangyu Chen 0006, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001 |
ECCV (18) | 3 |
| 2022 | A Transformer-Based Decoder for Semantic Segmentation with Multi-level Context Mining
Bowen Shi 0003, Dongsheng Jiang, Xiaopeng Zhang 0008, Wenrui Dai, Junni Zou, Hongkai Xiong, Qi Tian 0001 |
ECCV (28) | 3 |
| 2022 | Active Pointly-Supervised Instance Segmentation
Chufeng Tang, Lingxi Xie, Xiaopeng Zhang 0008, Qi Tian 0001, Xiaolin Hu 0001 |
ECCV (28) | 4 |
| 2022 | Bag of Instances Aggregation Boosts Self-supervised Distillation
Haohang Xu, Jiemin Fang, Xiaopeng Zhang 0008, Lingxi Xie, Xinggang Wang, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
ICLR | 3 |
| 2022 | Fast Dynamic Radiance Fields with Time-Aware Neural VoxelsabstractNeural radiance fields (NeRF) have shown great success in modeling 3D scenes and synthesizing novel-view images. However, most previous NeRF methods take much time to optimize one single scene. Explicit data structures, e.g. voxel features, show great potential to accelerate the training process. However, voxel features face two big challenges to be applied to dynamic scenes, i.e. modeling temporal information and capturing different scales of point motions. We propose a radiance field framework by representing scenes with time-aware voxel features, named as TiNeuVox. A tiny coordinate deformation network is introduced to model coarse motion trajectories and temporal information is further enhanced in the radiance network. A multi-distance interpolation method is proposed and applied on voxel features to model both small and large motions. Our framework significantly accelerates the optimization of dynamic radiance fields while maintaining high rendering quality. Empirical evaluation is performed on both synthetic and real scenes. Our TiNeuVox completes training with only 8 minutes and 8-MB storage cost while showing similar or even better rendering performance than previous dynamic NeRF methods. Code is available at https://github.com/hustvl/TiNeuVox. Jiemin Fang, Taoran Yi, Xinggang Wang, Lingxi Xie, Xiaopeng Zhang 0008, Wenyu Liu 0001, Matthias Nießner, Qi Tian 0001 |
SIGGRAPH Asia | 5 |
| 2022 | Weakly Supervised Region-Level Contrastive Learning for Efficient Object DetectionabstractSemi-supervised learning, which assigns pseudo labels with models trained using limited labeled data, has been widely used in object detection to reduce the labeling cost. However, the provided pseudo annotations inevitably suffer noise since the initial model is not perfect. To address this issue, this paper introduces contrastive learning into semi-supervised object detection, and we claim that contrastive loss, which inherently relies on data augmentations, is much more robust than traditional softmax regression for noisy labels. To take full advantage of it in the detection task, we incorporate labels prior to contrastive loss and leverage plenty of region proposals to enhance diversity, which is crucial for contrastive learning. In this way, the model is optimized to make the region-level features with the same class be translation and scale invariant. Furthermore, we redesign the negative memory bank in contrastive learning to make the training more efficient. As far as we know, we are the first attempt that introduces contrastive learning in semi-supervised object detection. Experimental results on detection benchmarks demonstrate the superiority of our method. Notably, our method achieves 79.9% accuracy on VOC, which is 6.2% better than the supervised baseline and 0.7% improvement compared with the state-of-the-art method. Yuang Deng, Yuhang Zhang 0012, Wenrui Dai, Xiaopeng Zhang 0008, Hongkai Xiong |
VCIP | 4 |
| 2022 | Feature Calibration Network for Occluded Pedestrian DetectionabstractPedestrian detection in the wild remains a challenging problem especially for scenes containing serious occlusion. In this paper, we propose a novel feature learning method in the deep learning framework, referred to as Feature Calibration Network (FC-Net), to adaptively detect pedestrians under various occlusions. FC-Net is based on the observation that the visible parts of pedestrians are selective and decisive for detection, and is implemented as a self-paced feature learning framework with a self-activation (SA) module and a feature calibration (FC) module. In a new self-activated manner, FC-Net learns features which highlight the visible parts and suppress the occluded parts of pedestrians. The SA module estimates pedestrian activation maps by reusing classifier weights, without any additional parameter involved, therefore resulting in an extremely parsimony model to reinforce the semantics of features, while the FC module calibrates the convolutional features for adaptive pedestrian representation in both pixel-wise and region-based ways. Experiments on CityPersons and Caltech datasets demonstrate that FC-Net improves detection performance on occluded pedestrians up to 10% while maintaining excellent performance on non-occluded instances. Tianliang Zhang 0003, Qixiang Ye, Baochang Zhang 0001, Jianzhuang Liu, Xiaopeng Zhang 0008, Qi Tian 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Answer Again: Improving VQA With Cascaded-Answering ModelabstractVisual Question Answering (VQA) is a very challenging task, which requires to understand visual images and natural language questions simultaneously. In the open-ended VQA task, most previous solutions focus on understanding the question and image contents, as well as their correlations. However, they mostly reason the answers in a one-stage way, which results in that the generated answers are significantly ignored. In this paper, we propose a novel approach, termed Cascaded-Answering Model (CAM), which extends the conventional one-stage VQA model to a two-stage model. Hence, the proposed model can fully explore the semantics embedded in the predicted answers. Specifically, CAM is composed of two cascaded answering modules: Candidate Answer Generation (CAG) module and Final Answer Prediction (FAP) module. In CAG module, we select multiple relevant candidates from the generated answers using a typical VQA approach with Co-Attention. While in FAP module, we integrate the information of question and image, together with the semantics explored from the selected candidate answers to predict the final answer. Experimental results demonstrate that the proposed model produces high-quality candidate answers and achieves the state-of-the-art performance on five large benchmark datasets, VQA-1.0, VQA-2.0, VQA-CP v2, TDIUC and COCO-QA. Yang Yang 0002, Xiaopeng Zhang 0008, Yanli Ji, Huimin Lu 0001, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Heterogeneous Contrastive Learning: Encoding Spatial Information for Compact Visual RepresentationsabstractUnsupervised pretraining is of great significance for visual representation. Especially, contrastive learning has achieved great success recently, but existing approaches mostly ignored spatial information which is often crucial for visual representation. Strong semantic embedding has an inherent advantage for classification, but dense prediction tasks require more spatial and low-level representation. This paper presentsheterogeneous contrastive learning(HCL), an effective approach that adds spatial information to the encoding stage to alleviate the learning inconsistency between the contrastive objective and strong data augmentation operations. We demonstrate the effectiveness of HCL by showing that (i) it achieves higher accuracy in instance discrimination, (ii) it surpasses existing pre-training methods in a series of downstream tasks (iii) and it shrinks the pre-training costs by half for almost 800 GPU-hours. More importantly, we show that our approach achieves higher efficiency in visual representations, and thus delivers a key message to inspire the future research of self-supervised visual representation learning. Xinyue Huo, Lingxi Xie, Longhui Wei, Xiaopeng Zhang 0008, Xin Chen 0033, Hao Li 0090, Zijie Yang, Wengang Zhou 0001, Houqiang Li, Qi Tian 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | Center-wise Local Image Mixture For Contrastive Representation Learning
Hao Li 0090, Xiaopeng Zhang 0008, Hongkai Xiong |
BMVC | 2 |
| 2021 | Rethinking Rotated Object Detection with Gaussian Wasserstein Distance LossabstractBoundary discontinuity and its inconsistency to the final detection metric have been the bottleneck for rotating detection regression loss design. In this paper, we propose a novel regression loss based on Gaussian Wasserstein distance as a fundamental approach to solve the problem. Specifically, the rotated bounding box is converted to a 2-D Gaussian distribution, which enables to approximate the indifferentiable rotational IoU induced loss by the Gaussian Wasserstein distance (GWD) which can be learned efficiently by gradient back-propagation. GWD can still be informative for learning even there is no overlapping between two rotating bounding boxes which is often the case for small object detection. Thanks to its three unique properties, GWD can also elegantly solve the boundary discontinuity and square-like problem regardless how the bounding box is defined. Experiments on five datasets using different detectors show the effectiveness of our approach, and codes are available at https://github.com/yangxue0827/RotationDetection. Xue Yang 0005, Junchi Yan, Qi Ming, Wentao Wang 0009, Xiaopeng Zhang 0008, Qi Tian 0001 |
ICML | 5 |
| 2021 | Partially-Connected Neural Architecture Search for Reduced Computational RedundancyabstractDifferentiable architecture search (DARTS) enables effective neural architecture search (NAS) using gradient descent, but suffers from high memory and computational costs. In this paper, we propose a novel approach, namely Partially-Connected DARTS (PC-DARTS), to achieve efficient and stable neural architecture search by reducing the channel and spatial redundancies of the super-network. In the channel level, partial channel connection is presented to randomly sample a small subset of channels for operation selection to accelerate the search process and suppress the over-fitting of the super-network. Side operation is introduced for bypassing (non-sampled) channels to guarantee the performance of searched architectures under extremely low sampling rates. In the spatial level, input features are down-sampled to eliminate spatial redundancy and enhance the efficiency of the mixed computation for operation selection. Furthermore, edge normalization is developed to maintain the consistency of edge selection based on channel sampling with the architectural parameters for edges. Theoretical analysis shows that partial channel connection and parameterized side operation are equivalent to regularizing the super-network on the weights and architectural parameters during bilevel optimization. Experimental results demonstrate that the proposed approach achieves higher search speed and training stability than DARTS. PC-DARTS obtains a top-1 error rate of 2.55 percent on CIFAR-10 with 0.07 GPU-days for architecture search, and a state-of-the-art top-1 error rate of 24.1 percent on ImageNet (under the mobile setting) within 2.8 GPU-days. Yuhui Xu 0002, Lingxi Xie, Wenrui Dai, Xiaopeng Zhang 0008, Xin Chen 0033, Guo-Jun Qi, Hongkai Xiong, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Central Similarity Quantization for Efficient Image and Video RetrievalabstractExisting data-dependent hashing methods usually learn hash functions from pairwise or triplet data relationships, which only capture the data similarity locally, and often suffer from low learning efficiency and low collision rate. In this work, we propose a new \emph{global} similarity metric, termed as \emph{central similarity}, with which the hash codes of similar data pairs are encouraged to approach a common center and those for dissimilar pairs to converge to different centers, to improve hash learning efficiency and retrieval accuracy. We principally formulate the computation of the proposed central similarity metric by introducing a new concept, i.e., \emph{hash center} that refers to a set of data points scattered in the Hamming space with a sufficient mutual distance between each other. We then provide an efficient method to construct well separated hash centers by leveraging the Hadamard matrix and Bernoulli distributions. Finally, we propose the Central Similarity Quantization (CSQ) that optimizes the central similarity between data points w.r.t.\ their hash centers instead of optimizing the local similarity. CSQ is generic and applicable to both image and video hashing scenarios. Extensive experiments on large-scale image and video retrieval tasks demonstrate that CSQ can generate cohesive hash codes for similar data pairs and dispersed hash codes for dissimilar pairs, achieving a noticeable boost in retrieval performance, i.e. 3\%-20\% in mAP over the previous state-of-the-arts. Li Yuan 0007, Tao Wang 0053, Xiaopeng Zhang 0008, Francis E. H. Tay, Zequn Jie, Wei Liu 0005, Jiashi Feng |
CVPR | 3 |
| 2020 | Circumventing Outliers of AutoAugment with Knowledge Distillation
Longhui Wei, An Xiao, Lingxi Xie, Xiaopeng Zhang 0008, Xin Chen 0033, Qi Tian 0001 |
ECCV (3) | 4 |
| 2020 | PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search
Yuhui Xu 0002, Lingxi Xie, Xiaopeng Zhang 0008, Xin Chen 0033, Guo-Jun Qi, Qi Tian 0001, Hongkai Xiong |
ICLR | 3 |
| 2020 | Attribute Mix: Semantic Data Augmentation for Fine Grained RecognitionabstractCollecting fine-grained labels usually requires expert-level domain knowledge and is prohibitive to scale up. In this paper, we propose Attribute Mix, a data augmentation strategy at attribute level to expand the fine-grained samples. The principle lies in that attribute features are shared among fine-grained sub-categories, and can be seamlessly transferred among images. Toward this goal, we propose an automatic attribute mining approach to discover attributes that belong to the same supercategory, and Attribute Mix is operated by mixing semantically meaningful attribute features from two images. Attribute Mix is a simple but effective data augmentation strategy that can significantly improve the recognition performance without increasing the inference budgets. Extensive experiments and ablation studies show that the proposed method consistently outperforms the state-of-the-art methods on challenging benchmarks including CUB-200-2011, FGVC-Aircraft, and Stanford Cars. Hao Li 0090, Xiaopeng Zhang 0008, Qi Tian 0001, Hongkai Xiong |
VCIP | 2 |
| 2020 | Depth Estimation From Light Field Using Graph-Based Structure-Aware AnalysisabstractExisting light field depth map estimation approaches only utilize partial angular views in occlusion areas and local spatial dependencies in the optimization. This paper proposes a novel two-stage light field depth estimation method via graph spectral analysis to exploit the complete correlations and dependencies within angular patches and spatial images. The initial depth map estimation leverages the undirected graph to jointly consider occluded and unoccluded views within each angular patch. The estimated depth minimizes the structural incoherence of its corresponding angular patch with the focused one by evaluating the highest graph frequency component. Subsequently, depth map refinement optimizes the initial depth map with the color consistency and smoothness formulated by weighted adjacency matrix. The structural constraints are efficiently employed using low-pass graph filtering with Chebyshev polynomial approximation. Experimental results demonstrate that the proposed method improves the depth map estimation, especially in the edge regions. Wenrui Dai, Mingxing Xu, Junni Zou, Xiaopeng Zhang 0008, Hongkai Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | PML-LocNet: Improving Object Localization With Prior-Induced Multi-View Learning NetworkabstractThis paper introduces a new model for Weakly Supervised Object Localization (WSOL) problems where only image-level supervision is provided. The key to solve such problems is to infer the object locations accurately. Previous methods usually model the missing object locations as latent variables, and alternate between updating their estimates and learning a detector accordingly. However, the performance of such alternative optimization is sensitive to the quality of the initial latent variables and the resulted localization model is prone to overfitting to improper localizations. To address these issues, we develop a Prior-induced Multi-view Learning Localization Network (PML-LocNet) which exploits both view diversity and sample diversity to improve object localization. In particular, the view diversity is imposed by a two-phase multi-view learning strategy, with which the complementarity among learned features from different views and the consensus among localized instances from each view are leveraged to benefit localization. The sample diversity is pursued by harnessing coarse-to-fine priors at both image and instance levels. With these priors, more emphasis would go to the reliable samples and the contributions of the unreliable ones would be decreased, such that the intrinsic characteristics of each sample can be exploited to make the model more robust during network learning. PML-LocNet can be easily combined with existing WSOL models to further improve the localization accuracy. Its effectiveness has been proved experimentally. Notably, it achieves 69.3% CorLoc and 50.4% mAP on PASCAL VOC 2007, surpassing the state-of-the-arts by a large margin. Xiaopeng Zhang 0008, Yang Yang 0002, Hongkai Xiong, Jiashi Feng |
IEEE Trans. Image Process. | 1 |
| 2019 | Learning to Localize Objects with Noisy Labeled InstancesabstractThis paper addresses Weakly Supervised Object Localization (WSOL) with only image-level supervision. We model the missing object locations as latent variables, and contribute a novel self-directed optimization strategy to infer them. With the strategy, our developed Self-Directed Localization Network (SD-LocNet) is able to localize object instance whose initial location is noisy. The self-directed inference hinges on an adaptive sampling method to identify reliable object instance via measuring its localization stability score. In this way, the resulted model is robust to noisy initialized object locations which we find is important in WSOL. Furthermore, we introduce a reliability induced prior propagation strategy to transfer object priors of the reliable instances to those unreliable ones by promoting their feature similarity, which effectively refines the unreliable object instances for better localization. The proposed SD-LocNet achieves 70.9% Cor-Loc and 51.3% mAP on PASCAL VOC 2007, surpassing the state-of-the-arts by a large margin. Xiaopeng Zhang 0008, Yang Yang 0002, Jiashi Feng |
AAAI | 1 |
| 2019 | Distilling Object Detectors With Fine-Grained Feature ImitationabstractState-of-the-art CNN based recognition models are often computationally prohibitive to deploy on low-end devices. A promising high level approach tackling this limitation is knowledge distillation, which let small student model mimic cumbersome teacher model's output to get improved generalization. However, related methods mainly focus on simple task of classification while do not consider complex tasks like object detection. We show applying the vanilla knowledge distillation to detection model gets minor gain. To address the challenge of distilling knowledge in detection model, we propose a fine-grained feature imitation method exploiting the cross-location discrepancy of feature response. Our intuition is that detectors care more about local near object regions. Thus the discrepancy of feature response on the near object anchor locations reveals important information of how teacher model tends to generalize. We design a novel mechanism to estimate those locations and let student model imitate the teacher on them to get enhanced performance. We first validate the idea on a developed lightweight toy detector which carries simplest notion of current state-of-the-art anchor based detection models on challenging KITTI dataset, our method generates up to 15% boost of mAP for the student model compared to the non-imitated counterpart. We then extensively evaluate the method with Faster R-CNN model under various scenarios with common object detection benchmark of Pascal VOC and COCO, imitation alleviates up to 74% performance drop of student model compared to teacher. Codes released at https://github.com/twangnh/Distilling-Object-Detectors. Tao Wang 0053, Li Yuan 0007, Xiaopeng Zhang 0008, Jiashi Feng |
CVPR | 3 |
| 2019 | Few-Shot Adaptive Faster R-CNNabstractTo mitigate the detection performance drop caused by domain shift, we aim to develop a novel few-shot adaptation approach that requires only a few target domain images with limited bounding box annotations. To this end, we first observe several significant challenges. First, the target domain data is highly insufficient, making most existing domain adaptation methods ineffective. Second, object detection involves simultaneous localization and classification, further complicating the model adaptation process. Third, the model suffers from over-adaptation (similar to overfitting when training with a few data example) and instability risk that may lead to degraded detection performance in the target domain. To address these challenges, we first introduce a pairing mechanism over source and target features to alleviate the issue of insufficient target domain samples. We then propose a bi-level module to adapt the source trained detector to the target domain: 1) the split pooling based image level adaptation module uniformly extracts and aligns paired local patch features over locations, with different scale and aspect ratio; 2) the instance level adaptation module semantically aligns paired object features while avoids inter-class confusion. Meanwhile, a source model feature regularization (SMFR) is applied to stabilize the adaptation process of the two modules. Combining these contributions gives a novel few-shot adaptive Faster-RCNN framework, termed FAFRCNN, which effectively adapts to target domain with a few labeled samples. Experiments with multiple datasets show that our model achieves new state-of-the-art performance under both the interested few-shot domain adaptation(FDA) and unsupervised domain adaptation(UDA) setting. Tao Wang 0053, Xiaopeng Zhang 0008, Li Yuan 0007, Jiashi Feng |
CVPR | 2 |
| 2019 | Weak to Strong Detector Learning for Simultaneous Classification and LocalizationabstractThis paper aims at learning discriminative part detectors with only image-level labels. To this end, we need to develop effective technologies for both pattern mining and detection learning. Different from previous methods, which train part detectors in one step, we divide the detector learning process into two stages and formulate it as a weak to strong learning framework. In particular, we first learn exemplar detectors from the unaligned patterns and perform a detector-based spectral clustering to produce weak detectors that are only responsible for a few discriminative patterns. In this way, the weak detectors are able to offer right initial patterns for strong detector learning. Second, we learn strong detectors with patterns discovered from the weak detectors, which we formulate as a confidence-loss sparse multiple instance learning (cls-MIL) task. The cls-MIL considers the diversity of positive samples while avoiding drifting away from the well localized ones by assigning a confidence value to each positive sample. The responses of the learned detectors produce an effective mid-level image representation for both image classification and object localization. Experiments conducted on benchmark data sets well demonstrate the superiority of our method over existing approaches. Xiaopeng Zhang 0008, Hongkai Xiong, Weiyao Lin, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Zigzag Learning for Weakly Supervised Object DetectionabstractThis paper addresses weakly supervised object detection with only image-level supervision at training stage. Previous approaches train detection models with entire images all at once, making the models prone to being trapped in sub-optimums due to the introduced false positive examples. Unlike them, we propose a zigzag learning strategy to simultaneously discover reliable object instances and prevent the model from overfitting initial seeds. Towards this goal, we first develop a criterion named mean Energy Accumulation Scores (mEAS) to automatically measure and rank localization difficulty of an image containing the target object, and accordingly learn the detector progressively by feeding examples with increasing difficulty. In this way, the model can be well prepared by training on easy examples for learning from more difficult ones and thus gain a stronger detection ability more efficiently. Furthermore, we introduce a novel masking regularization strategy over the high level convolutional feature maps to avoid overfitting initial samples. These two modules formulate a zigzag learning process, where progressive learning endeavors to discover reliable object instances, and masking regularization increases the difficulty of finding object instances properly. We achieve 47.6% mAP on PASCAL VOC 2007, surpassing the state-of-the-arts by a large margin. Xiaopeng Zhang 0008, Jiashi Feng, Hongkai Xiong, Qi Tian 0001 |
CVPR | 1 |
| 2018 | ML-LocNet: Improving Object Localization with Multi-view Learning Network
Xiaopeng Zhang 0008, Yang Yang 0002, Jiashi Feng |
ECCV (3) | 1 |
| 2017 | Webly-supervised visual concept learning with cardinality guided instance mining and clustered multitask refinementabstractConventional image classification and object detection methods depend on manual annotations, such as image-level labels and bounding boxes. However, the acquisition of such annotations for millions of images is trivial. This paper addresses the problem of webly-supervised visual concept learning, and develops an automatic algorithm using parallel text and visual corpora to discover informative visual patterns from the web images. Based on the mined patterns, a cardinality-guided multiple instance learning algorithm is designed to establish the link between the image patterns and the literal concepts. Furthermore, due to the diversity of visual concepts, we perform clustered multitask refinement on the learned concept classifiers to enhance their generalization capability via a clustered regularization. Experiments demonstrate the superiority of the proposed method over traditional approaches. Saijie Ni, Xiaopeng Zhang 0008, Hongkai Xiong |
ICME | 2 |
| 2017 | Picking Neural Activations for Fine-Grained RecognitionabstractIt is a challenging task to recognize fine-grained subcategories due to the highly localized and subtle differences among them. Different from most previous methods that rely on object/part annotations, this paper proposes an automatic fine-grained recognition approach, which is free of any object/part annotation at both training and testing stages. The key idea includes two steps of picking neural activations computed from the convolutional neural networks, one for localization, and the other for description. The first picking step is to find distinctive neurons that are sensitive to specific patterns significantly and consistently. Based on these picked neurons, we initialize positive samples and formulate the localization as a regularized multiple instance learning task, which aims at refining the detectors via iteratively alternating between new positive sample mining and part model retraining. The second picking step is to pool deep neural activations via a spatially weighted combination of Fisher Vectors coding. We conditionally select activations to encode them into the final representation, which considers the importance of each activation. Integrating the above techniques produces a powerful framework, and experiments conducted on several extensive fine-grained benchmarks demonstrate the superiority of our proposed algorithm over the existing methods. Xiaopeng Zhang 0008, Hongkai Xiong, Wengang Zhou 0001, Weiyao Lin, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |
| 2016 | Picking Deep Filter Responses for Fine-Grained Image RecognitionabstractRecognizing fine-grained sub-categories such as birds and dogs is extremely challenging due to the highly localized and subtle differences in some specific parts. Most previous works rely on object / part level annotations to build part-based representation, which is demanding in practical applications. This paper proposes an automatic fine-grained recognition approach which is free of any object / part annotation at both training and testing stages. Our method explores a unified framework based on two steps of deep filter response picking. The first picking step is to find distinctive filters which respond to specific patterns significantly and consistently, and learn a set of part detectors via iteratively alternating between new positive sample mining and part model retraining. The second picking step is to pool deep filter responses via spatially weighted combination of Fisher Vectors. We conditionally pick deep filter responses to encode them into the final representation, which considers the importance of filter responses themselves. Integrating all these techniques produces a much more powerful framework, and experiments conducted on CUB-200-2011 and Stanford Dogs demonstrate the superiority of our proposed algorithm over the existing methods. Xiaopeng Zhang 0008, Hongkai Xiong, Wengang Zhou 0001, Weiyao Lin, Qi Tian 0001 |
CVPR | 1 |
| 2016 | Fused One-vs-All Features With Semantic Alignments for Fine-Grained Visual CategorizationabstractFine-grained visual categorization is an emerging research area and has been attracting growing attention recently. Due to the large inter-class similarity and intra-class variance, it is extremely challenging to recognize objects in fine-grained domains. A traditional spatial pyramid matching model could obtain desirable results for the basic-level category classification by weak alignment, but may easily fail in fine-grained domains, since the discriminative features are extremely localized. This paper proposes a new framework for fine-grained visual categorization. First, an efficient part localization method incorporates semantic prior into geometric alignment. It detects the less deformable parts, such as the head of birds with a template-based model, and localizes other highly deformable parts with simple geometric alignment. Second, we learn one-vs-all features, which are simple and transplantable. The learned mid-level features are dimension friendly and more robust to outlier instances. Furthermore, in view that some subcategories are too similar to tell them apart easily, we fuse the subcategories iteratively according to their similarities, and learn fused one-vs-all features. Experimental results show the superior performance of our algorithms over the existing methods. Xiaopeng Zhang 0008, Hongkai Xiong, Wengang Zhou 0001, Qi Tian 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Multiscale Edge Coding and Adaptive Lifting for Depth Maps Coding in 3-D VideoabstractA depth image represents three-dimensional scene information and is commonly used for depth image based rendering (DIBR) to support 3-D video applications. Reconstruction errors around depth edges lead to image blurring in synthesized views. This paper proposes a new coding method to improve the depth edge preserving effect. It makes use of both edge regularity and grayscale regularity to model the depth map, where the edge regularity is modeled by multiscale beamlet representation and the grayscale regularity is modeled by edge adaptive wavelet transform. Xiaopeng Zhang 0008, Hongkai Xiong |
DCC | 1 |
| 2014 | Fused one-vs-all mid-level features for fine-grained visual categorizationabstractAs an emerging research topic, fine-grained visual categorization has been attracting growing attentions in recent years. Due to the large inter-class similarity and intra-class variance, recognizing objects in fine-grained domains is extremely challenging, and sometimes even humans can not recognize them accurately. Traditional bag-of-words model could obtain desirable results for basic-level category classification by weak alignment using spatial pyramid matching model, but may easily fail in fine-grained domains since the discriminative features are not only subtle but also extremely localized. The fine differences often get swamped by those irrelevant features, and it is virtually impossible to distinguish them. To address the problems above, we propose a new framework for fine-grained visual categorization. We strengthen the spatial correspondence among parts by including foreground segmentation and part localization. Based on the part representations of the images, we learn a large set of mid-level features which are more suitable for fine-grained tasks. Comparing with the low level features directly extracted from the images, the learned one-vs-all mid-level features enjoy the following advantages. First, the dimension of the mid-level features is relatively small. In order to obtain high classification accuracy, the dimension of the low level features usually reaches several thousand to tens of thousand, and becomes even larger when introducing spatial pyramid model. However, the dimension of our mid-level features is related to the number of classes, which is far less. Second, each entry of the proposed mid-level features is meaningful, which forms a more compact representation of the image. Third, the mid-level features are more robust than the low level ones, which is helpful for classification. Fourth, the learning process of the mid-level features is independent and can be easily combined with other techniques to boost the performance. We evaluate the proposed approach on the extensive fine-grained dataset CUB 200-2011 and Stanford Dogs, by learning the mid-level features based on the popular Fisher vectors and convolutional neural network, we boost the classification accuracy by a considerable margin and advance the state-of-the-art performance in fine-grained visual categorization. Xiaopeng Zhang 0008, Hongkai Xiong, Wengang Zhou 0001, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2014 | Spatio-temporal coherence for 3-D view synthesis with curve-based disparity warpingabstractView synthesis is dedicated to generating arbitrary views of the same scene from given inputs. As an alternative to depth-image-based rendering (DIBR), image warping based view synthesis approaches could automatically generate visually plausible virtual views in real-time. Recognizing that existing techniques would lead to temporal incoherence and shape distortions in synthesized videos, this paper proposes a novel video warping algorithm which motion saliency map and global motion from reference views are incorporated into motion-aware constraints to maintain temporal coherence in virtual views. Furthermore, a salient curve based disparity constraint is imposed to prevent shape deformations and avoid possible artifacts. Extensive experiments are validated by visual comparison, which demonstrates that the proposed algorithm outperforms existing warping-based methods. Hao Wang 0183, Xiaopeng Zhang 0008, Hongkai Xiong |
VCIP | 2 |
| 2013 | Illumination compensation via low rank matrix completion for multiview video codingabstractA multi-view video system would capture the same scene from different viewpoints, and suffer from significant illumination variations due to inaccurate camera calibration and varying light conditions. It deteriorates the inter-view correlations and the quality of synthesized views at decoder side. By converting the problem of illumination compensation to noise removing, this paper proposes a low-rank matrix completion algorithm to reduce the effect of illumination variation. Diverged from the existing work which chooses a central view as reference and keeps views consistent, it is dedicated to compensating all the views to match the low-rank structure of views. The discrepancies among views are regarded as mixed noise, and constructed as an incomplete matrix with low-rank. It is solved by a stable matrix completion and obtains a mapping function for color correction. It is robust to outliers since only plausible corresponding points are involved with completion. Experimental results show that the proposed algorithm can increase coding efficiency by up to 0.7 dB for luminance component and up to 2.1 dB for chrominance components. Xiaopeng Zhang 0008, Hongkai Xiong |
ICIP | 1 |