VLDB 2026 Research / reviewers in the wild / expert
Tao Chen 0003
dblp:69/510-3
· DBLP profile ↗
159ranked-venue papers
21as first author
125since 2021 · last 2026
0000-0002-0779-9818ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 92 · 16 first-author · 66 since 2021Artificial intelligence and machine learning · 88 · 4 first-author · 79 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 10 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion TransformersabstractWhile Diffusion Transformers (DiTs) have achieved breakthroughs in video generation, this long sequence generation task remains constrained by the quadratic complexity of attention mechanisms, resulting in significant inference latency. Through detailed analysis of attention maps in Video Diffusion Transformer (vDiT), we identify three recurring sparsity patterns: diagonal, multi-diagonal, and vertical-stripe structures. And even 3-6% attention heads can be skipped. Crucially, these patterns exhibit strong layer-depth and head-position correlations but show limited dependence on the input content. Leveraging these findings, we propose Sparse-vDiT, a sparsity acceleration framework for vDiT comprising: 1) Pattern-optimized sparse kernels that replace dense attention with computationally efficient implementations for each identified sparsity pattern. 2) An offline sparse diffusion search algorithm that selects the optimal sparse computation strategy per layer and head via hardware-aware cost modeling. After determining the optimal configuration, we fuse heads within the same layer that share the same attention strategy, enhancing inference efficiency. Integrated into state-of-the-art vDiT models (CogVideoX1.5, HunyuanVideo, and Wan2.1), Sparse-vDiT achieves 2.09×, 2.38×, and 1.67× theoretical FLOP reduction, and actual inference speedups of 1.76×, 1.85×, and 1.58×, respectively, while maintaining high visual fidelity, with PSNR values reaching 24.13, 27.09, and 22.59. Our work demonstrates that latent structural sparsity in vDiTs can be systematically exploited for long video synthesis. Pengtao Chen, Xianfang Zeng, Maosen Zhao, Mingzhu Shen, Gang Yu 0002, Tao Chen 0003 |
AAAI | 7 |
| 2026 | Mitigating Low-Quality Reasoning in MLLMs: Self-Driven Refined Multimodal CoT with Selective Thinking and Step-wise Visual EnhancementabstractCurrent Multimodal Chain-of-Thought (MCoT) methods suffer from low-quality multimodal reasoning, characterized by overthinking on simple queries and inefficient utilization of visual information, resulting in vast inefficient and ineffective computations. In this paper, we discover that Multimodal Large Language Models (MLLMs) possess inherent capabilities to distinguish between simple and difficult queries and enhance task-related visual information, which remain underutilized by existing approaches. Based on this insight, we propose Self-Driven Refined Multimodal CoT (SDR-MCoT), a training-free framework that mitigates these issues through two self-driven modules. First, our selective thinking module employs entropy-based confidence estimation to determine whether queries require detailed reasoning, preventing overthinking on simple questions. Second, our step-wise visual enhancement module strengthens attention to relevant visual regions at each reasoning step without inserting additional tokens, achieving fine-grained visual grounding and enhancement with minimal overhead. Moreover, SDR-MCoT can be seamlessly integrated into various MLLMs, offering a practical solution for improving multimodal reasoning. Comprehensive experiments across eight benchmarks from diverse domains (multimodal reasoning, visual understanding, hallucination, and mathematical reasoning) demonstrate that SDR-MCoT consistently outperforms existing MCoT methods on four different base models with reduced overhead. For instance, on Qwen2-VL-7B, our method improves average accuracy by over 6% while reducing token consumption by approximately 60% compared to zero-shot CoT. Chongjun Tu, Peng Ye 0006, Dongzhan Zhou, Tao Chen 0003, Wanli Ouyang |
AAAI | 4 |
| 2026 | A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven EnhancementabstractShengji Tang, Jianjian Cao, Weihao Lin, Jiale Hong, Bo Zhang, Shuyue Hu, Lei Bai, Tao Chen, Wanli Ouyang, Peng Ye. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shengji Tang, Jianjian Cao, Weihao Lin 0002, Jiale Hong, Bo Zhang 0069, Shuyue Hu, Lei Bai 0001, Tao Chen 0003, Wanli Ouyang, Peng Ye 0006 |
ACL (1) | 8 |
| 2026 | Δ-DiT: Accelerating Diffusion Transformers without Training via Denoising Property Alignment
Pengtao Chen, Mingzhu Shen, Peng Ye 0006, Jianjian Cao, Chongjun Tu, Christos-Savvas Bouganis, Tao Chen 0003 |
Int. J. Comput. Vis. | 8 |
| 2026 | Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs
Chongjun Tu, Peng Ye 0006, Dongzhan Zhou, Lei Bai 0001, Gang Yu 0002, Tao Chen 0003, Wanli Ouyang |
Int. J. Comput. Vis. | 6 |
| 2026 | Adapter-X: A general parameter-efficient fine-tuning framework for 2D and 3D vision
Peng Ye 0006, Lin Zhang 0055, Bizhe Bai, Tao Chen 0003 |
Neurocomputing | 5 |
| 2026 | Artifact-suppressed 3D retinal microvascular segmentation via multi-scale topology regulation
Ting Luo 0001, Jinxian Zhang, Tao Chen 0003, Zhouyan He, Yanda Meng, Jiong Zhang 0004, Dan Zhang 0026 |
Medical Image Anal. | 3 |
| 2026 | A knowledge-driven self-supervised learning method for enhancing EEG-based emotion recognition
Hanqi Wang, Peng Ye 0006, Kun Yang 0010, Jichuan Xiong, Tao Chen 0003 |
Neural Networks | 7 |
| 2026 | MADTP++: Bridge the Gap Between Token and Weight Pruning for Accelerating VLTsabstractVision-Language Transformers (VLTs) have achieved remarkable success, yet their high computational costs remain challenging due to numerous input tokens and large model parameters. Existing VLT compression methods primarily rely on single-modality-based token pruning or coarse-grained weight pruning techniques. However, these methods face significant obstacles, such as ignoring the critical alignment of different modalities and lacking layer-wise dynamic token pruning flexibility, exhibiting inevitable performance degradation due to coarsegrained weight pruning, and struggling with the simultaneous compression of both input tokens and model parameters. To address those limitations, we propose MADTP++, a novel approach that integrates custom-made token and weight pruning processes into a unified framework, achieving superior compression in both parameter counts and computational costs. Specifically, for the token pruning process, we introduce the Multi-modality Alignment Guidance (MAG) module and the Dynamic Token Pruning (DTP) module to align semantic features across different modalities and guide the dynamic elimination of redundant tokens based on different input instances. For the weight pruning process, we propose a Hardware-aware Weight Pruning (HWP) module that leverages the Sparse Tensor Cores across diverse hardware setups to enable fine-grained parameter pruning within VLTs. To further unify token and weight pruning, we also propose a Cooperative Optimization Training Strategy that automatically allocates GFLOPs and parameter reductions per branch before pruning and employs Knowledge Distillation Constraints to facilitate joint optimization of both pruning dimensions. Extensive experiments conducted on various VLT models and datasets demonstrate that MADTP++ can significantly reduce model parameters and computational costs while maintaining competitive performance. Jianjian Cao, Chong Yu 0001, Peng Ye 0006, Tao Chen 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | $\beta $-DARTS++: Bi-Level Regularization for Proxy-Robust Differentiable Architecture SearchabstractNeural Architecture Search (NAS) has attracted increasing attention in recent years because of its capability to design neural networks automatically. Among them, differential NAS approaches such as DARTS, have gained popularity for search efficiency. However, they still suffer from three main issues, that are, the weak stability due to the performance collapse, the poor generalization ability of the searched architectures, and the inferior robustness to different kinds of proxies (i.e., computationally reduced search configurations). To solve the search stability and searched architecture's generalization problems, a simple-but-effective regularization method, termed as Beta-Decay, is proposed to regularize the DARTS-based NAS searching process (referred as $\beta$β-DARTS). Specifically, Beta-Decay regularization can impose constraints to keep the value and variance of activated architecture parameters from being too large, thereby ensuring fair competition among architecture parameters and making the supernet less sensitive to the impact of input on the operation set. In-depth theoretical analyses on how it works and why it works are provided, and comprehensive experiments on a variety of search spaces and datasets validate that Beta-Decay regularization can help to stabilize the searching process and make the searched network more transferable across different datasets. To address the proxy robustness problem, we first benchmark differentiable NAS methods under a wide range of proxy data, proxy channels, proxy layers, and proxy epochs, since the robustness of NAS under different kinds of proxies has not been explored before. We then conclude some interesting findings and find that $\beta$β-DARTS always achieves the best result among all compared NAS methods under almost all proxy settings. We further introduce the novel flooding regularization to the weight optimization of $\beta$β-DARTS (termed as Bi-level regularization), and experimentally and theoretically verify its effectiveness for improving the proxy robustness of differentiable NAS. Peng Ye 0006, Tong He 0001, Baopu Li, Tao Chen 0003, Lei Bai 0001, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Bi3D++: Hybrid Bi-Domain Active Learning for Cross-Domain 3D Object DetectionabstractDomain adaptation has recently been widely explored for 3D detection. Previous works mainly use unsupervised domain adaptation (UDA) to address domain discrepancies. Despite notable improvements, their performance still largely trails models trained with fully annotated target data, due to larger domain gaps caused by different sensors and changing environments. In this paper, we exploit key characteristics of autonomous driving scenarios, including similar scenes and classimbalanced distributions, and explore a new task named active domain adaptation (ADA) for 3D object detection, which selects partial but important target data for annotation to further improve target-domain performance. Such a setting better reflects practical deployment in practice, where annotating all target-domain point clouds is prohibitively expensive while limited labels can substantially guide adaptation effectively. To this end, we propose a hybrid bi-domain active learning strategy, Bi3D++, to sample valuable data from both source and target domains and transfer source-domain knowledge to the target domain. Bi3D++ first samples target-like source data by measuring scene-level and instance-level similarity between domains, avoiding interference from irrelevant source data. Then, a hybrid active target sampling strategy selects target data by jointly considering rare-class similarity, intra-frame diversity, and inter-frame diversity, enabling diverse frames with diverse instances while emphasizing rare classes. Experiments on multiple cross-domain settings, including cross-beam and cross-location, show that Bi3D++ outperforms state-of-theart UDA methods with only 1% labeled target data and consistently improves performance as target annotations increase. Jiakang Yuan, Xiangchao Yan, Botian Shi, Bo Zhang 0069, Feng Xu 0001, Yu Qiao 0001, Tao Chen 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | A Unified multi-modality conditional latent diffusion model for point cloud generation
Yihang Yang, Zibo Zhao 0001, Fukun Yin, Wen Liu 0003, Yuhan Ding, Biao Jiang, Gang Yu 0002, Tao Chen 0003 |
Pattern Recognit. | 8 |
| 2026 | MTD-Net: A robust multi-task discriminative network for choroidal neovascularization segmentation
Dan Zhang 0026, Tao Chen 0003, Jianing Ying, Da Chen 0002, Baihua Li, Quanyong Yi, Jiong Zhang 0004 |
Pattern Recognit. | 3 |
| 2026 | Geometry-Aware Joint Attention for Efficient Native 3D Editing
Shuangkang Fang, Weicai Ye, Xuanyang Zhang, Yan-Pei Cao 0001, Gang Yu 0002, Tao Chen 0003 |
IEEE Signal Process. Lett. | 7 |
| 2026 | Temporally Coherent Dynamic Surface Reconstruction With Planar Gaussian SplattingabstractDynamic surface reconstruction requires not only photorealistic rendering but also temporally stable geometry. In this paper, we present DynaSurfGS, a dynamic surface reconstruction framework built on planar Gaussian splatting with a new temporal path consistency constraint. The key idea is to treat temporal coherence as a path-invariant property of deformation. For a given canonical primitive, its future state should be consistent, regardless of whether it is obtained by direct prediction or temporal propagation through an intermediate state. This trajectory-level formulation regularizes deformation in motion space rather than only supervising per-frame geometry, thereby suppressing temporal instability at its source. Built upon planar dynamic Gaussians and complemented by local geometric supervision, our framework produces both high-quality rendering and geometrically stable dynamic surfaces. Experiments demonstrate that DynaSurfGS achieves strong rendering performance while producing more coherent dynamic surfaces over time. Weicai Ye, Peng Ye 0006, Tong He 0001, Tao Chen 0003 |
IEEE Signal Process. Lett. | 5 |
| 2026 | FourierMask: Explain EEG-Based End-to-End Deep Learning Models in the Frequency DomainabstractThe rise of EEG-based end-to-end deep learning models has underscored the need to elucidate how these models process time-series raw EEG signals to generate predictions. The frequency domain provides a more suitable perspective for this task due to two key advantages: the strong correlation with cognitive states and the inherent capacity to model long-range temporal dependencies. However, this perspective remains underexplored in existing research. To bridge this gap, we propose FourierMask, the first mask perturbation framework specifically designed for frequency-domain explanation of EEG-based end-to-end models. Our method introduces three key innovations. First, the Fourier-based domain transformation enables direct manipulation of spectral components. Second, A learnable mask mechanism jointly models the spectral-spatial couplings relationship for EEG explanation. Third, a perturbation generator constrained by a target alignment loss ensures natural perturbations by minimizing distribution shift via cluster-aware regularization. We validate our method through experiments on an EEG benchmark dataset across EEGNet, TSCeption, and DeepConvNet models. Our method reaches a 36.0% average accuracy drop gap (vs. 8.6% for LIME and 6.6% for easyPEASI) at the group-level. And, it reaches a 17.8% average accuracy drop gap (vs. 8.9% for LIME and 9.9% for easyPEASI) at the instance-level. Our model-agnostic framework provides a plug-and-play solution for enhancing transparency of EEG-based end-to-end deep learning models. It links model decisions to frequency biomarkers, with potential applications in neuromedicine and brain-computer interfaces. Hanqi Wang, Kun Yang 0010, Jichuan Xiong, Tao Chen 0003 |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Rendered 2D Semantic and Generative Priors Guided 3D Multi-Object Groundingabstract3D multi-object visual grounding aims to identify and localize all objects in a 3D scene that correspond to a given text description. Unlike traditional single-object grounding, this task presents additional challenges as point clouds inherently lack fine-grained details, making it difficult to capture subtle object features and contextual information. Moreover, textual descriptions are inherently limited in perceiving and understanding complex 3D environments, especially in scenarios with high object similarity or intricate spatial arrangements. To tackle the above challenges, we propose SGMG, a Rendered 2D Semantic and Generative priors guided 3D Multi-object Grounding Framework. The SGMG framework introduces two key innovations that work cohesively to enhance grounding accuracy. First, a Generative-Assistant(GA) Module leverages the capabilities of a generative model to provide enriched scene prior information and capture the fine-grained scene details. Second, the Semantic-Augment Fusion(SAF) Module is designed to improve the representation of text features from the vision features, thereby boosting the accuracy of multimodal information interactions. Furthermore, we introduce a multi-level fusion mechanism, ensuring that semantic and spatial relationships between objects are preserved and effectively leveraged during the grounding process. Experimental results demonstrate that SGMG achieves state-of-the-art performance in multi-object 3D grounding and competitive results in traditional single-object tasks, highlighting its effectiveness in diverse scenarios. Peng Guo 0011, Hongyuan Zhu 0002, Hancheng Ye, Yihang Yang, Fukun Yin, Tao Chen 0003 |
IEEE Trans. Multim. | 6 |
| 2026 | Mamba-Based Progressive-Recovery Framework for Multimodal Low Light Image Enhancement
Mohammad Mahdizadeh, Jianjian Cao, Peng Ye 0006, Tao Chen 0003 |
IEEE Trans. Multim. | 4 |
| 2025 | All-in-One: Transferring Vision Foundation Models into Stereo MatchingabstractAs a fundamental vision task, stereo matching has made remarkable progress. While recent iterative optimization-based methods have achieved promising performance, their feature extraction capabilities still have room for improvement. Inspired by the ability of vision foundation models (VFMs) to extract general representations, in this work, we propose AIO-Stereo which can flexibly select and transfer knowledge from multiple heterogeneous VFMs to a single stereo matching model. To better reconcile features between heterogeneous VFMs and the stereo matching model and fully exploit prior knowledge from VFMs, we proposed a dual-level feature utilization mechanism that aligns heterogeneous features and transfers multi-level knowledge. Based on the mechanism, a dual-level selective knowledge transfer module is designed to selectively transfer knowledge and integrate the advantages of multiple VFMs. Experimental results show that AIO-Stereo achieves start-of-the-art performance on multiple datasets and ranks 1st on the Middlebury dataset and outperforms all the published work on the ETH3D benchmark. Jiakang Yuan, Peng Ye 0006, Tao Chen 0003, Hao Jiang 0013, Meiya Chen |
AAAI | 5 |
| 2025 | Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and FeedbackabstractJiakang Yuan, Xiangchao Yan, Bo Zhang, Tao Chen, Botian Shi, Wanli Ouyang, Yu Qiao, Lei Bai, Bowen Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiakang Yuan, Xiangchao Yan, Bo Zhang 0069, Tao Chen 0003, Botian Shi, Wanli Ouyang, Yu Qiao 0001, Lei Bai 0001, Bowen Zhou 0002 |
ACL (1) | 4 |
| 2025 | DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts ModelsabstractUpcycled Mixture-of-Experts (MoE) models have shown great potential in various tasks by converting the original Feed-Forward Network (FFN) layers in pre-trained dense models into MoE layers. However, these models still suffer from significant parameter inefficiency due to the introduction of multiple experts. In this work, we propose a novel DeRS (Decompose, Replace, and Synthesis) paradigm to overcome this shortcoming, which is motivated by our observations about the unique redundancy mechanisms of upcycled MoE experts. Specifically, DeRS decomposes the experts into one expert-shared base weight and multiple expert-specific delta weights, and subsequently represents these delta weights in lightweight forms. Our proposed DeRS paradigm can be applied to enhance parameter efficiency in two different scenarios, including: 1) DeRS Compression for inference stage, using sparsification or quantization to compress vanilla upcycled MoE models; and 2) DeRS Up-cycling for training stage, employing lightweight sparse or low-rank matrixes to efficiently upcycle dense models into MoE models. Extensive experiments across three different tasks show that the proposed methods can achieve extreme parameter efficiency while maintaining the performance for both training and compression of upcycled MoE models. Yongqi Huang, Peng Ye 0006, Chenyu Huang 0001, Jianjian Cao, Lin Zhang 0055, Baopu Li, Gang Yu 0002, Tao Chen 0003 |
CVPR | 8 |
| 2025 | Once-Tuning-Multiple-Variants: Tuning Once and Expanded as Multiple Vision-Language Model VariantsabstractVision-language model (VLM) is one of the most important models for multi-modal tasks. Real industrial applications often meet the challenge of adapting VLMs to different scenarios, such as varying hardware platforms or performance requirements. Traditional methods involve training or fine-tuning to adapt multiple unique VLMs or using model compression techniques to create multiple compact models. These approaches are complex and resource-intensive. This paper introduces a novel paradigm called Once-Tuning-Multiple-Variants (OTMV). OTMV requires only a single tuning process to inject dynamic weight expansion capacity into the original VLM structure. This tuned VLM can then be expanded into multiple variants tailored for different scenarios in inference. The tuning mechanism of OTMV is inspired by the mathematical series expansion theorem, which helps to reduce the parameter size and memory requirements while maintaining accuracy for VLM. Experiment results show that OTMV-tuned models achieve comparable accuracy to baseline VLMs across various visual-language tasks. The experiments also demonstrate the dynamic expansion capability of OTMV-tuned VLMs, outperforming traditional model compression and adaptation techniques in terms of accuracy and efficiency. Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001 |
CVPR | 2 |
| 2025 | Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-TuningabstractModel quantization reduces the bit-width of weights and activations, improving memory efficiency and inference speed in diffusion models. However, achieving 4-bit quantization remains challenging. Existing methods, primarily based on integer quantization and post-training quantization fine-tuning, struggle with inconsistent performance. Inspired by the success of floating-point (FP) quantization in large language models, we explore low-bit FP quantization for diffusion models and identify key challenges: the failure of signed FP quantization to handle asymmetric activation distributions, the insufficient consideration of temporal complexity in the denoising process during fine-tuning, and the misalignment between fine-tuning loss and quantization error. To address these challenges, we propose the mixup-sign floating-point quantization (MSFP) framework, first introducing unsigned FP quantization in model quantization, along with timestep-aware LoRA (TALoRA) and denoising-factor loss alignment (DFA), which ensure precise and stable fine-tuning. Extensive experiments show that we are the first to achieve superior performance in 4-bit FP quantization for diffusion models, outperforming existing PTQ fine-tuning methods in 4-bit INT quantization. Maosen Zhao, Pengtao Chen, Chong Yu 0001, Yan Wen 0005, Xudong Tan, Tao Chen 0003 |
CVPR | 6 |
| 2025 | Consistency-aware Self-Training for Iterative-based Stereo MatchingabstractIterative-based methods have become mainstream in stereo matching due to their high performance. However, these methods heavily rely on labeled data and face challenges with unlabeled real-world data. To this end, we propose a consistency-aware self-training framework for iterative-based stereo matching for the first time, leveraging real-world unlabeled data in a teacher-student manner. We first observe that regions with larger errors tend to exhibit more pronounced oscillation characteristics during model prediction. Based on this, we introduce a novel consistency-aware soft filtering module to evaluate the reliability of teacher-predicted pseudo-labels, which consists of a multi-resolution prediction consistency filter and an iterative prediction consistency filter to assess the prediction fluctuations of multiple resolutions and iterative optimization respectively. Further, we introduce a consistency-aware soft-weighted loss to adjust the weight of pseudo-labels accordingly, relieving the error accumulation and performance degradation problem due to incorrect pseudo-labels. Extensive experiments demonstrate that our method can improve the performance of various iterative-based stereo matching approaches in various scenarios. In particular, our method can achieve further enhancements over the current SOTA methods on several benchmark datasets. Peng Ye 0006, Jiakang Yuan, Rao Qiang, Yangchenxu Liu, Wu Cailin, Feng Xu 0001, Tao Chen 0003 |
CVPR | 9 |
| 2025 | Chimera: Improving Generalist Model with Domain-Specific ExpertsabstractRecent advancements in Large Multi-modal Models (LMMs) underscore the importance of scaling by increasing image-text paired data, achieving impressive performance on general tasks. Despite their effectiveness in broad applications, generalist models are primarily trained on web-scale datasets dominated by natural images, resulting in the sacrifice of specialized capabilities for domain-specific tasks that require extensive domain prior knowledge. Moreover, directly integrating expert models tailored for specific domains is challenging due to the representational gap and imbalanced optimization between the generalist model and experts. To address these challenges, we introduce Chimera, a scalable and low-cost multi-modal pipeline designed to boost the ability of existing LMMs with domain-specific experts. Specifically, we design a progressive training strategy to integrate features from expert models into the input of a generalist LMM. To address the imbalanced optimization caused by the well-aligned general visual encoder, we introduce a novel Generalist-Specialist Collaboration Masking (GSCM) mechanism. This results in a versatile model that excels across the chart, table, math, and document domains, achieving state-of-the-art performance on multi-modal reasoning and visual content extraction tasks, both of which are challenging tasks for assessing existing LMMs. Tianshuo Peng, Mingsheng Li, Jiakang Yuan, Hongbin Zhou, Renqiu Xia, Renrui Zhang, Lei Bai 0001, Song Mao, Bin Wang 0065, Aojun Zhou, Botian Shi, Tao Chen 0003, Bo Zhang 0069, Xiangyu Yue 0001 |
ICCV | 12 |
| 2025 | SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement LearningabstractWe propose SC-Captioner, a reinforcement learning framework that enables the self-correcting capability of image caption models. Our crucial technique lies in the design of the reward function to incentivize accurate caption corrections. Specifically, the predicted and reference captions are decomposed into object, attribute, and relation sets using scene-graph parsing algorithms. We calculate the set difference between sets of initial and self-corrected captions to identify added and removed elements. These elements are matched against the reference sets to calculate correctness bonuses for accurate refinements and mistake punishments for wrong additions and removals, thereby forming the final reward. For image caption quality assessment, we propose a set of metrics refined from CAPTURE that alleviate its incomplete precision evaluation and inefficient relation matching problems. Furthermore, we collect a fine-grained annotated image caption dataset, RefinedCaps, consisting of 6.5K diverse images from COCO dataset. Experiments show that applying SC-Captioner on large visual-language models can generate better image captions across various scenarios, significantly outperforming the direct preference optimization training strategy. Lin Zhang 0055, Xianfang Zeng, Kangcong Li, Gang Yu 0002, Tao Chen 0003 |
ICCV | 5 |
| 2025 | HiSplat: Hierarchical 3D Gaussian Splatting for Generalizable Sparse-View ReconstructionabstractReconstructing 3D scenes from multiple viewpoints is a fundamental task in stereo vision. Recently, advances in generalizable 3D Gaussian Splatting have enabled high-quality novel view synthesis for unseen scenes from sparse input views by feed-forward predicting per-pixel Gaussian parameters without extra optimization. However, existing methods typically generate single-scale 3D Gaussians, which lack representation of both large-scale structure and texture details, resulting in mislocation and artefacts. In this paper, we propose a novel framework, HiSplat, which introduces a hierarchical manner in generalizable 3D Gaussian Splatting to construct hierarchical 3D Gaussians via a coarse-to-fine strategy. Specifically, HiSplat generates large coarse-grained Gaussians to capture large-scale structures, followed by fine-grained Gaussians to enhance delicate texture details. To promote inter-scale interactions, we propose an Error Aware Module for Gaussian compensation and a Modulating Fusion Module for Gaussian repair. Our method achieves joint optimization of hierarchical representations, allowing for novel view synthesis using only two-view reference images. Comprehensive experiments on various datasets demonstrate that HiSplat significantly enhances reconstruction quality and cross-dataset generalization compared to prior single-scale methods. The corresponding ablation study and analysis of different-scale 3D Gaussians reveal the mechanism behind the effectiveness. Code is at https://github.com/Open3DVLab/HiSplat. Shengji Tang, Weicai Ye, Peng Ye 0006, Weihao Lin 0002, Tao Chen 0003, Wanli Ouyang |
ICLR | 6 |
| 2025 | GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-trainingabstractDespite their proficiency in general tasks, Multi-modal Large Language Models (MLLMs) struggle with automatic Geometry Problem Solving (GPS), which demands understanding diagrams, interpreting symbols, and performing complex reasoning. This limitation arises from their pre-training on natural images and texts, along with the lack of automated verification in the problem-solving process. Besides, current geometric specialists are limited by their task-specific designs, making them less effective for broader geometric problems. To this end, we present GeoX, a multi-modal large model focusing on geometric understanding and reasoning tasks. Given the significant differences between geometric diagram-symbol and natural image-text, we introduce unimodal pre-training to develop a diagram encoder and symbol decoder, enhancing the understanding of geometric images and corpora. Furthermore, we introduce geometry-language alignment, an effective pre-training paradigm that bridges the modality gap between unimodal geometric experts. We propose a Generator-And-Sampler Transformer (GS-Former) to generate discriminative queries and eliminate uninformative representations from unevenly distributed geometric signals. Finally, GeoX benefits from visual instruction tuning, empowering it to take geometric images and questions as input and generate verifiable solutions. Experiments show that GeoX outperforms both generalists and geometric specialists on publicly recognized benchmarks, such as GeoQA, UniGeo, Geometry3K, and PGPS9k. Our data and code will be released soon to accelerate future research on automatic GPS. Renqiu Xia, Mingsheng Li, Hancheng Ye, Hongbin Zhou, Jiakang Yuan, Tianshuo Peng, Xinyu Cai, Xiangchao Yan, Bin Wang 0065, Conghui He, Botian Shi, Tao Chen 0003, Junchi Yan, Bo Zhang 0069 |
ICLR | 13 |
| 2025 | Boost Embodied AI Models with Robust Compression BoundaryabstractThe rapid improvement of deep learning models with the integration of the physical world has dramatically improved embodied AI capabilities. Meanwhile, the powerful embodied AI models and their scales place an increasing burden on deployment efficiency. The efficiency issue is more apparent on embodied AI platforms than on data centers because they have more limited computational resources and memory bandwidth. Meanwhile, most embodied AI scenarios, like autonomous driving and robotics, are more sensitive to fast responses. Theoretically, the traditional model compression techniques can help embodied AI models with more efficient computation, lower memory and energy consumption, and reduced latency. Because the embodied AI models are expected to interact with the physical world, the corresponding compressed models are also expected to resist natural corruption caused by real-world events such as noise, blur, weather conditions, and even adversarial corruption. This paper explores the novel paradigm to boost the efficiency of the embodied AI models and the robust compression boundary. The efficacy of our method has been proven to find the optimal balance between accuracy, efficiency, and robustness in real-world conditions. Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001 |
IJCAI | 2 |
| 2025 | Cross-Modal Graph Learning for Perivascular Spaces Segmentation
Tao Chen 0003, Dan Zhang 0026, Xi Long 0001, Marcel Breeuwer, Svitlana Zinger, Peiyu Huang, Jiong Zhang 0004 |
MICCAI (4) | 1 |
| 2025 | DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
Zhende Song, Jiamu Sheng, Chi Zhang 0007, Shengji Tang, Jiayuan Fan 0001, Tao Chen 0003 |
ACM Multimedia | 7 |
| 2025 | PaceLLM: Brain-Inspired Large Language Models for Long-Context UnderstandingabstractWhile Large Language Models (LLMs) demonstrate strong performance across domains, their long-context capabilities are limited by transient neural activations causing information decay and unstructured feed-forward network (FFN) weights leading to semantic fragmentation. Inspired by the brain’s working memory and cortical modularity, we propose PaceLLM, featuring two innovations: (1) a Persistent Activity (PA) Mechanism that mimics prefrontal cortex (PFC) neurons’ persistent firing by introducing an activation-level memory bank to dynamically retrieve, reuse, and update critical FFN states, addressing contextual decay; and (2) Cortical Expert (CE) Clustering that emulates task-adaptive neural specialization to reorganize FFN weights into semantic modules, establishing cross-token dependencies and mitigating fragmentation. Extensive evaluations show that PaceLLM achieves 6% improvement on LongBench’s Multi-document QA and 12.5–17.5% performance gains on $\infty$-Bench tasks, while extending measurable context length to 200K tokens in Needle-In-A-Haystack (NIAH) tests. This work pioneers brain-inspired LLM optimization and is complementary to other works. Besides, it can be generalized to any model and enhance their long-context performance and interpretability without structural overhauls. Kangcong Li, Peng Ye 0006, Chongjun Tu, Lin Zhang 0055, Chunfeng Song, Qihao Zheng, Tao Chen 0003 |
NeurIPS | 9 |
| 2025 | FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion UnderstandingabstractMultimodal Large Language Models (MLLMs) have shown impressive video content understanding capabilities but struggle with fine-grained motion comprehension. To comprehensively assess the motion understanding ability of existing MLLMs, we introduce FAVOR-Bench, which comprises 1,776 videos from both ego-centric and third-person perspectives and enables assessment through both close-ended and open-ended tasks. For close-ended evaluation, we carefully design 8,184 multiple-choice question-answer pairs spanning six distinct sub-tasks. For open-ended evaluation, we employ the GPT-assisted evaluation and develop a novel cost-efficient LLM-free assessment method, where the latter can enhance benchmarking interpretability and accessibility. Comprehensive experiments with21 state-of-the-art MLLMs reveal significant limitations in their ability to comprehend and describe detailed temporal dynamics in video motions. To alleviate this limitation, we further build FAVOR-Train, a dataset of 17,152 videos with fine-grained motion annotations. Finetuning Qwen2.5-VL on FAVOR-Train yields consistent improvements on motion-related tasks across TVBench, MotionBenchand our FAVOR-Bench. Our assessment results demonstrate that the proposed FAVOR-Bench and FAVOR-Train provide valuable tools for the community to develop more powerful video understanding models. Chongjun Tu, Lin Zhang 0055, Pengtao Chen, Peng Ye 0006, Xianfang Zeng, Gang Yu 0002, Tao Chen 0003 |
NeurIPS | 8 |
| 2025 | Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta CompressionabstractWith the rise of the fine-tuned–pretrained paradigm, storing numerous fine-tuned models for multi-tasking creates significant storage overhead.
Delta compression alleviates this by storing only the pretrained model and the highly compressed delta weights (the differences between fine-tuned and pretrained model weights).
However, existing methods fail to maintain both high compression and performance, and often rely on data.
To address these challenges, we propose UltraDelta, the first data-free delta compression pipeline that achieves both ultra-high compression and strong performance.
UltraDelta is designed to minimize redundancy, maximize information, and stabilize performance across inter-layer, intra-layer, and global dimensions, using three key components:
(1) Variance-Based Mixed Sparsity Allocation assigns sparsity based on variance, giving lower sparsity to high-variance layers to preserve inter-layer information.
(2) Distribution-Aware Compression applies uniform quantization and then groups parameters by value, followed by group-wise pruning, to better preserve intra-layer distribution.
(3) Trace-Norm-Guided Rescaling uses the trace norm of delta weights to estimate a global rescaling factor, improving model stability under higher compression.
Extensive experiments across
(a) large language models (fine-tuned on LLaMA-2 7B and 13B) with up to 50$\times$ compression,
(b) general NLP models (RoBERTa-base, T5-base) with up to 224$\times$ compression,
(c) vision models (ViT-B/32, ViT-L/14) with up to 132$\times$ compression, and
(d) multi-modal models (BEiT-3) with 18$\times$ compression,
demonstrate that UltraDelta consistently outperforms existing methods, especially under ultra-high compression.
Code is available at https://github.com/xiaohuiwang000/UltraDelta. Peng Ye 0006, Chenyu Huang 0001, Shenghe Zheng, Bo Zhang 0069, Lei Bai 0001, Wanli Ouyang, Tao Chen 0003 |
NeurIPS | 8 |
| 2025 | PRING: Rethinking Protein-Protein Interaction Prediction from Pairs to GraphsabstractDeep learning-based computational methods have achieved promising results in predicting protein-protein interactions (PPIs). However, existing benchmarks predominantly focus on isolated pairwise evaluations, overlooking a model's capability to reconstruct biologically meaningful PPI networks, which is crucial for biology research. To address this gap, we introduce PRING, the first comprehensive benchmark that evaluates PRotein-protein INteraction prediction from a Graph-level perspective. PRING curates a high-quality, multi-species PPI network dataset comprising 21,484 proteins and 186,818 interactions, with well-designed strategies to address both data redundancy and leakage. Building on this golden-standard dataset, we establish two complementary evaluation paradigms: (1) topology-oriented tasks, which assess intra and cross-species PPI network construction, and (2) function-oriented tasks, including protein complex pathway prediction, GO module analysis, and essential protein justification. These evaluations not only reflect the model's capability to understand the network topology but also facilitate protein function annotation, biological module detection, and even disease mechanism analysis. Extensive experiments on four representative model categories, consisting of sequence similarity-based, naive sequence-based, protein language model-based, and structure-based approaches, demonstrate that current PPI models have potential limitations in recovering both structural and functional properties of PPI networks, highlighting the gap in supporting real-world biological applications. We believe PRING provides a reliable platform to guide the development of more effective PPI prediction models for the community. The dataset and source code of PRING are available at https://github.com/SophieSarceau/PRING. Xinzhe Zheng 0001, Fanding Xu, Jinzhe Li, Zhiyuan Liu 0001, Wenkang Wang, Tao Chen 0003, Wanli Ouyang, Stan Z. Li, Yan Lu 0001, Nanqing Dong, Yang Zhang 0094 |
NeurIPS | 7 |
| 2025 | Deep generative model for protein subcellular localization predictionabstractProtein sequence not only determines its structure but also provides important clues of its subcellular localization. Although a series of artificial intelligence models have been reported to predict protein subcellular localization, most of them provide only textual outputs. Here, we present deepGPS, a deep generative model for protein subcellular localization prediction. After training with protein primary sequences and fluorescence images, deepGPS shows the ability to predict cytoplasmic and nuclear localizations by reporting both textual labels and generative images as outputs. In addition, cell-type-specific deepGPS models can be developed by using distinct image datasets from different cell lines for comparative analyses. Moreover, deepGPS shows potential to be further extended for other specific organelles, such as vesicles and endoplasmic reticulum, even with limited volumes of training data. Finally, the openGPS website (https://bits.fudan.edu.cn/opengps) is constructed to provide a publicly accessible and user-friendly platform for studying protein subcellular localization and function. Guo-Hua Yuan, Jinzhe Li, Zejun Yang, Yao-Qi Chen, Zhonghang Yuan, Tao Chen 0003, Wanli Ouyang, Nanqing Dong |
Briefings Bioinform. | 6 |
| 2025 | ClipSAM: CLIP and SAM collaboration for zero-shot anomaly segmentation
Shengze Li, Jianjian Cao, Peng Ye 0006, Yuhan Ding, Chongjun Tu, Tao Chen 0003 |
Neurocomputing | 6 |
| 2025 | Learnable Bi-directional Data Augmentation for few-shot cross-domain point cloud classification
Lin Zhang 0055, Jiakang Yuan, Tao Chen 0003 |
Neurocomputing | 5 |
| 2025 | Neovascularization Segmentation via a Multilateral Interaction-Enhanced Graph Convolutional NetworkabstractChoroidal neovascularization (CNV), a primary characteristic of wet age-related macular degeneration (wet AMD), represents a leading cause of blindness worldwide. In clinical practice, optical coherence tomography angiography (OCTA) is commonly used for studying CNV-related pathological changes, due to its micron-level resolution and non-invasive nature. Thus, accurate segmentation of CNV regions and vessels in OCTA images is crucial for clinical assessment of wet AMD. However, challenges existed due to irregular CNV shapes and imaging limitations like projection artifacts, noises and boundary blurring. Moreover, the lack of publicly available datasets constraints the CNV analysis. To address these challenges, this paper constructs the first publicly accessible CNV dataset (CNVSeg), and proposes a novel multilateral graph convolutional interaction-enhanced CNV segmentation network (MTG-Net). This network integrates both region and vessel morphological information, exploring semantic and geometric duality constraints within the graph domain. Specifically, MTG-Net consists of a multi-task framework and two graph-based cross-task modules: Multilateral Interaction Graph Reasoning (MIGR) and Multilateral Reinforcement Graph Reasoning (MRGR). The multi-task framework encodes rich geometric features of lesion shapes and surfaces, decoupling the image into three task-specific feature maps. MIGR and MRGR iteratively reason about higher-order relationships across tasks through a graph mechanism, enabling complementary optimization for task-specific objectives. Additionally, an uncertainty-weighted loss is proposed to mitigate the impact of artifacts and noise on segmentation accuracy. Experimental results demonstrate that MTG-Net outperforms existing methods, achieving a Dice socre of 87.21% for region segmentation and 88.12% for vessel segmentation. Tao Chen 0003, Dan Zhang 0026, Da Chen 0002, Huazhu Fu, Shanshan Wang 0002, Laurent D. Cohen, Yitian Zhao, Quanyong Yi, Jiong Zhang 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | SPOT: Scalable 3D Pre-Training via Occupancy Prediction for Learning Transferable 3D RepresentationsabstractAnnotating 3D LiDAR point clouds for perception tasks is fundamental for many applications e.g. autonomous driving, yet it still remains notoriously labor-intensive. Pretraining-finetuning approach can alleviate the labeling burden by fine-tuning a pre-trained backbone across various downstream datasets as well as tasks. In this paper, we propose SPOT, namely Scalable Pre-training via Occupancy prediction for learning Transferable 3D representations under such a label-efficient fine-tuning paradigm. SPOT achieves effectiveness on various public datasets with different downstream tasks, showcasing its general representation power, cross-domain robustness and data scalability which are three key factors for real-world application. Specifically, we both theoretically and empirically show, for the first time, that general representations learning can be achieved through the task of occupancy prediction. Then, to address the domain gap caused by different LiDAR sensors and annotation methods, we develop a beam re-sampling technique for point cloud augmentation combined with class-balancing strategy. Furthermore, scalable pre-training is observed, that is, the downstream performance across all the experiments gets better with more pre-training data. Additionally, such pre-training strategy also remains compatible with unlabeled data. The hope is that our findings will facilitate the understanding of LiDAR points and pave the way for future advancements in LiDAR pre-training. Xiangchao Yan, Runjian Chen, Bo Zhang 0069, Hancheng Ye, Renqiu Xia, Jiakang Yuan, Hongbin Zhou, Xinyu Cai, Botian Shi, Wenqi Shao, Ping Luo 0002, Yu Qiao 0001, Tao Chen 0003, Junchi Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 13 |
| 2025 | Stimulative Training++: Go Beyond the Performance Limits of Residual NetworksabstractResidual networks have shown great success and become indispensable in recent deep neural network models. In this work, we aim to re-investigate the training process of residual networks from a novel perspective of loafing, and further propose a new training scheme as well as three improved strategies for boosting residual networks beyond their performance limits. Previous research has suggested that residual networks can be considered as ensembles of shallow networks, which implies that the final performance of a residual network is influenced by a group of subnetworks. Furthermore, we identify a previously overlooked problem, where subnetworks within a residual network are prone to exert less effort when working as part of a group compared to working alone. We define this problem as network loafing. Since network loafing may inevitably cause the sub-par performance of the residual network, we propose a novel training scheme called stimulative training, which randomly samples a residual subnetwork and calculates the KL divergence loss between the sampled subnetwork and the given residual network for extra supervision. In order to unleash the potential of stimulative training, we further propose three simple-yet-effective strategies, including a novel KL- loss that only aligns the network logits direction, random smaller inputs for subnetworks, and inter-stage sampling rules. Comprehensive experiments and analysis verify the effectiveness of stimulative training as well as its three improved strategies. For example, the proposed method can boost the performance of ResNet50 on ImageNet to 80.5% Top1 accuracy without using any extra data, model, trick, or changing the structure. With only uniform augment, the performance can be further improved to 81.0% Top1 accuracy, better than the best training recipes provided by Timm library and PyTorch official version. We also verify its superiority on various typical models, datasets, and tasks and give some theoretical analysis. As such, we advocate utilizing the proposed method as a general and next-generation technology to train residual networks. Peng Ye 0006, Tong He 0001, Shengji Tang, Baopu Li, Tao Chen 0003, Lei Bai 0001, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Taylor-Series-Expansion-Based Vision Transformer ModelsabstractTaylor-Series-Expansion (TSE) is a mathematics theorem. It proves that the expansion of the first few finite Taylor Series is a good approximation of a nonlinear function in most cases. Inspired by the TSE theorem, a brand-new TSE-based vision transformer is designed. TSE-based vision transformer uses the shared first-order TSE transformer block's weight (in analogy with the Taylor-Series first-order term), its finite multiple multiplications (in analogy with the Taylor-Series expanded high-order terms), and the corresponding learnable TSE coefficients to approximate the naive vision transformer. In this manner, the TSE-based vision model reduces the memory burden but keeps a similar accuracy as the naive counterpart. Derived from adding the Taylor skip mechanism in training, the TSE-based vision transformer has good dynamic expansion capability. Experiment results show TSE-based models can boost actual deployment latency by 1.30-1.36× on A100 GPU and 1.34-1.45× on AGX Orin with negligible accuracy degradation on ImageNet classification, COCO detection, and ADE20K segmentation benchmarking tasks. Moreover, TSE-based optimization is orthogonal to model compression. Combining with the state-of-the-art vision transformer compression method, it can boost actual deployment performance by 1.70-1.87× and 3.29-3.61× of latency and throughput on A100 GPU, and 1.67-1.74× and 2.76-2.94× improvement of latency and throughput on AGX Orin. Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | BridgeNet: Comprehensive and Effective Feature Interactions via Bridge Feature for Multi-Task Dense PredictionsabstractMulti-task dense prediction aims at handling multiple pixel-wise prediction tasks within a unified network simultaneously for visual scene understanding. However, cross-task feature interactions of current methods are still suffering from incomplete levels of representations, less discriminative semantics in feature participants, and inefficient pair-wise task interaction processes. To tackle these under-explored issues, we propose a novel BridgeNet framework, which extracts comprehensive and discriminative intermediate Bridge Features, and conducts interactions based on them. Specifically, a Task Pattern Propagation (TPP) module is first applied to ensure highly semantic task-specific feature participants are prepared for subsequent interactions, and a Bridge Feature Extractor (BFE) is specially designed to selectively integrate both high-level and low-level representations to generate the comprehensive bridge features. Then, instead of conducting heavy pair-wise cross-task interactions, a Task-Feature Refiner (TFR) is developed to efficiently take guidance from bridge features and form final task predictions. To the best of our knowledge, this is the first work considering the completeness and quality of feature participants in cross-task interactions. Extensive experiments are conducted on NYUD-v2, Cityscapes and PASCAL Context benchmarks, and the superior performance shows the proposed architecture is effective and powerful in promoting different dense prediction tasks simultaneously. Jingdong Zhang 0003, Jiayuan Fan 0001, Peng Ye 0006, Bo Zhang 0069, Hancheng Ye, Baopu Li, Yancheng Cai, Tao Chen 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | 3D microvascular reconstruction in retinal OCT angiography images via domain-adaptive learning
Jiong Zhang 0004, Yonghuai Liu, Dan Zhang 0026, Jianyang Xie, Tao Chen 0003, Yalin Zheng, Huazhu Fu, Yitian Zhao |
Pattern Recognit. | 6 |
| 2025 | Sparse-to-Dense Training: A Novel Training Scheme to Enhance Vision TransformersabstractAs Vision Transformers (ViTs) become increasingly popular in various vision tasks, one may question:if a new training scheme for ViTs exists that can improve performance without increasing training and inference computation cost?In this paper, we affirmatively answer this question with a novel Sparse-to-Dense (S2D) training scheme. Specifically, we decouple the training and inference phases of ViTs. During training, we replace some Feed-Forward Network (FFN) layers of ViTs with computationally efficient RUP-Mixture-of-FFN (RUP-MoF) layers, each comprising multiple FFN experts and allocating tokens to experts via Random Uniform Partition (RUP). Furthermore, an additional Experts Weights Averaging (EWA) update is performed specifically on these RUP-MoF layers after each gradient update. After training, we convert each RUP-MoF layer back to a single FFN layer by averaging the experts, transforming the training-time sparse model back to the original dense ViT model for inference. We further provide theoretical analysis to illustrate why and how it works. Comprehensive experiments across various 2D and 3D vision tasks, ViT architectures and datasets validate the effectiveness and generalization ability of the proposed S2D training scheme. Besides, we show that, S2D training scheme can also be applied to improve the performance of Transformer-based language models, and EWA update technique can also significantly improve the effectiveness of classic Mixture-of-Experts on various 2D vision small-scale datasets and 3D vision tasks. Yongqi Huang, Peng Ye 0006, Chongjun Tu, Tao Chen 0003, Tong He 0001, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Efficient Architecture Search via Bi-Level Data PruningabstractImproving the efficiency of Neural Architecture Search (NAS) is a challenging but significant task that has received much attention. Previous studies mainly adopt the Differentiable Architecture Search (DARTS) and improve its search strategies or modules to enhance search efficiency. Recently, some methods have started considering data reduction for speedup, but they are not tightly coupled with the architecture search process and cannot capture the training dynamics of DARTS well, resulting in sub-optimal performances. To this end, this work pioneers an exploration into the critical role of dataset characteristics in the bi-level optimization of DARTS, and then proposes a novel Bi-level Data Pruning (BDP) paradigm that targets the weights and architecture levels of DARTS to enhance efficiency from a data perspective. Specifically, we introduce a progressive bi-level data pruning strategy that utilizes supernet prediction dynamics as the metric to gradually prune unsuitable samples for DARTS during the search. An effective automatic class balance constraint is also integrated into BDP, to suppress potential class imbalances resulting from data-efficient algorithms. Comprehensive evaluations on the NAS-Bench-201 search space, DARTS search space, and MobileNet-like search space validate that BDP reduces search costs by over 50% while achieving superior performance when applied to the baseline DARTS. Besides, we demonstrate that BDP can harmoniously integrate with advanced DARTS variants, like P-DARTS, PC-DARTS, EG-NAS, and$\beta $-DARTS, offering an approximately$2\times $speedup with minimal performance compromise. Chongjun Tu, Peng Ye 0006, Weihao Lin 0002, Hancheng Ye, Chong Yu 0001, Tao Chen 0003, Baopu Li, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Dynamic Model Merging With Mixture of WeightsabstractThe pretrain-finetune paradigm brings about the release of numerous model weights. Under this background, model merging is becoming increasingly popular, as it enables a model to handle multiple tasks by fusing model weights from these tasks, without the need for labeled data, additional training, or high training costs. Though with great potential, model merging suffers from severe performance degradation due to the interference among model weights. And existing model merging methods (i.e., static merging) commonly provide a single set of merging coefficients for all the input samples and do not distinguish layers based on the severity of weight interference, which may not be the optimal solution. In this paper, we propose MoW-Merging, a dynamic model merging method based on Mixture of Weights. First, we apply a gating network to adaptively generate merging coefficients depending on the input samples, realizing sample-wisely dynamic merging and automated classifier selection. The gating network is lightweight and is trained with only a small number of unlabeled data. Further, we utilize a weight similarity metric to judge the severity of weight interference of each layer and apply suitable merging methods to different layers. The proposed MoW-Merging shows plug-and-play capabilities and can be seamlessly combined with various model merging methods to greatly boost their performance. The effectiveness of MoW-Merging is validated by comprehensive experiments on various classical and newly-established benchmarks under multiple settings. The code is available athttps://github.com/harveyhuang18/Mixture_of_Weights. Peng Ye 0006, Chenyu Huang 0001, Mingzhu Shen, Tao Chen 0003, Yongqi Huang, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Corrections to "DCNet: Large-Scale Point Cloud Semantic Segmentation With Discriminative and Efficient Feature Aggregation"abstractPresents corrections to the paper, “DCNet: Large-Scale Point Cloud Semantic Segmentation With Discriminative and Efficient Feature Aggregation”. Fukun Yin, Tao Chen 0003, Guozhong Luo, Gang Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | WI3D: Weakly Incremental 3D Detection via Vision Foundation ModelsabstractClass-incremental 3D object detection demands a 3D detector tolocateandrecognizenovel categories in a stream fashion while preserving its base detection ability. However, existing methods require delicate 3D annotations for learning novel categories, resulting in significant labeling costs. To this end, we explore a label-efficient approach calledWeaklyIncremental3DDetection (WI3D), which teaches a 3D detector to learn incrementally with off-the-shelf vision foundation models. We propose a novel dual-teaching framework incorporating both intra-modal and inter-modal knowledge from pseudo labels and feature space. Specifically, our framework features a class-agnostic pseudo-label refinement module, designed for the generation of high-quality 3D pseudo labels. This module is built on a lightweight transformer that models the spatial relationships between pseudo labels and their interactions with rich contextual information in point clouds. Additionally, we introduce a cross-modal knowledge transfer module to enhance the representation learning of novel classes, along with a reweighting knowledge distillation strategy that dynamically assesses and distills knowledge from previously learned categories. Extensive experiments show that our approach can efficiently learn novel concepts while preserving knowledge of base classes in WI3D scenarios, and surpass baseline approaches on both SUN-RGBD and ScanNet. Mingsheng Li, Sijin Chen, Shengji Tang, Hongyuan Zhu 0002, Yanyan Fang, Xin Chen 0040, Zhuoyuan Li 0006, Fukun Yin, Tao Chen 0003 |
IEEE Trans. Multim. | 9 |
| 2025 | ShapeGPT: 3D Shape Generation With a Unified Multi-Modal Language ModelabstractThe advent of large language models, which enable flexibility through instruction-driven approaches, has revolutionized many traditional generative tasks, but large models for 3D data, particularly in comprehensively handling 3D shapes with other modalities, are still under-explored. By achieving instruction-based shape generation, versatile multi-modal generative shape models can significantly benefit various fields, such as 3D virtual construction and network-aided design. In this article, we present ShapeGPT, a shape-included multi-modal framework to leverage strong pre-trained language models to address multiple shape-relevant tasks. Specifically, ShapeGPT employs a “word-sentence-paragraph” framework to discretize continuous shapes into shape words, further assembles these words into shape sentences, and integrates shape with instructional text for multi-modal paragraphs. To learn this shape-language model, we use a three-stage training scheme, including shape representation, multi-modal alignment, and instruction-based generation, to align shape-language codebooks and learn the intricate correlations among these modalities. Extensive experiments demonstrate that ShapeGPT achieves comparable performance across shape-relevant tasks, including text-to-shape, shape-to-text, shape completion, and shape editing. Fukun Yin, Xin Chen 0040, Chi Zhang 0007, Biao Jiang, Zibo Zhao 0001, Wen Liu 0003, Gang Yu 0002, Tao Chen 0003 |
IEEE Trans. Multim. | 8 |
| 2024 | Boosting Residual Networks with Group KnowledgeabstractRecent research understands the residual networks from a new perspective of the implicit ensemble model. From this view, previous methods such as stochastic depth and stimulative training have further improved the performance of the residual network by sampling and training of its subnets. However, they both use the same supervision for all subnets of different capacities and neglect the valuable knowledge generated by subnets during training. In this manuscript, we mitigate the significant knowledge distillation gap caused by using the same kind of supervision and advocate leveraging the subnets to provide diverse knowledge. Based on this motivation, we propose a group knowledge based training framework for boosting the performance of residual networks. Specifically, we implicitly divide all subnets into hierarchical groups by subnet-in-subnet sampling, aggregate the knowledge of different subnets in each group during training, and exploit upper-level group knowledge to supervise lower-level subnet group. Meanwhile, we also develop a subnet sampling strategy that naturally samples larger subnets, which are found to be more helpful than smaller subnets in boosting performance for hierarchical groups. Compared with typical subnet training and other methods, our method achieves the best efficiency and performance trade-offs on multiple datasets and network structures. The code is at https://github.com/tsj-001/AAAI24-GKT. Shengji Tang, Peng Ye 0006, Baopu Li, Weihao Lin 0002, Tao Chen 0003, Tong He 0001, Chong Yu 0001, Wanli Ouyang |
AAAI | 5 |
| 2024 | PM-INR: Prior-Rich Multi-Modal Implicit Large-Scale Scene Neural RepresentationabstractRecent advancements in implicit neural representations have contributed to high-fidelity surface reconstruction and photorealistic novel view synthesis. However, with the expansion of the scene scale, such as block or city level, existing methods will encounter challenges because traditional sampling cannot cope with the cubically growing sampling space. To alleviate the dependence on filling the sampling space, we explore using multi-modal priors to assist individual points to obtain more global semantic information and propose a priorrich multi-modal implicit neural representation network, Pm-INR, for the outdoor unbounded large-scale scene. The core of our method is multi-modal prior extraction and crossmodal prior fusion modules. The former encodes codebooks from different modality inputs and extracts valuable priors, while the latter fuses priors to maintain view consistency and preserve unique features among multi-modal priors. Finally, feature-rich cross-modal priors are injected into the sampling regions to allow each region to perceive global information without filling the sampling space. Extensive experiments have demonstrated the effectiveness and robustness of our method for outdoor unbounded large-scale scene novel view synthesis, which outperforms state-of-the-art methods in terms of PSNR, SSIM, and LPIPS. Fukun Yin, Wen Liu 0003, Jiayuan Fan 0001, Xin Chen 0040, Gang Yu 0002, Tao Chen 0003 |
AAAI | 7 |
| 2024 | MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language TransformerabstractVision-Language Transformers (VLTs) have shown great success recently, but are meanwhile accompanied by heavy computation costs, where a major reason can be attributed to the large number of visual and language tokens. Existing token pruning research for compressing VLTs mainly follows a single-modality-based scheme yet ignores the critical role of aligning different modalities for guiding the token pruning process, causing the important tokens for one modality to be falsely pruned in another modality branch. Meanwhile, existing VLT pruning works also lack the flexibility to dynamically compress each layer based on different input samples. To this end, we propose a novel framework named Multimodal Alignment-Guided Dynamic Token Pruning (MADTP) for accelerating various VLTs. Specifically, we first introduce a well-designed Multi-modality Alignment Guidance (MAG) module that can align features of the same semantic concept from different modalities, to ensure the pruned tokens are less important for all modalities. We further design a novel Dynamic Token Pruning (DTP) module, which can adaptively adjust the token compression ratio in each layer based on different input instances. Extensive experiments on various benchmarks demonstrate that MADTP significantly reduces the computational complexity of kinds of multimodal models while preserving competitive performance. Notably, when applied to the BLIP model in the NLVR2 dataset, MADTP can reduce the GFLOPs by 80% with less than 4% performance degradation. The code is available at https://github.com/double125IMADTP. Jianjian Cao, Peng Ye 0006, Shengze Li, Chong Yu 0001, Yansong Tang, Jiwen Lu, Tao Chen 0003 |
CVPR | 7 |
| 2024 | LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and PlanningabstractRecent progress in Large Multimodal Models (LMM) has opened up great possibilities for various applications in the field of human-machine interactions. However, developing LMMs that can comprehend, reason, and plan in complex and diverse 3D environments remains a challenging topic, especially considering the demand for understanding permutation-invariant point cloud representations of the 3D scene. Existing works seek help from multi-view images by projecting 2D features to 3D space, which inevitably leads to huge computational overhead and performance degradation. In this paper, we present LL3DA, a Large Language 3D Assistant that takes point cloud as the direct input and responds to both text instructions and visual interactions. The additional visual interaction enables LMMs to better comprehend human interactions with the 3D environment and further remove the ambiguities within plain texts. Experiments show that LL3DA achieves remarkable results and surpasses various 3D vision-language models on both 3D Dense Captioning and 3D Question Answering. Sijin Chen, Xin Chen 0040, Chi Zhang 0007, Mingsheng Li, Gang Yu 0002, Hao Fei 0001, Hongyuan Zhu 0002, Jiayuan Fan 0001, Tao Chen 0003 |
CVPR | 9 |
| 2024 | Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer CompressionabstractRecent Vision Transformer Compression (VTC) works mainly follow a two-stage scheme, where the importance score of each model unit is first evaluated or preset in each submodule, followed by the sparsity score evaluation ac-cording to the target sparsity constraint. Such a separate evaluation process induces the gap between importance and sparsity score distributions, thus causing high search costs for VTC. In this work, for the first time, we investigate how to integrate the evaluations of importance and sparsity scores into a single stage, searching the optimal subnets in an effi-cient manner. Specifically, we present OFB, a cost-efficient approach that simultaneously evaluates both importance and sparsity scores, termed Once for Both (OFB), for VTC. First, a bi-mask scheme is developed by entangling the importance score and the differentiable sparsity score to jointly deter-mine the pruning potential (prunability) of each unit. Such a bi-mask search strategy is further used together with a proposed adaptive one-hot loss to realize the progressive-and-efficient search for the most important subnet. Finally, Progressive Masked Image Modeling (PMIM) is proposed to regularize the feature space to be more representative during the search process, which may be degraded by the dimension reduction. Extensive experiments demonstrate that OFB can achieve superior compression performance over state-of-the-art searching-based and pruning-based methods under various Vision Transformer architectures, meanwhile pro-moting search efficiency significantly, e.g., costing one GPU search day for the compression of DeiT-S on ImageNet-1K. Hancheng Ye, Chong Yu 0001, Peng Ye 0006, Renqiu Xia, Yansong Tang, Jiwen Lu, Tao Chen 0003, Bo Zhang 0069 |
CVPR | 7 |
| 2024 | M3DBench: Towards Omni 3D Assistant with Interleaved Multi-modal Instructions
Mingsheng Li, Xin Chen 0040, Chi Zhang 0007, Sijin Chen, Hongyuan Zhu 0002, Fukun Yin, Zhuoyuan Li 0006, Gang Yu 0002, Tao Chen 0003 |
ECCV (58) | 9 |
| 2024 | Enhanced Sparsification via Stimulative Training
Shengji Tang, Weihao Lin 0002, Hancheng Ye, Peng Ye 0006, Chong Yu 0001, Baopu Li, Tao Chen 0003 |
ECCV (51) | 7 |
| 2024 | Reg-TTA3D: Better Regression Makes Better Test-Time Adaptive 3D Object Detection
Jiakang Yuan, Bo Zhang 0069, Kaixiong Gong, Xiangyu Yue 0001, Botian Shi, Yu Qiao 0001, Tao Chen 0003 |
ECCV (43) | 7 |
| 2024 | ReSimAD: Zero-Shot 3D Domain Transfer for Autonomous Driving with Source Reconstruction and Target SimulationabstractDomain shifts such as sensor type changes and geographical situation variations are prevalent in Autonomous Driving (AD), which poses a challenge since AD model relying on the previous domain knowledge can be hardly directly deployed to a new domain without additional costs. In this paper, we provide a new perspective and approach of alleviating the domain shifts, by proposing a Reconstruction-Simulation-Perception (ReSimAD) scheme. Specifically, the implicit reconstruction process is based on the knowledge from the previous old domain, aiming to convert the domain-related knowledge into domain-invariant representations, e.g., 3D scene-level meshes. Besides, the point clouds simulation process of multiple new domains is conditioned on the above reconstructed 3D meshes, where the target-domain-like simulation samples can be obtained, thus reducing the cost of collecting and annotating new-domain data for the subsequent perception process. For experiments, we consider different cross-domain situations such as Waymo-to-KITTI, Waymo-to-nuScenes, etc, to verify the zero-shot target-domain perception using ReSimAD. Results demonstrate that our method is beneficial to boost the domain generalization ability, even promising for 3D pre-training. Code and simulated points are available at: https://github.com/PJLab-ADG/3DTrans Bo Zhang 0069, Xinyu Cai, Jiakang Yuan, Donglin Yang, Jianfei Guo, Xiangchao Yan, Renqiu Xia, Botian Shi, Min Dou, Tao Chen 0003, Si Liu 0001, Junchi Yan, Yu Qiao 0001 |
ICLR | 10 |
| 2024 | Through the Real World Haze Scenes: Navigating the Synthetic-to-Real Gap in Challenging Image DehazingabstractDehazing real-world hazy images is challenging due to the complexity of natural haze, varying haze conditions, details preservation, and the risk of overexposure. Existing methods excel in synthetic hazy scenarios but struggle in the real world because they don’t use all available features. Classical dehazing techniques primarily focus on low-level dehazing enhancements, whereas deep learning-based methods extract more intricate weather-related features. However, both of these approaches exhibit limitations in effectively addressing the real-world dehazing. To address these challenges, we introduce an innovative approach that combines the strengths of both modalities to dehaze and enhance the visibility of real-world hazy scenes. Firstly, we extract both low-level and deep features and then employ a pre-trained vector quantization GAN to create well-detailed data patches. A decoder, with a normalized module, effectively utilizes these high-quality features. Additionally, we introduce a controllable operation to improve feature matching. To further enhance dehazing and generalizability, the decoder’s output undergoes a sequence of gamma-correction operations and generates a series of multi-exposure images that are combined to create a haze-free and higher-quality image. Our method effectively reduces haziness, enhances sharpness, preserves natural colors, and minimizes artifacts in challenging real-world scenarios. The approach surpasses five SOTA methods in both qualitative and quantitative evaluations across three key metrics, utilizing three synthetic and two real-world hazy datasets. Notably, it achieves a substantial improvement in real-world datasets over the second-best method, with 0.5702 and 0.129 in FADE metrics for the RTTS and Fattal datasets, respectively. Mohammad Mahdizadeh, Chong Yu 0001, Jiayuan Fan 0001, Tao Chen 0003 |
ICRA | 5 |
| 2024 | Spear: Evaluate the Adversarial Robustness of Compressed Neural Models
Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001, Jiayuan Fan 0001 |
IJCAI | 2 |
| 2024 | G-Former: A Grouping Transformer for Weakly Supervised Point Cloud SegmentationabstractRecent advancements in weakly supervised point cloud semantic segmentation have diminished the reliance on extensive annotations, thereby enhancing the efficacy of understanding the real-world environment. However, existing approaches, such as PSD [25], SQN [5] and OTOC [10], often overlook the valuable global class-related prior knowledge present in point clouds beyond the scope of labels. To fully leverage this prior knowledge, which suggests that points of the same class should be close in feature space and each class should have a representative feature, we propose G-Former, a grouping transformer model. G-Former incorporates the idea of clustering into the overall model by defining clusters aligned to classes and assigning learnable tensors as cluster centers. Points are then grouped into these clusters based on the similarity of their features to the cluster centers. The core components of G-Former include a Hierarchy Cluster Structure (HCS) and a Grouping Module (GM). The former consists of two sets of clusters, one for classes while the other serves as a middle layer to help class clusters handle large-scale point features. The latter facilitates grouping the point cloud into different clusters. With the help of the grouping transformer model, G-Former further proposes a series of cluster center constraints to augment inter-class distances and diminish intra-class distances to enhance the discriminability of points. Experimental results on ScanNet v2 and S3DIS datasets demonstrate that G-Former outperforms previous methods with limited labels (0.1% or 1%) by a significant margin and is even comparable to fully supervised methods. Zehan Huang, Fukun Yin, Jiayuan Fan 0001, Xin Chen 0040, Hongyuan Zhu 0002, Bin Wang 0008, Tao Chen 0003 |
IJCNN | 7 |
| 2024 | Multi-dimensional Search with Strip Convolution and R-Squared Loss for Lane DetectionabstractExisting deep learning methods for lane detection mainly rely on commonly used backbones such as ResNet, equipped with various hand-craft modules for feature extraction and modeling. Such manually designed backbones and modules heavily rely on human expert experience, and it is difficult to guarantee their extracted features always fit well with various lanes. Besides backbones, the imbalanced number of lanes with different curvature in datasets and the ways modeling lane lines incur a strong curvature bias. Therefore, inspired by the recent popularity of AutoML techniques such as Neural Architecture Search (NAS), we propose an automatic lane detection architecture design framework, namely StripLaneNet-NAS, to solve the above problem. To enable the searched model structure well capture various kinds of lane features, we propose a multi-dimensional search space equipped with specially designed lane-specific Strip Convolution Modules (SCM), and correspondingly propose an adaptive solution space regularization loss, to accelerate and optimize the multi-dimensional search process. A novel R-squared loss is further proposed in the optimization objective to alleviate the curvature bias problem as mentioned above. Experiments are conducted on two benchmarks, i.e., TuSimple and CULane, and results show that our method outperforms state-of-the-art lane detection baselines, achieving the fastest inference speed with a maximum of 74.0% parameter reduction over the baselines. Peng Ye 0006, Tao Chen 0003, Shengji Tang |
IJCNN | 3 |
| 2024 | CSD3D: Cross-Scale Distillation via Dual-Consistency Learning for Semi-Supervised 3D Object DetectionabstractSemi-supervised 3D object detection has gained significant attention due to its potential to mitigate the heavy reliance on extensive annotations in traditional 3D object detection methodologies. Most existing approaches leverage the teacher’s predictions to guide and refine the student’s predictions while discarding low-confidence predictions using a fixed threshold. However, the presence of imbalanced variance in object scale poses challenges as different objects often exhibit varying levels of detection difficulty. Current methods relying on pseudo labels struggle to comprehensively capture information pertaining to objects across diverse scales. To address these challenges, we propose CSD3D, a cross-scale distillation approach via dual-consistency learning. CSD3D encompasses cross-scale distillation between the teacher and student as well as within the student itself, thereby enhancing the algorithm’s resilience to scale variance. Moreover, by adopting a dual-consistency learning paradigm that incorporates supervision at both feature and prediction levels, our approach provides comprehensive guidance to the student model. This integration of dual-consistency learning within cross-scale conditions is conducive to comprehending cross-scale object features and maintaining scale-consistent predictions. Rigorous experiments performed on the ScanNet and SUN RGB-D benchmarks reveal that CSD3D attains state-of-the-art performance. By utilizing a mere 10% of labeled data on ScanNet, we observe absolute improvements of 3.8 and 3.5 in [email protected] and [email protected], respectively. Sikai Wu, Fukun Yin, Hancheng Ye, Tao Chen 0003 |
IJCNN | 4 |
| 2024 | S2HPruner: Soft-to-Hard Distillation Bridges the Discretization Gap in PruningabstractRecently, differentiable mask pruning methods optimize the continuous relaxation architecture (soft network) as the proxy of the pruned discrete network (hard network) for superior sub-architecture search. However, due to the agnostic impact of the discretization process, the hard network struggles with the equivalent representational capacity as the soft network, namely discretization gap, which severely spoils the pruning performance. In this paper, we first investigate the discretization gap and propose a novel structural differentiable mask pruning framework named S2HPruner to bridge the discretization gap in a one-stage manner. In the training procedure, SH2Pruner forwards both the soft network and its corresponding hard network, then distills the hard network under the supervision of the soft network. To optimize the mask and prevent performance degradation, we propose a decoupled bidirectional knowledge distillation. It blocks the weight updating from the hard to the soft network while maintaining the gradient corresponding to the mask. Compared with existing pruning arts, S2HPruner achieves surpassing pruning performance without fine-tuning on comprehensive benchmarks, including CIFAR-100, Tiny ImageNet, and ImageNet with a variety of network architectures. Besides, investigation and analysis experiments explain the effectiveness of S2HPruner. Codes will be released soon. Weihao Lin 0002, Shengji Tang, Chong Yu 0001, Peng Ye 0006, Tao Chen 0003 |
NeurIPS | 5 |
| 2024 | MeshXL: Neural Coordinate Field for Generative 3D Foundation ModelsabstractThe polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, given its unstructured graph representation, the direct generation of high-fidelity 3D meshes is challenging. Fortunately, with a pre-defined ordering strategy, 3D meshes can be represented as sequences, and the generation process can be seamlessly treated as an auto-regressive problem. In this paper, we validate Neural Coordinate Field (NeurCF), an explicit coordinate representation with implicit neural embeddings, is a simple-yet-effective representation for large-scale sequential mesh modeling. After that, we present MeshXL, a family of generative pre-trained auto-regressive models that addresses 3D mesh generation with modern large language model approaches. Extensive experiments show that MeshXL is able to generate high-quality 3D meshes, and can also serve as foundation models for various down-stream applications. Sijin Chen, Xin Chen 0040, Anqi Pang, Xianfang Zeng, Yijun Fu, Fukun Yin, Billzb Wang, Jingyi Yu 0001, Gang Yu 0002, Tao Chen 0003 |
NeurIPS | 12 |
| 2024 | FNP: Fourier Neural Processes for Arbitrary-Resolution Data AssimilationabstractData assimilation is a vital component in modern global medium-range weather forecasting systems to obtain the best estimation of the atmospheric state by combining the short-term forecast and observations. Recently, AI-based data assimilation approaches have attracted increasing attention for their significant advantages over traditional techniques in terms of computational consumption. However, existing AI-based data assimilation methods can only handle observations with a specific resolution, lacking the compatibility and generalization ability to assimilate observations with other resolutions. Considering that complex real-world observations often have different resolutions, we propose the Fourier Neural Processes (FNP) for arbitrary-resolution data assimilation in this paper. Leveraging the efficiency of the designed modules and flexible structure of neural processes, FNP achieves state-of-the-art results in assimilating observations with varying resolutions, and also exhibits increasing advantages over the counterparts as the resolution and the amount of observations increase. Moreover, our FNP trained on a fixed resolution can directly handle the assimilation of observations with out-of-distribution resolutions and the observational information reconstruction task without additional fine-tuning, demonstrating its excellent generalization ability across data resolutions as well as across tasks. Code is available at https://github.com/OpenEarthLab/FNP. Kun Chen 0004, Peng Ye 0006, Hao Chen 0045, Tao Han 0002, Wanli Ouyang, Tao Chen 0003, Lei Bai 0001 |
NeurIPS | 7 |
| 2024 | EMR-Merging: Tuning-Free High-Performance Model MergingabstractThe success of pretrain-finetune paradigm brings about the release of numerous model weights. In this case, merging models finetuned on different tasks to enable a single model with multi-task capabilities is gaining increasing attention for its practicability. Existing model merging methods usually suffer from (1) significant performance degradation or (2) requiring tuning by additional data or training. In this paper, we rethink and analyze the existing model merging paradigm. We discover that using a single model's weights can hardly simulate all the models' performance. To tackle this issue, we propose Elect, Mask & Rescale-Merging (EMR-Merging). We first (a) elect a unified model from all the model weights and then (b) generate extremely lightweight task-specific modulators, including masks and rescalers, to align the direction and magnitude between the unified model and each specific model, respectively. EMR-Merging is tuning-free, thus requiring no data availability or any additional training while showing impressive performance. We find that EMR-Merging shows outstanding performance compared to existing merging methods under different classical and newly-established settings, including merging different numbers of vision models (up to 30), NLP models, PEFT models, and multi-modal models. Chenyu Huang 0001, Peng Ye 0006, Tao Chen 0003, Tong He 0001, Xiangyu Yue 0001, Wanli Ouyang |
NeurIPS | 3 |
| 2024 | Bifröst: 3D-Aware Image Compositing with Language Instructions
Kaixiong Gong, Wei-Hong Li 0001, Xili Dai, Tao Chen 0003, Xiangyu Yue 0001 |
NeurIPS | 5 |
| 2024 | 3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object DetectionabstractTransformer-based architectures have been proven successful in detecting 3D objects from point clouds. However, the quadratic complexity of the attention mechanism struggles to encode rich information as point cloud resolution increases. Recently, state space models (SSM) such as Mamba have gained great attention due to their linear complexity and long sequence modeling ability for language understanding. To exploit the potential of Mamba on 3D scene-level perception, for the first time, we propose 3DET-Mamba, which is a novel SSM-based model designed for indoor 3d object detection. Specifically, we divide the point cloud into different patches and use a lightweight yet effective Inner Mamba to capture local geometric information. To observe the scene from a global perspective, we introduce a novel Dual Mamba module that models the point cloud in terms of spatial distribution and continuity. Additionally, we design a Query-aware Mamba module that decodes context features into object sets under the guidance of learnable queries. Extensive experiments demonstrate that 3DET-Mamba surpasses previous 3DETR on indoor 3D detection benchmarks such as ScanNet, improving AP25/AP50 from 65.0\%/47.0\% to 70.4\%/54.4\%, respectively. Mingsheng Li, Jiakang Yuan, Sijin Chen, Lin Zhang 0055, Anyu Zhu, Tao Chen 0003 |
NeurIPS | 7 |
| 2024 | Training-Free Adaptive Diffusion with Bounded Difference Approximation StrategyabstractDiffusion models have recently achieved great success in the synthesis of high-quality images and videos. However, the existing denoising techniques in diffusion models are commonly based on step-by-step noise predictions, which suffers from high computation cost, resulting in a prohibitive latency for interactive applications. In this paper, we propose AdaptiveDiffusion to relieve this bottleneck by adaptively reducing the noise prediction steps during the denoising process. Our method considers the potential of skipping as many noise prediction steps as possible while keeping the final denoised results identical to the original full-step ones. Specifically, the skipping strategy is guided by the third-order latent difference that indicates the stability between timesteps during the denoising process, which benefits the reusing of previous noise prediction results. Extensive experiments on image and video diffusion models demonstrate that our method can significantly speed up the denoising process while generating identical results to the original process, achieving up to an average 2-5x speedup without quality degradation. The code is available at https://github.com/UniModal4Reasoning/AdaptiveDiffusion Hancheng Ye, Jiakang Yuan, Renqiu Xia, Xiangchao Yan, Tao Chen 0003, Junchi Yan, Botian Shi, Bo Zhang 0069 |
NeurIPS | 5 |
| 2024 | Instruct Pix-to-3D: Instructional 3D object generation from a single image
Wen Liu 0003, Wanzhang Li, Zibo Zhao 0001, Fukun Yin, Xin Chen 0018, Lei Zhao 0035, Tao Chen 0003 |
Neurocomputing | 8 |
| 2024 | Revisiting 3D visual grounding with Context-aware Feature Aggregation
Peng Guo 0011, Hongyuan Zhu 0002, Hancheng Ye, Taihao Li, Tao Chen 0003 |
Neurocomputing | 5 |
| 2024 | Vote2Cap-DETR++: Decoupling Localization and Describing for End-to-End 3D Dense Captioningabstract3D dense captioning requires a model to translate its understanding of an input 3D scene into several captions associated with different object regions. Existing methods adopt a sophisticated "detect-then-describe" pipeline, which builds explicit relation modules upon a 3D detector with numerous hand-crafted components. While these methods have achieved initial success, the cascade pipeline tends to accumulate errors because of duplicated and inaccurate box estimations and messy 3D scenes. In this paper, we first propose Vote2Cap-DETR, a simple-yet-effective transformer framework that decouples the decoding process of caption generation and object localization through parallel decoding. Moreover, we argue that object localization and description generation require different levels of scene understanding, which could be challenging for a shared set of queries to capture. To this end, we propose an advanced version, Vote2Cap-DETR++, which decouples the queries into localization and caption queries to capture task-specific features. Additionally, we introduce the iterative spatial refinement strategy to vote queries for faster convergence and better localization performance. We also insert additional spatial information to the caption head for more accurate descriptions. Without bells and whistles, extensive experiments on two commonly used datasets, ScanRefer and Nr3D, demonstrate Vote2Cap-DETR and Vote2Cap-DETR++ surpass conventional "detect-then-describe" methods by a large margin. Sijin Chen, Hongyuan Zhu 0002, Mingsheng Li, Xin Chen 0040, Peng Guo 0011, Yinjie Lei, Gang Yu 0002, Taihao Li, Tao Chen 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2024 | Latency-Aware Neural Architecture Performance Predictor With Query-to-Tier TechniqueabstractNeural Architecture Search (NAS) is a powerful tool for automating effective image and video processing DNN designing. The ranking of the accuracy has been advocated to design an efficient performance predictor for NAS. The previous contrastive method solves the ranking problem by comparing pairs of architectures and predicting their relative performance. However, it only focuses on the rankings between the two involved architectures and neglects the overall quality distributions of the search space, which may suffer generalization issues. On the contrary, we propose to let the performance predictor concentrate on the global quality level of specific architecture, and learn the tier embeddings of the whole search space automatically with learnable queries. The proposed method, dubbed as Neural Architecture Ranker with Query-to-Tier technique (NARQ2T), explores the quality tiers of the search space globally and classifies each individual to the tier they belong to. Thus, the predictor gains knowledge of the performance distributions of the search space which helps to generalize its ranking ability to the datasets more easily. Thanks to the encoder-decoder design, our method is able to predict the latency of the searched model without deteriorating the performance prediction. Meanwhile, the global quality distribution facilitates the search phase by directly sampling candidates according to the statistics of quality tiers, which is free of training a search algorithm, e.g., Reinforcement Learning or Evolutionary Algorithm, thus it simplifies the NAS pipeline and saves the computational overheads. The proposed NARQ2T achieves state-of-the-art performance on two widely used datasets for NAS research. Moreover, extensive experiments have validated the efficacy of the designed method. Bicheng Guo, Lilin Xu, Tao Chen 0003, Peng Ye 0006, Shibo He, Haoyu Liu 0002, Jiming Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | DeNKD: Decoupled Non-Target Knowledge Distillation for Complementing Transformer-Based Unsupervised Domain AdaptationabstractThere is a growing need to explore the potential of transformers in Unsupervised Domain Adaptation (UDA) due to their increasing success in various vision tasks. However, the application of transformers in UDA has yet to be thoroughly investigated and requires further research. In this study, our primary focus is to design a novel pipeline specifically tailored for transformer-based UDA, to address a crucial challenge: the overemphasis on the transfer of target-oriented information, mainly caused by the self-attention blocks in transformers and the cross-domain adversarial learning scheme. First, we show that non-target information, including semantic contextual information such as background features and non-target classes, must be addressed in the domain adaptation process. Recognizing the importance of incorporating non-target knowledge, we propose a decoupled non-target knowledge distillation method called DeNKD. DeNKD decouples non-target information across domains at both feature and logit levels. This decoupling is achieved through a bi-directional knowledge distillation approach that facilitates the interaction and exchange of non-target knowledge to facilitate an effective transformer-based cross-domain knowledge transfer. We perform extensive evaluations on several well-established UDA benchmark datasets. The results consistently show that DeNKD outperforms other methods, achieving the best performance across the board. For example, on the Office-Home dataset, DeNKD achieves an accuracy of 85.54%, while on the VisDA-2017 dataset, it achieves an accuracy of 89.95%. These results highlight the effectiveness of DeNKD in transformer-based UDA and its potential for improving cross-domain adaptation performance. Peng Ye 0006, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Multi-View Vision Fusion Network: Can 2D Pre-Trained Model Boost 3D Point Cloud Data-Scarce Learning?abstractPoint cloud based 3D deep model has wide applications in many applications such as autonomous driving, house robot, etc. Inspired by the recent prompt learning in natural language processing, this work proposes a novel Multi-view Vision Fusion Network (MvNet) for few-shot 3D point cloud classification. MvNet investigates the possibility of leveraging the off-the-shelf 2D pre-trained models to achieve the few-shot classification, which can alleviate the over-dependence issue of the existing baseline models towards the large-scale annotated 3D point cloud data. Specifically, MvNet first encodes a 3D point cloud into multi-view image features for a number of different views. Then, a novel multi-view prompt fusion module is developed to fuse information from different views effectively to bridge the gap between 3D point cloud data and 2D pre-trained models. A set of 2D image prompts can then be derived to better describe the suitable prior knowledge for a large-scale pre-trained image model for few-shot 3D point cloud classification. Extensive experiments on ModelNet, ScanObjectNN, and ShapeNet datasets demonstrate that MvNet achieves new state-of-the-art performance for 3D few-shot point cloud image classification. The source code of this work is available at https://github.com/invictus717/MetaTransformer. Haoyang Peng, Baopu Li, Bo Zhang 0069, Xin Chen 0040, Tao Chen 0003, Hongyuan Zhu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Push-and-Pull: A General Training Framework With Differential Augmentor for Domain Generalized Point Cloud ClassificationabstractAs a fundamental task of 3D perception, point cloud recognition has shown significant progress in recent years. However, existing methods still face challenges when dealing with geometry differences, resulting in performance degradation when a distribution gap exists between the training and testing data, also known as domain generalization. In this work, we focus on this problem and propose a general training framework, named Push-and-Pull, aimed at effectively improving the generalization ability of models on unseen target domains. Specifically, our framework first introduces a learnable 3D data augmentor to generate new training point clouds, which helps to reduce the domain bias and enrich the source training set. Also, an adversarial training strategy is proposed topushthe augmented samples away from the original ones in the latent space and meanwhile keep the geometric structure. Second, based on the original and augmented samples, a dual-level consistency regularization strategy on logits and feature spaces is designed topullthe deviated representations back to their original space as close as possible, and promote discriminative and domain-agnostic representations. These two steps are iteratively optimized to enhance the overall performance. Extensive experiments on the PointDA-10 and Sim2Real benchmarks consistently demonstrate the effectiveness of our proposed framework. Xinzhu Ma, Lin Zhang 0055, Bo Zhang 0069, Tao Chen 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Few-Shot Cross-Domain Object Detection With Instance-Level Prototype-Based Meta-LearningabstractIn typical unsupervised domain adaptive object detection, it is assumed that extensive unlabeled training data from the target domain can be easily obtained. However, in some access-constrained scenarios, massive target data cannot be guaranteed, but acquiring only a few target samples and annotating them may costs less. Therefore, inspired by the meta-learning success in few-shot tasks, we propose an Instance-level Prototype learning Network (IPNet) for solving the domain adaptive object detection under the supervised few-shot scenario in this work. To compensate for the target domain data deficiency, we fuse cropped instances from labeled images in both domains to learn a representative prototype for each class, by enforcing features of the same class’s instances but from different domains to be as close as possible. These prototypes are further employed to discriminate various features’ salience in an image, and separate foreground and background regions for respective domain alignment. Extensive experiments are conducted on several cross-domain scenarios, and their results show the consistent accuracy gains of the IPNet over state-of-the-art methods, e.g., 10.4% mAP increase on Cityscapes-to-FoggyCityscapes setting and 3.0% mAP increase on Sim10k-to-Cityscapes setting. Lin Zhang 0055, Bo Zhang 0069, Botian Shi, Jiayuan Fan 0001, Tao Chen 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Joint Distribution Adaptive-Alignment for Cross-Domain Segmentation of High-Resolution Remote Sensing ImagesabstractAlthough existing unsupervised domain adaptation (UDA) methods have successfully applied to semantic segmentation tasks for high-resolution remote sensing (HRS) images, they still have some limitations that need to be addressed: 1) they mainly focus on aligning the marginal distributions while ignoring the interdomain differences in the conditional distributions, which may be suboptimal because they assume that the boundaries of category decision are identical across domains; and 2) they depend on self-supervised learning for easy-to-hard alignment, which may result in model learning erroneous knowledge from the pseudo labels. To address the above limitations, we propose a joint distribution adaptive-alignment framework (JDAF) to eliminate the distribution difference between the source and target domains, which is mainly composed of a marginal distribution alignment (MDA) module, a conditional distribution alignment (CDA) module, and an improved easy-to-hard adaptation strategy. The MDA module is used to narrow local semantic and global spatial differences between domains and first advance, and then, the CDA module that includes a category-invariant feature alignment (CFA) block and a dataset-level context aggregation (DCA) block is presented and designed, which can dynamically update and align the feature representations that are invariant to category change and adaptively incorporate dataset-level context into the features of source domain to enhance the pixel-level representation. An uncertainty-adaptive learning (UAL) method is, moreover, proposed to improve the easy-to-hard adaptation strategy by enabling the model to learn accurate knowledge from the pseudo labels, which can boost the adaptive performance of the whole JDAF. Comprehensive experiments with four cross-domain tasks on two benchmark datasets of aerospace HRS images demonstrate that the proposed JDAF achieves significant performance gains compared to the state-of-the-art cross-domain semantic segmentation methods. Our code is available at:https://github.com/maple-hx/JDAF. Baopu Li, Tao Chen 0003, Bin Wang 0008 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | U²ConvFormer: Marrying and Evolving Nested U-Net and Scale-Aware Transformer for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification plays an important role in the human exploration of the Earth. Recent research of deep learning-based HSI classification has been fast-growing, but still suffers from three obstacles: First, existing deep learning-based HSI works lack of extraction and utilization of multigrained multiscale information and multiscale local-to-global information. Second, most previous works have too fixed-sized receptive fields in their convolutional network parts to handle HSI classification problems, and pay no attention to the existence of asymmetries in the spectral-spatial dimension of the HSI data. Third, most networks for HSI classification are hand-craft. To this end, we propose a novel architecture in this article, which is the first to combine the advantages of nested U-Net and scale-aware Transformer, named U2ConvFormer. Specifically, the nested U-Net structure can fully extract and aggregate multiscale spectral-spatial features at both inter- and inner stage granularity. The scale-aware Transformer takes multiscale local spectral-spatial features from the encoder of nested U-Net and produces multiscale global spectral-spatial features for its decoder. After that, we design a novel plug-and-play searchable operation called asymmetric spectral-spatial convolution (A2SConv), where asymmetric spectral-spatial feature pooling and multiscale feature extraction can be concurrently searched. Furthermore, we develop a customized search strategy to automatically design U2ConvFormer, which uses advanced neural architecture search (NAS) methods to enable the customization of suitable models for different hyperspectral datasets. Experimental results on three benchmark datasets, including Indian Pines, Pavia University and Houston University 2018, validate the superiority of our proposed U2ConvFormer, which achieves new state-of-the-art performance across different benchmark datasets. Lin Zhan, Peng Ye 0006, Jiayuan Fan 0001, Tao Chen 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Exploring Multi-Timestep Multi-Stage Diffusion Features for Hyperspectral Image ClassificationabstractThe effectiveness of spectral-spatial feature learning is crucial for the hyperspectral image (HSI) classification task. Diffusion models, as a new class of groundbreaking generative models, have the ability to learn both contextual semantics and textual details from the distinct timestep dimension, enabling the modeling of complex spectral-spatial relations in HSIs. However, existing diffusion-based HSI classification methods only utilize manually selected single-timestep single-stage features, limiting the full exploration and exploitation of rich contextual semantics and textual information hidden in the diffusion model. To address this issue, we propose a novel diffusion-based feature learning framework that explores Multi-Timestep Multi-Stage Diffusion features for HSI classification for the first time, called MTMSD. Specifically, the diffusion model is first pretrained with unlabeled HSI patches to mine the connotation of unlabeled data, and then is used to extract the multi-timestep multi-stage diffusion features. To effectively and efficiently leverage multi-timestep multi-stage features, two strategies are further developed. One strategy is class & timestep-oriented multi-stage feature purification module with the inter-class and inter-timestep prior for reducing the redundancy of multi-stage features and alleviating memory constraints. The other one is selective timestep feature fusion module with the guidance of global features to adaptively select different timestep features for integrating texture and semantics. Both strategies facilitate the generality and adaptability of the MTMSD framework for diverse patterns of different HSI data. Extensive experiments are conducted on four public HSI datasets, and the results demonstrate that our method outperforms state-of-the-art methods for HSI classification, especially on the challenging Houston 2018 dataset. The codes are available at https://github.com/zjyaccount/MTMSD. Jiamu Sheng, Peng Ye 0006, Jiayuan Fan 0001, Tong He 0001, Bin Wang 0008, Tao Chen 0003 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Lightweight Model Pre-Training via Language Guided Knowledge DistillationabstractThis paper studies the problem of pre-training for small models, which is essential for many mobile devices. Current state-of-the-art methods on this problem transfer the representational knowledge of a large network (as a Teacher) into a smaller model (as a Student) using self-supervised distillation, improving the performance of the small model on downstream tasks. However, existing approaches are insufficient in extracting the crucial knowledge that is useful for discerning categories in downstream tasks during the distillation process. In this paper, for the first time, we introduce language guidance to the distillation process and propose a new method named Language-Guided Distillation (LGD) system, which uses category names of the target downstream task to help refine the knowledge transferred between the teacher and student. To this end, we utilize a pre-trained text encoder to extract semantic embeddings from language and construct a textual semantic space called Textual Semantics Bank (TSB). Furthermore, we design a Language-Guided Knowledge Aggregation (LGKA) module to construct the visual semantic space, also named Visual Semantics Bank (VSB). The task-related knowledge is transferred by driving a student encoder to mimic the similarity score distribution inferred by a teacher over TSB and VSB. Compared with other small models obtained by either ImageNet pre-training or self-supervised distillation, experiment results show that the distilled lightweight model using the proposed LGD method presents state-of-the-art performance and is validated on various downstream tasks, including classification, detection, and segmentation. Mingsheng Li, Lin Zhang 0055, Mingzhen Zhu, Gang Yu 0002, Jiayuan Fan 0001, Tao Chen 0003 |
IEEE Trans. Multim. | 7 |
| 2023 | Hyperscale Hardware Optimized Neural Architecture SearchabstractRecent advances in machine learning have leveraged dramatic increases in computational power, a trend expected to continue in the future. This paper introduces the first Hyperscale Hardware Optimized Neural Architecture Search (H2O-NAS) to automatically design accurate and performant machine learning models tailored to the underlying hardware architecture. H2O-NAS consists of three key components: a new massively parallel “one-shot” search algorithm with intelligent weight sharing, which can scale to search spaces of O(10280) and handle large volumes of production traffic; hardware-optimized search spaces for diverse ML models on heterogeneous hardware; and a novel two-phase hybrid performance model and a multi-objective reward function optimized for large scale deployments. Sheng Li 0007, Garrett Andersen, Tao Chen 0003, Liqun Cheng, Julian Grady, Quoc V. Le, Andrew Li, Xin Li 0082, Yang Li 0005, Yifeng Lu, Yun Ni, Ruoming Pang, Mingxing Tan, Martin Wicke, Shengqi Zhu 0003, Parthasarathy Ranganathan, Norman P. Jouppi |
ASPLOS (3) | 3 |
| 2023 | Executing your Commands via Motion Diffusion in Latent SpaceabstractWe study a challenging task, conditional human motion generation, which produces plausible human motion sequences according to various conditional inputs, such as action classes or textual descriptors. Since human motions are highly diverse and have a property of quite different distribution from conditional modalities, such as textual descriptors in natural languages, it is hard to learn a probabilistic mapping from the desired conditional modality to the human motion sequences. Besides, the raw motion data from the motion capture system might be redundant in sequences and contain noises; directly modeling the joint distribution over the raw motion sequences and conditional modalities would need a heavy computational over-head and might result in artifacts introduced by the captured noises. To learn a better representation of the various human motion sequences, we first design a powerful Variational AutoEncoder (VAE) and arrive at a representative and low-dimensional latent code for a human motion sequence. Then, instead of using a diffusion model to establish the connections between the raw motion sequences and the conditional inputs, we perform a diffusion process on the motion latent space. Our proposed Motion Latent-based Diffusion model (MLD) could produce vivid motion sequences conforming to the given conditional inputs and substantially reduce the computational overhead in both the training and inference stages. Extensive experiments on various human motion generation tasks demonstrate that our MLD achieves significant improvements over the state-of-the-art methods among extensive human motion generation tasks, with two orders of magnitude faster than previous diffusion models on raw motion sequences. Xin Chen 0040, Biao Jiang, Wen Liu 0003, Tao Chen 0003, Gang Yu 0002 |
CVPR | 6 |
| 2023 | End-to-End 3D Dense Captioning with Vote2Cap-DETRabstract3D dense captioning aims to generate multiple captions localized with their associated object regions. Existing methods follow a sophisticated “detect-then-describe” pipeline equipped with numerous hand-crafted components. However, these hand-crafted components would yield sub-optimal performance given cluttered object spatial and class distributions among different scenes. In this paper, we propose a simple-yet-effective transformer framework Vote2Cap-DETR based on recent popular DEtection TRansformer (DETR). Compared with prior arts, our framework has several appealing advantages: 1) Without resorting to numerous hand-crafted components, our method is based on a full transformer encoder-decoder architecture with a learnable vote query driven object decoder, and a caption decoder that produces the dense captions in a set-prediction manner. 2) In contrast to the two-stage scheme, our method can perform detection and captioning in one-stage. 3) Without bells and whistles, extensive experiments on two commonly used datasets, ScanRefer and Nr3D, demonstrate that our Vote2Cap-DETR surpasses current state-of-the-arts by 11.13% and 7.11% in [email protected], respectively. Codes will be released soon. Sijin Chen, Hongyuan Zhu 0002, Xin Chen 0040, Yinjie Lei, Gang Yu 0002, Tao Chen 0003 |
CVPR | 6 |
| 2023 | Boost Vision Transformer with GPU-Friendly Sparsity and QuantizationabstractThe transformer extends its success from the language to the vision domain. Because of the stacked self-attention and cross-attention blocks, the acceleration deployment of vision transformer on GPU hardware is challenging and also rarely studied. This paper thoroughly designs a compression scheme to maximally utilize the GPU-friendly 2:4 fine-grained structured sparsity and quantization. Specially, an original large model with dense weight parameters is first pruned into a sparse one by 2:4 structured pruning, which considers the GPU's acceleration of 2:4 structured sparse pattern with FP16 data type, then the floating-point sparse model is further quantized into a fixed-point one by sparse-distillation-aware quantization aware training, which considers GPU can provide an extra speedup of 2:4 sparse calculation with integer tensors. A mixed-strategy knowledge distillation is used during the pruning and quantization process. The proposed compression scheme is flexible to support supervised and unsupervised learning styles. Experiment results show GPUSQ-ViT scheme achieves state-of-the-art compression by reducing vision transformer models$\mathbf{6.4}-\mathbf{12.7}\times$on model size and$\mathbf{30.3}-\mathbf{62} \times$on FLOPs with negligible accuracy degradation on ImageNet classification, COCO detection and ADE20K segmentation benchmarking tasks. Moreover, GPUSQ-ViT can boost actual deployment performance by$\mathbf{1.39}-\mathbf{1.79}\times$and$\mathbf{3.22}-\mathbf{3.43}\times$of latency and throughput on A100 GPU, and$\mathbf{1.57}-\mathbf{1.69}\times$and$\mathbf{2.11}-\mathbf{2.51}\times$improvement of latency and throughput on AGX Orin. Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001, Jiayuan Fan 0001 |
CVPR | 2 |
| 2023 | Bi3D: Bi-Domain Active Learning for Cross-Domain 3D Object DetectionabstractUnsupervised Domain Adaptation (UDA) technique has been explored in 3D cross-domain tasks recently. Though preliminary progress has been made, the performance gap between the UDA-based 3D model and the supervised one trained with fully annotated target domain is still large. This motivates us to consider selecting partial-yet-important target data and labeling them at a minimum cost, to achieve a good trade-off between high performance and low annotation cost. To this end, we propose a Bi-domain active learning approach, namely Bi3D, to solve the cross-domain 3D object detection task. The Bi3D first develops a domainness-aware source sampling strategy, which identifies target-domain-like samples from the source domain to avoid the model being interfered by irrelevant source data. Then a diversity-based target sampling strategy is developed, which selects the most informative subset of target domain to improve the model adaptability to the target domain using as little annotation budget as possible. Experiments are conducted on typical cross-domain adaptation scenarios including cross-LiDAR-beam, cross-country, and cross-sensor, where Bi3D achieves a promising target-domain detection accuracy (89.63% on KITTI) compared with UDA-based work (84.29%), even surpassing the detector trained on the full set of the labeled target domain (88.98%). Our code is available at: https://github.com/PJLab-ADG/3DTrans. Jiakang Yuan, Bo Zhang 0069, Xiangchao Yan, Tao Chen 0003, Botian Shi, Yikang Li 0002, Yu Qiao 0001 |
CVPR | 4 |
| 2023 | Uni3D: A Unified Baseline for Multi-Dataset 3D Object DetectionabstractCurrent 3D object detection models follow a single dataset-specific training and testing paradigm, which often faces a serious detection accuracy drop when they are directly deployed in another dataset. In this paper, we study the task of training a unified 3D detector from multiple datasets. We observe that this appears to be a challenging task, which is mainly due to that these datasets present substantial data-level differences and taxonomy-level variations caused by different LiDAR types and data acquisition standards. Inspired by such observation, we present a Uni3D which leverages a simple data-level correction operation and a designed semantic-level coupling-and-recoupling module to alleviate the unavoidable data-level and taxonomy-level differences, respectively. Our method is simple and easily combined with many 3D object detection baselines such as PV-RCNN and Voxel-RCNN, enabling them to effectively learn from multiple off-the-shelf 3D datasets to obtain more discriminative and generalizable representations. Experiments are conducted on many dataset consolidation settings. Their results demonstrate that Uni3D exceeds a series of individual detectors trained on a single dataset, with a 1.04× parameter increase over a selected baseline detector. We expect this work will inspire the research of 3D generalization since it will push the limits of perceptual performance. Our code is available at: https://github.com/PJLab-ADG/3DTrans. Bo Zhang 0069, Jiakang Yuan, Botian Shi, Tao Chen 0003, Yikang Li 0002, Yu Qiao 0001 |
CVPR | 4 |
| 2023 | A Large-Scale Outdoor Multi-modal Dataset and Benchmark for Novel View Synthesis and Implicit Scene ReconstructionabstractNeural Radiance Fields (NeRF) [24] has achieved impressive results in single object scene reconstruction and novel view synthesis, as demonstrated on many single modality and single object focused indoor scene datasets like DTU [14], BMVS [42], and NeRF Synthetic [24]. However, the study of NeRF on large-scale outdoor scene reconstruction is still limited, as there is no unified outdoor scene dataset for large-scale NeRF evaluation due to expensive data acquisition and calibration costs. In this work, we propose a large-scale outdoor multi-modal dataset, OMMO dataset, containing complex objects and scenes with calibrated images, point clouds and prompt annotations. A new benchmark for several outdoor NeRF-based tasks is established, such as novel view synthesis, diverse 3D representation, and multi-modal NeRF. To create the dataset, we capture and collect a large number of real fly-view videos and select high-quality and high-resolution clips from them. Then we design a quality review module to refine images, remove low-quality frames and fail-to-calibrate scenes through a learning-based automatic evaluation plus manual review. Finally, volunteers are employed to label and review the prompt annotation for each scene and keyframe. Compared with existing NeRF datasets, our dataset contains abundant real-world urban and natural scenes with various scales, camera trajectories, and lighting conditions. Experiments show that our dataset can benchmark most state-of-the-art NeRF methods on different tasks. The dataset can be found at the following link: https://ommo.luchongshan.com/. Chongshan Lu, Fukun Yin, Xin Chen 0040, Wen Liu 0003, Tao Chen 0003, Gang Yu 0002, Jiayuan Fan 0001 |
ICCV | 5 |
| 2023 | Robust Geometry-Preserving Depth Estimation Using Differentiable RenderingabstractIn this study, we address the challenge of 3D scene structure recovery from monocular depth estimation. While traditional depth estimation methods leverage labeled datasets to directly predict absolute depth, recent advancements advocate for mix-dataset training, enhancing generalization across diverse scenes. However, such mixed dataset training yields depth predictions only up to an unknown scale and shift, hindering accurate 3D reconstructions. Existing solutions necessitate extra 3D datasets or geometry-complete depth annotations, constraints that limit their versatility. In this paper, we propose a learning framework that trains models to predict geometry-preserving depth without requiring extra data or annotations. To produce realistic 3D structures, we render novel views of the reconstructed scenes and design loss functions to promote depth estimation consistency across different views. Comprehensive experiments underscore our framework’s superior generalization capabilities, surpassing existing state-of-the-art methods on several benchmark datasets without leveraging extra training information. Moreover, our innovative loss functions empower the model to autonomously recover domain-specific scale-and-shift coefficients using solely unlabeled images. Chi Zhang 0007, Wei Yin 0006, Gang Yu 0002, Zhibin Wang 0004, Tao Chen 0003, Joey Tianyi Zhou, Chunhua Shen |
ICCV | 5 |
| 2023 | Adversarial Amendment is the Only Force Capable of Transforming an Enemy into a FriendabstractAdversarial attack is commonly regarded as a huge threat to neural networks because of misleading behavior. This paper presents an opposite perspective: adversarial attacks can be harnessed to improve neural models if amended correctly. Unlike traditional adversarial defense or adversarial training schemes that aim to improve the adversarial robustness, the proposed adversarial amendment (AdvAmd) method aims to improve the original accuracy level of neural models on benign samples. We thoroughly analyze the distribution mismatch between the benign and adversarial samples. This distribution mismatch and the mutual learning mechanism with the same learning ratio applied in prior art defense strategies is the main cause leading the accuracy degradation for benign samples. The proposed AdvAmd is demonstrated to steadily heal the accuracy degradation and even leads to a certain accuracy boost of common neural models on benign classification, object detection, and segmentation tasks. The efficacy of the AdvAmd is contributed by three key components: mediate samples (to reduce the influence of distribution mismatch with a fine-grained amendment), auxiliary batch norm (to solve the mutual learning mechanism and the smoother judgment surface), and AdvAmd loss (to adjust the learning ratios according to different attack vulnerabilities) through quantitative and ablation experiments. Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001 |
IJCAI | 2 |
| 2023 | RBGNet: Reliable Boundary-Guided Segmentation of Choroidal Neovascularization
Tao Chen 0003, Yitian Zhao, Lei Mou, Dan Zhang 0026, Xiayu Xu, Huazhu Fu, Jiong Zhang 0004 |
MICCAI (4) | 1 |
| 2023 | Rethinking Pseudo-Label-Based Unsupervised Person Re-ID with Hierarchical Prototype-based GraphabstractUnsupervised person re-identification (Re-ID) aims to match individuals without manual annotations. However, existing methods often struggle with intra-class variations due to differences in person poses and camera styles such as resolution and environment information. Additionally, clustering may produce incorrect pseudo-labels, compounding the issue. To address these challenges, we propose a novel hierarchical prototype-based graph network (HPG-Net) for unsupervised person Re-ID. Our approach uses a hierarchical prototype-based graph structure to describe person images by attributes of poses and camera styles, with each graph node representing the average of image features as a prototype. We then apply a hierarchical contrastive learning module to enhance the feature learning at each level, reducing the impact of intra-class differences caused by extraneous attributes. We also calculate the similarity between samples and each level of prototypes, maintaining prototype-based graph consistency with the mean-teacher network to mitigate the accumulation errors caused by pseudo-labels. Experimental results on three benchmarks show that our method outperforms state-of-the-art (SOTA) works. Moreover, we achieve promising performance on an occluded dataset. Ben Sha, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001 |
ACM Multimedia | 3 |
| 2023 | PDF: Point Diffusion Implicit Function for Large-scale Scene Neural RepresentationabstractRecent advances in implicit neural representations have achieved impressive results by sampling and fusing individual points along sampling rays in the sampling space. However, due to the explosively growing sampling space, finely representing and synthesizing detailed textures remains a challenge for unbounded large-scale outdoor scenes. To alleviate the dilemma of using individual points to perceive the entire colossal space, we explore learning the surface distribution of the scene to provide structural priors and reduce the samplable space and propose a Point Diffusion implicit Function, PDF, for large-scale scene neural representation. The core of our method is a large-scale point cloud super-resolution diffusion module that enhances the sparse point cloud reconstructed from several training images into a dense point cloud as an explicit prior. Then in the rendering stage, only sampling points with prior points within the sampling radius are retained. That is, the sampling space is reduced from the unbounded space to the scene surface. Meanwhile, to fill in the background of the scene that cannot be provided by point clouds, the region sampling based on Mip-NeRF 360 is employed to model the background representation. Expensive experiments have demonstrated the effectiveness of our method for large-scale scene novel view synthesis, which outperforms relevant state-of-the-art baselines. Yuhan Ding, Fukun Yin, Jiayuan Fan 0001, Xin Chen 0040, Wen Liu 0003, Chongshan Lu, Gang Yu 0002, Tao Chen 0003 |
NeurIPS | 9 |
| 2023 | MotionGPT: Human Motion as a Foreign LanguageabstractThough the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multimodal data, such as motion, remains challenging and untouched so far. Fortunately, human motion displays a semantic coupling akin to human language, often perceived as a form of body language. By fusing language data with large-scale motion models, motion-language pre-training that can enhance the performance of motion-related tasks becomes feasible. Driven by this insight, we propose MotionGPT, a unified, versatile, and user-friendly motion-language model to handle multiple motion-relevant tasks. Specifically, we employ the discrete vector quantization for human motion and transfer 3D motion into motion tokens, similar to the generation process of word tokens. Building upon this "motion vocabulary", we perform language modeling on both motion and text in a unified manner, treating human motion as a specific language. Moreover, inspired by prompt learning, we pre-train MotionGPT with a mixture of motion-language data and fine-tune it on prompt-based question-and-answer tasks. Extensive experiments demonstrate that MotionGPT achieves state-of-the-art performances on multiple motion tasks including text-driven motion generation, motion captioning, motion prediction, and motion in-between. Biao Jiang, Xin Chen 0040, Wen Liu 0003, Jingyi Yu 0001, Gang Yu 0002, Tao Chen 0003 |
NeurIPS | 6 |
| 2023 | AD-PT: Autonomous Driving Pre-Training with Large-scale Point Cloud DatasetabstractIt is a long-term vision for Autonomous Driving (AD) community that the perception models can learn from a large-scale point cloud dataset, to obtain unified representations that can achieve promising results on different tasks or benchmarks. Previous works mainly focus on the self-supervised pre-training pipeline, meaning that they perform the pre-training and fine-tuning on the same benchmark, which is difficult to attain the performance scalability and cross-dataset application for the pre-training checkpoint. In this paper, for the first time, we are committed to building a large-scale pre-training point-cloud dataset with diverse data distribution, and meanwhile learning generalizable representations from such a diverse pre-training dataset. We formulate the point-cloud pre-training task as a semi-supervised problem, which leverages the few-shot labeled and massive unlabeled point-cloud data to generate the unified backbone representations that can be directly applied to many baseline models and benchmarks, decoupling the AD-related pre-training process and downstream fine-tuning task. During the period of backbone pre-training, by enhancing the scene- and instance-level distribution diversity and exploiting the backbone's ability to learn from unknown instances, we achieve significant performance gains on a series of downstream perception benchmarks including Waymo, nuScenes, and KITTI, under different baseline models like PV-RCNN++, SECOND, CenterPoint. Jiakang Yuan, Bo Zhang 0069, Xiangchao Yan, Botian Shi, Tao Chen 0003, Yikang Li 0002, Yu Qiao 0001 |
NeurIPS | 5 |
| 2023 | Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationabstractWe present a novel alignment-before-generation approach to tackle the challenging task of generating general 3D shapes based on 2D images or texts. Directly learning a conditional generative model from images or texts to 3D shapes is prone to producing inconsistent results with the conditions because 3D shapes have an additional dimension whose distribution significantly differs from that of 2D images and texts. To bridge the domain gap among the three modalities and facilitate multi-modal-conditioned 3D shape generation, we explore representing 3D shapes in a shape-image-text-aligned space. Our framework comprises two models: a Shape-Image-Text-Aligned Variational Auto-Encoder (SITA-VAE) and a conditional Aligned Shape Latent Diffusion Model (ASLDM). The former model encodes the 3D shapes into the shape latent space aligned to the image and text and reconstructs the fine-grained 3D neural fields corresponding to given shape embeddings via the transformer-based decoder. The latter model learns a probabilistic mapping function from the image or text space to the latent shape space. Our extensive experiments demonstrate that our proposed approach can generate higher-quality and more diverse 3D shapes that better semantically conform to the visual or textural conditional inputs, validating the effectiveness of the shape-image-text-aligned space for cross-modality 3D shape generation. Zibo Zhao 0001, Wen Liu 0003, Xin Chen 0040, Xianfang Zeng, Rui Wang 0099, Tao Chen 0003, Gang Yu 0002, Shenghua Gao |
NeurIPS | 8 |
| 2023 | A Closer Look at Few-Shot 3D Point Cloud Classification
Chuangguan Ye, Hongyuan Zhu 0002, Bo Zhang 0069, Tao Chen 0003 |
Int. J. Comput. Vis. | 4 |
| 2023 | Performance-Aware Approximation of Global Channel Pruning for Multitask CNNsabstractGlobal channel pruning (GCP) aims to remove a subset of channels (filters) across different layers from a deep model without hurting the performance. Previous works focus on either single task model pruning or simply adapting it to multitask scenario, and still face the following problems when handling multitask pruning: 1) Due to the task mismatch, a well-pruned backbone for classification task focuses on preserving filters that can extract category-sensitive information, causing filters that may be useful for other tasks to be pruned during the backbone pruning stage; 2) For multitask predictions, different filters within or between layers are more closely related and interacted than that for single task prediction, making multitask pruning more difficult. Therefore, aiming at multitask model compression, we propose a Performance-Aware Global Channel Pruning (PAGCP) framework. We first theoretically present the objective for achieving superior GCP, by considering the joint saliency of filters from intra- and inter-layers. Then a sequentially greedy pruning strategy is proposed to optimize the objective, where a performance-aware oracle criterion is developed to evaluate sensitivity of filters to each task and preserve the globally most task-related filters. Experiments on several multitask datasets show that the proposed PAGCP can reduce the FLOPs and parameters by over 60% with minor performance drop, and achieves 1.2x ∼ 3.3x acceleration on both cloud and mobile platforms. Our code is available at http://www.github.com/HankYe/PAGCP.git. Hancheng Ye, Bo Zhang 0069, Tao Chen 0003, Jiayuan Fan 0001, Bin Wang 0008 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Automatic Loss Function Search for Adversarial Unsupervised Domain AdaptationabstractUnsupervised domain adaption (UDA) aims to reduce the domain gap between labeled source and unlabeled target domains. Many prior works exploit adversarial learning that leverages pre-designed discriminators to drive the network for aligning distributions between domains. However, most of them do not consider the degeneration of the domain discriminators caused by the gradually dominating gradients of aligned target samples during training, and they still suffer from the cross-domain semantic mismatch problem in the learned feature space. Hence, this paper attempts to understand and solve both issues from the lens of optimization loss and propose an automatic loss function search for adversarial domain adaptation (ALSDA). First, we extend the common adversarial loss by adding an adjustable hyper-parameter that can re-weight the gradients assigned to target samples, so that the domain discriminator can impose consecutive and influential driving forces for domain alignment. Meanwhile, we upgrade the traditional orthogonality loss with class-wisely adjustable hyper-parameters that can strengthen the cross-domain feature separation. Since manually determining the optimal loss functions requires expensive expert efforts, we leverage the popular AutoML to automatically search for the optimal loss functions from a pre-defined novel and unique search space for UDA. Further, to enable the loss function search when the target domain is unlabeled, we introduce a simple-but-effective entropy-guided search strategy with the aid of REINFORCE learning. Extensive experiments on various typical baselines and benchmark datasets such as Office-Home, Office-31, and Birds-31 have been conducted, and the results validate the generalization and superiority of the proposed ALSDA. Peng Ye 0006, Hancheng Ye, Baopu Li, Jinyang Guo 0002, Tao Chen 0003, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | DCNet: Large-Scale Point Cloud Semantic Segmentation With Discriminative and Efficient Feature AggregationabstractThe point cloud feature aggregation, which learns discriminative features from the disordered points, plays a key role for large-scale point cloud semantic segmentation. Most previous aggregation methods are based on sampling a representative point subset, i.e., by a carefully designed point density metric, facing expensive computation cost especially for large-scale point clouds. Even though speeding up the point sampling process is studied by several recent works, but the component points in the sampled subset are uncertain and may change randomly, thus leading to corrupted geometric structure and discarded edges for representing an object. Therefore, we propose the DCNet, which consists of a fast point random sampling based encoder-decoder structure and several fully connected layers for semantic segmentation. To overcome the key feature loss caused by random down-sampling, the DCNet develops two novel local feature aggregation schemes: Double attention and Consistent constraints, to learn features that are discriminative for the challenging scenarios as above. The former considers both topological and semantic similarity of neighboring points to generate attention features for discriminating classes with similar geometric structures. The latter develops class-consistent constraints between adjacent layers in the decoder stage, to guide each point to aggregate with high-level semantic features of points belonging to the same class from the previous layer, which is beneficial for distinguishing neighboring points of the same class on the boundary. We conduct experiments and compare the proposed DCNet with existing methods on two benchmarks S3DIS and Semantic3D. Experiments show that the mean Intersection-over-Union (mIoU) of our method outperforms state-of-the-art methods by 2-3%, based on the same fast random sampling, and is also comparable to latest sampling-slower but accuracy-higher methods. That is, our method achieves the optimal speed-accuracy trade-off in the field of point cloud segmentation. Fukun Yin, Tao Chen 0003, Guozhong Luo, Gang Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Pull & Push: Leveraging Differential Knowledge Distillation for Efficient Unsupervised Anomaly Detection and LocalizationabstractRecently, much attention has been paid to segmenting subtle unknown defect regions by knowledge distillation in an unsupervised setting. Most previous studies concentrated on guiding the student network to learn the same representations on the normality, neglecting the different behaviors of the abnormality. This leads to a high probability of false detection of subtle defects. To address such an issue, we propose to push representations on abnormal areas of the teacher and student network as far as possible while pulling representations on normal areas as close as possible. Based on this idea, we design an efficient teacher-student model for anomaly detection and localization, which maximizes pixel-wise discrepancies for anomalous regions approximated by data augmentation and simultaneously minimizes discrepancies for pixel-wise normal regions between these two networks. The explicit differential knowledge distillation enlarges the margin between normal representations and abnormal ones in favour of discriminating them. Then, the appropriate small student network is not only efficient, but more importantly, helps inhibit the generalization ability of anomalous patterns when learning normal patterns, facilitating the precise decision boundary. The experimental results on the MVTec AD, Fashion-MNIST, and CIFAR-10 datasets demonstrate that our proposed method achieves better performance than current state-of-the-art (SOTA) approaches. Especially, For the MVTec AD dataset with high resolution images, we achieve 98.1 AUROC% and 93.6 AUPRO% in anomaly localization, outperforming knowledge distillation based SOTA methods by 1.1 AUROC% and 1.5 AUPRO% with a lightweight model. Qihang Zhou, Shibo He, Haoyu Liu 0002, Tao Chen 0003, Jiming Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Rethinking Cross-Domain Pedestrian Detection: A Background-Focused Distribution Alignment Framework for Instance-Free One-Stage DetectorsabstractCross-domain pedestrian detection aims to generalize pedestrian detectors from one label-rich domain to another label-scarce domain, which is crucial for various real-world applications. Most recent works focus on domain alignment to train domain-adaptive detectors either at the instance level or image level. From a practical point of view, one-stage detectors are faster. Therefore, we concentrate on designing a cross-domain algorithm for rapid one-stage detectors that lacks instance-level proposals and can only perform image-level feature alignment. However, pure image-level feature alignment causes the foreground-background misalignment issue to arise, i.e., the foreground features in the source domain image are falsely aligned with background features in the target domain image. To address this issue, we systematically analyze the importance of foreground and background in image-level cross-domain alignment, and learn that background plays a more critical role in image-level cross-domain alignment. Therefore, we focus on cross-domain background feature alignment while minimizing the influence of foreground features on the cross-domain alignment stage. This paper proposes a novel framework, namely, background-focused distribution alignment (BFDA), to train domain adaptive one-stage pedestrian detectors. Specifically, BFDA first decouples the background features from the whole image feature maps and then aligns them via a novel long-short-range discriminator. Extensive experiments demonstrate that compared to mainstream domain adaptation technologies, BFDA significantly enhances cross-domain pedestrian detection performance for either one-stage or two-stage detectors. Moreover, by employing the efficient one-stage detector (YOLOv5), BFDA can reach 217.4 FPS ( 640×480 pixels) on NVIDIA Tesla V100 (7~12 times the FPS of the existing frameworks), which is highly significant for practical applications. The code from this study will be made publicly available. Yancheng Cai, Bo Zhang 0069, Baopu Li, Tao Chen 0003, Hongliang Yan, Jingdong Zhang 0003 |
IEEE Trans. Image Process. | 4 |
| 2023 | SpVOS: Efficient Video Object Segmentation With Triple Sparse ConvolutionabstractSemi-supervised video object segmentation (Semi-VOS), which requires only annotating the first frame of a video to segment future frames, has received increased attention recently. Among existing Semi-VOS pipelines, the memory-matching-based one is becoming the main research stream, as it can fully utilize the temporal sequence information to obtain high-quality segmentation results. Even though this type of method has achieved promising performance, the overall framework still suffers from heavy computation overhead, mainly caused by the per-frame dense convolution operations between high-resolution feature maps and each kernel filter. Therefore, we propose a sparse baseline of VOS named SpVOS in this work, which develops a novel triple sparse convolution to reduce the computation costs of the overall VOS framework. The designed triple gate, taking full consideration of both spatial and temporal redundancy between adjacent video frames, adaptively makes a triple decision to decide how to apply the sparse convolution on each pixel to control the computation overhead of each layer, while maintaining sufficient discrimination capability to distinguish similar objects and avoid error accumulation. A mixed sparse training strategy, coupled with a designed objective considering the sparsity constraint, is also developed to balance the VOS segmentation performance and computation costs. Experiments are conducted on two mainstream VOS datasets, including DAVIS and Youtube-VOS. Results show that, the proposed SpVOS achieves superior performance over other state-of-the-art sparse methods, and even maintains comparable performance, e.g., an 83.04% (79.29%) overall score on the DAVIS-2017 (Youtube-VOS) validation set, with the typical non-sparse VOS baseline (82.88% for DAVIS-2017 and 80.36% for Youtube-VOS) while saving up to 42% FLOPs, showing its application potential for resource-constrained scenarios. Weihao Lin 0002, Tao Chen 0003, Chong Yu 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Exploring Kernel-Based Texture Transfer for Pose-Guided Person Image GenerationabstractPose-guided person image generation that aims to transfer the pose of a given person to a target pose has recently received lots of research attention. Due to the spatial misalignment and occlusions of different local body parts by pose variations, this task is still challenging especially in maintaining high-fidelity textures and body structures in generated images. Besides, most works also suffer from the limited number of texture styles in the given person datasets, restricting the diversity of generated persons' appearances. To solve these problems, we design a Kernel-based Texture-Fusion Joint Refinement Network (TFJR-Net) to jointly refine the structure and texture information of generated images. First, we leverage a bone-map representation to guide the generation of human parsing maps, which has more structure priors and richer context information than traditional key-point maps, thus reduce the uncertainty of generated body structures. Next, a Texture-Kernel Injection Normalization module (TKIN) is proposed to inject the per-region texture-kernel into the corresponding semantic region from the human parsing map, which decouples the texture and shape information, and also preserves fine-grained features for complex textures. Furthermore, we are the first to introduce external texture patterns outside of the dataset in human semantic regions such as the upper clothes. We fuse the two texture domains in a shared texture space through our designed texture-fusion TKIN modules. Extensive experiments are conducted on the Deepfashion dataset, with the DTD dataset as an external texture source. The experimental results demonstrate the superiority of our proposed method in generating persons of better textures and structures than state-of-the-art works, and also show the generalization ability of our proposed method to absorb diversified external textures for generating person images. The source codes are available athttps://github.com/pilgrim00/TKIN. Jiaxiang Chen, Jiayuan Fan 0001, Hancheng Ye, Jie Li 0040, Yongbin Liao, Tao Chen 0003 |
IEEE Trans. Multim. | 6 |
| 2022 | β-DARTS: Beta-Decay Regularization for Differentiable Architecture SearchabstractNeural Architecture Search (NAS) has attracted increasingly more attention in recent years because of its capability to design deep neural network automatically. Among them, differential NAS approaches such as DARTS, have gained popularity for the search efficiency. However, they suffer from two main issues, the weak robustness to the performance collapse and the poor generalization ability of the searched architectures. To solve these two problems, a simple-but-efficient regularization method, termed as Beta-Decay, is proposed to regularize the DARTS-based NAS searching process. Specifically, Beta-Decay regularization can impose constraints to keep the value and variance of activated architecture parameters from too large. Furthermore, we provide in-depth theoretical analysis on how it works and why it works. Experimental results on NAS-Bench-201 show that our proposed method can help to stabilize the searching process and makes the searched network more transferable across different datasets. In addition, our search scheme shows an outstanding property of being less dependent on training time and data. Comprehensive experiments on a variety of search spaces and datasets validate the effectiveness of the proposed method. The code is available at https://github.com/Sunshine-Ye/Beta-DARTS. Peng Ye 0006, Baopu Li, Yikang Li 0002, Tao Chen 0003, Jiayuan Fan 0001, Wanli Ouyang |
CVPR | 4 |
| 2022 | TopFormer: Token Pyramid Transformer for Mobile Semantic SegmentationabstractAlthough vision transformers (ViTs) have achieved great success in computer vision, the heavy computational cost hampers their applications to dense prediction tasks such as semantic segmentation on mobile devices. In this paper, we present a mobile-friendly architecture named Token Pyramid Vision Transformer (TopFormer). The proposed TopFormer takes Tokens from various scales as input to produce scale-aware semantic features, which are then in-Jected into the corresponding tokens to augment the representation. Experimental results demonstrate that our method significantly outperforms CNN- and ViT-based networks across several semantic segmentation datasets and achieves a good trade-off between accuracy and latency. On the ADE20K dataset, TopFormer achieves 5% higher accuracy in mIoU than MobileNetV3 with lower latency on an ARM-based mobile device. Furthermore, the tiny version of TopFormer achieves real-time inference on an ARM-based mobile device with competitive results. The code and models are available at: https://github.com/hustvl/TopFormer. Guozhong Luo, Tao Chen 0003, Xinggang Wang, Wenyu Liu 0001, Gang Yu 0002, Chunhua Shen |
CVPR | 4 |
| 2022 | Generalized Global Ranking-Aware Neural Architecture Ranker for Efficient Image Classifier SearchabstractNeural Architecture Search (NAS) is a powerful tool for automating effective image processing DNN designing. The ranking has been advocated to design an efficient performance predictor for NAS. The previous contrastive method solves the ranking problem by comparing pairs of architectures and predicting their relative performance. However, it only focuses on the rankings between two involved architectures and neglects the overall quality distributions of the search space, which may suffer generalization issues. A predictor, namely Neural Architecture Ranker (NAR) which concentrates on the global quality tier of specific architecture, is proposed to tackle such problems caused by the local perspective. The NAR explores the quality tiers of the search space globally and classifies each individual to the tier they belong to according to its global ranking. Thus, the predictor gains the knowledge of the performance distributions of the search space which helps to generalize its ranking ability to the datasets more easily. Meanwhile, the global quality distribution facilitates the search phase by directly sampling candidates according to the statistics of quality tiers, which is free of training a search algorithm, e.g., Reinforcement Learning (RL) or Evolutionary Algorithm (EA), thus it simplifies the NAS pipeline and saves the computational overheads. The proposed NAR achieves better performance than the state-of-the-art methods on two widely used datasets for NAS research. On the vast search space of NAS-Bench-101, the NAR easily finds the architecture with top 0.01 performance only by sampling. It also generalizes well to different image datasets of NAS-Bench-201, i.e., CIFAR-10, CIFAR-100, and ImageNet-16-120 by identifying the optimal architectures for each of them. Bicheng Guo, Tao Chen 0003, Shibo He, Haoyu Liu 0002, Lilin Xu, Peng Ye 0006, Jiming Chen 0001 |
ACM Multimedia | 2 |
| 2022 | Learning Cross-Image Object Semantic Relation in Transformer for Few-Shot Fine-Grained Image ClassificationabstractFew-shot fine-grained learning aims to classify a query image into one of a set of support categories with fine-grained differences. Although learning different objects' local differences via Deep Neural Networks has achieved success, how to exploit the query-support cross-image object semantic relations in Transformer-based architecture remains under-explored in the few-shot fine-grained scenario. In this work, we propose a Transformer-based double-helix model, namely HelixFormer, to achieve the cross-image object semantic relation mining in a bidirectional and symmetrical manner. The HelixFormer consists of two steps: 1) Relation Mining Process (RMP) across different branches, and 2) Representation Enhancement Process (REP) within each individual branch. By the designed RMP, each branch can extract fine-grained object-level Cross-image Semantic Relation Maps (CSRMs) using information from the other branch, ensuring better cross-image interaction in semantically related local object regions. Further, with the aid of CSRMs, the developed REP can strengthen the extracted features for those discovered semantically-related local regions in each branch, boosting the model's ability to distinguish subtle feature differences of fine-grained objects. Extensive experiments conducted on five public fine-grained benchmarks demonstrate that HelixFormer can effectively enhance the cross-image object semantic relation matching for recognizing fine-grained objects, achieving much better performance over most state-of-the-art methods under 1-shot and 5-shot scenarios. Bo Zhang 0069, Jiakang Yuan, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001, Botian Shi |
ACM Multimedia | 4 |
| 2022 | Stimulative Training of Residual Networks: A Social Psychology Perspective of LoafingabstractResidual networks have shown great success and become indispensable in today’s deep models. In this work, we aim to re-investigate the training process of residual networks from a novel social psychology perspective of loafing, and further propose a new training strategy to strengthen the performance of residual networks. As residual networks can be viewed as ensembles of relatively shallow networks (i.e., unraveled view) in prior works, we also start from such view and consider that the final performance of a residual network is co-determined by a group of sub-networks. Inspired by the social loafing problem of social psychology, we find that residual networks invariably suffer from similar problem, where sub-networks in a residual network are prone to exert less effort when working as part of the group compared to working alone. We define this previously overlooked problem as network loafing. As social loafing will ultimately cause the low individual productivity and the reduced overall performance, network loafing will also hinder the performance of a given residual network and its sub-networks. Referring to the solutions of social psychology, we propose stimulative training, which randomly samples a residual sub-network and calculates the KL-divergence loss between the sampled sub-network and the given residual network, to act as extra supervision for sub-networks and make the overall goal consistent. Comprehensive empirical results and theoretical analyses verify that stimulative training can well handle the loafing problem, and improve the performance of a residual network by improving the performance of its sub-networks. The code is available at https://github.com/Sunshine-Ye/NIPS22-ST. Peng Ye 0006, Shengji Tang, Baopu Li, Tao Chen 0003, Wanli Ouyang |
NeurIPS | 4 |
| 2022 | Coordinates Are NOT Lonely - Codebook Prior Helps Implicit Neural 3D representationsabstractImplicit neural 3D representation has achieved impressive results in surface or scene reconstruction and novel view synthesis, which typically uses the coordinate-based multi-layer perceptrons (MLPs) to learn a continuous scene representation. However, existing approaches, such as Neural Radiance Field (NeRF) and its variants, usually require dense input views (i.e. 50-150) to obtain decent results. To relive the over-dependence on massive calibrated images and enrich the coordinate-based feature representation, we explore injecting the prior information into the coordinate-based network and introduce a novel coordinate-based model, CoCo-INR, for implicit neural 3D representation. The cores of our method are two attention modules: codebook attention and coordinate attention. The former extracts the useful prototypes containing rich geometry and appearance information from the prior codebook, and the latter propagates such prior information into each coordinate and enriches its feature representation for a scene or object surface. With the help of the prior information, our method can render 3D views with more photo-realistic appearance and geometries than the current methods using fewer calibrated images available. Experiments on various scene reconstruction datasets, including DTU and BlendedMVS, and the full 3D head reconstruction dataset, H3DS, demonstrate the robustness under fewer input views and fine detail-preserving capability of our proposed method. Fukun Yin, Wen Liu 0003, Tao Chen 0003, Gang Yu 0002 |
NeurIPS | 5 |
| 2022 | What Makes for Effective Few-shot Point Cloud Classification?abstractDue to the emergence of powerful computing resources and large-scale annotated datasets, deep learning has seen wide applications in our daily life. However, most current methods require extensive data collection and retraining when dealing with novel classes never seen before. On the other hand, we humans can quickly recognize new classes by looking at a few samples, which motivates the recent popularity of few-shot learning (FSL) in machine learning communities. Most current FSL approaches work on 2D image domain, however, its implication in 3D perception is relatively under-explored. Not only needs to recognize the unseen examples as in 2D domain, 3D few-shot learning is more challenging with unordered structures, high intra-class variances and subtle inter-class differences. Moreover, different architectures and learning algorithms make it difficult to study the effectiveness of existing 2D methods when migrating to the 3D domain.In this work, for the first time, we perform systematic and extensive studies of recent 2D FSL and 3D backbone networks for benchmarking few-shot point cloud classification, and we suggest a strong baseline and learning architectures for 3D FSL. Then, we propose a novel plug-and-play component called Cross-Instance Adaptation (CIA) module, to address the high intra-class variances and subtle inter-class differences issues, which can be easily inserted into current baselines with significant performance improvement. Extensive experiments on two newly introduced benchmark datasets, ModelNet40-FS and ShapeNet70-FS, demonstrate the superiority of our proposed network for 3D FSL. Chuangguan Ye, Hongyuan Zhu 0002, Yongbin Liao, Yanggang Zhang, Tao Chen 0003, Jiayuan Fan 0001 |
WACV | 5 |
| 2022 | Efficient Joint-Dimensional Search with Solution Space Regularization for Real-Time Semantic Segmentation
Peng Ye 0006, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001, Chen Lin 0003, Chongyan Zuo, Qinghua Chi, Wanli Ouyang |
Int. J. Comput. Vis. | 3 |
| 2022 | Point Cloud Instance Segmentation With Semi-Supervised Bounding-Box MiningabstractPoint cloud instance segmentation has achieved huge progress with the emergence of deep learning. However, these methods are usually data-hungry with expensive and time-consuming dense point cloud annotations. To alleviate the annotation cost, unlabeled or weakly labeled data is still less explored in the task. In this paper, we introduce the first semi-supervised point cloud instance segmentation framework (SPIB) using both labeled and unlabelled bounding boxes as supervision. To be specific, our SPIB architecture involves a two-stage learning procedure. For stage one, a bounding box proposal generation network is trained under a semi-supervised setting with perturbation consistency regularization (SPCR). The regularization works by enforcing an invariance of the bounding box predictions over different perturbations applied to the input point clouds, to provide self-supervision for network learning. For stage two, the bounding box proposals with SPCR are grouped into some subsets, and the instance masks are mined inside each subset with a novel semantic propagation module and a property consistency graph module. Moreover, we introduce a novel occupancy ratio guided refinement module to refine the instance masks. Extensive experiments on the challenging ScanNet v2 dataset demonstrate our method can achieve competitive performance compared with the recent fully-supervised methods. Yongbin Liao, Hongyuan Zhu 0002, Yanggang Zhang, Chuangguan Ye, Tao Chen 0003, Jiayuan Fan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Densely Semantic Enhancement for Domain Adaptive Region-Free DetectorsabstractUnsupervised domain adaptive object detection aims to adapt a well-trained detector from its original source domain with rich labeled data to a new target domain with unlabeled data. Previous works focus on improving the domain adaptability of region-based detectors,e.g., Faster-RCNN, through matching cross-domain instance-level features that are explicitly extracted from a region proposal network (RPN). However, this is unsuitable for region-free detectors such as single shot detector (SSD), which perform a dense prediction from all possible locations in an image and do not have the RPN to encode such instance-level features. As a result, they fail to align important image regions and crucial instance-level features between the domains of region-free detectors. In this work, we propose an adversarial module, namely, densely semantic enhancement module (DSEM), to strengthen the cross-domain matching of instance-level features for region-free detectors. Firstly, to emphasize the important regions of image, the DSEM learns to predict a transferable foreground enhancement mask that can be utilized to suppress the background disturbance in an image. Secondly, considering that region-free detectors recognize objects of different scales using multi-layer feature maps, the DSEM encodes multi-scale representations across different domains. Finally, the DSEM is pluggable into different region-free detectors, ultimately achieving the densely semantic feature matching via adversarial learning. Extensive experiments have been conducted on PASCAL VOC, Clipart, Comic, W atercolor, and FoggyCityscape benchmarks, and their results well demonstrate that the proposed approach not only improves the domain adaptability of region-free detectors but also outperforms existing domain adaptive region-based detectors under various domain shift settings. Bo Zhang 0069, Tao Chen 0003, Bin Wang 0008, Xiaofeng Wu 0003, Liming Zhang 0001, Jiayuan Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Curriculum-Style Local-to-Global Adaptation for Cross-Domain Remote Sensing Image SegmentationabstractAlthough domain adaptation has been extensively studied in natural image-based segmentation tasks, the research on cross-domain segmentation for very-high-resolution (VHR) remote sensing images (RSIs) still remains underexplored. The VHR RSI-based cross-domain segmentation mainly faces two critical challenges: 1) large area land covers with many diverse object categories bring severe local patch-level data distribution deviations, thus yielding different adaptation difficulties for different local patches and 2) different VHR sensor types or dynamically changing modes cause the VHR images to go through intensive data distribution differences even for the same geographical location, resulting in different global feature-level domain gaps. To address these challenges, we propose a curriculum-style local-to-global cross-domain adaptation framework for the segmentation of VHR RSIs. The proposed curriculum-style adaptation performs the adaptation process in an easy-to-hard way according to the adaptation difficulties that can be obtained using an entropy-based score for each patch of the target domain and, thus, well aligns the local patches in a domain image. The proposed local-to-global adaptation performs the feature alignment process from the locally semantic to globally structural feature discrepancies and consists of a semantic-level domain classifier and an entropy-level domain classifier that can reduce the above cross-domain feature discrepancies. Extensive experiments have been conducted in various cross-domain scenarios, including geographic location variations and imaging mode variations, and the experimental results demonstrate that the proposed method can significantly boost the domain adaptability of segmentation networks for VHR RSIs. Bo Zhang 0069, Tao Chen 0003, Bin Wang 0008 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | SC-EADNet: A Self-Supervised Contrastive Efficient Asymmetric Dilated Network for Hyperspectral Image ClassificationabstractUnsupervised and semisupervised feature learning has recently emerged as an effective way to reduce the reliance on expensive data collection and annotation for hyperspectral image (HSI) analysis. Existing unsupervised and semisupervised convolutional neural network (CNN)-based HSI classification works still face two challenges: underutilization of pixel-wise multiscale contextual information for feature learning and expensive computational cost, for example, large floating-point operations per seconds (FLOPs), due to the lack of lightweight design. To utilize the unlabeled pixels in the HSIs more efficiently, we propose a self-supervised contrastive efficient asymmetric dilated network (SC-EADNet) for HSI classification. There are two novelties in the SC-EADNet. First, a self-supervised multiscale pixel-wise contextual feature learning model is proposed, which generates multiple patches around each hyperspectral pixel and develops a contrastive learning framework to learn from these patches for HSI classification. Second, a lightweight feature extraction network EADNet, composed of multiple plug-and-play efficient asymmetric dilated convolution (EADC) blocks, is designed and inserted into the contrastive learning framework. The EADC block adopts different dilation rates to capture the spatial information of objects with varying shapes and sizes. Compared with other unsupervised, semisupervised, and supervised learning methods, our SC-EADNet provides competitive classification performance on four hyperspectral datasets, including Indian Pines, Pavia University, Salinas, and Houston 2013, but few FLOPs and fast computational speed. Mingzhen Zhu, Jiayuan Fan 0001, Qihang Yang 0002, Tao Chen 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Sample-Centric Feature Generation for Semi-Supervised Few-Shot LearningabstractSemi-supervised few-shot learning aims to improve the model generalization ability by means of both limited labeled data and widely-available unlabeled data. Previous works attempt to model the relations between the few-shot labeled data and extra unlabeled data, by performing a label propagation or pseudo-labeling process using an episodic training strategy. However, the feature distribution represented by the pseudo-labeled data itself is coarse-grained, meaning that there might be a large distribution gap between the pseudo-labeled data and the real query data. To this end, we propose a sample-centric feature generation (SFG) approach for semi-supervised few-shot image classification. Specifically, the few-shot labeled samples from different classes are initially trained to predict pseudo-labels for the potential unlabeled samples. Next, a semi-supervised meta-generator is utilized to produce derivative features centering around each pseudo-labeled sample, enriching the intra-class feature diversity. Meanwhile, the sample-centric generation constrains the generated features to be compact and close to the pseudo-labeled sample, ensuring the inter-class feature discriminability. Further, a reliability assessment (RA) metric is developed to weaken the influence of generated outliers on model learning. Extensive experiments validate the effectiveness of the proposed feature generation approach on challenging one- and few-shot image classification benchmarks. Bo Zhang 0069, Hancheng Ye, Gang Yu 0002, Bin Wang 0008, Yike Wu 0001, Jiayuan Fan 0001, Tao Chen 0003 |
IEEE Trans. Image Process. | 7 |
| 2022 | Joint Distribution Alignment via Adversarial Learning for Domain Adaptive Object DetectionabstractUnsupervised domain adaptive object detection aims to adapt a well-trained detector from its original source domain with rich labeled data to a new target domain with unlabeled data. Recently, mainstream approaches perform this task through adversarial learning, yet still suffer from two limitations. First, they mainly align marginal distribution by unsupervised cross-domain feature matching, and ignore each feature's categorical and positional information that can be exploited for conditional alignment; Second, they treat all classes as equally important for transferring cross-domain knowledge and ignore that different classes usually have different transferability. In this article, we propose a joint adaptive detection framework (JADF) to address the above challenges. First, an end-to-end joint adversarial adaptation framework for object detection is proposed, which aligns both marginal and conditional distributions between domains without introducing any extra hyper-parameter. Next, to consider the transferability of each object class, a metric for class-wise transferability assessment is proposed, which is incorporated into the JADF objective for domain adaptation. Further, an extended study from unsupervised domain adaptation (UDA) to unsupervised few-shot domain adaptation (UFDA) is conducted, where only a few unlabeled training images are available in unlabeled target domain. Extensive experiments validate that JADF is effective in both the UDA and UFDA settings, achieving significant performance gains over existing state-of-the-art cross-domain detection methods. Bo Zhang 0069, Tao Chen 0003, Bin Wang 0008, Ruoyao Li |
IEEE Trans. Multim. | 2 |
| 2021 | EADNet: Efficient Asymmetric Dilated Network For Semantic SegmentationabstractDue to real-time image semantic segmentation needs on power constrained edge devices, there has been an increasing desire to design lightweight semantic segmentation neural network, to simultaneously reduce computational cost and increase inference speed. In this paper, we propose an efficient asymmetric dilated semantic segmentation network, named EADNet, which consists of multiple developed asymmetric convolution branches with different dilation rates to capture the variable shapes and scales information of an image. Specially, a multi-scale multi-shape receptive field convolution (MMRFC) block with only a few parameters is designed to capture such information. Experimental results on the Cityscapes dataset demonstrate that our proposed EADNet achieves segmentation mIoU of 67.1% with smallest number of parameters (only 0.35M) among mainstream lightweight semantic segmentation networks. Qihang Yang 0002, Tao Chen 0003, Jiayuan Fan 0001, Chongyan Zuo, Qinghua Chi |
ICASSP | 2 |
| 2021 | HSEGAN: Hair Synthesis and Editing Using Structure-Adaptive Normalization on Generative Adversarial NetworkabstractHuman hair is a kind of special material with complex and varied high-frequency details. It is a challenging task to synthesize and edit realistic and fine-grained hair using deep learning methods. In this paper, we propose HSEGAN, a novel framework consisting of two condition modules encoding foreground hair and background respectively, followed by a hair synthesis generator that synthesizes the final result based on the encoded input. For the purpose of efficient and effective hair generation, we propose hair structure-adaptive normalization (HSAN) and use several HSAN residual blocks to build the hair synthesis generator. HSEGAN allows for explicit manipulation of hair at three different levels, including color, structure and shape. Extensive experiments on FFHQ dataset demonstrate our method can generate higher-quality hair images than state-of-the-art methods, yet consume less time in the inference stage. Wanling Fan, Jiayuan Fan 0001, Gang Yu 0002, Tao Chen 0003 |
ICIP | 5 |
| 2021 | Spcr: semi-supervised point cloud instance segmentation with perturbation consistency regularizationabstractPoint cloud instance segmentation is steadily improving with the development of deep learning. However, current progress is hindered by the expensive cost of collecting dense point cloud labels. To this end, we propose the first semi-supervised point cloud instance segmentation architecture, which is called semi-supervised point cloud instance segmentation with perturbation consistency regularization (SPCR). It is capable to alleviate the data-hungry bottleneck of existing strongly supervised methods. Specifically, SPCR enforces an invariance of the predictions over different perturbations applied to the input point clouds. We firstly introduce various perturbation schemes on inputs to force the network to be robust and easily generalized to the unseen and unlabeled data. Further, perturbation consistency regularization is then conducted on predicted instance masks from various transformed inputs to provide self-supervision for network learning. Extensive experiments on the challenging ScanNet v2 dataset demonstrate our method can achieve competitive performance compared with the state-of-the-art of fully supervised methods. Yongbin Liao, Hongyuan Zhu 0002, Tao Chen 0003, Jiayuan Fan 0001 |
ICIP | 3 |
| 2021 | Object-aware Long-short-range Spatial Alignment for Few-Shot Fine-Grained Image ClassificationabstractThe goal of few-shot fine-grained image classification is to recognize rarely seen fine-grained objects in the query set, given only a few samples of this class in the support set. Previous works focus on learning discriminative image features from a limited number of training samples for distinguishing various fine-grained classes, but ignore one important fact that spatial alignment of the discriminative semantic features between the query image with arbitrary changes and the support image, is also critical for computing the semantic similarity between each support-query pair. In this work, we propose an object-aware long-short-range spatial alignment approach, which is composed of a foreground object feature enhancement (FOE) module, a long-range semantic correspondence (LSC) module and a short-range spatial manipulation (SSM) module. The FOE is developed to weaken background disturbance and encourage higher foreground object response. To address the problem of long-range object feature misalignment between support-query image pairs, the LSC is proposed to learn the transferable long-range semantic correspondence by a designed feature similarity metric. Further, the SSM module is developed to refine the transformed support feature after the long-range step to align short-range misaligned features (or local details) with the query features. Extensive experiments have been conducted on four benchmark datasets, and the results show superior performance over most state-of-the-art methods under both 1-shot and 5-shot classification scenarios. Yike Wu 0001, Bo Zhang 0069, Gang Yu 0002, Weixi Zhang, Bin Wang 0008, Tao Chen 0003, Jiayuan Fan 0001 |
ACM Multimedia | 6 |
| 2021 | Coarse-to-Fine Gaze Redirection with Numerical and Pictorial GuidanceabstractGaze redirection aims at manipulating the gaze of a given face image with respect to a desired direction (i.e., a reference angle) and it can be applied to many real life scenarios, such as video-conferencing or taking group photos. However, previous work on this topic mainly suffers of two limitations: (1) Low-quality image generation and (2) Low redirection precision. In this paper, we propose to alleviate these problems by means of a novel gaze redirection framework which exploits both a numerical and a pictorial direction guidance, jointly with a coarse-to-fine learning strategy. Specifically, the coarse branch learns the spatial transformation which warps input image according to desired gaze. On the other hand, the fine-grained branch consists of a generator network with conditional residual image learning and a multi-task discriminator. This second branch reduces the gap between the previously warped image and the ground-truth image and recovers finer texture details. Moreover, we propose a numerical and pictorial guidance module (NPG) which uses a pictorial gazemap description and numerical angles as an extra guide to further improve the precision of gaze redirection. Extensive experiments on a benchmark dataset show that the proposed method outperforms the state-of-the-art approaches in terms of both image quality and redirection precision. The code is available at https://github.com/jingjingchen777/CFGR Jichao Zhang, Enver Sangineto, Tao Chen 0003, Jiayuan Fan 0001, Nicu Sebe |
WACV | 4 |
| 2020 | Cascade EF-GAN: Progressive Facial Expression Editing With Local FocusesabstractRecent advances in Generative Adversarial Nets (GANs) have shown remarkable improvements for facial expression editing. However, current methods are still prone to generate artifacts and blurs around expression-intensive regions, and often introduce undesired overlapping artifacts while handling large-gap expression transformations such as transformation from furious to laughing. To address these limitations, we propose Cascade Expression Focal GAN (Cascade EF-GAN), a novel network that performs progressive facial expression editing with local expression focuses. The introduction of the local focus enables the Cascade EF-GAN to better preserve identity-related features and details around eyes, noses and mouths, which further helps reduce artifacts and blurs within the generated facial images. In addition, an innovative cascade transformation strategy is designed by dividing a large facial expression transformation into multiple small ones in cascade, which helps suppress overlapping artifacts and produce more realistic editing while dealing with large-gap expression transformations. Extensive experiments over two publicly available facial expression datasets show that our proposed Cascade EF-GAN achieves superior performance for facial expression editing. Rongliang Wu, Gongjie Zhang, Shijian Lu, Tao Chen 0003 |
CVPR | 4 |
| 2020 | PIDNet: An Efficient Network for Dynamic Pedestrian Intrusion DetectionabstractVision-based dynamic pedestrian intrusion detection (PID), judging whether pedestrians intrude an area-of-interest (AoI) by a moving camera, is an important task in mobile surveillance. The dynamically changing AoIs and a number of pedestrians in video frames increase the difficulty and computational complexity of determining whether pedestrians intrude the AoI, which makes previous algorithms incapable of this task. In this paper, we propose a novel and efficient multi-task deep neural network, PIDNet, to solve this problem. PIDNet is mainly designed by considering two factors: accurately segmenting the dynamically changing AoIs from a video frame captured by the moving camera and quickly detecting pedestrians from the generated AoI-contained areas. Three efficient network designs are proposed and incorporated into PIDNet to reduce the computational complexity: 1) a special PID task backbone for feature sharing, 2) a feature cropping module for feature cropping, and 3) a lighter detection branch network for feature compression. In addition, considering there are no public datasets and benchmarks in this field, we establish a benchmark dataset to evaluate the proposed network and give the corresponding evaluation metrics for the first time. Experimental results show that PIDNet can achieve 67.1% PID accuracy and 9.6 fps inference speed on the proposed dataset, which serves as a good baseline for the future vision-based dynamic PID study. Jingchen Sun, Jiming Chen 0001, Tao Chen 0003, Jiayuan Fan 0001, Shibo He |
ACM Multimedia | 3 |
| 2020 | Fine-grained facial expression analysis using dimensional emotion model
Feng Zhou 0003, Shu Kong, Charless C. Fowlkes, Tao Chen 0003, Bai Ying Lei |
Neurocomputing | 4 |
| 2020 | BURSTS: A bottom-up approach for robust spotting of texts in scenes
Jiayuan Fan 0001, Tao Chen 0003, Feng Zhou 0003 |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | M$^3$Lung-Sys: A Deep Learning System for Multi-Class Lung Pneumonia Screening From CT ImagingabstractTo counter the outbreak of COVID-19, the accurate diagnosis of suspected cases plays a crucial role in timely quarantine, medical treatment, and preventing the spread of the pandemic. Considering the limited training cases and resources (e.g, time and budget), we propose a Multi-task Multi-slice Deep Learning System (M3Lung-Sys) for multi-class lung pneumonia screening from CT imaging, which only consists of two 2D CNN networks, i.e., slice- and patient-level classification networks. The former aims to seek the feature representations from abundant CT slices instead of limited CT volumes, and for the overall pneumonia screening, the latter one could recover the temporal information by feature refinement and aggregation between different slices. In addition to distinguish COVID-19 from Healthy, H1N1, and CAP cases, our M3Lung-Sys also be able to locate the areas of relevant lesions, without any pixel-level annotation. To further demonstrate the effectiveness of our model, we conduct extensive experiments on a chest CT imaging dataset with a total of 734 patients (251 healthy people, 245 COVID-19 patients, 105 H1N1 patients, and 133 CAP patients). The quantitative results with plenty of metrics indicate the superiority of our proposed model on both slice- and patient-level classification tasks. More importantly, the generated lesion location maps make our system interpretable and more valuable to clinicians. Xuelin Qian, Huazhu Fu, Weiya Shi, Tao Chen 0003, Yanwei Fu 0001, Xiangyang Xue 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | Comp-GAN: Compositional Generative Adversarial Network in Synthesizing and Recognizing Facial ExpressionabstractFacial expression is important in understanding our social interaction. Thus the ability to recognize facial expression enables the novel multimedia applications. With the advance of recent deep architectures, research on facial expression recognition has achieved great progress. However, these models are still suffering from the problems of lacking sufficient and diverse high quality training faces, vulnerability to the facial variations, and recognizing a limited number of basic types of emotions. To tackle these problems, this paper proposes a novel end-to-end Compositional Generative Adversarial Network (Comp-GAN) that is able to synthesize new face images with specified poses and desired facial expressions; and such synthesized images can be further utilized to help train a robust and generalized expression recognition model. Essentially, Comp-GAN can dynamically change the expression and pose of faces according to the input images while keeping the identity information. Specifically, the generator has two major components: one for generating images with desired expression and the other for changing the pose of faces. Furthermore, a face reconstruction learning process is applied to re-generate the input image and constrains the generator for preserving the key information such as facial identity. For the first time, various one/zero-shot facial expression recognition tasks have been created. We conduct extensive experiments to show that the images generated by Comp-GAN are helpful to improve the performance of one/zero-shot facial expression recognition. Wenxuan Wang 0003, Qiang Sun 0007, Yanwei Fu 0001, Tao Chen 0003, Chenjie Cao, Ziqi Zheng, Han Qiu 0002, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
ACM Multimedia | 4 |
| 2019 | SS-HCNN: Semi-Supervised Hierarchical Convolutional Neural Network for Image ClassificationabstractThe availability of large-scale annotated data and uneven separability of different data categories become two major impediments of deep learning for image classification. In this paper, we present a Semi-Supervised Hierarchical Convolutional Neural Network (SS-HCNN) to address these two challenges. A large-scale unsupervised maximum margin clustering technique is designed, which splits images into a number of hierarchical clusters iteratively to learn cluster-level CNNs at parent nodes and category-level CNNs at leaf nodes. The splitting uses the similarity of CNN features to group visually similar images into the same cluster, which relieves the uneven data separability constraint. With the hierarchical cluster-level CNNs capturing certain high-level image category information, the category-level CNNs can be trained with a small amount of labelled images, and this relieves the data annotation constraint. A novel cluster splitting criterion is also designed which automatically terminates the image clustering in the tree hierarchy. The proposed SS-HCNN has been evaluated on the CIFAR-100 and ImageNet classification datasets. Experiments show that the SS-HCNN trained using a portion of labelled training images can achieve comparable performance with other fully trained CNNs using all labelled images. Additionally, the SS-HCNN trained using all labelled images clearly outperforms other fully trained CNNs. Tao Chen 0003, Shijian Lu, Jiayuan Fan 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | S-CNN: Subcategory-Aware Convolutional Networks for Object DetectionabstractThe marriage between the deep convolutional neural network (CNN) and region proposals has made breakthroughs for object detection in recent years. While the discriminative object features are learned via a deep CNN for classification, the large intra-class variation and deformation still limit the performance of the CNN based object detection. We propose a subcategory-aware CNN (S-CNN) to solve the object intra-class variation problem. In the proposed technique, the training samples are first grouped into multiple subcategories automatically through a novel instance sharing maximum margin clustering process. A multi-component Aggregated Channel Feature (ACF) detector is then trained to produce more latent training samples, where each ACF component corresponds to one clustered subcategory. The produced latent samples together with their subcategory labels are further fed into a CNN classifier to filter out false proposals for object detection. An iterative learning algorithm is designed for the joint optimization of image subcategorization, multi-component ACF detector, and subcategory-aware CNN classifier. Experiments on INRIA Person dataset, Pascal VOC 2007 dataset and MS COCO dataset show that the proposed technique clearly outperforms the state-of-the-art methods for generic object detection. Tao Chen 0003, Shijian Lu, Jiayuan Fan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Superpixel Guided Deep-Sparse-Representation Learning for Hyperspectral Image ClassificationabstractThis paper presents a new technique for hyperspectral image (HSI) classification by using superpixel guided deep-sparse-representation learning. The proposed technique constructs a hierarchical architecture by exploiting the sparse coding to learn the HSI representation. Specifically, a multiple-layer architecture using different superpixel maps is designed, where each superpixel map is generated by downsampling the superpixels gradually along with enlarged spatial regions for labeled samples. In each layer, sparse representation of pixels within every spatial region is computed to construct a histogram via the sum-pooling with l1normalization. Finally, the representations (features) learned from the multiple-layer network are aggregated and trained by a support vector machine classifier. The proposed technique has been evaluated over three public HSI data sets, including the Indian Pines image set, the Salinas image set, and the University of Pavia image set. Experiments show superior performance compared with the state-of-the-art methods. Jiayuan Fan 0001, Tao Chen 0003, Shijian Lu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Robust Vehicle Detection and Viewpoint Estimation With Soft Discriminative Mixture ModelabstractVehicle detection and vehicle viewpoint estimation are both crucial for assistive and autonomous driving systems. In this paper, we propose a soft discriminative mixture of viewpoint (SDMoV) models for joint vehicle detection and vehicle viewpoint estimation. The proposed SDMoV model is learned in two steps. First, a discriminative viewpoint-specific component model, which aims to maximize vehicle viewpoint classification accuracy, is learned for each cluster of vehicle images with similar viewpoint. Second, a new soft margin objective function, which aims to maximize vehicle detection accuracy, is designed to retrain these component models into a mixture of viewpoint models. The proposed SDMoV model is capable of detecting vehicles and estimating their viewpoints simultaneously. Experiments on three state-of-the-art datasets show that the proposed SDMoV model achieves superior accuracy for both vehicle detection and vehicle viewpoint estimation tasks. Tao Chen 0003, Shijian Lu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Object-Level Motion Detection From Moving CamerasabstractIt is important for a moving observer to be able to identify his/her surrounding objects and determine whether these objects are moving or stationary, which is called object-level motion detection. Detecting object-level motion from moving cameras is a difficult problem to solve for collision-free navigation due to the dual motion introduced by the mixture of the camera motion and the object motion. This paper presents a novel technique that detects object-level motion from a freely moving camera using only two consecutive video frames. A context-aware motion descriptor (CMD) is designed based on the object’s moving speed and moving direction relative to that of the moving camera. The CMD employs the contextual information, e.g., the optical flow of the image background surrounding the moving object of interest, which describes the object motion behavior better than other contexts such as the camera’s GPS and direction. The inconsistency between the histogram of oriented optical flow of the object and its surrounding background is measured for the object-level motion detection. The proposed technique has been evaluated over two types of widely studied objects, i.e., vehicles and humans that are captured with different sizes, moving speeds, and image backgrounds using a moving camera. Experiments on challenging real-world videos show promising performance in object-level motion detection. Tao Chen 0003, Shijian Lu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Subcategory-Aware Feature Selection and SVM Optimization for Automatic Aerial Image-Based Oil Spill InspectionabstractOil spill inspection is critical to the marine and coastal ecosystems, and has been widely studied by various remote sensing technologies, such as synthetic aperture radar and hyperspectral. To improve the temporal resolution and the inspection flexibility, we propose a novel aerial image-based system that can find oil spills timely from images captured using onboard optical cameras installed in unmanned aerial vehicle or airplanes. In particular, a subcategory-aware feature selection (FS) and support vector machine (SVM) joint optimization technique is proposed to learn a discriminative model that can tell the existence of oil spills within an optical image of the marine surface. A set of color-based features is first extracted and concatenated together to characterize the oil spill incidence in an image, where a new color autocorrelogram is designed, which can better describe each color's spatial distribution in an image. Furthermore, subcategory-aware joint FS and SVM optimization technique is designed, which is capable of generating the optimal feature subset and SVM component models. Experiments on a set of real-world marine surface images show that the proposed technique outperforms the state-of-the-art techniques and achieves promising results for aerial image-based oil spill inspection. Tao Chen 0003, Shijian Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Unsupervised Feature Learning for Land-Use Scene RecognitionabstractThis paper proposes a novel unsupervised feature learning algorithm for land-use scene recognition on very high resolution remote sensing imagery. The proposed technique utilizes a multipath sparse coding architecture in order to capture multiple aspects of discriminative structures within complex remote sensing sceneries. Unlike the previous sparse coding and bag-of-visual-words-based techniques that rely on the handcrafted feature descriptors such as scale-invariant feature transform, the proposed technique extracts dense low-level features from the raw data, including the visual (RGB) data and near-infrared (NIR) data, using image patches of varying sizes at different layers. The proposed technique has been evaluated on three data sets, including the 21-category UC Merced landuse RGB data set with a 1-ft spatial resolution, the 9-category ground scene RGB-NIR data set, and the 10-category Singapore land-use RGB-NIR data set with a 0.5-m spatial resolution. The experimental results show that the proposed technique outperforms the state-of-the-art methods. Jiayuan Fan 0001, Tao Chen 0003, Shijian Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Landmark recognition with compact BoW histogram and ensemble ELM
Jiuwen Cao, Tao Chen 0003, Jiayuan Fan 0001 |
Multim. Tools Appl. | 2 |
| 2015 | DPM revisited: Utilizing root-part spatial distribution for vehicle viewpoint estimationabstractVehicle viewpoint estimation plays an important role for intelligent transportation systems. We present an effective vehicle viewpoint estimation technique by utilizing the spatial location information of root and part objects detected in vehicle images via deformable part-based model (DPM). The viewpoint-aware spatial distribution of each part relative to the root is learned using the Gaussian mixture model. The discriminative capability of each part for each viewpoint is then estimated through measuring the Kullback Leibler divergence between pairwise viewpoint-aware spatial distributions. The discriminative information is finally used to compute the likelihood that the detected vehicle belongs to each viewpoint. Experimental results on a benchmark dataset demonstrate the superior performance of the proposed vehicle viewpoint estimation technique. Tao Chen 0003, Shijian Lu |
ICIP | 1 |
| 2015 | Context-aware lane marking detection on urban roadsabstractAutomatic lane marking detection plays an important role in intelligent transportation systems. We present an effective lane marking detection technique that utilizes the context-aware information of lane marking on the urban roads. The proposed technique consists of two innovations. First, the context-aware color, texture and shape features which characterise both lane markings and their road context are designed to represent the lane markings on the road surface. Second, a hard negative mining technique is developed based on the Maximum Stable Extreme Region (MSER) detector and adaboost training. Experiments on a real world dataset demonstrate the superior performance of the proposed approach. Tao Chen 0003, Shijian Lu |
ICIP | 1 |
| 2015 | Reversible watermarking using enhanced local predictionabstractReversible watermarking has drawn extensive attentions in recent years due to its broad applications of digital forensics and data security. This paper proposes a novel reversible watermarking approach, which has the perfect reversibility of the embedded data and the original image. In order to increase the embedding capacity, the proposed approach utilizes the enhanced local prediction to reduce the prediction error of every pixel value. Different from traditional reversible watermarking algorithms considering all the image pixels equivalently before prediction, an enhanced image is first computed by multiplying the original image with its saliency map. By linearly formulating each pixel value in an original image as a weighted sum of the enhanced pixel values in its local neighborhood, the correlation coefficients as learned weights are then solved as a least squares solution. Based on two state-of-the-art datasets, the experimental results show that the proposed approach achieves large embedding capacity with relatively low visual distortion. Jiayuan Fan 0001, Tao Chen 0003 |
ICIP | 2 |
| 2015 | Vegetation coverage detection from very high resolution satellite imageryabstractAutomatic vegetation coverage detection plays a key role for monitoring and management of land usage, environmental variation, and urban planning. This paper presents a novel vegetation coverage detection technique for very high resolution multi-spectral satellite imagery. The proposed technique consists of two stages including a supervised patch-level scoring stage and an unsupervised pixel-level classification stage. In the first stage, a support vector regression (SVR) technique is developed which scores each image patch and generates a coarse patch-level vegetation map. In the second stage, an unsupervised pixel-level vegetation classification technique is developed, which produces a more detailed vegetation map by re-scoring those uncertain pixels based on the computed SVR scores. Experiments on very high resolution multi-spectral satellite images show that the proposed technique outperforms the state-of-the-art methods in both patch-level and pixel-level vegetation detection. Jiayuan Fan 0001, Tao Chen 0003, Shijian Lu |
VCIP | 2 |
| 2015 | Scene text extraction based on edges and support vector regression
Shijian Lu, Tao Chen 0003, Shangxuan Tian, Joo-Hwee Lim, Chew Lim Tan |
Int. J. Document Anal. Recognit. | 2 |
| 2015 | Context-aware vocabulary tree for mobile landmark recognition
Tao Chen 0003, Shijian Lu, Jiayuan Fan 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | Context-aware codebook learning for mobile landmark recognitionabstractThis paper presents a codebook learning based mobile landmark recognition technique based on context information that is acquired from mobile devices. Previous codebook learning methods are mainly developed on nonmobile platforms such as desktop PC, hence underutilize context features such as location and direction information as provided by the mobile devices. The proposed technique employs both the direction and location information to learn the codebook for mobile landmark recognition. A set of direction-aware leaf codewords are first generated by using direction data to decompose the leaf nodes of the original SVT. A visual word significance learning algorithm is then developed by considering location information to generate a compact codebook for image encoding. Experiments on the NTU50Landmark database show that the proposed method can achieve good recognition performance in mobile landmark recognition. Tao Chen 0003, Jiayuan Fan 0001, Shijian Lu |
ICIP | 1 |
| 2014 | Discriminative BoW Framework for Mobile Landmark RecognitionabstractThis paper proposes a new soft bag-of-words (BoW) method for mobile landmark recognition based on discriminative learning of image patches. Conventional BoW methods often consider the patches/regions in the images as equally important for learning. Amongst the few existing works that consider the discriminative information of the patches, they mainly focus on selecting the representative patches for training, and discard the others. This binary hard selection approach results in underutilization of the information available, as some discarded patches may still contain useful discriminative information. Further, not all the selected patches will contribute equally to the learning process. In view of this, this paper presents a new discriminative soft BoW approach for mobile landmark recognition. The main contribution of the method is that the representative and discriminative information of the landmark is learned at three levels: patches, images, and codewords. The patch discriminative information for each landmark is first learned and incorporated through vector quantization to generate soft BoW histograms. Coupled with the learned representative information of the images and codewords, these histograms are used to train an ensemble of classifiers using fuzzy support vector machine. Experimental results on two different datasets show that the proposed method is effective in mobile landmark recognition. Tao Chen 0003, Kim-Hui Yap |
IEEE Trans. Cybern. | 1 |
| 2014 | Discriminative Soft Bag-of-Visual Phrase for Mobile Landmark RecognitionabstractThis paper proposes a new bag-of-visual phrase (BoP) approach for mobile landmark recognition based on discriminative learning of category-dependent visual phrases. Many previous landmark recognition works adopt a bag-of-words (BoW) method which ignores the co-occurrence relationship between neighboring visual words in an image. Although some works that focus on visual phrase learning have appeared, they mainly construct a generalized phrase dictionary from all categories for recognition, which lacks descriptive capability for a specific category. Another shortcoming of these works is the hard assignment of numerous feature sets to a limited number of phrases, which causes some useful feature sets to be discarded, and yields information loss. In view of this, this paper presents a discriminative soft BoP approach for mobile landmark recognition. The candidate phrases defined as adjacent pairwise codewords are first generated for each category. The important candidates are then selected through a proposed discriminative visual phrase (DVP) selection approach to form the BoP dictionary. Finally, a soft encoding method is developed to quantize each image into a BoP histogram. The context information such as location and direction captured by mobile devices is also integrated with the proposed BoP-based content analysis for landmark recognition. Experimental results on two datasets show that the proposed method is effective in mobile landmark recognition. Tao Chen 0003, Kim-Hui Yap, Dajiang Zhang |
IEEE Trans. Multim. | 1 |
| 2013 | Context-Aware Discriminative Vocabulary Learning for Mobile Landmark RecognitionabstractThis paper proposes a discriminative vocabulary learning for landmark recognition based on the context information acquired from mobile devices. The vocabulary learning generates a set of discriminative codewords for image representation, which is important for landmark recognition. Many state-of-the-art methods use content analysis alone for vocabulary learning, which underutilizes the context information provided by mobile devices, such as location from the GPS positioner and direction from the digital compass. Although some works start to consider the images' location information for vocabulary learning, the location alone is insufficient since GPS data has significant errors in dense built-up areas. The context analysis techniques that use GPS to shortlist the geographically nearby landmark candidates for subsequent image matching are at times inadequate. In view of this, the paper proposes to employ both direction and location information to learn a discriminative compact vocabulary (DCV) for mobile landmark recognition. Direction information is first considered to supervise image feature clustering to construct direction-dependent scalable vocabulary trees (DSVTs). Location information is then incorporated into the proposed DCV learning algorithm, to select the discriminative codewords of the DSVT to form the DCV. An ImageRank technique and an iterative codeword selection algorithm are developed for DCV learning. Experimental results using the NTU50Landmark database show that the proposed approach achieves 4% improvement over the current method in mobile landmark recognition. Tao Chen 0003, Kim-Hui Yap |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | Discriminative bag-of-visual phrase learning for landmark recognitionabstractBag-of-visual phrase (BoP) has been proposed and developed for landmark recognition recently. However, existing BoP methods for landmark recognition have two major shortcomings: (i) they try to construct a universal phrase vocabulary for all object categories, which lacks specific descriptive capabilities for a particular category, and (ii) they often adopt simple criterion such as the frequency information to mine the visual phrases, which may cause the selected phrases to be less discriminative or representative for recognition. In view of this, this paper proposes a new discriminative BoP approach for landmark recognition. First, the candidate visual phrases defined as adjacent pairwise words are selected for each category. A phrase-level similarity measure at the latent space is proposed to evaluate the semantic similarity between pairwise phrases. This is then integrated with the phrase frequency information to shortlist the discriminative phrases for each category through a proposed phrase ranking algorithm. Finally, the BoP and bag-of-words (BoW) histograms are combined through a pyramid matching method for recognition. Experimental results on two different datasets demonstrate that the proposed method is effective in landmark recognition. Tao Chen 0003, Kim-Hui Yap, Dajiang Zhang |
ICASSP | 1 |
| 2011 | A discriminative learning technique for mobile landmark recognitionabstractThis paper proposes a discriminative learning bags-of-words (BoW) approach for mobile landmark recognition at patch and image levels. Conventional methods often treat the local patches and images equally important for recognition and do not differentiate their different importance. Although there exist several works that consider the patches' discrimination information, they mainly focus on which patches are to be retained for training and do not incorporate this information when generating the BoW histograms. In view of this, this paper proposes to learn the discriminative information for each landmark category at two levels: local patches and images. At patch level, the patches' discrimination information for each landmark is first discovered using an iterative learning approach. This information is then incorporated into the quantization process to generate the BoW histogram. At image level, the different importance of training images is estimated through a non-parametric density estimator. Finally, fuzzy SVM is used to train the classifier for each category. Experimental results on a landmark database consisting of 3622 training images and 534 testing images show that the proposed method is effective in mobile landmark recognition. Tao Chen 0003, Kim-Hui Yap, Lap-Pui Chau |
ICIP | 1 |
| 2011 | From universal bag-of-words to adaptive bag-of-phrases for mobile scene recognitionabstractThis paper proposes an adaptive bag-of-phrases (BoP) algorithm for mobile scene recognition based on bag-of- words approach. Conventional BoW methods do not consider the dependence and pairwise relationship among different codewords. However, these contextual relations between pairwise codewords play an important role for users to recognize an image. In light of this problem, this paper proposes an effective BoP technique to integrate both the spatial and contextual information between visual words for scene recognition. It first uses hierarchical k-means algorithm to construct a universal codebook for all categories. The contextual (dependence) relationship between pairwise words is then mined for each category based on the mutual information they contain. Subsequently, a visual phrase vocabulary is constructed which is then used to generate a BoP histogram through a proposed quantization method. Finally, support vector machine (SVM) is used to train these histograms into a classifier. Experimental results on the Scene 15 dataset show that the proposed method is effective for mobile scene recognition. Tao Chen 0003, Kim-Hui Yap, Lap-Pui Chau |
ICIP | 1 |
| 2011 | Integrated Content and Context Analysis for Mobile Landmark RecognitionabstractThis paper proposes a new approach for mobile landmark recognition based on integrated content and context analysis. Conventional scene/landmark recognition methods focus mainly on nonmobile desktop/PC platform, where content analysis alone is used to perform landmark recognition. These nonmobile systems, however, do not take unique features of mobile devices into consideration, e.g., limited computational power and fast response time requirement of mobile users. On the contrary, most existing context-aware content mobile landmark recognition methods mainly rely on global positioning system location information for context analysis. In view of this, this paper proposes an effective method that employs an integration of content and context analysis to perform landmark recognition using mobile devices. A new bags-of-words (BoW) framework is developed to perform content analysis. It is then integrated with context analysis involving fusion of location and direction information to perform mobile landmark recognition. Experimental results based on the NTU50Landmark database show that the proposed method can achieve good recognition performance in mobile landmark recognition. Tao Chen 0003, Kim-Hui Yap, Lap-Pui Chau |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | A Survey on Mobile Landmark Recognition for Information RetrievalabstractThe growing usage of mobile devices has led to proliferation of many mobile applications. A growing trend in mobile applications is centered on mobile landmark recognition. It is a new mobile application that recognizes a captured landmark using the mobile device and retrieves related information. This paper will present a survey on mobile landmark recognition for information retrieval. A general overview of existing mobile landmark recognition systems will be summarized. The techniques and algorithms used in the literatures, including content analysis of landmarks and classification methods for recognition, will be described. Tao Chen 0003, Kui Wu 0002, Kim-Hui Yap, Zhen Li 0047, Flora S. Tsai |
Mobile Data Management | 1 |
| 2004 | Transformation and combination of hiden Markov models for speaker selection training
Chao Huang 0011, Tao Chen 0003, Eric Chang |
INTERSPEECH | 2 |
| 2002 | Speaker selection training for large vocabulary continuous speech recognitionabstractAcoustic variability across speakers is one of the challenges of speaker independent (SI) speech recognition systems. As a powerful solution, dominant speaker adaptation technologies such as MLLR and MAP may become inefficient because of the lack of enough enrollment data. In this paper, we propose an adaptation method based on speaker selection training, which makes full use of statistics of training corpus. Relative error rate reduction of 5.31 % is achieved when only one utterance is available. We compare different speaker selection strategies, namely. PCA, HMM and GMM based methods. In addition, impacts of number of selected cohort speakers and number of utterances from target speaker are investigated. Furthermore, comparison and integration with MLLR adaptation are also shown. Finally, some ongoing work such as dynamicalJy varying number of selected speakers, measuring the relative contribution among the selected speakers and speeding up the computationally expensive procedure of re-estimation with model synthesis are also discussed. Chao Huang 0011, Tao Chen 0003, Eric Chang |
ICASSP | 2 |
| 2002 | On the use of Gaussian mixture model for speaker variability analysis
Tao Chen 0003, Chao Huang 0011, Eric Chang, Jingchun Wang |
INTERSPEECH | 1 |
| 2002 | Adaptive model combination for dynamic speaker selection training
Chao Huang 0011, Tao Chen 0003, Eric Chang |
INTERSPEECH | 2 |
| 2001 | Analysis of speaker variabilityabstractAnalysis and modeling of speaker variability, such as gender, accent, age, speech rate, and phones realizations, are important issues in speech recognition. It is known that existing feature representations describing speaker variations can be of very high dimension. In this paper, we introduce two powerful multivariate statistical analysis methods, namely, principal component analysis (PCA) and independent component analysis (ICA), as tools for analysis of such variability and extraction of low dimensional feature representation. Our findings are the following: (1) the first two principal components correspond to the gender and accent, respectively. The result that the second component corresponding to the accent has never been reported before, to the best of our knowledge. (2) It is shown that ICA based features yield better classification performance than PCA ones. Using 2dimensional ICA representation, we achieved about 6.1% and 13.3% error rate in gender and accent classification, respectively, for 980 speakers. Chao Huang 0011, Tao Chen 0003, Stan Z. Li, Eric Chang, Jian-Lai Zhou |
INTERSPEECH | 2 |