Xiao Bai 0001

dblp:99/4833-1 · DBLP profile ↗
← Back
154ranked-venue papers
21as first author
72since 2021 · last 2026
0000-0001-8561-9299ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 104 · 16 first-author · 57 since 2021Graphics, computer vision, multimedia, augmented reality and games · 69 · 7 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 2 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 SparseSurf: Sparse-View 3D Gaussian Splatting for Surface Reconstruction
abstract
Recent advances in optimizing Gaussian Splatting for scene geometry have enabled efficient reconstruction of detailed surfaces from images. However, when input views are sparse, such optimization is prone to overfitting, leading to suboptimal reconstruction quality. Existing approaches address this challenge by employing flattened Gaussian primitives to better fit surface geometry, combined with depth regularization to alleviate geometric ambiguities under limited viewpoints. Nevertheless, the increased anisotropy inherent in flattened Gaussians exacerbates overfitting in sparse-view scenarios, hindering accurate surface fitting and degrading novel view synthesis performance. In this paper, we propose SparseSurf, a method that reconstructs more accurate and detailed surfaces while preserving high-quality novel view rendering. Our key insight is to introduce Stereo Geometry-Texture Alignment, which bridges rendering quality and geometry estimation, thereby jointly enhancing both surface reconstruction and view synthesis. In addition, we present a Pseudo-Feature Enhanced Geometry Consistency that enforces multi-view geometric consistency by incorporating both training and unseen views, effectively mitigating overfitting caused by sparse supervision. Extensive experiments on the DTU, BlendedMVS, and Mip-NeRF360 datasets demonstrate that our method achieves the state-of-the-art performance.
Meiying Gu, Jiahe Li 0007, Xiaohan Yu 0001, Haonan Luo 0002, Xiao Bai 0001
AAAI7
2026 MTAttack: Multi-Target Backdoor Attacks Against Large Vision-Language Models
abstract
Recent advances in Large Visual Language Models (LVLMs) have demonstrated impressive performance across various vision-language tasks by leveraging large-scale image-text pretraining and instruction tuning. However, the security vulnerabilities of LVLMs have become increasingly concerning, particularly their susceptibility to backdoor attacks. Existing backdoor attacks focus on single-target attacks, i.e., targeting a single malicious output associated with a specific trigger. In this work, we uncover multi-target backdoor attacks, where multiple independent triggers corresponding to different attack targets are added in a single pass of training, posing a greater threat to LVLMs in real-world applications. Executing such attacks in LVLMs is challenging since there can be many incorrect trigger-target mappings due to severe feature interference among different triggers. To address this challenge, we propose MTAttack, the first multi-target backdoor attack framework for enforcing accurate multiple trigger-target mappings in LVLMs. The core of MTAttack is a novel optimization method with two constraints, namely Proxy Space Partitioning constraint and Trigger Prototype Anchoring constraint. It jointly optimizes multiple triggers in the latent space, with each trigger independently mapping clean images to a unique proxy class while at the same time guaranteeing their separability. Experiments on popular benchmarks demonstrate a high success rate of MTAttack for multi-target attacks, substantially outperforming existing attack methods. Furthermore, our attack exhibits strong generalizability across datasets and robustness against backdoor defense strategies. These findings highlight the vulnerability of LVLMs to multi-target backdoor attacks and underscore the urgent need for mitigating such threats.
Guansong Pang, Wenjun Miao, Xiao Bai 0001
AAAI5
2026 FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM
abstract
We present FoundationSLAM, a learning-based monocular dense SLAM system that addresses the absence of geometric consistency in previous flow-based approaches for accurate and robust tracking and mapping. Our core idea is to bridge flow estimation with geometric reasoning by leveraging the guidance from foundation depth models. To this end, we first develop a Hybrid Flow Network that produces geometry-aware correspondences, enabling consistent depth and pose inference across diverse keyframes. To enforce global consistency, we propose a Bi-Consistent Bundle Adjustment Layer that jointly optimizes keyframe pose and depth under multi-view constraints. Furthermore, we introduce a Reliability-Aware Refinement mechanism that dynamically adapts the flow update process by distinguishing between reliable and uncertain regions, forming a closed feedback loop between matching and optimization. Extensive experiments demonstrate that FoundationSLAM achieves superior trajectory accuracy and dense reconstruction quality across multiple challenging datasets, while running in real-time at 18 FPS, demonstrating strong generalization to various scenarios and practical applicability of our method.
Jiahe Li 0007, Fabio Tosi, Matteo Poggi, Xiao Bai 0001
AAAI6
2026 CurvLoc: Surface Curvature Prompted Gaussian Splatting for Visual Localization
Jiahe Li 0007, Botao Jiang, Zihang Wang 0002, Xiaohan Yu 0001, Xiao Bai 0001, Haonan Luo 0002
Int. J. Comput. Vis.8
2026 Dual-label noise filtering for weakly supervised person search
Huadong Lin, Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Chen Wang 0026
Neurocomputing4
2026 DNGaussian++: Improving Sparse-View Gaussian Radiance Fields With Depth Normalization
abstract
Synthesizing novel views from sparse views has achieved impressive advances with radiance fields, yet prevailing methods suffer from high consumption or insufficient refinement capability. This paper introduces DNGaussian, a depth-regularized framework based on 3D Gaussian Splatting, offering real-time and high-quality few-shot novel view synthesis at low costs. Our motivation stems from the remarkable advancement of recent 3D Gaussian Splatting, despite it will encounter a geometry degradation when input views decrease. In the Gaussian radiance fields, we find this degradation in scene geometry primarily lined to the positioning of Gaussian primitives and can be mitigated by depth constraint. Consequently, we propose a Hard and Soft Depth Regularization to restore accurate scene geometry under coarse monocular depth supervision while maintaining a fine-grained color appearance. To further refine detailed geometry, we introduce Global-Local Depth Normalization, enhancing the focus on small local depth changes. Although DNGaussian shows impressive performance, its patch-wise regularization obscures the inconsistency in cross-patch errors. Additionally, primitives can still be irreversibly trapped in local minima under sparse views, even if depth regularization is applied. In this paper, we propose an extended version, DNGaussian++. First, a Geometry Instance Regularizer is developed to enable depth regularization for continuous consistency by exploiting reliable instance-level depth cues. Leveraging the depth gradient guidance, we then propose a Depth-Guided Geometry Reorganization to address the aforementioned local minima problem with high representation efficiency. Extensive experiments show that DNGaussian++ exhibits state-of-the-art performance in multiple datasets and scenarios with high efficiency, and the broad applicability and effectiveness are verified on various backbones and tasks.
Jiahe Li 0007, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001, Lin Gu 0003
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Celebrating the Life and Research Work of Edwin Hancock
Xiao Bai 0001, Jun Zhou 0001, Richard C. Wilson 0001, Charlotte Davies, Josef Kittler
Pattern Recognit.1
2026 EIK-Nav: Boosting zero-shot object navigation with explicit and implicit knowledge
Botao Jiang, Haonan Luo 0002, Zihang Wang 0002, Jiahe Li 0007, Xiao Bai 0001
Pattern Recognit.7
2026 HybridEditDif: Text and exemplar guided image editing with diffusion models
Xuemei Fu, Long Cheng 0003, Jungong Han, Catarina Moreira, Xin Ning 0001, Xiao Bai 0001
Pattern Recognit.8
2026 OpenCIL: Benchmarking out-of-distribution detection in class incremental learning
Wenjun Miao, Guansong Pang, Trong-Tung Nguyen, Ruohuan Fang, Xiao Bai 0001
Pattern Recognit.6
2025 Visual Perturbation for Text-Based Person Search
abstract
Text-based person search aims at locating a person described by natural language in uncropped scene images. Recent works for TBPS mainly focus on aligning multi-granularity vision and language representations, neglecting a key discrepancy between training and inference where the former learns to unify vision and language features where the visual side covers all clues described by language, yet the latter matches image-text pairs where the images may capture only part of the described clues due to perturbations such as occlusions, background clutters and misaligned boundaries. To alleviate this issue, we present ViPer: a Visual Perturbation network that learns to match language descriptions with perturbed visual clues. On top of a CLIP-driven baseline, we design three visual perturbation modules: (1) Spatial ViPer that varies person proposals and produces visual features with misaligned boundaries, (2) Attentive ViPer that estimates visual attention on the fly and manipulates attentive visual tokens within a proposal to produce global features under visual perturbations, and (3) Fine-grained ViPer that learns to recover masked visual clues from detailed language descriptions to encourage matching language features with perturbed visual features at the fine granularity. This overall framework thus simulates real-world scenarios at the training stage to minimize the discrepancy and improve the generalization ability of the model. Experimental results demonstrate that the proposed method clearly surpasses previous TBPS methods on the PRW-TBPS and CUHK-SYSU-TBPS datasets.
Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001
AAAI3
2025 InsTaG: Learning Personalized 3D Talking Head from Few-Second Video
abstract
Despite exhibiting impressive performance in synthesizing lifelike personalized 3D talking heads, prevailing methods based on radiance fields suffer from high demands for training data and time for each new identity. This paper introduces InsTaG, a 3D talking head synthesis framework that allows a fast learning of realistic personalized 3D talking head from few training data. Built upon a lightweight 3DGS person-specific synthesizer with universal motion priors, InsTaG achieves high-quality and fast adaptation while preserving high-level personalization and efficiency. As preparation, we first propose an Identity-Free Pre-training strategy that enables the pre-training of the person-specific model and encourages the collection of universal motion priors from long-video data corpus. To fully exploit the universal motion priors to learn an unseen new identity, we then present a Motion-Aligned Adaptation strategy to adaptively align the target head to the pre-trained field, and constrain a robust dynamic head structure under few training data. Experiments demonstrate our outstanding performance and efficiency under various data scenarios to render high-quality personalized talking heads. Project page: https://fictionarry.github.io/InsTaG/.
Jiahe Li 0007, Xiao Bai 0001, Jun Zhou 0001, Lin Gu 0003
CVPR3
2025 Auxiliary Prompt Tuning of Vision-Language Models for Few-Shot Out-of-Distribution Detection
Wenjun Miao, Guansong Pang, Xiao Bai 0001
ICCV5
2025 Revisiting Continual Ultra-fine-grained Visual Recognition with Pre-trained Models
abstract
Continual ultra-fine-grained visual recognition (C-UFG) aims to continuously learn to categorize the increasing number of cultivates (VC-UFG) and consistently recognize crops across reproductive stages (HC-UFG), which is a fundamental goal of intelligent agriculture. Despite the progress made in general continual learning, C-UFG remains an underexplored issue. This work establishes the first comprehensive C-UFG benchmark using massive soy leaf data. By analyzing recent pre-trained model (PTM) based continual learning methods on the proposed benchmark, we propose two simple yet effective PTM-based methods to boost the performance of VC-UFG and HC-UFG, respectively. On top of those, we integrate the two methods into one unified framework and propose the first unified model, Unic, that is capable of tackling the C-UFG problem where VC-UFG and HC-UFG co-exist in a single continual learning sequence. To understand the effectiveness of the proposed methods, we first evaluate the models on VC-UFG and HC-UFG challenges and then test the proposed Unic on a unified C-UFG challenge. Experimental results demonstrate the proposed methods achieve superior performance for C-UFG. The code is available at https://github.com/PatrickZad/unicufg.
Pengcheng Zhang 0003, Xiaohan Yu 0001, Meiying Gu, Yongsheng Gao 0001, Xiao Bai 0001
IJCAI6
2025 View-aware Decomposition and Unification for Fast Ground-to-Aerial Person Search
abstract
Ground-to-aerial person search leverages cooperative efforts between unmanned aerial vehicles (UAV) and ground surveillance cameras to locate person individuals. Despite the progress made by recent works, the impact of the discrepancy between the two views is underestimated. This limits the overall person search performance when training the model in a view-agnostic way. To address this, we propose a view-aware decomposition and unification (VADU) framework for ground-to-aerial person search. Specifically, we decompose the person search model to learn view-oriented modules for image feature encoding and person proposal generation. The data sampling and retrieval feature learning are also composed to cope with the decomposed model. This decomposition improves both person detection and discriminative feature learning within each view. On top of the decomposition, we propose view-aware unification to produce unified cross-view person features. Cross-view prototypical contrastive learning is introduced to enhance the unification between different views, enhancing model robustness to retrieve a target person in cameras of a different view. As the decomposed parts of the model are deployed on different devices for inference, this overall framework adds no extra computation cost in real-world applications. Extensive experiments demonstrate that the proposed method achieves superior person search performance and guarantees the efficiency of inference. The source code is available at https://github.com/QFWang-11/vadu.
Qifei Wang, Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Yongsheng Gao 0001
IROS4
2025 GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction
abstract
Reconstructing accurate surfaces with radiance fields has achieved remarkable progress in recent years. However, prevailing approaches, primarily based on Gaussian Splatting, are increasingly constrained by representational bottlenecks. In this paper, we introduce GeoSVR, an explicit voxel-based framework that explores and extends the under-investigated potential of sparse voxels for achieving accurate, detailed, and complete surface reconstruction. As strengths, sparse voxels support preserving the coverage completeness and geometric clarity, while corresponding challenges also arise from absent scene constraints and locality in surface refinement. To ensure correct scene convergence, we first propose a Voxel-Uncertainty Depth Constraint that maximizes the effect of monocular depth cues while presenting a voxel-oriented uncertainty to avoid quality degradation, enabling effective and robust scene constraints yet preserving highly accurate geometries. Subsequently, Sparse Voxel Surface Regularization is designed to enhance geometric consistency for tiny voxels and facilitate the voxel-based formation of sharp and accurate surfaces. Extensive experiments demonstrate our superior performance compared to existing methods across diverse challenging scenarios, excelling in geometric accuracy, detail preservation, and reconstruction completeness while maintaining high efficiency. Code is available at https://github.com/Fictionarry/GeoSVR.
Jiahe Li 0007, Youmin Zhang 0005, Xiao Bai 0001, Xiaohan Yu 0001, Lin Gu 0003
NeurIPS4
2025 Eve3D: Elevating Vision Models for Enhanced 3D Surface Reconstruction via Gaussian Splatting
abstract
We present Eve3D, a novel framework for dense 3D reconstruction based on 3D Gaussian Splatting (3DGS). While most existing methods rely on imperfect priors derived from pre-trained vision models, Eve3D fully leverages these priors by jointly optimizing both them and the 3DGS backbone. This joint optimization creates a mutually reinforcing cycle: the priors enhance the quality of 3DGS, which in turn refines the priors, further improving the reconstruction. Additionally, Eve3D introduces a novel optimization step based on bundle adjustment, overcoming the limitations of the highly local supervision in standard 3DGS pipelines. Eve3D achieves state-of-the-art results in surface reconstruction and novel view synthesis on the Tanks & Temples, DTU, and Mip-NeRF360 datasets. while retaining fast convergence, highlighting an unprecedented trade-off between accuracy and speed.
Youmin Zhang 0005, Fabio Tosi, Meiying Gu, Jiahe Li 0007, Xiaohan Yu 0001, Xiao Bai 0001, Matteo Poggi
NeurIPS8
2025 Fully Decoupled End-to-End Person Search: An Approach without Conflicting Objectives
Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001, Edwin R. Hancock
Int. J. Comput. Vis.3
2025 3D human avatar reconstruction with neural fields: A recent survey
Meiying Gu, Jiahe Li 0007, Haonan Luo 0002, Xiao Bai 0001
Image Vis. Comput.6
2025 Investigating Synthetic-to-Real Transfer Robustness for Stereo Matching and Optical Flow Estimation
abstract
With advancements in robust stereo matching and optical flow estimation networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, their robustness can be seriously degraded when fine-tuning them in real-world scenarios. This paper investigates fine-tuning stereo matching and optical flow estimation networks without compromising their robustness to unseen domains. Specifically, we divide the pixels into consistent and inconsistent regions by comparing Ground Truth (GT) with Pseudo Label (PL) and demonstrate that the imbalance learning of consistent and inconsistent regions in GT causes robustness degradation. Based on our analysis, we propose the DKT framework, which utilizes PL to balance the learning of different regions in GT. The core idea is to utilize an exponential moving average (EMA) teacher to measure what the student network has learned and dynamically adjust the learning regions. We further propose the DKT++ framework, which improves target-domain performances and network robustness by applying slow-fast update teachers to generate more accurate PL, introducing the unlabeled data and synthetic data. We integrate our frameworks with state-of-the-art networks and evaluate their effectiveness on several real-world datasets. Extensive experiments show that our method effectively preserves the robustness of stereo matching and optical flow networks during fine-tuning.
Jiahe Li 0007, Lei Huang 0015, Haonan Luo 0002, Xiaohan Yu 0001, Lin Gu 0003, Xiao Bai 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 ALStereo: Active learning for stereo matching
Jiahe Li 0007, Meiying Gu, Xiaohan Yu 0001, Xiao Bai 0001, Edwin R. Hancock
Pattern Recognit.6
2025 Dual Guidance Enabled Fuzzy Inference for Enhanced Fine-Grained Recognition
abstract
In the field of fine-grained visual recognition (FGVR), the ability to resolve minute and often subtle differences between highly similar object categories is paramount. The advent of vision transformers (ViTs) has marked a significant advancement in this domain, primarily due to their capacity to model the intricate interdependencies among object parts represented as image patches. However, their inherent single-scale processing limitation hampers their effectiveness in FGVR tasks. Furthermore, the challenge of uncertainty inherent in FGVR tasks remains unresolved, necessitating the development of methods that bolster the robustness of these models, particularly across varying scales of visual features. We introduce a new plug-in module that can be seamlessly integrated into ViT, called dual guidance enabled fuzzy inference (DGEFI), which combines fuzzy inference with dual guidance mechanisms. Dual guidance includes scale-aware guidance and probability guidance. The former strengthens the model's focus on salient scales, and the latter refines the distinction between similar categories by optimizing intraclass compactness and interclass separability. Fuzzy inference enables the model to adaptively tweak the influence of distinct scales in the final decision-making phase, thereby enhancing the overall accuracy of recognition tasks. We demonstrate the versatility and efficacy of our DGEFI module by integrating it into several leading ViT backbones, including ViT, Swin, Mvitv2, and EVA-02. Empirical results exhibit exceptional performance gains, with the integration of DGEFI into EVA-02 remarkable accuracy improvements, reaching 93.6% on the CUB-200-2011 dataset and 94.5% on the NA-Birds dataset, respectively, improving over the state-of-the-art method 0.5% and 1.5%.
Qiupu Chen, Feng He 0008, Gang Wang 0023, Xiao Bai 0001, Long Cheng 0003, Xin Ning 0001
IEEE Trans. Fuzzy Syst.4
2025 Unsupervised Recognition of Unknown Objects for Open-World Object Detection
abstract
Open-world object detection (OWOD) extends object detection problem to a realistic and dynamic scenario, where a detection model is required to be capable of detecting both known and unknown objects and incrementally learning newly introduced knowledge. Current OWOD models detect the unknowns that exhibit similar features to the known objects, but they suffer from a severe label bias problem, i.e., they tend to detect all regions (including unknown object regions) that are dissimilar to the known objects as part of the background. To eliminate the label bias, this article proposes a novel module, namely reconstruction error-based Weibull (REW) model, that learns an unsupervised discriminative model for recognizing true unknown objects based on prior knowledge of object occurrence frequency via Weibull modeling. The resulting model can be further refined by another module of our method, called REW-enhanced object localization network (ROLNet), which iteratively extends pseudo-unknown objects to the unlabeled regions. Experimental results show that our method 1) significantly outperforms the prior SOTA in detecting unknown objects while maintaining competitive performance of detecting known object classes on the MS COCO dataset and 2) achieves better generalization ability on the LVIS and Objects365 datasets. Code is available at https://github.com/frh23333/mepu-owod.
Ruohuan Fang, Guansong Pang, Wenjun Miao, Xiao Bai 0001, Xin Ning 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Simple Image-Level Classification Improves Open-Vocabulary Object Detection
abstract
Open-Vocabulary Object Detection (OVOD) aims to detect novel objects beyond a given set of base categories on which the detection model is trained. Recent OVOD methods focus on adapting the image-level pre-trained vision-language models (VLMs), such as CLIP, to a region-level object detection task via, eg., region-level knowledge distillation, regional prompt learning, or region-text pre-training, to expand the detection vocabulary. These methods have demonstrated remarkable performance in recognizing regional visual concepts, but they are weak in exploiting the VLMs' powerful global scene understanding ability learned from the billion-scale image-level text descriptions. This limits their capability in detecting hard objects of small, blurred, or occluded appearance from novel/base categories, whose detection heavily relies on contextual information. To address this, we propose a novel approach, namely Simple Image-level Classification for Context-Aware Detection Scoring (SIC-CADS), to leverage the superior global knowledge yielded from CLIP for complementing the current OVOD models from a global perspective. The core of SIC-CADS is a multi-modal multi-label recognition (MLR) module that learns the object co-occurrence-based contextual information from CLIP to recognize all possible object categories in the scene. These image-level MLR scores can then be utilized to refine the instance-level detection scores of the current OVOD models in detecting those hard objects. This is verified by extensive empirical results on two popular benchmarks, OV-LVIS and OV-COCO, which show that SIC-CADS achieves significant and consistent improvement when combined with different types of OVOD models. Further, SIC-CADS also improves the cross-dataset generalization ability on Objects365 and OpenImages. Code is available at https://github.com/mala-lab/SIC-CADS.
Ruohuan Fang, Guansong Pang, Xiao Bai 0001
AAAI3
2024 Out-of-Distribution Detection in Long-Tailed Recognition with Calibrated Outlier Class Learning
abstract
Existing out-of-distribution (OOD) methods have shown great success on balanced datasets but become ineffective in long-tailed recognition (LTR) scenarios where 1) OOD samples are often wrongly classified into head classes and/or 2) tail-class samples are treated as OOD samples. To address these issues, current studies fit a prior distribution of auxiliary/pseudo OOD data to the long-tailed in-distribution (ID) data. However, it is difficult to obtain such an accurate prior distribution given the unknowingness of real OOD samples and heavy class imbalance in LTR. A straightforward solution to avoid the requirement of this prior is to learn an outlier class to encapsulate the OOD samples. The main challenge is then to tackle the aforementioned confusion between OOD samples and head/tail-class samples when learning the outlier class. To this end, we introduce a novel calibrated outlier class learning (COCL) approach, in which 1) a debiased large margin learning method is introduced in the outlier class learning to distinguish OOD samples from both head and tail classes in the representation space and 2) an outlier-class-aware logit calibration method is defined to enhance the long-tailed classification confidence. Extensive empirical results on three popular benchmarks CIFAR10-LT, CIFAR100-LT, and ImageNet-LT demonstrate that COCL substantially outperforms existing state-of-the-art OOD detection methods in LTR while being able to improve the classification accuracy on ID data. Code is available at https://github.com/mala-lab/COCL.
Wenjun Miao, Guansong Pang, Xiao Bai 0001
AAAI3
2024 DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization
abstract
Radiance fields have demonstrated impressive performance in synthesizing novel views from sparse input views, yet prevailing methods suffer from high training costs and slow inference speed. This paper introduces DNGaussian, a depth-regularized framework based on 3D Gaussian radiance fields, offering real-time and high-quality few-shot novel view synthesis at low costs. Our motivation stems from the highly efficient representation and surprising quality of the recent 3D Gaussian Splatting, despite it will encounter a geometry degradation when input views decrease. In the Gaussian radiance fields, we find this degradation in scene geometry primarily lined to the positioning of Gaussian primitives and can be mitigated by depth constraint. Consequently, we propose a Hard and Soft Depth Regularization to restore accurate scene geometry under coarse monocular depth supervision while maintaining a fine-grained color appearance. To further refine detailed geometry reshaping, we introduce Global-Local Depth Normalization, enhancing the focus on small local depth changes. Extensive experiments on LLFF, DTU, and Blender datasets demonstrate that DNGaussian outperforms state-of-the-art methods, achieving comparable or better results with significantly reduced memory cost, a 25 × reduction in training time, and over 3000 × faster rendering speed. Code is available at: https://github.com/Fictionarry/DNGaussian.
Jiahe Li 0007, Xiao Bai 0001, Xin Ning 0001, Jun Zhou 0001, Lin Gu 0003
CVPR3
2024 Learning Transferable Negative Prompts for Out-of-Distribution Detection
abstract
Existing prompt learning methods have shown certain capabilities in Out-of-Distribution (OOD) detection, but the lack of OOD images in the target dataset in their training can lead to mismatches between OOD images and In-Distribution (ID) categories, resulting in a high false positive rate. To address this issue, we introduce a novel OOD detection method, named ‘NegPrompt’, to learn a set of negative prompts, each representing a negative connotation of a given class label, for delineating the boundaries between ID and OOD images. It learns such negative prompts with ID data only, without any reliance on external out-lier data. Further, current methods assume the availability of samples of all ID classes, rendering them ineffective in open-vocabulary learning scenarios where the inference stage can contain novel ID classes not present during training. In contrast, our learned negative prompts are transferable to novel class labels. Experiments on various ImageNet benchmarks show that NegPrompt surpasses state-of-the-art prompt-learning-based OOD detection methods and maintains a consistent lead in hard OOD detection in closed- and open-vocabulary classification scenarios. Code is available at https://github.com/mala-lab/negprompt.
Guansong Pang, Xiao Bai 0001, Wenjun Miao
CVPR3
2024 Robust Synthetic-to-Real Transfer for Stereo Matching
abstract
With advancements in domain generalized stereo matching networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, few studies have investigated the robustness after fine-tuning them in real-world scenarios, during which the domain generalization ability can be seriously degraded. In this paper, we explore fine-tuning stereo matching networks without compromising their robustness to unseen domains. Our motivation stems from comparing Ground Truth (GT) versus Pseudo Label (PL) for fine-tuning: GT degrades, but PL preserves the domain generalization ability. Empirically, we find the difference between GT and PL implies valuable information that can regularize networks during fine-tuning. We also propose a framework to utilize this difference for fine-tuning, consisting of a frozen Teacher, an exponential moving average (EMA) Teacher, and a Student network. The core idea is to utilize the EMA Teacher to measure what the Student has learned and dynamically improve GT and PL for fine-tuning. We integrate our framework with state-of-the-art networks and evaluate its effectiveness on several real-world datasets. Extensive experiments show that our method effectively preserves the domain generalization ability during fine-tuning. Code is available at: https://github.com/jiaw-z/DKT-Stereo.
Jiahe Li 0007, Lei Huang 0015, Xiaohan Yu 0001, Lin Gu 0003, Xiao Bai 0001
CVPR7
2024 TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting
Jiahe Li 0007, Xiao Bai 0001, Xin Ning 0001, Jun Zhou 0001, Lin Gu 0003
ECCV (10)3
2024 CoR-GS: Sparse-View 3D Gaussian Splatting via Co-regularization
Jiahe Li 0007, Xiaohan Yu 0001, Lei Huang 0015, Lin Gu 0003, Xiao Bai 0001
ECCV (1)7
2024 Occluded Person Retrieval with Hierarchical Feature Optimization
abstract
Occluded person retrieval aims to match images from occluded pedestrians. It pushes forward progress of person retrieval towards applications in real-world scenarios, thus attracting increasing attention in recent years. A key challenge is to learn discriminative representation within limited informative regions due to obstacle or pedestrian occlusion. To that end, we propose a hierarchical feature optimization model (HFO) that jointly optimizes image-level, object-level and part-level features for improved occluded person retrieval. A hierarchical discriminative feature grouping (HDFG) module is developed to generate hierarchical object/part masks for comprehensive feature extraction. Via learning a set of part prototypes, HDFG localizes hierarchical informative object/parts by grouping intermediate feature vectors based on their similarity to these prototypes. The proposed HFO is trained in an end-to-end manner using only identity labels, making it a practical solution for occluded person retrieval. We verify the effectiveness of the proposed method on three challenging occluded datasets and two holistic datasets, i.e., Occluded-DukeMTMC, Occluded-REID, P-DukeMTMC-reID, Market1501, and DukeMTMC-reID. Extensive experiments and ablation studies demonstrate superior or comparable performance of the proposed method over the state-of-the-art methods. The code is available at https://github.com/Patrickzad/HFO.
Yang Zhao 0019, Pengcheng Zhang 0003, Xiaohan Yu 0001, Zhibin Liao, Johan Verjans, Xiao Bai 0001
FG6
2024 Prompting Continual Person Search
abstract
The development of person search techniques has been greatly promoted in recent years for its superior practicality and challenging goals. Despite their significant progress, existing person search models still lack the ability to continually learn from increasing real-world data and adaptively process input from different domains. To this end, this work introduces the continual person search task that sequentially learns on multiple domains and then performs person search on all seen domains. This requires balancing the stability and plasticity of the model to continually learn new knowledge without catastrophic forgetting. For this, we propose a Prompt-based Continual Person Search (PoPS) model in this paper. First, we design a compositional person search transformer to construct an effective pre-trained transformer without exhaustive pre-training from scratch on large-scale person search data. This serves as the fundamental for prompt-based continual learning. On top of that, we design a domain incremental prompt pool with a diverse attribute matching module. For each domain, we independently learn a set of prompts to encode the domain-oriented knowledge. Meanwhile, we jointly learn a group of diverse attribute projections and prototype embeddings to capture discriminative domain attributes. By matching an input image with the learned attributes across domains, the learned prompts can be properly selected for model inference. Extensive experiments are conducted to validate the proposed method for continual person search. The source code is available at https://github.com/PatrickZad/PoPS.
Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001
ACM Multimedia3
2024 Long-Tailed Out-of-Distribution Detection via Normalized Outlier Distribution Adaptation
abstract
One key challenge in Out-of-Distribution (OOD) detection is the absence of ground-truth OOD samples during training. One principled approach to address this issue is to use samples from external datasets as outliers ($\textit{i.e.}$, pseudo OOD samples) to train OOD detectors. However, we find empirically that the outlier samples often present a distribution shift compared to the true OOD samples, especially in Long-Tailed Recognition (LTR) scenarios, where ID classes are heavily imbalanced, $\textit{i.e.}$, the true OOD samples exhibit very different probability distribution to the head and tailed ID classes from the outliers. In this work, we propose a novel approach, namely $\textit{normalized outlier distribution adaptation}$ (AdaptOD), to tackle this distribution shift problem. One of its key components is $\textit{dynamic outlier distribution adaptation}$ that effectively adapts a vanilla outlier distribution based on the outlier samples to the true OOD distribution by utilizing the OOD knowledge in the predicted OOD samples during inference. Further, to obtain a more reliable set of predicted OOD samples on long-tailed ID data, a novel $\textit{dual-normalized energy loss}$ is introduced in AdaptOD, which leverages class- and sample-wise normalized energy to enforce a more balanced prediction energy on imbalanced ID samples. This helps avoid bias toward the head samples and learn a substantially better vanilla outlier distribution than existing energy losses during training. It also eliminates the need of manually tuning the sensitive margin hyperparameters in energy losses. Empirical results on three popular benchmarks for OOD detection in LTR show the superior performance of AdaptOD over state-of-the-art methods. Code is available at https://github.com/mala-lab/AdaptOD.
Wenjun Miao, Guansong Pang, Xiao Bai 0001
NeurIPS4
2024 Efficient Emotional Talking Head Generation via Dynamic 3D Gaussian Rendering
Jiahe Li 0007, Xiao Bai 0001
PRCV (6)3
2024 Exploring the Usage of Pre-trained Features for Stereo Matching
Lei Huang 0015, Xiao Bai 0001, Lin Gu 0003, Edwin R. Hancock
Int. J. Comput. Vis.3
2024 Consistent prototype contrastive learning for weakly supervised person search
Huadong Lin, Xiaohan Yu 0001, Pengcheng Zhang 0003, Xiao Bai 0001
J. Vis. Commun. Image Represent.4
2024 Learning From Human Attention for Attribute-Assisted Visual Recognition
abstract
With prior knowledge of seen objects, humans have a remarkable ability to recognize novel objects using shared and distinct local attributes. This is significant for the challenging tasks of zero-shot learning (ZSL) and fine-grained visual classification (FGVC), where the discriminative attributes of objects have played an important role. Inspired by human visual attention, neural networks have widely exploited the attention mechanism to learn the locally discriminative attributes for challenging tasks. Though greatly promoted the development of these fields, existing works mainly focus on learning the region embeddings of different attribute features and neglect the importance of discriminative attribute localization. It is also unclear whether the learned attention truly matches the real human attention. To tackle this problem, this paper proposes to employ real human gaze data for visual recognition networks to learn from human attention. Specifically, we design a unified Attribute Attention Network (A$^{2}$Net) that learns from human attention for both ZSL and FGVC tasks. The overall model consists of an attribute attention branch and a baseline classification network. On top of the image feature maps provided by the baseline classification network, the attribute attention branch employs attribute prototypes to produce attribute attention maps and attribute features. The attribute attention maps are converted to gaze-like attentions to be aligned with real human gaze attention. To guarantee the effectiveness of attribute feature learning, we further align the extracted attribute features with attribute-defined class embeddings. To facilitate learning from human gaze attention for the visual recognition problems, we design a bird classification game to collect real human gaze data using the CUB dataset via an eye-tracker device. Experiments on ZSL and FGVC tasks without/with real human gaze data validate the benefits and accuracy of our proposed model. This work supports the promising benefits of collecting human gaze datasets and automatic gaze estimation algorithms learning from human attention for high-level computer vision tasks.
Xiao Bai 0001, Pengcheng Zhang 0003, Xiaohan Yu 0001, Edwin R. Hancock, Jun Zhou 0001, Lin Gu 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Learning adversarial semantic embeddings for zero-shot recognition in open worlds
Guansong Pang, Xiao Bai 0001, Lei Zhou 0008, Xin Ning 0001
Pattern Recognit.3
2024 Joint discriminative representation learning for end-to-end person search
Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Chen Wang 0026, Xin Ning 0001
Pattern Recognit.3
2024 Towards effective person search with deep learning: A survey from systematic perspective
Pengcheng Zhang 0003, Xiaohan Yu 0001, Chen Wang 0026, Xin Ning 0001, Xiao Bai 0001
Pattern Recognit.6
2024 3D Person Re-Identification Based on Global Semantic Guidance and Local Feature Aggregation
abstract
Person re-identification (Re-ID) has played an extremely crucial role in ensuring social safety and has attracted considerable research attention. 3D shape information is an important clue to understand the posture and shape of pedestrians. However, most existing person Re-ID methods learn pedestrian feature representations from images, ignoring the real 3D human body structure and the spatial relationship between the pedestrians and interferents. To address this problem, our devise a new point cloud Re-ID network (PointReIDNet), designed to obtain 3D shape representations of pedestrians from point clouds of 3D scenes. The model consists of modules, namely global semantic guidance module and local feature extraction module. The global semantic guidance module is designed by enhancing the point cloud feature representation in similar feature neighborhoods and to reduce the interference caused by 3D shape reconstruction or noise. Further, to provide an efficient representation of point clouds, we propose space cover convolution (SC-Conv), which efficiently encodes information on human shapes in local point clouds by constructing anisotropic geometries in the coordinate neighborhoods. Extensive experiments are conducted on four holistic person Re-ID datasets, one occlusion person Re-ID dataset and one point cloud classification dataset. The results exhibit significant improvements over point-cloud-based person Re-ID methods. In particular, the proposed efficient PointReIDNet decreases the number of parameters from 2.30M to 0.35M with an insignificant drop in performance. The source code is available at: https://github.com/changshuowang/PointReIDNet.
Changshuo Wang 0001, Xin Ning 0001, Weijun Li 0002, Xiao Bai 0001, Xingyu Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Learning Aligned Vertex Convolutional Networks for Graph Classification
abstract
Graph convolutional networks (GCNs) are powerful tools for graph structure data analysis. One main drawback arising in most existing GCN models is that of the oversmoothing problem, i.e., the vertex features abstracted from the existing graph convolution operation have previously tended to be indistinguishable if the GCN model has many convolutional layers (e.g., more than two layers). To address this problem, in this article, we propose a family of aligned vertex convolutional network (AVCN) models that focus on learning multiscale features from local-level vertices for graph classification. This is done by adopting a transitive vertex alignment algorithm to transform arbitrary-sized graphs into fixed-size grid structures. Furthermore, we define a new aligned vertex convolution operation that can effectively learn multiscale vertex characteristics by gradually aggregating local-level neighboring aligned vertices residing on the original grid structures into a new packed aligned vertex. With the new vertex convolution operation to hand, we propose two architectures for the AVCN models to extract different hierarchical multiscale vertex feature representations for graph classification. We show that the proposed models can avoid iteratively propagating redundant information between specific neighboring vertices, restricting the notorious oversmoothing problem arising in most spatial-based GCN models. Experimental evaluations on benchmark datasets demonstrate the effectiveness.
Lixin Cui, Lu Bai 0001, Xiao Bai 0001, Yue Wang 0014, Edwin R. Hancock
IEEE Trans. Neural Networks Learn. Syst.3
2023 Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis
abstract
This paper presents ER-NeRF, a novel conditional Neural Radiance Fields (NeRF) based architecture for talking portrait synthesis that can concurrently achieve fast convergence, real-time rendering, and state-of-the-art performance with small model size. Our idea is to explicitly exploit the unequal contribution of spatial regions to guide talking portrait modeling. Specifically, to improve the accuracy of dynamic head reconstruction, a compact and expressive NeRF-based Tri-Plane Hash Representation is introduced by pruning empty spatial regions with three planar hash encoders. For speech audio, we propose a Region Attention Module to generate region-aware condition feature via an attention mechanism. Different from existing methods that utilize an MLP-based encoder to learn the cross-modal relation implicitly, the attention mechanism builds an explicit connection between audio features and spatial regions to capture the priors of local motions. Moreover, a direct and fast Adaptive Pose Encoding is introduced to optimize the head-torso separation problem by mapping the complex transformation of the head pose into spatial coordinates. Extensive experiments demonstrate that our method renders better high-fidelity and audio-lips synchronized talking portrait videos, with realistic details and high efficiency compared to previous methods. Code is available at https://github.com/Fictionarry/ER-NeRF.
Jiahe Li 0007, Xiao Bai 0001, Jun Zhou 0001, Lin Gu 0003
ICCV3
2023 Adaptive Cost Aggregation in Iterative Depth Estimation for Efficient Multi-view Stereo
Xiang Wang 0014, Xiao Bai 0001, Chen Wang 0026
ICIG (2)2
2023 A smoothing Group Lasso based interval type-2 fuzzy neural network for simultaneous feature selection and system identification
Tao Gao 0003, Chen Wang 0026, Guoqiang Wu, Xin Ning 0001, Xiao Bai 0001, Jian Wang 0010
Knowl. Based Syst.6
2023 Information bottleneck and selective noise supervision for zero-shot learning
Lei Zhou 0008, Yang Liu 0357, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Jun Zhou 0001, Yazhou Yao, Tatsuya Harada, Edwin R. Hancock
Mach. Learn.4
2023 Learning consistent region features for lifelong person re-identification
Jinze Huang, Xiaohan Yu 0001, Dong An 0001, Yaoguang Wei, Xiao Bai 0001, Chen Wang 0026, Jun Zhou 0001
Pattern Recognit.5
2023 Image manipulation detection by multiple tampering traces and edge artifact enhancement
Xun Lin, Shuai Wang 0049, Jiahao Deng, Ying Fu 0001, Xiao Bai 0001, Xinlei Chen, Xiaolei Qu, Wenzhong Tang
Pattern Recognit.5
2023 Hyper-sausage coverage function neuron model and learning algorithm for image classification
abstract
Recently, deep neural networks (DNNs) promote mainly by network architectures and loss functions; however, the development of neuron models has been quite limited. In this study, inspired by the mechanism of human cognition, a hyper-sausage coverage function (HSCF) neuron model possessing a high flexible plasticity. Then, a novel cross-entropy and volume-coverage (CE_VC) loss is defined, which compresses the volume of the hyper-sausage to the hilt, and helps alleviate confusion among different classes, thus ensuring the intra-class compactness of the samples. Finally, a divisive iteration method is introduced, which considers each neuron model as a weak classifier, and iteratively increases the number of weak classifiers. Thus, the optimal number of the HSCF neuron is adaptively determined and an end-to-end learning framework is constructed. In particular, to improve the classification performance, the HSCF neuron can be applied to classical DNNs. Comprehensive experiments on eight datasets in several domains demonstrate the effectiveness of the proposed method. The proposed method exhibits the feasibility of boosting DNNs with neuron plasticity and provides a novel perspective for further developments in DNNs. The source code is available at https://github.com/Tough2011/HSCFNet.git .
Xin Ning 0001, Weijuan Tian, Feng He 0008, Xiao Bai 0001, Le Sun 0003, Weijun Li 0002
Pattern Recognit.4
2023 Corrigendum to' HCFNN: High-order coverage function neural network for image classification' Pattern Recognition. Volume 131(2022) 108873
Xin Ning 0001, Weijuan Tian, Zaiyang Yu, Weijun Li 0002, Xiao Bai 0001, Yuebao Wang
Pattern Recognit.5
2023 Attribute subspaces for zero-shot learning
Lei Zhou 0008, Yang Liu 0357, Xiao Bai 0001, Na Li 0014, Xiaohan Yu 0001, Jun Zhou 0001, Edwin R. Hancock
Pattern Recognit.3
2023 Stereo Attention Cross-Decoupling Fusion-Guided Federated Neural Learning for Hyperspectral Image Classification
abstract
Federated learning is a promising solution in several industries for co-training models among distributed clients via centralized servers without leaving private user data on the devices. Thus, federated learning can be seen as a stimulus for the edge computing paradigm as it supports collaborative learning and model optimization. In view of the strict requirements for data security and system reliability of hyperspectral classification techniques for surveillance, aerospace, and military missions, this paper proposes a novel stereo attention cross-decoupling fusion-guided federated neural learning algorithm for hyperspectral image classification, which first trains client devices using a scalable federated learning approach consisting of master server, secure aggregator and edge client devices of a certain size.The distributed devices train local models of the neural network for classifying hyperspectral images and send them to the secure aggregator, which aggregates the local models using a weighted averaging strategy and sends them to the master server for iteration. In addition, the stereo attention cross-decoupling fusion module is used to mine the multidimensional spatial details of the hyperspectral images, specifically by first extracting the most discriminative features from different directions (horizontal, vertical, and spatial) using the attention mechanism, and then using the decoupling fusion strategy to classify the original feature map into three levels: significant, minor, and redundant, and use them to model the multidimensional spatial relationships, thus strengthening the capability to represent features. Extensive experiments on several public datasets have shown that the proposed method provides competitive performance and, more importantly, is effective in enhancing privacy and reliability for hyperspectral image classification.
Weiwei Cai 0001, Ming Gao 0026, Yao Ding 0010, Xin Ning 0001, Xiao Bai 0001, Pengjiang Qian
IEEE Trans. Geosci. Remote. Sens.5
2023 A Novel Hyperspectral Image Classification Model Using Bole Convolution With Three-Direction Attention Mechanism: Small Sample and Unbalanced Learning
abstract
Currently, the use of rich spectral and spatial information of hyperspectral images (HSIs) to classify ground objects is a research hotspot. However, the classification ability of existing models is significantly affected by its high data dimensionality and massive information redundancy. Therefore, we focus on the elimination of redundant information and the mining of promising features and propose a novel Bole convolution (BC) neural network with a tandem three-direction attention (TDA) mechanism (BTA-Net) for the classification of HSI. A new BC is proposed for the first time in this algorithm, whose core idea is to enhance effective features and eliminate redundant features through feature punishment and reward strategies. Considering that traditional attention mechanisms often assign weights in a one-direction manner, leading to a loss of the relationship between the spectra, a novel three-direction (horizontal, vertical, and spatial directions) attention mechanism is proposed, and an addition strategy and a maximization strategy are used to jointly assign weights to improve the context sensitivity of spatial–spectral features. In addition, we also designed a tandem TDA mechanism module and combined it with a multiscale BC output to improve classification accuracy and stability even when training samples are small and unbalanced. We conducted scene classification experiments on four commonly used hyperspectral datasets to demonstrate the superiority of the proposed model. The proposed algorithm achieves competitive performance on small samples and unbalanced data, according to the results of comparison and ablation experiments. The source code for BTA-Net can be found athttps://github.com/vivitsai/BTA-Net.
Weiwei Cai 0001, Xin Ning 0001, Guoxiong Zhou, Xiao Bai 0001, Yizhang Jiang, Wei Li 0032, Pengjiang Qian
IEEE Trans. Geosci. Remote. Sens.4
2023 Graph-Structured Convolution-Guided Continuous Context Threshold-Aware Networks for Hyperspectral Image Classification
abstract
Although convolutional neural networks (CNNs) have shown superior performance to traditional machine learning algorithms for hyperspectral image classification tasks, the ability of traditional CNNs to model remote dependencies in the spatial orientation of HSIs is still limited, and they always extract similar low-level features, leading to feature redundancy. To cope with this limitation, this paper proposes a novel multi-order statistical representation-guided graph convolution and continuous context threshold-aware network for the classification of hyperspectral images with limited training samples. Initially, the spectral spatial information is separately modeled using first-order features and second-order pooling operators. Secondly, we propose graph-structuring the patch’s features. By employing a random walk transition probability matrix, graph-structured convolution can mine more discriminative direction features. In addition, we design a continuous context threshold-aware network to model multidimensional spatial relationships, thereby enhancing the representation of graph features. Specifically, the cross-attention mechanism is used to calculate the attention weights in the vertical and horizontal directions, and the features are divided into two levels—important and secondary—by solving the cosine distance between feature vectors, and the former is retained and the latter is punished. Extensive experiments on multiple HSIs datasets demonstrated that the proposed method delivers competitive performance. The code will be available at: https://github.com/vivitsai/GSC-CCTA.
Weiwei Cai 0001, Pengjiang Qian, Yao Ding 0010, Meiqiao Bi, Xin Ning 0001, Danfeng Hong, Xiao Bai 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 Revisiting Domain Generalized Stereo Matching Networks from a Feature Consistency Perspective
abstract
Despite recent stereo matching networks achieving impressive performance given sufficient training data, they suffer from domain shifts and generalize poorly to unseen domains. We argue that maintaining feature consistency between matching pixels is a vital factor for promoting the generalization capability of stereo matching networks, which has not been adequately considered. Here we address this issue by proposing a simple pixel-wise contrastive learning across the viewpoints. The stereo contrastive feature loss function explicitly constrains the consistency between learned features of matching pixel pairs which are observations of the same 3D points. A stereo selective whitening loss is further introduced to better preserve the stereo feature consistency across domains, which decorrelates stereo features from stereo viewpoint-specific style information. Counter-intuitively, the generalization of feature consistency between two viewpoints in the same scene translates to the generalization of stereo matching performance to unseen domains. Our method is generic in nature as it can be easily embedded into existing stereo networks and does not require access to the samples in the target domain. When trained on synthetic data and generalized to four real-world testing sets, our method achieves superior performance over several state-of-the-art networks. The code is available online11https://github.com/jiaw-z/FCStereo.
Xiang Wang 0014, Xiao Bai 0001, Chen Wang 0026, Lei Huang 0015, Lin Gu 0003, Jun Zhou 0001, Tatsuya Harada, Edwin R. Hancock
CVPR3
2022 Where to Focus: Investigating Hierarchical Attention Relationship for Fine-Grained Visual Classification
Yang Liu 0357, Lei Zhou 0008, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Xiaohan Yu 0001, Jun Zhou 0001, Edwin R. Hancock
ECCV (24)4
2022 Unsupervised feature selection via adaptive autoencoder with redundancy control
Xiaoling Gong, Jian Wang 0010, Kai Zhang 0029, Xiao Bai 0001, Nikhil R. Pal
Neural Networks5
2022 A modified interval type-2 Takagi-Sugeno fuzzy neural network and its convergence analysis
Tao Gao 0003, Xiao Bai 0001, Chen Wang 0026, Liang Zhang 0044, Jian Wang 0010
Pattern Recognit.2
2022 HCFNN: High-order coverage function neural network for image classification
Xin Ning 0001, Weijuan Tian, Zaiyang Yu, Weijun Li 0002, Xiao Bai 0001, Yuebao Wang
Pattern Recognit.5
2022 Uncertainty estimation for stereo matching based on evidential deep learning
Chen Wang 0026, Xiang Wang 0014, Liang Zhang 0044, Xiao Bai 0001, Xin Ning 0001, Jun Zhou 0001, Edwin R. Hancock
Pattern Recognit.5
2022 Learning Discriminative Features by Covering Local Geometric Space for Point Cloud Analysis
abstract
At present, effectively aggregating and transferring the local features of point cloud is still an unresolved technological conundrum. In this study, we propose a new space-cover convolutional neural network (SC-CNN) for tasks such as point cloud classification and segmentation. The core of this network is space-cover convolution (SC-Conv), which implements depthwise separable convolution on the point cloud. In addition, a newly designed space-cover operator (SCOP) replaces depthwise convolution. The key to SC-Conv is constructing anisotropic spatial geometry in the local point cloud. The SCOP achieves this by utilizing the positional and feature relationships to learn the high-order relationship expression between points. First, data-driven adaptive learning from the 3-D coordinate relationship between the local points is used to determine the weight of the SCOP. Then, the edge feature of the neighboring point relative to the sampling point is used as the input of the SCOP. Finally, a deformable spatial geometry is constructed in the feature space between local points to aggregate the local high-order features. By stacking SC-Conv to construct SC-CNN with a hierarchical network structure for point cloud analysis, we can better perceive the shape information of point cloud and improve network robustness. Finally, we provide numerous experiments to verify that SC-CNN parallels or even outperforms advanced methods in shape classification, part segmentation, and large-scale indoor scene segmentation tasks. The open-source code was published athttps://github.com/changshuowang/SC-CNN.
Changshuo Wang 0001, Xin Ning 0001, Linjun Sun, Liping Zhang 0014, Weijun Li 0002, Xiao Bai 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Spectral-Spatial Boundary Detection in Hyperspectral Images
abstract
In this paper, we propose a novel method for boundary detection in close-range hyperspectral images. This method can effectively predict the boundaries of objects of similar colour but different materials. To effectively extract the material information in the image, the spatial distribution of the spectral responses of different materials or endmembers is first estimated by hyperspectral unmixing. The resulting abundance map represents the fraction of each endmember spectra at each pixel. The abundance map is used as a supportive feature such that the spectral signature and the abundance vector for each pixel are fused to form a new spectral feature vector. Then different spectral similarity measures are adopted to construct a sparse spectral-spatial affinity matrix that characterizes the similarity between the spectral feature vectors of neighbouring pixels within a local neighborhood. After that, a spectral clustering method is adopted to produce eigenimages. Finally, the boundary map is constructed from the most informative eigenimages. We created a new HSI dataset and use it to compare the proposed method with four alternative methods, one for hyperspectral image and three for RGB image. The results exhibit that our method outperforms the alternatives and can cope with several scenarios that methods based on colour images cannot handle.
Suhad Lateef Al-Khafaji, Jun Zhou 0001, Xiao Bai 0001, Yuntao Qian, Alan Wee-Chung Liew
IEEE Trans. Image Process.3
2022 Universal Adversarial Patch Attack for Automatic Checkout Using Perceptual and Attentional Bias
abstract
Adversarial examples are inputs with imperceptible perturbations that easily mislead deep neural networks (DNNs). Recently, adversarial patch, with noise confined to a small and localized patch, has emerged for its easy feasibility in real-world scenarios. However, existing strategies failed to generate adversarial patches with strong generalization ability due to the ignorance of the inherent biases of models. In other words, the adversarial patches are always input-specific and fail to attack images from all classes or different models, especially unseen classes and black-box models. To address the problem, this paper proposes a bias-based framework to generate universal adversarial patches with strong generalization ability, which exploits the perceptual bias and attentional bias to improve the attacking ability. Regarding the perceptual bias, since DNNs are strongly biased towards textures, we exploit the hard examples which convey strong model uncertainties and extract a textural patch prior from them by adopting the style similarities. The patch prior is closer to decision boundaries and would promote attacks across classes. As for the attentional bias, motivated by the fact that different models share similar attention patterns towards the same image, we exploit this bias by confusing the model-shared similar attention patterns. Thus, the generated adversarial patches can obtain stronger transferability among different models. Taking Automatic Check-out (ACO) as the typical scenario, extensive experiments including white-box/black-box settings in both digital-world (RPC, the largest ACO related dataset) and physical-world scenario (Taobao and JD, the world's largest online shopping platforms) are conducted. Experimental results demonstrate that our proposed framework outperforms state-of-the-art adversarial patch attack methods.
Jiakai Wang, Aishan Liu, Xiao Bai 0001, Xianglong Liu 0001
IEEE Trans. Image Process.3
2022 Beyond Triplet Loss: Person Re-Identification With Fine-Grained Difference-Aware Pairwise Loss
abstract
Person Re-IDentification (ReID) aims at re-identifying persons from different viewpoints across multiple cameras. Capturing the fine-grained appearance differences is often the key to accurate person ReID, because many identities can be differentiated only when looking into these fine-grained differences. However, most state-of-the-art person ReID approaches, typically driven by a triplet loss, fail to effectively learn the fine-grained features as they are focused more on differentiating large appearance differences. To address this issue, we introduce a novel pairwise loss function that enables ReID models to learn the fine-grained features by adaptively enforcing an exponential penalization on the images of small differences and a bounded penalization on the images of large differences. The proposed loss is generic and can be used as a plugin to replace the triplet loss to significantly enhance different types of state-of-the-art approaches. Experimental results on four benchmark datasets show that the proposed loss substantially outperforms a number of popular loss functions by large margins; and it also enables significantly improved data efficiency.
Guansong Pang, Xiao Bai 0001, Changhong Liu, Xin Ning 0001, Lin Gu 0003, Jun Zhou 0001
IEEE Trans. Multim.3
2022 Transductive Relation-Propagation With Decoupling Training for Few-Shot Learning
abstract
Few-shot learning, aiming to learn novel concepts from one or a few labeled examples, is an interesting and very challenging problem with many practical advantages. Existing few-shot methods usually utilize data of the same classes to train the feature embedding module and in a row, which is unable to learn adapting to new tasks. Besides, traditional few-shot models fail to take advantage of the valuable relations of the support-query pairs, leading to performance degradation. In this article, we propose a transductive relation-propagation graph neural network (GNN) with a decoupling training strategy (TRPN-D) to explicitly model and propagate such relations across support-query pairs, and empower the few-shot module the ability of transferring past knowledge to new tasks via the decoupling training. Our few-shot module, namely TRPN, treats the relation of each support-query pair as a graph node, named relational node, and resorts to the known relations between support samples, including both intraclass commonality and interclass uniqueness. Through relation propagation, the model could generate the discriminative relation embeddings for support-query pairs. To the best of our knowledge, this is the first work that decouples the training of the embedding network and the few-shot graph module with different tasks, which might offer a new way to solve the few-shot learning problem. Extensive experiments conducted on several benchmark datasets demonstrate that our method can significantly outperform a variety of state-of-the-art few-shot learning methods.
Yuqing Ma, Shihao Bai, Wei Liu 0005, Shuo Wang 0008, Yue Yu 0001, Xiao Bai 0001, Xianglong Liu 0001, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2021 Goal-Oriented Gaze Estimation for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen classes. Since semantic knowledge is built on attributes shared between different classes, which are highly local, strong prior for localization of object attribute is beneficial for visual-semantic embedding. Interestingly, when recognizing unseen images, human would also automatically gaze at regions with certain semantic clue. Therefore, we introduce a novel goal-oriented gaze estimation module (GEM) to improve the discriminative attribute localization based on the class-level attributes for ZSL. We aim to predict the actual human gaze location to get the visual attention regions for recognizing a novel object guided by attribute description. Specifically, the task-dependent attention is learned with the goal-oriented GEM, and the global image features are simultaneously optimized with the regression of local attribute features. Experiments on three ZSL benchmarks, i.e., CUB, SUN and AWA2, show the superiority or competitiveness of our proposed method against the state-of-the-art ZSL methods. The ablation analysis on real gaze data CUB-VWSW also validates the benefits and accuracy of our gaze estimation module. This work implies the promising benefits of collecting human gaze dataset and automatic gaze estimation algorithms on high-level computer vision tasks. The code is available at https://github.com/osierboy/GEM-ZSL.
Yang Liu 0357, Lei Zhou 0008, Xiao Bai 0001, Yifei Huang 0002, Lin Gu 0003, Jun Zhou 0001, Tatsuya Harada
CVPR3
2021 Occluded Person Re-Identification with Single-scale Global Representations
abstract
Occluded person re-identification (ReID) aims at re-identifying occluded pedestrians from occluded or holistic images taken across multiple cameras. Current state-of-the-art (SOTA) occluded ReID models rely on some auxiliary modules, including pose estimation, feature pyramid and graph matching modules, to learn multi-scale and/or part-level features to tackle the occlusion challenges. This unfortunately leads to complex ReID models that (i) fail to generalize to challenging occlusions of diverse appearance, shape or size, and (ii) become ineffective in handling non-occluded pedestrians. However, real-world ReID applications typically have highly diverse occlusions and involve a hybrid of occluded and non-occluded pedestrians. To address these two issues, we introduce a novel ReID model that learns discriminative single-scale global-level pedestrian features by enforcing a novel exponentially sensitive yet bounded distance loss on occlusion-based augmented data. We show for the first time that learning single-scale global features without using these auxiliary modules is able to outperform the SOTA multi-scale and/or part-level feature-based models. Further, our simple model can achieve new SOTA performance in both occluded and non-occluded ReID, as shown by extensive results on three occluded and two general ReID benchmarks. Additionally, we create a large-scale occluded person ReID dataset with various occlusions in different scenes, which is significantly larger and contains more diverse occlusions and pedestrian dressings than existing occluded ReID datasets, providing a more faithful occluded ReID benchmark. The dataset is available at: https://git.io/OPReID
Guansong Pang, Jile Jiao, Xiao Bai 0001, Xuetao Feng, Chunhua Shen
ICCV4
2021 Relation-Aware Reasoning with Graph Convolutional Network
Lei Zhou 0008, Yang Liu 0357, Xiao Bai 0001, Xiang Wang 0014, Chen Wang 0026, Liang Zhang 0044, Lin Gu 0003
ICIG (1)3
2021 A Joint Convolutional Neural Network for Simultaneous Despeckling and Classification of SAR Targets
abstract
Deep learning (DL) techniques recently have attracted much attention in the synthetic aperture radar (SAR) automatic target recognition (ATR). Due to the coherent imaging pattern, SAR images inherently suffer from the speckle noise. To mitigate its influence, this letter proposes a joint convolutional neural network (J-CNN) for simultaneous despeckling and classification of SAR targets. It integrates a two-step process in the CNN framework but without the pooling operation during the despeckling phase. Then, a new loss function is introduced, and its partial derivatives with respect to weights are given for the training of J-CNN. Finally, comparative experiments with some classical network models are carried out based on synthetic SAR target images. The results demonstrate that the proposed method not only significantly outperforms other models under strong speckle noise condition but also has an efficient architecture with fewer weight parameters.
Tong Zheng 0004, Jun Wang 0041, Xiao Bai 0001
IEEE Geosci. Remote. Sens. Lett.4
2021 Explainable deep learning for efficient and robust pattern recognition: A survey of recent developments
Xiao Bai 0001, Xiang Wang 0014, Xianglong Liu 0001, Qiang Liu 0001, Jingkuan Song, Nicu Sebe, Been Kim
Pattern Recognit.1
2021 Feature Refinement and Filter Network for Person Re-Identification
abstract
In the task of person re-identification, the attention mechanism and fine-grained information have been proved to be effective. However, it has been observed that models often focus on the extraction of features with strong discrimination, and neglect other valuable features. The extracted fine-grained information may include redundancies. In addition, current methods lack an effective scheme to remove background interference. Therefore, this paper proposes the feature refinement and filter network to solve the above problems from three aspects: first, by weakening the high response features, we aim to identify highly valuable features and extract the complete features of persons, thereby enhancing the robustness of the model; second, by positioning and intercepting the high response areas of persons, we eliminate the interference arising from background information and strengthen the response of the model to the complete features of persons; finally, valuable fine-grained features are selected using a multi-branch attention network for person re-identification to enhance the performance of the model. Our extensive experiments on the benchmark Market-1501, DukeMTMC-reID, CUHK03 and MSMT17 person re-identification datasets demonstrate that the performance of our method is comparable to that of state-of-the-art approaches.
Xin Ning 0001, Weijun Li 0002, Liping Zhang 0014, Xiao Bai 0001, Shengwei Tian
IEEE Trans. Circuits Syst. Video Technol.5
2021 Self-Supervised Multiscale Adversarial Regression Network for Stereo Disparity Estimation
abstract
Deep learning approaches have significantly contributed to recent progress in stereo matching. These deep stereo matching methods are usually based on supervised training, which requires a large amount of high-quality ground-truth depth map annotations that are expensive to collect. Furthermore, only a limited quantity of stereo vision training data are currently available, obtained either by active sensors (Lidar and ToF cameras) or through computer graphics simulations and not meeting requirements for deep supervised training. Here, we propose a novel deep stereo approach called the "self-supervised multiscale adversarial regression network (SMAR-Net)," which relaxes the need for ground-truth depth maps for training. Specifically, we design a two-stage network. The first stage is a disparity regressor, in which a regression network estimates disparity values from stacked stereo image pairs. Stereo image stacking method is a novel contribution as it not only contains the spatial appearances of stereo images but also implies matching correspondences with different disparity values. In the second stage, a synthetic left image is generated based on the left-right consistency assumption. Our network is trained by minimizing a hybrid loss function composed of a content loss and an adversarial loss. The content loss minimizes the average warping error between the synthetic images and the real ones. In contrast to the generative adversarial loss, our proposed adversarial loss penalizes mismatches using multiscale features. This constrains the synthetic image and real image as being pixelwise identical instead of just belonging to the same distribution. Furthermore, the combined utilization of multiscale feature extraction in both the content loss and adversarial loss further improves the adaptability of SMAR-Net in ill-posed regions. Experiments on multiple benchmark datasets show that SMAR-Net outperforms the current state-of-the-art self-supervised methods and achieves comparable outcomes to supervised methods. The source code can be accessed at: https://github.com/Dawnstar8411/SMAR-Net.
Chen Wang 0026, Xiao Bai 0001, Xiang Wang 0014, Xianglong Liu 0001, Jun Zhou 0001, Xinyu Wu 0001, Hongdong Li, Dacheng Tao
IEEE Trans. Cybern.2
2020 Adaptive Unimodal Cost Volume Filtering for Deep Stereo Matching
abstract
State-of-the-art deep learning based stereo matching approaches treat disparity estimation as a regression problem, where loss function is directly defined on true disparities and their estimated ones. However, disparity is just a byproduct of a matching process modeled by cost volume, while indirectly learning cost volume driven by disparity regression is prone to overfitting since the cost volume is under constrained. In this paper, we propose to directly add constraints to the cost volume by filtering cost volume with unimodal distribution peaked at true disparities. In addition, variances of the unimodal distributions for each pixel are estimated to explicitly model matching uncertainty under different contexts. The proposed architecture achieves state-of-the-art performance on Scene Flow and two KITTI stereo benchmarks. In particular, our method ranked the 1st place of KITTI 2012 evaluation and the 4th place of KITTI 2015 evaluation (recorded on 2019.8.20). The codes of AcfNet are available at: https://github.com/youmi-zym/AcfNet.
Youmin Zhang 0005, Xiao Bai 0001, Suihanjin Yu, Zhiwei Li 0006, Kuiyuan Yang
AAAI3
2020 Self-Trained Deep Ordinal Regression for End-to-End Video Anomaly Detection
abstract
Video anomaly detection is of critical practical importance to a variety of real applications because it allows human attention to be focused on events that are likely to be of interest, in spite of an otherwise overwhelming volume of video. We show that applying self-trained deep ordinal regression to video anomaly detection overcomes two key limitations of existing methods, namely, 1) being highly dependent on manually labeled normal training data; and 2) sub-optimal feature learning. By formulating a surrogate two-class ordinal regression task we devise an end-to-end trainable video anomaly detection approach that enables joint representation learning and anomaly scoring without manually labeled normal/abnormal data. Experiments on eight real-world video scenes show that our proposed method outperforms state-of-the-art methods that require no labeled training data by a substantial margin, and enables easy and accurate localization of the identified anomalies. Furthermore, we demonstrate that our method offers effective human-in-the-loop anomaly detection which can be critical in applications where anomalies are rare and the false-negative cost is high.
Guansong Pang, Chunhua Shen, Anton van den Hengel, Xiao Bai 0001
CVPR5
2020 Matrix Classifier On Dynamic Functional Connectivity For Mci Identification
abstract
One of the most popular method for Alzheimer's disease (AD) diagnosis is exploring the Brain functional connectivity (FC) from resting-state functional magnetic resonance imaging (RS-fMRI). To early prevent AD, it is crucial to distinguish AD and and its preclinical stage, mild cognitive impairment (MCI) and early MCI (eMCI). In many existing works, dynamic functional connectivity (dFC) which contains rich spatiotemporal information has been exploited for the MCI and eMCI identification. However, most of these dFC based methods only consider the correlation between discrete brain status while ignore the valuable spatiotemporal information contained in dFC. To overcome this limitation, we propose a matrix classifier based method on the dFC signal for MCI and eMCI identification. Specifically, we first represent the dFC correlations by matrix features which contain rich spatiotemporal information and then learn the support matrix machines (SMM) to classify AD and its preclinical stage. Experiments on 600 real people data provide by the Alzheimer's Disease Neuroimaging Initiative (ADNI) demonstrate that our proposed matrix classifier based method outperforms other FC and dFC based methods for both normal controls (NC)/MCI identification and NC/eMCI identification.
Lei Zhou 0008, Liang Zhang 0044, Xiao Bai 0001, Jun Zhou 0001
ICIP3
2020 Fast Subspace Clustering Based on the Kronecker Product
abstract
Subspace clustering is a useful technique for many computer vision applications in which the intrinsic dimension of high-dimensional data is often smaller than the ambient dimension. Spectral clustering, as one of the main approaches to subspace clustering, often takes on a sparse representation or a low-rank representation to learn a block diagonal self-representation matrix for subspace generation. However, existing methods require solving a large scale convex optimization problem with a large set of data, with computational complexity reaches O(N3) for N data points. Therefore, the efficiency and scalability of traditional spectral clustering methods can not be guaranteed for large scale datasets. In this paper, we propose a subspace clustering model based on the Kronecker product. Due to the property that the Kronecker product of a block diagonal matrix with any other matrix is still a block diagonal matrix, we can efficiently learn the representation matrix which is formed by the Kronecker product of k smaller matrices. By doing so, our model significantly reduces the computational complexity to O(kN3/k). Furthermore, our model is general in nature, and can be adapted to different regularization based subspace clustering methods. Experimental results on two public datasets show that our model significantly improves the efficiency compared with several state-of-the-art methods. Moreover, we have conducted experiments on synthetic data to verify the scalability of our model for large scale datasets.
Lei Zhou 0008, Xiao Bai 0001, Liang Zhang 0044, Jun Zhou 0001, Edwin R. Hancock
ICPR2
2020 HMFlow: Hybrid Matching Optical Flow Network for Small and Fast-Moving Objects
abstract
In optical flow estimation task, coarse-to-fine warping strategy is widely used to deal with the large displacement problem and provides efficiency and speed. However, limited by the small search range between the first images and warped second images, current coarse-to-fine optical flow networks fail to capture small and fast-moving objects which has disappeared at coarse resolution levels. To address this problem, we introduce a lightweight but effective Global Matching Component (GMC) to grab global matching features. We propose a new Hybrid Matching Optical Flow Network (HMFlow) by integrating GMC into existing coarse-to-fine networks seamlessly. Besides keeping in high accuracy and small model size, our proposed HMFlow can apply global matching features to guide the network to discover the small and fast-moving objects mismatched by local matching features. We also build a new dataset, named SFChairs, for evaluation. The experimental results show that our proposed network achieves considerable performance, especially at regions with small and fast-moving objects.
Suihanjin Yu, Youmin Zhang 0005, Chen Wang 0026, Xiao Bai 0001, Liang Zhang 0044, Edwin R. Hancock
ICPR4
2020 Binary neural networks: A survey
Haotong Qin, Ruihao Gong, Xianglong Liu 0001, Xiao Bai 0001, Jingkuan Song, Nicu Sebe
Pattern Recognit.4
2020 Multi-head enhanced self-attention network for novelty detection
abstract
One-class classification (OCC) is a classical problem in computer vision that can be described as the task of classifying outlier class samples (OC samples) from the OCC model trained on inlier class samples (IC samples) when datasets are highly biased toward one class due to the insufficient sample size of the other class. Currently, the adversarial learning OCC (ALOCC) method has been proven to significantly improve OCC performance. However, its drawbacks include instability issues and non-evident reconstruction between the IC and OC samples. Therefore, we propose multihead enhanced self-attention in the ALOCC network, thereby increasing the difference between the IC and OC samples and significantly increasing OCC accuracy compared with ALOCC accuracy. For training, we propose a new loss, called adversarial-balance loss, that effectively solves the training instability problem, further increasing OCC accuracy. The experiments show the effectiveness of the proposed method compared with state-of-art methods.
Yuxin Gong, Haogang Zhu, Xiao Bai 0001, Wenzhong Tang
Pattern Recognit.4
2020 Learning binary code for fast nearest subspace search
Lei Zhou 0008, Xiao Bai 0001, Xianglong Liu 0001, Jun Zhou 0001, Edwin R. Hancock
Pattern Recognit.2
2020 Local-global nested graph kernels using nested complexity traces
Lu Bai 0001, Lixin Cui, Luca Rossi 0004, Lixiang Xu, Xiao Bai 0001, Edwin R. Hancock
Pattern Recognit. Lett.5
2020 Special issue on recent advances in statistical, structural and syntactic pattern recognition
Xiao Bai 0001, Edwin R. Hancock, Richard C. Wilson 0001, Tin Kam Ho
Pattern Recognit. Lett.1
2020 Distributed Complementary Binary Quantization for Joint Hash Table Learning
abstract
Building multiple hash tables serves as a very successful technique for gigantic data indexing, which can simultaneously guarantee both the search accuracy and efficiency. However, most of existing multitable indexing solutions, without informative hash codes and strong table complementarity, largely suffer from the table redundancy. To address the problem, we propose a complementary binary quantization (CBQ) method for jointly learning multiple tables and the corresponding informative hash functions in a centralized way. Based on CBQ, we further design a distributed learning algorithm (D-CBQ) to accelerate the training over the large-scale distributed data set. The proposed (D-)CBQ exploits the power of prototype-based incomplete binary coding to well align the data distributions in the original space and the Hamming space and further utilizes the nature of multi-index search to jointly reduce the quantization loss. (D-)CBQ possesses several attractive properties, including the extensibility for generating long hash codes in the product space and the scalability with linear training time. Extensive experiments on two popular large-scale tasks, including the Euclidean and semantic nearest neighbor search, demonstrate that the proposed (D-)CBQ enjoys efficient computation, informative binary quantization, and strong table complementarity, which together help significantly outperform the state of the arts, with up to 57.76% performance gains relatively.
Xianglong Liu 0001, Qiang Fu 0006, Deqing Wang 0001, Xiao Bai 0001, Xinyu Wu 0001, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.4
2019 Discriminative Features Matter: Multi-layer Bilinear Pooling for Camera Localization
Xiang Wang 0014, Chen Wang 0026, Xiao Bai 0001, Jing Wu 0004, Edwin R. Hancock
BMVC4
2019 Accelerating Deep Convnets via Sparse Subspace Clustering
Shengge Shi, Xiao Bai 0001, Xueni Zhang
ICIG (2)3
2019 A One-step Pruning-recovery Framework for Acceleration of Convolutional Neural Networks
abstract
Acceleration of convolutional neural network has received increasing attention during the past several years. Among various acceleration techniques, filter pruning has its inherent merit by effectively reducing the number of convolution filters. However, most filter pruning methods resort to tedious and time-consuming layer-by-layer pruning-recovery strategy to avoid a significant drop of accuracy. In this paper, we present an efficient filter pruning framework to solve this problem. Our method accelerates the network in one-step pruning-recovery manner with a novel optimization objective function, which achieves higher accuracy with much less cost compared with existing pruning methods. Furthermore, our method allows network compression with global filter pruning. Given a global pruning rate, it can adaptively determine the pruning rate for each single convolutional layer, while these rates are often set as hyper-parameters in previous approaches. Evaluated on VGG- 16 and ResNet-50 using ImageNet, our approach outperforms several state-of-the-art methods with less accuracy drop under the same and even much fewer floating-point operations (FLOPs).
Xiao Bai 0001, Lei Zhou 0008, Jun Zhou 0001
ICTAI2
2019 Hyperspectral Image Classification Based on Non-Local Neural Networks
abstract
Deep convolutional neural network has been used for pixel-wise hyperspectral image classification. However, convolutional operations only extract features from local neighborhood at a time, which is inefficient to capture long-range dependencies. On the other hand, the lack of training samples often leads to over-fitting problem. In this paper, we proposed a neural network which is formed by sequential local and non-local operation blocks. The proposed network takes hyperspectral image as input and outputs the class inference of each pixel. The local operation module extracts local spatial and spectral features. The non-local operation module computes the response at a position as a weighted sum of the features at all positions. So it can capture long-range dependencies without stacking deep layers. Experiments on two public datasets show that our proposed method outperforms several state-of-the-art methods using limited number of training samples.
Chen Wang 0031, Xiao Bai 0001, Lei Zhou 0008, Jun Zhou 0001
IGARSS2
2019 Latent Distribution Preserving Deep Subspace Clustering
abstract
Subspace clustering is a useful technique for many computer vision applications in which the intrinsic dimension of high-dimensional data is smaller than the ambient dimension. Traditional subspace clustering methods often rely on the self-expressiveness property, which has proven effective for linear subspace clustering. However, they perform unsatisfactorily on real data with complex nonlinear subspaces. More recently, deep autoencoder based subspace clustering methods have achieved success owning to the more powerful representation extracted by the autoencoder network. Unfortunately, these methods only considering the reconstruction of original input data can hardly guarantee the latent representation for the data distributed in subspaces, which inevitably limits the performance in practice. In this paper, we propose a novel deep subspace clustering method based on a latent distribution-preserving autoencoder, which introduces a distribution consistency loss to guide the learning of distribution-preserving latent representation, and consequently enables strong capacity of characterizing the real-world data for subspace clustering. Experimental results on several public databases show that our method achieves significant improvement compared with the state-of-the-art subspace clustering methods.
Lei Zhou 0008, Xiao Bai 0001, Xianglong Liu 0001, Jun Zhou 0001, Edwin R. Hancock
IJCAI2
2019 Deep Hashing by Discriminating Hard Examples
abstract
This paper tackles a rarely explored but critical problem within learning to hash, i.e., to learn hash codes that effectively discriminate hard similar and dissimilar examples, to empower large-scale image retrieval. Hard similar examples refer to image pairs from the same semantic class that demonstrate some shared appearance but have different fine-grained appearance. Hard dissimilar examples are image pairs that come from different semantic classes but exhibit similar appearance. These hard examples generally have a small distance due to the shared appearance. Therefore, effective encoding of the hard examples can well discriminate the relevant images within a small Hamming distance, enabling more accurate retrieval in the top-ranked returned images. However, most existing hashing methods cannot capture this key information as their optimization is dominated byeasy examples, i.e., distant similar/dissimilar pairs that share no or limited appearance. To address this problem, we introduce a novel Gamma distribution-enabled and symmetric Kullback-Leibler divergence-based loss, which is dubbed dual hinge loss because it works similarly as imposing two smoothed hinge losses on the respective similar and dissimilar pairs. Specifically, the loss enforces exponentially variant penalization on the hard similar (dissimilar) examples to emphasize and learn their fine-grained difference. It meanwhile imposes a bounding penalization on easy similar (dissimilar) examples to prevent the dominance of the easy examples in the optimization while preserving the high-level similarity (dissimilarity). This enables our model to well encode the key information carried by both easy and hard examples. Extensive empirical results on three widely-used image retrieval datasets show that (i) our method consistently and substantially outperforms state-of-the-art competing methods using hash codes of the same length and (ii) our method can use significantly (e.g., 50%-75%) shorter hash codes to perform substantially better than, or comparably well to, the competing methods.
Guansong Pang, Xiao Bai 0001, Chunhua Shen, Jun Zhou 0001, Edwin R. Hancock
ACM Multimedia3
2019 Privacy-preserved community discovery in online social networks
Xu Zheng 0001, Zhipeng Cai 0001, Guangchun Luo, Ling Tian, Xiao Bai 0001
Future Gener. Comput. Syst.5
2019 Deep depth-based representations of graphs through deep learning networks
Lu Bai 0001, Lixin Cui, Xiao Bai 0001, Edwin R. Hancock
Neurocomputing3
2019 Cross-modal hashing with semantic deep embedding
Xiao Bai 0001, Shuai Wang 0049, Jun Zhou 0001, Edwin R. Hancock
Neurocomputing2
2019 Multiscale Visual Attention Networks for Object Detection in VHR Remote Sensing Images
abstract
Object detection plays an active role in remote sensing applications. Recently, deep convolutional neural network models have been applied to automatically extract features, generate region proposals, and predict corresponding object class. However, these models face new challenges in VHR remote sensing images due to the orientation and scale variations and the cluttered background. In this letter, we propose an end-to-end multiscale visual attention networks (MS-VANs) method. We use skip-connected encoder-decoder model to extract multiscale features from a full-size image. For feature maps in each scale, we learn a visual attention network, which is followed by a classification branch and a regression branch, so as to highlight the features from object region and suppress the cluttered background. We train the MS-VANs model by a hybrid loss function which is a weighted sum of attention loss, classification loss, and regression loss. Experiments on a combined data set consisting of Dataset for Object Detection in Aerial Images and NWPU VHR-10 show that the proposed method outperforms several state-of-the-art approaches.
Chen Wang 0026, Xiao Bai 0001, Shuai Wang 0049, Jun Zhou 0001, Peng Ren 0001
IEEE Geosci. Remote. Sens. Lett.2
2019 Self-Supervised deep homography estimation with invertibility constraints
Chen Wang 0026, Xiang Wang 0014, Xiao Bai 0001, Yun Liu 0014, Jun Zhou 0001
Pattern Recognit. Lett.3
2019 Deep supervised hashing using symmetric relative entropy
Xueni Zhang, Lei Zhou 0008, Xiao Bai 0001, Xiushu Luan, Jie Luo 0004, Edwin R. Hancock
Pattern Recognit. Lett.3
2018 Discriminate Cross-modal Quantization for Efficient Retrieval
abstract
Efficient cross-modal retrieval involves searching similar items across different modalities, e.g., using an image(text) to search for texts(images). To speed up cross-modal retrieval, hashing-based methods threshold continuous embeddings into binary codes, inducing substantial loss of accuracy retrieval. To further improve retrieval performance, several quantization-based methods quantize embeddings into real-valued codewords to maximumlly preserve inter-modal and intra-modal similarity relation, while the discrimination between dissimilar data is ignored. To address these challenges, we propose, for the first time, a novel discriminate cross-modal quantization(DCMQ) which nonlinearly maps different modalities into a common space where ir-relevant data points are semantically separable: the points belonging to a class lie in a cluster that is not overlapped with other clusters corresponding to other classes. An effective optimization algorithm is developed for the proposed method to jointly learn the modality-specific mapping functions, the sharing codebooks, the unified binary codes and a linear classifier. Experimental comparison with state-of-the-art algorithms over three benchmark datasets demonstrates that DCMQ achieves significant improvement in search accuracy.
Shuai Wang 0049, Xiao Bai 0001
ICPR4
2018 Binary Coding by Matrix Classifier for Efficient Subspace Retrieval
abstract
Fast retrieval in large-scale database with high-dimensional subspaces is an important task in many applications, such as image retrieval, video retrieval and visual recognition. This can be facilitated by approximate nearest subspace (ANS) retrieval which requires effective subspace representation. Most of the existing methods for this problem represent subspace by point in the Euclidean space or the Grassmannian space before applying the approximate nearest neighbor (ANN) search. However, the efficiency of these methods can not be guaranteed because the subspace representation step can be very time consuming when coping with high dimensional data. Moreover, the transforming process for subspace to point will cause subspace structural information loss which influence the retrieval accuracy. In this paper, we present a new approach for hashing-based ANS retrieval. The proposed method learns the binary codes for given subspace set following a similarity preserving criterion. It simultaneously leverages the learned binary codes to train matrix classifiers as hash functions. This method can directly binarize a subspace without transforming it into a vector. Therefore, it can efficiently solve the large-scale and high-dimensional multimedia data retrieval problem. Experiments on face recognition and video retrieval show that our method outperforms several state-of-the-art methods in both efficiency and accuracy.
Lei Zhou 0008, Xiao Bai 0001, Xianglong Liu 0001, Jun Zhou 0001
ICMR2
2018 Adaptive hash retrieval with kernel based similarity
Xiao Bai 0001, Haichuan Yang, Lu Bai 0001, Jun Zhou 0001, Edwin R. Hancock
Pattern Recognit.1
2018 Material based salient object detection from hyperspectral images
Jie Liang 0003, Jun Zhou 0001, Xiao Bai 0001, Bin Wang 0041
Pattern Recognit.4
2017 Deep Residual Convolutional Neural Network for Hyperspectral Image Super-Resolution
Chen Wang 0026, Yun Liu 0014, Xiao Bai 0001, Wenzhong Tang, Jun Zhou 0001
ICIG (3)3
2017 Heterogeneous face recognition via grassmannian based nearest subspace search
abstract
Heterogeneous face recognition involves matching faces in different image modalities, such as near infrared images to visible images or sketch images to photos. This challenging task has attracted increasing attention in recent years. This paper presents, for the first time, a subspace based method to tackle the problem of face recognition between visible images (VIS) and near infrared (NIR) images. Subspace is used to extract essential attributes from VIS and NIR images. We adopt Grassmannian radial basis function (RBF) kernel to keep the relationship between subspaces, and use kernel canonical correlation analysis (KCCA) to handle correlation mapping between VIS and NIR domains. After mapping both VIS and NIR images to the common space, the heterogeneous face recognition problem can be easily completed by the nearest search. We evaluate the proposed method on the CASIA NIR-VIS 2.0 dataset. The experimental results demonstrate that our method is very effective for NIR-VIS face recognition.
Xiao Bai 0001, Jun Zhou 0001
ICIP3
2017 Non-local similarity based tensor decomposition for hyperspectral image denoising
abstract
Compared to traditional color or grayscale images, hyperspectral image (HSI) can help deliver more faithful representation of ground objects and enhance the performance of many computer vision tasks. However, an HSI is often corrupted by various noises, which has serious impact on the subsequent processing. Considering the non-local similarity across spatial domain and global similarity along spectral domain, a novel denoising method based on tensor decomposition is proposed in this paper. Firstly, 3D full band patches extracted from the HSI are grouped to form a 4th-order tensor by utilizing the non-local similarity in a proper window size. Then the task of hyperspectral image denoising is transformed into a high order tensor approximation problem, which can be efficiently solved by alternating optimization. An iterative denoising strategy is adopted for better effect in practice. Experimental results on simulated and real HSI data show that the proposed algorithm outperforms several state-of-the-art methods.
Xiao Bai 0001, Jun Zhou 0001
ICIP2
2017 Quantum kernels for unattributed graphs using discrete-time quantum walks
Lu Bai 0001, Luca Rossi 0004, Lixin Cui, Zhihong Zhang 0001, Peng Ren 0001, Xiao Bai 0001, Edwin R. Hancock
Pattern Recognit. Lett.6
2017 On the Sampling Strategy for Evaluation of Spectral-Spatial Methods in Hyperspectral Image Classification
abstract
Spectral-spatial processing has been increasingly explored in remote sensing hyperspectral image classification. While extensive studies have focused on developing methods to improve the classification accuracy, experimental setting and design for method evaluation have drawn little attention. In the scope of supervised classification, we find that traditional experimental designs for spectral processing are often improperly used in the spectral-spatial processing context, leading to unfair or biased performance evaluation. This is especially the case when training and testing samples are randomly drawn from the same image - a practice that has been commonly adopted in the experiments. Under such setting, the dependence caused by overlap between the training and testing samples may be artificially enhanced by some spatial information processing methods, such as spatial filtering and morphological operation. Such enhancement of dependence in return amplifies the classification accuracy, leading to an improper evaluation of spectral-spatial classification techniques. Therefore, the widely adopted pixel-based random sampling strategy is not always suitable to evaluate spectral-spatial classification algorithms, because it is difficult to determine whether the improvement of classification accuracy is caused by incorporating spatial information into classifier or by increasing the overlap between training and testing samples. To tackle this problem, we propose a novel controlled random sampling strategy for spectral-spatial methods. It can greatly reduce the overlap between training and testing samples and provides more objective and accurate evaluation.
Jie Liang 0003, Jun Zhou 0001, Yuntao Qian, Lian Wen, Xiao Bai 0001, Yongsheng Gao 0001
IEEE Trans. Geosci. Remote. Sens.5
2016 Bilinear Discriminant Analysis Hashing: A Supervised Hashing Approach for High-Dimensional Data
Yanzhen Liu, Xiao Bai 0001, Jun Zhou 0001
ACCV (5)2
2016 Shape classification with a vertex clustering graph kernel
abstract
Graph kernels are powerful tools for structural analysis in computer vision.Unfortunately, most existing stateof-the-art graph kernels ignore the locational or structural correspondence information between graphs, based on the visual background.This drawback influences the performance of existing kernels for computer vision based classification problems, e.g., classification of shapes, point clouds and digital images.The aim of this paper is to address the problem with existing kernels, by developing a novel vertex clustering graph kernel.We show that this kernel not only overcomes the shortcoming of ignoring correspondence information between isomorphic substructures that arises in most existing graph kernels, but also guarantees the transitivity between the correspondence information.Our kernel can easily outperform state-of-the-art graph kernels in terms of classification accuracy on standard shape based graph datasets.
Lu Bai 0001, Lixin Cui, Yue Wang 0014, Xin Jin 0007, Xiao Bai 0001, Edwin R. Hancock
ICPR5
2016 Discriminative weighted band selection via one-class SVM for hyperspectral imagery
abstract
In the task of hyperspectral image classification, band selection is often adopted to select a subset of informative bands to reduce the computation and storage cost. We propose a supervised band selection method which allows calculation of a discriminative weight for each band. Specifically, we consider discriminative bands as those that contribute more positive scores to a one-class classifier than those for other classes during the training stage. Based on this observation, we learn discriminative a band weight vector for each class, then bands with larger discriminative weights can be selected. Our method can be efficiently solved in one-class SVM framework. Experimental results demonstrate the effectiveness of our method.
Enlong Fan, Xiao Bai 0001, Jun Zhou 0001
IGARSS4
2016 Describing and learning of related parts based on latent structural model in big data
Xiao Bai 0001, Huigang Zhang, Jun Zhou 0001, Wenzhong Tang
Neurocomputing2
2016 Band Weighting via Maximizing Interclass Distance for Hyperspectral Image Classification
abstract
We present a novel band weighting strategy that exploits multiple binary support vector machines (SVMs) to maximize interclass spectral distances for multiclass hyperspectral remote image classification. Specifically, we commence by training binary SVMs based on the original training samples. We then balance the bands of training samples by maximizing the modified classification scores for SVMs. This balance scheme enlarges the distances between individual training samples and the SVM hyperplane. For each class, we reformulate the binary SVM objective function based on the balanced training samples, resulting in a weighting vector that associates a weight to each spectral band for the class. For a testing sample, we weight it and then classify it by using the binary SVM, both with respect to every individual class. The classification result is obtained from the classifier with the greatest score. Experiments on two benchmark data sets show the effectiveness of the proposed strategy.
Xiao Bai 0001, Peng Ren 0001, Lu Bai 0001, Wenzhong Tang, Jun Zhou 0001
IEEE Geosci. Remote. Sens. Lett.2
2016 Discriminative sparse neighbor coding
Xiao Bai 0001, Peng Ren 0001, Lu Bai 0001, Jun Zhou 0001
Multim. Tools Appl.1
2016 Maximum margin hashing with supervised information
Haichuan Yang, Xiao Bai 0001, Yanzhen Liu, Lu Bai 0001, Jun Zhou 0001, Wenzhong Tang
Multim. Tools Appl.2
2016 Nonnegative-Matrix-Factorization-Based Hyperspectral Unmixing With Partially Known Endmembers
abstract
Hyperspectral unmixing is an important technique for estimating fractions of various materials from remote sensing imagery. Most unmixing methods make the assumption that no prior knowledge of endmembers is available before the estimation. This is, however, not true for some unmixing tasks for which part of the endmember signatures may be known in advance. In this paper, we address the hyperspectral unmixing problem with partially known endmembers. We extend nonnegative-matrix-factorization-based unmixing algorithms to incorporate prior information into their models. The proposed approach uses the spectral signature of known endmembers as a constraint, among others, in the unmixing model, and propagates the knowledge by an optimization process which minimizes the difference between the image data and the prior knowledge. Results on both synthetic and real data have validated the effectiveness of the proposed method and have shown that it has outperformed several state-of-the-art methods that use or do not use prior knowledge of endmembers.
Jun Zhou 0001, Yuntao Qian, Xiao Bai 0001, Yongsheng Gao 0001
IEEE Trans. Geosci. Remote. Sens.4
2015 Online sketching hashing
abstract
Recently, hashing based approximate nearest neighbor (ANN) search has attracted much attention. Extensive new algorithms have been developed and successfully applied to different applications. However, two critical problems are rarely mentioned. First, in real-world applications, the data often comes in a streaming fashion but most of existing hashing methods are batch based models. Second, when the dataset becomes huge, it is almost impossible to load all the data into memory to train hashing models. In this paper, we propose a novel approach to handle these two problems simultaneously based on the idea of data sketching. A sketch of one dataset preserves its major characters but with significantly smaller size. With a small size sketch, our method can learn hash functions in an online fashion, while needs rather low computational complexity and storage space. Extensive experiments on two large scale benchmarks and one synthetic dataset demonstrate the efficacy of the proposed method.
Cong Leng, Jiaxiang Wu 0001, Jian Cheng 0001, Xiao Bai 0001, Hanqing Lu
CVPR4
2015 Multilayer manifold and sparsity constrainted nonnegative matrix factorization for hyperspectral unmixing
abstract
Given a hyperspectral image, unmixing tries to estimate the spectral responses of the latent constituent materials and their corresponding fractions. Recently, Nonnegative Matrix Factorization (NMF) has been widely applied to solve the hyper-spectral unmixing problem because of its plausible physical interpretation. In this paper, we propose a novel method, Multilayer Manifold and Sparsity constrained Nonnegative Matrix Factorization (MMSNMF), for hyperspectral unmixing. In this approach, Multilayer NMF decomposes a hyperspectral image iteratively at several layers. In order to consider both the manifold structure of hyperspectral image and the sparsity of abundance matrix, we impose a graph regularization term and a sparsity regularization term on both the spectral signature matrix and the abundance matrix. Experimental results on both synthetic and real data validate the effectiveness of the proposed method in hyperspectral unmixing.
Zhenqiu Shu, Jun Zhou 0001, Xiao Bai 0001, Chunxia Zhao
ICIP4
2015 Band weighting and selection based on hyperplane margin maximization for hyperspectral image classification
abstract
Band selection is an effective solutions for dimensionality reduction in hyperspectral imagery. In this paper, a novel band weighting and selection method is proposed based on maximizing margin in support vector machine (SVM). The goal is to reduce high dimensionality if hyperspectral data while achieving accuracy classification performance. This method computes the weights of the samples to maximize the margin between the samples and the hyperplane in SVM. Bands are selected if they can enlarge the differences between classes and improve the classification performance. Experiments on two public benchmark hyperspectral datasets show the effectiveness of our method.
Xiao Bai 0001, Jun Zhou 0001
IGARSS2
2015 A Graph Kernel Based on the Jensen-Shannon Representation Alignment
Lu Bai 0001, Zhihong Zhang 0001, Chaoyan Wang, Xiao Bai 0001, Edwin R. Hancock
IJCAI4
2015 An incremental structured part model for object recognition
Xiao Bai 0001, Peng Ren 0001, Huigang Zhang, Jun Zhou 0001
Neurocomputing1
2015 Object Classification via Feature Fusion Based Marginalized Kernels
abstract
Various types of features can be extracted from very high resolution remote sensing images for object classification. It has been widely acknowledged that the classification performance can benefit from proper feature fusion. In this letter, we propose a softmax regression-based feature fusion method by learning distinct weights for different features. Our fusion method enables the estimation of object-to-class similarity measures and the conditional probabilities that each object belongs to different classes. Moreover, we introduce an approximate method for calculating the class-to-class similarities between different classes. Finally, the obtained fusion and similarity information are integrated into a marginalized kernel to build a support vector machine classifier. The advantages of our method are validated on QuickBird imagery.
Xiao Bai 0001, Chuntian Liu, Peng Ren 0001, Jun Zhou 0001, Huijie Zhao
IEEE Geosci. Remote. Sens. Lett.1
2014 Adaptive Object Retrieval with Kernel Reconstructive Hashing
abstract
Hashing is very useful for fast approximate similarity search on large database. In the unsupervised settings, most hashing methods aim at preserving the similarity defined by Euclidean distance. Hash codes generated by these approaches only keep their Hamming distance corresponding to the pairwise Euclidean distance, ignoring the local distribution of each data point. This objective does not hold for k-nearest neighbors search. In this paper, we firstly propose a new adaptive similarity measure which is consistent with k-NN search, and prove that it leads to a valid kernel. Then we propose a hashing scheme which uses binary codes to preserve the kernel function. Using low-rank approximation, our hashing framework is more effective than existing methods that preserve similarity over arbitrary kernel. The proposed kernel function, hashing framework, and their combination have demonstrated significant advantages compared with several state-of-the-art methods.
Haichuan Yang, Xiao Bai 0001, Jun Zhou 0001, Peng Ren 0001, Zhihong Zhang 0001, Jian Cheng 0001
CVPR2
2014 Semi-randomized hashing for large scale data retrieval
abstract
In information retrieval, efficient accomplishing the nearest neighbor search on large scale database is a great challenge. Hashing based indexing methods represent each data instance as a binary string to retrieve the approximate nearest neighbors. In this paper, we present a semi-randomized hashing approach to preserve the Euclidean distance by binary codes. Euclidean distance preserving is a classic research problem in hashing. Most hashing methods used purely randomized or optimized learning strategy to achieve this goal. Our method, on the other hand, combines both randomized and optimized strategies. It starts from generating multiple random vectors, and then approximates them by a single projection vector. In the quantization step, it uses the orthogonal transformation to minimize an upper bound of the deviation between real-valued vectors and binary codes. The proposed method overcomes the problem that randomized hash functions are isolated from the data distribution. What's more, our method supports an arbitrary number of hash functions, which is beneficial in building better hashing methods. The experiments show that our approach outperforms the alternative state-of-the-art methods for retrieval on the large scale dataset.
Haichuan Yang, Xiao Bai 0001, Jun Zhou 0001, Peng Ren 0001, Jian Cheng 0001, Lu Bai 0001
DSAA2
2014 Regularized Hierarchical Feature Learning with Non-negative Sparsity and Selectivity for Image Classification
abstract
Recently, many deep networks are proposed to learn hierarchical image representation to replace traditional hand-designed features. To enhance the ability of the generative model to tackle discriminative computer vision tasks (e.g. image classification), we propose a hierarchical deconvolutional network with two biologically inspired properties incorporated, i.e., non-negative sparsity and selectivity. First, we propose a single layer deconvolutional model with a raw image as input, attempting to decompose the input as a weighted sum of feature maps convolving with filters. Here, the filters are the model parameters common to all the inputs, while the feature maps and the summing weights are specific to the input. The non-negative sparsity is formulated as the /i-norm regularizer on the feature map, which is used to generate feature representations for image classification. And the selectivity is forced on the filters to make different filters active different inputs, through requiring the sparsity on the summing weights specifically. The two properties are summarized into an overall cost function, which can be solved with an alternatively iterative algorithm. Then, we build multiple layer deconvolutional network by stacking the single models, where the next-layer inputs are the results of a 3D max-pooling operation on the inferred feature maps of the front layer, and train the network in a greedy layer wise scheme. Finally, we explore the feature maps of each layer to generate the image representations and input them to a SVM classifier for the classification task. Experiments on two image benchmark datasets of Caltech-101 and Caltech-256 demonstrate the encouraging performance of our model compared with other deep feature learning models as well as some hand-designed features.
Bingyuan Liu, Jing Liu 0001, Xiao Bai 0001, Hanqing Lu
ICPR3
2014 Discriminative Context Models for Collective Activity Recognition
abstract
Context information has been widely studied for recognizing collective activities. Most existing works assume that all individuals in a single image share the same activity label. However, in many cases, multiple activities can be coexisted and serve as the context for each other in real-world scenarios. Based on this observation, we propose a novel approach to model both the intra-class and inter-class behavior interactions among persons in the scenario. By introducing the intra-class and inter-class context descriptors, we propose a unified discriminative model to jointly capture the individual appearance information and the context patterns around the focal person in a max-margin framework. Finally, a greedy forward search method is utilized to optimally label the activities in the testing scene. Experimental results demonstrate the superiority of our approach in activity recognition.
Chaoyang Zhao, Jinqiao Wang, Xiao Bai 0001, Qingshan Liu 0001, Hanqing Lu
ICPR4
2014 Learning Binary Codes with Bagging PCA
Cong Leng, Jian Cheng 0001, Xiao Bai 0001, Hanqing Lu
ECML/PKDD (2)4
2014 Optimized graph-based segmentation for ultrasound images
Qinghua Huang, Xiao Bai 0001, Yingguang Li, Xuelong Li 0001
Neurocomputing2
2014 VHR Object Detection Based on Structural Feature Extraction and Query Expansion
abstract
Object detection is an important task in very high-resolution remote sensing image analysis. Traditional detection approaches are often not sufficiently robust in dealing with the variations of targets and sometimes suffer from limited training samples. In this paper, we tackle these two problems by proposing a novel method for object detection based on structural feature description and query expansion. The feature description combines both local and global information of objects. After initial feature extraction from a query image and representative samples, these descriptors are updated through an augmentation process to better describe the object of interest. The object detection step is implemented using a ranking support vector machine (SVM), which converts the detection task to a ranking query task. The ranking SVM is first trained on a small subset of training data with samples automatically ranked based on similarities to the query image. Then, a novel query expansion method is introduced to update the initial object model by active learning with human inputs on ranking of image pairs. Once the query expansion process is completed, which is determined by measuring entropy changes, the model is then applied to the whole target data set in which objects in different classes shall be detected. We evaluate the proposed method on high-resolution satellite images and demonstrate its clear advantages over several other object detection methods.
Xiao Bai 0001, Huigang Zhang, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.1
2014 Data-Dependent Hashing Based on p-Stable Distribution
abstract
The p-stable distribution is traditionally used for data-independent hashing. In this paper, we describe how to perform data-dependent hashing based on p-stable distribution. We commence by formulating the Euclidean distance preserving property in terms of variance estimation. Based on this property, we develop a projection method, which maps the original data to arbitrary dimensional vectors. Each projection vector is a linear combination of multiple random vectors subject to p-stable distribution, in which the weights for the linear combination are learned based on the training data. An orthogonal matrix is then learned data-dependently for minimizing the thresholding error in quantization. Combining the projection method and orthogonal matrix, we develop an unsupervised hashing scheme, which preserves the Euclidean distance. Compared with data-independent hashing methods, our method takes the data distribution into consideration and gives more accurate hashing results with compact hash codes. Different from many data-dependent hashing methods, our method accommodates multiple hash tables and is not restricted by the number of hash functions. To extend our method to a supervised scenario, we incorporate a supervised label propagation scheme into the proposed projection method. This results in a supervised hashing scheme, which preserves semantic similarity of data. Experimental results show that our methods have outperformed several state-of-the-art hashing approaches in both effectiveness and efficiency.
Xiao Bai 0001, Haichuan Yang, Jun Zhou 0001, Peng Ren 0001, Jian Cheng 0001
IEEE Trans. Image Process.1
2013 A hypergraph based semi-supervised band selection method for hyperspectral image classification
abstract
Band selection is a fundamental problem in hyperspectral data processing. In this paper, we present a semi-supervised learning approach and a hypergraph model to select useful bands based on few labeled object information. The contributions of this paper are two-fold. Firstly, the hypergraph model captures multiple relationships between hyperspectral image samples. Secondly, the semi-supervised learning method not only utilizes unlabeled samples in the learning process to improve model performance, but also requires little labeled samples which can significantly reduce large amount of human labor and costs. The proposed approach is evaluated on AVIRIS and APHI datasets, which demonstrate its advantages over several other band selection methods.
Zhouxiao Guo, Xiao Bai 0001, Zhihong Zhang 0001, Jun Zhou 0001
ICIP2
2013 Salient object detection in hyperspectral imagery
abstract
Object detection in hyperspectral images is an important task for many applications. While most traditional methods are pixel-based, many recent efforts have been put on extracting spatial-spectral features. In this paper, we introduce Itti's visual saliency model into the spectral domain for object detection. This enables the extraction of salient spectral features, which is related to the material property and spatial layout of objects, in the scale space. To our knowledge, this is the first attempt to combine hyperspectral data with salient object detection. Three methods have been implemented and compared to show how color component in the traditional saliency model can be replaced by spectral information. We have performed experiments on selected images from three online hyperspectral datasets, and show the effectiveness of the proposed methods.
Jie Liang 0003, Jun Zhou 0001, Xiao Bai 0001, Yuntao Qian
ICIP3
2013 Label propagation hashing based on p-stable distribution and coordinate descent
abstract
Hashing is a useful tool for contents-based image retrieval on large scale database. This paper presents an unsupervised data-dependent hashing method which learns similarity preserving binary codes. It uses p-stable distribution and coordinate descent method to achieve a good approximate solution for an acknowledged objective of hashing. This method consists of two steps. Firstly, it uses p-stable distribution properties to generate an initial partial hashing solution. Next, coordinate descent method is used to extend this partial solution to be complete. Our approach combines the advantages of both data-independent and data-dependent methods, which makes full use of the training data, requires reduced training time, and is easy to implement. Experiments show that our method outperforms several other state-of-the-art methods.
Haichuan Yang, Xiao Bai 0001, Chuntian Liu, Jun Zhou 0001
ICIP2
2013 Semi-supervised hyperspectral band selection via sparse linear regression and hypergraph models
abstract
Band selection is an important step towards effective and efficient object classification in hyperspectral imagery. In this paper, we propose a semi-supervised learning method for band selection based on a sparse linear regression model. This model uses a least absolute shrinkage and selection operator to compute the regression coefficients from both labeled and unlabeled samples. These coefficients are then used to compute a contribution score for each band, which allows bands with high scores being selected for the testing step. During this process, unlabeled samples also contribute to the coefficients calculation. In order to propagate the labels to these samples, a hypergraph is first built to describe the relationship between labeled and unlabeled samples. This leads to an adjacency matrix whose entries are the sum of corresponding weights of hyperedges. Then matrix subspace learning method is used to estimate the labels of unlabeled samples. The proposed method is evaluated on the APHI dataset. Comparison with several baseline methods has shown the advantages of the proposed method on the pixel-level classification.
Zhouxiao Guo, Haichuan Yang, Xiao Bai 0001, Zhihong Zhang 0001, Jun Zhou 0001
IGARSS3
2013 Marginalized kernel-based feature fusion method for VHR object classification
abstract
Many image features can be extracted from very high resolution remote sensing images for object classification. Proper feature combination is a step towards better classification performance. In this paper, we propose a logistic regression-based feature fusion method which assigns different weights to different features. This method considers the probability that two images belongs to the same classes and the image-to-class similarity to define the similarity between two objects. This similarity is used as a marginalized kernel for the final classifier construction. Experiments on remote sensing images suggest that this approach is effective in various feature combination, and has outperformed the SVM baseline method.
Chuntian Liu, Xiao Bai 0001, Jun Zhou 0001
IGARSS3
2013 Hierarchical Remote Sensing Image Analysis via Graph Laplacian Energy
abstract
Segmentation and classification are important tasks in remote sensing image analysis. Recent research shows that images can be described in hierarchical structure or regions. Such hierarchies can produce the state-of-the-art segmentations and can be used in the classification. However, they often contain more levels and regions than required for an efficient image description, which may cause increased computational complexity. In this letter, we propose a new hierarchical segmentation method that applies graph Laplacian energy as a generic measure for segmentation. It reduces the redundancy in the hierarchy by an order of magnitude with little or no loss of performance. In the classification stage, we apply local self-similarity feature to capture the internal geometric layouts of regions in an image. By incorporating advantages from both semantic hierarchical segmentation and local geometric region description, we have achieved better performance than those from the methods being compared. In the experimental section, we validate the effectiveness of our method by showing results on QuickBird and GeoEye-1 image data sets.
Huigang Zhang, Xiao Bai 0001, Huaxin Zheng, Huijie Zhao, Jun Zhou 0001, Jian Cheng 0001, Hanqing Lu
IEEE Geosci. Remote. Sens. Lett.2
2013 Asymmetric propagation based batch mode active learning for image retrieval
Biao Niu, Jian Cheng 0001, Xiao Bai 0001, Hanqing Lu
Signal Process.3
2013 Object Detection Via Structural Feature Selection and Shape Model
abstract
In this paper, we propose an approach for object detection via structural feature selection and part-based shape model. It automatically learns a shape model from cluttered training images without need to explicitly use bounding boxes on objects. Our approach first builds a class-specific codebook of local contour features, and then generates structural feature descriptors by combining context shape information. These descriptors are robust to both within-class variations and scale changes. Through exploring pairwise image matching using fast earth mover's distance, feature weights can be iteratively updated. Those discriminative foreground features are assigned high weights and then selected to build a part-based shape model. Finally, object detection is performed by matching each testing image with this model. Experiments show that the proposed method is very effective. It has achieved comparable performance to the state-of-the-art shape-based detection methods, but requires much less training information.
Huigang Zhang, Xiao Bai 0001, Jun Zhou 0001, Jian Cheng 0001, Huijie Zhao
IEEE Trans. Image Process.2
2012 Object detection via foreground contour feature selection and part-based shape model
Huigang Zhang, Junxiu Wang, Xiao Bai 0001, Jun Zhou 0001, Jian Cheng 0001, Huijie Zhao
ICPR3
2012 Hypergraph Spectra for Semi-supervised Feature Selection
Zhihong Zhang 0001, Edwin R. Hancock, Xiao Bai 0001
ECML/PKDD (1)3
2012 Discriminative features for image classification and retrieval
Xiao Bai 0001
Pattern Recognit. Lett.2
2011 Discriminative Features for Image Classification and Retrieval
abstract
In this paper, we present a new method to improve the performance of current bag-of-word based image classification process. After feature extraction, we introduce a pair wise image matching scheme to select the discriminative features. Only the label information from the raining-sets is used to update the feature weights via an iterative matching processing. The selected features correspond to the foreground content of the images thus highlight the high level category knowledge of images. "Visual words" are constructed on these selected features. Our method can be used as a refinement step for current image classification and retrieval process. We prove the efficiency of our methods in three tasks: supervised image classification, semi-supervised image classification and image retrieval. In the experiment part, we use two canonical datasets Caltech 256 and MSRC-v2 to validate our method. The results show that the share of discriminative features increases significantly to 87% by applying our selection. In four group of contract tests, supervised classification accuracies are shown to improve up to 21% while semi supervised classification accuracies are 15%. Image retrieval precision is also remarkable enhanced by our method.
Xiao Bai 0001
ICIG2
2011 High-resolution satellite image classification and segmentation using Laplacian graph energy
abstract
Many segmentation algorithms describe images in terms of a hierarchy of regions. Although such hierarchies can produce state of the art segmentations and can be used in the classification, they often contain more data than is required for an efficient description which cause increased complexity and time cost. In this paper, we proposed a new hierarchical segmentation method which apply Laplacian graph energy as a generic measure to reduce the number of levels and regions in the hierarchy by an order of magnitude with little or no loss in performance. We apply our method in remote sensing image analysis.
Xiao Bai 0001
IGARSS2
2011 A novel approach for satellite image classification using local self-similarity
abstract
Extracting man-made objects in satellite images which are generated from the meter to sub-meter resolution plays an important role in remote satellite image analysis. However, spectral characteristics of urban land objects are so similar. So the classification accuracies are far from satisfactory by using only spectral information. As a result, researchers turn to incorporate geometrical information into satellite image classification. In this paper, we introduce a new local feature, namely local self-similarity(LSS) which captures internal geometric layouts of local self-similarities, into high spatial resolution images classification application. Our method captures self-similarity of color, edges, repetitive patterns and complex textures in a single unified way. With the help of Bag-of-Visual Words and SVMs, the proposed method performs well. Experimental results on Quickbird-image data set show that the proposed local self-similarity representation yields better classification performance than the low-level features, such as the spectral and texture features.
Huaxin Zheng, Xiao Bai 0001, Huijie Zhao
IGARSS2
2011 Query expansion for VHR image detection
abstract
In order to detect the objects of interest, many different approaches have been proposed. One kind of popular approaches are based on template matching, which use a template of the object class to match the image at different positions. The matching can be computed using similarity measures such as the correlation coefficient. These approaches, although easy and robust, has the limitation of not containing to much variability of the object class, especially for shape information. Despite the statistical variation in each kind of object, collecting enough training samples is another problem which is time consuming. Inspire by template matching and incremental learning, a new object-oriented object detection methodology for very high resolution remote sensing images is proposed in this paper. We obtain the first initial query results via the bag-of-visual-words method. Then we introduce two query expansion baseline expansion and PAS expansion to obtain a new incremental model for re-query. In the experiment part, we compare and evaluate the performance of our proposed methods.
Huaxin Zheng, Huigang Zhang, Xiao Bai 0001, Huijie Zhao
IGARSS3
2011 Learning invariant structure for object identification by using graph methods
Xiao Bai 0001, Yi-Zhe Song, Peter Hall 0001
Comput. Vis. Image Underst.1
2011 In Search of Perceptually Salient Groupings
abstract
Finding meaningful groupings of image primitives has been a long-standing problem in computer vision. This paper studies how salient groupings can be produced using established theories in the field of visual perception alone. The major contribution is a novel definition of the Gestalt principle of Prägnanz, based upon Koffka's definition that image descriptions should be both stable and simple. Our method is global in the sense that it operates over all primitives in an image at once. It works regardless of the type of image primitives and is generally independent of image properties such as intensity, color, and texture. A novel experiment is designed to quantitatively evaluate the groupings outputs by our method, which takes human disagreement into account and is generic to outputs of any grouper. We also demonstrate the value of our method in an image segmentation application and quantitatively show that segmentations deliver promising results when benchmarked using the Berkeley Segmentation Dataset (BSDS).
Yi-Zhe Song, Xiao Bai 0001, Peter Hall 0001, Liang Wang 0001
IEEE Trans. Image Process.2
2010 Manifold embedding for shape analysis
Xiao Bai 0001, Edwin R. Hancock, Hang Yu 0001
Neurocomputing1
2010 Geometric characterization and clustering of graphs using heat kernel embeddings
Xiao Bai 0001, Edwin R. Hancock, Richard C. Wilson 0001
Image Vis. Comput.1
2009 A generative model for graph matching and embedding
Xiao Bai 0001, Edwin R. Hancock, Richard C. Wilson 0001
Comput. Vis. Image Underst.1
2009 Graph characteristics from the heat kernel trace
Xiao Bai 0001, Edwin R. Hancock, Richard C. Wilson 0001
Pattern Recognit.1
2008 Object recognition using graph spectral invariants
abstract
Graph structures have been proved important in high level-vision since they can be used to represent structural and relational arrangements of objects in a scene. One of the problems that arises in the analysis of structural abstractions of object is graph clustering. In this paper, we explore how permutation invariants computed from the trace of the heat kernel can be used to characterize graphs for the purposes of measuring similarity and clustering. We explore three different approaches to characterize the heat kernel trace as a function of time. These are the heat kernel trace moments, heat content invariants and symmetric polynomials with Laplacian eigenvalues as inputs. Experiments on the COIL 100 and Caltech 256 databases reveal that the proposed invariants are effective and outperform the tradition methods.
Xiao Bai 0001, Richard C. Wilson 0001, Edwin R. Hancock
ICPR1
2008 Isotree: Tree clustering via metric embedding
Xiao Bai 0001, Andrea Torsello, Edwin R. Hancock
Neurocomputing1
2007 Learning Object Classes from Structure
abstract
The problem of identifying the class of an object from its visual appearance has received significant attention recently. Most of the work to date is premised on photometric measures, often building codebooks made from interest regions. All of it has been tested only on photographs, so far as we know. Our approach differs in two significant ways. First, we do not build a codebook of interest regions but instead make use of a hierarchical description of an image based on a watershed transform. Root nodes in the hierarchy are putative objects to be classified. Second, we classify these putative objects using a vector of fixed length that represents the structure of the hierarchy below the node. This allows us to classify not just photographs, but also paintings and drawings of visual objects. 1
Xiao Bai 0001, Yi-Zhe Song, Peter Hall 0001
BMVC1
2005 Characterising Graphs using the Heat Kernel
abstract
The heat-kernel of a graph is computed by exponentiating the Laplacian eigen-system with time. In this paper, we study the heat kernel mapping of the nodes of a graph into a vector-space. Specifically, we investigate whether the resulting point distribution can be used for the purposes of graphclustering. Our characterisation is based on the covariance matrix of the point distribution. We explore the relationship between the covariance matrix and the heat kernel, and demonstrate the eigenvalues of the covariance matrix are found be exponentiating the Laplacian eigenvalues with time. We apply the technique to images from the COIL database, and demonstrate that it leads to well defined graph clusters.
Xiao Bai 0001, Richard C. Wilson 0001, Edwin R. Hancock
BMVC1
2005 Clustering shapes using heat content invariants
abstract
In this paper, we investigate the use of invariants derived from the heat kernel as a means of clustering graphs. We turn to the heat-content, i.e. the sum of the elements of the heat kernel. The heat content can be expanded as a polynomial in time, and the coefficients of the polynomial are known to be permutation invariants. We demonstrate how the polynomial coefficients can be computed from the Laplacian eigen-system. Graph-clustering is performed by applying principal components analysis to vectors constructed from the polynomial coefficients. We experiment with the resulting algorithm on the COIL database, where it is demonstrated to outperform the use of Laplacian eigenvalues.
Xiao Bai 0001, Edwin R. Hancock
ICIP (1)1
2004 Graph Matching using Spectral Embedding and Semidefinite Programming
abstract
This paper describes how graph-spectral methods can be used to transform the node correspondence problem into one of point-set alignment. We commence by using the ISOMAP algorithm to embed the nodes of a graph in a low-dimensional Euclidean space. With the nodes in the graph transformed to points in a metric space, we can recast the problem of graph-matching into that of aligning the points. Here we use semidefinite programming to develop a variant of the Scott and Longuet-Higgins algorithm to find point correspondences. We experiment with the resulting algorithm on a number of real-world problems. 1
Xiao Bai 0001, Hang Yu 0001, Edwin R. Hancock
BMVC1
2003 Graph Clustering using Symmetric Polynomials and Local Linear Embedding
abstract
Although graph structures have proved useful in high level vision for object recognition and matching, they can prove computationally cumbersome because of the need to establish reliable correspondences between nodes. Hence, standard pattern recognition techniques can not be easily applied to graphs since feature vectors and not easily contructed. To overcome this problem, in this paper we turn to the spectral matrix. We show how the elements of this matrix can be used to construct symmetric polynomials that are permutation invariants. The co-efficients of these polynomials can be used as graph-features which can be encoded in a vectorial manner. We demonstrate that these vectors can be embedded in a low dimensional space using locally linear embedding, and that the embedding results in well defined graph clusters.
Edwin R. Hancock, Richard C. Wilson 0001, Xiao Bai 0001
BMVC3