VLDB 2026 Research / reviewers in the wild / expert
Xiaoqin Zhang 0002
dblp:z/XiaoqinZhang-2
· DBLP profile ↗
190ranked-venue papers
44as first author
110since 2021 · last 2026
0000-0003-0958-7285ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 115 · 24 first-author · 64 since 2021Graphics, computer vision, multimedia, augmented reality and games · 92 · 21 first-author · 49 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 18 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 first-authorComputer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MuSASplat: Efficient Sparse-View 3D Gaussian Splats via Lightweight Multi-Scale AdaptationabstractSparse-view 3D Gaussian splatting seeks to render high-quality novel views of 3D scenes from a limited set of input images. While recent pose-free feed-forward methods leveraging pre-trained 3D priors have achieved impressive results, most of them rely on full fine-tuning of large Vision Transformer (ViT) backbones and incur substantial GPU costs. In this work, we introduce MuSASplat, a novel framework that dramatically reduces the computational burden of training pose-free feed-forward 3D Gaussian splats models with little compromise of rendering quality. Central to our approach is a lightweight Multi-Scale Adapter that enables efficient fine-tuning of ViT-based architectures with only a small fraction of training parameters. This design avoids the prohibitive GPU overhead associated with previous full-model adaptation techniques while maintaining high fidelity in novel view synthesis, even with very sparse input views. In addition, we introduce a Feature Fusion Aggregator that integrates features across input views effectively and efficiently. Unlike widely adopted memory banks, the Feature Fusion Aggregator ensures consistent geometric integration across input views and meanwhile mitigates the memory usage, training complexity, and computational costs significantly. Extensive experiments across diverse datasets show that MuSASplat achieves state-of-the-art rendering quality but has significantly reduced parameters and training resource requirements as compared with existing methods. Muyu Xu, Fangneng Zhan, Xiaoqin Zhang 0002, Ling Shao 0001, Shijian Lu |
AAAI | 3 |
| 2026 | Shape-aware and feature fused power line detection network
Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, LinLin Shen, Jun Zhang 0011 |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Prototype-based latent space distance optimization on vehicle re-identification
Sixian Chan 0001, Jiaao Cui, Zheng Wang 0059, Xiaolong Zhou 0001, Xiaoqin Zhang 0002 |
Expert Syst. Appl. | 7 |
| 2026 | Multi-Modal Primitive Retrieval for Compositional Zero-Shot Learning
Chenchen Jing, Haozhe Zhang 0002, Junbo Lu, Yang Liu 0357, Hao Chen 0041, Xiaoqin Zhang 0002, Chunhua Shen |
Int. J. Comput. Vis. | 6 |
| 2026 | Scalable and Generalizable Correspondence Pruning via Geometry-Consistent Pre-TrainingabstractTwo-view correspondence pruning aims to identify reliable correspondences for camera pose estimation, serving as a fundamental step in many 3D vision tasks. Existing methods rely on geometric consistency to seek true correspondences (inliers) from numerous false correspondences (outliers). In this learning paradigm, outliers severely affect the representation learning of inliers, resulting in models that are neither robust nor generalizable. To address this issue, we propose a geometry-consistent pre-training paradigm that sculpts scalable and generalizable representations free from outlier interference. The paradigm features two appealing properties. 1) Implementation of geometry-consistent pre-training. We introduce masked inlier reconstruction as a pretext task and develop a simple yet effective pre-training framework based on a masked autoencoder. Specifically, due to the irregular and unordered nature of correspondences, which lack explicit positional information, we adopt a dual-branch structure that separately reconstructs the keypoints of two images. This enables indirect reconstruction of 4D correspondences, where keypoints from the paired image provide positional prompts. 2) Unified correspondence encoder. We propose a simple dual-stream encoder with built-in consensus interaction, providing a unified, extensible architecture that enhances representation learning. Extensive experiments demonstrate that our method, GeneralPruner, consistently outperforms state-of-the-art approaches in terms of robustness and generalization across various downstream tasks. Specifically, our method achieves 10.76%, 11.84%, and 8.65% performance gains in camera pose estimation, visual localization, and 3D registration, respectively. To the best of our knowledge, we are the first work to introduce a pre-training framework tailored for correspondence pruning, offering a more universal and scalable solution. Tangfei Liao, Xiaoqin Zhang 0002, Tao Wang 0052, Min Li 0052, Guobao Xiao, Mang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | S3Mamba: Region-aware spatio-semantic Mamba for remote sensing change detection
Yuan Wang 0032, Sixian Chan 0001, Guoyu Yang, Tianyang Dong, Xiaoqin Zhang 0002 |
Pattern Recognit. | 7 |
| 2026 | Motion2Motion: Learning human pose refining in videos without ground truth label
Zhenyu Wen, Zhen Hong, Haoran Duan 0001, Xiaoqin Zhang 0002 |
Pattern Recognit. | 7 |
| 2026 | Wavelet-based physically guided normalization network for real-time traffic dehazing
Shengdong Zhang, Xiaoqin Zhang 0002, LinLin Shen, Shaohua Wan 0001, Wenqi Ren |
Pattern Recognit. | 2 |
| 2026 | Hierarchical Multi-Modal Enhancement for Robust Transmission Line Detection
Shengdong Zhang, Xiaoqin Zhang 0002, Shaohua Wan 0001, Yujing M. Jiang, Wujie Zhou, LinLin Shen, Wenqi Ren |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | Corrections to "Exploring Fuzzy Priors From Multimapping GAN for Robust Image Dehazing"
Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, Li Zhao 0005, En Fan, Feng Huang 0007 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2026 | Enhancing the Interpretation of Skin Lesion Diagnosis: Concept Adaptive Fine-Tuning of Vision-Language ModelsabstractSignificant progress has been made in applying deep learning for the automatic diagnosis of skin lesions. However, most models remain unexplainable, which severely hinders their application in clinical settings. Concept-based ante-hoc interpretable models have the potential to clarify the decision-making process of diagnosis by learning high-level, human-understandable concepts, while they can only provide numerical values of conceptual contributions. Pre-trained Vision-Language Models (VLMs) can learn rich vision-language correlations from large-scale image-text pairs. Fine-tuning pre-trained VLMs for specific downstream tasks is an effective way to reduce data requirements. Nevertheless, when there is a substantial disparity between the pre-trained model and the target task, existing tuning methods frequently struggle to generalize, necessitating substantial training data to fully adapt VLMs to specialized medical tasks. In this work, we propose a concept adaptive fine-tuning (CptAFT) method based on the pre-trained VLM, BiomedCLIP, to develop a concept-based multi-modal interpretable skin lesion diagnosis model. By incorporating medical texts, such as reports and conceptual terms, our model can recognize fine-grained features and provide robust, natural language-driven interpretability. Moreover, our concept-adaptive method that reconstructs images using concept logits and imposes a consistency loss with the original image, enabling the VLM to quickly adapt to the task with a small amount of training data. Extensive experimental results demonstrate that our approach outperforms state-of-the-art closed box and interpretable models in both classification performance and medically relevant interpretability. In particular, after fine-tuning with a small amount of data, our model outperforms MONET, a model trained on the large Skin Disease Image-Report dataset, by 8.28% in concept recognition ability, demonstrating the interpretability of our model. Yating Zhu, Xiaoyan Wang 0007, Ming Xia 0005, Pan Mu, Haigen Hu, Xiaoqin Zhang 0002 |
IEEE J. Biomed. Health Informatics | 7 |
| 2026 | Patch Matter: Dual Modality Patch Contrastive for Non-Stationary Radio SignalsabstractThe emergence of abundant non-stationary radio signal (NSRS) data presents significant opportunities for applications in wireless communications, radar systems, remote sensing, and healthcare. While deep learning models have shown promise in capturing sequence dependencies, deriving generic and fine-grained representations of NSRS data remains challenging due to its complex, dynamic nature and the scarcity of labeled data. The NSRS data are often frequency-sensitive and exhibit minuscule inter-class distances, posing significant challenges for precise classification. To address these issues, we propose a novelDualModalityPatchContrastive (DMPC) framework. This framework leverages a stochastic patching paradigm for diverse local pattern extraction and a time-frequency cross-view optimization for frequency-sensitive feature mining. Furthermore, an Attentive Patch Aggregation (APA) mechanism enhances fine-grained inference under few-shot conditions through patch-level feature voting. Extensive experiments demonstrate the effectiveness of our approach in addressing the unique challenges of NSRS data. Jie Su 0001, Yuheng Ye, Zhenyu Wen, Taotao Li, Shibo He, Xiaoqin Zhang 0002, Rajiv Ranjan 0001 |
IEEE Trans. Mob. Comput. | 7 |
| 2026 | GVLTrack: Global Vision-Language Tracking with Multi-Stage Modal FusionabstractIn general, local Visual-Language (VL) trackers search targets around the previous bounding box by initial VL annotations. However, there is an inherent contradiction between the local searching perspective of the tracker and the orientation descriptions in language conducted under the global perspective. Furthermore, most methods only fuse modality information in a single stage, which tends to an insufficient relation modeling. To address these issues, we propose a Global Vision-Language Tracker (GVLTrack) with multi-stage modal fusion. First, it tracks the target in the entire image instead of local tracking based on previous results to resolve the above contradiction. Second, GVLTrack incorporates three modal interaction modules: Consistent Relationship Modeling (CRM), VL-Guided Query (VLQ) Initialization, and Recurrent Cross-Modal Decoder (RC-Decoder) to fuse vision-language modality and refine the bounding box progressively comprehensively. We conduct extensive experiments on several benchmarks and achieve competitive performance, demonstrating the effectiveness of our approach. The code will be made publicly available as soon as it is accepted. Sixian Chan 0001, Cong Bai, Xiaoqin Zhang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | SGFormer: Semantic-Geometry Fusion Transformer for Multi-modal 3D Panoptic SegmentationabstractModern methods for autonomous driving perception widely adopt multi-modal fusion to enhance 3D scene understanding. However, existing methods suffer from inferior semantic extraction in image encoders that treat all pixels equally, ignoring contextual differences. The generated multi-modal representations also typically lack comprehensive semantic and spatial geometry information, which is crucial for the 3D panoptic segmentation task. In this paper, we propose a novel Semantic-Geometry Fusion Transformer (SGFormer) that extracts adaptive semantic contexts, aggregates geometric information and captures the semantic-geometry fusion. First, in the Image Branch, we tailor semantic contexts for each pixel with context-guided attention and spatial context alignment to refine semantic details. Second, we transform image and voxel features into point-pixel geometry representations, simultaneously learning semantic category priors as embeddings to better represent scene geometry and semantics. Finally, to aggregate semantic information with related geometry, we design a semantic-geometry fusion that combines the transformer, effectively capturing semantic-geometry relationships into multi-modal panoptic representations. Notably, SGFormer achieves the state-of-the-art (SOTA) results on the nuScenes and SemanticPOSS, as well as yielding competitive performance on the SemanticKITTI. Moreover, SGFormer exhibits superior robustness compared to leading methods, marking an improvement of 2% to 10%. Hongqi Yu, Sixian Chan 0001, Xiaolong Zhou 0001, Xiaoqin Zhang 0002 |
AAAI | 4 |
| 2025 | SMStracker: Tri-Path Score Mask Sigma Fusion for Multi-Modal Tracking
Sixian Chan 0001, Zedong Li, Shijian Lu, Chunhua Shen, Xiaoqin Zhang 0002 |
ICCV | 6 |
| 2025 | Spatial Preference Rewarding for MLLMs Spatial Understanding
Han Qiu 0008, Peng Gao 0007, Lewei Lu, Xiaoqin Zhang 0002, Ling Shao 0001, Shijian Lu |
ICCV | 4 |
| 2025 | PacGDC: Label-Efficient Generalizable Depth Completion with Projection Ambiguity and ConsistencyabstractGeneralizable depth completion enables the acquisition of dense metric depth maps for unseen environments, offering robust perception capabilities for various downstream tasks. However, training such models typically requires large-scale datasets with metric depth labels, which are often labor-intensive to collect. This paper presents PacGDC, a label-efficient technique that enhances data diversity with minimal annotation effort for generalizable depth completion. PacGDC builds on novel insights into inherent ambiguities and consistencies in object shapes and positions during 2D-to-3D projection, allowing the synthesis of numerous pseudo geometries for the same visual scene. This process greatly broadens available geometries by manipulating scene scales of the corresponding depth maps. To leverage this property, we propose a new data synthesis pipeline that uses multiple depth foundation models as scale manipulators. These models robustly provide pseudo depth labels with varied scene scales, affecting both local objects and global layouts, while ensuring projection consistency that supports generalization. To further diversify geometries, we incorporate interpolation and relocation strategies, as well as unlabeled images, extending the data coverage beyond the individual use of foundation models. Extensive experiments show that PacGDC achieves remarkable generalizability across multiple benchmarks, excelling in diverse scene semantics/scales and depth sparsity/patterns under both zero-shot and few-shot settings. Code: https://github.com/Wang-xjtu/PacGDC. Aoran Xiao, Xiaoqin Zhang 0002, Shijian Lu |
ICCV | 3 |
| 2025 | PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-RegularizationsabstractCOLMAP-free 3D Gaussian Splatting (3D-GS) has recently attracted increasing attention due to its remarkable performance in reconstructing high-quality 3D scenes from unposed images or videos. However, it often struggles to handle scenes with complex camera trajectories as featured by drastic rotation and translation across adjacent camera views, leading to degraded estimation of camera poses and further local minima in joint optimization of camera poses and 3D-GS. We propose PCR-GS, an innovative COLMAP-free 3DGS technique that achieves superior 3D scene modeling and camera pose estimation via camera pose co-regularization. PCR-GS achieves regularization from two perspectives. The first is feature reprojection regularization which extracts view-robust DINO features from adjacent camera views and aligns their semantic information for camera pose regularization. The second is wavelet-based frequency regularization which exploits discrepancy in high-frequency details to further optimize the rotation matrix in camera poses. Extensive experiments over multiple real-world scenes show that the proposed PCR-GS achieves superior pose-free 3D-GS scene modeling under dramatic changes of camera trajectories. Xiaoqin Zhang 0002, Ling Shao 0001, Shijian Lu |
ICCV | 3 |
| 2025 | SAM-TTT: Segment Anything Model via Reverse Parameter Configuration and Test-Time Training for Camouflaged Object DetectionabstractThis paper introduces a new Segment Anything Model (SAM) that leverages reverse parameter configuration and test-time training to enhance its performance on Camouflaged Object Detection (COD), named SAM-TTT. While most existing SAM-based COD models primarily focus on enhancing SAM by extracting favorable features and amplifying its advantageous parameters, a crucial gap is identified: insufficient attention to adverse parameters that impair SAM's semantic understanding in downstream tasks. To tackle this issue, the Reverse SAM Parameter Configuration Module is proposed to effectively mitigate the influence of adverse parameters in a train-free manner by configuring SAM's parameters. Building on this foundation, the T-Visioner Module is unveiled to strengthen advantageous parameters by integrating Test-Time Training layers, originally developed for language tasks, into vision tasks. Test-Time Training layers represent a new class of sequence modeling layers characterized by linear complexity and an expressive hidden state. By integrating two modules, SAM-TTT simultaneously suppresses adverse parameters while reinforcing advantageous ones, significantly improving SAM's semantic understanding in COD task. Our experimental results on various COD benchmarks demonstrate that the proposed approach achieves state-of-the-art performance, setting a new benchmark in the field. The code will be available at https://github.com/guobaoxiao/SAM-TTT. Zhenni Yu, Li Zhao 0005, Guobao Xiao, Xiaoqin Zhang 0002 |
ACM Multimedia | 4 |
| 2025 | Physics-Guided Diffusion Model for Unpaired Real-World Dehazing
Hanqi Wang, Chenxiang Fan, Haigen Hu, Li Zhao 0005, Xiaoqin Zhang 0002 |
PRCV (9) | 5 |
| 2025 | COMPrompter: reconceptualized segment anything model with multiprompt network for camouflaged object detection
Xiaoqin Zhang 0002, Zhenni Yu, Li Zhao 0005, Deng-Ping Fan, Guobao Xiao |
Sci. China Inf. Sci. | 1 |
| 2025 | An Experimental Study on Exploring Strong Lightweight Vision Transformers via Masked Image Modeling Pre-training
Shubo Lin, Shaoru Wang, Yutong Kou, Congxuan Zhang, Xiaoqin Zhang 0002, Yizheng Wang, Weiming Hu 0004 |
Int. J. Comput. Vis. | 8 |
| 2025 | Visual Instruction Tuning towards General-Purpose Multimodal Large Language Model: A Survey
Jiaxing Huang 0001, Jingyi Zhang 0005, Kai Jiang 0001, Han Qiu 0008, Xiaoqin Zhang 0002, Ling Shao 0001, Shijian Lu, Dacheng Tao |
Int. J. Comput. Vis. | 5 |
| 2025 | Joint Global-Local Frames Modeling to enhance semantic alignment for zero-shot long video editing
Zewen Yu, Pengchong Qiao, Xiaoqin Zhang 0002 |
Neurocomputing | 4 |
| 2025 | DiFusionSeg: Diffusion-driven semantic segmentation with multi-modal image fusion for enhanced perception
Defeng He, Li Zhao 0005, Yayu Zheng, Xiaoqin Zhang 0002 |
Knowl. Based Syst. | 6 |
| 2025 | Global-local feature-mixed network with template update for visual tracking
Li Zhao 0005, Chenxiang Fan, Min Li 0052, Zhonglong Zheng, Xiaoqin Zhang 0002 |
Pattern Recognit. Lett. | 5 |
| 2025 | Exploring Fuzzy Priors From Multimapping GAN for Robust Image DehazingabstractSingle image dehazing has been extensively studied. While convolutional neural networks (CNNs) have driven notable progress in single image dehazing, their performance remains fundamentally constrained by the limited local receptive fields of convolutional operations, which impede the capture of global structural dependencies. In contrast, generative adversarial networks (GANs) have demonstrated exceptional capabilities in image synthesis, offering global insights into structure, texture, and color. The fuzzy prior, a probabilistic knowledge acquired through adversarial training in GANs, plays a pivotal role in robust dehazing. Motivated by this, we propose the fuzzy prior guided dehazing network (FPGDN). Our framework begins with a novel module that distills the fuzzy prior by translating an edge map into a color image, simultaneously capturing global structural, local textural, and color information. Subsequently, a dehazing network is constructed, leveraging this fuzzy prior. While the fuzzy prior captures rich color and texture features, the generated images may exhibit color shifts relative to the original scene. To remedy this, a CNN network is employed to capture local nuances and refine the dehazing outcome. Extensive experiments substantiate that the proposed FPGDN achieves superior dehazing performance on a variety of real and synthetic hazy images. Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, Li Zhao 0005, En Fan, Feng Huang 0007 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2025 | PTH-Net: Dynamic Facial Expression Recognition Without Face Detection and AlignmentabstractPyramid Temporal Hierarchy Network (PTH-Net) is a new paradigm for dynamic facial expression recognition, applied directly to raw videos, without face detection and alignment. Unlike the traditional paradigm, which focus only on facial areas and often overlooks valuable information like body movements, PTH-Net preserves more critical information. It does this by distinguishing between backgrounds and human bodies at the feature level, offering greater flexibility as an end-to-end network. Specifically, PTH-Net utilizes a pre-trained backbone to extract multiple general features of video understanding at various temporal frequencies, forming a temporal feature pyramid. It then further expands this temporal hierarchy through differentiated parameter sharing and downsampling, ultimately refining emotional information under the supervision of expression temporal-frequency invariance. Additionally, PTH-Net features an efficient Scalable Semantic Distinction layer that enhances feature discrimination, helping to better identify target expressions versus non-target ones in the video. Finally, extensive experiments demonstrate that PTH-Net performs excellently in eight challenging benchmarks, with lower computational costs compared to previous methods. The source code is available at https://github.com/lm495455/PTH-Net. Min Li 0052, Xiaoqin Zhang 0002, Tangfei Liao, Guobao Xiao |
IEEE Trans. Image Process. | 2 |
| 2025 | Towards Gradient Equalization and Feature Diversification for Long-Tailed Multi-Label Image RecognitionabstractMulti-label image recognition with convolutional neural networks has achieved remarkable progress in the past few years. However, most existing multi-label image recognition methods suffer from the long-tailed data distribution problem,i.e., head categories occupy most training samples, while tailed classes have few samples. This work firstly studies the influence of long-tailed data distribution on existing multi-label image recognition methods. Based on this, two crucial issues of the existing methods are identified: 1) severe gradient imbalance between head and tailed categories, even though re-balancing strategies are adopted; 2) the lack of diversity of tail category training samples. To tackle the first issue, this paper proposes a group sampling strategy to create group-wise balanced data distribution. Meanwhile, a dynamic gradient balancing loss is proposed to equalize the gradient for all categories. To tackle the second issue, this paper proposes a diversity enhancement module to fuse the information across all categories, preventing the network from overfitting tail classes. Furthermore, it also balances the gradient, promoting the discriminability of learned classifiers. Our method significantly outperforms the baseline method and achieves competitive performance with state-of-the-art methods on VOC-LT and COCO-LT datasets. Extensive ablation studies are conducted to verify the effectiveness of the essential proposals. Quan Cui, Xiaoqin Zhang 0002, Ruoxi Deng, Chaoqun Xia, Shijian Lu |
IEEE Trans. Multim. | 3 |
| 2025 | PrimePSegter: Progressively Combined Diffusion for 3D Panoptic Segmentation With Multi-Modal BEV RefinementabstractEffective and robust 3D panoptic segmentation is crucial for scene perception in autonomous driving. Modern methods widely adopt multi-modal fusion based simple feature concatenation to enhance 3D scene understanding, resulting in generated multi-modal representations typically lack comprehensive semantic and geometry information. These methods focused on panoptic prediction in a single step also limit the capability to progressively refine panoptic predictions under varying noise levels, which is essential for enhancing model robustness. To address these limitations, we first utilize BEV space to unify semantic-geometry perceptual representation, allowing for a more effective integration of LiDAR and camera data. Then, we propose PrimePSegter, a progressively combined diffusion 3D panoptic segmentation model that is conditioned on BEV maps to iteratively refine predictions by denoising samples generated from Gaussian distribution. PrimePSegter adopts a conditional encoder-decoder architecture for fine-grained panoptic predictions. Specifically, a multi-modal conditional encoder is equipped with BEV fusion network to integrate semantic and geometric information from LiDAR and camera streams into unified BEV space. Additionally, a diffusion transformer decoder operates on multi-modal BEV features with varying noise levels to guide the training of diffusion model, refining the BEV panoptic representations enriched with semantics and geometry in a progressive way. PrimePSegter achieves state-of-the-art performance on the nuScenes and competitive results on the SemanticKITTI, respectively. Moreover, PrimePSegter demonstrates superior robustness towards various scenarios, outperforming leading methods. Hongqi Yu, Sixian Chan 0001, Xiaolong Zhou 0001, Xiaoqin Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2025 | SyNet: A Synergistic Network for 3D Object Detection Through Geometric-Semantic-Based Multi-Interaction FusionabstractDriven by rising demands in autonomous driving, robotics,etc., 3D object detection has recently achieved great advancement by fusing optical images and LiDAR point data. On the other hand, most existing optical-LiDAR fusion methods straightly overlay RGB images and point clouds without adequately exploiting the synergy between them, leading to suboptimal fusion and 3D detection performance. Additionally, they often suffer from limited localization accuracy without proper balancing of global and local object information. To address this issue, we design a synergistic network (SyNet) that fuses geometric information, semantic information, as well as global and local information of objects for robust and accurate 3D detection. The SyNet captures synergies between optical images and LiDAR point clouds from three perspectives. The first is geometric, which derives high-quality depth by projecting point clouds onto multi-view images, enriching optical RGB images with 3D spatial information for a more accurate interpretation of image semantics. The second is semantic, which voxelizes point clouds and establishes correspondences between the derived voxels and image pixels, enriching 3D point clouds with semantic information for more accurate 3D detection. The third is balancing local and global object information, which introduces deformable self-attention and cross-attention to process the two types of complementary information in parallel for more accurate object localization. Extensive experiments show that SyNet achieves 70.7% mAP and 73.5% NDS on the nuScenes test set, demonstrating its effectiveness and superiority as compared with the state-of-the-art. Xiaoqin Zhang 0002, Kenan Bi, Sixian Chan 0001, Shijian Lu, Xiaolong Zhou 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Semantic-Spatial Attention for Refined Object Placement in Text-to-Image SynthesisabstractSolely based on given prompts, text-guided diffusion models have enjoyed a unique capability in generating diverse and creative images. Nevertheless, the conveyance of image information through text presents a series of challenges, particularly in controlling the positioning of objects in synthesized images. Despite attempts of recent efforts in exploring alternative conditions, such as bounding box/mask-image pairs, the requirement of a substantial amount of paired data and time-consuming fine-tuning emerge as new issues. Given the observations that not only prompt-related cross-attention maps reveal the spatial arrangement and centroid positions of the objects, but also out-of-prompt markers enjoy rich semantic information, we thus engineer a weighted optimization loss. Specifically, three spatial sub-losses, namely inner box reinforcement loss, outer box attenuation loss, and centroid loss, are devised and seamlessly integrated into the sampling step of current vanilla diffusion models. Without any annotations of layout data required, the final approach runs in a training-free fashion. Extensive experiments with new performance scores demonstrate that our proposal not only successfully addresses the issue of object positioning but also boosts the capabilities of most current models, such as Stable Diffusion and GLIGEN, in high-quality synthesis and coverage of various concepts. Moreover, the proposed mechanism plays a plug-and-play role. Jianwei Zheng 0001, Ni Xu, Wei Li 0034, Jiawei Jiang 0002, Xiaoqin Zhang 0002 |
IEEE Trans. Multim. | 5 |
| 2024 | VSFormer: Visual-Spatial Fusion Transformer for Correspondence PruningabstractCorrespondence pruning aims to find correct matches (inliers) from an initial set of putative correspondences, which is a fundamental task for many applications. The process of finding is challenging, given the varying inlier ratios between scenes/image pairs due to significant visual differences. However, the performance of the existing methods is usually limited by the problem of lacking visual cues (e.g., texture, illumination, structure) of scenes. In this paper, we propose a Visual-Spatial Fusion Transformer (VSFormer) to identify inliers and recover camera poses accurately. Firstly, we obtain highly abstract visual cues of a scene with the cross attention between local features of two-view images. Then, we model these visual cues and correspondences by a joint visual-spatial fusion module, simultaneously embedding visual cues into correspondences for pruning. Additionally, to mine the consistency of correspondences, we also design a novel module that combines the KNN-based graph and the transformer, effectively capturing both local and global contexts. Extensive experiments have demonstrated that the proposed VSFormer outperforms state-of-the-art methods on outdoor and indoor benchmarks. Our code is provided at the following repository: https://github.com/sugar-fly/VSFormer. Tangfei Liao, Xiaoqin Zhang 0002, Li Zhao 0005, Tao Wang 0047, Guobao Xiao |
AAAI | 2 |
| 2024 | Masked AutoDecoder is Effective Multi-Task Vision GeneralistabstractInspired by the success of general-purpose models in NLP, recent studies attempt to unify different vision tasks in the same sequence format and employ autoregressive Transformers for sequence prediction. They apply uni-directional attention to capture sequential dependencies and generate task sequences recursively. However, such autoregressive Transformers may not fit vision tasks well, as vision task sequences usually lack the sequential dependencies typically observed in natural languages. In this work, we design Masked AutoDecoder (MAD), an effective multitask vision generalist. MAD consists of two core designs. First, we develop a parallel decoding framework that introduces bi-directional attention to capture contextual dependencies comprehensively and decode vision task sequences in parallel. Second, we design a masked sequence modeling approach that learns rich task contexts by masking and reconstructing task sequences. In this way, MAD handles all the tasks by a single network branch and a simple cross-entropy loss with minimal task-specific designs. Extensive experiments demonstrate the great potential of MAD as a new paradigm for unifying various vision tasks. MAD achieves superior performance and inference efficiency compared to autoregressive counterparts while obtaining competitive accuracy with task-specific models. Code will be released at https://github.com/hanqiu-hq/MAD. Han Qiu 0008, Jiaxing Huang 0001, Peng Gao 0007, Lewei Lu, Xiaoqin Zhang 0002, Shijian Lu |
CVPR | 5 |
| 2024 | Weakly Supervised Monocular 3D Detection with a Single-View ImageabstractMonocular 3D detection (M3D) aims for precise 3D object localization from a single-view image which usually involves labor-intensive annotation of 3D detection boxes. Weakly supervised M3D has recently been studied to obviate the 3D annotation process by leveraging many existing 2D annotations, but it often requires extra training data such as LiDAR point clouds or multi-view images which greatly degrades its applicability and usability in various applications. We propose SKD-WM3D, a weakly supervised monocular 3D detection framework that exploits depth information to achieve M3D with a single-view image exclusively without any 3D annotations or other training data. One key design in SKD-WM3D is a self-knowledge distillation framework, which transforms image features into 3D-like representations by fusing depth information and effectively mitigates the inherent depth ambiguity in monocular scenarios with little computational overhead in inference. In addition, we design an uncertainty-aware distillation loss and a gradient-targeted transfer modulation strategy which facilitate knowledge acquisition and knowledge transfer, respectively. Extensive experiments show that SKD-WM3D surpasses the state-of-the-art clearly and is even on par with many fully supervised methods. Xueying Jiang, Sheng Jin 0002, Lewei Lu, Xiaoqin Zhang 0002, Shijian Lu |
CVPR | 4 |
| 2024 | CAT-SAM: Conditional Tuning for Few-Shot Adaptation of Segment Anything Model
Aoran Xiao, Weihao Xuan, Heli Qi, Yun Xing 0001, Ruijie Ren, Xiaoqin Zhang 0002, Ling Shao 0001, Shijian Lu |
ECCV (40) | 6 |
| 2024 | Handling the Non-smooth Challenge in Tensor SVD: A Multi-objective Tensor Recovery Framework
Wanglong Lu, Wenzhe Wang, Yankai Cao, Xiaoqin Zhang 0002, Xianta Jiang |
ECCV (14) | 5 |
| 2024 | Label Decoupling and Reconstruction: A Two-Stage Training Framework for Long-tailed Multi-label Medical Image RecognitionabstractDeep learning has made significant advancements and breakthroughs in medical image recognition. However, the clinical reality is complex and multifaceted, with patients often suffering from multiple intertwined diseases, not all of which are equally common, leading to medical datasets that are frequently characterized by multi-labels and a long-tailed distribution. In this paper, we propose a method involving label decoupling and reconstruction (LDRNet) to address these two specific challenges. The label decoupling utilizes the fusion of semantic information from both categories and images to capture the class-aware features across different labels. This process not only integrates semantic information from labels and images to improve the model's ability to recognize diseases, but also captures comprehensive features across various labels to facilitate a deeper understanding of disease characteristics within the dataset. Following this, our label reconstruction method uses the class-aware features to reconstruct the label distribution. This step generates a diverse array of virtual features for tail categories, promoting unbiased learning for the classifier and significantly enhancing the model's generalization ability and robustness. Extensive experiments conducted on three multi-label long-tailed medical image datasets, including the Axial Spondyloarthritis Dataset, NIH Chest X-ray 14 Dataset, and ODIR-5K Dataset, have demonstrated that our approach achieves state-of-the-art performance, showcasing its effectiveness in handling the complexities associated with multi-label and long-tailed distributions in medical image recognition. Xiaoqin Zhang 0002, Yisu Ge, Lusi Ye, Guodao Zhang, Huiling Chen 0001 |
ACM Multimedia | 3 |
| 2024 | Exploring Deeper! Segment Anything Model with Depth Perception for Camouflaged Object DetectionabstractThis paper introduces a new Segment Anything Model with Depth Perception (DSAM) for Camouflaged Object Detection (COD). DSAM exploits the zero-shot capability of SAM to realize precise segmentation in the RGB-D domain. It consists of the Prompt-Deeper Module and the Finer Module. The Prompt-Deeper Module utilizes knowledge distillation and the Bias Correction Module to achieve the interaction between RGB features and depth features, especially using depth features to correct erroneous parts in RGB features. Then, the interacted features are combined with the box prompt in SAM to create a prompt with depth perception. The Finer Module explores the possibility of accurately segmenting highly camouflaged targets from a depth perspective. It uncovers depth cues in areas missed by SAM through mask reversion, self-filtering, and self-attention operations, compensating for its defects in the COD domain. DSAM represents the first step towards the SAM-based RGB-D COD model. It maximizes the utilization of depth features while synergizing with RGB features to achieve multimodal complementarity, thereby overcoming the segmentation limitations of SAM and improving its accuracy in COD. Experimental results on COD benchmarks demonstrate that DSAM achieves excellent segmentation performance and reaches the state-of-the-art (SOTA) on COD benchmarks with less consumption of training resources. The code will be available at https://github.com/guobaoxiao/DSAM. Zhenni Yu, Xiaoqin Zhang 0002, Li Zhao 0005, Yi Bin, Guobao Xiao |
ACM Multimedia | 2 |
| 2024 | Alias-Free Mamba Neural OperatorabstractBenefiting from the booming deep learning techniques, neural operators (NO) are considered as an ideal alternative to break the traditions of solving Partial Differential Equations (PDE) with expensive cost.
Yet with the remarkable progress, current solutions concern little on the holistic function features--both global and local information-- during the process of solving PDEs.
Besides, a meticulously designed kernel integration to meet desirable performance often suffers from a severe computational burden, such as GNO with $O(N(N-1))$, FNO with $O(NlogN)$, and Transformer-based NO with $O(N^2)$.
To counteract the dilemma, we propose a mamba neural operator with $O(N)$ computational complexity, namely MambaNO.
Functionally, MambaNO achieves a clever balance between global integration, facilitated by state space model of Mamba that scans the entire function, and local integration, engaged with an alias-free architecture. We prove a property of continuous-discrete equivalence to show the capability of
MambaNO in approximating operators arising from universal PDEs to desired accuracy. MambaNOs are evaluated on a diverse set of benchmarks with possibly multi-scale solutions and set new state-of-the-art scores, yet with fewer parameters and better efficiency. Jianwei Zheng 0001, Wei Li 0034, Ni Xu, Xiaoxu Lin, Xiaoqin Zhang 0002 |
NeurIPS | 6 |
| 2024 | MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked AutoencodersabstractMonocular 3D object detection aims for precise 3D localization and identification of objects from a single-view image. Despite its recent progress, it often struggles while handling pervasive object occlusions that tend to complicate and degrade the prediction of object dimensions, depths, and orientations. We design MonoMAE, a monocular 3D detector inspired by Masked Autoencoders that addresses the object occlusion issue by masking and reconstructing objects in the feature space. MonoMAE consists of two novel designs. The first is depth-aware masking that selectively masks certain parts of non-occluded object queries in the feature space for simulating occluded object queries for network training. It masks non-occluded object queries by balancing the masked and preserved query portions adaptively according to the depth information. The second is lightweight query completion that works with the depth-aware masking to learn to reconstruct and complete the masked object queries. With the proposed feature-space occlusion and completion, MonoMAE learns enriched 3D representations that achieve superior monocular 3D detection performance qualitatively and quantitatively for both occluded and non-occluded objects. Additionally, MonoMAE learns generalizable representations that can work well in new domains. Xueying Jiang, Sheng Jin 0002, Xiaoqin Zhang 0002, Ling Shao 0001, Shijian Lu |
NeurIPS | 3 |
| 2024 | Historical Test-time Prompt Tuning for Vision Foundation ModelsabstractTest-time prompt tuning, which learns prompts online with unlabelled test samples during the inference stage, has demonstrated great potential by learning effective prompts on-the-fly without requiring any task-specific annotations. However, its performance often degrades clearly along the tuning process when the prompts are continuously updated with the test data flow, and the degradation becomes more severe when the domain of test samples changes continuously. We propose HisTPT, a Historical Test-time Prompt Tuning technique that memorizes the useful knowledge of the learnt test samples and enables robust test-time prompt tuning with the memorized knowledge. HisTPT introduces three types of knowledge banks, namely, local knowledge bank, hard-sample knowledge bank, and global knowledge bank, each of which works with different mechanisms for effective knowledge memorization and test-time prompt optimization. In addition, HisTPT features an adaptive knowledge retrieval mechanism that regularizes the prediction of each test sample by adaptively retrieving the memorized knowledge. Extensive experiments show that HisTPT achieves superior prompt tuning performance consistently while handling different visual recognition tasks (e.g., image classification, semantic segmentation, and object detection) and test samples from continuously changing domains. Jingyi Zhang 0005, Jiaxing Huang 0001, Xiaoqin Zhang 0002, Ling Shao 0001, Shijian Lu |
NeurIPS | 3 |
| 2024 | DcnnGrasp: towards accurate grasp pattern recognition with adaptive regularizer learning
Xiaoqin Zhang 0002, Xianta Jiang |
Sci. China Inf. Sci. | 1 |
| 2024 | Photo realistic synthetic dataset and multi-scale attention dehazing network
Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, LinLin Shen, Li Zhao 0005, Jun Zhang 0011 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Dual-STI: Dual-path spatial-temporal interaction learning for dynamic facial expression recognition
Min Li 0052, Xiaoqin Zhang 0002, Chenxiang Fan, Tangfei Liao, Guobao Xiao |
Inf. Sci. | 2 |
| 2024 | GAN-based dehazing network with knowledge transferring
Shengdong Zhang, Xiaoqin Zhang 0002, LinLin Shen, En Fan |
Multim. Tools Appl. | 2 |
| 2024 | DECNet: Dense embedding contrast for unsupervised semantic segmentation
Xiaoqin Zhang 0002, Xiaolong Zhou 0001, Sixian Chan 0001 |
Neural Networks | 1 |
| 2024 | T-Net++: Effective Permutation-Equivariance Network for Two-View Correspondence PruningabstractWe propose a conceptually novel, flexible, and effective framework (named T-Net++) for the task of two-view correspondence pruning. T-Net++ comprises two unique structures: the "-'' structure and the "|'' structure. The "-'' structure utilizes an iterative learning strategy to process correspondences, while the "|'' structure integrates all feature information of the "-'' structure and produces inlier weights. Moreover, within the "|'' structure, we design a new Local-Global Attention Fusion module to fully exploit valuable information obtained from concatenating features through channel-wise and spatial-wise relationships. Furthermore, we develop a Channel-Spatial Squeeze-and-Excitation module, a modified network backbone that enhances the representation ability of important channels and correspondences through the squeeze-and-excitation operation. T-Net++ not only preserves the permutation-equivariance manner for correspondence pruning, but also gathers rich contextual information, thereby enhancing the effectiveness of the network. Experimental results demonstrate that T-Net++ outperforms other state-of-the-art correspondence pruning methods on various benchmarks and excels in two extended tasks. Guobao Xiao, Xin Liu 0091, Xiaoqin Zhang 0002, Jiayi Ma 0001, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | A Survey of Label-Efficient Deep Learning for 3D Point CloudsabstractIn the past decade, deep neural networks have achieved significant progress in point cloud learning. However, collecting large-scale precisely-annotated point clouds is extremely laborious and expensive, which hinders the scalability of existing point cloud datasets and poses a bottleneck for efficient exploration of point cloud data in various tasks and applications. Label-efficient learning offers a promising solution by enabling effective deep network training with much-reduced annotation efforts. This paper presents the first comprehensive survey of label-efficient learning of point clouds. We address three critical questions in this emerging research field: i) the importance and urgency of label-efficient learning in point cloud processing, ii) the subfields it encompasses, and iii) the progress achieved in this area. To this end, we propose a taxonomy that organizes label-efficient learning methods based on the data prerequisites provided by different types of labels. We categorize four typical label-efficient learning approaches that significantly reduce point cloud annotation efforts: data augmentation, domain transfer learning, weakly-supervised learning, and pretrained foundation models. For each approach, we outline the problem setup and provide an extensive literature review that showcases relevant progress and challenges. Finally, we share our views on the current research challenges and potential future directions. Aoran Xiao, Xiaoqin Zhang 0002, Ling Shao 0001, Shijian Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Semantics-Guided Contrastive Network for Zero-Shot Object DetectionabstractZero-shot object detection (ZSD), the task that extends conventional detection models to detecting objects from unseen categories, has emerged as a new challenge in computer vision. Most existing approaches tackle the ZSD task with a strict mapping-transfer strategy that may lead to suboptimal ZSD results: 1) the learning process of these models neglects the available semantic information on unseen classes, which can easily bias towards the seen categories; 2) the original visual feature space is not well-structured for the ZSD task due to the lack of discriminative information. To address these issues, we develop a novel Semantics-Guided Contrastive Network for ZSD, named ContrastZSD, a detection framework that first brings contrastive learning mechanism into the realm of zero-shot detection. Particularly, ContrastZSD incorporates two semantics-guided contrastive learning subnets that contrast between region-category and region-region pairs respectively. The pairwise contrastive tasks take advantage of supervision signals derived from both the ground truth label and class similarity information. By performing supervised contrastive learning over those explicit semantic supervision, the model can learn more knowledge about unseen categories to avoid the bias problem to seen concepts, while optimizing the visual data structure to be more discriminative for better visual-semantic alignment. Extensive experiments are conducted on two popular benchmarks for ZSD, i.e., PASCAL VOC and MS COCO. Results show that our method outperforms the previous state-of-the-art on both ZSD and generalized ZSD tasks. Caixia Yan, Xiaojun Chang, Minnan Luo, Huan Liu 0012, Xiaoqin Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Restoring vision in hazy weather with hierarchical contrastive learning
Tao Wang 0052, Guangpin Tao, Wanglong Lu, Kaihao Zhang, Wenhan Luo, Xiaoqin Zhang 0002, Tong Lu 0002 |
Pattern Recognit. | 6 |
| 2024 | A novel non-pretrained deep supervision network for polyp segmentation
Zhenni Yu, Li Zhao 0005, Tangfei Liao, Xiaoqin Zhang 0002, Geng Chen 0001, Guobao Xiao |
Pattern Recognit. | 4 |
| 2024 | Adversarial Attacks on Video Object Segmentation With Hard Region DiscoveryabstractVideo object segmentation has been applied to various computer vision tasks, such as video editing, autonomous driving, and human-robot interaction. However, the methods based on deep neural networks are vulnerable to adversarial examples, which are the inputs attacked by almost human-imperceptible perturbations, and the adversary (i.e., attacker) will fool the segmentation model to make incorrect pixel-level predictions. This will rise the security issues in highly-demanding tasks because small perturbations to the input video will result in potential attack risks. Though adversarial examples have been extensively used for classification, it is rarely studied in video object segmentation. Existing related methods in computer vision either require prior knowledge of categories or cannot be directly applied due to the special design for certain tasks, failing to consider the pixel-wise region attack. Hence, this work develops an object-agnostic adversary that has adversarial impacts on VOS by first-frame attacking via hard region discovery. Particularly, the gradients from the segmentation model are exploited to discover the easily confused region, in which it is difficult to identify the pixel-wise objects from the background in a frame. This provides a hardness map that helps to generate perturbations with a stronger adversarial power for attacking the first frame. Empirical studies on three benchmarks indicate that our attacker significantly degrades the performance of several state-of-the-art video object segmentation models. Ping Li 0006, Li Yuan 0007, Jian Zhao 0006, Xianghua Xu, Xiaoqin Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Transformer-Based Multimodal Emotional Perception for Dynamic Facial Expression Recognition in the WildabstractDynamic expression recognition in the wild is a challenging task due to various obstacles, including low light condition, non-positive face, and face occlusion. Purely vision-based approaches may not suffice to accurately capture the complexity of human emotions. To address this issue, we propose a Transformer-based Multimodal Emotional Perception (T-MEP) framework capable of effectively extracting multimodal information and achieving significant augmentation. Specifically, we design three transformer-based encoders to extract modality-specific features from audio, image, and text sequences, respectively. Each encoder is carefully designed to maximize its adaptation to the corresponding modality. In addition, we design a transformer-based multimodal information fusion module to model cross-modal representation among these modalities. The unique combination of self-attention and cross-attention in this module enhances the robustness of output-integrated features in encoding emotion. By mapping the information from audio and textual features to the latent space of visual features, this module aligns the semantics of the three modalities for cross-modal information augmentation. Finally, we evaluate our method on three popular datasets (MAFW, DFEW, and AFEW) through extensive experiments, which demonstrate its state-of-the-art performance. This research offers a promising direction for future studies to improve emotion recognition accuracy by exploiting the power of multimodal features. Xiaoqin Zhang 0002, Min Li 0052, Guobao Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Multi-Prior Driven Network for RGB-D Salient Object DetectionabstractMost existing RGB-D salient object detection (SOD) methods rely on high-quality depth images. However, their performance is limited when processing low-quality depth maps. This paper exploits more complementary image priors to guide the model to learn on variable depth maps, and a novel multi-prior driven network called MPDNet is proposed for RGB-D SOD. MPDNet utilizes four processing pipelines to process RGB images and other priors, which include an RGB image processing pipeline, a depth map processing pipeline, a fine-grained and gradient prior processing pipeline, and an edge learning pipeline. Specifically, fine-grained and gradient priors are input to the same processing pipeline. For the depth maps, fine-grained and gradient priors, a prior channel attention module utilizes the channel attention mechanism to filter noises and highlights the salient cues. The RGB image processing pipeline uses a multi-feature progressive enhancement module to fuse and enhance features from depth maps. And a multi-feature prediction decoder decodes initial salient masks. In the edge learning pipeline, edge prior serves as an edge label and is captured by an edge capture module. Finally, the clear salient masks are obtained by fusing the salient information from the four pipelines. The experimental results on six benchmarks indicate that the proposed method outperforms thirteen state-of-the-art methods in six evaluation metrics. Xiaoqin Zhang 0002, Yuewang Xu, Tao Wang 0052, Tangfei Liao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Video-Based Multi-Camera Vehicle Tracking via Appearance-Parsing Spatio-Temporal Trajectory Matching NetworkabstractMulti-camera vehicle tracking is a fundamental task for city traffic management to count traffic flow or monitor roads. This paper focuses on multi-camera tracking on the highway, which is more challenging compared with city streets in some problems such as fast-moving vehicles, tiny similar vehicles in appearance, longer tracking distance, and lighting intensity changes in the dark tunnels. In this paper, we propose a practical Appearance-Parsing Spatio-Temporal Trajectory Matching Network (ASTM-Net) based on the global appearance matching of local trajectory for addressing the cross-camera tracking tasks on the highway. Specifically, considering that the environmental disturbance and small vehicles have a similar appearance, we propose a multiple appearance-attribute parsing (MAP) module consisting of a Bi-propagation top-down (Bi-TD) block and appearance re-identification (ARe-ID) block to obtain salient global appearance-attribute features through given a video sequence. To address discrete tracking fragments caused by occlusion, we develop an appearance-joint-tracking (AJT) mechanism to merge the isolated tracklets with target interaction and occlusion handling. We then exploit an appearance-informed spatio-temporal matching (ASTM) module to achieve multi-camera tracklet-totarget assignment, which employs spatio-temporal consistency relation for intra-camera trajectory correction and coarse intercamera tracklet correlation and aggregate appearance matrix of local trajectories for assigning global trajectory ID. Finally, in order to evaluate our proposed ASTM-Net, a new dataset, named HST, collected on the highway is established.We verify the ASTM-Net on the HST and the other three public datasets,i.e., CityFlow, UA-DETRAC, and Synthehicle, whose experimental results demonstrate the effectiveness and robustness of the proposed method. Xiaoqin Zhang 0002, Hongqi Yu, Xiaolong Zhou 0001, Sixian Chan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Generative Adversarial and Self-Supervised Dehazing NetworkabstractOwing to the fast developments of economics, a lot of devices and objects have been connected and have formed the Internet of Things (IoT). Visual sensors have been applied in vehicle navigation, traffic situational awareness, and traffic safety management. However, the particles in the air degrade the imaging quality, which affects the performance of vehicle navigation, traffic situational awareness, and traffic safety management. Deep-learning-based dehazing methods were proposed to address this issue. However, these methods are trained with simulated hazy images and cannot generalize to natural haze images well. To address the domain shift problem, some methods resort to zero-shot learning or domain adaption to boost the generalization of the model on natural haze images. However, the relevance between dehazed results and clean images is ignored by zero-shot dehazing methods. Domain-adaption-based dehazing methods ignore the relationship between the dehazed results and the hazy images. To overcome these issues, a generative adversarial and self-supervised dehazing network is introduced to boost the dehazing performance on real haze images. First, generative adversarial is employed to construct the relevance between dehazed results and haze-free images, which can boost the natural appearance of dehazed results. Second, self-supervised learning is employed to construct the relevance between the dehazed results and hazy images, which can restrict the solution space of dehazing. To show the effectiveness of the proposed model, we conduct extensive experiments on real and simulated haze images. Compared with state-of-the-art methods, the proposed model achieves state-of-the-art dehazing performance. Shengdong Zhang, Xiaoqin Zhang 0002, Shaohua Wan 0001, Wenqi Ren, Liping Zhao 0005, LinLin Shen |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Domain Adaptive LiDAR Point Cloud Segmentation via Density-Aware Self-TrainingabstractDomain adaptive LiDAR point cloud segmentation aims to learn a target segmentation model from labeled source point clouds and unlabelled target point clouds, which has recently attracted increasing attention due to various challenges in point cloud annotation. However, its performance is still very constrained as most existing studies did not well capture data-specific characteristics of LiDAR point clouds. Inspired by the observation that the domain discrepancy of LiDAR point clouds is highly correlated with point density, we design a density-aware self-training (DAST) technique that introduces point density into the self-training framework for domain adaptive point cloud segmentation. DAST consists of two novel and complementary designs. The first is density-aware pseudo labelling that introduces point density for accurate pseudo labelling of target data and effective self-supervised network retraining. The second is density-aware consistency regularization that encourages to learn density-invariant representations by enforcing target predictions to be consistent across points of different densities. Extensive experiments over multiple large-scale public datasets show that DAST achieves superior domain adaptation performance as compared with the state-of-the-art. Aoran Xiao, Jiaxing Huang 0001, Kangcheng Liu, Dayan Guan, Xiaoqin Zhang 0002, Shijian Lu |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Efficient Multiview Representation Learning With Correntropy and Anchor GraphabstractGraph-based multiview clustering methods have attracted much attention because of their ability to mine nonlinear structural information among instances. Although they perform well in many scenarios, they consume a lot of computational resources when dealing with large-scale multiview scenarios. To address this issue, we present a new insight into the anchor graph mechanism and propose a novel Nonnegative Anchor Graph Reconstruction (NAGR) model. NAGR introduces the sparse similarity graph into the symmetric matrix factorization and gets the nonnegative representation that retains the graph structural information. Thereafter, we develop a novel Efficient Multiview nonnegative Representation learning framework with Correntropy and Anchor graph (EMR-CA), which integrates multiview anchor graph reconstruction and consensus nonnegative representation learning into a unified framework. EMR-CA uses multiview anchor graph reconstruction to learn consensus nonnegative representation, where correntropy rather than F-norm is used as the approximation measurement criterion. Specifically, normalized anchor graphs of different views are decomposed into a consensus nonnegative representation and multiple view-specific representations, where the consensus representation retains the neighbor graph information between multiview instances and representative anchors on different views. Finally, the effectiveness of the proposed EMR-CA framework is verified by theoretical analysis and experimental results on large-scale realistic multiview scenarios. Nan Zhang 0014, Xiaoqin Zhang 0002, Shiliang Sun |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Domain Adaptive LiDAR Point Cloud Segmentation With 3D Spatial ConsistencyabstractDomain adaptive LiDAR point cloud segmentation aims to learn an effective target segmentation model from labelled source data and unlabelled target data, which has attracted increasing attention in recent years due to the difficulty in point-cloud annotation. It remains a very open research challenge as point clouds of different domains often have clear distribution discrepancies with variations in LiDAR sensor configurations, environmental conditions, occlusions, etc. We design a simple yet effective spatial consistency training framework that can learn superior domain-invariant feature representations from unlabelled target point clouds. The framework exploits three types of spatial consistency, namely, geometric-transform consistency, sparsity consistency, and mixing consistency which capture the semantic invariance of point clouds with respect to viewpoint changes, sparsity changes, and local context changes, respectively. With a concise mean teacher learning strategy, our experiments show that the proposed spatial consistency training outperforms the state-of-the-art significantly and consistently across multiple public benchmarks. Aoran Xiao, Dayan Guan, Xiaoqin Zhang 0002, Shijian Lu |
IEEE Trans. Multim. | 3 |
| 2024 | Tensor Recovery With Weighted Tensor Average RankabstractIn this article, a curious phenomenon in the tensor recovery algorithm is considered: can the same recovered results be obtained when the observation tensors in the algorithm are transposed in different ways? If not, it is reasonable to imagine that some information within the data will be lost for the case of observation tensors under certain transpose operators. To solve this problem, a new tensor rank called weighted tensor average rank (WTAR) is proposed to learn the relationship between different resulting tensors by performing a series of transpose operators on an observation tensor. WTAR is applied to three-order tensor robust principal component analysis (TRPCA) to investigate its effectiveness. Meanwhile, to balance the effectiveness and solvability of the resulting model, a generalized model that involves the convex surrogate and a series of nonconvex surrogates are studied, and the corresponding worst case error bounds of the recovered tensor is given. Besides, a generalized tensor singular value thresholding (GTSVT) method and a generalized optimization algorithm based on GTSVT are proposed to solve the generalized model effectively. The experimental results indicate that the proposed method is effective. Xiaoqin Zhang 0002, Li Zhao 0005, Zhengyuan Zhou, Zhouchen Lin |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | UniDAformer: Unified Domain Adaptive Panoptic Segmentation Transformer via Hierarchical Mask CalibrationabstractDomain adaptive panoptic segmentation aims to miti-gate data annotation challenge by leveraging off-the-shelf annotated data in one or multiple related source domains. However, existing studies employ two separate networks for instance segmentation and semantic segmentation which lead to excessive network parameters as well as complicated and computationally intensive training and inference processes. We design UniDAformer, a unified domain adaptive panoptic segmentation transformer that is simple but can achieve domain adaptive instance segmentation and semantic segmentation simultaneously within a single network. UniDAformer introduces Hierarchical Mask Calibration (HMC) that rectifies inaccurate predictions at the level of regions, superpixels and pixels via online self-training on the fly. It has three unique features: 1) it enables unified domain adaptive panoptic adaptation; 2) it mitigates false predictions and improves domain adaptive panoptic segmentation effectively; 3) it is end-to-end trainable with a much simpler training and inference pipeline. Exten-sive experiments over multiple public benchmarks show that UniDAformer achieves superior domain adaptive panoptic segmentation as compared with the state-of-the-art. Jingyi Zhang 0005, Jiaxing Huang 0001, Xiaoqin Zhang 0002, Shijian Lu |
CVPR | 3 |
| 2023 | FAC: 3D Representation Learning via Foreground Aware Feature ContrastabstractContrastive learning has recently demonstrated great potential for unsupervised pre-training in 3D scene understanding tasks. However, most existing work randomly selects point features as anchors while building contrast, leading to a clear bias toward background points that often dominate in 3D scenes. Also, object awareness and foreground-to-background discrimination are neglected, making contrastive learning less effective. To tackle these issues, we propose a general foreground-aware feature contrast (FAC) framework to learn more effective point cloud representations in pre-training. FAC consists of two novel contrast designs to construct more effective and informative contrast pairs. The first is building positive pairs within the same foreground segment where points tend to have the same semantics. The second is that we prevent over-discrimination between 3D segments/objects and encourage foreground-to-background distinctions at the segment level with adaptive feature learning in a Siamese correspondence network, which adaptively learns feature correlations within and across point cloud views effectively. Visualization with point activation maps shows that our contrast pairs capture clear correspondences among fore-ground regions during pre-training. Quantitative experiments also show that FAC achieves superior knowledge transfer and data efficiency in various downstream 3D semantic segmentation and object detection tasks. All codes, data, and models are available. Kangcheng Liu, Aoran Xiao, Xiaoqin Zhang 0002, Shijian Lu, Ling Shao 0001 |
CVPR | 3 |
| 2023 | DA-DETR: Domain Adaptive Detection Transformer with Information FusionabstractThe recent detection transformer (DETR) simplifies the object detection pipeline by removing hand-crafted designs and hyperparameters as employed in conventional two-stage object detectors. However, how to leverage the simple yet effective DETR architecture in domain adaptive object detection is largely neglected. Inspired by the unique DETR attention mechanisms, we design DA-DETR, a domain adaptive object detection transformer that introduces information fusion for effective transfer from a labeled source domain to an unlabeled target domain. DA-DETR introduces a novel CNN-Transformer Blender (CTBlender) that fuses the CNN features and Transformer features ingeniously for effective feature alignment and knowledge transfer across domains. Specifically, CTBlender employs the Transformer features to modulate the CNN features across multiple scales where the high-level semantic information and the low-level spatial information are fused for accurate object identification and localization. Extensive experiments show that DA-DETR achieves superior detection performance consistently across multiple widely adopted domain adaptation benchmarks. Jingyi Zhang 0005, Jiaxing Huang 0001, Gongjie Zhang, Xiaoqin Zhang 0002, Shijian Lu |
CVPR | 5 |
| 2023 | Towards Efficient Use of Multi-Scale Features in Transformer-Based Object DetectorsabstractMulti-scale features have been proven highly effective for object detection but often come with huge and even prohibitive extra computation costs, especially for the recent Transformer-based detectors. In this paper, we propose Iterative Multi-scale Feature Aggregation (IMFA) - a generic paradigm that enables efficient use of multi-scale features in Transformer-based object detectors. The core idea is to exploit sparse multi-scale features from just a few crucial locations, and it is achieved with two novel designs. First, IMFA rearranges the Transformer encoder-decoder pipeline so that the encoded features can be iteratively updated based on the detection predictions. Second, IMFA sparsely samples scale-adaptive features for refined detection from just a few keypoint locations under the guidance of prior detection predictions. As a result, the sampled multi-scale features are sparse yet still highly beneficial for object detection. Extensive experiments show that the proposed IMFA boosts the performance of multiple Transformer-based object detectors significantly yet with only slight computational overhead. Gongjie Zhang, Zichen Tian, Jingyi Zhang 0005, Xiaoqin Zhang 0002, Shijian Lu |
CVPR | 5 |
| 2023 | WaveNeRF: Wavelet-based Generalizable Neural Radiance FieldsabstractNeural Radiance Field (NeRF) has shown impressive performance in novel view synthesis via implicit scene representation. However, it usually suffers from poor scalability as requiring densely sampled images for each new scene. Several studies have attempted to mitigate this problem by integrating Multi-View Stereo (MVS) technique into NeRF while they still entail a cumbersome fine-tuning process for new scenes. Notably, the rendering quality will drop severely without this fine-tuning process and the errors mainly appear around the high-frequency features. In the light of this observation, we design WaveNeRF, which integrates wavelet frequency decomposition into MVS and NeRF to achieve generalizable yet high-quality synthesis without any per-scene optimization. To preserve high-frequency information when generating 3D feature volumes, WaveNeRF builds Multi-View Stereo in the Wavelet domain by integrating the discrete wavelet transform into the classical cascade MVS, which disentangles high-frequency information explicitly. With that, disentangled frequency features can be injected into classic NeRF via a novel hybrid neural renderer to yield faithful high-frequency details, and an intuitive frequency-guided sampling strategy can be designed to suppress artifacts around high-frequency regions. Extensive experiments over three widely studied benchmarks show that WaveNeRF achieves superior generalizable radiance field modeling when only given three images as input. Muyu Xu, Fangneng Zhan, Yingchen Yu, Xiaoqin Zhang 0002, Christian Theobalt, Ling Shao 0001, Shijian Lu |
ICCV | 5 |
| 2023 | Pose-Free Neural Radiance Fields via Implicit Pose RegularizationabstractPose-free neural radiance fields (NeRF) aim to train NeRF with unposed multi-view images and it has achieved very impressive success in recent years. Most existing works share the pipeline of training a coarse pose estimator with rendered images at first, followed by a joint optimization of estimated poses and neural radiance field. However, as the pose estimator is trained with only rendered images, the pose estimation is usually biased or inaccurate for real images due to the domain gap between real images and rendered images, leading to poor robustness for the pose estimation of real images and further local minima in joint optimization. We design IR-NeRF, an innovative pose-free NeRF that introduces implicit pose regularization to refine pose estimator with unposed real images and improve the robustness of the pose estimation for real images. With a collection of 2D images of a specific scene, IR-NeRF constructs a scene codebook that stores scene features and captures the scene-specific pose distribution implicitly as priors. Thus, the robustness of pose estimation can be promoted with the scene priors according to the rationale that a 2D real image can be well reconstructed from the scene codebook only when its estimated pose lies within the pose distribution. Extensive experiments show that IR-NeRF achieves superior novel view synthesis and outperforms the state-of-the-art consistently across multiple synthetic and real datasets. Fangneng Zhan, Yingchen Yu, Kunhao Liu, Rongliang Wu, Xiaoqin Zhang 0002, Ling Shao 0001, Shijian Lu |
ICCV | 6 |
| 2023 | A Closer Look at Self-Supervised Lightweight Vision TransformersabstractSelf-supervised learning on large-scale Vision Transformers (ViTs) as pre-training methods has achieved promising downstream performance. Yet, how much these pre-training paradigms promote lightweight ViTs' performance is considerably less studied. In this work, we develop and benchmark several self-supervised pre-training methods on image classification tasks and some downstream dense prediction tasks. We surprisingly find that if proper pre-training is adopted, even vanilla lightweight ViTs show comparable performance to previous SOTA networks with delicate architecture design. It breaks the recently popular conception that vanilla ViTs are not suitable for vision tasks in lightweight regimes. We also point out some defects of such pre-training, e.g., failing to benefit from large-scale pre-training data and showing inferior performance on data-insufficient downstream tasks. Furthermore, we analyze and clearly show the effect of such pre-training by analyzing the properties of the layer representation and attention maps for related models. Finally, based on the above analyses, a distillation strategy during pre-training is developed, which leads to further downstream performance improvement for MAE-based pre-training. Code is available at https://github.com/wangsr126/mae-lite. Shaoru Wang, Xiaoqin Zhang 0002, Weiming Hu 0004 |
ICML | 4 |
| 2023 | RIME: A physics-based optimization
Dong Zhao 0006, Ali Asghar Heidari, Lei Liu 0048, Xiaoqin Zhang 0002, Majdi M. Mafarja, Huiling Chen 0001 |
Neurocomputing | 5 |
| 2023 | Unsupervised Point Cloud Representation Learning With Deep Neural Networks: A SurveyabstractPoint cloud data have been widely explored due to its superior accuracy and robustness under various adverse situations. Meanwhile, deep neural networks (DNNs) have achieved very impressive success in various applications such as surveillance and autonomous driving. The convergence of point cloud and DNNs has led to many deep point cloud models, largely trained under the supervision of large-scale and densely-labelled point cloud data. Unsupervised point cloud representation learning, which aims to learn general and useful point cloud representations from unlabelled point cloud data, has recently attracted increasing attention due to the constraint in large-scale point cloud labelling. This paper provides a comprehensive review of unsupervised point cloud representation learning using DNNs. It first describes the motivation, general pipelines as well as terminologies of the recent studies. Relevant background including widely adopted point cloud datasets and DNN architectures is then briefly presented. This is followed by an extensive discussion of existing unsupervised point cloud representation learning methods according to their technical approaches. We also quantitatively benchmark and discuss the reviewed methods over multiple widely adopted point cloud datasets. Finally, we share our humble opinion about several challenges and problems that could be pursued in the future research in unsupervised point cloud representation learning. Aoran Xiao, Jiaxing Huang 0001, Dayan Guan, Xiaoqin Zhang 0002, Shijian Lu, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Structured Sparsity Optimization With Non-Convex Surrogates of $\ell _{2,0}$ℓ2,0-Norm: A Unified Algorithmic FrameworkabstractIn this paper, we present a general optimization framework that leverages structured sparsity to achieve superior recovery results. The traditional method for solving the structured sparse objectives based on$\ell _{2,0}$-norm is to use the$\ell _{2,1}$-norm as a convex surrogate. However, such an approximation often yields a large performance gap. To tackle this issue, we first provide a framework that allows for a wide range of surrogate functions (including non-convex surrogates), which exhibits better performance in harnessing structured sparsity. Moreover, we develop a fixed point algorithm that solves a key underlying non-convex structured sparse recovery optimization problem to global optimality with a guaranteed super-linear convergence rate. Building on this, we consider three specific applications, i.e., outlier pursuit, supervised feature selection, and structured dictionary learning, which can benefit from the proposed structured sparsity optimization framework. In each application, how the optimization problem can be formulated and thus be relaxed under a generic surrogate function is explained in detail. We conduct extensive experiments on both synthetic and real-world data and demonstrate the effectiveness and efficiency of the proposed framework. Xiaoqin Zhang 0002, Di Wang 0008, Guiying Tang, Zhengyuan Zhou, Zhouchen Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Audio-driven talking face generation with diverse yet realistic facial animationsabstractAudio-driven talking face generation, which aims to synthesize talking faces with realistic facial animations (including accurate lip movements, vivid facial expression details and natural head poses) corresponding to the audio, has achieved rapid progress in recent years. However, most existing work focuses on generating lip movements only without handling the closely correlated facial expressions, which degrades the realism of the generated faces greatly. This paper presents DIRFA, a novel method that can generate talking faces with diverse yet realistic facial animations from the same driving audio. To accommodate fair variation of plausible facial animations for the same audio, we design a transformer-based probabilistic mapping network that can model the variational facial animation distribution conditioned upon the input audio and autoregressively convert the audio signals into a facial animation sequence. In addition, we introduce a temporally-biased mask into the mapping network, which allows to model the temporal dependency of facial animations and produce temporally smooth facial animation sequence. With the generated facial animation sequence and a source image, photo-realistic talking faces can be synthesized with a generic generation network. Extensive experiments show that DIRFA can generate talking faces with realistic facial animations effectively. Rongliang Wu, Yingchen Yu, Fangneng Zhan, Xiaoqin Zhang 0002, Shijian Lu |
Pattern Recognit. | 5 |
| 2023 | SGA-Net: A Sparse Graph Attention Network for Two-View Correspondence LearningabstractEstablishing reliable correspondences between two images is a fundamental and important task in computer vision. This paper proposes a novel network called Sparse Graph Attention Network (SGA-Net), to capture rich contextual information of sparse graphs for feature matching task. Specifically, a graph attention block is proposed to enhance the representational ability of graph-structured features. The proposed block introduces a novel normalization technique for graph-structured features to embed global information into each edge feature, and it adopts the squeeze-and-excitation mechanism to capture graph-wise contextual information. Meanwhile, to further obtain interesting structural information of sparse graphs, a novel sparse graph transformer is developed based on multi-headed self-attention mechanism, while maintaining permutation-equivariance. Additionally, considering that the graph contexts in shallow layers are not fully exploited, a simple graph-context fusion block is introduced to adaptively capture topological information from different layers by implicitly modeling the interdependence between these graph contexts. The proposed SGA-Net can search dependable candidates among the putative correspondences and simultaneously estimate accurate camera poses for two-view geometry estimation. Extensive experiments on outlier removal and camera pose estimation tasks have demonstrated that the proposed SGA-Net outperforms state-of-the-art methods on both outdoor and indoor benchmarks (i.e., YFCC100M and SUN3D). SGA-Net achieves a mAP5° of 58.88% without RANSAC on the outdoor dataset, and it achieves a precision increase of 13.45% and 7.34% compared with the state-of-the-art result on outdoor and indoor datasets, respectively. Tangfei Liao, Xiaoqin Zhang 0002, Yuewang Xu, Ziwei Shi, Guobao Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Semantic-Aware Dehazing Network With Adaptive Feature FusionabstractDespite that convolutional neural networks (CNNs) have shown high-quality reconstruction for single image dehazing, recovering natural and realistic dehazed results remains a challenging problem due to semantic confusion in the hazy scene. In this article, we show that it is possible to recover textures faithfully by incorporating semantic prior into dehazing network since objects in haze-free images tend to show certain shapes, textures, and colors. We propose a semantic-aware dehazing network (SDNet) in which the semantic prior is taken as a color constraint for dehazing, benefiting the acquisition of a reasonable scene configuration. In addition, we design a densely connected block to capture global and local information for dehazing and semantic prior estimation. To eliminate the unnatural appearance of some objects, we propose to fuse the features from shallow and deep layers adaptively. Experimental results demonstrate that our proposed model performs favorably against the state-of-the-art single image dehazing approaches. Shengdong Zhang, Wenqi Ren, Xin Tan 0002, Zhi-Jie Wang 0009, Yong Liu 0018, Jingang Zhang, Xiaoqin Zhang 0002, Xiaochun Cao |
IEEE Trans. Cybern. | 7 |
| 2023 | Random Reconstructed Unpaired Image-to-Image TranslationabstractThe goal of unpaired image-to-image translation is to learn a mapping from a source domain to a target domain without using any labeled examples of paired images. This problem can be solved by learning the conditional distribution of source images in the target domain. A major limitation of existing unpaired image-to-image translation algorithms is that they generate untruthful images which are overcolored and lack details, while the translation of realistic images must be rich in details. To address this limitation, in this article, we propose a random reconstructed unpaired image-to-image translation (RRUIT) framework by generative adversarial network, which uses random reconstruction to preserve the high-level features in the source and adopts an adversarial strategy to learn the distribution in the target. We update the proposed objective function with two loss functions. The auxiliary loss guides the generator to create a coarse image, while the coarse-to-fine block next to the generator block produces an image that obeys the distribution of the target domain. The coarse-to-fine block contains two submodules based on the densely connected atrous spatial pyramid pooling, which enriches the details of generated images. We conduct extensive experiments on photorealistic stylization and artistic stylization. The experimental results confirm the superiority of the proposed RRUIT. Xiaoqin Zhang 0002, Chenxiang Fan, Zhiheng Xiao, Li Zhao 0005, Huiling Chen 0001, Xiaojun Chang |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Infrared and Visible Image Fusion via Interactive Compensatory Attention Adversarial LearningabstractThe existing generative adversarial fusion methods generally concatenate source images or deep features, and extract local features through convolutional operations without considering their global characteristics, which tends to produce a limited fusion performance. Toward this end, we propose a novel interactive compensatory attention fusion network, termed ICAFusion. In particular, in the generator, we construct a multi-level encoder-decoder network with a triple path, and design infrared and visible paths to provide additional intensity and gradient information for the concatenating path. Moreover, we develop the interactive and compensatory attention modules to communicate their pathwise information, and model their long-range dependencies through a cascading channel-spatial model. The generated attention maps can more focus on infrared target perception and visible detail characterization, and are used to reconstruct the fusion image. Therefore, the generator takes full advantage of local and global features to further increase the representation ability of feature extraction and feature reconstruction. Extensive experiments illustrate that our ICAFusion obtains superior fusion performance and better generalization ability, which precedes other advanced methods in the subjective visual description and objective metric evaluation. Our codes will be public athttps://github.com/Zhishe-Wang/ICAFusion. Zhishe Wang, Wenyu Shao, Jiawei Xu 0004, Xiaoqin Zhang 0002 |
IEEE Trans. Multim. | 5 |
| 2022 | Handling Slice Permutations Variability in Tensor RecoveryabstractThis work studies the influence of slice permutations on tensor recovery, which is derived from a reasonable assumption about algorithm, i.e. changing data order should not affect the effectiveness of the algorithm. However, as we will discussed in this paper, this assumption is not satisfied by tensor recovery under some cases. We call this interesting problem as Slice Permutations Variability (SPV) in tensor recovery. In this paper, we discuss SPV of several key tensor recovery problems theoretically and experimentally. The obtained results show that there is a huge gap between results by tensor recovery using tensor with different slices sequences. To overcome SPV in tensor recovery, we develop a novel tensor recovery algorithm by Minimum Hamiltonian Circle for SPV (TRSPV) which exploits a low dimensional subspace structures within data tensor more exactly. To the best of our knowledge, this is the first work to discuss and effectively solve the SPV problem in tensor recovery. The experimental results demonstrate the effectiveness of the proposed algorithm in eliminating SPV in tensor recovery. Xiaoqin Zhang 0002, Wenzhe Wang, Xianta Jiang |
AAAI | 2 |
| 2022 | Multi-scale Residual Interaction for RGB-D Salient Object Detection
Mingjun Hu, Xiaoqin Zhang 0002, Li Zhao 0005 |
ACCV (3) | 2 |
| 2022 | VMRF: View Matching Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) has demonstrated very impressive performance in novel view synthesis via implicitly modelling 3D representations from multi-view 2D images. However, most existing studies train NeRF models with either reasonable camera pose initialization or manually-crafted camera pose distributions which are often unavailable or hard to acquire in various real-world data. We design VMRF, an innovative view matching NeRF that enables effective NeRF training without requiring prior knowledge in camera poses or camera pose distributions. VMRF introduces a view matching scheme, which exploits unbalanced optimal transport to produce a feature transport plan for mapping a rendered image with randomly initialized camera pose to the corresponding real image. With the feature transport plan as the guidance, a novel pose calibration technique is designed which rectifies the initially randomized camera poses by predicting relative pose transformations between the pair of rendered and real images. Extensive experiments over a number of synthetic and real datasets show that the proposed VMRF outperforms the state-of-the-art qualitatively and quantitatively by large margins. Fangneng Zhan, Rongliang Wu, Yingchen Yu, Song Bai 0001, Xiaoqin Zhang 0002, Shijian Lu |
ACM Multimedia | 7 |
| 2022 | Referring Segmentation in Images and Videos With Cross-Modal Self-Attention NetworkabstractWe consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in the image or video. In this paper, we propose a cross-modal self-attention (CMSA) module to utilize fine details of individual words and the input image or video, which effectively captures the long-range dependencies between linguistic and visual features. Our model can adaptively focus on informative words in the referring expression and important regions in the visual input. We further propose a gated multi-level fusion (GMLF) module to selectively integrate self-attentive cross-modal features corresponding to different levels of visual features. This module controls the feature fusion of information flow of features at different levels with high-level and low-level semantic information related to different attentive words. Besides, we introduce cross-frame self-attention (CFSA) module to effectively integrate temporal information in consecutive frames which extends our method in the case of referring segmentation in videos. Experiments on benchmark datasets of four referring image datasets and two actor and action video segmentation datasets consistently demonstrate that our proposed approach outperforms existing state-of-the-art methods. Linwei Ye, Mrigank Rochan, Zhi Liu 0003, Xiaoqin Zhang 0002, Yang Wang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Online multiple object tracking using joint detection and embedding network
Sixian Chan 0001, Yangwei Jia, Xiaolong Zhou 0001, Cong Bai, Shengyong Chen, Xiaoqin Zhang 0002 |
Pattern Recognit. | 6 |
| 2022 | UNFusion: A Unified Multi-Scale Densely Connected Network for Infrared and Visible Image FusionabstractInfrared image retains typical thermal targets while visible image preserves rich texture details, image fusion aims to reconstruct a synthesized image containing prominent targets and abundant texture details. Most of deep learning-based methods mainly focus on convolution operation to extract the local features, but do not fully consider their multi-scale characteristics and global dependencies, which may cause loss of target regions and texture details in the fused image. Towards this goal, we present a unified multi-scale densely connected fusion network in this paper, named as UNFusion. We carefully design a multi-scale encoder-decoder architecture that can efficiently extract and reconstruct multi-scale deep features. Dense skip connections are employed in both encoder and decoder sub-networks to reuse all the intermediate features of different layers and scales for fusion tasks. In the fusion layer,$L_{p} $normalized attention models, which include three kinds of different norms, are proposed to highlight and combine these deep features from spatial and channel dimensions, and the combined spatial and channel attention maps are used to reconstruct a final fused image. We conduct extensive experiments on the public TNO and Roadscene datasets, and the results demonstrate that our UNFusion can simultaneously preserve high brightness of typical thermal targets and abundant texture details to obtain superior scene representation and better visual perception. Besides, our UNFusion achieves better fusion performance and transcends other state-of-the-art methods in terms of qualitative and quantitative comparisons. Our code is available athttps://github.com/Zhishe-Wang/UNFusion. Zhishe Wang, Junyao Wang 0002, Jiawei Xu 0004, Xiaoqin Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Multi-Attention Convolutional Neural Network for Video DeblurringabstractVideo deblurring, which aims at restoring the sharp video from blurry video, is drawing increasing attention in the field of computer vision. In this paper, a method called Multi-Attention Convolutional Neural Network (MACNN) consisting of the temporal-spatial attention module, the frame channel attention module, and the feature extraction-reconstruction module is proposed. First, we use the temporal-spatial attention module and the frame channel attention module to capture features with temporal and spatial information existing across neighboring frames. Then, these captured features are fused and reconstructed to restore the sharp frame. Last but not least, we train MACNN together with a content loss and a perceptual loss in an end-to-end manner to recover realistic video details. Both quantitative and qualitative evaluation results on standard benchmarks demonstrate the proposed MACNN is superior to the state-of-the-art methods in terms of accuracy, efficiency, and visual effect. Xiaoqin Zhang 0002, Tao Wang 0052, Runhua Jiang, Li Zhao 0005, Yuewang Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Single Image Haze Removal Based on a Simple Additive Model With Haze Smoothness PriorabstractSingle image haze removal, which is to recover the clear version of a hazy image, is a challenging task in computer vision. In this paper, an additive haze model is proposed to approximate the hazy image formation process. In contrast with the traditional optical model, it regards the haze as an additive layer to a clean image. The model thus avoids estimating the medium transmission rate and the global atmospherical light. In addition, based on a critical observation that haze changes gradually and smoothly across the image, a haze smoothness prior is proposed to constrain this model. This prior assumes that the haze layer is much smoother than the clear image. Benefiting from this prior, we can directly separate the clean image from a single hazy image. Experimental results and comparisons with synthetic images and real-world images demonstrate that the proposed method outperforms state-of-the-art single image haze removal algorithms. Xiaoqin Zhang 0002, Tao Wang 0052, Guiying Tang, Li Zhao 0005, Yuewang Xu, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Hierarchical Feature Fusion With Mixed Convolution Attention for Single Image DehazingabstractSingle image dehazing, which aims at restoring a haze-free image from its correspondingly unconstrained hazy scene, is a fundamental yet challenging task and has gained immense popularity recently. However, the images recovered by some existing haze-removal methods often contain haze, artifacts, and color distortions, which severely degrade the visual quality and have negative impacts on subsequent computer vision tasks. To this end, we propose a network combining multi-scale hierarchical feature fusion and mixed convolution attention to progressively and adaptively enhance the dehazing performance. The haze levels and image structure information are accurately estimated by fusing multi-scale hierarchical features, thus the model restores images with less remaining haze. The proposed mixed convolution attention mechanism is capable of reducing feature redundancy, learning compact and effective internal representations and highlighting task-relevant features, thus, it can further help the model estimate images with sharper textural details and more vivid colors. Furthermore, a deep semantic loss is also proposed to highlight essential semantic information in deep features. The experimental results show that the proposed method outperforms state-of-the-art haze removal algorithms. Xiaoqin Zhang 0002, Tao Wang 0052, Runhua Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Bioinspired Scene Classification by Deep Active Learning With Remote Sensing ApplicationsabstractAccurately classifying sceneries with different spatial configurations is an indispensable technique in computer vision and intelligent systems, for example, scene parsing, robot motion planning, and autonomous driving. Remarkable performance has been achieved by the deep recognition models in the past decade. As far as we know, however, these deep architectures are incapable of explicitly encoding the human visual perception, that is, the sequence of gaze movements and the subsequent cognitive processes. In this article, a biologically inspired deep model is proposed for scene classification, where the human gaze behaviors are robustly discovered and represented by a unified deep active learning (UDAL) framework. More specifically, to characterize objects' components with varied sizes, an objectness measure is employed to decompose each scenery into a set of semantically aware object patches. To represent each region at a low level, a local-global feature fusion scheme is developed which optimally integrates multimodal features by automatically calculating each feature's weight. To mimic the human visual perception of various sceneries, we develop the UDAL that hierarchically represents the human gaze behavior by recognizing semantically important regions within the scenery. Importantly, UDAL combines the semantically salient region detection and the deep gaze shifting path (GSP) representation learning into a principled framework, where only the partial semantic tags are required. Meanwhile, by incorporating the sparsity penalty, the contaminated/redundant low-level regional features can be intelligently avoided. Finally, the learned deep GSP features from the entire scene images are integrated to form an image kernel machine, which is subsequently fed into a kernel SVM to classify different sceneries. Experimental evaluations on six well-known scenery sets (including remote sensing images) have shown the competitiveness of our approach. Ge Su, Jianwei Yin, Ying Li 0001, Qiuru Lin, Xiaoqin Zhang 0002, Ling Shao 0001 |
IEEE Trans. Cybern. | 6 |
| 2022 | Infrared Small Target Detection via Dynamic Image Structure EvolutionabstractInfrared small target detection (IRSTD) is a challenging task, due to the scarce target feature, complex background interferences, and poor image quality. The existing studies handle the IRSTD problem by making specific distribution assumptions of target and background, which incurs two issues. First, the specific assumptions of background cannot always hold. Second, the distorted target distributions contaminated by heavy clutters and noise are difficult to model statically. This paper handles the problems from the discriminative and dynamic perspectives, and proposes an original mechanism named dynamic image structure evolution (DISE) and an DISE-derived single-frame IRSTD framework. First, we resolve IRSTD by a discriminative model without assuming specific background distributions, which is based on a mathematical definition of structure singularity modeling the ideal target structure. Second, to decouple the target distributions from interferences, DISE guides the distorted target to reveal potential structure singularity while suppress the interference signal through iterative procedures of structure collapse, intensity settlement, and collapse convergence. The three functional procedures of DISE perform their own duties regarding target enhancement and background suppression. Moreover, an innovative chain mechanism is introduced to propagate the structure field. Experiments on real data sets demonstrate the superiority of DISE against the state-of-the-art IRSTD methods. Chaoqun Xia, Shuhan Chen, Xiaoqin Zhang 0002, Zhiyong Pan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Siamese Implicit Region Proposal Network With Compound Attention for Visual TrackingabstractRecently, siamese-based trackers have achieved significant successes. However, those trackers are restricted by the difficulty of learning consistent feature representation with the object. To address the above challenge, this paper proposes a novel siamese implicit region proposal network with compound attention for visual tracking. First, an implicit region proposal (IRP) module is designed by combining a novel pixel-wise correlation method. This module can aggregate feature information of different regions that are similar to the pre-defined anchor boxes in Region Proposal Network. To this end, the adaptive feature receptive fields then can be obtained by linear fusion of features from different regions. Second, a compound attention module including a channel and non-local attention is raised to assist the IRP module to perform a better perception of the scale and shape of the object. The channel attention is applied for mining the discriminative information of the object to handle the background clutters of the template, while non-local attention is trained to aggregate the contextual information to learn the semantic range of the object. Finally, experimental results demonstrate that the proposed tracker achieves state-of-the-art performance on six challenging benchmark tests, including VOT-2018, VOT-2019, OTB-100, GOT-10k, LaSOT, and TrackingNet. Further, our obtained results demonstrate that the proposed approach can be run at an average speed of 72 FPS in real time. Sixian Chan 0001, Xiaolong Zhou 0001, Cong Bai, Xiaoqin Zhang 0002 |
IEEE Trans. Image Process. | 5 |
| 2022 | SST: Spatial and Semantic Transformers for Multi-Label Image RecognitionabstractMulti-label image recognition has attracted considerable research attention and achieved great success in recent years. Capturing label correlations is an effective manner to advance the performance of multi-label image recognition. Two types of label correlations were principally studied, i.e., the spatial and semantic correlations. However, in the literature, previous methods considered only either of them. In this work, inspired by the great success of Transformer, we propose a plug-and-play module, named the Spatial and Semantic Transformers (SST), to simultaneously capture spatial and semantic correlations in multi-label images. Our proposal is mainly comprised of two independent transformers, aiming to capture the spatial and semantic correlations respectively. Specifically, our Spatial Transformer is designed to model the correlations between features from different spatial positions, while the Semantic Transformer is leveraged to capture the co-existence of labels without manually defined rules. Other than methodological contributions, we also prove that spatial and semantic correlations complement each other and deserve to be leveraged simultaneously in multi-label image recognition. Benefitting from the Transformer's ability to capture long-range correlations, our method remarkably outperforms state-of-the-art methods on four popular multi-label benchmark datasets. In addition, extensive ablation studies and visualizations are provided to validate the essential components of our method. Quan Cui, Borui Zhao, Renjie Song, Xiaoqin Zhang 0002, Osamu Yoshie |
IEEE Trans. Image Process. | 5 |
| 2022 | The Improvement of Road Driving Safety Guided by Visual Inattentional BlindnessabstractThe computational modeling of human visual attention has received much attention in recent decades. In advanced industrial applications, it has been demonstrated that computational visual attention models (CVAMs) can predict visual attention very similarly to human visual attention. However, it is controversial whether the driver’s eye fixation location (EFL) or the predicted eye fixation location of computational visual attention models is more reliable and helpful for actual driving. To address this issue, an open database of videos taken under the most common 18 driving conditions in everyday driving has been established. In experiments using this database, expert drivers found that it was not sufficient for drivers to rely on only one of the two EFLs. Based on this finding, a hybrid EFL recommendation strategy is proposed for improving driving safety. By extracting visual characteristics from human dynamic vision, the performance of the proposed recommendation method demonstrates its potential value in these collected driving tasks. In addition, the visual comfort of driving is further addressed to enhance the safety of driving. From the results of experiments on 108 driving video clips taken of the most common 18 real driving conditions, it is confirmed that the proposed EFL recommendation achieves an experience rating of driving comfort between 88.1 and 92.7 out of 100. Jiawei Xu 0004, Seop Hyeong Park, Xiaoqin Zhang 0002, Jie Hu 0041 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | The Alleviation of Perceptual Blindness During Driving in Urban Areas Guided by Saccades RecommendationabstractIn advanced industrial applications, computational visual attention models (CVAMs) could predict visual attention very similarly to actual human attention allocation. This has been used as a very important component of technology in advanced driver assistance systems (ADAS). Given that the biological inspiration of the driving-related CVAMs could be obtained from skilled drivers in complex driving conditions, in which the driver’s attention is constantly directed at various salient and informative visual stimuli by alternating the eye fixations via saccades to drive safely, this paper proposes a saccade recommendation strategy to enhance the driving safety under urban road environment, particularly when the driver’s vision is often impaired by the visual crowding. The altered and directed saccades are collected and optimized by extracting four innate features from human dynamic vision. A neural network isdesigned to classify preferable saccades to reduce perceptual blindness due to visual crowding under urban scenes. A state-of-the-art CVAM is firstly adopted to localize the predicted eye fixation locations (EFLs) in driving video clips. Besides, human subjects’ gaze at the recommended EFLs is measured via an eye-tracker. The time delays between the predicted EFLs and drivers’ EFLs are analyzed under different driving conditions, followed by the time delays between the predicted EFLs and the driver’s hand control. The visually safe margin is then measured by mediating the driving speed and the total delay. Experimental results demonstrate that the recommended saccades can effectively reduce the amount of perceptual blindness, which is known to be of help to further improve road driving safety. Jiawei Xu 0004, Xiaoqin Zhang 0002, Seop Hyeong Park, Kun Guo 0004 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Adaptive Data Structure Regularized Multiclass Discriminative Feature SelectionabstractFeature selection (FS), which aims to identify the most informative subset of input features, is an important approach to dimensionality reduction. In this article, a novel FS framework is proposed for both unsupervised and semisupervised scenarios. To make efficient use of data distribution to evaluate features, the framework combines data structure learning (as referred to as data distribution modeling) and FS in a unified formulation such that the data structure learning improves the results of FS and vice versa. Moreover, two types of data structures, namely the soft and hard data structures, are learned and used in the proposed FS framework. The soft data structure refers to the pairwise weights among data samples, and the hard data structure refers to the estimated labels obtained from clustering or semisupervised classification. Both of these data structures are naturally formulated as regularization terms in the proposed framework. In the optimization process, the soft and hard data structures are learned from data represented by the selected features, and then, the most informative features are reselected by referring to the data structures. In this way, the framework uses the interactions between data structure learning and FS to select the most discriminative and informative features. Following the proposed framework, a new semisupervised FS (SSFS) method is derived and studied in depth. Experiments on real-world data sets demonstrate the effectiveness of the proposed method. Mingyu Fan, Xiaoqin Zhang 0002, Jie Hu 0041, Nannan Gu, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Massive-Scale Aerial Photo Categorization by Cross-Resolution Visual Perception EnhancementabstractCategorizing aerial photographs with varied weather/lighting conditions and sophisticated geomorphic factors is a key module in autonomous navigation, environmental evaluation, and so on. Previous image recognizers cannot fulfill this task due to three challenges: 1) localizing visually/semantically salient regions within each aerial photograph in a weakly annotated context due to the unaffordable human resources required for pixel-level annotation; 2) aerial photographs are generally with multiple informative attributes (e.g., clarity and reflectivity), and we have to encode them for better aerial photograph modeling; and 3) designing a cross-domain knowledge transferal module to enhance aerial photograph perception since multiresolution aerial photographs are taken asynchronistically and are mutually complementary. To handle the above problems, we propose to optimize aerial photograph's feature learning by leveraging the low-resolution spatial composition to enhance the deep learning of perceptual features with a high resolution. More specifically, we first extract many BING-based object patches (Cheng et al., 2014) from each aerial photograph. A weakly supervised ranking algorithm selects a few semantically salient ones by seamlessly incorporating multiple aerial photograph attributes. Toward an interpretable aerial photograph recognizer indicative to human visual perception, we construct a gaze shifting path (GSP) by linking the top-ranking object patches and, subsequently, derive the deep GSP feature. Finally, a cross-domain multilabel SVM is formulated to categorize each aerial photograph. It leverages the global feature from low-resolution counterparts to optimize the deep GSP feature from a high-resolution aerial photograph. Comparative results on our compiled million-scale aerial photograph set have demonstrated the competitiveness of our approach. Besides, the eye-tracking experiment has shown that our ranking-based GSPs are over 92% consistent with the real human gaze shifting sequences. Xiaoqin Zhang 0002, Mingliang Xu 0001, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Learning to Decode Contextual Information for Efficient Contour DetectionabstractContour detection plays an important role in both academic research and real-world applications. As the basic building block of many applications, its accuracy and efficiency highly influence the subsequent stages. In this work, we propose a novel lightweight system for contour detection that achieves state-of-the-art performance while keeps ultra-slim model size. The proposed method is built on an efficient encoder in a bottom-up/top-down fashion. Specially, we propose a novel decoder that compresses side features from an encoder and effectively decodes compact contextual information for high-accurate boundary localization. Besides, we propose a novel loss function that is able to assist a model to produce crisp object boundaries. Ruoxi Deng, Shengjun Liu 0002, Huibing Wang, Hanli Zhao, Xiaoqin Zhang 0002 |
ACM Multimedia | 6 |
| 2021 | Multiple object tracking: A literature review
Wenhan Luo, Junliang Xing, Anton Milan, Xiaoqin Zhang 0002, Wei Liu 0005, Tae-Kyun Kim 0001 |
Artif. Intell. | 4 |
| 2021 | Video Deblurring via Spatiotemporal Pyramid Network and Adversarial Gradient Prior
Tao Wang 0052, Xiaoqin Zhang 0002, Runhua Jiang, Li Zhao 0005, Huiling Chen 0001, Wenhan Luo |
Comput. Vis. Image Underst. | 2 |
| 2021 | Self-filtering image dehazing with self-supporting module
Pengcheng Huang 0002, Li Zhao 0005, Runhua Jiang, Tao Wang 0052, Xiaoqin Zhang 0002 |
Neurocomputing | 5 |
| 2021 | Haze concentration adaptive network for image dehazing
Tao Wang 0052, Li Zhao 0005, Pengcheng Huang 0002, Xiaoqin Zhang 0002, Jiawei Xu 0004 |
Neurocomputing | 4 |
| 2021 | Attention-based interpolation network for video deblurring
Xiaoqin Zhang 0002, Runhua Jiang, Tao Wang 0052, Pengcheng Huang 0002, Li Zhao 0005 |
Neurocomputing | 1 |
| 2021 | Robust feature learning for adversarial defense via hierarchical feature alignment
Xiaoqin Zhang 0002, Tao Wang 0052, Runhua Jiang, Jiawei Xu 0004, Li Zhao 0005 |
Inf. Sci. | 1 |
| 2021 | Orthogonal learning covariance matrix for defects of grey wolf optimizer: Insights, balance, diversity, and feature selection
Jiao Hu, Huiling Chen 0001, Ali Asghar Heidari, Mingjing Wang, Xiaoqin Zhang 0002, Ying Chen 0023, Zhifang Pan |
Knowl. Based Syst. | 5 |
| 2021 | Evolutionary biogeography-based whale optimization methods with communication structure: Towards measuring the balance
Jiaze Tu, Huiling Chen 0001, Jiacong Liu, Ali Asghar Heidari, Xiaoqin Zhang 0002, Mingjing Wang, Rukhsana Ruby, Quoc-Viet Pham |
Knowl. Based Syst. | 5 |
| 2021 | Robust Low-Rank Tensor Recovery with Rectification and AlignmentabstractLow-rank tensor recovery in the presence of sparse but arbitrary errors is an important problem with many practical applications. In this work, we propose a general framework that recovers low-rank tensors, in which the data can be deformed by some unknown transformations and corrupted by arbitrary sparse errors. We give a unified presentation of the surrogate-based formulations that incorporate the features of rectification and alignment simultaneously, and establish worst-case error bounds of the recovered tensor. In this context, the state-of-the-art methods 'RASL' and 'TILT' can be viewed as two special cases of our work, and yet each only performs part of the function of our method. Subsequently, we study the optimization aspects of the problem in detail by deriving two algorithms, one based on the alternating direction method of multipliers (ADMM) and the other based on proximal gradient. We provide convergence guarantees for the latter algorithm, and demonstrate the performance of the former through in-depth simulations. Finally, we present extensive experimental results on public datasets to demonstrate the effectiveness and efficiency of the proposed framework and algorithms. Xiaoqin Zhang 0002, Di Wang 0008, Zhengyuan Zhou, Yi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Recursive Neural Network for Video DeblurringabstractVideo deblurring is still a challenging low-level vision task since spatio-temporal characteristics across both the spatial and temporal domains are difficult to model. In this article, to model the temporal information, we develop a non-local block which estimates inter-frame similarity and inter-frame difference. Specially, for modeling the spatial characteristics and restoring sharp frame details, we propose a recursive block that iteratively refines feature maps generated at the last iteration. In addition, a novel temporal loss function is introduced to ensure the temporal consistency of generated frames. Experimental results on public datasets demonstrate that our method achieves state-of-the-art performance both quantitatively and qualitatively. Xiaoqin Zhang 0002, Runhua Jiang, Tao Wang 0052 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Multi-Level Fusion and Attention-Guided CNN for Image DehazingabstractIn this paper, we tackle the problem of single image dehazing with a convolutional neural network. Within this network, we develop a multi-level fusion module to utilize both low-level and high-level features. The low-level features help to recover finer details, and the high-level features discover abstract semantics. They are complementary in the restoring of clear images. Moreover, a Residual Mixed-convolution Attention Module (RMAM) with an attention block is proposed to guide the network to focus on important features in the learning process. In this RMAM, group convolution, depth-wise convolution, and point-wise convolution are mixed, and thus it is much faster than its counterparts. With these two modules, we thus have an end-to-end network without explicitly estimating the atmospheric light intensity and the transmission map in the classical atmosphere scattering model. Both qualitative and quantitative experimental studies are carried out on public datasets including RESIDE, DCPDN-TestA, and the real-world dataset. The extensive results demonstrate both the effectiveness and efficiency of the proposed solution to single image dehazing. Xiaoqin Zhang 0002, Tao Wang 0052, Wenhan Luo, Pengcheng Huang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | HCE: Hierarchical Context Embedding for Region-Based Object DetectionabstractState-of-the-art two-stage object detectors apply a classifier to a sparse set of object proposals, relying on region-wise features extracted by RoIPool or RoIAlign as inputs. The region-wise features, in spite of aligning well with the proposal locations, may still lack the crucial context information which is necessary for filtering out noisy background detections, as well as recognizing objects possessing no distinctive appearances. To address this issue, we present a simple but effective Hierarchical Context Embedding (HCE) framework, which can be applied as a plug-and-play component, to facilitate the classification ability of a series of region-based detectors by mining contextual cues. Specifically, to advance the recognition of context-dependent object categories, we propose an image-level categorical embedding module which leverages the holistic image-level context to learn object-level concepts. Then, novel RoI features are generated by exploiting hierarchically embedded context information beneath both whole images and interested regions, which are also complementary to conventional RoI features. Moreover, to make full use of our hierarchical contextual RoI features, we propose the early-and-late fusion strategies (i.e., feature fusion and confidence fusion), which can be combined to boost the classification accuracy of region-based detectors. Comprehensive experiments demonstrate that our HCE framework is flexible and generalizable, leading to significant and consistent improvements upon various region-based detectors, including FPN, Cascade R-CNN, Mask R-CNN and PA-FPN. With simple modification, our HCE framework can be conveniently adapted to fit the structure of one-stage detectors, and achieve improved performance for SSD, RetinaNet and EfficientDet. Xin Jin 0023, Borui Zhao, Xiaoqin Zhang 0002, Yanwen Guo 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Single Image Dehazing via Dual-Path Recurrent NetworkabstractAn image can be decomposed into two parts: the basic content and details, which usually correspond to the low-frequency and high-frequency information of the image. For a hazy image, these two parts are often affected by haze in different levels, e.g., high-frequency parts are often affected more serious than low-frequency parts. In this paper, we approach the single image dehazing problem as two restoration problems of recovering basic content and image details, and propose a Dual-Path Recurrent Network (DPRN) to simultaneously tackle these two problems. Specifically, the core structure of DPRN is a dual-path block, which uses two parallel branches to learn the characteristics of the basic content and details of hazy images. Each branch consists of several Convolutional LSTM blocks and convolution layers. Moreover, a parallel interaction function is incorporated into the dual-path block, thus enables each branch to dynamically fuse the intermediate features of both the basic content and image details. In this way, both branches can benefit from each other, and recover the basic content and image details alternately, therefore alleviating the color distortion problem in the dehazing process. Experimental results show that the proposed DPRN outperforms state-of-the-art image dehazing methods in terms of both quantitative accuracy and qualitative visual effect. Xiaoqin Zhang 0002, Runhua Jiang, Tao Wang 0052, Wenhan Luo |
IEEE Trans. Image Process. | 1 |
| 2021 | Self-weighted Robust LDA for Multiclass Classification with Edge ClassesabstractLinear discriminant analysis (LDA) is a popular technique to learn the most discriminative features for multi-class classification. A vast majority of existing LDA algorithms are prone to be dominated by the class with very large deviation from the others, i.e., edge class, which occurs frequently in multi-class classification. First, the existence of edge classes often makes the total mean biased in the calculation of between-class scatter matrix. Second, the exploitation of ℓ2-norm based between-class distance criterion magnifies the extremely large distance corresponding to edge class. In this regard, a novel self-weighted robust LDA with ℓ2,1-norm based pairwise between-class distance criterion, called SWRLDA, is proposed for multi-class classification especially with edge classes. SWRLDA can automatically avoid the optimal mean calculation and simultaneously learn adaptive weights for each class pair without setting any additional parameter. An efficient re-weighted algorithm is exploited to derive the global optimum of the challenging ℓ2,1-norm maximization problem. The proposed SWRLDA is easy to implement and converges fast in practice. Extensive experiments demonstrate that SWRLDA performs favorably against other compared methods on both synthetic and real-world datasets while presenting superior computational efficiency in comparison with other techniques. Caixia Yan, Xiaojun Chang, Minnan Luo, Xiaoqin Zhang 0002, Zhihui Li 0001, Feiping Nie 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2021 | Multiple Lane Detection via Combining Complementary Structural ConstraintsabstractMany studies have been conducted on single lane detection, but multi-lane detection is rarely addressed. The latter is more advantageous for applications such as autonomous navigation, unmanned vehicles, departure warning, and cruise control. In this paper, we propose a novel and robust multiple lane detection algorithm based on the road structure information, which contains five complementary constraints: length constraint, parallel constraint, distribution constraint, pair constraint and uniform width constraint. All the five constraints are incorporated into a Hough transform (HT) based unified framework to select lane candidates. Nearly 99% of the false alarm candidates in HT space can be removed. Moreover, a dynamic programming strategy is proposed to find the most rational solutions among the remaining candidates. This strategy can effectively deal with combination complexity and interferences introduced by multi-lane detection. Experimental results on the benchmark dataset and other collected data demonstrate that the proposed method can outperform the state-of-the-art approaches in both accuracy and efficiency. Sheng Luo 0003, Xiaoqin Zhang 0002, Jie Hu 0041, Jinghua Xu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Top-k Feature Selection Framework Using Robust 0-1 Integer ProgrammingabstractFeature selection (FS), which identifies the relevant features in a data set to facilitate subsequent data analysis, is a fundamental problem in machine learning and has been widely studied in recent years. Most FS methods rank the features in order of their scores based on a specific criterion and then select the k top-ranked features, where k is the number of desired features. However, these features are usually not the top- k features and may present a suboptimal choice. To address this issue, we propose a novel FS framework in this article to select the exact top- k features in the unsupervised, semisupervised, and supervised scenarios. The new framework utilizes thel0,2-norm as the matrix sparsity constraint rather than its relaxations, such as thel1,2-norm. Since thel0,2-norm constrained problem is difficult to solve, we transform the discretel0,2-norm-based constraint into an equivalent 0-1 integer constraint and replace the 0-1 integer constraint with two continuous constraints. The obtained top- k FS framework with two continuous constraints is theoretically equivalent to thel0,2-norm constrained problem and can be optimized by the alternating direction method of multipliers (ADMM). Unsupervised and semisupervised FS methods are developed based on the proposed framework, and extensive experiments on real-world data sets are conducted to demonstrate the effectiveness of the proposed FS framework. Xiaoqin Zhang 0002, Mingyu Fan, Di Wang 0008, Peng Zhou 0006, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Multi-level Feature Fusion Network for Single Image Super-ResolutionabstractRecently, deep convolution neural networks have achieved remarkable performance in the task of single image super-resolution (SISR). However, effectiveness of existing networks highly relies on their receptive field, which always increases with the depth of the network. In this work, we propose a novel module, named as residual group, to effectively learn feature maps by using dynamic receptive field. This residual group firstly uses a selective kernel convolution layer to dynamically learn multi-scale information from its input features. Then, several residual blocks are employed to further refine the learned feature. In addition, we also propose a selective feature fusion module to fuse appearance information in multi-level features. Within this module, the low-level features and high-level features are selectively fused to complement the high-level ones. Finally, by combining these two methods, we introduce a multi-level feature fusion network (MLFFN) for single image super-resolution (SISR). Through comprehensive experiments, we demonstrate that the proposed MLFFN achieves state-of-the-art performance both quantitatively and qualitatively. Xinxia Zhang, Xiaoqin Zhang 0002, Li Zhao 0005, Runhua Jiang, Pengcheng Huang 0002, Jiawei Xu 0004 |
IEEE BigData | 2 |
| 2020 | Pyramid Channel-based Feature Attention Network for image dehazing
Xiaoqin Zhang 0002, Tao Wang 0052, Guiying Tang, Li Zhao 0005 |
Comput. Vis. Image Underst. | 1 |
| 2020 | Improvement of viewing experience on stereoscopic image guided by human stereo vision
Jiawei Xu 0004, Seop Hyeong Park, Xiaoqin Zhang 0002 |
Multim. Tools Appl. | 3 |
| 2020 | Exemplar-Based Denoising: A Unified Low-Rank Recovery FrameworkabstractExemplar-based image denoising algorithms have shown great potential for image restoration with a multitude of existing models. In this paper, we interpret nonlocal similar patch-based denoising as a problem of low-rank recovery. This offers a physically plausible model and unifies several existing techniques in a single low-rank recovery framework. The framework can handle complex noise models, such as zero-mean Gaussian noise, impulse noise, and any other noise that can be approximated by mixing these two kinds of noise. Moreover, we introduce a new nonconvex surrogate for the $l_{0}$ -norm and find the optimal solution of the optimization problems when the new norm is applied to low-rank recovery. The experimental results with different kinds of noise confirm the effectiveness of the proposed low-rank recovery framework and the new norm. Xiaoqin Zhang 0002, Di Wang 0008, Li Zhao 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | A Temporally Irreversible Visual Attention Model Inspired by Motion Sensitive NeuronsabstractWhen a human perceives videos composed of the same images in various orders, such as normal order, reverse order, and random order, the human visual attention system perceives them as different visual inputs. This means that the temporal change in the image sequence exerts a considerable influence on the human visual system. However, most state-of-the-art computational visual attention models have not considered the temporal cues adequately. Motivated by this deficiency, we propose a novel temporally irreversible visual attention model considering the following three aspects. First, the central bias of human dynamic vision is incorporated into the model to manifest this tendency. Second, the depth and directional motion-sensitive neurons are fused to discern different motion patterns. Third, the rarity factor is integrated into the model to mimic the attention shift when human observer perceives new emerging motion cues. The proposed model demonstrates its competitiveness to select attentive events in our experiments, in both laboratory setting and real driving video clips. When compared with recent visual attention models, the proposed model achieves the highest score in similarity with human dynamic vision. The proposed model could be one of the fundamental building blocks for any visual attention systems coping with dynamic scenes. Jiawei Xu 0004, Seop Hyeong Park, Xiaoqin Zhang 0002 |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Self-Taught Semisupervised Dictionary Learning With Nonnegative ConstraintabstractThis paper investigates classification by dictionary learning. A novel unified framework termed self-taught semisupervised dictionary learning with nonnegative constraint is proposed for simultaneously optimizing the components of a dictionary and a graph Laplacian. Specifically, an atom graph Laplacian regularization is built by using sparse coefficients to effectively capture the underlying manifold structure. It is more robust to noisy samples and outliers because atoms are more concise and representative than training samples. A nonnegative constraint imposed on the sparse coefficients guarantees that each sample is in the middle of its related atoms. In this way, the dependency between samples and atoms is made explicit. Furthermore, a self-taught mechanism is introduced to effectively feed back the manifold structure induced by atom graph Laplacian regularization and the supervised information hidden in unlabeled samples in order to learn a better dictionary. An efficient algorithm, combining a block coordinate descent method with the alternating direction method of multipliers, is derived to optimize the unified framework. Experimental results on several benchmark datasets show the effectiveness of the proposed model. Xiaoqin Zhang 0002, Di Wang 0008, Li Zhao 0005, Nannan Gu, Stephen J. Maybank |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Pair-based Uncertainty and Diversity Promoting Early Active Learning for Person Re-identificationabstractThe effective training of supervised Person Re-identification (Re-ID) models requires sufficient pairwise labeled data. However, when there is limited annotation resource, it is difficult to collect pairwise labeled data. We consider a challenging and practical problem called Early Active Learning, which is applied to the early stage of experiments when there is no pre-labeled sample available as references for human annotating. Previous early active learning methods suffer from two limitations for Re-ID. First, these instance-based algorithms select instances rather than pairs, which can result in missing optimal pairs for Re-ID. Second, most of these methods only consider the representativeness of instances, which can result in selecting less diverse and less informative pairs. To overcome these limitations, we propose a novel pair-based active learning for Re-ID. Our algorithm selects pairs instead of instances from the entire dataset for annotation. Besides representativeness, we further take into account the uncertainty and the diversity in terms of pairwise relations. Therefore, our algorithm can produce the most representative, informative, and diverse pairs for Re-ID data annotation. Extensive experimental results on five benchmark Re-ID datasets have demonstrated the superiority of the proposed pair-based early active learning algorithm. Wenhe Liu, Xiaojun Chang, Ling Chen 0006, Dinh Q. Phung, Xiaoqin Zhang 0002, Yi Yang 0001, Alex Hauptmann 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2019 | Zero-Shot Object Detection with Textual DescriptionsabstractObject detection is important in real-world applications. Existing methods mainly focus on object detection with sufficient labelled training data or zero-shot object detection with only concept names. In this paper, we address the challenging problem of zero-shot object detection with natural language description, which aims to simultaneously detect and recognize novel concept instances with textual descriptions. We propose a novel deep learning framework to jointly learn visual units, visual-unit attention and word-level attention, which are combined to achieve word-proposal affinity by an element-wise multiplication. To the best of our knowledge, this is the first work on zero-shot object detection with textual descriptions. Since there is no directly related work in the literature, we investigate plausible solutions based on existing zero-shot object detection for a fair comparison. We conduct extensive experiments on three challenging benchmark datasets. The extensive experimental results confirm the superiority of the proposed model. Zhihui Li 0001, Lina Yao 0001, Xiaoqin Zhang 0002, Xianzhi Wang 0001, Salil S. Kanhere, Huaxiang Zhang 0001 |
AAAI | 3 |
| 2019 | Multi-View Subspace Clustering based on Tensor Schatten-p NormabstractIn this paper, we focus on the multi-view clustering problem. A novel multi-view clustering framework, called multi-view subspace clustering based on tensor Schatten-p norm (MVSC-TSP), is proposed for clustering task. In our method, the tensor Schatten-p norm, which is based on tensor singular value decomposition, is utilized to explore the global low-rank structure of multi-view self-representations. Since 0<; p<; 1, using tensor Schatten-p norm to relax the tensor multi-rank is more effective than the commonly used tensor nuclear norm. Furthermore, we present a new generalized tensor soft thresholding algorithm to solve the tensor Schatten-p norm minimization problem. Based on this, the proposed non-convex optimal problem can be efficiently solved by the alternating direction method of multipliers. Experimental results on image clustering demonstrate that the proposed method is superior to the state-of-the-art methods in term of various evaluation metrics. Yongli Liu, Xiaoqin Zhang 0002, Guiying Tang, Di Wang 0008 |
IEEE BigData | 2 |
| 2019 | Structural Dictionary Learning based on Non-convex Surrogate of ℓ₂, ₁ Norm for ClassificationabstractRecently, group sparse representation which is based on a hypothesis about correlation of coefficient variables has attracted much attention due to its effectiveness and robustness in dictionary learning. Traditional group sparse representation methods use ℓ2,1norm to enforce the estimation of models with joint sparsity patterns, which often leads to over-punishment phenomenon. To solve this issue, we replace ℓ2,1with non-convex surrogate of ℓ2,1, and give a general solver for the corresponding optimization algorithm. Experimental results confirm the effectiveness of our proposed method. Xiaoju Lu, Guiying Tang, Di Wang 0008, Xiaoqin Zhang 0002 |
IEEE BigData | 4 |
| 2019 | Single Image Dehazing via Lightweight Multi-scale NetworksabstractSingle image haze removal is a challenging ill-posed problem in computer vision. Instead of leveraging the traditional model or handcrafted image priors, an end-to-end multi-scale convolutional neural network is proposed for single image haze removal task by directly mapping the hazy image to its corresponding haze-free image. To better retain the coarse and fine information, a multi-scale block is elaborated and embedded into the proposed architecture. This block can extract the feature at varying scales with a model size that is as small as possible. The global skip connection is adopted to promote the model performance. Extensive experiment results demonstrate that the proposed network outperforms the state-of-the-art single image haze removal algorithms on both synthetical and real-world images. In addition, the size of the model in this paper dominates among the high performance methods based on convolutional neural networks. Guiying Tang, Li Zhao 0005, Runhua Jiang, Xiaoqin Zhang 0002 |
IEEE BigData | 4 |
| 2019 | Single-Image Dehazing Using Color Attenuation Prior Based on Haze-LinesabstractIn this paper, we propose a new single-image dehazing method for synthetic and real-world hazy images. Based on the color attenuation prior, this proposed dehazing method improves it in two aspects. First, we estimate the atmospheric light with the haze-lines prior, which is based on the observation that pixel values of a hazy image can be modeled as lines in the RGB color space that intersects at the air-light. Second, the dynamic scattering coefficient, which is an exponential function of image depth, is proposed to replace the constant scattering coefficient. Experimental results demonstrate that the dehazed image of proposed algorithm is clearer and more natural than that of the color attenuation prior. The proposed algorithm can effectively improve the effect of dehazing. Qianru Wang, Li Zhao 0005, Guiying Tang, Hanli Zhao, Xiaoqin Zhang 0002 |
IEEE BigData | 5 |
| 2019 | A bio-inspired motion sensitive model and its application to estimating human gaze positions under classified driving conditions
Jiawei Xu 0004, Seop Hyeong Park, Xiaoqin Zhang 0002 |
Neurocomputing | 3 |
| 2019 | Enhanced Moth-flame optimizer with mutation strategy for global optimization
Yueting Xu, Huiling Chen 0001, Jie Luo 0002, Qian Zhang 0049, Shan Jiao, Xiaoqin Zhang 0002 |
Inf. Sci. | 6 |
| 2019 | Semantic Segmentation of Remote Sensing Images Using Multiscale Decoding NetworkabstractIn this letter, we propose a practical convolutional neural network architecture for semantic pixelwise segmentation of remote sensing images, named Multiscale Decoding Network. The proposed method is built on the success of fully convolutional networks (FCNs) and the transfer of pretrained networks. The decoding network of our architecture utilizes the combination of three paths, namely, unpooling path, transposed convolution path, and dilated convolution path, in the form of an inception module. The whole network is trained in the end-to-end manner and the parameters of the three paths are learned automatically. Since the proposed method transfers the feature of pretrained networks and has three simplified decoding paths with fewer parameters, it requires less training data and training time. Compared with the classical networks FCN, SegNet, and U-net, our network shows better performance on remote sensing images segmentation. Xiaoqin Zhang 0002, Zhiheng Xiao, Mingyu Fan, Li Zhao 0005 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | Constructive Neural Network LearningabstractIn this paper, we aim at developing scalable neural network-type learning systems. Motivated by the idea of constructive neural networks in approximation theory, we focus on constructing rather than training feed-forward neural networks (FNNs) for learning, and propose a novel FNNs learning system called the constructive FNN (CFN). Theoretically, we prove that the proposed method not only overcomes the classical saturation problem for constructive FNN approximation, but also reaches the optimal learning rate when the regression function is smooth, while the state-of-the-art learning rates established for traditional FNNs are only near optimal (up to a logarithmic factor). A series of numerical simulations are provided to show the efficiency and feasibility of CFN. Shaobo Lin, Jinshan Zeng, Xiaoqin Zhang 0002 |
IEEE Trans. Cybern. | 3 |
| 2018 | Semi-Supervised Dictionary Learning Based on Atom Graph RegularizationabstractIn this paper, we propose a novel unified optimization framework for semi-supervised dictionary learning, which optimizes a graph Laplacian component and the dictionary simultaneously. In the framework, the graph Laplacian is defined on the atoms and the corresponding sparse codings. Since the atoms are more concise and representative than the original training samples, the constructed graph Laplacian can not only effectively capture the manifold structure of training samples, but also be more robust to noise and outliers. Moreover, the dictionary and the graph Laplacian can facilitate each other during the learning iterations. We derive an efficient algorithm by combining the block coordinate descent method with the alternating direction method of multipliers to solve the unified optimization problem. Extensive experimental evaluation on several challenging datasets demonstrates the superior performance of the proposed method. Xiaoqin Zhang 0002, Di Wang 0008, Jie Hu 0041, Nannan Gu, Tianhao Wang 0006 |
IEEE BigData | 1 |
| 2018 | Simultaneous Learning of Affinity Matrix and Laplacian Regularized Least Squares for Semi-Supervised ClassificationabstractGraph based Semi-Supervised Learning (G-SSL) methods usually include the stages of the construction of affinity matrix and the mechanism of inferring unknown labels. However, solving each of the stages individually does not fully exploit the potential relationship between the affinity matrix and the labels of samples. In this paper, we formulate the global self-expressiveness induced affinity and Laplacian Regularized Least Squares (LapRLS) into a single optimization model, called as Self-Taught LapRLS (ST-LapRLS). In the unified model, both the given labels and the estimated labels are used to build a better affinity matrix and to facilitate the LapRLS classifer. The proposed ST-LapRLS classifier is explicit, and can be easily extended to deal with out-of-sample problem. We propose an efficient algorithm which combines the alternating direction method of multiplier and LapRLS to solve the unified optimization problem. Experiments on several Benchmark datasets show the superior performance of our method in classification applications. Di Wang 0008, Xiaoqin Zhang 0002, Nannan Gu, Mingyu Fan |
ICIP | 3 |
| 2018 | Semi-Supervised Learning Through Label Propagation on GeodesicsabstractGraph-based semi-supervised learning (SSL) has attracted great attention over the past decade. However, there are still several open problems in this paper, including: 1) how to construct an effective graph over data with complex distribution and 2) how to define and effectively use pair-wise similarity for robust label propagation. In this paper, we utilize a simple and effective graph construction method to construct the graph over data lying on multiple data manifolds. The method can guarantee the connectiveness between pair-wise data points. Then, the global pair-wise data similarity is naturally characterized by geodesic distance-based joint probability, where the geodesic distance is approximated by the graph distance. The new data similarity is much more effective than previous Euclidean distance-based similarities. To apply data structure for robust label propagation, Kullback-Leibler divergence is utilized to measure the inconsistency between the input pair-wise similarity and the output similarity. In order to further consider intraclass and interclass variances, a novel regularization term on sample-wise margins is introduced to the objective function. This enables the proposed method fully utilizes the input data structure and the label information for classification. An efficient optimization method and the convergence analysis have been proposed for our problem. Besides, out-of-sample extension is discussed and addressed. Comparisons with the state-of-the-art SSL methods on image classification tasks have been presented to show the effectiveness of the proposed method. Mingyu Fan, Xiaoqin Zhang 0002, Liang Du 0003, Liang Chen 0012, Dacheng Tao |
IEEE Trans. Cybern. | 2 |
| 2017 | Top-k Supervise Feature Selection via ADMM for Integer ProgrammingabstractRecently, structured sparsity inducing based feature selection has become a hot topic in machine learning and pattern recognition. Most of the sparsity inducing feature selection methods are designed to rank all features by certain criterion and then select the k top ranked features, where k is an integer. However, the k top features are usually not the top k features and therefore maybe a suboptimal result. In this paper, we propose a novel supervised feature selection method to directly identify the top k features. The new method is formulated as a classic regularized least squares regression model with two groups of variables. The problem with respect to one group of the variables turn out to be a 0-1 integer programming, which had been considered very hard to solve. To address this, we utilize an efficient optimization method to solve the integer programming, which first replaces the discrete 0-1 constraints with two continuous constraints and then utilizes the alternating direction method of multipliers to optimize the equivalent problem. The obtained result is the top subset with k features under the proposed criterion rather than the subset of k top features. Experiments have been conducted on benchmark data sets to show the effectiveness of proposed method. Mingyu Fan, Xiaojun Chang, Xiaoqin Zhang 0002, Di Wang 0008, Liang Du 0003 |
IJCAI | 3 |
| 2017 | Segmentation and Quantification for Angle-Closure Glaucoma Assessment in Anterior Segment OCTabstractAngle-closure glaucoma is a major cause of irreversible visual impairment and can be identified by measuring the anterior chamber angle (ACA) of the eye. The ACA can be viewed clearly through anterior segment optical coherence tomography (AS-OCT), but the imaging characteristics and the shapes and locations of major ocular structures can vary significantly among different AS-OCT modalities, thus complicating image analysis. To address this problem, we propose a data-driven approach for automatic AS-OCT structure segmentation, measurement, and screening. Our technique first estimates initial markers in the eye through label transfer from a hand-labeled exemplar data set, whose images are collected over different patients and AS-OCT modalities. These initial markers are then refined by using a graph-based smoothing method that is guided by AS-OCT structural information. These markers facilitate segmentation of major clinical structures, which are used to recover standard clinical parameters. These parameters can be used not only to support clinicians in making anatomical assessments, but also to serve as features for detecting anterior angle closure in automatic glaucoma screening algorithms. Experiments on Visante AS-OCT and Cirrus high-definition-OCT data sets demonstrate the effectiveness of our approach. Huazhu Fu, Yanwu Xu 0001, Stephen Lin 0001, Xiaoqin Zhang 0002, Damon Wing Kee Wong, Jiang Liu 0001, Alejandro F. Frangi, Mani Baskaran, Tin Aung |
IEEE Trans. Medical Imaging | 4 |
| 2016 | Semi-Supervised Dictionary Learning via Structural Sparse PreservingabstractWhile recent techniques for discriminative dictionary learning have attained promising results on the classification tasks, their performance is highly dependent on the number of labeled samples available for training. However, labeling samples is expensive and time consuming due to the significant human effort involved. In this paper, we present a novel semi- supervised dictionary learning method which utilizes the structural sparse relationships between the labeled and unlabeled samples. Specifically, by connecting the sparse reconstruction coefficients on both the original samples and dictionary, the unlabeled samples can be automatically grouped to the different labeled samples, and the grouped samples share a small number of atoms in the dictionary via mixed l2p- norm regularization. This makes the learned dictionary more representative and discriminative since the shared atoms are learned by using the labeled and unlabeled samples potentially from the same class. Minimizing the derived objective function is a challenging task because it is non-convex and highly non-smooth. We propose an efficient optimization algorithm to solve the problem based on the block coordinate descent method. Moreover, we have a rigorous proof of the convergence of the algorithm. Extensive experiments are presented to show the superior performance of our method in classification applications. Di Wang 0008, Xiaoqin Zhang 0002, Mingyu Fan, Xiuzi Ye |
AAAI | 2 |
| 2016 | Axial Alignment for Anterior Segment Swept Source Optical Coherence Tomography via Robust Low-Rank Tensor RecoveryabstractWe present a one-step approach based on low-rank tensor recovery for axial alignment in 360-degree anterior chamber optical coherence tomography. Achieving translational alignment and rotation correction of cross-sections simultaneously, this technique obtains a better anterior segment topographical representation and improves quantitative measurement accuracy and reproducibility of disease related parameters. Through its use of global information, the proposed method is more robust compared to using only individual or paired slices, and less sensitive to noise and motion artifacts. In angle closure analysis on 30 patient eyes, the preliminary results indicate that the proposed axial alignment method can not only facilitate manual qualitative analysis with more distinct landmark representation and much less human labor, but also can improve the accuracy of automatic quantitative assessment by 2.9 %, which demonstrates that the proposed approach is promising for a wide range of clinical applications. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Yanwu Xu 0001, Lixin Duan, Huazhu Fu, Xiaoqin Zhang 0002, Damon Wing Kee Wong, Mani Baskaran, Tin Aung, Jiang Liu 0001 |
MICCAI (3) | 4 |
| 2016 | Efficient isometric multi-manifold learning based on the self-organizing method
Mingyu Fan, Xiaoqin Zhang 0002, Hong Qiao, Bo Zhang 0006 |
Inf. Sci. | 2 |
| 2016 | Hierarchical mixing linear support vector machines for nonlinear classification
Di Wang 0008, Xiaoqin Zhang 0002, Mingyu Fan, Xiuzi Ye |
Pattern Recognit. | 2 |
| 2015 | Local Subspace Collaborative TrackingabstractSubspace models have been widely used for appearance based object tracking. Most existing subspace based trackers employ a linear subspace to represent object appearances, which are not accurate enough to model large variations of objects. To address this, this paper presents a local subspace collaborative tracking method for robust visual tracking, where multiple linear and nonlinear subspaces are learned to better model the nonlinear relationship of object appearances. First, we retain a set of key samples and compute a set of local subspaces for each key sample. Then, we construct a hyper sphere to represent the local nonlinear subspace for each key sample. The hyper sphere of one key sample passes the local key samples and also is tangent to the local linear subspace of the specific key sample. In this way, we are able to represent the nonlinear distribution of the key samples and also approximate the local linear subspace near the specific key sample, so that local distributions of the samples can be represented more accurately. Experimental results on challenging video sequences demonstrate the effectiveness of our method. Xiaoqin Zhang 0002, Weiming Hu 0004, Junliang Xing, Jiwen Lu, Jie Zhou 0001 |
ICCV | 2 |
| 2015 | An Efficient Classifier Based on Hierarchical Mixing Linear Support Vector Machines
Di Wang 0008, Xiaoqin Zhang 0002, Mingyu Fan, Xiuzi Ye |
IJCAI | 2 |
| 2015 | Multi-Modality Tracker Aggregation: From Generative to Discriminative
Xiaoqin Zhang 0002, Wei Li 0034, Mingyu Fan, Di Wang 0008, Xiuzi Ye |
IJCAI | 1 |
| 2015 | A Robust Tracking System for Low Frame Rate Video
Xiaoqin Zhang 0002, Weiming Hu 0004, Nianhua Xie, Hujun Bao, Stephen J. Maybank |
Int. J. Comput. Vis. | 1 |
| 2015 | Erratum to: A Robust Tracking System for Low Frame Rate Video
Xiaoqin Zhang 0002, Weiming Hu 0004, Nianhua Xie, Hujun Bao, Stephen J. Maybank |
Int. J. Comput. Vis. | 1 |
| 2015 | Robust hand tracking via novel multi-cue integration
Xiaoqin Zhang 0002, Wei Li 0034, Xiuzi Ye, Stephen J. Maybank |
Neurocomputing | 1 |
| 2015 | Single and Multiple Object Tracking Using a Multi-Feature Joint Sparse RepresentationabstractIn this paper, we propose a tracking algorithm based on a multi-feature joint sparse representation. The templates for the sparse representation can include pixel values, textures, and edges. In the multi-feature joint optimization, noise or occlusion is dealt with using a set of trivial templates. A sparse weight constraint is introduced to dynamically select the relevant templates from the full set of templates. A variance ratio measure is adopted to adaptively adjust the weights of different features. The multi-feature template set is updated adaptively. We further propose an algorithm for tracking multi-objects with occlusion handling based on the multi-feature joint sparse reconstruction. The observation model based on sparse reconstruction automatically focuses on the visible parts of an occluded object by using the information in the trivial templates. The multi-object tracking is simplified into a joint Bayesian inference. The experimental results show the superiority of our algorithm over several state-of-the-art tracking algorithms. Weiming Hu 0004, Wei Li 0034, Xiaoqin Zhang 0002, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | An Efficient Semi-Supervised Classifier Based on Block-Polynomial MappingabstractIn this paper, we propose a block-polynomial mapping for image feature learning, which can be efficiently represented by the matrix Khatri-Rao product. The block-polynomial mapping not only captures the local discriminative information within the image structure, but is also much more efficient than the traditional kernel mapping. Moreover, we embed the proposed mapping into the manifold regularization framework for semi-supervised image classification. Experimental results demonstrate that, while maintaining a comparable classification accuracy, the proposed algorithm performs much more efficient than the state-of-the-art methods. Di Wang 0008, Xiaoqin Zhang 0002, Mingyu Fan, Xiuzi Ye |
IEEE Signal Process. Lett. | 2 |
| 2014 | Hybrid Singular Value Thresholding for Tensor CompletionabstractIn this paper, we study the low-rank tensor completion problem, where a high-order tensor with missing entries is given and the goal is to complete the tensor. We propose to minimize a new convex objective function, based on log sum of exponentials of nuclear norms, that promotes the low-rankness of unfolding matrices of the completed tensor. We show for the first time that the proximal operator to this objective function is readily computable through a hybrid singular value thresholding scheme. This leads to a new solution to high-order (low-rank) tensor completion via convex relaxation. We show that this convex relaxation and the resulting solution are much more effective than existing tensor completion methods (including those also based on minimizing ranks of unfolding matrices). The hybrid singular value thresholding scheme can be applied to any problem where the goal is to minimize the maximum rank of a set of low-rank matrices. Xiaoqin Zhang 0002, Zhengyuan Zhou, Di Wang 0008, Yi Ma 0001 |
AAAI | 1 |
| 2014 | A Regularized Approach for Geodesic-Based Semisupervised Multimanifold LearningabstractGeodesic distance, as an essential measurement for data dissimilarity, has been successfully used in manifold learning. However, most geodesic distance-based manifold learning algorithms have two limitations when applied to classification: 1) class information is rarely used in computing the geodesic distances between data points on manifolds and 2) little attention has been paid to building an explicit dimension reduction mapping for extracting the discriminative information hidden in the geodesic distances. In this paper, we regard geodesic distance as a kind of kernel, which maps data from linearly inseparable space to linear separable distance space. In doing this, a new semisupervised manifold learning algorithm, namely regularized geodesic feature learning algorithm, is proposed. The method consists of three techniques: a semisupervised graph construction method, replacement of original data points with feature vectors which are built by geodesic distances, and a new semisupervised dimension reduction method for feature vectors. Experiments on the MNIST, USPS handwritten digit data sets, MIT CBCL face versus nonface data set, and an intelligent traffic data set show the effectiveness of the proposed algorithm. Mingyu Fan, Xiaoqin Zhang 0002, Zhouchen Lin, Zhongfei Zhang, Hujun Bao |
IEEE Trans. Image Process. | 2 |
| 2014 | Human Pose Estimation and Tracking via Parsing a Tree Structure Based Human ModelabstractHuman pose estimation and tracking is the task of determining the states (location, orientation, and scale) of each body part over time. It is important for many vision understanding applications, such as visual interactive gaming, immersive virtual reality, visual surveillance, and content-based image retrieval. However, it remains a challenging task due to unknown image background, presence of clutter and especially the high dimensional state space (usually 30+ dimensions). In this paper, we contribute to human pose estimation and tracking in two aspects. First, we design two efficient Markov Chain dynamics under the data-driven Markov Chain Monte Carlo framework to effectively explore the high dimensional state space. Second, we parse the tree structure state space into a lexicographic order according to the image observations and body topology, and the optimization process is conducted in this order. This realizes a much more efficient exploration of the state space than the sampling based search or exhaustive search, and thus achieves a tremendous speed-up. Experimental results demonstrate the efficiency and effectiveness of the proposed method in estimating and tracking various kinds of human poses, even against cluttered backgrounds, in poor illumination or under partial self-occlusion. Xiaoqin Zhang 0002, Weiming Hu 0004, Xiaofeng Tong, Stephen J. Maybank, Yimin Zhang 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2013 | Distance Map of Various Weights: A new feature for adaptive object trackingabstractIn this paper, we propose a new feature, Distance Map of Various Weights (DMVW) based on distances between rows' textures, to perform tracking. The proposed new feature provides an effective object appearance model which is both illumination-invariant and robust to occlusion. We also develop a 2D PCA based method to effectively evaluate the new feature. We demonstrate the validity of the rows' or column's weights in computing 2D PCA subspaces. To balance the importance of local and global information, we define a coefficient to revise the locality extent of the proposed feature. A new method based on entropy of candidate state evaluation is proposed to select the most discriminative coefficient. Experimental results on challenging video sequences demonstrated the effectiveness of our method. Junliang Xing, Xiaoqin Zhang 0002, Weiming Hu 0004 |
ICASSP | 3 |
| 2013 | Adaptive cooperative tracking based on multi-graph embedding and Markov Random FieldabstractAppearance model is of fundamental importance in a tracking algorithm. In this paper, we propose a new tracking method based on a cooperative object appearance model which incorporates both the discriminative and generative information. We represent the discriminative information with graph embedding (GE). To represent the local object appearance effectively, we divide the object and nearby background into patches. As the discriminative conditions around the 4 object boundaries are different, we divide the patches into 4 groups and perform GE for each group. Markov Random Filed (MRF) is designed to represent the generative information. We propose a novel MRF based method which not only considers the single patch's appearance but also the appearance relations between neighbor patches (not the relations between neighbor patches' states). The proposed cooperative appearance model can represent the object appearance's variation effectively and meanwhile discriminate the object from background robustly. Experimental results on challenging test sequences demonstrated the effectiveness of our method. Junliang Xing, Xiaoqin Zhang 0002, Weiming Hu 0004 |
ICASSP | 3 |
| 2013 | Simultaneous Rectification and Alignment via Robust Recovery of Low-rank TensorsabstractIn this work, we propose a general method for recovering low-rank three-order tensors, in which the data can be deformed by some unknown transformation and corrupted by arbitrary sparse errors. Since the unfolding matrices of a tensor are interdependent, we introduce auxiliary variables and relax the hard equality constraints by the augmented Lagrange multiplier method. To improve the computational efficiency, we introduce a proximal gradient step to the alternating direction minimization method. We have provided proof for the convergence of the linearized version of the problem which is the inner loop of the overall algorithm. Both simulations and experiments show that our methods are more efficient and effective than previous work. The proposed method can be easily applied to simultaneously rectify and align multiple images or videos frames. In this context, the state-of-the-art algorithms RASL'' and "TILT'' can be viewed as two special cases of our work, and yet each only performs part of the function of our method." Xiaoqin Zhang 0002, Di Wang 0008, Zhengyuan Zhou, Yi Ma 0001 |
NIPS | 1 |
| 2013 | Dimension estimation of image manifolds by minimal cover approximation
Mingyu Fan, Xiaoqin Zhang 0002, Shengyong Chen, Hujun Bao, Stephen J. Maybank |
Neurocomputing | 2 |
| 2013 | Block covariance based l1 tracker with a subtle template dictionary
Xiaoqin Zhang 0002, Wei Li 0034, Weiming Hu 0004, Haibin Ling, Stephen J. Maybank |
Pattern Recognit. | 1 |
| 2013 | Robust Head Tracking Based on Multiple Cues Fusion in the Kernel-Bayesian FrameworkabstractThis paper presents a robust head tracking algorithm based on multiple cues fusion in a kernel-Bayesian framework. In this algorithm, the object to be tracked is characterized using a spatial-constraint mixture of the Gaussians-based appearance model and a multichannel chamfer matching-based shape model. These two models complement each other and their combination is discriminative in distinguishing the object from the background. A selective updating technique for the appearance model is employed to accommodate appearance and illumination changes. Meantime, the kernel method-mean shift algorithm is embedded into the Bayesian framework to give a heuristic prediction in the hypotheses generation process. This alleviates the great computational load suffered by conventional Bayesian trackers. Experimental results demonstrate that the proposed algorithm is effective. Xiaoqin Zhang 0002, Weiming Hu 0004, Hujun Bao, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Active Contour-Based Visual Tracking by Integrating Colors, Shapes, and MotionsabstractIn this paper, we present a framework for active contour-based visual tracking using level sets. The main components of our framework include contour-based tracking initialization, color-based contour evolution, adaptive shape-based contour evolution for non-periodic motions, dynamic shape-based contour evolution for periodic motions, and the handling of abrupt motions. For the initialization of contour-based tracking, we develop an optical flow-based algorithm for automatically initializing contours at the first frame. For the color-based contour evolution, Markov random field theory is used to measure correlations between values of neighboring pixels for posterior probability estimation. For adaptive shape-based contour evolution, the global shape information and the local color information are combined to hierarchically evolve the contour, and a flexible shape updating model is constructed. For the dynamic shape-based contour evolution, a shape mode transition matrix is learnt to characterize the temporal correlations of object shapes. For the handling of abrupt motions, particle swarm optimization is adopted to capture the global motion which is applied to the contour in the current frame to produce an initial contour in the next frame. Weiming Hu 0004, Wei Li 0034, Wenhan Luo, Xiaoqin Zhang 0002, Stephen J. Maybank |
IEEE Trans. Image Process. | 5 |
| 2012 | Displacement Template with Divide-&-Conquer Algorithm for Significantly Improving Descriptor Based Face Recognition Approaches
Liang Chen 0012, Yonghuai Liu, Lixin Gao 0004, Xiaoqin Zhang 0002 |
ECCV (5) | 5 |
| 2012 | Isometric Multi-manifold Learning for Feature ExtractionabstractManifold learning is an important topic in pattern recognition and computer vision. However, most manifold learning algorithms implicitly assume the data are aligned on a single manifold, which is too strict in actual applications. Isometric feature mapping (Isomap), as a promising manifold learning method, fails to work on data which distribute on clusters in a single manifold or manifolds. In this paper, we propose a new multi-manifold learning algorithm (M-Isomap). The algorithm first discovers the data manifolds and then reduces the dimensionality of the manifolds separately. Meanwhile, a skeleton representing the global structure of whole data set is built and kept in low-dimensional space. Secondly, by referring to the low-dimensional representation of the skeleton, the embeddings of the manifolds are relocated to a global coordinate system. Compared with previous methods, these algorithms can keep both of the intra and inter manifolds geodesics faithfully. The features and effectiveness of the proposed multi-manifold learning algorithms are demonstrated and compared through experiments. Mingyu Fan, Hong Qiao, Bo Zhang 0006, Xiaoqin Zhang 0002 |
ICDM | 4 |
| 2012 | Geodesic Based Semi-supervised Multi-manifold Feature ExtractionabstractManifold learning is an important feature extraction approach in data mining. This paper presents a new semi-supervised manifold learning algorithm, called Multi-Manifold Discriminative Analysis (Multi-MDA). The proposed method is designed to explore the discriminative information hidden in geodesic distances. The main contributions of the proposed method are: 1) we propose a semi-supervised graph construction method which can effectively capture the multiple manifolds structure of the data, 2) each data point is replaced with an associated feature vector whose elements are the graph distances from it to the other data points. Information of the nonlinear structure is contained in the feature vectors which are helpful for classification, 3) we propose a new semi-supervised linear dimension reduction method for feature vectors which introduces the class information into the manifold learning process and establishes an explicit dimension reduction mapping. Experiments on benchmark data sets are conducted to show the effectiveness of the proposed method. Mingyu Fan, Xiaoqin Zhang 0002, Zhouchen Lin, Zhongfei Zhang, Hujun Bao |
ICDM | 2 |
| 2012 | Multiple sample group pairs' graph embedding for trackingabstractThis paper presents a new method which uses graph embedding and foreground-background patch pairs to perform object tracking. We first use particle filter to sample some particles. Then we evaluate each particle based on graph embedding and foreground-background patch pairs. For each particle, we use a two-layer model to represent the object, i.e. the inner layer (object layer) and the outer layer (background layer). Both the two layers are divided into patches. We cluster the foreground patches to several classes. Each class forms one sample group pair with the background patches. We perform graph embedding on multiple sample group pairs to discriminate the foreground and the background. Experimental results showed that our method tracked the objects efficiently. Weiming Hu 0004, Xiaoqin Zhang 0002 |
ICIP | 3 |
| 2012 | Single and Multiple Object Tracking Using Log-Euclidean Riemannian Subspace and Block-Division Appearance ModelabstractObject appearance modeling is crucial for tracking objects, especially in videos captured by nonstationary cameras and for reasoning about occlusions between multiple moving objects. Based on the log-euclidean Riemannian metric on symmetric positive definite matrices, we propose an incremental log-euclidean Riemannian subspace learning algorithm in which covariance matrices of image features are mapped into a vector space with the log-euclidean Riemannian metric. Based on the subspace learning algorithm, we develop a log-euclidean block-division appearance model which captures both the global and local spatial layout information about object appearances. Single object tracking and multi-object tracking with occlusion reasoning are then achieved by particle filtering-based Bayesian state inference. During tracking, incremental updating of the log-euclidean block-division appearance model captures changes in object appearance. For multi-object tracking, the appearance models of the objects can be updated even in the presence of occlusions. Experimental results demonstrate that the proposed tracking algorithm obtains more accurate results than six state-of-the-art tracking algorithms. Weiming Hu 0004, Xi Li 0001, Wenhan Luo, Xiaoqin Zhang 0002, Stephen J. Maybank, Zhongfei Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | Efficient block-division model for robust multiple object trackingabstractTracking multiple objects under occlusion is one of the most challenging issues in computer vision. Occlusion results in mistaken match when finding the most similar candidate. Adapting to the change of objects is essential for tracking as objects often undergo intrinsic changes, but noise is unavoidably introduced during updating of the object, and this further confuses the tracker. In order to address these problems, a block-division appearance model is introduced to efficiently handle occlusion. In this model, spatial information is introduced to avoid the mistaken match between object and candidate. Based on this model, a selective updating strategy is proposed to incrementally learn the change of the object, avoiding introducing noise when updating. At the same time occlusion is deduced by monitoring the variation of each block. Experimental results in various videos validate the effectiveness of our algorithm in tracking multiple objects under occlusion. Wenhan Luo, Xiaoqin Zhang 0002, Yang Liu 0020, Xi Li 0001, Weiming Hu 0004, Wei Li 0034 |
ICASSP | 2 |
| 2011 | Multi-cue based multi-target tracking using online random forestsabstractDiscriminative tracking has become popular tracking methods due to their descriptive power for foreground/background separation. Among these methods, online random forest is recently proposed and received a large amount of research attention due to its advantages such as efficiency and robust ness to noise, etc. However, the fact that only one kind of features is used limits the discriminative performance of this tracker. Additionally, the standard online forest tracker works only for a single target object. In this paper, we introduce a novel tracking method that integrates multiple cues capturing both geometric structures and edge-based shape information. Compared with the current online random forest based tracking algorithm, the proposed multi-cue tracker is more robust thanks to the complimentary information provided from these hybrid cues. Furthermore, the new tracker can track multiple targets as well as single target object. The effectiveness of the proposed tracker is validated using five public sequences. Xinchu Shi, Xiaoqin Zhang 0002, Yang Liu 0020, Weiming Hu 0004, Haibin Ling |
ICASSP | 2 |
| 2011 | RKOF: Robust Kernel-Based Local Outlier Detection
Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Ou Wu 0001 |
PAKDD (2) | 4 |
| 2011 | Incremental Tensor Subspace Learning and Its Applications to Foreground Segmentation and TrackingabstractAppearance modeling is very important for background modeling and object tracking. Subspace learning-based algorithms have been used to model the appearances of objects or scenes. Current vector subspace-based algorithms cannot effectively represent spatial correlations between pixel values. Current tensor subspace-based algorithms construct an offline representation of image ensembles, and current online tensor subspace learning algorithms cannot be applied to background modeling and object tracking. In this paper, we propose an online tensor subspace learning algorithm which models appearance changes by incrementally learning a tensor subspace representation through adaptively updating the sample mean and an eigenbasis for each unfolding matrix of the tensor. The proposed incremental tensor subspace learning algorithm is applied to foreground segmentation and object tracking for grayscale and color image sequences. The new background models capture the intrinsic spatiotemporal characteristics of scenes. The new tracking algorithm captures the appearance characteristics of an object during tracking and uses a particle filter to estimate the optimal object state. Experimental evaluations against state-of-the-art algorithms demonstrate the promise and effectiveness of the proposed incremental tensor subspace learning algorithm, and its applications to foreground segmentation and object tracking. Weiming Hu 0004, Xi Li 0001, Xiaoqin Zhang 0002, Xinchu Shi, Stephen J. Maybank, Zhongfei Zhang |
Int. J. Comput. Vis. | 3 |
| 2011 | Visual tracking via dynamic tensor analysis with mean update
Xiaoqin Zhang 0002, Xinchu Shi, Weiming Hu 0004, Xi Li 0001, Stephen J. Maybank |
Neurocomputing | 1 |
| 2011 | Adaptive learning codebook for action recognition
Yu Kong 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, Yunde Jia |
Pattern Recognit. Lett. | 2 |
| 2010 | Occlusion Handling with ℓ1-Regularized Sparse Reconstruction
Wei Li 0034, Bing Li 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, Hanzi Wang, Guan Luo |
ACCV (4) | 3 |
| 2010 | Use bin-ratio information for category and scene classificationabstractIn this paper we propose using bin-ratio information, which is collected from the ratios between bin values of histograms, for scene and category classification. To use such information, a new histogram dissimilarity, bin-ratio dissimilarity (BRD), is designed. We show that BRD provides several attractive advantages for category and scene classification tasks: First, BRD is robust to cluttering, partial occlusion and histogram normalization; Second, BRD captures rich co-occurrence information while enjoying a linear computational complexity; Third, BRD can be easily combined with other dissimilarity measures, such as L1and χ2, to gather complimentary information. We apply the proposed methods to category and scene classification tasks in the bag-of-words framework. The experiments are conducted on several widely tested datasets including PASCAL 2005, PASCAL 2008, Oxford flowers, and Scene-15 dataset. In all experiments, the proposed methods demonstrate excellent performance in comparison with previously reported solutions. Nianhua Xie, Haibin Ling, Weiming Hu 0004, Xiaoqin Zhang 0002 |
CVPR | 4 |
| 2010 | Compact visual codebook for action recognitionabstractVisual codebook has been popular in object classification as well as action analysis. However, its performance is often sensitive to the codebook size that is usually predefined. Moreover, the codebook generated by unsupervised methods, e.g., K-means, often suffers from the problem of ambiguity and weak efficiency. In other words, the visual codebook contains a lot of noisy and/or ambiguous words. In this paper, we propose a novel method to address these issues by constructing a compact but effective visual codebook using sparse reconstruction. Given a large codebook generated by K-means, we reformulate it in a sparse manner, and learn the weight of each word in the original visual codebook. Since the weights are sparse, they naturally introduce a new compact codebook. We apply this compact codebook to action recognition tasks and verify it on the widely used Weizmann action database. The experimental results show clearly the benefits of the proposed solution. Qingdi Wei, Xiaoqin Zhang 0002, Yu Kong 0001, Weiming Hu 0004, Haibin Ling |
ICIP | 2 |
| 2010 | Discriminative Level Set for Contour TrackingabstractConventional contour tracking algorithms with level set often use generative models to construct the energy function. For tracking through cluttered and noisy background, however, a generative model may not be discriminative enough. In this paper we integrate the discriminative methods into a level set framework when constructing the level set energy function. We train a set of weak classifiers to distinguish the object from the background. Each weak classifier is designed to select the most discriminative feature space and integrated via AdaBoost according to their training errors. We also introduce a novel interaction term to explore the correlation between pixels near the object edge. This term together with the discriminative model both enhance the discriminative power of the level set. The experimental results show that the contour tracked by our approach is more accurate than the conventional algorithms with the generative model. Our algorithm successfully tracks the object contour even in a cluttered environment. Wei Li 0034, Xiaoqin Zhang 0002, Weiming Hu 0004, Haibin Ling |
ICPR | 2 |
| 2010 | Video Scene Segmentation Using Time Constraint Dominant-Set Clustering
Xianglin Zeng, Xiaoqin Zhang 0002, Weiming Hu 0004, Wanqing Li 0001 |
MMM | 2 |
| 2010 | Multiple Object Tracking Via Species-Based Particle Swarm OptimizationabstractMultiple object tracking is particularly challenging when many objects with similar appearances occlude one another. Most existing approaches concatenate the states of different objects, view the multi-object tracking as a joint motion estimation problem and search for the best state of the joint motion in a rather high dimensional space. However, this centralized framework suffers from a high computational load. We bring a new view to the tracking problem from a swarm intelligence perspective. In analogy with the foraging behavior of bird flocks, we propose a species-based particle swarm optimization algorithm for multiple object tracking, in which the global swarm is divided into many species according to the number of objects, and each species searches for its object and maintains track of it. The interaction between different objects is modeled as species competition and repulsion, and the occlusion relationship is implicitly deduced from the “power” of each species, which is a function of the image observations. Therefore, our approach decentralizes the joint tracker to a set of individual trackers, each of which tries to maximize its visual evidence. Experimental results demonstrate the efficiency and effectiveness of our method. Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Learning Group Activity in Soccer Videos from Local Motion
Yu Kong 0001, Weiming Hu 0004, Xiaoqin Zhang 0002, Hanzi Wang, Yunde Jia |
ACCV (1) | 3 |
| 2009 | A Smarter Particle Filter
Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank |
ACCV (2) | 1 |
| 2009 | Efficient human pose estimation via parsing a tree structure based human modelabstractHuman pose estimation is the task of determining the states (location, orientation and scale) of each body part. It is important for many vision understanding applications, e.g. visual interactive gaming, immersive virtual reality, content-based image retrieval, etc. However, it remains a challenging task because of unknown image background, presence of clutter, partial occlusion and especially the high dimensional state space (usually 30+ dimensions). In this paper, we contribute to human pose estimation in two aspects. First, we design two efficient Markov Chain dynamics under the data-driven Markov Chain Monte Carlo (DDMCMC) framework to effectively explore the complex solution space. Second, we parse the tree structure state space into a lexicographic order according to the image observations and body topology, and the optimization process is conducted in this order. This realizes a much more efficient exploration than the sampling based search and exhaustive search, and thus achieves a tremendous speed-up. Experimental results demonstrate the efficiency and effectiveness of the proposed method in estimating various kinds of human poses, even with cluttered background , poor illumination or partial self-occlusion. Xiaoqin Zhang 0002, Xiaofeng Tong, Weiming Hu 0004, Stephen J. Maybank, Yimin Zhang 0002 |
ICCV | 1 |
| 2009 | Contour tracking with abrupt motionabstractTraditional contour tracking methods can not handle abrupt motion or low frame rate video. This is because the basis of the traditional tracking lies in the assumption that the motion is smooth between consecutive frames. However, the abrupt motion destroys the foundation of this assumption. In this paper, we integrate the stochastic search into the level set evolution to reinstitute the continuity. Our approach can be viewed as a two-layer hierarchical level set-based tracking framework in which Particle Swarm Optimization (PSO) and level set evolution are fused seamlessly. In the first layer, the PSO is adopted to capture the global motion of the object. The coarse contour is obtained by applying the global motion to the contour in the previous frame. For the second layer, the level set evolution based on the coarse contour is carried out to track the local deformation, which results in the actual contour. The promising experimental results for numerous real videos reveal the effectiveness of our approach. Wei Li 0034, Xiaoqin Zhang 0002, Weiming Hu 0004 |
ICIP | 2 |
| 2009 | Adaptive Distributed Intrusion Detection Using Parametric ModelabstractDue to the increasing demands for network security, distributed intrusion detection has become a hot research topic in computer science. However, the design and maintenance of the intrusion detection system (IDS) is still a challenging task due to its dynamic, scalability, and privacy properties. In this paper, we propose a distributed IDS framework which consists of the individual and global models. Specifically, the individual model for the local unit derives from Gaussian Mixture Model based on online Adaboost algorithm, while the global model is constructed through the PSO-SVM fusion algorithm. Experimental results demonstrate that our approach can achieve a good detection performance while being trained online and consuming little traffic to communicate between local units. Weiming Hu 0004, Xiaoqin Zhang 0002, Xi Li 0001 |
Web Intelligence | 3 |
| 2008 | Trajectory-Based Video Retrieval Using Dirichlet Process Mixture ModelsabstractIn this paper, we present a trajectory-based video retrieval framework using Dirichlet process mixture models. The main contribution of this framework is four-fold. (1) We apply a Dirichlet process mixture model (DPMM) to unsupervised trajectory learning. DPMM is a countably infinite mixture model with its components growing by itself. (2) We employ a time-sensitive Dirichlet process mixture model (tDPMM) to learn trajectories ’ time-series characteristics. Furthermore, a novel likelihood estimation algorithm for tDPMM is proposed for the first time. (3) We develop a tDPMM-based probabilistic model matching scheme, which is empirically shown to be more error-tolerating and is able to deliver higher retrieval accuracy than the peer methods in the literature. (4) The framework has a nice scalability and adaptability in the sense that when new cluster data are presented, the framework automatically identifies the new cluster information without having to redo the training. Theoretic analysis and experimental evaluations against the state-of-the-art methods demonstrate the promise and effectiveness of the framework. 1 Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Guan Luo |
BMVC | 4 |
| 2008 | Visual tracking via incremental Log-Euclidean Riemannian subspace learningabstractRecently, a novel Log-Euclidean Riemannian metric is proposed for statistics on symmetric positive definite (SPD) matrices. Under this metric, distances and Riemannian means take a much simpler form than the widely used affine-invariant Riemannian metric. Based on the Log-Euclidean Riemannian metric, we develop a tracking framework in this paper. In the framework, the covariance matrices of image features in the five modes are used to represent object appearance. Since a nonsingular covariance matrix is a SPD matrix lying on a connected Riemannian manifold, the Log-Euclidean Riemannian metric is used for statistics on the covariance matrices of image features. Further, we present an effective online Log-Euclidean Riemannian subspace learning algorithm which models the appearance changes of an object by incrementally learning a low-order Log-Euclidean eigenspace representation through adaptively updating the sample mean and eigenbasis. Tracking is then led by the Bayesian state inference framework in which a particle filter is used for propagating sample distributions over the time. Theoretic analysis and experimental evaluations demonstrate the promise and effectiveness of the proposed framework. Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Mingliang Zhu, Jian Cheng 0002 |
CVPR | 4 |
| 2008 | Sequential particle swarm optimization for visual trackingabstractVisual tracking usually involves an optimization process for estimating the motion of an object from measured images in a video sequence. In this paper, a new evolutionary approach, PSO (particle swarm optimization), is adopted for visual tracking. Since the tracking process is a dynamic optimization problem which is simultaneously influenced by the object state and the time, we propose a sequential particle swarm optimization framework by incorporating the temporal continuity information into the traditional PSO algorithm. In addition, the parameters in PSO are changed adaptively according to the fitness values of particles and the predicted motion of the tracked object, leading to a favourable performance in tracking applications. Furthermore, we show theoretically that, in a Bayesian inference view, the sequential PSO framework is in essence a multilayer importance sampling based particle filter. Experimental results demonstrate that, compared with the state-of-the-art particle filter and its variation - the unscented particle filter, the proposed tracking algorithm is more robust and effective, especially when the object has an arbitrary motion or undergoes large appearance changes. Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank, Xi Li 0001, Mingliang Zhu |
CVPR | 1 |
| 2008 | Robust Visual Tracking Based on an Effective Appearance Model
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002 |
ECCV (4) | 4 |
| 2008 | Key-frame extraction using dominant-set clusteringabstractKey frames play an important role in video abstraction. Clustering is a popular approach for key-frame extraction. In this paper, we propose a novel method for key-frame extraction based on dominant-set clustering. Compared with the existing clustering-based methods, the proposed method dynamically decides the number of key frames depending on the complexity of video shots, produces key frames in a progressive manner and requires less computation. Experimental results on different types of video shots have verified the effectiveness of the method. Xianglin Zeng, Weiming Hu 0004, Wanqing Li 0001, Xiaoqin Zhang 0002 |
ICME | 4 |
| 2008 | Group action recognition in soccer videosabstractGroup action recognition in soccer videos is a challenging problem due to the difficulties of group action representation and camera motion estimation. This paper presents a novel approach for recognizing group action with a moving camera. In our approach, ego-motion is estimated by the Kanade-Lucas-Tomasi feature sets on successive frames. The optical flow is then computed on compensated frames. Due to the inaccurate ego-motion estimation, the optical flow can not reflect accurate motion of objects. In this paper, we propose a new motion descriptor which treats the optical flow as spatial patterns and extracts accurate global motion from the noisy optical flow. The latent-dynamic conditional random field model is employed to recognize group action. Experimental results show that our approach is promising. Yu Kong 0001, Xiaoqin Zhang 0002, Qingdi Wei, Weiming Hu 0004, Yunde Jia |
ICPR | 2 |
| 2008 | Boosted cannabis image recognitionabstractWith the large number of Web sites promoting the use of illicit drugs, it has become important to screen these sites for the protection of children on the Internet. Conventional keyword-based approaches are not sufficient because these Web sites often have lots of images and little meaningful words than prices. We propose an AdaBoost-based algorithm for cannabis image recognition. This is the first known attempt at computerized detection of illicit drug Web contents using images. The main technical contributions of our work are two-fold. First, we introduce a novel weak classifier which considers the inherently structural property or ldquoself-similarityrdquo of the cannabis plants. The self-correlation structural characteristics of cannabis can be used as a discriminative property for the purpose of cannabis image recognition. Second, we propose a rapid weak classifier finder, which can efficiently select discriminative weak classifiers from the weak classifier space with little degradation to the classification accuracy. Experiments on real world images have demonstrated improved performance of our method over other methods. Nianhua Xie, Xi Li 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, James Z. Wang 0001 |
ICPR | 3 |
| 2008 | SVD based Kalman particle filter for robust visual trackingabstractObject tracking is one of the most important tasks in computer vision. The unscented particle filter algorithm has been extensively used to tackle this problem and achieved a great success, because it uses the UKF (unscented Kalman filter) to generate a sophisticated proposal distributions which incorporates the newest observations into the state transition distribution and thus overcomes the sample impoverishment problem suffered by the particle filter. However, UKF often encounters the ill-conditioned problem when solving the square root of the covariance matrix in practice. In this paper, we propose a novel Kalman particle filter based on SVD (singular value decomposition), and apply it for visual tracking. Experimental results demonstrate that, compared with the particle filter and the unscented particle filter, the proposed algorithm is more robust in tracking performance. Xiaoqin Zhang 0002, Weiming Hu 0004, Zixiang Zhao, Yanguo Wang, Xi Li 0001, Qingdi Wei |
ICPR | 1 |
| 2008 | User oriented link function classificationabstractCurrently most link-related applications treat all links in the same web page to be identical. One link-related application usually requires one certain property of hyperlinks but actually not all links have this property or they have this property on different levels. Based on a study of how human users judge the links, the idea of the link function classification (LFC) is introduced in this paper. The link functions reflect the purpose that links are created by web page designers and the way they are used by viewers. Links in a certain function class imply one certain relationship between the adjacent pages, and thus they can be assumed to have similar properties. An algorithm is proposed to analyze the link functions based on both vision and structure features which simulates the reaction on the links of human users. Current applications can be enhanced by LFC with a more accurate modeling of the web graph. New mining methods can be also developed by making more and stronger assumptions on links within each function class due to the purer property set they share. Mingliang Zhu, Weiming Hu 0004, Ou Wu 0001, Xi Li 0001, Xiaoqin Zhang 0002 |
WWW | 5 |
| 2007 | Kernel-Bayesian Framework for Object Tracking
Xiaoqin Zhang 0002, Weiming Hu 0004, Guan Luo, Stephen J. Maybank |
ACCV (1) | 1 |
| 2007 | Robust Visual Tracking Based on Incremental Tensor Subspace LearningabstractMost existing subspace analysis-based tracking algorithms utilize a flattened vector to represent a target, resulting in a high dimensional data learning problem. Recently, subspace analysis is incorporated into the multilinear framework which offline constructs a representation of image ensembles using high-order tensors. This reduces spatio-temporal redundancies substantially, whereas the computational and memory cost is high. In this paper, we present an effective online tensor subspace learning algorithm which models the appearance changes of a target by incrementally learning a low-order tensor eigenspace representation through adaptively updating the sample mean and eigenbasis. Tracking then is led by the state inference within the framework in which a particle filter is used for propagating sample distributions over the time. A novel likelihood function, based on the tensor reconstruction error norm, is developed to measure the similarity between the test image and the learned tensor subspace model during the tracking. Theoretic analysis and experimental evaluations against a state-of-the-art method demonstrate the promise and effectiveness of this algorithm. Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Guan Luo |
ICCV | 4 |
| 2007 | Graph Based Discriminative Learning for Robust and Efficient Object TrackingabstractObject tracking is viewed as a two-class 'one-versus-rest' classification problem, in which the sample distribution of the target is approximately Gausian while the background samples are often multimodal. Based on these special properties, we propose a graph embedding based discriminative learning method, in which the topology structures of graphs are carefully designed to reflect the properties of the sample distributions. This method can simultaneously learn the subspace of the target and its local discriminative structure against the background. Moreover, a heuristic negative sample selection scheme is adopted to make the classification more effective. In tracking procedure, the graph based learning is embedded into a Bayesian inference framework cascaded with hierarchical motion estimation, which significantly improves the accuracy and efficiency of the localization. Furthermore, an incremental updating technique for the graphs is developed to capture the changes in both appearance and illumination. Experimental results demonstrate that, compared with two state-of-the-art methods, the proposed tracking algorithm is more efficient and effective, especially in dynamically changing and clutter scenes. Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank, Xi Li 0001 |
ICCV | 1 |
| 2007 | Dominant Sets-Based Action Recognition using Image Sequence MatchingabstractAction recognition is one of the most active research fields in computer vision. In this paper, we propose a novel method for classifying human actions in a series of image sequences containing certain actions. Human action in image sequences can be recognized by a time-varying contour of human body. We first extract shape context of each contour to form the feature space. Then the dominant sets approach is used for feature clustering and classification to obtain the labeled sequences. Finally, we use a smoothing algorithm upon the labeled sequences to recognize human actions. The proposed dominant sets-based approach has been tested in comparison to three classical methods: K-means, mean shift, and fuzzy-C-mean. Experimental results demonstrate that the dominant sets-based approach achieves the best recognition performance. Moreover, our method is robust to non-rigid deformations, significant scale changes, high action irregularities, and low quality video. Qingdi Wei, Weiming Hu 0004, Xiaoqin Zhang 0002, Guan Luo |
ICIP (6) | 3 |
| 2007 | A Robust Multiple Cues Fusion based Bayesian TrackerabstractThis paper presents an efficient and robust tracking algorithm based on multiple cues fusion in the Bayesian framework. This method characterizes the object to be tracked using a MOG (mixture of Gaussians) based appearance model and a chamfer-matching based shape model. A selective updating technique for the models is employed to accommodate for appearance and illumination changes. Meantime, the mean shift algorithm is embedded as the prior information into the Bayesian framework to give a heuristic prediction in the hypotheses generation process, which also alleviates the great computational load suffered by the conventional Bayesian tracker. Experimental results demonstrate that, compared with some existing works, the proposed algorithm has a better adaptability to changes of the object as well as the environments. Xiaoqin Zhang 0002, Zhiyong Liu 0001, Hong Qiao |
ICRA | 1 |
| 2006 | Multi-Information Fusion for Scale Selection in Robot TrackingabstractMean shift, for its simplicity and efficiency, has achieved a considerable success in robot tracking. For the mean shift based tracking algorithm, the scale of the mean-shift kernel bandwidth is a crucial parameter which reflects the size of tracking window. However, in literature how to properly update or select the bandwidth remains a tough task as the size of the object under consideration changes. In this paper, a weighted average integral projection approach is proposed to extract the local information of the object, and then a multiinformation fusion strategy is suggested for the scale selection, which combines both the global and local information of the sample weight image. Moreover, a coarse-to-fine approximate approach is employed to accelerate the procedure. Experimental results demonstrate that, compared to some existing works, the strategy proposed has a better adaptability as the size of the object changes in clutter environments. Xiaoqin Zhang 0002, Hong Qiao, Zhiyong Liu 0001 |
IROS | 1 |