Xudong Jiang 0001

dblp:11/2494 · DBLP profile ↗
← Back
218ranked-venue papers
18as first author
105since 2021 · last 2026
0000-0002-9104-2315ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 140 · 7 first-author · 57 since 2021Artificial intelligence and machine learning · 101 · 9 first-author · 63 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 7 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
abstract
Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitigation methods effectively reduce hallucinations in photographic images, they largely overlook the potential risks posed by stylized images, which play crucial roles in critical scenarios such as game scene understanding, art education, and medical analysis. In this work, we first construct a dataset comprising photographic images and their corresponding stylized versions with carefully annotated caption labels. We then conduct head-to-head comparisons on both discriminative and generative tasks by benchmarking 13 advanced LVLMs on the collected datasets. Our findings reveal that stylized images tend to induce significantly more hallucinations than their photographic counterparts. To address this issue, we propose Style-Aware Visual Early Revision (SAVER), a novel mechanism that dynamically adjusts LVLMs' final outputs based on the token-level visual attention patterns, leveraging early-layer feedback to mitigate hallucinations caused by stylized images. Extensive experiments demonstrate that SAVER achieves state-of-the-art performance in hallucination mitigation across various models, datasets, and tasks.
Zhaoxu Li, Chenqi Kong, Yi Yu 0011, Qiangqiang Wu, Xinghao Jiang, Ngai-Man Cheung, Bihan Wen, Alex Chichung Kot, Xudong Jiang 0001
AAAI9
2026 From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge
abstract
Large-scale Video Foundation Models (VFMs) have significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs). However, the open accessibility of VFMs also introduces critical security risks, as adversaries can exploit full knowledge of the VFMs to launch potent attacks. This paper investigates a novel and practical adversarial threat scenario: attacking downstream models or MLLMs fine-tuned from open-source VFMs, without requiring access to the victim task, training data, model query, and architecture. In contrast to conventional transfer-based attacks that rely on task-aligned surrogate models, we demonstrate that adversarial vulnerabilities can be exploited directly from the VFMs. To this end, we propose the Transferable Video Attack (TVA), a temporal-aware adversarial attack method that leverages the temporal representation dynamics of VFMs to craft effective perturbations. TVA integrates a bidirectional contrastive learning mechanism to maximize the discrepancy between the clean and adversarial features, and introduces a temporal consistency loss that exploits motion cues to enhance the sequential impact of perturbations. TVA avoids the need to train expensive surrogate models or access to domain-specific data, thereby offering a more practical and efficient attack strategy. Extensive experiments across 24 video-related tasks demonstrate the efficacy of TVA against downstream models and MLLMs, revealing a previously underexplored security vulnerability in the deployment of video models.
Yi Yu 0011, Song Xia, Deepu Rajan, Boon Poh Ng, Alex Chichung Kot, Xudong Jiang 0001
AAAI8
2026 SLAN: A state-space linear attention network with meta guidance and chunk-wise fusion for long-term time series forecasting
Jihong Guan, Xudong Jiang 0001, Mingshan Loo, Hanchen Yang 0002, Wengen Li, Yichao Zhang 0001, Shuigeng Zhou
Expert Syst. Appl.2
2026 PiFormer: Towards Subseasonal SST Prediction with Spatial-Patched Inverted Transformer
Hanchen Yang 0002, Wengen Li, Xudong Jiang 0001, Jihong Guan, Yichao Zhang 0001, Shuigeng Zhou
Expert Syst. Appl.4
2026 GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation
Henghui Ding, Chang Liu 0072, Shuting He, Xudong Jiang 0001, Yu-Gang Jiang 0001
Int. J. Comput. Vis.4
2026 Curvilinear structure-preserving unpaired cross-domain medical image translation
Yi Zhou 0024, Xudong Jiang 0001, Li Chen 0011, Leopold Schmetterer, Bingyao Tan, Jun Cheng 0003
Neurocomputing3
2026 Adaptive memory refinement and perception enhancement for exo-to-ego video generation
Weipeng Hu, Jiun Tian Hoe, Ping Hu 0001, Xudong Jiang 0001, Yap-Peng Tan
Neurocomputing6
2026 Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
abstract
Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks. However, effectively integrating image segmentation into these models remains a significant challenge.In this work, we propose a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process. Our key innovation is semantic descriptors, a new textual representation of segmentation masks where each image patch is mapped to its corresponding text label. We first introduce image-wise semantic descriptors, a patch-aligned textual representation of segmentation masks that integrates naturally into the language modeling pipeline. To enhance efficiency, we introduce the Row-wise Run-Length Encoding (R-RLE), which compresses redundant text sequences, reducing the length of semantic descriptorsby 74% and accelerating inference by $3\times$3×, without compromising performance. Building upon this, our initial framework Text4Segachieves strong segmentation performance across a wide range of vision tasks. To further improve granularity and compactness, we propose box-wise semantic descriptors, which localizes regions of interest using bounding boxes and represents region masks via structured mask tokens called semantic bricks. This leads to our refined model, Text4Seg++, which formulates segmentation as a next-brick prediction task, combining precision, scalability, and generative efficiency. Comprehensive experiments on natural and remote sensing datasets show that Text4Seg++consistently outperforms state-of-the-art models across diverse benchmarks without any task-specific fine-tuning, while remaining compatible with existing MLLM backbones. Our work highlights the effectiveness, scalability, and generalizability of text-driven image segmentation within the MLLM framework.
Mengcheng Lan, Chaofeng Chen, Jiaxing Xu, Zongrui Li 0001, Yiping Ke, Xudong Jiang 0001, Yingchen Yu, Yunqing Zhao, Song Bai 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Concave Cut: Analyzing the role of concave functions in clustering
Shenfei Pei, Yuanchen Sun, Zhongqi Lin, Feiping Nie 0001, Jitao Lu, Xudong Jiang 0001, Canyu Zhang 0001, Zengwei Zheng
Pattern Recognit.6
2026 Reasoning step by step via a neural-symbolic geometry problem solver
Yaxian Wang, Bifan Wei, Yinghong Ma, Xudong Jiang 0001, Henghui Ding, Zhongmin Cai, Jun Liu 0002
Pattern Recognit.4
2026 Transferable Adversarial Attack on Referring Video Object Segmentation
abstract
Referring video object segmentation (RVOS) is an emerging task that aims to segment the text-referred objects in the given video sequence. This capability plays a critical role in some real-world safety-critical applications such as autonomous driving. However, advanced RVOS models predominantly leverage deep neural networks that are inherently vulnerable to adversarial perturbations, which raises serious safety concerns. Although some studies have explored adversarial attacks on video object segmentation (VOS), the robustness and security of RVOS models against such attacks remain insufficiently investigated. This work thus, for the first time, comprehensively investigates the adversarial robustness of RVOS models. Distinct from other VOS tasks, RVOS is more challenging due to its multi-modal nature and high dependence on spatial-temporal information. Considering that, we propose a cross-prompt Multimodal attack with Inter-Clip Momentum (xM-ICM) to effectively mislead RVOS models under both white-box and black-box scenarios. The proposed xM jointly corrupts visual and textual embeddings and integrates a cross-prompt strategy during iterative optimization to enhance generalization across diverse linguistic queries. The ICM module harnesses the spatial-temporal dependencies across sequence clips via two momentum banks to preserve the perturbation coherence throughout the whole video and stabilize the adversarial optimization. Experimental results on three benchmarks and five prevalent RVOS models demonstrate the superior white-box attack performance and strong black-box transferability of our proposed method.
Meiwen Ding, Song Xia, Yi Yu 0011, Shuting He, Xudong Jiang 0001
IEEE Trans. Inf. Forensics Secur.5
2026 LP2DH: A Locality-Preserving Pixel-Difference Hashing Framework for Dynamic Texture Recognition
abstract
Spatiotemporal Local Binary Pattern (STLBP) is a widely used dynamic texture descriptor, but it suffers from extremely high dimensionality. To tackle this, STLBP features are often extracted on three orthogonal planes, which sacrifice inter-plane correlation. In this work, we propose a Locality-Preserving Pixel-Difference Hashing (LP2DH) framework that jointly encodes pixel differences in the full spatiotemporal neighborhood. LP2DH transforms Pixel-Difference Vectors (PDVs) into compact binary codes with maximal discriminative power. Furthermore, we incorporate a locality-preserving embedding to maintain the PDVs' local structure before and after hashing. Then, a curvilinear search strategy is utilized to jointly optimize the hashing matrix and binary codes via gradient descent on the Stiefel manifold. After hashing, dictionary learning is applied to encode the binary vectors into codewords, and the resulting histogram is utilized as the final feature representation. The proposed LP2DH achieves state-of-the-art performance on three major dynamic texture recognition benchmarks: 99.80% against DT-GoogleNet's 98.93% on UCLA, 98.52% against HoGF3D's 97.63% on DynTex++, and 96.19% compared to STS's 95.00% on YUPENN. The source code is available at: https://github.com/drx770/LP2DH.
Ruxin Ding, Jianfeng Ren, Heng Yu 0001, Jiawei Li 0001, Xudong Jiang 0001
IEEE Trans. Image Process.5
2026 Identity-Compensated Style Distillation for Visible-Infrared Person Re-Identification
abstract
Visible-Infrared Person Re-Identification (VI-ReID) that matches pedestrian images across visible and infrared modalities suffers from substantial modality discrepancies and intra-class variations. While existing methods typically address the modality gap via style alignment, they often lose identity-relevant semantics and overlook fine-grained inter-class nuances, such as body part contours and structural cues around the head, shoulders, or feet. To tackle these challenges, we propose an Identity-Compensated Style Distillation (ICSD) network that enforces cross-modality style consistency and enhances the discriminative power of modality-invariant features. Specifically, ICSD comprises two core components: (1) a Style Knowledge Distillation (SKD) module, which integrates Style Discrepancy Reduction (SDR) and Identity Knowledge Compensation (IKC) to align modality styles while preserving identity-relevant semantics; (2) an Identity Discrimination Amplification (IDA) module, which captures and enhances subtle inter-class differences by refining identity-specific cues, thereby facilitating more accurate discrimination between different pedestrians. Extensive experiments on three public benchmarks-SYSU-MM01, RegDB, and LLCM-demonstrate that ICSD consistently outperforms state-of-the-art methods, validating the effectiveness and complementarity of its components.
Yongguo Ling, Zihao Hu, Nan Pu, Zhun Zhong, Xudong Jiang 0001
IEEE Trans. Image Process.5
2026 A Greedy Strategy for Graph Cut
abstract
We propose a novel Greedy Graph Cut (GGC) algorithm to address the graph partitioning problem. The algorithm begins by treating each data point as an individual cluster and iteratively merges cluster pairs that maximize the reduction in the global objective function until the desired number of clusters is achieved. We provide a theoretical proof of the monotonic convergence of the objective function values throughout this process. To improve computational efficiency, the algorithm restricts merging operations to adjacent clusters, resulting in a computational complexity that scales nearly linearly with the sample size. A significant advantage of our greedy approach is its deterministic nature, which ensures consistent results across multiple runs. This stands in contrast to many existing algorithms that are sensitive to random initialization effects. We demonstrate the effectiveness of the proposed algorithm by applying it to the Normalized Cut (N-Cut) problem, a well-studied variant of graph partitioning. Extensive experimental results show that GGC consistently outperforms the conventional two-stage optimization approach-which involves eigendecomposition followed by k-means clustering-in solving the N-Cut problem. Furthermore, comparative analyses reveal that GGC achieves superior performance compared to several state-of-the-art clustering algorithms.
Shenfei Pei, Huijuan Dong, Nianci Guan, Zhongqi Lin, Feiping Nie 0001, Xudong Jiang 0001, Zengwei Zheng
IEEE Trans. Image Process.6
2026 Open-Set Anomaly Segmentation in Complex Scenarios
abstract
Precise segmentation of out-of-distribution (OoD) objects, herein referred to as anomalies, is crucial for the reliable deployment of semantic segmentation models in open-set, safety-critical applications, such as autonomous driving. Current anomalous segmentation benchmarks predominantly focus on favorable weather conditions, resulting in untrustworthy evaluations that overlook the risks posed by diverse meteorological conditions in open-set environments, such as low illumination, dense fog, and heavy rain. To bridge this gap, this paper introduces the ComSAmy, a Complex Scenarios Anomaly segmentation benchmark. ComSAmy encompasses a wide spectrum of adverse weather conditions, dynamic driving environments, and diverse anomaly types to comprehensively evaluate the model performance in realistic open-world scenarios. Our extensive evaluation of several state-of-the-art anomalous segmentation models reveals that existing methods demonstrate significant deficiencies in such challenging scenarios, highlighting their serious safety risks for real-world deployment. To solve that, we propose a novel energy-entropy learning (EEL) strategy that integrates the complementary information from energy and entropy to bolster the robustness of anomaly segmentation under complex open-world environments. Additionally, a diffusion-based anomalous training data synthesizer is proposed to generate diverse and high-quality anomalous images to enhance the existing copy-paste training data synthesizer. Extensive experimental results on both public and ComSAmy benchmarks demonstrate that our proposed diffusion-based synthesizer with energy and entropy learning (DiffEEL) framework serves as an effective and generalizable plug-and-play method to enhance existing models, yielding an average improvement of around 4.96% in AUPRC and 9.87% in $\rm {FPR}_{95}$ .
Song Xia, Yi Yu 0011, Henghui Ding, Wenhan Yang, Shifei Liu, Alex Chichung Kot, Xudong Jiang 0001
IEEE Trans. Image Process.7
2026 Predictive Reasoning With Augmented Anomaly Contrastive Learning for Compositional Visual Relations
abstract
While visual reasoning for simple analogies has received significant attention, compositional visual relations (CVR) remain relatively unexplored due to their greater complexity. To solve CVR tasks, we propose Predictive Reasoning with Augmented Anomaly Contrastive Learning (PR-A$^{2}$CL), i.e., to identify an outlier image given three other images that follow the same compositional rules. To address the challenge of modelling abundant compositional rules, an Augmented Anomaly Contrastive Learning is designed to distil discriminative and generalizable features by maximizing similarity among normal instances while minimizing similarity between normal and anomalous outliers. More importantly, a predict-and-verify paradigm is introduced for rule-based reasoning, in which a series of Predictive Anomaly Reasoning Blocks (PARBs) iteratively leverage features from three out of the four images to predict those of the remaining one. Throughout the subsequent verification stage, the PARBs progressively pinpoint the specific discrepancies attributable to the underlying rules. Experimental results on SVRT, CVR and MC$^{2}$R datasets show that PR-A$^{2}$CL significantly outperforms state-of-the-art reasoning models.
Chengtai Li, Yuting He 0002, Jianfeng Ren, Ruibin Bai, Yitian Zhao, Heng Yu 0001, Xudong Jiang 0001
IEEE Trans. Multim.7
2026 GeoTree: A Dynamic Tree-Based Geometry Problem Solver Through LLM-Symbolic Reasoning
abstract
Geometry problem solving (GPS) requires high-level symbolic and logical reasoning based on geometry theorem knowledge to arrive at the answer. Despite the remarkable advances achieved by Large Language Models (LLMs) in various problem-solving tasks, they still struggle to perform rigorous multi-step geometry reasoning, which is essential for GPS. In this paper, we propose a dynamic tree-based geometry problem solver named GeoTree, which combines a knowledgeable LLM with a rigorous symbolic solver to perform geometry reasoning cooperatively. Specifically, an iterative multi-step geometry reasoning process is performed dynamically based on a tree-like structure, thereby emulating divergent and deliberate human problem-solving thinking. Each geometry reasoning step is completed collaboratively through four components, consisting ofTheorem Seeker,Symbolic Solver,Evaluator, andController. First,Theorem Seekerprompts LLMs to seek out candidate theorems with their inherent geometry theorem knowledge. Subsequently,Symbolic Solverapplies the theorems on the known conditions to obtain new additional conditions. Then,Evaluatorassesses the availability of the theorems and prompts LLMs to judge the usefulness of these new conditions for the problem target, which serves as the heuristic guidance for subsequent reasoning. Finally,Controllerdetermines the termination state, which decides whether to continue invoking the other three components for further attempts. Extensive experiments on Geometry3K demonstrate the superiority of GeoTree in accuracy, efficiency, and explainability.
Yaxian Wang, Bifan Wei, Yinghong Ma, Lingling Zhang 0005, Xudong Jiang 0001, Henghui Ding, Jun Liu 0002
IEEE Trans. Multim.5
2025 EvHDR-GS: Event-guided HDR Video Reconstruction with 3D Gaussian Splatting
abstract
High Dynamic Range (HDR) video reconstruction seeks to accurately restore the extensive dynamic range present in real-world scenes and is widely employed in downstream applications. Existing methods typically operate on one or a small number of consecutive frames, which often leads to inconsistent brightness across the video due to their limited perspective on the video sequence. Moreover, supervised learning-based approaches are susceptible to data bias, resulting in reduced effectiveness when confronted with test inputs exhibiting a domain gap relative to the training data. To address these limitations, we present an event-guided HDR video reconstruction method through building 3D Gaussian Splatting (3DGS), to ensure consistent brightness imposed by 3D consistency. We introduce HDR 3D Gaussians capable of simultaneously representing HDR and low-dynamic-range (LDR) colors. Furthermore, we incorporate a learnable HDR-to-LDR transformation optimized by input event streams and LDR frames to eliminate the data bias. Experimental results on both synthetic and real-world datasets demonstrate that the proposed method achieves state-of-the-art performance.
Zhan Lu, De Ma, Huajin Tang, Xudong Jiang 0001, Gang Pan 0001
AAAI5
2025 DARR: A Dual-Branch Arithmetic Regression Reasoning Framework for Solving Machine Number Reasoning
abstract
Abstract visual reasoning (AVR) is a critical ability of humans, and it has been widely studied, but arithmetic visual reasoning, a unique task in AVR to reason over number sense, is less studied in the literature. To facilitate this research, we construct a Machine Number Reasoning (MNR) dataset to assess the model's ability in arithmetic visual reasoning over number sense and spatial layouts. To solve the MNR tasks, we propose a Dual-branch Arithmetic Regression Reasoning (DARR) framework, which includes an Intra-Image Arithmetic Regression Reasoning (IIARR) module and a Cross-Image Arithmetic Regression Reasoning (CIARR) module. The IIARR includes a set of Intra-Image Regression Blocks to identify the correct number orders and the underlying arithmetic rules within individual images, and an Order Gate to determine the correct number order. The CIARR establishes the arithmetic relations across different images through a `3-to-1' regressor and a set of `2-to-1' regressors, with a Selection Gate to select the most suitable `2-to-1' regressor and a gated fusion to combine the two kinds of regressors. Experiments on the MNR dataset show that the DARR outperforms state-of-the-art models for arithmetic visual reasoning.
Chengtai Li, Yee Yang Tan, Yuting He 0002, Jianfeng Ren, Ruibin Bai, Yitian Zhao, Heng Yu 0001, Xudong Jiang 0001
AAAI8
2025 ERL-MPP: Evolutionary Reinforcement Learning with Multi-head Puzzle Perception for Solving Large-scale Jigsaw Puzzles of Eroded Gaps
abstract
Solving jigsaw puzzles has been extensively studied. While most existing models focus on solving either small-scale puzzles or puzzles with no gap between fragments, solving large-scale puzzles with gaps presents distinctive challenges in both image understanding and combinatorial optimization. To tackle these challenges, we propose a framework of Evolutionary Reinforcement Learning with Multi-head Puzzle Perception (ERL-MPP) to derive a better set of swapping actions for solving the puzzles. Specifically, to tackle the challenges of perceiving the puzzle with gaps, a Multi-head Puzzle Perception Network (MPPN) with a shared encoder is designed, where multiple puzzlet heads comprehensively perceive the local assembly status, and a discriminator head provides a global assessment of the puzzle. To explore the large swapping action space efficiently, an Evolutionary Reinforcement Learning (EvoRL) agent is designed, where an actor recommends a set of suitable swapping actions from a large action space based on the perceived puzzle status, a critic updates the actor using the estimated rewards and the puzzle status, and an evaluator coupled with evolutionary strategies evolves the actions aligning with the historical assembly experience. The proposed ERL-MPP is comprehensively evaluated on the JPLEG-5 dataset with large gaps and the MIT dataset with large-scale puzzles. It significantly outperforms all state-of-the-art models on both datasets.
Xingke Song, Chenglin Yao, Jianfeng Ren, Ruibin Bai, Xin Chen 0003, Xudong Jiang 0001
AAAI7
2025 Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension
abstract
In this work, we address the challenging task of Generalized Referring Expression Comprehension (GREC). Compared to the classic Referring Expression Comprehension (REC) that focuses on single-target expressions, GREC extends the scope to a more practical setting by further encompassing no-target and multi-target expressions. Existing REC methods face challenges in handling the complex cases encountered in GREC, primarily due to their fixed output and limitations in multi-modal representations. To address these issues, we propose a Hierarchical Alignment-enhanced Adaptive Grounding Network (HieA2G) for GREC, which can flexibly deal with various types of referring expressions. First, a Hierarchical Multi-modal Semantic Alignment (HMSA) module is proposed to incorporate three levels of alignments, including word-object, phrase-object, and text-image alignment. It enables hierarchical cross-modal interactions across multiple levels to achieve comprehensive and robust multi-modal understanding, greatly enhancing grounding ability for complex cases. Then, to address the varying number of target objects in GREC, we introduce an Adaptive Grounding Counter (AGC) to dynamically determine the number of output targets. Additionally, an auxiliary contrastive loss is employed in AGC to enhance object-counting ability by pulling in multi-modal features with the same counting and pushing away those with different counting. Extensive experimental results show that HieA2G achieves new state-of-the-art performance on the challenging GREC task and also the other 4 tasks, including REC, Phrase Grounding, Referring Expression Segmentation (RES), and Generalized Referring Expression Segmentation (GRES), demonstrating the remarkable superiority and generalizability of the proposed HieA2G.
Yaxian Wang, Henghui Ding, Shuting He, Xudong Jiang 0001, Bifan Wei, Jun Liu 0002
AAAI4
2025 Exploiting Temporal State Space Sharing for Video Semantic Segmentation
abstract
Video semantic segmentation (VSS) plays a vital role in understanding the temporal evolution of scenes. Traditional methods often segment videos frame-by-frame or in a short temporal window, leading to limited temporal context, redundant computations, and heavy memory requirements. To this end, we introduce a Temporal Video State Space Sharing (TV3S) architecture to leverage Mamba state space models for temporal feature sharing. Our model features a selective gating mechanism that efficiently propagates relevant information across video frames, eliminating the need for a memory-heavy feature pool. By processing spatial patches independently and incorporating shifted operation, TV3S supports highly parallel computation in both training and inference stages, which reduces the delay in sequential state space processing and improves the scalability for long video sequences. Moreover, TV3S incorporates information from prior frames during inference, achieving long-range temporal coherence and superior adaptability to extended sequences. Evaluations on the VSPW and Cityscapes datasets reveal that our approach outperforms current state-of-the-art methods, establishing a new standard for VSS with consistent results across long video sequences. By achieving a good balance between accuracy and efficiency, TV3S shows a significant advancement in spatiotemporal modeling, paving the way for efficient video analysis. The code is publicly available at https://github.com/Ashesham/TV3S.git.
Syed Ariff Syed Hesham, Yun Liu 0011, Guolei Sun, Henghui Ding, Ender Konukoglu, Xue Geng, Xudong Jiang 0001
CVPR8
2025 Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference Systems
abstract
By locally encoding raw data into intermediate features, collaborative inference enables end users to leverage powerful deep learning models without exposure of sensitive raw data to cloud servers. However, recent studies have revealed that these intermediate features may not sufficiently preserve privacy, as information can be leaked and raw data can be reconstructed via model inversion attacks (MIAs). Obfuscation-based methods, such as noise corruption, adversarial representation learning, and information filters, enhance the inversion robustness by obfuscating the task-irrelevant redundancy empirically. However, methods for quantifying such redundancy remain elusive, and the explicit mathematical relation between this redundancy minimization and inversion robustness enhancement has not yet been established. To address that, this work first theoretically proves that the conditional entropy of inputs given intermediate features provides a guaranteed lower bound on the reconstruction mean square error (MSE) under any MIA. Then, we derive a differentiable and solvable measure for bounding this conditional entropy based on the Gaussian mixture estimation and propose a conditional entropy maximization (CEM) algorithm to enhance the inversion robustness. Experimental results on four datasets demonstrate the effectiveness and adaptability of our proposed CEM; without compromising feature utility and computing efficiency, plugging the proposed CEM into obfuscation-based defense mechanisms consistently boosts their inversion robustness, achieving average gains ranging from 12.9% to 48.2%. Code is available at https://github.com/xiasong0501/CEM.
Song Xia, Yi Yu 0011, Wenhan Yang, Meiwen Ding, Zhuo Chen 0006, Ling-Yu Duan, Alex Chichung Kot, Xudong Jiang 0001
CVPR8
2025 DBCR: Exploiting Both Intra-cluster and Extra-cluster Relations for Compositional Reasoning
abstract
Most existing models for abstract visual reasoning perform poorly in compositional visual reasoning (CVR), due to complex nature of compositional rules and difficulties in distinguishing tiny rule differences between outliers and normal images. To tackle the challenges, we propose a Dual-Branch Compositional Reasoning (DBCR) model, exploiting both intra-cluster relations among the cluster of normal images and extra-cluster relations between normal images and outliers. Specifically, we design one branch of Intra-Cluster Regression Reasoning Blocks (ICR2Bs) to encapsulate common relations among normal images through hierarchical regressing reasoning, and the other branch of Contrastive Attention Reasoning Blocks (CARBs) to exploit extra-cluster differences between normal images and outliers through self-attention. Simultaneously minimizing the regression errors in ICR2Bs and maximizing the extra-cluster differences in CARBs help identify the correct cluster of normal images. Experimental results on two CVR datasets show that the proposed DBCR consistently outperforms state-of-the-art models. The code is available at https://github.com/He1mont/DBCR.
Chengtai Li, Guosheng Su, Jianfeng Ren, Ruibin Bai, Yitian Zhao, Xudong Jiang 0001
ICASSP6
2025 Jointly Optimizing Data Discretization and Naive Bayes Classifier via Multi-Objective Optimization
abstract
Data discretization plays a critical role in enhancing the performance of the naive Bayes classifier. Traditional data discretization methods often utilize a two-stage framework, where data discretization and classification are optimized separately, leading to sub-optimal performance. To tackle the issue, we propose a novel multi-objective optimization framework that incorporates the optimization of the naive Bayes classifier into the objective function of optimizing data discretization. To solve this problem, we employ an alternative optimization method to jointly optimize both data discretization and classification. Additionally, to further enhance the optimization process, we leverage a genetic algorithm to explore and exploit a larger solution space. Experimental results on 20 datasets demonstrate that our method outperforms state-of-the-art methods.
Jiacheng Tu, Haiyan Yu 0003, Ruxin Ding, Shihe Wang, Jianfeng Ren, Xudong Jiang 0001
ICASSP6
2025 GCA-SUNet: A Gated Context-Aware Swin-UNet for Exemplar-Free Counting
abstract
Exemplar-Free Counting aims to count objects of interest without intensive annotations of objects or exemplars. To achieve this, we propose a Gated Context-Aware Swin-UNet (GCA-SUNet) to directly map an input image to the density map of countable objects. Specifically, a set of Swin transformers form an encoder to derive a robust feature representation, and a Gated Context-Aware Modulation block is designed to suppress irrelevant objects or background through a gate mechanism and exploit the attentive support of objects of interest through a self-similarity matrix. The gate strategy is also incorporated into the bottleneck network and the decoder of the Swin-UNet to highlight the features most relevant to objects of interest. By explicitly exploiting the attentive support among countable objects and eliminating irrelevant features through the gate mechanisms, the proposed GCA-SUNet focuses on and counts objects of interest without relying on predefined categories or exemplars. Experimental results on the real-world datasets such as FSC-147 and CARPK demonstrate that GCA-SUNet significantly and consistently outperforms state-of-the-art methods. The code is available at https://github.com/Amordia/GCA-SUNet.
Yipeng Xu, Jialu Zhang 0003, Jianfeng Ren, Xudong Jiang 0001
ICME6
2025 CEARI: Co-Evolutionary Agents for Reassembling and Inpainting Puzzles with Gaps and Missing Pieces
abstract
Puzzle solving has recently become a popular research topic. Existing solvers often overlook puzzles with missing pieces. The missing pieces, together with gaps between pieces, pose significant challenges, amplified by a large solution space. To tackle the challenges, we propose Co-Evolutionary Agents for Reassembling and Inpainting (CEARI), one agent to inpaint missing contents and the other to reassemble the puzzle, with a shared perception network to perceive the puzzle status. The reassembly agent utilizes an evolutionary algorithm to explore the large solution space, to discover a sequence of fragment-swapping actions to efficiently reassemble the puzzle, while the inpainting agent evolves from using a local outpainting network at the early stage to using a global inpainting network at the latter stage. Furthermore, a co-evolutionary training paradigm is designed to iteratively evolve the two agents in a coherent and collaborative manner, improving reassembly accuracy and inpainting quality simultaneously. Experimental results on three datasets show that CEARI largely outperforms state-of-the-art methods in terms of both reassembly accuracy and inpainting quality.
Xingke Song, Jianxu Shangguan, Yiran Li 0003, Jialu Zhang 0003, Jianfeng Ren, Ruibin Bai, Xin Chen 0003, Xudong Jiang 0001
ACM Multimedia8
2025 DSRF: A Dynamic and Scalable Reasoning Framework for Solving RPMs
abstract
Abstract Visual Reasoning (AVR) entails discerning latent patterns in visual data and inferring underlying rules. Existing solutions often lack scalability and adaptability, as deep architectures tend to overfit training data, and static neural networks fail to dynamically capture diverse rules. To tackle the challenges, we propose a Dynamic and Scalable Reasoning Framework (DSRF) that greatly enhances the reasoning ability by widening the network instead of deepening it, and dynamically adjusting the reasoning network to better fit novel samples instead of a static network. Specifically, we design a Multi-View Reasoning Pyramid (MVRP) to capture complex rules through layered reasoning to focus features at each view on distinct combinations of attributes, widening the reasoning network to cover more attribute combinations analogous to complex reasoning rules. Additionally, we propose a Dynamic Domain-Contrast Prediction (DDCP) block to handle varying task-specific relationships dynamically by introducing a Gram matrix to model feature distributions, and a gate matrix to capture subtle domain differences between context and target features. Extensive experiments on six AVR tasks demonstrate DSRF’s superior performance, achieving state-of-the-art results under various settings. Code is available here: https://github.com/UNNCRoxLi/DSRF.
Chengtai Li, Yuting He 0002, Jianfeng Ren, Ruibin Bai, Yitian Zhao, Xudong Jiang 0001
NeurIPS6
2025 Multi-Scale Finetuning for Encoder-based Time Series Foundation Models
abstract
Time series foundation models (TSFMs) demonstrate impressive zero-shot performance for time series forecasting. However, an important yet underexplored challenge is how to effectively finetune TSFMs on specific downstream tasks. While naive finetuning can yield performance gains, we argue that it falls short of fully leveraging TSFMs' capabilities, often resulting in overfitting and suboptimal performance. Given the diverse temporal patterns across sampling scales and the inherent multi-scale forecasting capabilities of TSFMs, we adopt a causal perspective to analyze finetuning process, through which we highlight the critical importance of explicitly modeling multiple scales and reveal the shortcomings of naive approaches. Focusing on encoder-based TSFMs, we propose Multiscale finetuning (MSFT), a simple yet general framework that explicitly integrates multi-scale modeling into the finetuning process. Experimental results on three different backbones (Moirai, Moment and Units) demonstrate that TSFMs finetuned with MSFT not only outperform naive and typical parameter efficient finetuning methods but also surpass state-of-the-art deep learning methods. Codes are available at https://github.com/zqiao11/MSFT.
Zhongzheng Qiao, Ming Jin 0005, Quang Pham, Qingsong Wen, Ponnuthurai N. Suganthan, Xudong Jiang 0001, Savitha Ramasamy
NeurIPS8
2025 Spatiotemporal Attention Network for Chl-a Prediction With Sparse Multifactor Observations
abstract
Chlorophyll-a (Chl-a) is a critical indicator of water quality, and accurate Chl-a prediction is essential for marine ecosystem protection. However, existing methods for Chl-a prediction cannot adequately uncover the correlations between Chl-a and other environmental factors, e.g., SST and PAR. In addition, it is also difficult for these methods to learn the burst distributions of Chl-a data, i.e., increasing sharply for certain short periods of time and remaining stable for the rest of time. Furthermore, as original Chl-a, SST, and PAR data are often of high sparsity, most approaches rely on complete reanalysis data, which can incur accumulated error accumulation and degrade prediction performance. To address these three issues, we proposed a Spatio-Temporal Attention Network entitled SMO-STANet for Chl-a prediction. Concretely, the multi-branch spatio-temporal embedding module and spatio-temporal attention module are developed to learn the correlations between Chl-a and the two external factors, i.e., SST and PAR, thus facilitating the learning of the underlying spatio-temporal distribution of Chl-a. In addition, we designed a scaled loss function to enable SMO-STANet to adapt to the burst distributions of Chl-a. Last but not least, we develops a sparse observation data completion module to address the issue of data sparsity. According to the experimental results on two real datasets, SMO-STANet outperforms existing methods for Chl-a prediction by a large margin. The code is available at https://github.com/ADMIS-TONGJI/SMO-STANet.
Xudong Jiang 0001, Wengen Li, Jihong Guan
IEEE Geosci. Remote. Sens. Lett.1
2025 MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
abstract
This paper proposes a large-scale multi-modal dataset for referring motion expression video segmentation, focusing on segmenting and tracking target objects in videos based on language description of objects' motions. Existing referring video segmentation datasets often focus on salient objects and use language expressions rich in static attributes, potentially allowing the target object to be identified in a single frame. Such datasets underemphasize the role of motion in both videos and languages. To explore the feasibility of using motion expressions and motion reasoning clues for pixel-level video understanding, we introduce MeViS, a dataset containing 33,072 human-annotated motion expressions in both text and audio, covering 8,171 objects in 2,006 videos of complex scenarios. We benchmark 15 existing methods across 4 tasks supported by MeViS, including 6 referring video object segmentation (RVOS) methods, 3 audio-guided video object segmentation (AVOS) methods, 2 referring multi-object tracking (RMOT) methods, and 4 video captioning methods for the newly introduced referring motion expression generation (RMEG) task. The results demonstrate weaknesses and limitations of existing methods in addressing motion expression-guided video understanding. We further analyze the challenges and propose an approach LMPM++ for RVOS/AVOS/RMOT that achieves new state-of-the-art results. Our dataset provides a platform that facilitates the development of motion expression-guided video understanding algorithms in complex video scenes.
Henghui Ding, Chang Liu 0072, Shuting He, Kaining Ying, Xudong Jiang 0001, Chen Change Loy, Yu-Gang Jiang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Cascaded Dynamic Memory Refinement and Semantic Alignment for Exo-to-Ego Cross-View Video Generation
abstract
Cross-view video generation from exocentric (third-person) to egocentric (first-person) perspectives poses a challenging task, due to the significant viewpoint gap and limited overlap between these two views. Previous methods exhibit limitations in capturing long-range temporal context and overlook egocentric semantic priors, leading to degraded performance in cross-view synthesis. To address these challenges, we propose a cue-free video-based approach termed cascaded Dynamic memory Refinement and Semantic Alignment (DRSA), which integrates temporal knowledge over extended periods and learns egocentric semantic information to generate videos. The Dynamic Memory Refinement (DMR) exploits long horizon temporal dynamics to learn salient information that compensates for the limited overlap between views. Specifically, we devise a dynamic memory that serves as a knowledge repository, and utilize a sliding window to locate the corresponding long-term temporal information, which is subsequently processed with adaptive weighting and cross-attention transformer to refine feature representations. Furthermore, aware of the considerable viewpoint divergence that hinder semantic learning of target view, we propose Viewpoint-aware Semantic Alignment (VSA) with dual encoder-decoder learning and semantic alignment, which transfer egocentric semantic details from the egocentric synthesis pipeline to the exocentric synthesis pipeline. In particular, the VSA module narrows the semantic gap between views, further promoting long-range temporal modeling in DMR under alignment constraints. By extending this into a cascaded fashion, the Cascaded Alignment and Refinement (CAR) progressively aligns semantic features and performs feature refinement to facilitate viewpoint learning at different levels of granularity. To overcome the limitations of existing databases known for their limited static scenes and scarcity of interacting objects, we create a new dataset with dynamic exocentric scenes and rich interacting objects to further promote the task. Thorough experimental analysis reveals that our method surpasses current state-of-the-art techniques in terms of both quantitative metrics and qualitative evaluations.
Weipeng Hu, Jiun Tian Hoe, Haifeng Hu 0001, Xudong Jiang 0001, Yap-Peng Tan
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Revisiting One-Stage Deep Uncalibrated Photometric Stereo via Fourier Embedding
abstract
This paper introduces a one-stage deep uncalibrated photometric stereo (UPS) network, namely Fourier Uncalibrated Photometric Stereo Network (FUPS-Net), for non-Lambertian objects under unknown light directions. It departs from traditional two-stage methods that first explicitly learn lighting information and then estimate surface normals. Two-stage methods were deployed because the interplay of lighting with shading cues presents challenges for directly estimating surface normals without explicit lighting information. However, these two-stage networks are disjointed and separately trained so that the error in explicit light calibration will propagate to the second stage and cannot be eliminated. In contrast, the proposed FUPS-Net utilizes an embedded Fourier transform network to implicitly learn lighting features by decomposing inputs, rather than employing a disjointed light estimation network. Our approach is motivated from observations in the Fourier domain of photometric stereo images: lighting information is mainly encoded in amplitudes, while geometry information is mainly associated with phases. Leveraging this property, our method "decomposes" geometry and lighting in the Fourier domain as guidance, via the proposed Fourier Embedding Extraction (FEE) block and Fourier Embedding Aggregation (FEA) block, which generate lighting and geometry features for the FUPS-Net to implicitly resolve the geometry-lighting ambiguity. Furthermore, we propose a Frequency-Spatial Weighted (FSW) block that assigns weights to combine features extracted from the frequency domain and those from the spatial domain for enhancing surface reconstructions. FUPS-Net overcomes the limitations of two-stage UPS methods, offering better training stability, a concise end-to-end structure, and avoiding accumulated errors in disjointed networks. Experimental results on synthetic and real datasets demonstrate the superior performance of our approach, and its simpler training setup, potentially paving the way for a new strategy in deep learning-based UPS methods.
Yakun Ju, Boxin Shi, Bihan Wen, Kin-Man Lam 0001, Xudong Jiang 0001, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Revisiting Supervised Learning-Based Photometric Stereo Networks
abstract
Deep learning has significantly propelled the development of photometric stereo by handling the challenges posed by unknown reflectance and global illumination effects. However, how supervised learning-based photometric stereo networks resolve these challenges remains to be elucidated. In this paper, we aim to reveal how existing methods address these challenges by revisiting their deep features, deep feature encoding strategies, and network architectures. Based on the insights gained from our analysis, we propose ESSENCE-Net, which effectively encodes deep shading features with an easy-first-encoding strategy, enhances shading features with shading supervision, and accurately decodes normal with spatial context-aware attention. The experimental results verify that the proposed method outperforms state-of-the-art methods on three benchmark datasets, whether with dense or sparse inputs.
Xiaoyao Wei, Zongrui Li 0001, Binjie Ding, Boxin Shi, Xudong Jiang 0001, Gang Pan 0001, Yanlong Cao
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Radar gait recognition using Dual-branch Swin Transformer with Asymmetric Attention Fusion
abstract
Video-based gait recognition suffers from potential privacy issues and performance degradation due to dim environments, partial occlusions, or camera view changes. Radar has recently become increasingly popular and overcome various challenges presented by vision sensors. To capture tiny differences in radar gait signatures of different people, a dual-branch Swin Transformer is proposed, where one branch captures the time variations of the radar micro-Doppler signature and the other captures the repetitive frequency patterns in the spectrogram. Unlike natural images where objects can be translated, rotated, or scaled, the spatial coordinates of spectrograms and CVDs have unique physical meanings, and there is no affine transformation for radar targets in these synthetic images. The patch splitting mechanism in Vision Transformer makes it ideal to extract discriminant information from patches, and learn the attentive information across patches, as each patch carries some unique physical properties of radar targets. Swin Transformer consists of a set of cascaded Swin blocks to extract semantic features from shallow to deep representations, further improving the classification performance. Lastly, to highlight the branch with larger discriminant power, an Asymmetric Attention Fusion is proposed to optimally fuse the discriminant features from the two branches. To enrich the research on radar gait recognition, a large-scale NTU-RGR dataset is constructed, containing 45,768 radar frames of 98 subjects. The proposed method is evaluated on the NTU-RGR dataset and the MMRGait-1.0 database. It consistently and significantly outperforms all the compared methods on both datasets. The codes are available at: https://github.com/wentaoheunnc/NTU-RGR . • The proposed method could well extract complementary information from both spectrograms and CVDs. • The proposed Swin-T could extract discriminant features with physical meanings. • The proposed asymmetric attention fusion could effectively combine features with known importance. • A large-scale benchmark dataset, NTU-RGR dataset, is developed to advance the radar gait recognition.
Jianfeng Ren, Ruibin Bai, Xudong Jiang 0001
Pattern Recognit.4
2025 Two-stage Rule-induction visual reasoning on RPMs with an application to video prediction
abstract
Raven's Progressive Matrices (RPMs) are frequently used in evaluating human's visual reasoning ability. Researchers have made considerable efforts in developing systems to automatically solve the RPM problem, often through a black-box end-to-end convolutional neural network for both visual recognition and logical reasoning tasks. Based on the intrinsic natures of RPM problem, we propose a Two-stage Rule-Induction Visual Reasoner (TRIVR), which consists of a perception module and a reasoning module, to tackle the challenges of real-world visual recognition and subsequent logical reasoning tasks, respectively. For the reasoning module, we further propose a “2+1” formulation that models human's thinking in solving RPMs and significantly reduces the model complexity. It derives a reasoning rule from each RPM sample, which is not feasible for existing methods. As a result, the proposed reasoning module is capable of yielding a set of reasoning rules modeling human in solving the RPM problems. To validate the proposed method on real-world applications, an RPM-like Video Prediction (RVP) dataset is constructed, where visual reasoning is conducted on RPMs constructed using real-world video frames. Experimental results on various RPM-like datasets demonstrate that the proposed TRIVR achieves a significant and consistent performance gain compared with state-of-the-art models.
Jianfeng Ren, Ruibin Bai, Xudong Jiang 0001
Pattern Recognit.4
2025 Adaptive Graph K-Means
Shenfei Pei, Yuanchen Sun, Feiping Nie 0001, Xudong Jiang 0001, Zengwei Zheng
Pattern Recognit.4
2025 Hierarchical Relation Learning for Few-Shot Semantic Segmentation in Remote Sensing Images
abstract
Few-shot semantic segmentation (FSS) aims to segment specific semantic classes in a query image using only a few annotated support samples. While FSS has gained significant attention in natural image processing, it remains underexplored in the more challenging domain of remote sensing images (RSIs). Existing FSS approaches for RSIs primarily focus on enhancing feature representations of support or query images through hierarchical/multi-level feature fusion. However, unlike fully supervised segmentation that relies on feature extraction and optimization, FSS requires segmenting the query image based on its relations with annotated support images. To address this need, we propose the concept of Hierarchical Relation Learning (HRL) to explore the intrinsic support-query relations, allowing for the direct refinement of target object appearances in the query image. Specifically, we propose a Hierarchical Relation Network (HRNet), which performs single-scale relation extraction at each network hierarchy and multi-scale relation aggregation across hierarchies. In addition, we construct a Bidirectional Hierarchical Loss (BHLoss) to guide HRNet training, providing targeted supervision at each hierarchy in both top-down and bottom-up directions, thus facilitating robust multi-scale relation learning across hierarchies. Comprehensive experiments on the iSAID-5i, DLRSD-5i, and LoveDA-2i datasets demonstrate the superiority of the proposed HRL. The code will be available at https://github.com/XinnHe/HRL.
Xin He 0024, Yun Liu 0011, Yong Zhou 0003, Henghui Ding, Jiaqi Zhao 0001, Bing Liu 0016, Xudong Jiang 0001
IEEE Trans. Geosci. Remote. Sens.7
2025 STDMamba: Spatiotemporal Decomposition Mamba for Long-Term Fine-Grained SST Prediction
abstract
Long-term prediction of Sea Surface Temperature (SST) is a pervasive issue in ocean science, particularly for understanding climate changes and improving marine disaster risk assessment. However, existing approaches typically focus either on short-term or long-term coarse-grained prediction, due to their limited ability to overcome noise interference and model complex spatio-temporal dependencies in long SST sequences. To overcome these limitations, we proposed a novel spatio-temporal decomposition Mamba model, termed STDMamba, for long-term fine-grained SST prediction. First, we introduce a Gaussian-weighted series decomposition module with smoothing mechanisms to decompose SST sequences into trend and fluctuation components, thereby mitigating noise interference. Then, we design a dual spatio-temporal representation learning module which utilizes Mamba2 to effectively capture the long-term spatio-temporal dependencies in the fluctuation component, and employs a Temporal Convolutional Network (TCN) to learn the spatio-temporal feature representation of the trend component. Finally, the dual representations are fused and passed through a prediction layer to generate the long-term fine-grained SST prediction results. Experiments on real-world datasets demonstrate that STDMamba significantly outperforms state-of-the-art prediction models. The code of STDMamba is available at https://github.com/ADMIS-TONGJI/STDMamba.
Xudong Jiang 0001, Wengen Li, Hanchen Yang 0002, Jihong Guan, Yichao Zhang 0001, Shuigeng Zhou
IEEE Trans. Geosci. Remote. Sens.1
2025 Line-of-Sight Depth Attention for Panoptic Parsing of Distant Small-Faint Instances
abstract
Current scene parsers have effectively distilled abstract relationships among refined instances, while overlooking the discrepancies arising from variations in scene depth. Hence, their potential to imitate the intrinsic 3D perception ability of humans is constrained. In accordance with the principle of perspective, we advocate first grading the depth of the scenes into several slices, and then digging semantic correlations within a slice or between multiple slices. Two attention-based components, namely the Scene Depth Grading Module (SDGM) and the Edge-oriented Correlation Refining Module (EoCRM), comprise our framework, the Line-of-Sight Depth Network (LoSDN). SDGM grades scene into several slices by calculating depth attention tendencies based on parameters with explicit physical meanings, e.g., albedo, occlusion, specular embeddings. This process allocates numerous multi-scale instances to each scene slice based on their line-of-sight extension distance, establishing a solid groundwork for ordered association mining in EoCRM. Since the primary step in distinguishing distant faint targets is boundary delineation, EoCRM implements edge-wise saliency quantification and association digging. Quantitative and diagnostic experiments on Cityscapes, ADE20K, and PASCAL Context datasets reveal the competitiveness of LoSDN and the individual contribution of each highlight. Visualizations display that our strategy offers clear benefits in detecting distant, faint targets.
Zhongqi Lin, Xudong Jiang 0001, Zengwei Zheng
IEEE Trans. Image Process.2
2025 Spatial Frequency Modulation Network for Efficient Image Dehazing
abstract
Currently, two main research lines in efficient context modeling for image dehazing are tailoring effective feature modulation mechanisms and utilizing the Fourier transform more precisely. The former is usually based on self-scale features that ignore complementary cross-scale/level features, and the latter tends to overlook regions with pronounced haze degradation and intricate structures. This paper introduces a novel spatial and frequency modulation perspective to synergistically investigate contextual feature modeling for efficient image dehazing. Specifically, we delicately develop a Spatial Frequency Modulator (SFM) equipped with a Cross-Scale Modulator (CSM) and Frequency Modulator (FM) to implement intra-block feature modulation. The CSM progressively aggregates hierarchical features across different scales, employing them for spatial self-modulation, and the FM subsequently adopts a dual-branch design to focus more on the crucial areas with severe haze and complex structures for reconstruction. Further, we propose a Cross-Level Modulator (CLM) to facilitate inter-block feature mutual modulation, enhancing seamless interaction between features at different depths and layers. Integrating the above-developed modules into the U-Net architecture, we construct a two-stage spatial frequency modulation network (SFMN). Extensive quantitative and qualitative evaluations showcase the superior performance and efficiency of the proposed SFMN over recent state-of-the-art image dehazing methods. The source code can be found in https://github.com/it-hao/SFMN.
Hao Shen 0006, Henghui Ding, Yulun Zhang 0001, Zhong-Qiu Zhao, Xudong Jiang 0001
IEEE Trans. Image Process.5
2024 Pano-NeRF: Synthesizing High Dynamic Range Novel Views with Geometry from Sparse Low Dynamic Range Panoramic Images
abstract
Panoramic imaging research on geometry recovery and High Dynamic Range (HDR) reconstruction becomes a trend with the development of Extended Reality (XR). Neural Radiance Fields (NeRF) provide a promising scene representation for both tasks without requiring extensive prior data. How- ever, in the case of inputting sparse Low Dynamic Range (LDR) panoramic images, NeRF often degrades with under-constrained geometry and is unable to reconstruct HDR radiance from LDR inputs. We observe that the radiance from each pixel in panoramic images can be modeled as both a signal to convey scene lighting information and a light source to illuminate other pixels. Hence, we propose the irradiance fields from sparse LDR panoramic images, which increases the observation counts for faithful geometry recovery and leverages the irradiance-radiance attenuation for HDR reconstruction. Extensive experiments demonstrate that the irradiance fields outperform state-of-the-art methods on both geometry recovery and HDR reconstruction and validate their effectiveness. Furthermore, we show a promising byproduct of spatially-varying lighting estimation. The code is available at https://github.com/Lu-Zhan/Pano-NeRF.
Zhan Lu, Boxin Shi, Xudong Jiang 0001
AAAI4
2024 InteractDiffusion: Interaction Control in Text-to-Image Diffusion Models
abstract
Large-scale text-to-image (T2I) diffusion models have showcased incredible capabilities in generating coherent images based on textual descriptions, enabling vast applications in content generation. While recent advancements have introduced control over factors such as object localization, posture, and image contours, a crucial gap remains in our ability to control the interactions between objects in the generated content. Well-controlling interactions in generated images could yield meaningful applications, such as creating realistic scenes with interacting characters. In this work, we study the problems of conditioning T2I diffusion models with Human-Object Interaction (HOI) information, consisting of a triplet label (person, action, object) and corresponding bounding boxes. We propose a pluggable interaction control model, called InteractDiffusion that extends existing pre-trained T2I diffusion models to enable them being better conditioned on interactions. Specifically, we tokenize the HOI information and learn their relationships via interaction embeddings. A conditioning self-attention layer is trained to map HOI tokens to visual tokens, thereby conditioning the visual tokens better in existing T2I diffusion models. Our model attains the ability to control the interaction and location on existing T2I diffusion models, which outperforms existing baselines by a large margin in HOI detection score, as well as fidelity in FID and KID. Project page: https://jiuntian.github.io/interactdiffusion.
Jiun Tian Hoe, Xudong Jiang 0001, Chee Seng Chan, Yap-Peng Tan, Weipeng Hu
CVPR2
2024 Spin-UP: Spin Light for Natural Light Uncalibrated Photometric Stereo
abstract
Natural Light Uncalibrated Photometric Stereo (NaUPS) relieves the strict environment and light assumptions in classical Uncalibrated Photometric Stereo (UPS) methods. However, due to the intrinsic ill-posedness and high-dimensional ambiguities, addressing NaUPS is still an open question. Existing works impose strong assumptions on the environment lights and objects' material, restricting the effectiveness in more general scenarios. Alternatively, some methods leverage supervised learning with intricate models while lacking interpretability, resulting in a biased estimation. In this work, we propose Spin Light Uncalibrated Photometric Stereo (Spin-UP), an unsupervised method to tackle NaUPS in various environment lights and objects. The proposed method uses a novel setup that captures the object's images on a rotatable platform, which mitigates NaUPS's ill-posedness by reducing unknowns and provides reliable priors to alleviate NaUPS's ambiguities. Leveraging neural inverse rendering and the proposed training strategies, Spin-UP recovers surface normals, environment light, and isotropic reflectance under complex natural light with low computational cost. Experiments have shown that Spin-UP outperforms other supervised / unsupervised NaUPS meth-ods and achieves state-of-the-art performance on synthetic and real-world datasets. Codes and data are available at https://github.com/LMozart/CVPR2024-SpinUP.
Zongrui Li 0001, Zhan Lu, Haojie Yan, Boxin Shi, Gang Pan 0001, Xudong Jiang 0001
CVPR7
2024 🤖 SegPoint: Segment Any Point Cloud via Large Language Model
Shuting He, Henghui Ding, Xudong Jiang 0001, Bihan Wen
ECCV (22)3
2024 Connecting Consistency Distillation to Score Distillation for Text-to-3D Generation
Zongrui Li 0001, Minghui Hu 0001, Xudong Jiang 0001
ECCV (43)4
2024 Mitigating the Curse of Dimensionality for Certified Robustness via Dual Randomized Smoothing
abstract
Randomized Smoothing (RS) has been proven a promising method for endowing an arbitrary image classifier with certified robustness. However, the substantial uncertainty inherent in the high-dimensional isotropic Gaussian noise imposes the curse of dimensionality on RS. Specifically, the upper bound of ${\ell_2}$ certified robustness radius provided by RS exhibits a diminishing trend with the expansion of the input dimension $d$, proportionally decreasing at a rate of $1/\sqrt{d}$. This paper explores the feasibility of providing ${\ell_2}$ certified robustness for high-dimensional input through the utilization of dual smoothing in the lower-dimensional space. The proposed Dual Randomized Smoothing (DRS) down-samples the input image into two sub-images and smooths the two sub-images in lower dimensions. Theoretically, we prove that DRS guarantees a tight ${\ell_2}$ certified robustness radius for the original input and reveal that DRS attains a superior upper bound on the ${\ell_2}$ robustness radius, which decreases proportionally at a rate of $(1/\sqrt m + 1/\sqrt n )$ with $m+n=d$. Extensive experiments demonstrate the generalizability and effectiveness of DRS, which exhibits a notable capability to integrate with established methodologies, yielding substantial improvements in both accuracy and ${\ell_2}$ certified robustness baselines of RS on the CIFAR-10 and ImageNet datasets. Code is available at https://github.com/xiasong0501/DRS.
Song Xia, Yi Yu 0011, Xudong Jiang 0001, Henghui Ding
ICLR3
2024 Regression Residual Reasoning with Pseudo-labeled Contrastive Learning for Uncovering Multiple Complex Compositional Relations
Chengtai Li, Yuting He 0002, Jianfeng Ren, Ruibin Bai, Yitian Zhao, Heng Yu 0001, Xudong Jiang 0001
IJCAI7
2024 Continual Learning for Robust Gate Detection under Dynamic Lighting in Autonomous Drone Racing
abstract
In autonomous and mobile robotics, a principal challenge is resilient real-time environmental perception, particularly in situations characterized by unknown and dynamic elements, as exemplified in the context of autonomous drone racing. This study introduces a perception technique for detecting drone racing gates under illumination variations, which is common during high-speed drone flights. The proposed technique relies upon a lightweight neural network backbone augmented with capabilities for continual learning. The envisaged approach amalgamates predictions of the gates' positional coordinates, distance, and orientation, encapsulating them into a cohesive pose tuple. A comprehensive number of tests serve to underscore the efficacy of this approach in confronting diverse and challenging scenarios, specifically those involving variable lighting conditions. The proposed methodology exhibits notable robustness in the face of illumination variations, thereby substantiating its effectiveness.
Zhongzheng Qiao, Xuan Huy Pham, Savitha Ramasamy, Xudong Jiang 0001, Erdal Kayacan, Andriy Sarabakha
IJCNN4
2024 Generalized Few-Shot 3D Point Cloud Segmentation
abstract
Few-Shot 3D Point Cloud Semantic Segmentation (3D-FS) mitigates the issues of insufficient data annotation and emerging new classes in real-world scenarios, but it totally ignores the performance on base classes. In this paper, we address a more practical task named Generalized Few-Shot 3D Point Cloud Semantic Segmentation (3D-GFS), which aims to perform segmentation on both the base classes with adequate samples and the novel classes with few samples simultaneously. Based on the prototypical Base Model for the 3D-GFS task, we propose an Adaptive Support Enrichment module and a Query Aware Representation module to utilize the contextual information of semantic segmentation. The former exploits the essential co-relationship between base and novel classes in support samples while the latter mines the semantic information from individual query samples. Besides, considering the different embedding spaces, we propose a new training strategy to get a better representation of prototypes for further performance improvement. Extensive experiments on S3DIS and ScanNet show that our proposed method outperforms our Base Model and the conventional 3D-FS methods.
Shuqian Yang, Henhui Ding, Xudong Jiang 0001
ISCAS3
2024 Class-incremental Learning for Time Series: Benchmark and Evaluation
abstract
Real-world environments are inherently non-stationary, frequently introducing new classes over time. This is especially common in time series classification, such as the emergence of new disease classification in healthcare or the addition of new activities in human activity recognition. In such cases, a learning system is required to assimilate novel classes effectively while avoiding catastrophic forgetting of the old ones, which gives rise to the Class-incremental Learning (CIL) problem. However, despite the encouraging progress in the image and language domains, CIL for time series data remains relatively understudied. Existing studies suffer from inconsistent experimental designs, necessitating a comprehensive evaluation and benchmarking of methods across a wide range of datasets. To this end, we first present an overview of the Time Series Class-incremental Learning (TSCIL) problem, highlight its unique challenges, and cover the advanced methodologies. Further, based on standardized settings, we develop a unified experimental framework that supports the rapid development of new algorithms, easy integration of new datasets, and standardization of the evaluation process. Using this framework, we conduct a comprehensive evaluation of various generic and time-series-specific CIL methods in both standard and privacy-sensitive scenarios. Our extensive experiments not only provide a standard baseline to support future research but also shed light on the impact of various design factors such as normalization layers or memory budget thresholds. Codes are available at https://github.com/zqiao11/TSCIL.
Zhongzheng Qiao, Quang Pham, Hoang H. Le, Ponnuthurai N. Suganthan, Xudong Jiang 0001, Savitha Ramasamy
KDD6
2024 Visual-linguistic Cross-domain Feature Learning with Group Attention and Gamma-correct Gated Fusion for Extracting Commonsense Knowledge
abstract
Acquiring commonsense knowledge about entity-pairs from images is crucial across diverse applications. Distantly supervised learning has made significant advancements by automatically retrieving images containing entity pairs and summarizing commonsense knowledge from the bag of images. However, the retrieved images may not always cover all possible relations, and the informative features across the bag of images are often overlooked. To address these challenges, a Multi-modal Cross-domain Feature Learning framework is proposed to incorporate the general domain knowledge from a large vision-text foundation model, ViT-GPT2, to handle unseen relations and exploit complementary information from multiple sources. Then, a Group Attention module is designed to exploit the attentive information from other instances of the same bag to boost the informative features of individual instances. Finally, a Gamma-corrected Gated Fusion is designed to select a subset of informative instances for a comprehensive summarization of commonsense entity relations. Extensive experimental results demonstrate the superiority of the proposed method over state-of-the-art models for extracting commonsense knowledge.
Jialu Zhang 0003, Chenglin Yao, Jianfeng Ren, Xudong Jiang 0001
ACM Multimedia5
2024 Event-ID: Intrinsic Decomposition Using an Event Camera
abstract
Reconstructing 3D scenes from multi-view images is challenging, especially under extreme scenarios. We propose Event-ID, an event-based intrinsic decomposition framework that leverages events and images for stable decomposition under extreme scenarios. Our method is based on two observations: event cameras maintain good imaging quality under blurry or poorly exposed scenarios, and event signals from different viewpoints exhibit similarity in diffuse regions while varying in specular regions. We establish an event-based reflectance model and introduce an event-based warping method to extract specular clues. Our two-stage framework constructs a radiance field and decomposes the scene into normal, material, and lighting. Experimental results demonstrate superior performance compared to state-of-the-art methods. Our project can be found at https://zehaoc.github.io/EventID.github.io/
Zhan Lu, De Ma, Huajin Tang, Xudong Jiang 0001, Gang Pan 0001
ACM Multimedia5
2024 Hierarchical Perceptual and Predictive Analogy-Inference Network for Abstract Visual Reasoning
abstract
Advances in computer vision research enable human-like high-dimensional perceptual induction over analogical visual reasoning problems, such as Raven's Progressive Matrices (RPMs). In this paper, we propose a Hierarchical Perception and Predictive Analogy-Inference network (HP^2AI), consisting of three major components that tackle key challenges of RPM problems. Firstly, in view of the limited receptive fields of shallow networks in most existing RPM solvers, a perceptual encoder is proposed, consisting of a series of hierarchically coupled Patch Attention and Local Context (PALC) blocks, which could capture local attributes at early stages and capture the global panel layout at deep stages. Secondly, most methods seek for object-level similarities to map the context images directly to the answer image, while failing to extract the underlying analogies. The proposed reasoning module, Predictive Analogy-Inference (PredAI), consists of a set of Analogy-Inference Blocks (AIBs) to model and exploit the inherent analogical reasoning rules instead of object similarity. Lastly, the Squeeze-and-Excitation Channel-wise Attention (SECA) in the proposed PredAI discriminates essential attributes and analogies from irrelevant ones. Extensive experiments over four benchmark RPM datasets show that the proposed HP^2AI achieves significant performance gains over all the state-of-the-art methods consistently on all four datasets.
Jianfeng Ren, Ruibin Bai, Xudong Jiang 0001
ACM Multimedia4
2024 Transferable Adversarial Attacks on SAM and Its Downstream Models
abstract
The utilization of large foundational models has a dilemma: while fine-tuning downstream tasks from them holds promise for making use of the well-generalized knowledge in practical applications, their open accessibility also poses threats of adverse usage. This paper, for the first time, explores the feasibility of adversarial attacking various downstream models fine-tuned from the segment anything model (SAM), by solely utilizing the information from the open-sourced SAM. In contrast to prevailing transfer-based adversarial attacks, we demonstrate the existence of adversarial dangers even without accessing the downstream task and dataset to train a similar surrogate model. To enhance the effectiveness of the adversarial attack towards models fine-tuned on unknown datasets, we propose a universal meta-initialization (UMI) algorithm to extract the intrinsic vulnerability inherent in the foundation model, which is then utilized as the prior knowledge to guide the generation of adversarial perturbations. Moreover, by formulating the gradient difference in the attacking process between the open-sourced SAM and its fine-tuned downstream models, we theoretically demonstrate that a deviation occurs in the adversarial update direction by directly maximizing the distance of encoded feature embeddings in the open-sourced SAM. Consequently, we propose a gradient robust loss that simulates the associated uncertainty with gradient-based noise augmentation to enhance the robustness of generated adversarial examples (AEs) towards this deviation, thus improving the transferability. Extensive experiments demonstrate the effectiveness of the proposed universal meta-initialized and gradient robust adversarial attack (UMI-GRAT) toward SAMs and their downstream models. Code is available at https://github.com/xiasong0501/GRAT.
Song Xia, Wenhan Yang, Yi Yu 0011, Xun Lin, Henghui Ding, Ling-Yu Duan, Xudong Jiang 0001
NeurIPS7
2024 Progressively-orthogonally-mapped EfficientNet for action recognition on time-range-Doppler signature
abstract
Although 2D radar signal representations, such as spectrograms and range-Doppler maps have been widely used for target recognition, 3D time-range-Doppler (TRD) has been less studied, partially because of the difficulties in extracting features from the TRD representation, i.e., shallow 3D neural networks have limited discriminant power, but repeatedly applying 3D convolutions will lead to an oversized 3D network. A hybrid 3D–2D network architecture, Progressively-Orthogonally-Mapped EfficientNet (POMEN), is proposed to address these challenges. More specifically, the proposed POMEN utilizes 3D convolutions in the earlier stages to capture the information embedded in the sparse 3D TRD representation, and to avoid the oversized feature map caused by excessively applying 3D convolutions, we propose to progressively map the 3D features into three sets of 2D features corresponding to the range-time signature, range-Doppler map and time-Doppler signature (spectrogram), respectively. Subsequently, 2D EfficientNet blocks were designed to extract discriminant information from the three sets of 2D feature maps. This hybrid 3D–2D network design effectively extracts features from the 3D TRD representation, thereby avoiding oversized features from full-sized 3D networks and the information loss of 2D networks on 2D representations. Finally, a homogeneous gated fusion network was designed to fuse the three sets of 2D features. The proposed method was evaluated on the UGRS, MIMOGR, and mmWRWD datasets. The experimental results for all datasets demonstrate that the proposed POMEN significantly and consistently outperforms the state-of-the-art models in both 2D and 3D representations.
Chenglin Yao, Jianfeng Ren, Ruibin Bai, Heshan Du, Jiang Liu 0001, Xudong Jiang 0001
Expert Syst. Appl.6
2024 Towards Open Vocabulary Learning: A Survey
abstract
In the field of visual scene understanding, deep neural networks have made impressive advancements in various core tasks like segmentation, tracking, and detection. However, most approaches operate on the close-set assumption, meaning that the model can only identify pre-defined categories that are present in the training set. Recently, open vocabulary settings were proposed due to the rapid progress of vision language pre-training. These new approaches seek to locate and recognize categories beyond the annotated label space. The open vocabulary approach is more general, practical, and effective than weakly supervised and zero-shot settings. This paper thoroughly reviews open vocabulary learning, summarizing and analyzing recent developments in the field. In particular, we begin by juxtaposing open vocabulary learning with analogous concepts such as zero-shot learning, open-set recognition, and out-of-distribution detection. Subsequently, we examine several pertinent tasks within the realms of segmentation and detection, encompassing long-tail problems, few-shot, and zero-shot settings. As a foundation for our method survey, we first elucidate the fundamental principles of detection and segmentation in close-set scenarios. Next, we examine various contexts where open vocabulary learning is employed, pinpointing recurring design elements and central themes. This is followed by a comparative analysis of recent detection and segmentation methodologies in commonly used datasets and benchmarks. Our review culminates with a synthesis of insights, challenges, and discourse on prospective research trajectories. To our knowledge, this constitutes the inaugural exhaustive literature review on open vocabulary learning.
Jianzong Wu, Xiangtai Li, Shilin Xu 0001, Haobo Yuan, Henghui Ding, Xia Li 0005, Jiangning Zhang, Yunhai Tong, Xudong Jiang 0001, Bernard Ghanem, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.10
2024 A coarse-to-fine pattern parser for mitigating the issue of drastic imbalance in pixel distribution
Zhongqi Lin, Xudong Jiang 0001, Zengwei Zheng
Pattern Recognit.2
2024 A Max-Relevance-Min-Divergence criterion for data discretization with applications on naive Bayes
abstract
In many classification models, data is discretized to better estimate its distribution. Existing discretization methods often target at maximizing the discriminant power of discretized data, while overlooking the fact that the primary target of data discretization in classification is to improve the generalization performance. As a result, the data tend to be over-split into many small bins since the data without discretization retain the maximal discriminant information. Thus, we propose a Max-Dependency-Min-Divergence (MDmD) criterion that maximizes both the discriminant information and generalization ability of the discretized data. More specifically, the Max-Dependency criterion maximizes the statistical dependency between the discretized data and the classification variable while the Min-Divergence criterion explicitly minimizes the JS-divergence between the training data and the validation data for a given discretization scheme. The proposed MDmD criterion is technically appealing, but it is difficult to reliably estimate the high-order joint distributions of attributes and the classification variable. We hence further propose a more practical solution, Max-Relevance-Min-Divergence (MRmD) discretization scheme, where each attribute is discretized separately, by simultaneously maximizing the discriminant information and the generalization ability of the discretized data. The proposed MRmD is compared with the state-of-the-art discretization algorithms under the naive Bayes classification framework on 45 benchmark datasets. It significantly outperforms all the compared methods on most of the datasets.
Shihe Wang, Jianfeng Ren, Ruibin Bai, Yuan Yao 0007, Xudong Jiang 0001
Pattern Recognit.5
2024 Multiple Complementary Priors for Multispectral Image Compressive Sensing Reconstruction
abstract
Compressive sensing (CS) techniques using a few compressed measurements have drawn considerable interest in reconstructing multispectral imagery (MSI). Nonlocal-based tensor methods have been widely used for MSI-CS reconstruction, which employ the nonlocal self-similarity (NSS) property of MSI to obtain satisfactory results. However, such methods only consider the internal priors of MSI while ignoring important external image information, for example deep-driven priors learned from a corpus of natural image datasets. Meanwhile, they usually suffer from annoying ringing artifacts due to the aggregation of overlapping patches. In this article, we propose a novel approach for highly effective MSI-CS reconstruction using multiple complementary priors (MCPs). The proposed MCP jointly exploits nonlocal low-rank and deep image priors under a hybrid plug-and-play framework, which contains multiple pairs of complementary priors, namely, internal and external, shallow and deep, and NSS and local spatial priors. To make the optimization tractable, a well-known alternating direction method of multiplier (ADMM) algorithm based on the alternating minimization framework is developed to solve the proposed MCP-based MSI-CS reconstruction problem. Extensive experimental results demonstrate that the proposed MCP algorithm outperforms many state-of-the-art CS techniques in MSI reconstruction. The source code of the proposed MCP-based MSI-CS reconstruction algorithm is available at: https://github.com/zhazhiyuan/MCP_MSI_CS_Demo.git.
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Xudong Jiang 0001, Ce Zhu
IEEE Trans. Cybern.6
2024 Spatial-Frequency Adaptive Remote Sensing Image Dehazing With Mixture of Experts
abstract
The feature modulation mechanism has been demonstrated to be particularly well-suited for efficient network design and is rarely explored in remote sensing dehazing tasks. Moreover, we observe distinct patterns in haze distribution across the low-frequency (LF) and high-frequency (HF) components of haze images from various datasets. However, existing research rarely investigated the potential solution in the frequency domain. In response, we propose a novel spatial-frequency adaptive network (SFAN), which is mainly built by the proposed mixture of modulation experts (MoME) and decoupled frequency learning block (DFLB). Different from the fixed feature modulation design used in other tasks, the MoME adopts the mixture-of-expert mechanism to dynamically learn diverse contextual features of various granularities and scales in a sample-adaptive manner and then utilize them to perform elementwise local feature modulation. This pure convolution architecture enables our network to have superior performance and efficiency tradeoffs. Furthermore, the DFLB is devised to facilitate the LF global haze removal and reconstruction of HF local texture information. At the micro level, we first utilize a mask extractor (ME) to generate the frequency mask from the input hazy image, then employ a dual-branch decoupled learning unit to boost frequency learning, and finally develop a mixture of fusion experts (MoFE) to achieve HF and LF feature interaction. Extensive experiments on publicly available dehazing datasets demonstrate that our network performs superior performance while incurring lower computational costs. Compared to the state-of-the-art approach (DEA-Net), SFAN achieves, an average, 0.83-dB PSNR improvement on five remote sensing datasets but consumes only 51% of the FLOPs. The code will be available athttps://github.com/it-hao/SFAN.
Hao Shen 0006, Henghui Ding, Yulun Zhang 0001, Xiaofeng Cong, Zhong-Qiu Zhao, Xudong Jiang 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Noise Separation and Discriminative Feature Learning for Partial Discharge Recognition
abstract
Developing intelligent methods for partial discharge (PD) diagnosis, capable of handling various types of insulation defects in switchgear, has garnered significant attention in recent years. Certain PD signals exhibit similar characteristics, often leading to their confusion with noisy signals during data acquisition. To mitigate noise interference and enhance the precision of PD recognition, this article introduces a novel framework for separating PD signals from noise and acquiring discriminative features for identifying different types of PDs. Specifically, the proposed approach incorporates an adaptive frequency sampling strategy to extract effective and efficient features for the separation of PD signals and noise, followed by the clustering of the captured signals. Phase Resolved PD (PRPD) patterns are then generated for each clustered signal group, forming the PRPD pattern database. In order to identify the informative region within the PRPD patterns, we introduce spatial correlation attention and discriminative feature learning modules. These modules aim to reduce intraclass variance and increase interclass differences in the PRPD patterns. To evaluate the effectiveness of the proposed method in separating PD signals from noise and recognizing different PD patterns, we constructed a PD recognition dataset that encompasses noise as well as three types of PDs: 1) corona, 2) internal, and 3) surface. By conducting experiments and comparing the results with state-of-the-art methods, we demonstrate the performance of our method in achieving accurate PD recognition with a notable improvement of 1.9% on the constructed PD dataset.
Jinsheng Ji, Wensong Wang, Hongqun Li, Kai Xian Lai, Yuanjin Zheng, Xudong Jiang 0001
IEEE Trans. Ind. Informatics7
2024 VGSG: Vision-Guided Semantic-Group Network for Text-Based Person Search
abstract
Text-based Person Search (TBPS) aims to retrieve images of target pedestrian indicated by textual descriptions. It is essential for TBPS to extract fine-grained local features and align them crossing modality. Existing methods utilize external tools or heavy cross-modal interaction to achieve explicit alignment of cross-modal fine-grained features, which is inefficient and time-consuming. In this work, we propose a Vision-Guided Semantic-Group Network (VGSG) for text-based person search to extract well-aligned fine-grained visual and textual features. In the proposed VGSG, we develop a Semantic-Group Textual Learning (SGTL) module and a Vision-guided Knowledge Transfer (VGKT) module to extract textual local features under the guidance of visual local clues. In SGTL, in order to obtain the local textual representation, we group textual features from the channel dimension based on the semantic cues of language expression, which encourages similar semantic patterns to be grouped implicitly without external tools. In VGKT, a vision-guided attention is employed to extract visual-related textual features, which are inherently aligned with visual cues and termed vision-guided textual features. Furthermore, we design a relational knowledge transfer, including a vision-language similarity transfer and a class probability transfer, to adaptively propagate information of the vision-guided textual features to semantic-group textual features. With the help of relational knowledge transfer, VGKT is capable of aligning semantic-group textual features with corresponding visual features without external tools and complex pairwise interaction. Experimental results on two challenging benchmarks demonstrate its superiority over state-of-the-art methods.
Shuting He, Hao Luo 0004, Wei Jiang 0009, Xudong Jiang 0001, Henghui Ding
IEEE Trans. Image Process.4
2024 Learning Relation in Crowd Using Gated Graph Convolutional Networks for DRL-Based Robot Navigation
abstract
Deep reinforcement learning (DRL) frameworks have shown their remarkable effectiveness in learning navigation policy for the mobile robot navigating in a human crowded environment. Moreover, attention mechanisms coupled with DRL allows the robot to identify neighbors with different level of influence and incorporate them into the robot’s decision. However, as the crowd density increases, attention mechanisms may fail to identify critical neighbors which can lead to significant drops in navigation efficiency. In this work, we aim to address this limitation by encoding both human-human and human-robot interaction using a special class of Graph Convolutional Networks (GCN) known as Message-Passing GCN (MP-GCN). In contrast to existing methods, where attention between robot and humans are encoded uniformly, the proposed approach named MP-GatedGCN-RL encodes asymmetric interactions using the combination of novel message-passing function and edge-wise gating mechanisms. We evaluate our approach on the simulated environments of ETH/UCY pedestrians datasets consisting of different scenarios like collision avoidance, group forming, diverging, crossing, and so on. Experimental results demonstrate that our proposed method outperforms the conventional benchmark dynamic avoidance method ORCA with a 20.6% increase in success rate and a 9.1% reduction in navigation time. Moreover, we also achieve a 5.5% enhancement in success rate compared to other state-of-the-art DRL-based methods without any additional labeled expert data nor prior supervised learning.
Haoge Jiang, Niraj Bhujel, Zhuoyi Lin, Kong-Wah Wan, Jun Li 0005, J. Senthilnath 0001, Xudong Jiang 0001
IEEE Trans. Intell. Transp. Syst.7
2024 Reinforcement Learning for Blast Furnace Ironmaking Operation With Safety and Partial Observation Considerations
abstract
Making proper decision online in complex environment during the blast furnace (BF) operation is a key factor in achieving long-term success and profitability in the steel manufacturing industry. Regulatory lags, ore source uncertainty, and continuous decision requirement make it a challenging task. Recently, reinforcement learning (RL) has demonstrated state-of-the-art performance in various sequential decision-making problems. However, the strict safety requirements make it impossible to explore optimal decisions through online trial and error. Therefore, this article proposes a novel offline RL approach designed to ensure safety, maximize return, and address issues of partially observed states. Specifically, it utilizes an off-policy actor-critic framework to infer the optimal decision from expert operation trajectories. The "actor" in this framework is jointly trained by the supervision and evaluation signals to make decision with low risk and high return. Furthermore, we investigate a recurrent version of the actor and critic networks to better capture the complete observations, which solves the partially observed Markov decision process (POMDP) arising from sensor limitations. Verification within the BF smelting process demonstrates the improvements of the proposed algorithm in performance, i.e., safety and return.
Zhaohui Jiang 0001, Xudong Jiang 0001, Yongfang Xie, Weihua Gui 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 EAD-GAN: A Generative Adversarial Network for Disentangling Affine Transforms in Images
abstract
This article proposes a generative adversarial network called explicit affine disentangled generative adversarial network (EAD-GAN), which explicitly disentangles affine transform in a self-supervised manner. We propose an affine transform regularizer to force the InfoGAN to have explicit properties of affine transform. To facilitate training an affine transform encoder, we decompose the affine matrix into two separate matrices and infer the explicit transform parameters by the least-squares method. Unlike the existing approaches, representations learned by the proposed EAD-GAN have clear physical meaning, where transforms, such as rotation, horizontal and vertical zooms, skews, and translations, are explicitly learned from training data. Thus, we set different values of each transform parameter individually to generate specifically affine transformed data by the learned network. We show that the proposed EAD-GAN successfully disentangles these attributes on the MNIST, CelebA, and dSprites datasets. EAD-GAN achieves higher disentanglement scores with a large margin compared to the state-of-the-art methods on the dSprites dataset. For example, on the dSprites dataset, EAD-GAN achieves the MIG and DCI score of 0.59 and 0.96 respectively, compared to 0.37 and 0.71, respectively, for the state-of-the-art methods.
Letao Liu, Xudong Jiang 0001, Martin Saerbeck, Justin Dauwels
IEEE Trans. Neural Networks Learn. Syst.2
2023 Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical Reasoning
abstract
Raven’s Progressive Matrices (RPMs) have been widely used to evaluate the visual reasoning ability of humans. To tackle the challenges of visual perception and logic reasoning on RPMs, we propose a Hierarchical ConViT with Attention-based Relational Reasoner (HCV-ARR). Traditional solution methods often apply relatively shallow convolution networks to visually perceive shape patterns in RPM images, which may not fully model the long-range dependencies of complex pattern combinations in RPMs. The proposed ConViT consists of a convolutional block to capture the low-level attributes of visual patterns, and a transformer block to capture the high-level image semantics such as pattern formations. Furthermore, the proposed hierarchical ConViT captures visual features from multiple receptive fields, where the shallow layers focus on the image fine details while the deeper layers focus on the image semantics. To better model the underlying reasoning rules embedded in RPM images, an Attention-based Relational Reasoner (ARR) is proposed to establish the underlying relations among images. The proposed ARR well exploits the hidden relations among question images through the developed element-wise attentive reasoner. Experimental results on three RPM datasets demonstrate that the proposed HCV-ARR achieves a significant performance gain compared with the state-of-the-art models. The source code is available at: https://github.com/wentaoheunnc/HCV-ARR.
Jialu Zhang 0003, Jianfeng Ren, Ruibin Bai, Xudong Jiang 0001
AAAI5
2023 GRES: Generalized Referring Expression Segmentation
abstract
Referring Expression Segmentation (RES) aims to generate a segmentation mask for the object described by a given language expression. Existing classic RES datasets and methods commonly support single-target expressions only, i.e., one expression refers to one target object. Multitarget and no-target expressions are not considered. This limits the usage of RES in practice. In this paper, we introduce a new benchmark called Generalized Referring Expression Segmentation (GRES), which extends the classic RES to allow expressions to refer to an arbitrary number of target objects. Towards this, we construct the first largescale GRES dataset called gRefCOCO that contains multitarget, no-target, and single-target expressions. GRES and gRefCOCO are designed to be well-compatible with RES, facilitating extensive experiments to study the performance gap of the existing RES methods on the GRES task. In the experimental study, we find that one of the big challenges of GRES is complex relationship modeling. Based on this, we propose a region-based GRES baseline ReLA that adaptively divides the image into regions with subinstance clues, and explicitly models the region-region and region-language dependencies. The proposed approach ReLA achieves new state-of-the-art performance on the both newly proposed GRES and classic RES tasks. The proposed gRefCOCO dataset and method are available at https://henghuiding.github.io/GRES.
Chang Liu 0072, Henghui Ding, Xudong Jiang 0001
CVPR3
2023 DANI-Net: Uncalibrated Photometric Stereo by Differentiable Shadow Handling, Anisotropic Reflectance Modeling, and Neural Inverse Rendering
abstract
Uncalibrated photometric stereo (UPS) is challenging due to the inherent ambiguity brought by the unknown light. Although the ambiguity is alleviated on non-Lambertian objects, the problem is still difficult to solve for more general objects with complex shapes introducing irregular shadows and general materials with complex reflectance like anisotropic reflectance. To exploit cues from shadow and reflectance to solve UPS and improve performance on general materials, we propose DANI-Net, an inverse rendering framework with differentiable shadow handling and anisotropic reflectance modeling. Unlike most previous methods that use non-differentiable shadow maps and assume isotropic material, our network benefits from cues of shadow and anisotropic reflectance through two differentiable paths. Experiments on multiple real-world datasets demonstrate our superior and robust performance.
Zongrui Li 0001, Boxin Shi, Gang Pan 0001, Xudong Jiang 0001
CVPR5
2023 Dual-Stream Siamese Vision Transformer With Mutual Attention For Radar Gait Verification
abstract
The inconspicuousness of human gait characteristic in radar signal makes it hard to differentiate different identities. In this work, a Dual-stream Siamese Vision Transformer with Mutual Attention is proposed to verify whether a pair of radar gait sequences originate from the same person or not. The proposed Siamese Vision Transformer extracts pairwise discriminant spectral information from spectrograms and cadence velocity diagrams (CVDs). The proposed Mutual Attention scheme extracts the discriminant information from each stream through a self-attention mechanism and discovers the complement information cross the two streams through a cross-attention mechanism. The proposed method is evaluated on a large benchmark radar gait verification dataset. It significantly outperforms state-of-the-art solutions.
Jiarui Li 0001, Jianfeng Ren, Xudong Jiang 0001
ICASSP5
2023 Face Recognition on Point Cloud with Cgan-Top for Denoising
abstract
Face recognition using 3D point clouds is gaining growing interest, while raw point clouds often contain a significant amount of noise due to imperfect sensors. In this paper, an end-to-end 3D face recognition on a noisy point cloud is proposed, which synergistically integrates the denoising and recognition modules. Specifically, a Conditional Generative Adversarial Network on Three Orthogonal Planes (cGAN-TOP) is designed to effectively remove the noise in the point cloud, and recover the underlying features for subsequent recognition. A Linked Dynamic Graph Convolutional Neural Network (LDGCNN) is then adapted to recognize faces from the processed point cloud, which hierarchically links both the local point features and neighboring features of multiple scales. The proposed method is validated on the Bosphorus dataset. It significantly improves the recognition accuracy under all noise settings, with a maximum gain of 14.81%.
Junyu Liu, Jianfeng Ren, Xudong Jiang 0001
ICASSP4
2023 Solving Jigsaw Puzzle of Large Eroded Gaps Using Puzzlet Discriminant Network
abstract
Solving Jigsaw puzzles has recently become an emerging research topic. Traditionally, boundary similarities are utilized for puzzle reassembly. In this paper, we solve Jigsaw Puzzles of Large Eroded Gaps (JPLEG), where boundary similarities are weak and image semantics are the only feasible clues. Inspired by human strategy in solving a puzzle, we introduce the concept of puzzlet, where fragments are gradually combined to form puzzlets of different sizes until the completion of the puzzle. Two sets of Puzzlet Discriminant Networks are designed to visually perceive whether these puzzlets are correctly reassembled. The puzzle reassembly is then formulated as a combinatorial optimization problem, and solved using a genetic algorithm. The proposed method is evaluated on two large datasets, which shows that it significantly outperforms the state-of-the-art methods for puzzle solving.
Xingke Song, Jianfeng Ren, Ruibin Bai, Xudong Jiang 0001
ICASSP5
2023 Class-Incremental Learning on Multivariate Time Series Via Shape-Aligned Temporal Distillation
abstract
Class-incremental learning (CIL) on multivariate time series (MTS) is an important yet understudied problem. Based on practical privacy-sensitive circumstances, we propose a novel distillation-based strategy using a single-headed classifier without saving historical samples. We propose to exploit Soft-Dynamic Time Warping (Soft-DTW) for knowledge distillation, which aligns the feature maps along the temporal dimension before calculating the discrepancy. Compared with Euclidean distance, Soft-DTW shows its advantages in overcoming catastrophic forgetting and balancing the stability-plasticity dilemma. We construct two novel MTS-CIL benchmarks for comprehensive experiments. Combined with a prototype augmentation strategy, our framework demonstrates significant superiority over other prominent exemplar-free algorithms.
Zhongzheng Qiao, Minghui Hu 0001, Xudong Jiang 0001, Ponnuthurai N. Suganthan, Savitha Ramasamy
ICASSP3
2023 MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions
abstract
This paper strives for motion expressions guided video segmentation, which focuses on segmenting objects in video content based on a sentence describing the motion of the objects. Existing referring video object datasets typically focus on salient objects and use language expressions that contain excessive static attributes that could potentially enable the target object to be identified in a single frame. These datasets downplay the importance of motion in video content for language-guided video object segmentation. To investigate the feasibility of using motion expressions to ground and segment objects in videos, we propose a large-scale dataset called MeViS, which contains numerous motion expressions to indicate target objects in complex environments. We benchmarked 5 existing referring video object segmentation (RVOS) methods and conducted a comprehensive comparison on the MeViS dataset. The results show that current RVOS methods cannot effectively address motion expression-guided video segmentation. We further analyze the challenges and propose a baseline approach for the proposed MeViS dataset. The goal of our benchmark is to provide a platform that enables the development of effective language-guided video segmentation algorithms that leverage motion expressions as a primary cue for object segmentation in complex video scenes. The proposed MeViS dataset has been released at https://henghuiding.github.io/MeViS.
Henghui Ding, Chang Liu 0072, Shuting He, Xudong Jiang 0001, Chen Change Loy
ICCV4
2023 MOSE: A New Dataset for Video Object Segmentation in Complex Scenes
abstract
Video object segmentation (VOS) aims at segmenting a particular object throughout the entire video clip sequence. The state-of-the-art VOS methods have achieved excellent performance (e.g., 90+% $\mathcal{J}$ & $\mathcal{F}$) on existing datasets. However, since the target objects in these existing datasets are usually relatively salient, dominant, and isolated, VOS under complex scenes has rarely been studied. To revisit VOS and make it more applicable in the real world, we collect a new VOS dataset called coMplex video Object SEgmentation (MOSE) to study the tracking and segmenting objects in complex scenarios. MOSE contains 2,149 video clips and 5,200 objects from 36 categories, with 431,725 high-quality object segmentation masks. The most notable feature of MOSE dataset is complex scenes with crowded and occluded objects. The target objects in the videos are commonly occluded by others and disappear in some frames. To analyze the proposed MOSE dataset, we benchmark 18 existing VOS methods under 4 different settings on the proposed MOSE dataset and conduct comprehensive comparisons. The experiments show that current VOS algorithms cannot well perceive objects in complex scenes. For example, under the semi-supervised VOS setting, the highest $\mathcal{J}$ & $\mathcal{F}$ by existing state-of-the-art VOS methods is only 59.4% on MOSE, much lower than their ∼90% $\mathcal{J}$ & $\mathcal{F}$ performance on DAVIS. The results reveal that although excellent performance has been achieved on existing benchmarks, there are unresolved challenges under complex scenes and more efforts are desired to explore these challenges in the future.
Henghui Ding, Chang Liu 0072, Shuting He, Xudong Jiang 0001, Philip Torr 0001, Song Bai 0001
ICCV4
2023 Video Question Answering Using Clip-Guided Visual-Text Attention
abstract
Cross-modal learning of video and text plays a key role in Video Question Answering (VideoQA). In this paper, we propose a visual-text attention mechanism to utilize the Contrastive Language-Image Pre-training (CLIP) trained on lots of general domain language-image pairs to guide the cross-modal learning for VideoQA. Specifically, we first extract video features using a TimeSformer and text features using a BERT from the target application domain, and utilize CLIP to extract a pair of visual-text features from the general-knowledge domain through the domain-specific learning. We then propose a Cross-domain Learning to extract the attention information between visual and linguistic features across the target domain and general domain. The set of CLIP-guided visual-text features are integrated to predict the answer. The proposed method is evaluated on MSVD-QA and MSRVTTQA datasets and outperforms state-of-the-art methods.
Shuhong Ye, Weikai Kong, Chenglin Yao, Jianfeng Ren, Xudong Jiang 0001
ICIP5
2023 VLT: Vision-Language Transformer and Query Generation for Referring Segmentation
abstract
We propose a Vision-Language Transformer (VLT) framework for referring segmentation to facilitate deep interactions among multi-modal information and enhance the holistic understanding to vision-language features. There are different ways to understand the dynamic emphasis of a language expression, especially when interacting with the image. However, the learned queries in existing transformer works are fixed after training, which cannot cope with the randomness and huge diversity of the language expressions. To address this issue, we propose a Query Generation Module, which dynamically produces multiple sets of input-specific queries to represent the diverse comprehensions of language expression. To find the best among these diverse comprehensions, so as to generate a better mask, we propose a Query Balance Module to selectively fuse the corresponding responses of the set of queries. Furthermore, to enhance the model's ability in dealing with diverse language expressions, we consider inter-sample learning to explicitly endow the model with knowledge of understanding different language expressions to the same object. We introduce masked contrastive learning to narrow down the features of different expressions for the same target object while distinguishing the features of different objects. The proposed approach is lightweight and achieves new state-of-the-art referring segmentation results consistently on five datasets.
Henghui Ding, Chang Liu 0072, Suchen Wang, Xudong Jiang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 PAR$^{2}$2Net: End-to-End Panoramic Image Reflection Removal
abstract
In this article, we investigate the problem of panoramic image reflection removal to relieve the content ambiguity between the reflection layer and the transmission scene. Although a partial view of the reflection scene is attainable in the panoramic image and provides additional information for reflection removal, it is not trivial to directly apply this for getting rid of undesired reflections due to its misalignment with the reflection-contaminated image. We propose an end-to-end framework to tackle this problem. By resolving misalignment issues with adaptive modules, the high-fidelity recovery of reflection layer and transmission scenes is accomplished. We further propose a new data generation approach that considers the physics-based formation model of mixture images and the in-camera dynamic range clipping to diminish the domain gap between synthetic and real data. Experimental results demonstrate the effectiveness of the proposed method and its applicability for mobile devices and industrial applications.
Yuchen Hong, Lingran Zhao, Xudong Jiang 0001, Alex Chichung Kot, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Fast-SNN: Fast Spiking Neural Network by Converting Quantized ANN
abstract
Spiking neural networks (SNNs) have shown advantages in computation and energy efficiency over traditional artificial neural networks (ANNs) thanks to their event-driven representations. SNNs also replace weight multiplications in ANNs with additions, which are more energy-efficient and less computationally intensive. However, it remains a challenge to train deep SNNs due to the discrete spiking function. A popular approach to circumvent this challenge is ANN-to-SNN conversion. However, due to the quantization error and accumulating error, it often requires lots of time steps (high inference latency) to achieve high performance, which negates SNN's advantages. To this end, this paper proposes Fast-SNN that achieves high performance with low latency. We demonstrate the equivalent mapping between temporal quantization in SNNs and spatial quantization in ANNs, based on which the minimization of the quantization error is transferred to quantized ANN training. With the minimization of the quantization error, we show that the sequential error is the primary cause of the accumulating error, which is addressed by introducing a signed IF neuron model and a layer-wise fine-tuning mechanism. Our method achieves state-of-the-art performance and low latency on various computer vision tasks, including image classification, object detection, and semantic segmentation. Codes are available at: https://github.com/yangfan-hu/Fast-SNN.
Yangfan Hu, Xudong Jiang 0001, Gang Pan 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Self-regularized prototypical network for few-shot semantic segmentation
Henghui Ding, Hui Zhang 0100, Xudong Jiang 0001
Pattern Recognit.3
2023 GAN-in-GAN for Monaural Speech Enhancement
abstract
Some generative adversarial networks (GANs) have been developed to remove background noise in real-world audio recordings. MetricGAN and its variants focus on generating a clean spectrogram from a noisy one, but the final audio quality can't be guaranteed. SEGAN and its variants directly generate an enhanced audio from a noisy one, but their over-long input representations make it less effective in identifying and removing audio noise. In this paper, a novel GAN-in-GAN framework is proposed, where the inner GAN conducts spectrogram-to-spectrogram recovery under the supervision of metric discriminators to effectively clean the audio noise, and the outer GAN conducts an audio-to-audio recovery under the supervision of multi-resolution discriminators to optimize the final audio quality. To tackle the challenges of utilizing multiple adversarial losses for training the proposed GAN-in-GAN simultaneously, a novel gradient balancing scheme is proposed to facilitate a coherent training. The proposed method is compared with state-of-the-art methods on the VoiceBank+DEMAND dataset for audio denoising. It outperforms all the compared methods.
Yicun Duan, Jianfeng Ren, Heng Yu 0001, Xudong Jiang 0001
IEEE Signal Process. Lett.4
2023 Mask Attack Detection Using Vascular-Weighted Motion-Robust rPPG Signals
abstract
Detecting 3D mask attacks to a face recognition system is challenging. Although genuine faces and 3D face masks show significantly different remote photoplethysmography (rPPG) signals, rPPG-based face anti-spoofing methods often suffer from performance degradation due to unstable face alignment in the video sequence and weak rPPG signals. To enhance the rPPG signal in a motion-robust way, a landmark-anchored face stitching method is proposed to align the faces robustly and precisely at the pixel-wise level by using both SIFT keypoints and facial landmarks. To better encode the rPPG signal, a weighted spatial-temporal representation is proposed, which emphasizes the face regions with rich blood vessels. In addition, characteristics of rPPG signals in different color spaces are jointly utilized. To improve the generalization capability, a lightweight EfficientNet with a Gated Recurrent Unit (GRU) is designed to extract both spatial and temporal features from the rPPG spatial-temporal representation for classification. The proposed method is compared with the state-of-the-art methods on five benchmark datasets under both intra-dataset and cross-dataset evaluations. The proposed method shows a significant and consistent improvement in performance over other state-of-the-art rPPG-based methods for face spoofing detection.
Chenglin Yao, Jianfeng Ren, Ruibin Bai, Heshan Du, Jiang Liu 0001, Xudong Jiang 0001
IEEE Trans. Inf. Forensics Secur.6
2023 Prototype Adaption and Projection for Few- and Zero-Shot 3D Point Cloud Semantic Segmentation
abstract
In this work, we address the challenging task of few-shot and zero-shot 3D point cloud semantic segmentation. The success of few-shot semantic segmentation in 2D computer vision is mainly driven by the pre-training on large-scale datasets like imagenet. The feature extractor pre-trained on large-scale 2D datasets greatly helps the 2D few-shot learning. However, the development of 3D deep learning is hindered by the limited volume and instance modality of datasets due to the significant cost of 3D data collection and annotation. This results in less representative features and large intra-class feature variation for few-shot 3D point cloud segmentation. As a consequence, directly extending existing popular prototypical methods of 2D few-shot classification/segmentation into 3D point cloud segmentation won't work as well as in 2D domain. To address this issue, we propose a Query-Guided Prototype Adaption (QGPA) module to adapt the prototype from support point clouds feature space to query point clouds feature space. With such prototype adaption, we greatly alleviate the issue of large feature intra-class variation in point cloud and significantly improve the performance of few-shot 3D segmentation. Besides, to enhance the representation of prototypes, we introduce a Self-Reconstruction (SR) module that enables prototype to reconstruct the support mask as well as possible. Moreover, we further consider zero-shot 3D point cloud semantic segmentation where there is no support sample. To this end, we introduce category words as semantic information and propose a semantic-visual projection model to bridge the semantic and visual spaces. Our proposed method surpasses state-of-the-art algorithms by a considerable 7.90% and 14.82% under the 2-way 1-shot setting on S3DIS and ScanNet benchmarks, respectively.
Shuting He, Xudong Jiang 0001, Wei Jiang 0009, Henghui Ding
IEEE Trans. Image Process.2
2023 Multi-Modal Mutual Attention and Iterative Interaction for Referring Image Segmentation
abstract
We address the problem of referring image segmentation that aims to generate a mask for the object specified by a natural language expression. Many recent works utilize Transformer to extract features for the target object by aggregating the attended visual regions. However, the generic attention mechanism in Transformer only uses the language input for attention weight calculation, which does not explicitly fuse language features in its output. Thus, its output feature is dominated by vision information, which limits the model to comprehensively understand the multi-modal information, and brings uncertainty for the subsequent mask decoder to extract the output mask. To address this issue, we propose Multi-Modal Mutual Attention (M3Att) and Multi-Modal Mutual Decoder (M3Dec) that better fuse information from the two input modalities. Based on M3Dec, we further propose Iterative Multi-modal Interaction (IMI) to allow continuous and in-depth interactions between language and vision features. Furthermore, we introduce Language Feature Reconstruction (LFR) to prevent the language information from being lost or distorted in the extracted feature. Extensive experiments show that our proposed approach significantly improves the baseline and outperforms state-of-the-art referring image segmentation methods on RefCOCO series datasets consistently.
Chang Liu 0072, Henghui Ding, Yulun Zhang 0001, Xudong Jiang 0001
IEEE Trans. Image Process.4
2023 Spatial Context-Aware Object-Attentional Network for Multi-Label Image Classification
abstract
Multi-label image classification is a fundamental but challenging task in computer vision. To tackle the problem, the label-related semantic information is often exploited, but the background context and spatial semantic information of related objects are not fully utilized. To address these issues, a multi-branch deep neural network is proposed in this paper. The first branch is designed to extract the discriminant information from regions of interest to detect target objects. In the second branch, a spatial context-aware approach is proposed to better capture the contextual information of an object in its surroundings by using an adaptive patch expansion mechanism. It helps the detection of small objects that are easily lost without the support of context information. The third one, the object-attentional branch, exploits the spatial semantic relations between the target object and its related objects, to better detect partially occluded, small or dim objects with the support of those easily detectable objects. To better encode such relations, an attention mechanism jointly considering the spatial and semantic relations between objects is developed. Two widely used benchmark datasets for multi-labeling classification, MS COCO and PASCAL VOC, are used to evaluate the proposed framework. The experimental results demonstrate that the proposed method outperforms the state-of-the-art methods for multi-label image classification.
Jialu Zhang 0003, Jianfeng Ren, Qian Zhang 0018, Jiang Liu 0001, Xudong Jiang 0001
IEEE Trans. Image Process.5
2023 Chromosome Detection in Metaphase Cell Images Using Morphological Priors
abstract
Reliable chromosome detection in metaphase cell (MC) images can greatly alleviate the workload of cytogeneticists for karyotype analysis and the diagnosis of chromosomal disorders. However, it is still an extremely challenging task due to the complicated characteristics of chromosomes, e.g., dense distributions, arbitrary orientations, and various morphologies. In this article, we propose a novel rotated-anchor-based detection framework, named DeepCHM, for fast and accurate chromosome detection in MC images. Our framework has three main innovations: 1) A deep saliency map representing chromosomal morphological features is learned end-to-end with semantic features. This not only enhances the feature representations for anchor classification and regression but also guides the anchor setting to significantly reduce redundant anchors. This accelerates the detection and improves the performance; 2) A hardness-aware loss weights the contribution of positive anchors, which effectively reinforces the model to identify hard chromosomes; 3) A model-driven sampling strategy addresses the anchor imbalance issue by adaptively selecting hard negative anchors for model training. In addition, a large-scale benchmark dataset with a total of 624 images and 27,763 chromosome instances was built for chromosome detection and segmentation. Extensive experimental results demonstrate that our method outperforms most state-of-the-art (SOTA) approaches and successfully handles chromosome detection, with an AP score of 93.53%.
Jun Wang 0072, Chengfeng Zhou, Songchang Chen, Jianwu Hu, Minghui Wu 0001, Xudong Jiang 0001, Dahong Qian
IEEE J. Biomed. Health Informatics6
2023 Instance-Specific Feature Propagation for Referring Segmentation
abstract
Referring segmentation aims to generate a segmentation mask for the target instance indicated by a natural language expression. There are typically two kinds of existing methods: one-stage methods that directly perform segmentation on the fused vision and language features; and two-stage methods that first utilize an instance segmentation model for instance proposal and then select one of these instances via matching them with language features. In this work, we propose a novel framework that simultaneously detects the target-of-interest via feature propagation and generates a fine-grained segmentation mask. In our framework, each instance is represented by an Instance-Specific Feature (ISF), and the target-of-referring is identified by exchanging information among all ISFs using our proposed Feature Propagation Module (FPM). Our instance-aware approach learns the relationship among all objects, which helps to better locate the target-of-interest than one-stage methods. Comparing to two-stage methods, our approach collaboratively and interactively utilizes both vision and language information for synchronous identification and segmentation. In the experimental tests, our method outperforms previous state-of-the-art methods on all three RefCOCO series datasets.
Chang Liu 0072, Xudong Jiang 0001, Henghui Ding
IEEE Trans. Multim.2
2022 A Dueling Twin Delayed DDPG Architecture for mobile robot navigation
abstract
Collision-free path planning is challenging for mobile robot navigation tasks. Recently, the deep reinforcement learning method provided a more effective way to derive safe and efficient velocity commands directly from raw sensor information. Twin Delayed Deep Deterministic policy gradient (TD3) is an efficient approach for DRL navigation. However, original TD3 exists issues such as inefficiency learning and slow convergence speed which may influence the model to derive an ideal action for mobile robot navigation. To address these issues, we proposed a novel dueling architecture model, dueling deep deterministic policy gradient (Dueling-TD3), we compose the dueling network architecture into the critic network to increase the Q-value estimate precision. The results demonstrate that our proposed model outperforms the original model in terms of route planning capabilities.
Haoge Jiang, Kong-Wah Wan, Han Wang 0001, Xudong Jiang 0001
ICARCV4
2022 Attention-Based Dual-Stream Vision Transformer for Radar Gait Recognition
abstract
Radar gait recognition is robust to light variations and less infringement on privacy. Previous studies often utilize either spectrograms or cadence velocity diagrams. While the former shows the time-frequency patterns, the latter encodes the repetitive frequency patterns. In this work, a dual-stream net-work with attention-based fusion is proposed to fully aggregate the discriminant information from these two representations. Both streams are analyzed through the Vision Trans-former, which well captures the gait characteristics embedded in these representations. The proposed method is validated on a large benchmark dataset for radar gait recognition, showing that it significantly outperforms state-of-the-art solutions.
Shiliang Chen, Jianfeng Ren, Xudong Jiang 0001
ICASSP4
2022 Boosting the Discriminant Power of Naive Bayes
abstract
Naive Bayes has been widely used in many applications because of its simplicity and ability in handling both numerical data and categorical data. However, lack of modeling of correlations between features limits its performance. In addition, noise and outliers in the real-world dataset also greatly degrade the classification performance. In this paper, we propose a feature augmentation method employing a stack auto-encoder to reduce the noise in the data and boost the discriminant power of naive Bayes. The proposed stack auto-encoder consists of two auto-encoders for different purposes. The first encoder shrinks the initial features to derive a compact feature representation in order to remove the noise and redundant information. The second encoder boosts the discriminant power of the features by expanding them into a higher-dimensional space so that different classes of samples could be better separated in the higher-dimensional space. By integrating the proposed feature augmentation method with the regularized naive Bayes, the discrimination power of the model is greatly enhanced. The proposed method is evaluated on a set of machine-learning benchmark datasets. The experimental results show that the proposed method significantly and consistently outperforms the state-of-the-art naive Bayes classifiers.
Shihe Wang, Jianfeng Ren, Xiaoyu Lian, Ruibin Bai, Xudong Jiang 0001
ICPR5
2022 iTD3-CLN: Learn to navigate in dynamic scene through Deep Reinforcement Learning
Haoge Jiang, Mahdi Abolfazli Esfahani, Keyu Wu 0002, Kong-Wah Wan, Kuan-kian Heng, Han Wang 0001, Xudong Jiang 0001
Neurocomputing7
2022 Spatial feature mapping for 6DoF object pose estimation
Jianhan Mei, Xudong Jiang 0001, Henghui Ding
Pattern Recognit.2
2022 Deep Interactive Image Matting With Feature Propagation
abstract
Image matting has attracted growing interest in recent years for its wide applications in numerous vision tasks. Most previous image matting methods rely on trimaps as auxiliary input to define the foreground, background and unknown region. However, trimaps involve fussy manual annotation efforts and are expensive to be obtained in practice. Thus, it is hard and inflexible to update user's input or achieve real-time interaction with trimaps. Although some automatic matting approaches discard trimaps, they can only be applied to some certain scenarios, like human matting, which limits their versatility. In this work, we employ clicks as interactive behaviours for image matting, to indicate the user-defined foreground, background and unknown region, and propose a click-based deep interactive image matting (DIIM) approach. Compared with trimaps, clicks provide sparse information and are much easier and more flexible, especially for novice users. Based on clicks, users can perform interactive operations and gradually correct the errors until they are satisfied with the prediction. What's more, we propose a recurrent alpha feature propagation and a full-resolution extraction module to enhance the alpha matte estimation from high-level and low-level respectively. Experimental results show that the proposed click-based deep interactive image matting approach achieves promising performance on image matting datasets.
Henghui Ding, Hui Zhang 0100, Chang Liu 0072, Xudong Jiang 0001
IEEE Trans. Image Process.4
2022 Multistage Spatio-Temporal Networks for Robust Sketch Recognition
abstract
Sketch recognition relies on two types of information, namely, spatial contexts like the local structures in images and temporal contexts like the orders of strokes. Existing methods usually adopt convolutional neural networks (CNNs) to model spatial contexts, and recurrent neural networks (RNNs) for temporal contexts. However, most of them combine spatial and temporal features with late fusion or single-stage transformation, which is prone to losing the informative details in sketches. To tackle this problem, we propose a novel framework that aims at the multi-stage interactions and refinements of spatial and temporal features. Specifically, given a sketch represented by a stroke array, we first generate a temporal-enriched image (TEI), which is a pseudo-color image retaining the temporal order of strokes, to overcome the difficulty of CNNs in leveraging temporal information. We then construct a dual-branch network, in which a CNN branch and a RNN branch are adopted to process the stroke array and the TEI respectively. In the early stages of our network, considering the limited ability of RNNs in capturing spatial structures, we utilize multiple enhancement modules to enhance the stroke features with the TEI features. While in the last stage of our network, we propose a spatio-temporal enhancement module that refines stroke features and TEI features in a joint feature space. Furthermore, a bidirectional temporal-compatible unit that adaptively merges features in opposite temporal orders, is proposed to help RNNs tackle abrupt strokes. Comprehensive experimental results on QuickDraw and TU-Berlin demonstrate that the proposed method is a robust and efficient solution for sketch recognition.
Xudong Jiang 0001, Boliang Guan, Ruomei Wang 0001, Nadia Magnenat-Thalmann
IEEE Trans. Image Process.2
2021 Panoramic Image Reflection Removal
abstract
This paper studies the problem of panoramic image reflection removal, aiming at reliving the content ambiguity between reflection and transmission scenes. Although a partial view of the reflection scene is included in the panoramic image, it cannot be utilized directly due to its misalignment with the reflection-contaminated image. We propose a two-step approach to solve this problem, by first accomplishing geometric and photometric alignment for the reflection scene via a coarse-to-fine strategy, and then restoring the transmission scene via a recovery network. The proposed method is trained with a synthetic dataset and verified quantitatively with a real panoramic image dataset. The effectiveness of the proposed method is validated by the significant performance advantage over single image-based reflection removal methods and generalization capacity to limited-FoV scenarios captured by conventional camera or mobile phone users.
Yuchen Hong, Lingran Zhao, Xudong Jiang 0001, Alex Chichung Kot, Boxin Shi
CVPR4
2021 Single Image Reflection Removal With Absorption Effect
abstract
In this paper, we consider the absorption effect for the problem of single image reflection removal. We show that the absorption effect can be numerically approximated by the average of refractive amplitude coefficient map. We then reformulate the image formation model and propose a two-step solution that explicitly takes the absorption effect into account. The first step estimates the absorption effect from a reflection-contaminated image, while the second step recovers the transmission image by taking a reflection-contaminated image and the estimated absorption effect as the input. Experimental results on four public datasets show that our two-step solution not only successfully removes reflection artifact, but also faithfully restores the intensity distortion caused by the absorption effect. Our ablation studies further demonstrate that our method achieves superior performance on the recovery of overall intensity and has good model generalization capacity. The code is available at https://github.com/q-zh/absorption.
Boxin Shi, Jinnan Chen, Xudong Jiang 0001, Ling-Yu Duan, Alex Chichung Kot
CVPR4
2021 Vision-Language Transformer and Query Generation for Referring Segmentation
abstract
In this work, we address the challenging task of referring segmentation. The query expression in referring segmentation typically indicates the target object by describing its relationship with others. Therefore, to find the target one among all instances in the image, the model must have a holistic understanding of the whole image. To achieve this, we reformulate referring segmentation as a direct attention problem: finding the region in the image where the query language expression is most attended to. We introduce transformer and multi-head attention to build a network with an encoder-decoder attention mechanism architecture that "queries" the given image with the language expression. Furthermore, we propose a Query Generation Module, which produces multiple sets of queries with different attention weights that represent the diversified comprehensions of the language expression from different aspects. At the same time, to find the best way from these diversified comprehensions based on visual clues, we further propose a Query Balance Module to adaptively select the output features of these queries for a better mask generation. Without bells and whistles, our approach is light-weight and achieves new state-of-the-art performance consistently on three referring segmentation datasets, RefCOCO, RefCOCO+, and G-Ref. Our code is available at https://github.com/henghuiding/Vision-Language-Transformer.
Henghui Ding, Chang Liu 0072, Suchen Wang, Xudong Jiang 0001
ICCV4
2021 Interaction via Bi-directional Graph of Semantic Region Affinity for Scene Parsing
abstract
In this work, we devote to address the challenging problem of scene parsing. It is well known that pixels in an image are highly correlated with each other, especially those from the same semantic region, while treating pixels independently fails to take advantage of such correlations. In this work, we treat each respective region in an image as a whole, and capture the structure topology as well as the affinity among different regions. To this end, we first divide the entire feature maps to different regions and extract respective global features from them. Next, we construct a directed graph whose nodes are regional features, and the bi-directional edges connecting every two nodes are the affinities between the regional features they represent. After that, we transfer the affinity-aware nodes in the directed graph back to corresponding regions of the image, which helps to model the region dependencies and mitigate unrealistic results. In addition, to further boost the correlation among pixels, we propose a region-level loss that evaluates all pixels in a region as a whole and motivates the network to learn the exclusive regional feature per class. With the proposed approach, we achieves new state-of-the-art segmentation results on PASCAL-Context, ADE20K, and COCO-Stuff consistently.
Henghui Ding, Hui Zhang 0100, Jun Liu 0036, Zijian Feng, Xudong Jiang 0001
ICCV6
2021 Efficient Sketch Recognition Via Compact Spatial Embedding Graph Neural Networks
abstract
Sketches are descriptive, high-level visual media in many systems and applications. However, current methods for sketch recognition are mainly based on large neural networks, which have millions of parameters and are too cumbersome to be deployed on edge devices. Besides, convolutional neural networks are not optimal, since many areas in sketches are blank and without any information. Hence, this paper aims at designing an efficient network that maintains the state-of-the-art accuracy. Our solution is a novel graph neural network that utilizes densely connected grouped convolutions on the graph representation of sketches. It allows us to extract and aggregate spatio-temporal features efficiently. Moreover, a compact spatial embedding module is introduced to explore spatial contexts, and consequently facilitates recognition. In this way, our network is small (about 10% parameters of ResNet- 18) and efficient (inference speed in 800 ~ 2,000 FPS), meanwhile has the competitive accuracy on the QuickDraw dataset.
Xudong Jiang 0001, Boliang Guan, Nadia Magnenat-Thalmann
ICME2
2021 Towards Enhancing Fine-grained Details for Image Matting
abstract
In recent years, deep natural image matting has been rapidly evolved by extracting high-level contextual features into the model. However, most current methods still have difficulties with handling tiny details, like hairs or furs. In this paper, we argue that recovering these microscopic de-tails relies on low-level but high-definition texture features. However, these features are downsampled in a very early stage in current encoder-decoder-based models, resulting in the loss of microscopic details. To address this issue, we design a deep image matting model to enhance fine-grained details. Our model consists of two parallel paths: a conventional encoder-decoder Semantic Path and an independent downsampling-free Textural Compensate Path (TCP). The TCP is proposed to extract fine-grained details such as lines and edges in the original image size, which greatly enhances the fineness of prediction. Meanwhile, to lever-age the benefits of high-level context, we propose a feature fusion unit(FFU) to fuse multi-scale features from the se-mantic path and inject them into the TCP. In addition, we have observed that poorly annotated trimaps severely affect the performance of the model. Thus we further propose a novel term in loss function and a trimap generation method to improve our model's robustness to the trimaps. The experiments show that our method outperforms previous start-of-the-art methods on the Composition-1k dataset.
Chang Liu 0072, Henghui Ding, Xudong Jiang 0001
WACV3
2021 Variational posterior approximation using stochastic gradient ascent with adaptive stepsize
Kart-Leong Lim, Xudong Jiang 0001
Pattern Recognit.2
2021 Complex common spatial patterns on time-frequency decomposed EEG for brain-computer interface
Vasilisa Mishuhina, Xudong Jiang 0001
Pattern Recognit.2
2021 A three-step classification framework to handle complex data distribution for radar UAV detection
Jianfeng Ren, Xudong Jiang 0001
Pattern Recognit.2
2021 Knowledge-aware deep framework for collaborative skin lesion segmentation and melanoma recognition
Xiaohong Wang 0003, Xudong Jiang 0001, Henghui Ding, Jun Liu 0036
Pattern Recognit.2
2021 Joint Feature Optimization and Fusion for Compressed Action Recognition
abstract
Recent methods including CoViAR and DMC-Net provide a new paradigm for action recognition since they are directly targeted at compressed videos (e.g., MPEG4 files). It avoids the cumbersome decoding procedure of traditional methods, and leverages the pre-encoded motion vectors and residuals in compressed videos to complete recognition efficiently. However, motion vectors and residuals are noisy, sparse and highly correlated information, which cannot be effectively exploited by plain and separated networks. To tackle these issues, we propose a joint feature optimization and fusion framework that better utilizes motion vectors and residuals in the following three aspects. (i) We model the feature optimization problem as a reconstruction process that represents features by a set of bases, and propose a joint feature optimization module that extracts bases in the both modalities. (ii) A low-rank non-local attention module, which combines the non-local operation with the low-rank constraint, is proposed to tackle the noise and sparsity problem during the feature reconstruction process. (iii) A lightweight feature fusion module and a self-adaptive knowledge distillation method are introduced, which use motion vectors and residuals to generate predictions similar to those from networks with optical flows. With these proposed components embedded in a baseline network, the proposed network not only achieves the state-of-the-art performance on HMDB-51 and UCF-101, but also maintains its advantage in computational complexity.
Xudong Jiang 0001, Boliang Guan, Raymond Rui Ming Tan, Ruomei Wang 0001, Nadia Magnenat-Thalmann
IEEE Trans. Image Process.2
2020 What Does Plate Glass Reveal About Camera Calibration?
abstract
This paper aims to calibrate the orientation of glass and the field of view of the camera from a single reflection-contaminated image. We show how a reflective amplitude coefficient map can be used as a calibration cue. Different from existing methods, the proposed solution is free from image contents. To reduce the impact of a noisy calibration cue estimated from a reflection-contaminated image, we propose two strategies: an optimization-based method that imposes part of though reliable entries on the map and a learning-based method that fully exploits all entries. We collect a dataset containing 320 samples as well as their camera parameters for evaluation. We demonstrate that our method not only facilitates a general single image camera calibration method that leverages image contents but also contributes to improving the performance of single image reflection removal. Furthermore, we show our byproduct output helps alleviate the ill-posed problem of estimating the panorama from a single image.
Jinnan Chen, Zhan Lu, Boxin Shi, Xudong Jiang 0001, Kim-Hui Yap, Ling-Yu Duan, Alex Chichung Kot
CVPR5
2020 PhraseClick: Toward Achieving Flexible Interactive Segmentation by Phrase and Click
Henghui Ding, Scott Cohen, Brian L. Price, Xudong Jiang 0001
ECCV (3)4
2020 Temporal Distinct Representation Learning for Action Recognition
Junwu Weng, Donghao Luo 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Xudong Jiang 0001, Junsong Yuan 0001
ECCV (7)8
2020 Skin lesion segmentation via generative adversarial networks with dual discriminators
Bai Ying Lei, Zaimin Xia, Xudong Jiang 0001, ZongYuan Ge, Yanwu Xu 0001, Jie Du 0001, Siping Chen, Tianfu Wang 0001, Shuqiang Wang
Medical Image Anal.4
2020 Feature Boosting Network For 3D Pose Estimation
abstract
In this paper, a feature boosting network is proposed for estimating 3D hand pose and 3D body pose from a single RGB image. In this method, the features learned by the convolutional layers are boosted with a new long short-term dependence-aware (LSTD) module, which enables the intermediate convolutional feature maps to perceive the graphical long short-term dependency among different hand (or body) parts using the designed Graphical ConvLSTM. Learning a set of features that are reliable and discriminatively representative of the pose of a hand (or body) part is difficult due to the ambiguities, texture and illumination variation, and self-occlusion in the real application of 3D pose estimation. To improve the reliability of the features for representing each body part and enhance the LSTD module, we further introduce a context consistency gate (CCG) in this paper, with which the convolutional feature maps are modulated according to their consistency with the context representations. We evaluate the proposed method on challenging benchmark datasets for 3D hand pose estimation and 3D full body pose estimation. Experimental results show the effectiveness of our method that achieves state-of-the-art performance on both of the tasks.
Jun Liu 0036, Henghui Ding, Amir Shahroudy, Ling-Yu Duan, Xudong Jiang 0001, Gang Wang 0012, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 Deep Clustering With Variational Autoencoder
abstract
An autoencoder that learns a latent space in an unsupervised manner has many applications in signal processing. However, the latent space of an autoencoder does not pursue the same clustering goal as Kmeans or GMM. A recent work proposes to artificially re-align each point in the latent space of an autoencoder to its nearest class neighbors during training (Song et al. 2013). The resulting new latent space is found to be much more suitable for clustering, since clustering information is used. Inspired by previous works (Song et al. 2013), in this letter we propose several extensions to this technique. First, we propose a probabilistic approach to generalize Song's approach, such that Euclidean distance in the latent space is now represented by KL divergence. Second, as a consequence of this generalization we can now use probability distributions as inputs rather than points in the latent space. Third, we propose using Bayesian Gaussian mixture model for clustering in the latent space. We demonstrated our proposed method on digit recognition datasets, MNIST, USPS and SHVN as well as scene datasets, Scene15 and MIT67 with interesting findings.
Kart-Leong Lim, Xudong Jiang 0001, Chenyu Yi
IEEE Signal Process. Lett.2
2020 Early Action Recognition With Category Exclusion Using Policy-Based Reinforcement Learning
abstract
The goal of early action recognition is to predict action label when the sequence is partially observed. The existing methods treat the early action recognition task as sequential classification problems on different observation ratios of an action sequence. Since these models are trained by differentiating positive category from all negative classes, the diverse information of different negative categories is ignored, which we believe can be collected to help improve the recognition performance. In this paper, we step towards to a new direction by introducing category exclusion to early action recognition. We model the exclusion as a mask operation on the classification probability output of a pre-trained early action recognition classifier. Specifically, we use policy-based reinforcement learning to train an agent. The agent generates a series of binary masks to exclude interfering negative categories during action execution and hence help improve the recognition accuracy. The proposed method is evaluated on three benchmark recognition datasets, NTU-RGBD, First-Person Hand Action, as well as UCF-101. The proposed method enhances the recognition accuracy consistently over all different observation ratios on the three datasets, where the accuracy improvements on the early stages are especially significant.
Junwu Weng, Xudong Jiang 0001, Wei-Long Zheng, Junsong Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Semantic Segmentation With Context Encoding and Multi-Path Decoding
abstract
Semantic image segmentation aims to classify every pixel of a scene image to one of many classes. It implicitly involves object recognition, localization, and boundary delineation. In this paper, we propose a segmentation network called CGBNet to enhance the paring results by context encoding and multi-path decoding. We first propose a context encoding module that generates context contrasted local feature to make use of the informative context and the discriminative local information. This context encoding module greatly improves the segmentation performance, especially for inconspicuous objects. Furthermore, we propose a scale-selection scheme to selectively fuse the parsing results from different-scales of features at every spatial position. It adaptively selects appropriate score maps from rich scales of features. To improve the parsing results of boundary, we further propose a boundary delineation module that encourages the location-specific very-low-level feature near the boundaries to take part in the final prediction and suppresses them far from the boundaries. Without bells and whistles, the proposed segmentation network achieves very competitive performance in terms of all three different evaluation metrics consistently on the four popular scene segmentation datasets, Pascal Context, SUN-RGBD, Sift Flow, and COCO Stuff.
Henghui Ding, Xudong Jiang 0001, Bing Shuai, Ai Qun Liu, Gang Wang 0012
IEEE Trans. Image Process.2
2020 Bi-Directional Dermoscopic Feature Learning and Multi-Scale Consistent Decision Fusion for Skin Lesion Segmentation
abstract
Accurate segmentation of skin lesion from dermoscopic images is a crucial part of computer-aided diagnosis of melanoma. It is challenging due to the fact that dermoscopic images from different patients have non-negligible lesion variation, which causes difficulties in anatomical structure learning and consistent skin lesion delineation. In this paper, we propose a novel bi-directional dermoscopic feature learning (biDFL) framework to model the complex correlation between skin lesions and their informative context. By controlling feature information passing through two complementary directions, a substantially rich and discriminative feature representation is achieved. Specifically, we place biDFL module on the top of a CNN network to enhance high-level parsing performance. Furthermore, we propose a multi-scale consistent decision fusion (mCDF) that is capable of selectively focusing on the informative decisions generated from multiple classification layers. By analysis of the consistency of the decision at each position, mCDF automatically adjusts the reliability of decisions and thus allows a more insightful skin lesion delineation. The comprehensive experimental results show the effectiveness of the proposed method on skin lesion segmentation, achieving state-of-the-art performance consistently on two publicly available dermoscopic image databases.
Xiaohong Wang 0003, Xudong Jiang 0001, Henghui Ding, Jun Liu 0036
IEEE Trans. Image Process.2
2019 Semantic Correlation Promoted Shape-Variant Context for Segmentation
abstract
Context is essential for semantic segmentation. Due to the diverse shapes of objects and their complex layout in various scene images, the spatial scales and shapes of contexts for different objects have very large variation. It is thus ineffective or inefficient to aggregate various context information from a predefined fixed region. In this work, we propose to generate a scale- and shape-variant semantic mask for each pixel to confine its contextual region. To this end, we first propose a novel paired convolution to infer the semantic correlation of the pair and based on that to generate a shape mask. Using the inferred spatial scope of the contextual region, we propose a shape-variant convolution, of which the receptive field is controlled by the shape mask that varies with the appearance of input. In this way, the proposed network aggregates the context information of a pixel from its semantic-correlated region instead of a predefined fixed region. Furthermore, this work also proposes a labeling denoising model to reduce wrong predictions caused by the noisy low-level features. Without bells and whistles, the proposed segmentation network achieves new state-of-the-arts consistently on the six public segmentation datasets.
Henghui Ding, Xudong Jiang 0001, Bing Shuai, Ai Qun Liu, Gang Wang 0012
CVPR2
2019 Boundary-Aware Feature Propagation for Scene Segmentation
abstract
In this work, we address the challenging issue of scene segmentation. To increase the feature similarity of the same object while keeping the feature discrimination of different objects, we explore to propagate information throughout the image under the control of objects' boundaries. To this end, we first propose to learn the boundary as an additional semantic class to enable the network to be aware of the boundary layout. Then, we propose unidirectional acyclic graphs (UAGs) to model the function of undirected cyclic graphs (UCGs), which structurize the image via building graphic pixel-by-pixel connections, in an efficient and effective way. Furthermore, we propose a boundary-aware feature propagation (BFP) module to harvest and propagate the local features within their regions isolated by the learned boundaries in the UAG-structured image. The proposed BFP is capable of splitting the feature propagation into a set of semantic groups via building strong connections among the same segment region but weak connections between different segment regions. Without bells and whistles, our approach achieves new state-of-the-art segmentation performance on three challenging semantic segmentation datasets, i.e., PASCAL-Context, CamVid, and Cityscapes.
Henghui Ding, Xudong Jiang 0001, Ai Qun Liu, Nadia Magnenat-Thalmann, Gang Wang 0012
ICCV2
2019 SPLINE-Net: Sparse Photometric Stereo Through Lighting Interpolation and Normal Estimation Networks
abstract
This paper solves the Sparse Photometric stereo through Lighting Interpolation and Normal Estimation using a generative Network (SPLINE-Net). SPLINE-Net contains a lighting interpolation network to generate dense lighting observations given a sparse set of lights as inputs followed by a normal estimation network to estimate surface normals. Both networks are jointly constrained by the proposed symmetric and asymmetric loss functions to enforce isotropic constrain and perform outlier rejection of global illumination effects. SPLINE-Net is verified to outperform existing methods for photometric stereo of general BRDFs by using only ten images of different lights instead of using nearly one hundred images.
Yiming Jia, Boxin Shi, Xudong Jiang 0001, Ling-Yu Duan, Alex Chichung Kot
ICCV4
2019 Dermoscopic Image Segmentation Through the Enhanced High-Level Parsing and Class Weighted Loss
abstract
Accurate skin lesion segmentation plays an important role in the computer aided analysis of melanoma. It is a challenging task due to the variation of skin lesion appearance, the low contrast with background, and the existence of the artifacts in dermoscopic images. In this paper, we try to boost the skin lesion segmentation performance based on the fully convolutional neural network. To this end, we first propose an enhanced high-level parsing (EHP) module to generate meaningful feature representation for skin lesion and make more precise delineation of the detailed lesion structure. Furthermore, to handle the imbalance data distribution of skin lesion and background, we propose a class weighted loss (CWL) to achieve more consistent lesion prediction. Experiment results evaluated on ISBI 2017 database demonstrate the effectiveness and robustness of the proposed architecture on skin lesion segmentation, achieving new state-of-the-art prediction performance.
Xiaohong Wang 0003, Henghui Ding, Xudong Jiang 0001
ICIP3
2019 Denoising Adversarial Networks for Rain Removal and Reflection Removal
abstract
This paper presents a novel adversarial scheme to perform image denoising for the tasks of rain streak removal and reflection removal. Similar to several previous works, the proposed method first estimates a prior image and then uses it to guide the inference of noise-free image. The novelty of our approach is to jointly learn the gradient and noise-free image based on an adversarial scheme. More specifically, we use the gradient map as the prior image. The inferred noise-free image guided by an estimated gradient is regarded as a negative sample, while the noise-free image guided by the ground truth of a gradient is taken as a positive sample. With the anchor defined by the ground truth of noise-free image, we play a min-max game to jointly train two optimizers for the estimation of the gradient and the inference of noise-free images. We show that both prior image and noise-free image can be accurately obtained under this adversarial scheme. Our state-of-the-art performance achieved on two public benchmark datasets validate the effectiveness of our approach.
Boxin Shi, Xudong Jiang 0001, Ling-Yu Duan, Alex Chichung Kot
ICIP3
2019 Learning local feature representation from matching, clustering and spatial transform
Jianhan Mei, Xudong Jiang 0001, Jianfei Cai 0001
J. Vis. Commun. Image Represent.2
2019 Palmprint identification using sparse and dense hybrid representation
Somaya Al-Máadeed, Xudong Jiang 0001, Imad Rida, Ahmed Bouridane
Multim. Tools Appl.2
2019 DeepDeblur: text image recovery from blur to sharp
Jianhan Mei, Ziming Wu, Yu Qiao 0001, Henghui Ding, Xudong Jiang 0001
Multim. Tools Appl.6
2019 Blood vessel segmentation from fundus image by a cascade classification framework
Xiaohong Wang 0003, Xudong Jiang 0001, Jianfeng Ren
Pattern Recognit.2
2019 Retinal vessel segmentation by a divide-and-conquer funnel-structured classification framework
Xiaohong Wang 0003, Xudong Jiang 0001
Signal Process.2
2019 Toward Achieving Robust Low-Level and High-Level Scene Parsing
abstract
In this paper, we address the challenging task of scene segmentation. We first discuss and compare two widely used approaches to retain detailed spatial information from pretrained CNN - "dilation" and "skip". Then, we demonstrate that the parsing performance of "skip" network can be noticeably improved by modifying the parameterization of skip layers. Furthermore, we introduce a "dense skip" architecture to retain a rich set of low-level information from pre-trained CNN, which is essential to improve the low-level parsing performance. Meanwhile, we propose a convolutional context network (CCN) and place it on top of pre-trained CNNs, which is used to aggregate contexts for high-level feature maps so that robust high-level parsing can be achieved. We name our segmentation network enhanced fully convolutional network (EFCN) based on its significantly enhanced structure over FCN. Extensive experimental studies justify each contribution separately. Without bells and whistles, EFCN achieves state-of-the-arts on segmentation datasets of ADE20K, Pascal Context, SUN-RGBD and Pascal VOC 2012.
Bing Shuai, Henghui Ding, Ting Liu 0009, Gang Wang 0012, Xudong Jiang 0001
IEEE Trans. Image Process.5
2019 Corrections to "Accurate Cervical Cell Segmentation From Overlapping Clumps in Pap Smear Images"
abstract
In [1], Baiying Lei was indicated as the corresponding author. Tianfu Wang and Baiying Lei should have been indicated as the corresponding authors.
Youyi Song, Ee-Leng Tan, Xudong Jiang 0001, Jie-Zhi Cheng, Bai Ying Lei, Tianfu Wang 0001
IEEE Trans. Medical Imaging3
2018 Context Contrasted Feature and Gated Multi-Scale Aggregation for Scene Segmentation
abstract
Scene segmentation is a challenging task as it need label every pixel in the image. It is crucial to exploit discriminative context and aggregate multi-scale features to achieve better segmentation. In this paper, we first propose a novel context contrasted local feature that not only leverages the informative context but also spotlights the local information in contrast to the context. The proposed context contrasted local feature greatly improves the parsing performance, especially for inconspicuous objects and background stuff. Furthermore, we propose a scheme of gated sum to selectively aggregate multi-scale features for each spatial position. The gates in this scheme control the information flow of different scale features. Their values are generated from the testing image by the proposed network learnt from the training data so that they are adaptive not only to the training data, but also to the specific testing image. Without bells and whistles, the proposed approach achieves the state-of-the-arts consistently on the three popular scene segmentation datasets, Pascal Context, SUN-RGBD and COCO Stuff.
Henghui Ding, Xudong Jiang 0001, Bing Shuai, Ai Qun Liu, Gang Wang 0012
CVPR2
2018 Deformable Pose Traversal Convolution for 3D Action and Gesture Recognition
Junwu Weng, Mengyuan Liu 0004, Xudong Jiang 0001, Junsong Yuan 0001
ECCV (7)3
2018 An Ensemble Learning Method Based on Random Subspace Sampling for Palmprint Identification
abstract
Palmprint recognition is an important and widely used biometric modality with high reliability, stability and user acceptability. In this paper we propose a simple and effective ensemble learning method for palmprint identification based on Random Subspace Sampling (RSS). To achieve it, we rely on 2D-PCA to build the random subspaces. As 2D-PCA is an unsurpevised technique, features are extracted in each subspace using 2D-LDA. A simple 1-Nearest Neighbor classifier is associated to each subspace, the final decision rule being obtained by majority voting rule. The experimental results on multispectral and PolyU palmprint datasets show very encouraging performances compared to state-of-the-art techniques.
Imad Rida, Somaya Al-Máadeed, Xudong Jiang 0001, Lunke Fei, Abdelaziz Bensrhair
ICASSP3
2018 3D Convolutional Generative Adversarial Networks for Detecting Temporal Irregularities in Videos
abstract
In this work, we introduce a novel method for video temporal irregularity detection using the discriminative framework of 3D convolutional generative adversarial networks (3D-GANs). Temporal irregularities indicate unusual video segments. Detecting such irregularities is essential to video analysis applications like video anomaly detection and video summarization. To detect temporal irregularities in videos we need to address two problems: 1) temporal irregularities are difficult to define, different situations have different irregularities, and 2) irregularities are scarce in videos. Therefore, we formulate video temporal irregularity detection as fake data detection via the discriminative framework of a designed 3D-GAN. This new formulation only employs regular videos during the training phase and detects irregularities according to the deviation estimated by the discriminator of 3D-GAN. We take regular videos as real data and construct a 3D-GAN to learn the distribution of regular videos during the training phase. Since testing data contain irregular videos or fake data, whose distribution is different from regular videos or real data, the trained discriminator of our networks is able to detect temporal regularities and irregularities. Experiments show that 3D-GANs outperforms 2D-GANs in temporal irregularity detection, and demonstrate the effectiveness and competitive performance of our approach on anomaly detection datasets.
Mengjia Yan 0003, Xudong Jiang 0001, Junsong Yuan 0001
ICPR2
2018 Color space construction by optimizing luminance and chrominance components for face recognition
Ze Lu, Xudong Jiang 0001, Alex Chichung Kot
Pattern Recognit.2
2018 Feature fusion with covariance matrix regularization in face recognition
abstract
The fusion of multiple features is important for achieving state-of-the-art face recognition results. This has been proven in both traditional and deep learning approaches. Existing feature fusion methods either reduce the dimensionality of each feature first and then concatenate all low-dimensional feature vectors, named as DR-Cat, or the vice versa, named as Cat-DR. However, DR-Cat ignores the correlation information between different features which is useful for classification. In Cat-DR, on the other hand, the correlation information estimated from the training data may not be reliable especially when the number of training samples is limited. We propose a covariance matrix regularization (CMR) technique to solve problems of DR-Cat and Cat-DR. It works by assigning weights to cross-feature covariances in the covariance matrix of training data . Thus the feature correlation estimated from training data is regularized before being used to train the feature fusion model. The proposed CMR is applied to 4 feature fusion schemes: fusion of pixel values from 3 color channels, fusion of LBP features from 3 color channels, fusion of pixel values and LBP features from a single color channel, and fusion of CNN features extracted by 2 deep models. Extensive experiments of face recognition and verification are conducted on databases including MultiPIE, Georgia Tech, AR and LFW. Results show that the proposed CMR technique significantly and consistently outperforms the best single feature , DR-Cat and Cat-DR.
Ze Lu, Xudong Jiang 0001, Alex Chichung Kot
Signal Process.2
2018 Off-Feature Information Incorporated Metric Learning for Face Recognition
abstract
Distance metric learning suppresses the intraclass variation while preserving the inter-class variation between two feature vectors. However, these two types of information are mixed in the feature vectors that need to be separated based on learning from the training data. The limited training data may not be able to well separate these two types of information and hence limits the effectiveness of metric learning. This letter proposes to exploit off-feature information to help suppress the intraclass variation of the feature vectors. For face recognition, some identity-independent information such as pose, expression, and occlusion is extracted from source images and utilized as the off-feature information to enhance the performance of distance metric learning. In training, the algorithm learns how to incorporate the off-feature information to suppress the intraclass variation of features. In recognition, the similarity score of a face image pair is determined by its distance in feature space and that of off-feature space. Extensive experiments demonstrate that the proposed off-feature information incorporated metric learning is helpful to suppress the intraclass variation of feature vectors, which visibly enhances the existing metric learning algorithms.
Renjie Huang, Xudong Jiang 0001
IEEE Signal Process. Lett.2
2018 Deep Coupled ResNet for Low-Resolution Face Recognition
abstract
Face images captured by surveillance cameras are often of low resolution (LR), which adversely affects the performance of their matching with high-resolution (HR) gallery images. Existing methods including super resolution, coupled mappings (CMs), multidimensional scaling, and convolutional neural network yield only modest performance. In this letter, we propose the deep coupled ResNet (DCR) model. It consists of one trunk network and two branch networks. The trunk network, trained by face images of three significantly different resolutions, is used to extract discriminative features robust to the resolution change. Two branch networks, trained by HR images and images of the targeted LR, work as resolution-specific CMs to transform HR and corresponding LR features to a space where their difference is minimized. Model parameters of branch networks are optimized using our proposed CM loss function, which considers not only the discriminability of HR and LR features, but also the similarity between them. In order to deal with various possible resolutions of probe images, we train multiple pairs of small branch networks while using the same trunk network. Thorough evaluation on LFW and SCface databases shows that the proposed DCR model achieves consistently and considerably better performance than the state of the arts.
Ze Lu, Xudong Jiang 0001, Alex Chichung Kot
IEEE Signal Process. Lett.2
2018 Feature Weighting and Regularization of Common Spatial Patterns in EEG-Based Motor Imagery BCI
abstract
Electroencephalography signals have very low spatial resolution and electrodes capture signals that are overlapping each other. To extract the discriminative features and alleviate overfitting problem for motor imagery brain-computer interface (BCI), spatial filtering is widely applied but often only very few common spatial patterns (CSP) are selected as features while ignoring all others. However, using only few CSP features, though alleviates overfitting problem, loses the discriminating information, which limits the BCI performance. This letter proposes a novel feature weighting and regularization (FWR) method that utilizes all CSP features to avoid information loss. The proposed method can be applied in all CSP-based approaches. Experiments of this letter show the effect of the proposed method applied in the standard CSP and its two extensions, common spatio-spectral patterns and regularized CSP. Results on BCI Competition III Dataset IIIa and IV Dataset IIa demonstrate that the proposed FWR method enhances the classification accuracy comparing to the conventional feature selection approaches.
Vasilisa Mishuhina, Xudong Jiang 0001
IEEE Signal Process. Lett.2
2017 A novel LBP-based Color descriptor for face recognition
abstract
LBP-based color features have shown excellent performance for color face recognition tasks, such as Color LBP, CLBP and LCVBP. However, existing methods encode the inter-channel information on pairs of color channels by applying the same spatial structure as that used in the intra-channel encoding. This results in a very high dimensional feature vector yet ineffective in encoding inter-channel information. Moreover, the difference of pixel values across color channels may not be a proper measure if they are not quantitatively comparable. To tackle these problems of existing methods, we propose a novel LBP-based color feature, Ternary-Color LBP (TCLBP), to encode the inter-channel information more effectively and efficiently. Extensive experiments on 4 public face databases, Color FERET, Georgia Tech, FRGC and LFW, are conducted to verify the effectiveness of the proposed TCLBP color feature for face recognition. Results show that the proposed TCLBP leads to visibly better face recognition performance than Color LBP, CLBP and LCVBP consistently over the 4 databases.
Ze Lu, Xudong Jiang 0001, Alex Chichung Kot
ICASSP2
2017 Laplace gradient based Discriminative and Contrast Invertible descriptor
abstract
The performance of local descriptors such as SIFT drops under severe illumination changes. In this paper, we propose a Discriminative and Contrast Invertible (DCI) local feature descriptor. In order to increase the discriminative ability of the descriptor under illumination changes, a Laplace gradient based histogram is proposed. Moreover, a robust contrast flipping estimate is proposed based on the divergence of a local region. Experiments on fine-grained object recognition and retrieval applications demonstrate the superior performance of the DCI descriptor to others.
Zhenwei Miao, Kim-Hui Yap, Xudong Jiang 0001, Subbhuraam Sinduja
ICASSP3
2017 Enhancing retinal vessel segmentation by color fusion
abstract
Accurate segmentation of retinal vessel plays an important role in the computer-aided diagnosis of eye diseases. Existing supervised methods extract features only from green channel due to its much higher contrast between vessel and background than in red and blue channels. However, red and blue channels also contain useful information for distinguishing vessel from background. This work investigates various ways of combining information in all 3 color channels to enhance the segmentation performance, based on which an effective color fusion scheme is proposed in this paper. Its performance is evaluated on two publicly available databases DRIVE and STARE. Results demonstrate that the proposed feature fusion with dimensionality reduction by asymmetric PCA visibly enhances the segmentation performance consistently on both databases, rendering better performance than state-of-the-art methods in dealing with healthy and pathological retinal images.
Xiaohong Wang 0003, Xudong Jiang 0001
ICASSP2
2017 Multi-modal and multi-layout discriminative learning for placental maturity staging
Bai Ying Lei, Wanjun Li, Yuan Yao 0007, Xudong Jiang 0001, Ee-Leng Tan, Harry Qin, Siping Chen, Dong Ni 0001, Tianfu Wang 0001
Pattern Recognit.4
2017 Regularized 2-D complex-log spectral analysis and subspace reliability analysis of micro-Doppler signature for UAV detection
Jianfeng Ren, Xudong Jiang 0001
Pattern Recognit.2
2017 LBP-Structure Optimization With Symmetry and Uniformity Regularizations for Scene Classification
abstract
Local binary pattern (LBP) and its variants have been widely used in many visual recognition tasks. Most existing approaches utilize predefined LBP structures to extract LBP features. Recently, data-driven LBP structures have shown promising results. However, due to the limited number of training samples, data-driven structures may overfit the training samples, hence could not generalize well on the novel testing samples. To address this problem, we propose two structural regularization constraints for LBP-structure optimization: symmetry constraint and uniformity constraint. These two constraints are inspired by predefined LBP structures, which convey the human prior knowledge on designing LBP structures. The LBP-structure optimization is casted as a binary quadratic programming problem and solved efficiently via the branch-and-bound algorithm. The evaluation on two scene-classification datasets demonstrates the superior performance of the proposed approach compared with both predefined LBP structures and unconstrained data-driven LBP structures.
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001
IEEE Signal Process. Lett.2
2017 Discovering Class-Specific Spatial Layouts for Scene Recognition
abstract
Scene image is a spatial composition of objects and background contexts and finding discriminative spatial layouts is critical for scene recognition. In this letter, we propose an ℓ1-regularized max-margin formulation to discover class-specific spatial layouts by jointly learning the image classifier and the class-specific spatial layouts for scene recognition. Unlike previous methods that classify images into different categories either without considering the spatial layouts explicitly or only using class generic spatial layout, our proposed method can discover a sparse combination of class-specific spatial layouts for different scenes and boost the recognition performance. Experiments on scene-15, landuse-21, and MIT indoor-67 datasets validate the advantages of our proposed algorithm.
Chaoqun Weng, Hongxing Wang 0001, Junsong Yuan 0001, Xudong Jiang 0001
IEEE Signal Process. Lett.4
2017 Accurate Cervical Cell Segmentation from Overlapping Clumps in Pap Smear Images
abstract
Accurate segmentation of cervical cells in Pap smear images is an important step in automatic pre-cancer identification in the uterine cervix. One of the major segmentation challenges is overlapping of cytoplasm, which has not been well-addressed in previous studies. To tackle the overlapping issue, this paper proposes a learning-based method with robust shape priors to segment individual cell in Pap smear images to support automatic monitoring of changes in cells, which is a vital prerequisite of early detection of cervical cancer. We define this splitting problem as a discrete labeling task for multiple cells with a suitable cost function. The labeling results are then fed into our dynamic multi-template deformation model for further boundary refinement. Multi-scale deep convolutional networks are adopted to learn the diverse cell appearance features. We also incorporated high-level shape information to guide segmentation where cell boundary might be weak or lost due to cell overlapping. An evaluation carried out using two different datasets demonstrates the superiority of our proposed method over the state-of-the-art methods in terms of segmentation accuracy.
Youyi Song, Ee-Leng Tan, Xudong Jiang 0001, Jie-Zhi Cheng, Dong Ni 0001, Siping Chen, Bai Ying Lei, Tianfu Wang 0001
IEEE Trans. Medical Imaging3
2017 Sound-Event Classification Using Robust Texture Features for Robot Hearing
abstract
Sound-event classification often utilizes time-frequency analysis, which produces an image-like spectrogram. Recent approaches such as spectrogram image features and subband power distribution image features extract the image local statistics such as mean and variance from the spectrogram. They have demonstrated good performance. However, we argue that such simple image statistics cannot well capture the complex texture details of the spectrogram. Thus, we propose to extract the local binary pattern (LBP) from the logarithm of the Gammatone-like spectrogram. However, the LBP feature is sensitive to noise. After analyzing the spectrograms of sound events and the audio noise, we find that the magnitude of pixel differences, which is discarded by the LBP feature, carries important information for sound-event classification. We thus propose a multichannel LBP feature via pixel difference quantization to improve the robustness to the audio noise. In view of the differences between spectrograms and natural images, and the reliability issues of LBP features, we propose two projection-based LBP features to better capture the texture information of the spectrogram. To validate the proposed multichannel projection-based LBP features for robot hearing, we have built a new sound-event classification database, the NTU-SEC database, in the context of social interaction between human and robot. It is publicly available to promote research on sound-event classification in a social context. The proposed approaches are compared with the state of the art on the RWCP database and the NTU-SEC database. They consistently demonstrate superior performance under various noise conditions.
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001, Nadia Magnenat-Thalmann
IEEE Trans. Multim.2
2016 Shadow detection using double-threshold pulse coupled neural networks
abstract
A novel double-threshold pulse coupled neural networks (DT-PCNN) is proposed and applied to shadow detection. It attempts to reduce the false detection of shadows in a single image where the hue and brightness of some non-shadow regions are similar to or even lower than those of shadows. Shadows whose intensity and hue fall in between those of the scene and objectives are often viewed as non-shadows by the single dynamic threshold of PCNN. Moreover, entities with similar or darker hue and intensity may be wrongly classified as shadows. To solve this problem, two different dynamic thresholds that iteratively alter are designed. The upper and lower limits of detecting shadows are determined respectively by a higher threshold that decreases iteratively and a lower one that increases iteratively. The detection result is obtained by a fusion of two detection components. Experimental results demonstrate that compared to other tested methods, the misclassifications are significantly reduced and the shadows are more accurately extracted.
Xudong Jiang 0001
ICASSP2
2016 An effective color space for face recognition
abstract
The three color components specifying a color can be defined in various ways leading to significantly different classification abilities. Several effective color spaces including RQCr, DCS and ZRG have been proposed to achieve better face recognition performance. However, their performance is not consistent on different databases. What's more, the framework of effective color spaces has not been thoroughly studied yet. In this paper, we propose an effective color space LC\C2 based on a framework of effective color spaces. LC\C2 consists of one discriminant luminance component L and two discriminant chrominance components C\C2. To find the discriminant luminance component, 4 luminance components from existing effective color models are compared. After that, the weighted color space normalization technique (WCSN) is applied on the DCS color space to generate two complementary and discriminative chrominance components. Experiments conducted on three databases (FRGC, AR and CMU Multi-PIE) show that the proposed color space LC\C2 achieves the best face recognition performance consistently.
Ze Lu, Xudong Jiang 0001, Alex Chichung Kot
ICASSP2
2016 Human Body Part Selection by Group Lasso of Motion for Model-Free Gait Recognition
abstract
Gait recognition is an emerging biometric technology that identifies people through the analysis of the way they walk. The challenge of model-free based gait recognition is to cope with various intra-class variations such as clothing variations, carrying conditions and angle variations that adversely affect the recognition performance. This paper proposes a method to select the most discriminative human body part based on group Lasso of motion to reduce the intra-class variation so as to improve the recognition performance. The proposed method is evaluated using CASIA Gait Dataset B. Experimental results demonstrate that the proposed technique gives promising results.
Imad Rida, Xudong Jiang 0001, Gian Luca Marcialis
IEEE Signal Process. Lett.2
2016 Classwise Sparse and Collaborative Patch Representation for Face Recognition
abstract
Sparse representation has shown its merits in solving some classification problems and delivered some impressive results in face recognition. However, the unsupervised optimization of the sparse representation may result in undesired classification outcome if the variations of the data population are not well represented by the training samples. In this paper, a method of class-wise sparse representation (CSR) is proposed to tackle the problems of the conventional sample-wise sparse representation and applied to face recognition. It seeks an optimum representation of the query image by minimizing the class-wise sparsity of the training data. To tackle the problem of the uncontrolled training data, this paper further proposes a collaborative patch (CP) framework, together with the proposed CSR, named CSR-CP. Different from the conventional patch-based methods that optimize each patch representation separately, the CSR-CP approach optimizes all patches together to seek a CP groupwise sparse representation by putting all patches of an image into a group. It alleviates the problem of losing discriminative information in the training data caused by the partition of the image into patches. Extensive experiments on several benchmark face databases demonstrate that the proposed CSR-CP significantly outperforms the sparse representation-related holistic and patch-based approaches.
Jian Lai, Xudong Jiang 0001
IEEE Trans. Image Process.2
2016 Contrast Invariant Interest Point Detection by Zero-Norm LoG Filter
abstract
The Laplacian of Gaussian (LoG) filter is widely used in interest point detection. However, low-contrast image structures, though stable and significant, are often submerged by the high-contrast ones in the response image of the LoG filter, and hence are difficult to be detected. To solve this problem, we derive a generalized LoG filter, and propose a zero-norm LoG filter. The response of the zero-norm LoG filter is proportional to the weighted number of bright/dark pixels in a local region, which makes this filter be invariant to the image contrast. Based on the zero-norm LoG filter, we develop an interest point detector to extract local structures from images. Compared with the contrast dependent detectors, such as the popular scale invariant feature transform detector, the proposed detector is robust to illumination changes and abrupt variations of images. Experiments on benchmark databases demonstrate the superior performance of the proposed zero-norm LoG detector in terms of the repeatability and matching score of the detected points as well as the image recognition rate under different conditions.
Zhenwei Miao, Xudong Jiang 0001, Kim-Hui Yap
IEEE Trans. Image Process.2
2015 Quantized fuzzy LBP for face recognition
abstract
Face recognition under large illumination variations is challenging. Local binary pattern (LBP) is robust to illumination variation, but sensitive to noise. Fuzzy LBP (FLBP) partially solves the noise-sensitivity problem by incorporating fuzzy logic in the representation of local binary patterns. The fuzzy membership function is determined by both sign and magnitude of the pixel difference. However, the magnitude is easily altered by noise, hence could be unreliable. Thus, we propose to determine the fuzzy membership function by its sign only. We name the proposed approach as Quantized Fuzzy LBP (QFLBP). On two challenging face recognition datasets, it is shown more robust to noise, and demonstrates a superior performance to FLBP and many other LBP variants.
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001
ICASSP2
2015 Sparse and Dense Hybrid Representation via Dictionary Decomposition for Face Recognition
abstract
Sparse representation provides an effective tool for classification under the conditions that every class has sufficient representative training samples and the training data are uncorrupted. These conditions may not hold true in many practical applications. Face identification is an example where we have a large number of identities but sufficient representative and uncorrupted training images cannot be guaranteed for every identity. A violation of the two conditions leads to a poor performance of the sparse representation-based classification (SRC). This paper addresses this critic issue by analyzing the merits and limitations of SRC. A sparse- and dense-hybrid representation (SDR) framework is proposed in this paper to alleviate the problems of SRC. We further propose a procedure of supervised low-rank (SLR) dictionary decomposition to facilitate the proposed SDR framework. In addition, the problem of the corrupted training data is also alleviated by the proposed SLR dictionary decomposition. The application of the proposed SDR-SLR approach in face recognition verifies its effectiveness and advancement to the field. Extensive experiments on benchmark face databases demonstrate that it consistently outperforms the state-of-the-art sparse representation based approaches and the performance gains are significant in most cases.
Xudong Jiang 0001, Jian Lai
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Learning LBP structure by maximizing the conditional mutual information
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001
Pattern Recognit.2
2015 A Color Channel Fusion Approach for Face Recognition
abstract
Due to high dimensionality of images or generated color features, different color channels are usually processed separately and then concatenated together into a feature vector for classification. This makes channel fusion a crucial step in color face recognition (FR) systems. However, existing methods simply concatenate channel-wise color features without identifying the importance or reliability of features in different color channels. In this paper, we propose a color channel fusion (CCF) approach using jointly dimension reduction algorithms to select more features from reliable and discriminative channels. Experiments using two different dimension reduction approaches, two different types of features on three image datasets show that CCF achieves consistently better performance than color channel concatenation (CCC) method which deals with different color channels equally.
Ze Lu, Xudong Jiang 0001, Alex Chichung Kot
IEEE Signal Process. Lett.2
2015 LBP Encoding Schemes Jointly Utilizing the Information of Current Bit and Other LBP Bits
abstract
Local binary pattern (LBP) is sensitive to image noise. Noise-resistant LBP (NRLBP) improves the robustness to noise by incorporating the prior knowledge of images and information of other LBP bits into encoding process. However, it encodes the small pixel difference in such a way that its sign and magnitude are ignored. Although the small pixel difference may be easily distorted by noise, some of its information is still useful for LBP encoding. In this letter, we propose two enhanced NRLBPs that jointly utilize the sign and the magnitude of the current pixel difference, and also the information of other LBP bits. The proposed approaches are validated on two benchmark databases and demonstrate a superior performance compared with NRLBP and other LBP variants. The performance gain is significant when the noise level is high.
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001
IEEE Signal Process. Lett.2
2015 A Chi-Squared-Transformed Subspace of LBP Histogram for Visual Recognition
abstract
Local binary pattern (LBP) and its variants have been widely used in many recognition tasks. Subspace approaches are often applied to the LBP feature in order to remove unreliable dimensions, or to derive a compact feature representation. It is well-known that subspace approaches utilizing up to the second-order statistics are optimal only when the underlying distribution is Gaussian. However, due to its nonnegative and simplex constraints, the LBP feature deviates significantly from Gaussian distribution. To alleviate this problem, we propose a chi-squared transformation (CST) to transfer the LBP feature to a feature that fits better to Gaussian distribution. The proposed CST leads to the formulation of a two-class classification problem. Due to its asymmetric nature, we apply asymmetric principal component analysis (APCA) to better remove the unreliable dimensions in the CST feature space. The proposed CST-APCA is evaluated extensively on spatial LBP for face recognition, protein cellular classification, and spatial-temporal LBP for dynamic texture recognition. All experiments show that the proposed feature transformation significantly enhances the recognition accuracy.
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001
IEEE Trans. Image Process.2
2014 Learning Discriminative and Shareable Features for Scene Classification
Zhen Zuo, Gang Wang 0012, Bing Shuai, Lifan Zhao, Qingxiong Yang, Xudong Jiang 0001
ECCV (1)6
2014 Supervised trace lasso for robust face recognition
abstract
In this paper, we address the robust face recognition problem. Recently, trace lasso was introduced as an adaptive norm based on the training data. It uses the correlation among the training samples to tackle the instability problem of sparse representation coding. Trace lasso naturally clusters the highly correlated data together. However, the face images with similar variations, such as illumination or expression, often have higher correlation than those from the same class. In this case, the result of trace lasso is contradictory to the goal of recognition, which is to cluster the samples according to their identities. Therefore, trace lasso is not a good choice for face recognition task. In this work, we propose a supervised trace lasso (STL) framework by employing the class label information. To represent the query sample, the proposed STL approach seeks the sparsity of the number of classes instead of the number of training samples. This directly coincides with the objective of the classification. Furthermore, an efficient algorithm to solve the optimization problem of proposed method is given. The extensive experimental results have demonstrated the effectiveness of the proposed framework.
Jian Lai, Xudong Jiang 0001
ICME2
2014 Additive and exclusive noise suppression by iterative trimmed and truncated mean algorithm
Zhenwei Miao, Xudong Jiang 0001
Signal Process.2
2014 Optimizing LBP Structure For Visual Recognition Using Binary Quadratic Programming
abstract
Local binary pattern (LBP) and its variants have shown promising results in visual recognition applications. However, most existing approaches rely on a pre-defined structure to extract LBP features. We argue that the optimal LBP structure should be task-dependent and propose a new method to learn discriminative LBP structures. We formulate it as a point selection problem: Given a set of point candidates, the goal is to select an optimal subset to compose the LBP structure. In view of the problems of current feature selection algorithms, we propose a novel Maximal Joint Mutual Information criterion. Then, the point selection is converted into a binary quadratic programming problem and solved efficiently via the branch and bound algorithm. The proposed LBP structures demonstrate superior performance to the state-of-the-art approaches on classifying both spatial patterns in scene recognition and spatial-temporal patterns in dynamic texture recognition.
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001, Gang Wang 0012
IEEE Signal Process. Lett.2
2014 Human Detection by Quadratic Classification on Subspace of Extended Histogram of Gradients
abstract
This paper proposes a quadratic classification approach on the subspace of Extended Histogram of Gradients (ExHoG) for human detection. By investigating the limitations of Histogram of Gradients (HG) and Histogram of Oriented Gradients (HOG), ExHoG is proposed as a new feature for human detection. ExHoG alleviates the problem of discrimination between a dark object against a bright background and vice versa inherent in HG. It also resolves an issue of HOG whereby gradients of opposite directions in the same cell are mapped into the same histogram bin. We reduce the dimensionality of ExHoG using Asymmetric Principal Component Analysis (APCA) for improved quadratic classification. APCA also addresses the asymmetry issue in training sets of human detection where there are much fewer human samples than non-human samples. Our proposed approach is tested on three established benchmarking data sets--INRIA, Caltech, and Daimler--using a modified Minimum Mahalanobis distance classifier. Results indicate that the proposed approach outperforms current state-of-the-art human detection methods.
Amit Satpathy, Xudong Jiang 0001, How-Lung Eng
IEEE Trans. Image Process.2
2014 LBP-Based Edge-Texture Features for Object Recognition
abstract
This paper proposes two sets of novel edge-texture features, Discriminative Robust Local Binary Pattern (DRLBP) and Ternary Pattern (DRLTP), for object recognition. By investigating the limitations of Local Binary Pattern (LBP), Local Ternary Pattern (LTP) and Robust LBP (RLBP), DRLBP and DRLTP are proposed as new features. They solve the problem of discrimination between a bright object against a dark background and vice-versa inherent in LBP and LTP. DRLBP also resolves the problem of RLBP whereby LBP codes and their complements in the same block are mapped to the same code. Furthermore, the proposed features retain contrast information necessary for proper representation of object contours that LBP, LTP, and RLBP discard. Our proposed features are tested on seven challenging data sets: INRIA Human, Caltech Pedestrian, UIUC Car, Caltech 101, Caltech 256, Brodatz, and KTH-TIPS2-a. Results demonstrate that the proposed features outperform the compared approaches on most data sets.
Amit Satpathy, Xudong Jiang 0001, How-Lung Eng
IEEE Trans. Image Process.2
2013 Robust face recognition using trimmed linear regression
abstract
In this work, we focus on the problem of partially occluded face recognition. Using a robust estimator, we detect and trim the contaminated pixels from query sample. The corresponding pixels in the training samples are trimmed as well. The linear regression is applied to the trimmed images. Finally, the query image is labeled to the class with minimum normalized reconstruction error. Extensive experiments on benchmark face datasets demonstrate that the proposed approach is much more robust than state-of-the-art methods in dealing with occluded faces.
Jian Lai, Xudong Jiang 0001
ICASSP2
2013 A vote of confidence based interest point detector
abstract
In this paper, a vote of confidence (VC) based detector is proposed to detect bright and dark regions from images. Whether a local region is bright or dark is voted by all the pixels in this region. Compared to the contrast based detectors, such as the popular SIFT detector, the VC detector is invariant to illumination change and robust to abrupt variations. Experiments are conducted on benchmark databases to verify the superior performance of the VC detector in terms of the repeatability and matching score. The proposed detector is also evaluated in the application of face recognition.
Zhenwei Miao, Xudong Jiang 0001
ICASSP2
2013 Dynamic texture recognition using enhanced LBP features
abstract
This paper addresses the challenge of recognizing dynamic textures based on spatial-temporal descriptors. Dynamic textures are composed of both spatial and temporal features. The histogram of local binary pattern (LBP) has been used in dynamic texture recognition. However, its performance is limited by the reliability issues of the LBP histograms. In this paper, two learning-based approaches are proposed to remove the unreliable information in LBP features by utilizing Principal Histogram Analysis. Furthermore, a super histogram is proposed to improve the reliability of the LBP histograms. The temporal information is partially transferred to the super histogram. The proposed approaches are evaluated on two widely used benchmark databases: UCLA and Dyntex++ databases. Superior performance is demonstrated compared with the state of the arts.
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001
ICASSP2
2013 Human detection using Discriminative and Robust Local Binary Pattern
abstract
Despite superior performance of Local Binary Pattern (LBP) in texture classification and face detection, its performance in human detection has been limited for two reasons. Firstly, LBP differentiates a bright human from a dark background and vice-versa. This increases the intra-class variation of humans. Secondly, LBP is contrast and illumination invariant. It does not discriminate between weak contrast local regions and similar strong contrast ones, resulting in a similar feature representation. Non-Redundant LBP (NRLBP) has been proposed to solve the first issue of LBP. However, an inherent limitation of NRLBP is that LBP codes and their complements in the same block are mapped to the same code. Furthermore, NRLBP, like LBP, is also contrast and illumination invariant. In this paper, we propose a novel edge-texture feature, Discriminative Robust Local Binary Pattern (DRLBP), for human detection. DRLBP alleviates the problems of LBP and NRLBP by considering the weighted sum and absolute difference of a LBP code and its complement. Our experimental results show that DRLBP consistently outperforms LBP and NRLBP for human detection.
Amit Satpathy, Xudong Jiang 0001, How-Lung Eng
ICASSP2
2013 Discriminative sparsity preserving embedding for face recognition
abstract
Over the past few years, sparse representation (SR) becomes a hotspot and applied in many research fields. Sparsity preserving projections (SPP) utilizes SR to dimensionality reduction (DR) for face classification. However, as the original framework of SR is unsupervised, SPP can not employ the class information, which is very crucial for classification. To address this problem, we propose an algorithm, namely supervised SR (SSR), to cooperate with label information. Furthermore, we also propose a DR method, discriminative sparsity preserving embedding (DSPE), in this paper. DSPE learns the discriminative sparse structure with SSR and finds the low dimensional subspace that reduces the within class distances and keeps the between class distances. Compared with the related state-of-the-art methods, experimental results on benchmark face databases verify the advancement of the proposed method.
Jian Lai, Xudong Jiang 0001
ICIP2
2013 Learning binarized pixel-difference pattern for scene recognition
abstract
Local binary pattern (LBP) and its variants have been used in scene recognition. However, most existing approaches rely on a pre-defined LBP structure to extract features. Those pre-defined structures can be generalized as the patterns constructed from the binarized pixel differences in a local neighborhood. Instead of using a handcraft structure, we propose to learn binarized pixel-difference patterns (BPP). We cast the problem as a feature selection problem and solve it by an incremental search via the criterion of minimum-redundancy-maximum-relevance. Then, BPP features are extracted based on the structures derived. On two challenging scene recognition databases, the proposed approach significantly outperforms the state of the arts.
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001
ICIP2
2013 Relaxed local ternary pattern for face recognition
abstract
Local binary pattern (LBP) is sensitive to noise. Local ternary pattern (LTP) partially solves this problem by encoding the small pixel difference into a third state. The small pixel difference may be easily overwhelmed by noise. Thus, it is difficult to precisely determine its sign and magnitude. In this paper, we propose the concept of uncertain state to encode the small pixel difference. We do not care its sign and magnitude, and encode it as both 0 and 1 with equal probability. The proposed Relaxed LTP is tested on the CMU-PIE database, the extended Yale B database and the O2FN mobile face database. Superior performance is demonstrated compared with LBP and LTP.
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001
ICIP2
2013 Visual object detection by parts-based modeling using extended histogram of gradients
abstract
In this paper, we present a parts-based modeling framework using Extended Histogram of Gradients (ExHoG) for object detection. Visual object detection is a challenging issue in computer vision where objects need to be detected in varying illumination and contrast environments. Furthermore, objects belonging to the same class exhibit large intra-class variations. Here, we propose using ExHoG with the discriminatively trained deformable part models of Felzenszwalb et. al. [1]. This framework is based on mixtures of multiscale deformable part models. ExHoG is a novel feature proposed earlier for the purpose of human detection and has shown promising results against other state-of-the-art approaches. The proposed approach is tested on INRIA Human data set and the PASCAL VOC 2007 data set. Results demonstrate superior performance on INRIA compared to existing state-of-the-art approaches and improved performance on PASCAL VOC 2007.
Amit Satpathy, Xudong Jiang 0001, How-Lung Eng
ICIP2
2013 Fully automatic face recognition framework based on local and global features
Cong Geng, Xudong Jiang 0001
Mach. Vis. Appl.2
2013 Interest point detection using rank order LoG filter
Zhenwei Miao, Xudong Jiang 0001
Pattern Recognit.2
2013 A complete and fully automated face verification system on mobile devices
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001
Pattern Recognit.2
2013 Noise-Resistant Local Binary Pattern With an Embedded Error-Correction Mechanism
abstract
Local binary pattern (LBP) is sensitive to noise. Local ternary pattern (LTP) partially solves this problem. Both LBP and LTP, however, treat the corrupted image patterns as they are. In view of this, we propose a noise-resistant LBP (NRLBP) to preserve the image local structures in presence of noise. The small pixel difference is vulnerable to noise. Thus, we encode it as an uncertain state first, and then determine its value based on the other bits of the LBP code. It is widely accepted that most of the image local structures are represented by uniform codes and noise patterns most likely fall into the non-uniform codes. Therefore, we assign the value of an uncertain bit hence as to form possible uniform codes. Thus, we develop an error-correction mechanism to recover the distorted image patterns. In addition, we find that some image patterns such as lines are not captured in uniform codes. Those line patterns may appear less frequently than uniform codes, but they represent a set of important local primitives for pattern recognition. Thus, we propose an extended noise-resistant LBP (ENRLBP) to capture line patterns. The proposed NRLBP and ENRLBP are more resistant to noise compared with LBP, LTP, and many other variants. On various applications, the proposed NRLBP and ENRLBP demonstrate superior performance to LBP/LTP variants.
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001
IEEE Trans. Image Process.2
2012 Face alignment based on the multi-scale local features
abstract
Many face recognition algorithms depend on careful positioning of face images into the same canonical pose. Currently, this positioning is usually done by detecting the locations of eyes. And the face images are transformed to the same positions according to the eye coordinates detected. In this paper, we describe a method based on multi-scale local features to achieve face alignment automatically not just dependent on the localizations of two eyes. Given an unaligned face image resulting from a face detector and a set of aligned face images in the data set, we build an automatic transformation mechanism, under which the unaligned face image can be precisely aligned for the following recognition process. Our alignment method improves performance on face recognition tasks, over images aligned by many other algorithms.
Cong Geng, Xudong Jiang 0001
ICASSP2
2012 A novel rank order LoG filter for interest point detection
abstract
This paper proposes a novel non-linear filter, named rank order LoG (ROLG) filter, and a new interest point detector, named ROLG detector. The ROLG filter is a weighted rank order filter. It is used to detect image structures whose significant majority of pixels are brighter (or darker) than the significant majority of pixels in their corresponding surroundings. The ROLG detector is built on this filter. Compared to linear filter based detectors, the proposed rank order filter based detector is more robust to abrupt variations of images. Experiments on the benchmark databases demonstrate that the ROLG detector achieves superior performance compared to four state-of-the-art detectors. Evaluation experiments are also conducted on face recognition. The results further demonstrate that the ROLG detector has better performance compared to other detectors.
Zhenwei Miao, Xudong Jiang 0001
ICASSP2
2012 Modular Weighted Global Sparse Representation for Robust Face Recognition
abstract
This work proposes a novel framework of robust face recognition based on the sparse representation. Image is first divided into modules and each module is processed separately to determine its reliability. A reconstructed image from the modules weighted by their reliability is formed for the robust recognition. We propose to use the modular sparsity and residual jointly to determine the modular reliability. The proposed framework advances both the modular and global sparse representation approaches, especially in dealing with disguise, large illumination variations and expression changes. Compared with the related state-of-the-art methods, experimental results on benchmark face databases verify the advancement of the proposed method.
Jian Lai, Xudong Jiang 0001
IEEE Signal Process. Lett.2
2012 Iterative Truncated Arithmetic Mean Filter and Its Properties
abstract
The arithmetic mean and the order statistical median are two fundamental operations in signal and image processing. They have their own merits and limitations in noise attenuation and image structure preservation. This paper proposes an iterative algorithm that truncates the extreme values of samples in the filter window to a dynamic threshold. The resulting nonlinear filter shows some merits of both the fundamental operations. Some dynamic truncation thresholds are proposed that guarantee the filter output, starting from the mean, to approach the median of the input samples. As a by-product, this paper unveils some statistics of a finite data set as the upper bounds of the deviation of the median from the mean. Some stopping criteria are suggested to facilitate edge preservation and noise attenuation for both the long- and short-tailed distributions. Although the proposed iterative truncated mean (ITM) algorithm is not aimed at the median, it offers a way to estimate the median by simple arithmetic computing. Some properties of the ITM filters are analyzed and experimentally verified on synthetic data and real images.
Xudong Jiang 0001
IEEE Trans. Image Process.1
2012 Binarization of Low-Quality Barcode Images Captured by Mobile Phones Using Local Window of Adaptive Location and Size
abstract
It is difficult to directly apply existing binarization approaches to the barcode images captured by mobile device due to their low quality. This paper proposes a novel scheme for the binarization of such images. The barcode and background regions are differentiated by the number of edge pixels in a search window. Unlike existing approaches that center the pixel to be binarized with a window of fixed size, we propose to shift the window center to the nearest edge pixel so that the balance of the number of object and background pixels can be achieved. The window size is adaptive either to the minimum distance to edges or minimum element width in the barcode. The threshold is calculated using the statistics in the window. Our proposed method has demonstrated its capability in handling the nonuniform illumination problem and the size variation of objects. Experimental results conducted on 350 images captured by five mobile phones achieve about 100% of recognition rate in good lighting conditions, and about 95% and 83% in bad lighting conditions. Comparisons made with nine existing binarization methods demonstrate the advancement of our proposed scheme.
Huijuan Yang, Alex Chichung Kot, Xudong Jiang 0001
IEEE Trans. Image Process.3
2011 Face recognition based on the multi-scale local image structures
Cong Geng, Xudong Jiang 0001
Pattern Recognit.2
2010 Knowledge guided adaptive binarization for 2D barcode images captured by mobile phones
abstract
In this paper, we investigate the problem of binarization of the 2D barcode images captured by mobile devices. The poor quality of the images due to noise, nonuniform illumination and inherent limitation and distortion of the camera makes the task of binarization more challenging. The global thresholding techniques employed by most existing approaches do not work well for the barcode images captured under uncontrolled conditions. Hence, we propose an adaptive binarization technique by taking the prior knowledge such as the edge structure and maximum element width of the 2D barcode into consideration. Incrementally updating the elements in a moving window is proposed to expediate the process for mobile phone applications. Comparisons with other binarization methods showthe superior performance of our method, which achieves about 96.6% recognition rate on a dataset of 787 images.
Huijuan Yang, Alex Chichung Kot, Xudong Jiang 0001
ICASSP3
2010 Extended Histogram of Gradients feature for human detection
abstract
Unsigned Histogram of Gradients (UHoG) is a popular feature used for human detection. Despite its superior performance as reported in recent literature, an inherent limitation of UHoG is that gradients of opposite directions in a cell are mapped into the same histogram bin. This is undesirable as it will produce the same UHoG feature for two different patterns. To address this problem, we propose a new feature named the Extended Histogram of Gradients (ExHoG) in this paper. It comprises two components: UHoG and a histogram of absolute bin value differences of opposite gradient directions computed from Histogram of Gradients (HoG). Our experimental results show that the proposed ExHoG consistently outperforms the standard HoG and UHoG for human detection.
Amit Satpathy, Xudong Jiang 0001, How-Lung Eng
ICIP2
2010 Accurate localization of four extreme corners for barcode images captured by mobile phones
abstract
In this paper, we propose a novel method to locate the four extreme corners of barcodes in the images captured by mobile phones. To achieve this goal, the two nearly-parallel outer boundary lines are firstly localized by utilizing the prior knowledge of the relative distances and angles between the lines, which are subsequently employed to obtain the initially localized corners. A novel post-localization process based on edge tracing is also proposed to further validate the initially localized corners. This is achieved by setting the constraints in maximum direction change when tracing from the current to the next candidate corners. The distinctive feature of the proposed algorithm lies in the capability of handling those curved barcode images taken by the mobile phones. Experiments conducted on a data set of 1410 images show the accuracy of the corner localization is about 92.3% and 90.5% for indoor and outdoor barcode images, respectively. The processing time taken is about 0.8 second using Matlab.
Huijuan Yang, Xudong Jiang 0001, Alex Chichung Kot
ICIP2
2010 Dynamic window construction for the binarization of barcode images captured by mobile phones
abstract
It is difficult to directly apply existing binarization methods to the barcode images captured by mobile devices under uncontrolled lighting conditions. Employing a fix-sized window to locally binarize the image cannot handle the situation when small and large objects co-exist in an image. This paper proposes a novel scheme to dynamically determine the size of the binarization window, which changes with the presence of the high gradient pixels in a search window centering the processing pixel. The proposed method has demonstrated its capability in handling objects of different sizes and the uneven illumination problem. Further, it is not constrained to certain barcode types. Experimental results conducted on 330 images captured using five mobile phones in indoor and outdoor environments achieve about 93% recognition rate evaluated using a barcode decoder. Comparisons with some existing approaches show its superior binarization performance.
Huijuan Yang, Alex Chichung Kot, Xudong Jiang 0001
ICIP3
2010 Image watermarking using dual-tree complex wavelet by coefficients swapping and group of coefficients quantization
abstract
In this paper, we present two watermarking schemes for images using the low-pass frequency coefficients and high-pass complex frequency coefficients of the Dual-Tree Complex Wavelet Transform (DT-CWT) of the image, the distinctive characteristics of the two schemes can lead to different applications. We investigate the problem of embedding binary watermark sequence in DT-CWT domain, which is a challenging problem. The coefficients swapping is done by carefully selecting those 2×2 blocks with high-energy in the low frequency and swap the coefficients that lie in the median range in the block to embed the watermark. Whereas the group of coefficients quantization quantizes a group of high-pass complex frequency coefficients to make the quantized coefficients lie in the middle of the quantization range and distribute the changes among the coefficients. Experimental results conducted on 100 standard images achieve about 96% and 92% detection rate for the low-pass and high-pass frequency coefficients-based schemes, respectively. Robustness tests against some common signal processing attacks such as additive noise, median filtering and lossy JPEG compression confirm the superior performance of the proposed schemes.
Huijuan Yang, Xudong Jiang 0001, Alex Chichung Kot
ICME2
2010 Special issue on: Recent advances and future directions in biometrics personal identification
Muhammad Khurram Khan, Mohamed S. Kamel, Xudong Jiang 0001
J. Netw. Comput. Appl.3
2010 Two-Dimensional Polar Harmonic Transforms for Invariant Image Representation
abstract
This paper introduces a set of 2D transforms, based on a set of orthogonal projection bases, to generate a set of features which are invariant to rotation. We call these transforms Polar Harmonic Transforms (PHTs). Unlike the well-known Zernike and pseudo-Zernike moments, the kernel computation of PHTs is extremely simple and has no numerical stability issue whatsoever. This implies that PHTs encompass the orthogonality and invariance advantages of Zernike and pseudo-Zernike moments, but are free from their inherent limitations. This also means that PHTs are well suited for application where maximal discriminant information is needed. Furthermore, PHTs make available a large set of features for further feature selection in the process of seeking for the best discriminative or representative features for a particular application.
Pew-Thian Yap, Xudong Jiang 0001, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Image quality assessment by discrete orthogonal moments
Chong-Yaw Wee, Raveendran Paramesran, Ramakrishnan Mukundan, Xudong Jiang 0001
Pattern Recognit.4
2010 Prediction of eigenvalues and regularization of eigenfeatures for human face verification
Bappaditya Mandal, Xudong Jiang 0001, How-Lung Eng, Alex Chichung Kot
Pattern Recognit. Lett.2
2009 Face recognition using sift features
abstract
Scale Invariant Feature Transform (SIFT) has shown to be a powerful technique for general object recognition/detection. In this paper, we propose two new approaches: Volume-SIFT (VSIFT) and Partial-Descriptor-SIFT (PDSIFT) for face recognition based on the original SIFT algorithm. We compare holistic approaches: Fisherface (FLDA), the null space approach (NLDA) and Eigenfeature Regularization and Extraction (ERE) with feature based approaches: SIFT and PDSIFT. Experiments on the ORL and AR databases show that the performance of PDSIFT is significantly better than the original SIFT approach. Moreover, PDSIFT can achieve comparable performance as the most successful holistic approach ERE and significantly outperforms FLDA and NLDA.
Cong Geng, Xudong Jiang 0001
ICIP2
2009 Fast eye localization based on pixel differences
abstract
A novel fast eye localization algorithm based on pixel differences is presented, which is suitable for face recognition system on mobile device. It is based on the fact that eyeball is dark and round. A binary eye map is obtained by choosing those pixels darker than surrounding; then it is filtered by a rank order filter; connected regions in the eye map are then labeled by their geometric centers; best suitable eyeball pair is selected based on a set of geometric constraints. If no eyeball pair is detected, the algorithm is repeated iteratively until one pair is found. The algorithm is fast since it converts the gray level image to a binary eye map at the beginning. The algorithm is tested on our own face database, which consists of 4095 images of size 250×200. Detection rate is 93.04% when the tolerance is 0.7 times of eyeball width.
Jianfeng Ren, Xudong Jiang 0001
ICIP2
2009 Complete discriminant evaluation and feature extraction in kernel space for face recognition
Xudong Jiang 0001, Bappaditya Mandal, Alex Chichung Kot
Mach. Vis. Appl.1
2009 Asymmetric Principal Component and Discriminant Analyses for Pattern Classification
abstract
This paper studies the roles of the principal component and discriminant analyses in the pattern classification and explores their problems with the asymmetric classes and/or the unbalanced training data. An asymmetric principal component analysis (APCA) is proposed to remove the unreliable dimensions more effectively than the conventional PCA. Targeted at the two-class problem, an asymmetric discriminant analysis in the APCA subspace is proposed to regularize the eigenvalue that is, in general, a biased estimate of the variance in the corresponding dimension. These efforts facilitate a reliable and discriminative feature extraction for the asymmetric classes and/or the unbalanced training data. The proposed approach is validated in the experiments by comparing it with the related methods. It consistently achieves the highest classification accuracy among all tested methods in the experiments.
Xudong Jiang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 A multi-prototype clustering algorithm
Manhua Liu, Xudong Jiang 0001, Alex Chichung Kot
Pattern Recognit.2
2008 Verification of human faces using predicted eigenvalues
abstract
To alleviate the conventional problems of LDA and its variants, we propose a procedure of predicting eigenvalues using few reliable eigenvalues from the range space. Partitioning of entire eigenspace is performed using two control points, however, the effective low dimensional discriminative vectors are extracted from the whole eigenspace. This prediction strategy enables to perform discriminant evaluation in the full eigenspace. The proposed method is evaluated and compared with 8 popular subspace based methods for face verification task. Experimental results on popular face databases show that our method consistently outperforms others.
Bappaditya Mandal, Xudong Jiang 0001, Alex Chichung Kot
ICPR2
2008 Boosted complex moments for discriminant rotation invariant object recognition
abstract
This paper proposes a method for constructing a discriminative rotation invariant object recognition system from the set of complex moments by using a multi-class boosting algorithm. Experimental results show that a large of number images can be discriminated accurately with only a small number of features. This basically means economy of computational effort in feature acquisition and also possibility of higher speed in recognition task.
Pew-Thian Yap, Xudong Jiang 0001, Alex Chichung Kot
ICPR2
2008 An interactive and secure user authentication scheme for mobile devices
abstract
Graphical password (i.e., image based authentication) is considered as a promising alternative to traditional textual password for mobile devices, to achieve better tradeoff between usability and security. However, previous proposals of graphical password have the limitation of limited entropy. In this paper, we propose a new scheme incorporating user face based authentication into the association-based graphical password solution we proposed before, aiming at achieving higher security without compromising user-friendliness for mobile application scenarios. System performance analysis and comparisons with other schemes are presented to validate our scheme.
Qibin Sun, Zhi Li 0001, Xudong Jiang 0001, Alex Chichung Kot
ISCAS3
2008 Eigenfeature Regularization and Extraction in Face Recognition
abstract
This work proposes a subspace approach that regularizes and extracts eigenfeatures from the face image. Eigenspace of the within-class scatter matrix is decomposed into three subspaces: a reliable subspace spanned mainly by the facial variation, an unstable subspace due to noise and finite number of training samples and a null subspace. Eigenfeatures are regularized differently in these three subspaces based on an eigenspectrum model to alleviate problems of instability, over-fitting or poor generalization. This also enables the discriminant evaluation performed in the whole space. Feature extraction or dimensionality reduction occurs only at the final stage after the discriminant assessment. These efforts facilitate a discriminative and stable low-dimensional feature representation of the face image. Experiments comparing the proposed approach with some other popular subspace methods on the FERET, ORL, AR and GT databases show that our method consistently outperforms others.
Xudong Jiang 0001, Bappaditya Mandal, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 Face Recognition Based on Discriminant Evaluation in the Whole Space
abstract
This paper proposes a face recognition approach that performs linear discriminant analysis in the whole eigenspace. It decomposes the eigenspace into two subspaces: a reliable subspace spanned mainly by the facial variation and an unstable subspace due to finite number of training samples. Eigenvalues in the unstable subspace are replaced by a constant. This alleviates the over-fitting problem and enables the discriminant evaluation in the whole space. Feature extraction or dimensionality reduction occurs only at the final stage after the discriminant assessment. These efforts facilitate a discriminative and stable low-dimensional feature representation of the face image. Experimental results comparing some popular subspace methods on FERET and ORL databases show that our approach consistently outperforms others.
Xudong Jiang 0001, Bappaditya Mandal, Alex Chichung Kot
ICASSP (2)1
2007 Database Clustering Based on Multi-Prototype Representation of Cluster
abstract
Clustering is a useful technique to provide the organization of multimedia database. Using single prototype to represent each cluster may not adequately model the different types of clusters and hence limits the clustering performance on the complex data structure. This paper proposes a clustering algorithm based on multi-prototype representation of cluster. The square-error clustering is used to produce a number of prototypes to locate the regions of high density. The prototypes are organized into a given number of clusters in agglomerative method based on a proposed separation measure. New prototypes are iteratively added to improve the poor cluster boundaries. As a result, the proposed algorithm can discover the clusters of complex structure. Experimental results demonstrate the effectiveness of the proposed clustering algorithm.
Manhua Liu, Xudong Jiang 0001, Alex Chichung Kot
ICME2
2007 High Capacity Data Hiding for Binary Images in Morphological Wavelet Transform Domain
abstract
This paper investigates the problem of data hiding for binary images authentication in morphological binary wavelet transform domain. Directly using the detail coefficients as a location map to determine the data hiding locations is difficult to achieve blind watermark extraction. Hence, we view flipping an edge pixel as shifting the edge location one pixel horizontally and vertically. Based on this observation, we propose an interlaced transform to track the shifted edges. Different from existing block-based approach, in which the block size is constrained by not less than 3 x 3, we process the image in 2 x 2 blocks. The two processing cases that the flippantly condition of one is not affected by flipping the candidates of another are combined such that an extremely large capacity can be achieved, which slightly sacrifices the visual quality.
Huijuan Yang, Alex Chichung Kot, Susanto Rahardja, Xudong Jiang 0001
ICME4
2007 Extracting image orientation feature by using integration operator
Xudong Jiang 0001
Pattern Recognit.1
2007 Efficient fingerprint search based on database clustering
Manhua Liu, Xudong Jiang 0001, Alex Chichung Kot
Pattern Recognit.2
2006 Fingerprint Retrieval for Identification
abstract
This paper presents a front-end filtering algorithm for fingerprint identification, which uses orientation field and dominant ridge distance as retrieval features. We propose a new distance measure that better quantifies the similarity evaluation between two orientation fields than the conventional Euclidean and Manhattan distance measures. Furthermore, fingerprints in the data base are clustered to facilitate a fast retrieval process that avoids exhaustive comparisons of an input fingerprint with all fingerprints in the data base. This makes the proposed approach applicable to large databases. Experimental results on the National Institute of Standards and Technology data base-4 show consistent better retrieval performance of the proposed approach compared to other continuous and exclusive fingerprint classification methods as well as minutia-based indexing schemes
Xudong Jiang 0001, Manhua Liu, Alex Chichung Kot
IEEE Trans. Inf. Forensics Secur.1
2005 Model-guided deformable hand shape recognition without positioning aids
Kar-Ann Toh, Weiyun Yau, Xudong Jiang 0001
Pattern Recognit.4
2004 Fingerprint image quality analysis
abstract
Fingerprint image quality analysis is crucial in eliminating poor fingerprint images, which will affect the performance of the automatic fingerprint identification system. In this article, two types of new quality measures will be introduced: ridge and valley clarity and global orientation flow to calculate the overall image quality score that can be used to quantitatively determine the quality of the fingerprint image. In order to evaluate the performance of the proposed algorithm, the quality measure is used to rank the performance of a fingerprint recognition system and the ranking is compared with the quality measure rated manually. The result shows that the proposed scheme will return a score that ensures its reliability to indicate the quality of a given fingerprint image.
Tai Pang Chen, Xudong Jiang 0001, Weiyun Yau
ICIP2
2004 Fingerprint image quality analysis
abstract
This paper discusses methods in evaluating fingerprint image quality on a local level. Feature vectors covering directional strength, sinusoidal local ridge/valley pattern, ridge/valley uniformity and core occurrences are first extracted from fingerprint image subblocks. Each subblock is then assigned a quality level through pattern classification. Three different classifiers are employed to compare each of its different effectiveness. Positive results have been obtained based on our database.
Eyung Lim, Kar-Ann Toh, P. N. Saganthan, Xudong Jiang 0001, Weiyun Yau
ICIP4
2004 A reduced multivariate polynomial model for multimodal biometrics and classifiers fusion
abstract
The multivariate polynomial model provides an effective way to describe complex nonlinear input-output relationships since it is tractable for optimization, sensitivity analysis, and prediction of confidence intervals. However, for high-dimensional and high-order problems, multivariate polynomial regression becomes impractical due to its huge number of product terms. This is especially true for the case of a full interaction model. In this paper, we propose a reduced multivariate polynomial model to circumvent the dimensionality problem with some compromise in its approximation capability. In multimodal biometrics and many classifiers fusion applications, as individual classifiers to be combined would have attained a certain level of classification accuracy, this reduced multivariate polynomial model can be used to combine these classifiers in the next level of classification taking their outputs as the inputs to the reduced multivariate polynomial model. The model is first applied to a well-known pattern classification problem to illustrate its classification capability. The reduced multivariate polynomial model is then applied to combine two biometric verification systems with improved receiver operating characteristics performance as compared to an optimal weighing method and a few commonly used classifiers.
Kar-Ann Toh, Weiyun Yau, Xudong Jiang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2003 Constructing and training feed-forward neural networks for pattern classification
Xudong Jiang 0001, Alvin Harvey Kam
Pattern Recognit.1
2002 Effective and efficient fingerprint image postprocessing
abstract
Minutiae extraction is a crucial step in an automatic fingerprint identification system. However, the presence of noise in poor-quality images causes a large number of extraction errors, including the dropping of true minutiae and production of false minutiae. A study on these errors reveals that postprocessing is effective in removing false minutiae while keeping true ones. Furthermore, the overall processing efficiency could be improved because of the reduction in total minutia number. In this paper, we present a novel fingerprint image postprocessing algorithm. It is developed based on several rules, which are generalized through a study on the errors that commonly occur in minutiae extraction and their effects on the overall verification performance. Thorough experimental tests demonstrate the proposed postprocessing algorithm to be both effective and efficient.
Haiping Lu, Xudong Jiang 0001, Weiyun Yau
ICARCV2
2002 Fingerprint quality and validity analysis
abstract
Discusses methods to estimate the quality as well as validity of a fingerprint image. Orientation certainty is used to certify the localized texture pattern of the fingerprint images while ridge to valley structure is analyzed to detect invalid images. Global uniformity and continuity ensures that the image is valid as a whole. 150 images with various qualities are evaluated using the proposed algorithm and quality benchmark we defined. A monotonic relationship is found indicating that the proposed algorithm is feasible in detecting low quality as well as invalid fingerprint images.
Eyung Lim, Xudong Jiang 0001, Weiyun Yau
ICIP (1)2
2002 Online Fingerprint Template Improvement
abstract
This work proposes a technique that improves fingerprint templates by merging and averaging minutiae of multiple fingerprints. The weighted averaging scheme enables the template to change gradually with time in line with changes of the skin and imaging conditions. The recursive nature of the algorithm greatly reduces the storage and computation requirements of this technique. As a result, the proposed template improvement procedure can be performed online during the fingerprint verification process. Extensive experimental studies demonstrate the feasibility of the proposed algorithm.
Xudong Jiang 0001, Wee Ser
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 A study of fingerprint image filtering
abstract
Fingerprint image enhancement is a crucial step in automatic fingerprint recognition. A large number of approaches for filtering the fingerprint image have been suggested. Most of them perform oriented band pass filtering. However, such filters, for example, Gabor filters, may create spurious ridge structure information. This is harmful for the feature extraction and therefore is harmful for the automatic fingerprint recognition. This paper examines the properties of applying the Gabor filter in the fingerprint image enhancement. It shows that the nonsinusoidal-shaped ridge structure, the ridge frequency estimation error, and the small filter size are the causes of creating spurious ridge structures. As a solution, we suggest an adaptive oriented low pass filter instead of the Gabor filter to avoid producing undesired harmful side effects for the automatic fingerprint recognition.
Xudong Jiang 0001
ICIP (3)1
2001 Minutiae data synthesis for fingerprint identification applications
abstract
In this paper, we address the false rejection problem due to the small solid state sensor area available for fingerprint image capture. We propose a minutiae data synthesis approach to circumvent this problem. The main advantages of this approach over the existing image mosaicing approach include low memory storage requirements and low computational complexity. Moreover, the possible matching search overhead due to data redundancy can be reduced. Extensive experiments are conducted to determine the best transformation suitable for minutiae alignment. Among the three transformations presented, affine transformation is found to be most suited for minutiae alignment. We demonstrate the idea of synthesis with an example using physical fingerprint images. The proposed synthesis system is also shown to reduce the number of false rejects caused by the use of different fingerprint regions for matching.
Kar-Ann Toh, Weiyun Yau, Xudong Jiang 0001, Tai Pang Chen, Juwei Lu, Eyung Lim
ICIP (3)3
2001 Detecting the fingerprint minutiae by adaptive tracing the gray-level ridge
Xudong Jiang 0001, Weiyun Yau, Wee Ser
Pattern Recognit.1
2000 Fundamental frequency estimation by higher order spectrum
abstract
This paper proposes a new approach to estimate the fundamental frequency of periodic non-sinusoidal signals by higher order spectrum (HOS). The power of a periodic non-sinusoidal signal is distributed to its fundamental frequency and harmonics. Properly defined higher order spectrum can enhance the fundamental frequency component of spectrum by using the harmonics and therefore has the signal more easily detected from the noise. A specific form of higher order spectrum other than the traditional bispectrum or trispectrum is proposed to improve the reliability of the fundamental frequency estimation for unknown, periodic non-sinusoidal signals. The results of applying the proposed method to weak, high noise signals are surprisingly good. The performances of the various techniques are visually, quantitatively and statistically compared for some signals with different signal-to-noise ratios.
Xudong Jiang 0001
ICASSP1
2000 Fingerprint Image Ridge Frequency Estimation by Higher Order Spectrum
abstract
This paper proposes a new approach for estimating the ridgeline frequency of fingerprint images by using higher order spectrum. The higher order spectrum enhances the fundamental frequency component of the ridgeline by using its harmonics and therefore suppresses the noise. A specific form of higher order spectrum other than the traditional bispectrum or trispectrum is employed to improve the reliability of the fingerprint ridge frequency estimation. This approach provides a reliable ridgeline frequency estimation for heavy noised fingerprint images. Experimental results are presented to compare this approach with the traditional approaches of power spectrum, bispectrum and trispectrum.
Xudong Jiang 0001
ICIP1
2000 Fingerprint Minutiae Matching Based on the Local and Global Structures
abstract
Proposes a fingerprint minutia matching technique, which matches the fingerprint minutiae by using both the local and global structures of minutiae. The local structure of a minutia describes a rotation and translation invariant feature of the minutia in its neighborhood. It is used to find the correspondence of two minutiae sets and increase the reliability of the global matching. The global structure of minutiae reliably determines the uniqueness of fingerprint. Therefore, the local and global structures of minutiae together provide a solid basis for reliable and robust minutiae matching. The proposed minutiae matching scheme is suitable for an online processing due to its high processing speed. Experimental results show the performance of the proposed technique.
Xudong Jiang 0001, Weiyun Yau
ICPR1
1999 Minutiae Extraction by Adaptive Tracing the Gray Level Ridge of the Fingerprint Image
abstract
This paper presents an improved approach of minutiae detection that adaptively traces the gray level ridge of the originals fingerprint image with piecewise linear lines of different length. While tracing the ridges, the fingerprint image is smoothed with an oriented smoothing filter only at selected pixels where smoothing is necessary. After tracing all the ridges, a piece-wise linear skeleton image is obtained. Each ridge in the skeleton is labeled with a number so that each minutiae is associated with one or two ridge numbers. The post-processing is based not only on the location relationship of the minutiae, but also the associated ridge relationship and the certainty level of the minutiae. The performance of this approach is objectively assessed by using two large fingerprint databases.
Xudong Jiang 0001, Weiyun Yau, Wee Ser
ICIP (2)1