VLDB 2026 Research / reviewers in the wild / expert
Jianguo Zhang 0001
dblp:90/6415-1
· DBLP profile ↗
86ranked-venue papers
7as first author
37since 2021 · last 2026
0000-0001-9317-0268ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 6 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 52 · 1 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text EditingabstractScene Text Editing (STE) aims to naturally modify text in images while preserving visual consistency, the decisive factors of which can be divided into three parts, i.e., text style, text content, and background. Previous methods have struggled with incomplete disentanglement of editable attributes, typically addressing only one aspect—such as editing text content—thus limiting controllability and visual consistency. To overcome these limitations, we propose TripleFDS, a novel framework for STE with disentangled modular attributes, and an accompanying dataset called SCB Synthesis. SCB Synthesis provides robust training data for triple feature disentanglement by utilizing the "SCB Group", a novel construct that combines three attributes per image to generate diverse, disentangled training groups. Leveraging this construct as a basic training unit, TripleFDS first disentangles triple features, ensuring semantic accuracy through inter-group contrastive regularization and preventing redundancy through intra-sample multi-feature orthogonality. In the synthesis phase, TripleFDS performs feature remapping to prevent "shortcut" phenomena during reconstruction and mitigate potential feature leakage. Trained on 125,000 SCB Groups, TripleFDS achieves state-of-the-art image fidelity (SSIM of 44.54) and text accuracy (ACC of 93.58%) on the mainstream STE benchmarks. Besides superior performance, the more flexible editing of TripleFDS supports new operations such as style replacement and background transfer. Yuchen Bao, Wenjian Huang 0001, Haowei Wang 0001, Shen Chen 0004, Taiping Yao, Shouhong Ding, Jianguo Zhang 0001 |
AAAI | 8 |
| 2026 | Video Understanding With Large Language Models: A SurveyabstractWith the rapid growth of online video platforms and the escalating volume of video content, the need for proficient video understanding tools has increased significantly. Given the remarkable capabilities of large language models (LLMs) in language and multimodal tasks, this survey provides a detailed overview of recent advances in video understanding that harness the power of LLMs (Vid-LLMs). The emergent capabilities of Vid-LLMs are surprisingly advanced, particularly their ability for open-ended multi-granularity (abstract, temporal, and spatiotemporal) reasoning combined with common-sense knowledge, suggesting a promising path for future video understanding. We examine the unique characteristics and capabilities of Vid-LLMs, categorizing the approaches into three main types:Video Analyzer × LLM, Video Embedder × LLM, and (Analyzer + Embedder) × LLM. We identify five subtypes based on the functions of LLMs in Vid-LLMs:LLMas Summarizer,LLMas Manager,LLMas Text Decoder,LLMas Regressor, andLLMas Hidden Layer. This survey also presents a comprehensive study of the tasks, datasets, benchmarks, and evaluation methods for Vid-LLMs. Additionally, it explores the extensive applications of Vid-LLMs in various domains, highlighting their remarkable scalability and versatility in real-world video understanding challenges. Additionally, it summarizes the limitations of existing Vid-LLMs and outlines directions for future research. For more information, readers are encouraged to visit the repository at https://github.com/yunlong10/Awesome-LLMs-for-Video-Understanding. Yunlong Tang 0002, Jing Bi 0002, Siting Xu, Luchuan Song, Susan Liang, Teng Wang 0007, Daoan Zhang, Jie An 0002, Rongyi Zhu, Ali Vosoughi, Chao Huang 0033, Zeliang Zhang 0001, Pinxin Liu, Mingqian Feng, Feng Zheng 0001, Jianguo Zhang 0001, Ping Luo 0002, Jiebo Luo 0001, Chenliang Xu |
IEEE Trans. Circuits Syst. Video Technol. | 17 |
| 2025 | Enhancing Low-Light Images: A Synthetic Data Perspective on Practical and Generalizable SolutionsabstractRecently, deep neural networks (DNNs) have emerged as the leading approach for low-light image enhancement (LLIE). However, training these models generally requires large-scale paired datasets, which are challenging to obtain due to the labor-intensive and time-consuming nature of real-world data collection. To alleviate this issue, synthetic data are often combined with real-captured data for training. However, most existing low-light image synthesis methods are simply performed in the sRGB domain using Gamma correction or manual adjustments via Lightroom, which fail to incorporate the physical imaging prior through the image signal processing (ISP) pipeline and thus result in limited dataset size and degradation space. Consequently, LLIE methods trained on such data often exhibit some drawbacks in the results, such as inaccurate white balance and abnormal enhancement artifacts, which limit their practicality and generalizability. In this paper, we propose a practical low-light image synthesis pipeline capable of generating unlimited paired training data. Our pipeline starts with a reverse ISP model that converts sRGB images back to the unprocessed RAW domain, where we then simulate low-light degradation, noise degradation, and white balance adjustments. Finally, the degraded RAW images are processed through a forward ISP model to produce low-light sRGB images. The pipeline further employs multiple tone mapping curves and color correction matrices (CCMs) to expand the degradation space. Hence, trained with our proposed synthetic data, existing state-of-the-art (SOTA) LLIE deep models are expected to improve their performance. Extensive experiments across various datasets demonstrate that our synthetic data can indeed effectively enhance existing LLIE deep models, improving both their practicality and generalizability. Qinghua Lin, Zhihua Wang 0002, Jianguo Zhang 0001, Yuming Fang 0001 |
AAAI | 5 |
| 2025 | SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space ModelsabstractKnown as low energy consumption networks, spiking neural networks (SNNs) have gained a lot of attention within the past decades. While SNNs are increasing competitive with artificial neural networks (ANNs) for vision tasks, they are rarely used for long sequence tasks, despite their intrinsic temporal dynamics. In this work, we develop spiking state space models (SpikingSSMs) for long sequence learning by leveraging on the sequence learning abilities of state space models (SSMs). Inspired by dendritic neuron structure, we hierarchically integrate neuronal dynamics with the original SSM block, meanwhile realizing sparse synaptic computation. Furthermore, to solve the conflict of event-driven neuronal dynamics with parallel computing, we propose a light-weight surrogate dynamic network which accurately predicts the after-reset membrane potential and compatible to learnable thresholds, enabling orders of acceleration in training speed compared with conventional iterative methods. On the long range arena benchmark task, SpikingSSM achieves competitive performance to state-of-the-art SSMs meanwhile realizing on average 90% of network sparsity. On language modeling, our network significantly surpasses existing spiking large language models (spikingLLMs) on the WikiText-103 dataset with only a third of the model size, demonstrating its potential as backbone architecture for low computation cost LLMs. Shuaijie Shen, Renzhuo Huang, Yan Zhong 0001, Qinghai Guo, Zhichao Lu, Jianguo Zhang 0001, Luziwei Leng |
AAAI | 7 |
| 2025 | ReferDINO: Referring Video Object Segmentation with Visual Grounding FoundationsabstractReferring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language understanding, pixel-level dense prediction and spatiotemporal reasoning. Despite notable progress in recent years, existing methods still exhibit a noticeable gap when considering all these aspects. In this work, we propose \textbf{ReferDINO}, a strong RVOS model that inherits region-level vision-language alignment from foundational visual grounding models, and is further endowed with pixel-level dense perception and cross-modal spatiotemporal reasoning. In detail, ReferDINO integrates two key components: 1) a grounding-guided deformable mask decoder that utilizes location prediction to progressively guide mask prediction through differentiable deformation mechanisms; 2) an object-consistent temporal enhancer that injects pretrained time-varying text features into inter-frame interaction to capture object-aware dynamic changes. Moreover, a confidence-aware query pruning strategy is designed to accelerate object decoding without compromising model performance. Extensive experimental results on five benchmarks demonstrate that our ReferDINO significantly outperforms previous methods (e.g., +3.9% (\mathcal{J}&\mathcal{F}) on Ref-YouTube-VOS) with real-time inference speed (51 FPS). Tianming Liang, Kun-Yu Lin, Chaolei Tan, Jianguo Zhang 0001, Wei-Shi Zheng 0001, Jianfang Hu |
ICCV | 4 |
| 2025 | Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized ConstraintsabstractPart-level features are crucial for image understanding, but few studies focus on them because of the lack of fine-grained labels. Although unsupervised part discovery can eliminate the reliance on labels, most of them cannot maintain robustness across various categories and scenarios, which restricts their application range. To overcome this limitation, we present a more effective paradigm for unsupervised part discovery, named Masked Part Autoencoder (MPAE). It first learns part descriptors as well as a feature map from the inputs and produces patch features from a masked version of the original images. Then, the masked regions are filled with the learned part descriptors based on the similarity between the local features and descriptors. By restoring these masked patches using the part descriptors, they become better aligned with their part shapes, guided by appearance features from unmasked patches. Finally, MPAE robustly discovers meaningful parts that closely match the actual object shapes, even in complex scenarios. Moreover, several looser yet more effective constraints are proposed to enable MPAE to identify the presence of parts across various scenarios and categories in an unsupervised manner. This provides the foundation for addressing challenges posed by occlusion and for exploring part similarity across multiple categories. Extensive experiments demonstrate that our method robustly discovers meaningful parts across various categories and scenarios. The code is available at the project https://github.com/Jiahao-UTS/MPAE. Jiahao Xia 0001, Yike Wu 0001, Wenjian Huang 0001, Jianguo Zhang 0001, Jian Zhang 0002 |
ICCV | 4 |
| 2025 | Open-Det: An Efficient Learning Framework for Open-Ended DetectionabstractOpen-Ended object Detection (OED) is a novel and challenging task that detects objects and generates their category names in a free-form manner, without requiring additional vocabularies during inference. However, the existing OED models, such as GenerateU, require large-scale datasets for training, suffer from slow convergence, and exhibit limited performance. To address these issues, we present a novel and efficient Open-Det framework, consisting of four collaborative parts. Specifically, Open-Det accelerates model training in both the bounding box and object name generation process by reconstructing the Object Detector and the Object Name Generator. To bridge the semantic gap between Vision and Language modalities, we propose a Vision-Language Aligner with V-to-L and L-to-V alignment mechanisms, incorporating with the Prompts Distiller to transfer knowledge from the VLM into VL-prompts, enabling accurate object name generation for the LLM. In addition, we design a Masked Alignment Loss to eliminate contradictory supervision and introduce a Joint Loss to enhance classification, resulting in more efficient training. Compared to GenerateU, Open-Det, using only 1.5% of the training data (0.077M vs. 5.077M), 20.8% of the training epochs (31 vs. 149), and fewer GPU resources (4 V100 vs. 16 A100), achieves even higher performance (+1.0% in APr). The source codes are available at: https://github.com/Med-Process/Open-Det. Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang |
ICML | 5 |
| 2025 | DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object DetectionabstractPopular transformer detectors have achieved promising performance through query-based learning using attention mechanisms. However, the roles of existing decoder query types (e.g., content query and positional query) are still underexplored. These queries are generally predefined with a fixed number (fixed-query), which limits their flexibility. We find that the learning of these fixed-query is impaired by Recurrent Opposing in Teractions (ROT) between two attention operations: Self-Attention (query-to-query) and Cross-Attention (query-to-encoder), thereby degrading decoder efficiency. Furthermore, "query ambiguity" arises when shared-weight decoder layers are processed with both one-to-one and one-to-many label assignments during training, violating DETR's one-to-one matching principle. To address these challenges, we propose DS-Det, a more efficient detector capable of detecting a flexible number of objects in images. Specifically, we reformulate and introduce a new unified Single-Query paradigm for decoder modeling, transforming the fixed-query into flexible. Furthermore, we propose a simplified decoder framework through attention disentangled learning: locating boxes with Cross-Attention (one-to-many process), deduplicating predictions with Self-Attention (one-to-one process), addressing ''query ambiguity'' and ''ROT'' issues directly, and enhancing decoder efficiency. We further introduce a unified PoCoo loss that leverages box size priors to prioritize query learning on hard samples such as small objects. Extensive experiments across five different backbone models on COCO2017 and WiderPerson datasets demonstrate the general effectiveness and superiority of DS-Det. The source codes are available at https://github.com/Med-Process/DS-Det/. Guiping Cao, Xiangyuan Lan, Wenjian Huang 0001, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001 |
ACM Multimedia | 4 |
| 2025 | Mitigating Knowledge Discrepancies among Multiple Datasets for Task-agnostic Unified Face AlignmentabstractAbstract Despite the similar structures of human faces, existing face alignment methods cannot learn unified knowledge from multiple datasets with different landmark annotations. The limited training samples in a single dataset commonly result in fragile robustness in this field. To mitigate knowledge discrepancies among different datasets and train a task-agnostic unified face alignment (TUFA) framework, this paper presents a strategy to unify knowledge from multiple datasets. Specifically, we calculate a mean face shape for each dataset. To explicitly align these mean shapes on an interpretable plane based on their semantics, each shape is then incorporated with a group of semantic alignment embeddings. The 2D coordinates of these aligned shapes can be viewed as the anchors of the plane. By encoding them into structure prompts and further regressing the corresponding facial landmarks using image features, a mapping from the plane to the target faces is finally established, which unifies the learning target of different datasets. Consequently, multiple datasets can be utilized to boost the generalization ability of the model. The successful mitigation of discrepancies also enhances the efficiency of knowledge transferring to a novel dataset, significantly boosts the performance of few-shot face alignment. Additionally, the interpretable plane endows TUFA with a task-agnostic characteristic, enabling it to locate landmarks unseen during training in a zero-shot manner. Extensive experiments are carried on seven benchmarks and the results demonstrate an impressive improvement in face alignment brought by knowledge discrepancies mitigation. The code is available at https://github.com/Jiahao-UTS/TUFA Jiahao Xia 0001, Min Xu 0001, Wenjian Huang 0001, Jianguo Zhang 0001, Haimin Zhang 0001, Chunxia Xiao |
Int. J. Comput. Vis. | 4 |
| 2025 | H-Calibration: Rethinking Classifier Recalibration With Probabilistic Error-Bounded ObjectiveabstractDeep neural networks have demonstrated remarkable performance across numerous learning tasks but often suffer from miscalibration, resulting in unreliable probability outputs. This has inspired many recent works on mitigating miscalibration, particularly through post-hoc recalibration methods that aim to obtain calibrated probabilities without sacrificing the classification performance of pre-trained models. In this study, we summarize and categorize previous works into three general strategies: intuitively designed methods, binning-based methods, and methods based on formulations of ideal calibration. Through theoretical and practical analysis, we highlight ten common limitations in previous approaches. To address these limitations, we propose a probabilistic learning framework for calibration called $h$h-calibration, which theoretically constructs an equivalent learning formulation for canonical calibration with boundedness. On this basis, we design a simple yet effective post-hoc calibration algorithm. Our method not only overcomes the ten identified limitations but also achieves markedly better performance than traditional methods, as validated by extensive experiments. We further analyze, both theoretically and experimentally, the relationship and advantages of our learning objective compared to traditional proper scoring rule. In summary, our probabilistic framework derives an approximately equivalent differentiable objective for learning error-bounded calibrated probabilities, elucidating the correspondence and convergence properties of computational statistics with respect to theoretical bounds in canonical calibration. The theoretical effectiveness is verified on standard post-hoc calibration benchmarks by achieving state-of-the-art performance. This research offers valuable reference for learning reliable likelihood in related fields. Wenjian Huang 0001, Guiping Cao, Jiahao Xia 0001, Jingkun Chen, Hao Wang 0230, Jianguo Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Addressing Inconsistent Labeling With Cross Image Matching for Scribble-Based Medical Image SegmentationabstractIn recent years, there has been a notable surge in the adoption of weakly-supervised learning for medical image segmentation, utilizing scribble annotation as a means to potentially reduce annotation costs. However, the inherent characteristics of scribble labeling, marked by incompleteness, subjectivity, and a lack of standardization, introduce inconsistencies into the annotations. These inconsistencies become significant challenges for the network's learning process, ultimately affecting the performance of segmentation. To address this challenge, we propose creating a reference set to guide pixel-level feature matching, constructed from class-specific tokens and pixel-level features extracted from variously images. Serving as a repository showcasing diverse pixel styles and classes, the reference set becomes the cornerstone for a pixel-level feature matching strategy. This strategy enables the effective comparison of unlabeled pixels, offering guidance, particularly in learning scenarios characterized by inconsistent and incomplete scribbles. The proposed strategy incorporates smoothing and regression techniques to align pixel-level features across different images. By leveraging the diversity of pixel sources, our matching approach enhances the network's ability to learn consistent patterns from the reference set. This, in turn, mitigates the impact of inconsistent and incomplete labeling, resulting in improved segmentation outcomes. Extensive experiments conducted on three publicly available datasets demonstrate the superiority of our approach over state-of-the-art methods in terms of segmentation accuracy and stability. The code will be made publicly available at https://github.com/jingkunchen/scribble-medical-segmentation. Jingkun Chen, Wenjian Huang 0001, Jianguo Zhang 0001, Kurt Debattista, Jungong Han |
IEEE Trans. Image Process. | 3 |
| 2025 | A Lesion-Fusion Neural Network for Multi-View Diabetic Retinopathy GradingabstractAs the most common complication of diabetes, diabetic retinopathy (DR) is one of the main causes of irreversible blindness. Automatic DR grading plays a crucial role in early diagnosis and intervention, reducing the risk of vision loss in people with diabetes. In these years, various deep-learning approaches for DR grading have been proposed. Most previous DR grading models are trained using the dataset of single-field fundus images, but the entire retina cannot be fully visualized in a single field of view. There are also problems of scattered location and great differences in the appearance of lesions in fundus images. To address the limitations caused by incomplete fundus features, and the difficulty in obtaining lesion information. This work introduces a novel multi-view DR grading framework, which solves the problem of incomplete fundus features by jointly learning fundus images from multiple fields of view. Furthermore, the proposed model combines multi-view inputs such as fundus images and lesion snapshots. It utilizes heterogeneous convolution blocks (HCB) and scalable self-attention classes (SSAC), which enhance the ability of the model to obtain lesion information. The experimental results show that our proposed method performs better than the benchmark methods on the large-scale dataset. Xiaoling Luo 0001, Qihao Xu, Zhihua Wang 0002, Chao Huang 0008, Chengliang Liu 0003, Xiaopeng Jin, Jianguo Zhang 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Cross-DINO: Cross the Deep MLP and Transformer for Small Object DetectionabstractSmall Object Detection (SOD) poses significant challenges due to limited information and the model's low class prediction score. While Transformer-based detectors have shown promising performance, their potential for SOD remains largely unexplored. In typical DETR-like frameworks, the CNN backbone network, specialized in aggregating local information, struggles to capture the necessary contextual information for SOD. The multiple attention layers in the Transformer Encoder face difficulties in effectively attending to small objects and can also lead to blurring of features. Furthermore, the model's lower class prediction score of small objects compared to large objects further increases the difficulty of SOD. To address these challenges, we introduce a novel approach calledCross-DINO. This approach incorporates the deep MLP network to aggregate initial feature representations with both short and long range information for SOD. Then, a new Cross Coding Twice Module (CCTM) is applied to integrate these initial representations to the Transformer Encoder feature, enhancing the details of small objects. Additionally, we introduce a new kind of soft label named Category-Size (CS), integrating the Category and Size of objects. By treating CS as new ground truth, we propose a new loss function called Boost Loss to improve the class prediction score of the model. Extensive experimental results on COCO, WiderPerson, VisDrone, AI-TOD, and SODA-D datasets demonstrate that Cross-DINO efficiently improves the performance of DETR-like models on SOD. Specifically, our model achieves36.4%AP$_{S}$on COCO for SOD with only 45M parameters, outperforming the DINO by+4.4%AP$_{S}$(36.4% vs. 32.0%) with fewer parameters and FLOPs, under 12 epochs training setting. Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Efficient Deep Spiking Multilayer Perceptrons With Multiplication-Free InferenceabstractAdvancements in adapting deep convolution architectures for spiking neural networks (SNNs) have significantly enhanced image classification performance and reduced computational burdens. However, the inability of multiplication-free inference (MFI) to align with attention and transformer mechanisms, which are critical to superior performance on high-resolution vision tasks, imposes limitations on these gains. To address this, our research explores a new pathway, drawing inspiration from the progress made in multilayer perceptrons (MLPs). We propose an innovative spiking MLP architecture that uses batch normalization (BN) to retain MFI compatibility and introduce a spiking patch encoding (SPE) layer to enhance local feature extraction capabilities. As a result, we establish an efficient multistage spiking MLP network that blends effectively global receptive fields with local feature extraction for comprehensive spike-based computation. Without relying on pretraining or sophisticated SNN training techniques, our network secures a top-one accuracy of 66.39% on the ImageNet-1K dataset, surpassing the directly trained spiking ResNet-34 by 2.67%. Furthermore, we curtail computational costs, model parameters, and simulation steps. An expanded version of our network compares with the performance of the spiking VGG-16 network with a 71.64% top-one accuracy, all while operating with a model capacity 2.1 times smaller. Our findings highlight the potential of our deep SNN architecture in effectively integrating global and local learning abilities. Interestingly, the trained receptive field in our network mirrors the activity patterns of cortical cells. Boyan Li 0001, Luziwei Leng, Shuaijie Shen, Jianguo Zhang 0001, Jianxing Liao, Ran Cheng 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection
Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001 |
IJCAI | 4 |
| 2024 | Unsupervised Part Discovery via Dual Representation AlignmentabstractObject parts serve as crucial intermediate representations in various downstream tasks, but part-level representation learning still has not received as much attention as other vision tasks. Previous research has established that Vision Transformer can learn instance-level attention without labels, extracting high-quality instance-level representations for boosting downstream tasks. In this paper, we achieve unsupervised part-specific attention learning using a novel paradigm and further employ the part representations to improve part discovery performance. Specifically, paired images are generated from the same image with different geometric transformations, and multiple part representations are extracted from these paired images using a novel module, named PartFormer. These part representations from the paired images are then exchanged to improve geometric transformation invariance. Subsequently, the part representations are aligned with the feature map extracted by a feature map encoder, achieving high similarity with the pixel representations of the corresponding part regions and low similarity in irrelevant regions. Finally, the geometric and semantic constraints are applied to the part representations through the intermediate results in alignment for part-specific attention learning, encouraging the PartFormer to focus locally and the part representations to explicitly include the information of the corresponding parts. Moreover, the aligned part representations can further serve as a series of reliable detectors in the testing phase, predicting pixel masks for part discovery. Extensive experiments are carried out on four widely used datasets, and our results demonstrate that the proposed method achieves competitive performance and robustness due to its part-specific attention. Jiahao Xia 0001, Wenjian Huang 0001, Min Xu 0001, Jianguo Zhang 0001, Haimin Zhang 0001, Ziyu Sheng, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Dynamic contrastive learning guided by class confidence and confusion degree for medical image segmentation
Jingkun Chen, Changrui Chen, Wenjian Huang 0001, Jianguo Zhang 0001, Kurt Debattista, Jungong Han |
Pattern Recognit. | 4 |
| 2023 | Rethinking Alignment and Uniformity in Unsupervised Image Semantic SegmentationabstractUnsupervised image segmentation aims to match low-level visual features with semantic-level representations without outer supervision. In this paper, we address the critical properties from the view of feature alignments and feature uniformity for UISS models. We also make a comparison between UISS and image-wise representation learning. Based on the analysis, we argue that the existing MI-based methods in UISS suffer from representation collapse. By this, we proposed a robust network called Semantic Attention Network(SAN), in which a new module Semantic Attention(SEAT) is proposed to generate pixel-wise and semantic features dynamically. Experimental results on multiple semantic segmentation benchmarks show that our unsupervised segmentation framework specializes in catching semantic representations, which outperforms all the unpretrained and even several pretrained methods. Daoan Zhang, Haoquan Li, Wenjian Huang 0001, Lingyun Huang, Jianguo Zhang 0001 |
AAAI | 6 |
| 2023 | Feature Alignment and Uniformity for Test Time AdaptationabstractTest time adaptation (TTA) aims to adapt deep neural networks when receiving out of distribution test domain samples. In this setting, the model can only access online unlabeled test samples and pretrained models on the training domains. We first address TTA as a feature revision problem due to the domain gap between source domains and target domains. After that, we follow the two measurements alignment and uniformity to discuss the test time feature revision. For test time feature uniformity, we propose a test time self-distillation strategy to guarantee the consistency of uniformity between representations of the current batch and all the previous batches. For test time feature alignment, we propose a memorized spatial local clustering strategy to align the representations among the neighborhood samples for the upcoming batch. To deal with the common noisy label problem, we propound the entropy and consistency filters to select and drop the possible noisy labels. To prove the scalability and efficacy of our method, we conduct experiments on four domain generalization bench marks and four medical image segmentation tasks with various backbones. Experiment results show that our method not only improves baseline stably but also outperforms existing state-of-the-art test time adaptation methods. Shuai Wang 0048, Daoan Zhang, Zipei Yan, Jianguo Zhang 0001 |
CVPR | 4 |
| 2023 | Strip-MLP: Efficient Token Interaction for Vision MLPabstractToken interaction operation is one of the core modules in MLP-based models to exchange and aggregate information between different spatial locations. However, the power of token interaction on the spatial dimension is highly dependent on the spatial resolution of the feature maps, which limits the model’s expressive ability, especially in deep layers where the feature are down-sampled to a small spatial size. To address this issue, we present a novel method called Strip-MLP to enrich the token interaction power in three ways. Firstly, we introduce a new MLP paradigm called Strip MLP layer that allows the token to interact with other tokens in a cross-strip manner, enabling the tokens in a row (or column) to contribute to the information aggregations in adjacent but different strips of rows (or columns). Secondly, a Cascade Group Strip Mixing Module (CGSMM) is proposed to overcome the performance degradation caused by small spatial feature size. The module allows tokens to interact more effectively in the manners of within-patch and cross-patch, which is independent to the feature spatial size. Finally, based on the Strip MLP layer, we propose a novel Local Strip Mixing Module (LSMM) to boost the token interaction power in the local region. Extensive experiments demonstrate that Strip-MLP significantly improves the performance of MLP-based models on small datasets and obtains comparable or even better results on ImageNet. In particular, Strip-MLP models achieve higher average Top-1 accuracy than existing MLP-based models by +2.44% on Caltech-101 and +2.16% on CIFAR-100. The source codes will be available at https://github.com/Med-Process/Strip_MLP. Guiping Cao, Shengda Luo, Wenjian Huang 0001, Xiangyuan Lan, Dongmei Jiang, Yaowei Wang 0001, Jianguo Zhang 0001 |
ICCV | 7 |
| 2023 | Cross Contrasting Feature Perturbation for Domain GeneralizationabstractDomain generalization (DG) aims to learn a robust model from source domains that generalize well on unseen target domains. Recent studies focus on generating novel domain samples or features to diversify distributions complementary to source domains. Yet, these approaches can hardly deal with the restriction that the samples synthesized from various domains can cause semantic distortion. In this paper, we propose an online one-stage Cross Contrasting Feature Perturbation (CCFP) framework to simulate domain shift by generating perturbed features in the latent space while regularizing the model prediction against domain shift. Different from the previous fixed synthesizing strategy, we design modules with learnable feature perturbations and semantic consistency constraints. In contrast to prior work, our method does not use any generative-based models or domain labels. We conduct extensive experiments on a standard DomainBed benchmark with a strict evaluation protocol for a fair comparison. Comprehensive experiments show that our method outperforms the previous state-of-the-art, and quantitative analyses illustrate that our approach can alleviate the domain shift problem in out-of-distribution (OOD) scenarios. https://github.com/hackmebroo/CCFP Daoan Zhang, Wenjian Huang 0001, Jianguo Zhang 0001 |
ICCV | 4 |
| 2023 | Class attention to regions of lesion for imbalanced medical image recognition
Jiaxin Zhuang, Jiabin Cai, Jianguo Zhang 0001, Wei-Shi Zheng 0001 |
Neurocomputing | 3 |
| 2023 | Egocentric Action Recognition by Automatic Relation ModelingabstractEgocentric videos, which record the daily activities of individuals from a first-person point of view, have attracted increasing attention during recent years because of their growing use in many popular applications, including life logging, health monitoring and virtual reality. As a fundamental problem in egocentric vision, one of the tasks of egocentric action recognition aims to recognize the actions of the camera wearers from egocentric videos. In egocentric action recognition, relation modeling is important, because the interactions between the camera wearer and the recorded persons or objects form complex relations in egocentric videos. However, only a few of existing methods model the relations between the camera wearer and the interacting persons for egocentric action recognition, and moreover they require prior knowledge or auxiliary data to localize the interacting persons. In this work, we consider modeling the relations in a weakly supervised manner, i.e., without using annotations or prior knowledge about the interacting persons or objects, for egocentric action recognition. We form a weakly supervised framework by unifying automatic interactor localization and explicit relation modeling for the purpose of automatic relation modeling. First, we learn to automatically localize the interactors, i.e., the body parts of the camera wearer and the persons or objects that the camera wearer interacts with, by learning a series of keypoints directly from video data to localize the action-relevant regions with only action labels and some constraints on these keypoints. Second, more importantly, to explicitly model the relations between the interactors, we develop an ego-relational LSTM (long short-term memory) network with several candidate connections to model the complex relations in egocentric videos, such as the temporal, interactive, and contextual relations. In particular, to reduce human efforts and manual interventions needed to construct an optimal ego-relational LSTM structure, we search for the optimal connections by employing a differentiable network architecture search mechanism, which automatically constructs the ego-relational LSTM network to explicitly model different relations for egocentric action recognition. We conduct extensive experiments on egocentric video datasets to illustrate the effectiveness of our method. Haoxin Li, Wei-Shi Zheng 0001, Jianguo Zhang 0001, Haifeng Hu 0001, Jiwen Lu, Jian-Huang Lai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Robust Face Alignment via Inherent Relation Learning and Uncertainty EstimationabstractHuman tends to locate the facial landmarks with heavy occlusion by their relative position to the easily identified landmarks. The clue is defined as the landmark inherent relation while it is ignored by most existing methods. In this paper, we present Dynamic Sparse Local Patch Transformer (DSLPT), a novel face alignment framework for the inherent relation learning and uncertainty estimation. Unlike most existing methods that regress facial landmarks directly from global features, the DSLPT first generates a rough representation of each landmark from a local patch cropped from the feature map and then adaptively aggregates them by a case dependent inherent relation. Finally, the DSLPT predicts the coordinate and uncertainty of each landmark by regressing their probability distribution from the output features. Moreover, we introduce a coarse-to-fine framework to incorporate with DSLPT for an improved result. In the framework, the position and size of each patch are determined by the probability distribution of the corresponding landmark predicted in the previous stage. The dynamic patches will ensure a fine-grained landmark representation for inherent relation learning so that a rough prediction result can gradually converge to the target facial landmarks. We integrate the coarse-to-fine model into an end-to-end training pipeline and carry out experiments on the mainstream benchmarks. The results demonstrate that the DSLPT achieves state-of-the-art performance with much less computational complexity. The codes and models are available at https://github.com/Jiahao-UTS/DSLPT. Jiahao Xia 0001, Min Xu 0001, Haimin Zhang 0001, Jianguo Zhang 0001, Wenjian Huang 0001, Hu Cao, Shiping Wen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Toward a blind image quality evaluator in the wild by learning beyond human opinion scoresabstractNowadays, most existing blind image quality assessment (BIQA) models i n t h e w i l d heavily rely on human ratings, which are extraordinarily labor-expensive to collect. Here, we propose an o p i n i o n − f r e e BIQA method that learns from multiple annotators to assess the perceptual quality of images captured in the wild. Specifically, we first synthesize distorted images based on the pristine counterparts. We then randomly assemble a set of image pairs from the synthetic images, and use a group of IQA models to assign pseudo-binary labels for each pair indicating which image has higher quality as the supervisory signal. Based on the newly established pseudo-labeled dataset, we train a deep neural network (DNN)-based BIQA model to rank the perceptual quality, optimized for consistency with the binary rank labels. Since there exists domain shift, e.g., distortion shift and content shift, between the synthetic and in-the-wild images, we leverage two ways to alleviate this issue. First, the simulated distortions should be similar to authentic distortions as much as possible. Second, an unsupervised domain adaptation (UDA) module is further applied to encourage learning domain-invariant features between two domains. Extensive experiments demonstrate the effectiveness of our proposed o p i n i o n − f r e e BIQA model, yielding SOTA performance in terms of correlation with human opinion scores, as well as gMAD competition. Our code is available at: https://github.com/wangzhihua520/OF_BIQA . Zhihua Wang 0002, Jianguo Zhang 0001, Yuming Fang 0001 |
Pattern Recognit. | 3 |
| 2023 | Corrigendum to 'Toward a Blind Image Quality Evaluator in the Wild by Learning beyond Human Opinion Scores' Pattern Recognition. Volume 137 (2023) 109296
Zhihua Wang 0002, Jianguo Zhang 0001, Yuming Fang 0001 |
Pattern Recognit. | 3 |
| 2023 | Semi-Supervised Unpaired Medical Image Segmentation Through Task-Affinity ConsistencyabstractDeep learning-based semi-supervised learning (SSL) algorithms are promising in reducing the cost of manual annotation of clinicians by using unlabelled data, when developing medical image segmentation tools. However, to date, most existing semi-supervised learning (SSL) algorithms treat the labelled images and unlabelled images separately and ignore the explicit connection between them; this disregards essential shared information and thus hinders further performance improvements. To mine the shared information between the labelled and unlabelled images, we introduce a class-specific representation extraction approach, in which a task-affinity module is specifically designed for representation extraction. We further cast the representation into two different views of feature maps; one is focusing on low-level context, while the other concentrates on structural information. The two views of feature maps are incorporated into the task-affinity module, which then extracts the class-specific representations to aid the knowledge transfer from the labelled images to the unlabelled images. In particular, a task-affinity consistency loss between the labelled images and unlabelled images based on the multi-scale class-specific representations is formulated, leading to a significant performance improvement. Experimental results on three datasets show that our method consistently outperforms existing state-of-the-art methods. Our findings highlight the potential of consistency between class-specific knowledge for semi-supervised medical image segmentation. The code and models are to be made publicly available at https://github.com/jingkunchen/TAC. Jingkun Chen, Jianguo Zhang 0001, Kurt Debattista, Jungong Han |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Sparse Local Patch Transformer for Robust Face Alignment and Landmarks Inherent Relation LearningabstractHeatmap regression methods have dominated face alignment area in recent years while they ignore the inherent relation between different landmarks. In this paper, we propose a Sparse Local Patch Transformer (SLPT) for learning the inherent relation. The SLPT generates the representation of each single landmark from a local patch and aggregates them by an adaptive inherent relation based on the attention mechanism. The subpixel coordinate of each landmark is predicted independently based on the aggregated feature. Moreover, a coarse-to-fine framework is further introduced to incorporate with the SLPT, which enables the initial landmarks to gradually converge to the target facial landmarks using fine-grained features from dynamically resized local patches. Extensive experiments carried out on three popular benchmarks, including WFLW, 300W and COFW, demonstrate that the proposed method works at the state-of-the-art level with much less computational complexity by learning the inherent relation between facial landmarks. The code is available at the project website11https://github.com/Jiahao-UTS/SLPT-master. Jiahao Xia 0001, Weiwei Qu, Wenjian Huang 0001, Jianguo Zhang 0001, Min Xu 0001 |
CVPR | 4 |
| 2022 | Discrete time convolution for fast event-based stereoabstractInspired by biological retina, dynamical vision sensor transmits events of instantaneous changes of pixel intensity, giving it a series of advantages over traditional frame-based camera, such as high dynamical range, high temporal resolution and low power consumption. However, extracting information from highly asynchronous event data is a challenging task. Inspired by continuous dynamics of biological neuron models, we propose a novel encoding method for sparse events-continuous time convolution (CTC)-which learns to model the spatial feature of the data with intrinsic dynamics. Adopting channel-wise parameterization, temporal dynamics of the model is synchronized on the same feature map and diverges across different ones, enabling it to embed data in a variety of temporal scales. Abstracted from CTC, we further develop discrete time convolution (DTC) which accelerates the process with lower computational cost. We apply these methods to event-based multi- view stereo matching where they surpass state-of-the-art methods on benchmark criteria of the MVSEC dataset. Spatially sparse event data often leads to inaccurate estimation of edges and local contours. To address this problem, we propose a dual-path architecture in which the feature map is complemented by underlying edge information from original events extracted with spatially-adaptive denormal-ization. We demonstrate the superiority of our model in terms of speed (up to 110 FPS), accuracy and robustness, showing a great potential for real-time fast depth estimation. Finally, we perform experiments on the recent DSEC dataset to demonstrate the general usage of our model. Kaiwei Che, Jianguo Zhang 0001, Qinghai Guo, Luziwei Leng |
CVPR | 3 |
| 2022 | TransVLAD: Focusing on Locally Aggregated Descriptors for Few-Shot Learning
Haoquan Li, Laoming Zhang, Daoan Zhang, Lang Fu, Peng Yang 0008, Jianguo Zhang 0001 |
ECCV (20) | 6 |
| 2022 | Domain-Adaptive 3D Medical Image Synthesis: An Efficient Unsupervised Approach
Qingqiao Hu, Hongwei Li 0004, Jianguo Zhang 0001 |
MICCAI (6) | 3 |
| 2022 | Differentiable hierarchical and surrogate gradient search for spiking neural networksabstractSpiking neural network (SNN) has been viewed as a potential candidate for the next generation of artificial intelligence with appealing characteristics such as sparse computation and inherent temporal dynamics. By adopting architectures of deep artificial neural networks (ANNs), SNNs are achieving competitive performances in benchmark tasks such as image classification. However, successful architectures of ANNs are not necessary ideal for SNN and when tasks become more diverse effective architectural variations could be critical. To this end, we develop a spike-based differentiable hierarchical search (SpikeDHS) framework, where spike-based computation is realized on both the cell and the layer level search space. Based on this framework, we find effective SNN architectures under limited computation cost. During the training of SNN, a suboptimal surrogate gradient function could lead to poor approximations of true gradients, making the network enter certain local minima. To address this problem, we extend the differential approach to surrogate gradient search where the SG function is efficiently optimized locally. Our models achieve state-of-the-art performances on classification of CIFAR10/100 and ImageNet with accuracy of 95.50%, 76.25% and 68.64%. On event-based deep stereo, our method finds optimal layer variation and surpasses the accuracy of specially designed ANNs meanwhile with 26$\times$ lower energy cost ($6.7\mathrm{mJ}$), demonstrating the advantage of SNN in processing highly sparse and dynamic signals. Codes are available at \url{https://github.com/Huawei-BIC/SpikeDHS}. Kaiwei Che, Luziwei Leng, Jianguo Zhang 0001, Qinghu Meng, Qinghai Guo, Jianxing Liao |
NeurIPS | 4 |
| 2022 | Density-driven Regularization for Out-of-distribution DetectionabstractDetecting out-of-distribution (OOD) samples is essential for reliably deploying deep learning classifiers in open-world applications. However, existing detectors relying on discriminative probability suffer from the overconfident posterior estimate for OOD data. Other reported approaches either impose strong unproven parametric assumptions to estimate OOD sample density or develop empirical detectors lacking clear theoretical motivations. To address these issues, we propose a theoretical probabilistic framework for OOD detection in deep classification networks, in which two regularization constraints are constructed to reliably calibrate and estimate sample density to identify OOD. Specifically, the density consistency regularization enforces the agreement between analytical and empirical densities of observable low-dimensional categorical labels. The contrastive distribution regularization separates the densities between in distribution (ID) and distribution-deviated samples. A simple and robust implementation algorithm is also provided, which can be used for any pre-trained neural network classifiers. To the best of our knowledge, we have conducted the most extensive evaluations and comparisons on computer vision benchmarks. The results show that our method significantly outperforms state-of-the-art detectors, and even achieves comparable or better performance than methods utilizing additional large-scale outlier exposure datasets. Wenjian Huang 0001, Hao Wang 0005, Jiahao Xia 0001, Chengyan Wang, Jianguo Zhang 0001 |
NeurIPS | 5 |
| 2022 | Partial Least Square Regression via Three-Factor SVD-Type Manifold Optimization for EEG Decoding
Wanguang Yin, Zhichao Liang, Jianguo Zhang 0001, Quanying Liu |
PRCV (1) | 3 |
| 2021 | Sign-Agnostic Implicit Learning of Surface Self-Similarities for Shape Modeling and Reconstruction From Raw Point CloudsabstractShape modeling and reconstruction from raw point clouds of objects stand as a fundamental challenge in vision and graphics research. Classical methods consider analytic shape priors; however, their performance is degraded when the scanned points deviate from the ideal conditions of cleanness and completeness. Important progress has been recently made by data-driven approaches, which learn global and/or local models of implicit surface representations from auxiliary sets of training shapes. Motivated from a universal phenomenon that self-similar shape patterns of local surface patches repeat across the entire surface of an object, we aim to push forward the data-driven strategies and propose to learn a local implicit surface network for a shared, adaptive modeling of the entire surface for a direct surface reconstruction from raw point cloud; we also enhance the leveraging of surface self-similarities by improving correlations among the optimized latent codes of individual surface patches. Given that orientations of raw points could be unavailable or noisy, we extend signagnostic learning into our local implicit model, which enables our recovery of signed implicit fields of local surfaces from the unsigned inputs. We term our framework as Sign-Agnostic Implicit Learning of Surface Self-Similarities (SAIL-S3). With a global post-optimization of local sign flipping, SAIL-S3 is able to directly model raw, un-oriented point clouds and reconstruct high-quality object surfaces. Experiments show its superiority over existing methods. Wenbin Zhao, Jiabao Lei, Yuxin Wen, Jianguo Zhang 0001, Kui Jia |
CVPR | 4 |
| 2021 | Imbalance-Aware Self-supervised Learning for 3D Radiomic Representations
Hongwei Li 0004, Fei-Fei Xue, Krishna Chaitanya, Shengda Luo, Ivan Ezhov, Benedikt Wiestler, Jianguo Zhang 0001, Bjoern Menze |
MICCAI (2) | 7 |
| 2021 | Continual Representation Learning via Auto-Weighted Latent Embeddings on Person ReID
Tianjun Huang, Weiwei Qu, Jianguo Zhang 0001 |
PRCV (3) | 3 |
| 2020 | TEA: Temporal Excitation and Aggregation for Action RecognitionabstractTemporal modeling is key for action recognition in videos. It normally considers both short-range motions and long-range aggregations. In this paper, we propose a Temporal Excitation and Aggregation (TEA) block, including a motion excitation (ME) module and a multiple temporal aggregation (MTA) module, specifically designed to capture both short- and long-range temporal evolution. In particular, for short-range motion modeling, the ME module calculates the feature-level temporal differences from spatiotemporal features. It then utilizes the differences to excite the motion-sensitive channels of the features. The long-range temporal aggregations in previous works are typically achieved by stacking a large number of local temporal convolutions. Each convolution processes a local temporal window at a time. In contrast, the MTA module proposes to deform the local convolution to a group of sub-convolutions, forming a hierarchical residual architecture. Without introducing additional parameters, the features will be processed with a series of sub-convolutions, and each frame could complete multiple temporal aggregations with neighborhoods. The final equivalent receptive field of temporal dimension is accordingly enlarged, which is capable of modeling the long-range temporal relationship over distant frames. The two components of the TEA block are complementary in temporal modeling. Finally, our approach achieves impressive results at low FLOPs on several action recognition benchmarks, such as Kinetics, Something-Something, HMDB51, and UCF101, which confirms its effectiveness and efficiency. Yan Li 0043, Xintian Shi, Jianguo Zhang 0001, Bin Kang, Limin Wang 0002 |
CVPR | 4 |
| 2020 | Deep Class-Specific Affinity-Guided Convolutional Network for Multimodal Unpaired Image Segmentation
Jingkun Chen, Wenqi Li 0001, Hongwei Li 0004, Jianguo Zhang 0001 |
MICCAI (4) | 4 |
| 2020 | Abnormality Detection in Chest X-Ray Images Using Uncertainty Prediction Autoencoders
Feifei Xue, Jianguo Zhang 0001, Wei-Shi Zheng 0001, Hongmei Liu 0001 |
MICCAI (6) | 4 |
| 2020 | Deep kNN for Medical Image Classification
Jiaxin Zhuang, Jiabin Cai, Jianguo Zhang 0001, Wei-Shi Zheng 0001 |
MICCAI (1) | 4 |
| 2019 | Progressive Teacher-Student Learning for Early Action PredictionabstractThe goal of early action prediction is to recognize actions from partially observed videos with incomplete action executions, which is quite different from action recognition. Predicting early actions is very challenging since the partially observed videos do not contain enough action information for recognition. In this paper, we aim at improving early action prediction by proposing a novel teacherstudent learning framework. Our framework involves a teacher model for recognizing actions from full videos, a student model for predicting early actions from partial videos, and a teacher-student learning block for distilling progressive knowledge from teacher to student, crossing different tasks. Extensive experiments on three public action datasets show that the proposed progressive teacher-student learning framework can consistently improve performance of early action prediction model. We have also reported the state-of-the-art performances for early action prediction on all of these sets. Xionghui Wang, Jianfang Hu, Jian-Huang Lai, Jianguo Zhang 0001, Wei-Shi Zheng 0001 |
CVPR | 4 |
| 2019 | DiamondGAN: Unified Multi-modal Generative Adversarial Networks for MRI Sequences Synthesis
Hongwei Li 0004, Johannes C. Paetzold, Anjany Sekuboyina, Florian Kofler, Jianguo Zhang 0001, Jan Kirschke, Benedikt Wiestler, Bjoern Menze |
MICCAI (4) | 5 |
| 2019 | Early Action Prediction by Soft RegressionabstractWe propose a novel approach for predicting on-going action with the assistance of a low-cost depth camera. Our approach introduces a soft regression-based early prediction framework. In this framework, we estimate soft labels for the subsequences at different progress levels, jointly learned with an action predictor. Our formulation of soft regression framework 1) overcomes a usual assumption in existing early action prediction systems that the progress level of on-going sequence is given in the testing stage; and 2) presents a theoretical framework to better resolve the ambiguity and uncertainty of subsequences at early performing stage. The proposed soft regression framework is further enhanced in order to take the relationships among subsequences and the discrepancy of soft labels over different classes into consideration, so that a Multiple Soft labels Recurrent Neural Network (MSRNN) is finally developed. For real-time performance, we also introduce a new RGB-D feature called "local accumulative frame feature (LAFF)", which can be computed efficiently by constructing an integral feature map. Our experiments on three RGB-D benchmark datasets and an unconstrained RGB action set demonstrate that the proposed regression-based early action prediction model outperforms existing models significantly and also show that the early action prediction on RGB-D sequence is more accurate than that on RGB channel. Jianfang Hu, Wei-Shi Zheng 0001, Lianyang Ma, Gang Wang 0012, Jian-Huang Lai, Jianguo Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | Mixed Supervised Object Detection with Robust Objectness TransferabstractIn this paper, we consider the problem of leveraging existing fully labeled categories to improve the weakly supervised detection (WSD) of new object categories, which we refer to as mixed supervised detection (MSD). Different from previous MSD methods that directly transfer the pre-trained object detectors from existing categories to new categories, we propose a more reasonable and robust objectness transfer approach for MSD. In our framework, we first learn domain-invariant objectness knowledge from the existing fully labeled categories. The knowledge is modeled based on invariant features that are robust to the distribution discrepancy between the existing categories and new categories; therefore the resulting knowledge would generalize well to new categories and could assist detection models to reject distractors (e.g., object parts) in weakly labeled images of new categories. Under the guidance of learned objectness knowledge, we utilize multiple instance learning (MIL) to model the concepts of both objects and distractors and to further improve the ability of rejecting distractors in weakly labeled images. Our robust objectness transfer approach outperforms the existing MSD methods, and achieves state-of-the-art results on the challenging ILSVRC2013 detection dataset and the PASCAL VOC datasets. Yan Li 0043, Junge Zhang, Kaiqi Huang, Jianguo Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Standardized Assessment of Automatic Segmentation of White Matter Hyperintensities and Results of the WMH Segmentation ChallengeabstractQuantification of cerebral white matter hyperintensities (WMH) of presumed vascular origin is of key importance in many neurological research studies. Currently, measurements are often still obtained from manual segmentations on brain MR images, which is a laborious procedure. The automatic WMH segmentation methods exist, but a standardized comparison of the performance of such methods is lacking. We organized a scientific challenge, in which developers could evaluate their methods on a standardized multi-center/-scanner image dataset, giving an objective comparison: the WMH Segmentation Challenge. Sixty T1 + FLAIR images from three MR scanners were released with the manual WMH segmentations for training. A test set of 110 images from five MR scanners was used for evaluation. The segmentation methods had to be containerized and submitted to the challenge organizers. Five evaluation metrics were used to rank the methods: 1) Dice similarity coefficient; 2) modified Hausdorff distance (95th percentile); 3) absolute log-transformed volume difference; 4) sensitivity for detecting individual lesions; and 5) F1-score for individual lesions. In addition, the methods were ranked on their inter-scanner robustness; 20 participants submitted their methods for evaluation. This paper provides a detailed analysis of the results. In brief, there is a cluster of four methods that rank significantly better than the other methods, with one clear winner. The inter-scanner robustness ranking shows that not all the methods generalize to unseen scanners. The challenge remains open for future submissions and provides a public platform for method evaluation. Hugo J. Kuijf, Adrià Casamitjana, D. Louis Collins, Mahsa Dadar, Achilleas Georgiou, Mohsen Ghafoorian, Dakai Jin, April Khademi, Jesse Knight, Hongwei Li 0004, Xavier Lladó, J. Matthijs Biesbroek, Miguel Luna, Qaiser Mahmood, Richard McKinley, Alireza Mehrtash, Sébastien Ourselin, Bo-yong Park, Hyunjin Park, Simon Pezold, Élodie Puybareau, Jeroen de Bresser, Letícia Rittner, Carole H. Sudre, Sergi Valverde, Verónica Vilaplana, Roland Wiest, Yongchao Xu, Ziyue Xu 0004, Guodong Zeng, Jianguo Zhang 0001, Guoyan Zheng, Rutger Heinen, Christopher Li Hsian Chen, Wiesje M. van der Flier, Frederik Barkhof, Max A. Viergever, Geert Jan Biessels, Simon Andermatt, Mariana P. Bento, Matt Berseth, Mikhail Belyaev, Manuel Jorge Cardoso |
IEEE Trans. Medical Imaging | 32 |
| 2018 | Discriminative Learning of Latent Features for Zero-Shot RecognitionabstractZero-shot learning (ZSL) aims to recognize unseen image categories by learning an embedding space between image and semantic representations. For years, among existing works, it has been the center task to learn the proper mapping matrices aligning the visual and semantic space, whilst the importance to learn discriminative representations for ZSL is ignored. In this work, we retrospect existing methods and demonstrate the necessity to learn discriminative representations for both visual and semantic instances of ZSL. We propose an end-to-end network that is capable of 1) automatically discovering discriminative regions by a zoom network; and 2) learning discriminative semantic representations in an augmented space introduced for both user-defined and latent attributes. Our proposed method is tested extensively on two challenging ZSL datasets, and the experiment results show that the proposed method significantly outperforms state-of-the-art methods. Yan Li 0043, Junge Zhang, Jianguo Zhang 0001, Kaiqi Huang |
CVPR | 3 |
| 2018 | Deep Bilinear Learning for RGB-D Action Recognition
Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Jianguo Zhang 0001 |
ECCV (7) | 5 |
| 2018 | Global-Local Temporal Saliency Action PredictionabstractAction prediction on a partially observed action sequence is a very challenging task. To address this challenge, we first design a global-local distance model, where a global-temporal distance compares subsequences as a whole and local-temporal distance focuses on individual segment. Our distance model introduces temporal saliency for each segment to adapt its contribution. Finally, a global-local temporal action prediction model is formulated in order to jointly learn and fuse these two types of distances. Such a prediction model is capable of recognizing action of: 1) an on-going sequence and 2) a sequence with arbitrarily frames missing between the beginning and end (known as gap-filling). Our proposed model is tested and compared with related action prediction models on BIT, UCF11, and HMDB data sets. The results demonstrated the effectiveness of our proposal. In particular, we showed the benefit of our proposed model on predicting unseen action types and the advantage on addressing the gapfilling problem as compared with recently developed action prediction models. Shaofan Lai, Wei-Shi Zheng 0001, Jianfang Hu, Jianguo Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Structure Prediction for Gland Segmentation With Hand-Crafted and Deep Convolutional FeaturesabstractWe present a novel method to segment instances of glandular structures from colon histopathology images. We use a structure learning approach which represents local spatial configurations of class labels, capturing structural information normally ignored by sliding-window methods. This allows us to reveal different spatial structures of pixel labels (e.g., locations between adjacent glands, or far from glands), and to identify correctly neighboring glandular structures as separate instances. Exemplars of label structures are obtained via clustering and used to train support vector machine classifiers. The label structures predicted are then combined and post-processed to obtain segmentation maps. We combine hand-crafted, multi-scale image features with features computed by a deep convolutional network trained to map images to segmentation maps. We evaluate the proposed method on the public domain GlaS data set, which allows extensive comparisons with recent, alternative methods. Using the GlaS contest protocol, our method achieves the overall best performance. Siyamalan Manivannan, Wenqi Li 0001, Jianguo Zhang 0001, Emanuele Trucco, Stephen J. McKenna |
IEEE Trans. Medical Imaging | 3 |
| 2017 | A Multi-Task Deep Network for Person Re-IdentificationabstractPerson re-identification (ReID) focuses on identifying people across different scenes in video surveillance, which is usually formulated as a binary classification task or a ranking task in current person ReID approaches. In this paper, we take both tasks into account and propose a multi-task deep network (MTDnet) that makes use of their own advantages and jointly optimize the two tasks simultaneously for person ReID. To the best of our knowledge, we are the first to integrate both tasks in one network to solve the person ReID. We show that our proposed architecture significantly boosts the performance. Furthermore, deep architecture in general requires a sufficient dataset for training, which is usually not met in person ReID. To cope with this situation, we further extend the MTDnet and propose a cross-domain architecture that is capable of using an auxiliary set to assist training on small target sets. In the experiments, our approach outperforms most of existing person ReID algorithms on representative datasets including CUHK03, CUHK01, VIPeR, iLIDS and PRID2011, which clearly demonstrates the effectiveness of the proposed approach. Xiaotang Chen, Jianguo Zhang 0001, Kaiqi Huang |
AAAI | 3 |
| 2017 | Beyond Triplet Loss: A Deep Quadruplet Network for Person Re-identificationabstractPerson re-identification (ReID) is an important task in wide area video surveillance which focuses on identifying people across different cameras. Recently, deep learning networks with a triplet loss become a common framework for person ReID. However, the triplet loss pays main attentions on obtaining correct orders on the training set. It still suffers from a weaker generalization capability from the training set to the testing set, thus resulting in inferior performance. In this paper, we design a quadruplet loss, which can lead to the model output with a larger inter-class variation and a smaller intra-class variation compared to the triplet loss. As a result, our model has a better generalization ability and can achieve a higher performance on the testing set. In particular, a quadruplet deep network using a margin-based online hard negative mining is proposed based on the quadruplet loss for the person ReID. In extensive experiments, the proposed network outperforms most of the state-of-the-art algorithms on representative datasets which clearly demonstrates the effectiveness of our proposed method. Xiaotang Chen, Jianguo Zhang 0001, Kaiqi Huang |
CVPR | 3 |
| 2017 | Fully connected CRF with data-driven prior for multi-class brain tumor segmentationabstractGrid conditional random fields (CRFs) are widely applied in both natural and medical image segmentation tasks. However, they only consider the label coherence in neighborhood pixels or regions, which limits their ability to model long-range connections within the image and generally results in excessive smoothing of tumor boundaries. In this paper, we present a novel method for brain tumor segmentation in MR images based on fully-connected CRF (FC-CRF) model that establishes pairwise potentials on all pairs of pixels in the images. We employ a hierarchical approach to differentiate different structures of tumor and further formulate a FC-CRF model with learned data-driven prior knowledge of tumor core. The methods were evaluated on the testing and leaderboard set of Brain Tumor Image Segmentation Benchmark (BRATS) 2013 challenge. The precision of segmented tumor boundaries is improved significantly and the results are competitive compared to the start-of-the-arts. Haocheng Shen, Jianguo Zhang 0001 |
ICIP | 2 |
| 2017 | Efficient symmetry-driven fully convolutional network for multimodal brain tumor segmentationabstractIn this paper, we present a novel and efficient method for brain tumor (and sub regions) segmentation in multimodal MR images based on a fully convolutional network (FCN) that enables end-to-end training and fast inference. Our structure consists of a downsampling path and three upsampling paths, which extract multi-level contextual information by concatenating hierarchical feature representation from each upsam-pling path. Meanwhile, we introduce a symmetry-driven FCN by the proposal of using symmetry difference images. The model was evaluated on Brain Tumor Image Segmentation Benchmark (BRATS) 2013 challenge dataset and achieved the state-of-the-art results while the computational cost is less than competitors. Haocheng Shen, Jianguo Zhang 0001, Wei-Shi Zheng 0001 |
ICIP | 2 |
| 2017 | Boundary-Aware Fully Convolutional Network for Brain Tumor Segmentation
Haocheng Shen, Jianguo Zhang 0001, Stephen J. McKenna |
MICCAI (2) | 3 |
| 2017 | Jointly Learning Heterogeneous Features for RGB-D Activity RecognitionabstractIn this paper, we focus on heterogeneous features learning for RGB-D activity recognition. We find that features from different channels (RGB, depth) could share some similar hidden structures, and then propose a joint learning model to simultaneously explore the shared and feature-specific components as an instance of heterogeneous multi-task learning. The proposed model formed in a unified framework is capable of: 1) jointly mining a set of subspaces with the same dimensionality to exploit latent shared features across different feature channels, 2) meanwhile, quantifying the shared and feature-specific components of features in the subspaces, and 3) transferring feature-specific intermediate transforms (i-transforms) for learning fusion of heterogeneous features across datasets. To efficiently train the joint model, a three-step iterative optimization algorithm is proposed, followed by a simple inference model. Extensive experimental results on four activity datasets have demonstrated the efficacy of the proposed method. A new RGB-D activity dataset focusing on human-object interaction is further contributed, which presents more challenges for RGB-D activity benchmarking. Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Jianguo Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | HEp-2 specimen classification via deep CNNs and pattern histogramabstractAutomatic classification of Human Epithelial Type-2 (HEp-2) specimen patterns is an important yet challenging problem in medical image analysis. Most prior works have primarily focused on cells images classification problem which is one of the early essential steps in the system pipeline, while less attention has been paid to the classification of whole-specimen ones. In this work, a specimen pattern recognition system combining convolutional neural networks (CNNs) and pattern histogram was proposed. The pattern histograms were obtained based on the prediction of each single cell inside the specimens. Two strategies were designed to predicted the pattern of a whole specimen: 1) the most dominant cell pattern in pattern histogram was represented as the specimen pattern, 2) the pattern histograms were employed as bags of patterns and then were trained and predicted separately by a SVM classifier. Experimental results show that the proposed system is effective and achieves high classification accuracy on public benchmark datasets. We further evaluate the robustness of the proposed framework by testing trained CNNs on another different dataset, demonstrating that the system is robust to inter-lab data. Hongwei Li 0004, Wei-Shi Zheng 0001, Xiaohua Xie, Jianguo Zhang 0001 |
ICPR | 5 |
| 2016 | An automated pattern recognition system for classifying indirect immunofluorescence images of HEp-2 cells and specimensabstractImmunofluorescence antinuclear antibody tests are important for diagnosis and management of autoimmune conditions; a key step that would benefit from reliable automation is the recognition of subcellular patterns suggestive of different diseases. We present a system to recognize such patterns, at cellular and specimen levels, in images of HEp-2 cells. Ensembles of SVMs were trained to classify cells into six classes based on sparse encoding of texture features with cell pyramids, capturing spatial, multi-scale structure. A similar approach was used to classify specimens into seven classes. Software implementations were submitted to an international contest hosted by ICPR 2014 (Performance Evaluation of Indirect Immunofluorescence Image Analysis Systems). Mean class accuracies obtained on heldout test data sets were 87.1% and 88.5% for cell and specimen classification respectively. These were the highest achieved in the competition, suggesting that our methods are state-of-the-art. We provide detailed descriptions and extensive experiments with various features and encoding methods. Siyamalan Manivannan, Wenqi Li 0001, Shazia Akbar, Jianguo Zhang 0001, Stephen J. McKenna |
Pattern Recognit. | 5 |
| 2016 | Cross-Scenario Transfer Person ReidentificationabstractPerson reidentification (Re-ID) matches images of the same person captured in disjoint camera views and at different times. To obtain a reliable similarity measurement between images, manually annotating a large amount of pairwise cross-camera-view person images is deemed necessary. However, this kind of annotation is both costly and impractical for efficiently deploying a Re-ID system to a completely new scenario, a new setting of nonoverlapping camera views between which person images are to be matched. To solve this problem, we consider utilizing other existing person images captured in other scenarios to help the Re-ID system in a target (new) scenario, provided that a few samples are captured under the new scenario. More specifically, we tackle this problem by jointly learning the similarity measurements for Re-ID in different scenarios in an asymmetric way. To model the joint learning, we consider that the Re-ID models share certain component across tasks. A distinct consideration in our multitask modeling is to extract the discriminant shared component that reduces the cross-task data overlap in the shared latent space during the joint learning, so as to enhance the target inter-class separation in the shared latent space. For this purpose, we propose to maximize the cross-task data discrepancy on the shared component during asymmetric multitask learning (MTL), along with maximizing the local inter-class variation and minimizing local intra-class variation on all tasks. We call our proposed method the constrained asymmetric multitask discriminant component analysis (cAMT-DCA). We show that cAMT-DCA can be solved by a simple eigen decomposition with a closed form, getting rid of any iterative learning used in most conventional MTL analyses. The experimental results show that the proposed transfer model gains a clear improvement against the related nontransfer and general multitask person Re-ID models. Wei-Shi Zheng 0001, Xiang Li 0032, Jianguo Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Data Separation of L1-minimization for Real-time Motion Detection
Yu Liu 0008, Huaxin Xiao, Zheng Zhang 0011, Wei Xu 0019, Maojun Zhang, Jianguo Zhang 0001 |
BMVC | 6 |
| 2015 | Jointly learning heterogeneous features for RGB-D activity recognitionabstractIn this paper, we focus on heterogeneous feature learning for RGB-D activity recognition. Considering that features from different channels could share some similar hidden structures, we propose a joint learning model to simultaneously explore the shared and feature-specific components as an instance of heterogenous multi-task learning. The proposed model in an unified framework is capable of: 1) jointly mining a set of subspaces with the same dimensionality to enable the multi-task classifier learning, and 2) meanwhile, quantifying the shared and feature-specific components of features in the subspaces. To efficiently train the joint model, a three-step iterative optimization algorithm is proposed, followed by two inference models. Extensive results on three activity datasets have demonstrated the efficacy of the proposed method. In addition, a novel RGB-D activity dataset focusing on human-object interaction is collected for evaluating the proposed method, which will be made available to the community for RGB-D activity benchmarking and analysis. Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Jianguo Zhang 0001 |
CVPR | 4 |
| 2015 | Multiple Instance Cancer Detection by Boosting Regularised Trees
Wenqi Li 0001, Jianguo Zhang 0001, Stephen J. McKenna |
MICCAI (1) | 2 |
| 2015 | Discriminating dysplasia: Optical tomographic texture analysis of colorectal polyps
Wenqi Li 0001, Maria Coats, Jianguo Zhang 0001, Stephen J. McKenna |
Medical Image Anal. | 3 |
| 2015 | Efficient Video Stitching Based on Fast Structure DeformationabstractIn computer vision, video stitching is a very challenging problem. In this paper, we proposed an efficient and effective wide-view video stitching method based on fast structure deformation that is capable of simultaneously achieving quality stitching and computational efficiency. For a group of synchronized frames, firstly, an effective double-seam selection scheme is designed to search two distinct but structurally corresponding seams in the two original images. The seam location of the previous frame is further considered to preserve the interframe consistency. Secondly, along the double seams, 1-D feature detection and matching is performed to capture the structural relationship between the two adjacent views. Thirdly, after feature matching, we propose an efficient algorithm to linearly propagate the deformation vectors to eliminate structure misalignment. At last, image intensity misalignment is corrected by rapid gradient fusion based on the successive over relaxation iteration (SORI) solver. A principled solution to the initialization of the SORI significantly reduced the number of iterations required. We have compared favorably our method with seven state-of-the-art image and video stitching algorithms as well as traditional ones. Experimental results show that our method outperforms the existing ones compared in terms of overall stitching quality and computational efficiency. Jing Li 0014, Wei Xu 0019, Jianguo Zhang 0001, Maojun Zhang, Zhengming Wang, Xuelong Li 0001 |
IEEE Trans. Cybern. | 3 |
| 2015 | Learning Person-Person Interaction in Collective Activity RecognitionabstractCollective activity is a collection of atomic activities (individual person's activity) and can hardly be distinguished by an atomic activity in isolation. The interactions among people are important cues for recognizing collective activity. In this paper, we concentrate on modeling the person-person interactions for collective activity recognition. Rather than relying on hand-craft description of the person-person interaction, we propose a novel learning-based approach that is capable of computing the class-specific person-person interaction patterns. In particular, we model each class of collective activity by an interaction matrix, which is designed to measure the connection between any pair of atomic activities in a collective activity instance. We then formulate an interaction response (IR) model by assembling all these measurements and make the IR class specific and distinct from each other. A multitask IR is further proposed to jointly learn different person-person interaction patterns simultaneously in order to learn the relation between different person-person interactions and keep more distinct activity-specific factor for each interaction at the same time. Our model is able to exploit discriminative low-rank representation of person-person interaction. Experimental results on two challenging data sets demonstrate our proposed model is comparable with the state-of-the-art models and show that learning person-person interactions plays a critical role in collective activity recognition. Xiaobin Chang, Wei-Shi Zheng 0001, Jianguo Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | Gait Based Gender Recognition Using Sparse Spatio Temporal Features
Matthew Collins, Paul Miller 0003, Jianguo Zhang 0001 |
MMM (2) | 3 |
| 2014 | Using a Discrete Hidden Markov Model Kernel for lip-based biometric identification
Carlos Manuel Travieso-González, Jianguo Zhang 0001, Paul Miller 0003, Jesús B. Alonso |
Image Vis. Comput. | 2 |
| 2013 | Learning from Partially Annotated OPT Images by Contextual Relevance Ranking
Wenqi Li 0001, Jianguo Zhang 0001, Wei-Shi Zheng 0001, Maria Coats, Frank A. Carey, Stephen J. McKenna |
MICCAI (3) | 2 |
| 2013 | Angle consistency for registration between catadioptric omni-images and orthorectified aerial imagesabstractRegistration between catadioptric omni‐images and orthorectified aerial images is the key step to integrate them to achieve three‐dimensional urban construction. This problem becomes very challenging because of the non‐linearity of the imaging model of catadioptric omni‐cameras. In this study, the authors attempt to address this problem. The authors first study the properties of horizontal line structure under catadioptric omni‐cameras to prove and extend the theorem of catadioptric distance, and then present angle consistency of horizontal lines between a catadioptric omni‐image and an orthorectified aerial image. The authors further employ them to achieve registration between catadioptric omni‐images and orthorectified aerial images. To the best of authors’ knowledge, this study has not been done before. Experimental results on both simulated data and real scene images confirm the effectiveness of this approach. Wei Xu 0019, Jianguo Zhang 0001, Maojun Zhang |
IET Image Process. | 3 |
| 2013 | "Pattern Recognition" special issue: Sparse representation for event recognition in video surveillance
Huiyu Zhou 0001, Jianguo Zhang 0001, Liang Wang 0001, Zhengyou Zhang, Lisa M. Brown |
Pattern Recognit. | 2 |
| 2012 | Relevance feedback for real-world human action retrieval
Ling Shao 0001, Jianguo Zhang 0001, Yan Liu 0004 |
Pattern Recognit. Lett. | 3 |
| 2012 | Human action segmentation and recognition via motion and shape analysis
Ling Shao 0001, Ling Ji, Yan Liu 0004, Jianguo Zhang 0001 |
Pattern Recognit. Lett. | 4 |
| 2011 | Age classification using Radon transform and entropy based scaling SVMabstractThis paper mainly addresses the problem of age classification. Image fe atures can be extracted using a difference of Gaussian filter followed by Radon tran sform. The relevance and importance of these features are determined in a scaling support vector machine classifier, where zero weights are assigned to irrelevant varia bles. To enhance the quality of feature selection, we introduce entropy estimation to the scaling c lassifier. Experimental results demonstrate that the proposed algorithm leads to better recognition accuracy than the state of the art. Huiyu Zhou 0001, Paul Miller 0003, Jianguo Zhang 0001 |
BMVC | 3 |
| 2011 | Modeling and representing events in multimediaabstractThis paper presents an overview of the Joint Workshop on Modeling and Representing Events (JMRE), which is held as part of ACM Multimedia 2011. JMRE is concerned with the understanding of events from multimedia, and with using events in order to better organize and consume multimedia. Vasileios Mezaris, Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli, Huiyu Zhou 0001, Jianguo Zhang 0001, Liang Wang 0001, Zhengyou Zhang |
ACM Multimedia | 6 |
| 2011 | Bimodal biometric verification based on face and lips
Carlos Manuel Travieso-González, Jianguo Zhang 0001, Paul Miller 0003, Jesús B. Alonso, Miguel A. Ferrer |
Neurocomputing | 2 |
| 2010 | Intelligent Sensor Information System For Public Transport - To Safely GoabstractThe Intelligent Sensor Information System (ISIS) is described. ISIS is an active CCTV approach to reducing crime and anti-social behavior on public transport systems such as buses. Key to the system is the idea of event composition, in which directly detected atomic events are combined to infer higher-level events with semantic meaning. Video analytics are described that profile the gender of passengers and track them as they move about a 3-D space. The overall system architecture is described which integrates the on-board event recognition with the control room software over a wireless network to generate a real-time alert. Data from preliminary data-gathering trial is presented. Paul Miller 0003, Weiru Liu, Chris Fowler, Huiyu Zhou 0001, Jiali Shen, Jianbing Ma, Jianguo Zhang 0001, Wei Qi Yan 0001, Kieran McLaughlin, Sakir Sezer |
AVSS | 7 |
| 2010 | Inter-frame contextual modelling for visual speech recognitionabstractIn this paper, we present a new approach to visual speech recognition which improves contextual modelling by combining Inter-Frame Dependent and Hidden Markov Models. This approach captures contextual information in visual speech that may be lost using a Hidden Markov Model alone. We apply contextual modelling to a large speaker independent isolated digit recognition task, and compare our approach to two commonly adopted feature based techniques for incorporating speech dynamics. Results are presented from baseline feature based systems and the combined modelling technique. We illustrate that both of these techniques achieve similar levels of performance when used independently. However significant improvements in performance can be achieved through a combination of the two. In particular we report an improvement in excess of 17% relative Word Error Rate in comparison to our best baseline system. Adrian Pass, Ji Ming, Philip Hanna 0001, Jianguo Zhang 0001, Darryl Stewart |
ICIP | 4 |
| 2010 | AN investigation into features for multi-view lipreadingabstractFor the first time in this paper we present results showing the effect of speaker head pose angle on automatic lip-reading performance over a wide range of closely spaced angles. We analyse the effect head pose has upon the features themselves and show that by selecting coefficients with minimum variance w.r.t. pose angle, recognition performance can be improved when train-test pose angles differ. Experiments are conducted using the initial phase of a unique multi view Audio-Visual database designed specifically for research and development of pose-invariant lip-reading systems. We firstly show that it is the higher order horizontal spatial frequency components that become most detrimental as the pose deviates. Secondly we assess the performance of different feature selection masks across a range of pose angles including a new mask based on Minimum Cross-Pose Variance coefficients. We report a relative improvement of 50% in Word Error Rate when using our selection mask over a common energy based selection during profile view lip-reading. Adrian Pass, Jianguo Zhang 0001, Darryl Stewart |
ICIP | 2 |
| 2010 | Feature selection for pose invariant lip biometricsabstractFor the first time in this paper we present results showing the effect of out of plane speaker head pose variation on a lip based speaker verification system. Using appearance DCT based features, we adopt a Mutual Information analysis technique to highlight the class discriminant DCT components most robust to changes in out of plane pose. Experiments are conducted using the initial phase of a new multi view Audio-Visual database designed for research and development of pose-invariant speech and speaker recognition. We show that verification performance can be improved by substituting higher order horizontal DCT components for vertical, particularly in the case of a train/test pose angle mismatch. We further show that the best performance can be achieved by combining this alternative feature selection with multi view training, reporting a relative 45% Equal Error Rate reduction over a common energy based selection. © 2010 ISCA. Adrian Pass, Jianguo Zhang 0001, Darryl Stewart |
INTERSPEECH | 2 |
| 2010 | Action categorization by structural probabilistic latent semantic analysis
Jianguo Zhang 0001, Shaogang Gong |
Comput. Vis. Image Underst. | 1 |
| 2010 | Action categorization with modified hidden conditional random field
Jianguo Zhang 0001, Shaogang Gong |
Pattern Recognit. | 1 |
| 2009 | People detection in low-resolution video with non-stationary background
Jianguo Zhang 0001, Shaogang Gong |
Image Vis. Comput. | 1 |
| 2007 | Local Features and Kernels for Classification of Texture and Object Categories: A Comprehensive Study
Jianguo Zhang 0001, Marcin Marszalek, Svetlana Lazebnik, Cordelia Schmid |
Int. J. Comput. Vis. | 1 |
| 2003 | Affine invariant classification and retrieval of texture images
Jianguo Zhang 0001, Tieniu Tan |
Pattern Recognit. | 1 |
| 2002 | Brief review of invariant texture analysis methods
Jianguo Zhang 0001, Tieniu Tan |
Pattern Recognit. | 1 |
| 2001 | Affine invariant texture signaturesabstractWe develop a new approach for texture classification independent of affine transforms. Based on a spectral representation of texture images under affine transform, anisotropic scale invariant signatures of the orientation spectrum distribution are extracted. A peaks distribution vector (PDV) obtained on the distribution of these signatures captures texture properties invariant to affine distortion. The PDV is used to measure the similarity between textures. Experimental results show the efficiency of the PDV for affine invariant texture classification. Jianguo Zhang 0001, Tieniu Tan |
ICIP (2) | 1 |