Xiantong Zhen

dblp:78/10651 · DBLP profile ↗
← Back
144ranked-venue papers
20as first author
72since 2021 · last 2026
0000-0001-5213-0462ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 81 · 8 first-author · 32 since 2021Artificial intelligence and machine learning · 76 · 12 first-author · 44 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 UniAD: Unified cross-modal prompt regularization for zero-shot anomaly detection across domains
Xuezhi Xiang, Songran Luo, Yingjun Du, Lei Zhang 0093, Xiantong Zhen
Neurocomputing5
2026 HySaDe-Mamba: A Mamba-Based Network for Hyperspectral Salient Object Detection
abstract
Hyperspectral salient object detection (HSOD) aims to identify visually and spectrally distinctive regions in hyperspectral images (HSIs). However, existing HSOD methods often suffer from spectral redundancy and inefficient spatial-spectral modeling, which hinder their scalability and accuracy in complex scenes. To tackle these challenges, we propose HySaDe-Mamba, a novel HSOD framework built upon the Mamba architecture. Specifically, to address information redundancy in HSI, we design a spatial-enhanced spectral-embedding (SeSe) module, which maps high-dimensional data into a more compact but effective representation. On the compact SeSe representation features, we further propose a Bi-scale spatial and Bi-directional spectral (BsBd) Mamba module, performing the selective scanning mechanism in a spatial-spectral hybrid, end-to-end way, which not only facilitates comprehensive spatial structural interaction across both global and local scales, but also effectively exploits the underlying spectral semantic correlation. Extensive experiments on two public HSOD datasets demonstrate that our HySaDe-Mamba achieves state-of-the-art detection accuracy across seven metrics, while maintaining an efficient inference speed of 40.22 FPS. The source code is publicly available at https://github.com/Leezl/HySaDe-Mamba.
Lei Zhang 0109, Xiaoyan Luo, Liheng Bian, Xiantong Zhen
IEEE Trans. Circuits Syst. Video Technol.5
2026 Diagnosing and Improving Vector-Quantization-Based Blind Image Restoration
abstract
Vector-Quantization (VQ) based discrete generative models are widely used to learn powerful high-quality (HQ) priors for blind image restoration (BIR). In this paper, we diagnose the side-effects of discrete VQ process essential to VQ-based BIR methods: 1) confining the representation capacity of HQ codebook, 2) being error-prone for code index prediction on low-quality (LQ) images, and 3) under-valuing the importance of input LQ image. These motivate us to learn continuous feature representation of HQ codebook for better restoration performance than using discrete VQ process. To further improve the restoration fidelity, we propose a new Self-in-Cross-Attention (SinCA) module to augment the HQ codebook with the feature of input LQ image, and perform cross-attention between LQ feature and input-augmented codebook. By this way, our SinCA leverages the input LQ image to enhance the representation of codebook for restoration fidelity. Experiments on four typical VQ-based BIR methods demonstrate that, by replacing the VQ process with a transformer using our SinCA, they achieve better quantitative and qualitative performance on blind image super-resolution and blind face restoration. The code and pre-trained models are publicly released at https://github.com/lhy-85/SinCA.
Zengyou Wang, Xiantong Zhen, Ran Gu, David Zhang 0001, Jun Xu 0019
IEEE Trans. Image Process.4
2025 HYDEN: Hyperbolic Density Representations for Medical Images and Reports
abstract
In light of the inherent entailment relations between images and text, embedding point vectors in hyperbolic space has been employed to leverage its hierarchical modeling advantages for visual semantic representation learning. However, point vector embeddings struggle to address semantic uncertainty, where an image may have multiple interpretations, and text may correspond to different images—a challenge especially prevalent in the medical domain. Therefor, we propose HYDEN, a novel hyperbolic density embedding based image-text representation learning approach tailored for specific medical domain data. This method integrates text-aware local features with global features from images, mapping image-text features to density features in hyperbolic space via using hyperbolic pseudo-Gaussian distributions. An encapsulation loss function is employed to model the partial order relations between image-text density distributions. Experimental results demonstrate the interpretability of our approach and its superior performance compared to the baseline methods across various zero-shot tasks and fine-tuning task on different datasets.
Linbin Han, Xiantong Zhen, Jia-Hong Gao 0001
COLING3
2025 LKA-ReID: Vehicle Re-Identification with Large Kernel Attention
abstract
With the rapid development of intelligent transportation systems and the popularity of smart city infrastructure, Vehicle Re-ID technology has become an important research field. The vehicle Re-ID task faces an important challenge, which is the high similarity between different vehicles. Existing methods use additional detection or segmentation models to extract differentiated local features. However, these methods either rely on additional annotations or greatly increase the computational cost. Using attention mechanism to capture global and local features is crucial to solve the challenge of high similarity between classes in vehicle Re-ID tasks. In this paper, we propose LKA-ReID with large kernel attention. Specifically, the large kernel attention (LKA) utilizes the advantages of self-attention and also benefits from the advantages of convolution, which can extract the global and local features of the vehicle more comprehensively. We also introduce hybrid channel attention (HCA), which combines channel attention with spatial information, so that the model can better focus on channels and feature regions, and ignore background and other disturbing information. Experiments on VeRi-776 dataset demonstrated the effectiveness of LKA-ReID, with mAP reaches 86.65% and Rank-1 reaches 98.03%.
Xuezhi Xiang, Zhushan Ma, Lei Zhang 0093, Denis Ombati, Himaloy Himu, Xiantong Zhen
ICASSP6
2025 Mamba-SF: Monocular Scene Flow Learning with State Space Models
abstract
Monocular scene flow estimation has been a long-standing problem in computer vision. Methods based on the RAFT architecture are currently the mainstream approaches, while often overlooking the long-range dependencies in motion and texture features and fail to fully utilize the spatial information in texture features. In this paper, we consider that using Transformers introduces high computational complexity. Therefore, we propose the Mamba Motion Module based on State Space Models design, which first models long-range dependencies in motion and texture features and then fully leverages the spatial information in texture features to enhance the motion features, generating global motion features while maintaining low computational complexity. Additionally, texture features play a significant role in constraining motion boundaries. Therefore, we propose the Enhanced Texture Module, which predicts a set of channel weights from the global motion features to enrich the channel properties of the texture features and concatenates texture features with the global motion features along the channels, thereby improving scene flow accuracy. Experimental results show that our method achieves highly competitive results on the KITTI 2015 and Eigen Split datasets, increasing by 18.82% and 2.15% compared to the baseline, respectively.
Xuezhi Xiang, Xianye Ben, Insha Hassan, Mingliang Zhai, Lei Zhang 0093, Xiantong Zhen
ICIP7
2025 Refining and adaption: multi-modal learning in wear debris analysis
abstract
Wear debris analysis(WDA) is a critical predictive maintenance technique that provides key data for diagnosing wear faults in mechanical equipment. This technique helps ensure long-term stable operation and fault prediction for machinery. However, recent approaches in WDA rely heavily on data-driven methods and often neglect the most important expert knowledge, such as the morphology, size, and color of debris. This knowledge is essential for fault diagnosis, but traditional networks struggle to represent and integrate it effectively, in contrast, large language models (LLMs) and vision-language models (VLMs), with their exceptional generalization capabilities, are rapidly evolving into ideal tools for representing and integrating expert knowledge. Inspired by this, we propose a novel multimodal learning method that combines VLMs and LLMs to integrate expert knowledge into image information, extending general models to specialized industrial domains. Specifically, we use an LLMs to regenerate textual descriptions from expert knowledge and refine CLIP’s Layer Normalization parameters to enhance the model’s transferability to specialized datasets, such as ferrography images, for zero-shot learning. Additionally, we introduce an adaptive meta-learning strategy that allows the model to quickly adapt to ferrography images with limited samples, further improving performance. Experimental results on industrial datasets show that the proposed method effectively addresses the challenge of applying VLMs in data-limited specialized domains, significantly improving the accuracy and robustness of WDA.
Yongcai Chen, Fang Lei, Xiantong Zhen, Xin Li 0100, Lei Zhang 0093
IJCNN3
2025 EEG-DINO: Learning EEG Foundation Models via Hierarchical Self-distillation
Xujia Wang, Xuhui Liu, Qian Si, Zhaoliang Xu, Yang Li 0010, Xiantong Zhen
MICCAI (1)7
2025 Variational Task Vector Composition
abstract
Task vectors capture how a model changes during fine-tuning by recording the difference between pre-trained and task-specific weights. The composition of task vectors, a key operator in task arithmetic, enables models to integrate knowledge from multiple tasks without incurring significant additional inference costs. In this paper, we propose variational task vector composition (VTVC), where composition coefficients are taken as latent variables and estimated in a Bayesian inference framework. Unlike previous methods that operate at the task level, our framework focuses on sample-specific composition. Motivated by the observation of structural redundancy in task vectors, we introduce a Spike-and-Slab prior that promotes sparsity and aims to preserve the most informative components. To further address the high variance and sampling inefficiency in sparse, high-dimensional spaces, we develop a gated sampling mechanism that constructs a controllable posterior by filtering the composition coefficients based on both uncertainty and importance. This yields a more stable and interpretable variational framework by deterministically selecting reliable task components, reducing sampling variance while improving transparency and generalization. Experimental results demonstrate that our method achieves state-of-the-art average performance across a diverse range of benchmarks, including image classification and natural language understanding. These findings highlight the practical value of our approach, offering a new, efficient, and effective framework for task vector composition.
Yingjun Du, Xiantong Zhen, Ling Shao 0001
NeurIPS3
2025 GeneralizeFormer: Layer-Adaptive Model Generation Across Test-Time Distribution Shifts
abstract
We consider the problem of test-time domain generalization, where a model is trained on several source domains and adjusted on target domains never seen during training. Different from the common methods that fine-tune the model or adjust the classifier parameters online, we propose to generate multiple layer parameters on the fly during inference by a lightweight meta-learned transformer, which we call GeneralizeFormer: The layer-wise parameters are generated per target batch without fine-tuning or online adjustment. By doing so, our method is more effective in dynamic scenarios with multiple target distributions and also avoids forgetting valuable source distribution characteristics. Moreover, by considering layer-wise gradients, the proposed method adapts itself to various distribution shifts. To reduce the computational and time cost, we fix the convolutional parameters while only generating parameters of the Batch Normalization layers and the linear classifier. Experiments on six widely used domain generalization datasets demonstrate the benefits and abilities of the proposed method to efficiently handle various distribution shifts, generalize in dynamic scenarios, and avoid forgetting. Our code is available: https://github.com/ambekarsameer96/generalizeformer
Sameer Ambekar, Zehao Xiao, Xiantong Zhen, Cees Snoek
WACV3
2025 Learning temporal-aware representation for controllable interventional radiology imaging
Wei Si, Zhaolin Zheng, Zhewei Huang, Ximing Xu 0002, Ruijue Wang, Ji-Gang Bao, Xiantong Zhen, Jun Xu 0019
Comput. Vis. Image Underst.8
2025 Learning to Generalize Heterogeneous Representation for Cross-Modality Image Synthesis via Multiple Domain Interventions
Yawen Huang, Huimin Huang 0002, Hao Zheng 0008, Yuexiang Li, Feng Zheng 0001, Xiantong Zhen, Yefeng Zheng 0001
Int. J. Comput. Vis.6
2025 Few-Shot Referring Video Single- and Multi-Object Segmentation Via Cross-Modal Affinity with Instance Sequence Matching
Heng Liu 0002, Mingqi Gao 0003, Xiantong Zhen, Feng Zheng 0001, Yang Wang 0023
Int. J. Comput. Vis.4
2025 D3T: Dual-Domain Diffusion Transformer in Triplanar Latent Space for 3D Incomplete-View CT Reconstruction
Xuhui Liu, Hong Li 0016, Yawen Huang, Xiantong Zhen, Baochang Zhang 0001
Int. J. Comput. Vis.8
2025 Deep scene flow learning from point cloud with Transformer
Xuezhi Xiang, Rokia Abdein, Lei Zhang 0093, Xiantong Zhen
Neurocomputing5
2025 PVFT-Net: A point-voxel fusion method for self-supervised scene flow estimation with transformer
Xuezhi Xiang, Xiaoheng Li, Xiankun Zhou, Lei Zhang 0093, Xiantong Zhen
Neurocomputing6
2025 Joint super-resolution and inverse tone-mapping: A feature decomposition aggregation network and a new benchmark
Xiantong Zhen, Jun Xu 0019
Neurocomputing4
2025 De-noising mask transformer for referring image segmentation
Yehui Wang, Fang Lei, Baoyan Wang, Xiantong Zhen, Lei Zhang 0093
Image Vis. Comput.5
2025 Vehicle re-identification with large separable kernel attention and hybrid channel attention
Xuezhi Xiang, Zhushan Ma, Xiaoheng Li, Lei Zhang 0093, Xiantong Zhen
Image Vis. Comput.5
2025 Self-supervised monocular depth estimation with large kernel attention and dynamic scene perception
Xuezhi Xiang, Xiaoheng Li, Lei Zhang 0093, Xiantong Zhen
J. Vis. Commun. Image Represent.5
2024 DiffuX2CT: Diffusion Learning to Reconstruct CT Images from Biplanar X-Rays
Xuhui Liu, Runkun Liu, Hong Li 0016, Xiantong Zhen, Baochang Zhang 0001
ECCV (43)6
2024 Deep Optical Flow Learning With Deformable Large-Kernel Cross-Attention
abstract
Optical flow estimation from image sequences is a fundamental problem in computer vision. In recent years, some methods have utilized Transformer to model global dependencies and improve optical flow, achieving impressive performance. However, in these methods, Transformers typically treat two-dimensional image features as one-dimensional sequences. While position encoding partially mitigates the loss of position information between different feature patches, Transformer still lacks inherent biases for modeling local visual patterns and tend to overlook channel characteristics in image features. Therefore, this paper introduces a deformable large kernel attention module, combining the strengths of convolution and attention mechanisms, which can preserve feature channel adaptability while modeling global dependencies without compromising the two-dimensional structure of features, significantly enhancing optical flow estimation. Additionally, the introduced deformable mechanism allows the model to adapt appropriately to different data patterns. Experimental results demonstrate that our optical flow estimation method achieves competitive results on publicly available benchmarks such as Sintel and KITTI.
Xuezhi Xiang, Denis Ombati, Lei Zhang 0093, Xiantong Zhen
ICIP5
2024 Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing Backpropagation
abstract
Fine-tuning pretrained large models to downstream tasks is an important problem, which however suffers from huge memory overhead due to large-scale parameters. This work strives to reduce memory overhead in fine-tuning from perspectives of activation function and layer normalization. To this end, we propose the Approximate Backpropagation (Approx-BP) theory, which provides the theoretical feasibility of decoupling the forward and backward passes. We apply our Approx-BP theory to backpropagation training and derive memory-efficient alternatives of GELU and SiLU activation functions, which use derivative functions of ReLUs in the backward pass while keeping their forward pass unchanged. In addition, we introduce a Memory-Sharing Backpropagation strategy, which enables the activation memory to be shared by two adjacent layers, thereby removing activation memory usage redundancy. Our method neither induces extra computation nor reduces training efficiency. We conduct extensive experiments with pretrained vision and language models, and the results demonstrate that our proposal can reduce up to $\sim$$30\%$ of the peak memory usage. Our code is released at [github](https://github.com/yyyyychen/LowMemoryBP).
Yingdong Shi, Cheems Wang, Xiantong Zhen, Jun Xu 0019
ICML4
2024 Multi-Constraint Transferable Generative Adversarial Networks for Cross-Modal Brain Image Synthesis
Yawen Huang, Hao Zheng 0008, Yuexiang Li, Feng Zheng 0001, Xiantong Zhen, Guo-Jun Qi, Ling Shao 0001, Yefeng Zheng 0001
Int. J. Comput. Vis.5
2024 Lightweight improved residual network for efficient inverse tone mapping
Liqi Xue, Yongbao Song, Yan Liu 0004, Lei Zhang 0093, Xiantong Zhen, Jun Xu 0019
Multim. Tools Appl.6
2024 MetaKernel: Learning Variational Random Features With Limited Labels
abstract
Few-shot learning deals with the fundamental and challenging problem of learning from a few annotated samples, while being able to generalize well on new tasks. The crux of few-shot learning is to extract prior knowledge from related tasks to enable fast adaptation to a new task with a limited amount of data. In this paper, we propose meta-learning kernels with random Fourier features for few-shot learning, we call MetaKernel. Specifically, we propose learning variational random features in a data-driven manner to obtain task-specific kernels by leveraging the shared knowledge provided by related tasks in a meta-learning setting. We treat the random feature basis as the latent variable, which is estimated by variational inference. The shared knowledge from related tasks is incorporated into a context inference of the posterior, which we achieve via a long-short term memory module. To establish more expressive kernels, we deploy conditional normalizing flows based on coupling layers to achieve a richer posterior distribution over random Fourier bases. The resultant kernels are more informative and discriminative, which further improves the few-shot learning. To evaluate our method, we conduct extensive experiments on both few-shot image classification and regression tasks. A thorough ablation study demonstrates that the effectiveness of each introduced component in our method. The benchmark results on fourteen datasets demonstrate MetaKernel consistently delivers at least comparable and often better performance than state-of-the-art alternatives.
Yingjun Du, Haoliang Sun, Xiantong Zhen, Jun Xu 0019, Yilong Yin, Ling Shao 0001, Cees Snoek
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 On the Number of Linear Regions of Convolutional Neural Networks With Piecewise Linear Activations
abstract
One fundamental problem in deep learning is understanding the excellent performance of deep Neural Networks (NNs) in practice. An explanation for the superiority of NNs is that they can realize a large family of complicated functions, i.e., they have powerful expressivity. The expressivity of a Neural Network with Piecewise Linear activations (PLNN) can be quantified by the maximal number of linear regions it can separate its input space into. In this paper, we provide several mathematical results needed for studying the linear regions of Convolutional Neural Networks with Piecewise Linear activations (PLCNNs), and use them to derive the maximal and average numbers of linear regions for one-layer PLCNNs. Furthermore, we obtain upper and lower bounds for the number of linear regions of multi-layer PLCNNs. Our results suggest that deeper PLCNNs have more powerful expressivity than shallow PLCNNs, while PLCNNs have more expressivity than fully-connected PLNNs per parameter, in terms of the number of linear regions.
Huan Xiong, Lei Huang 0015, Wenston J. T. Zang, Xiantong Zhen, Guosen Xie, Bin Gu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Gaze Estimation by Attention-Induced Hierarchical Variational Auto-Encoder
abstract
Appearance-based gaze estimation has been widely studied recently with promising performance. The majority of appearance-based gaze estimation methods are developed under the deterministic frameworks. However, the deterministic gaze estimation methods suffer from large performance drop upon challenging eye images in low-resolution, darkness, partial occlusions, etc. To alleviate this problem, in this article, we alternatively reformulate the appearance-based gaze estimation problem under a generative framework. Specifically, we propose a variational inference model, that is, variational gaze estimation network (VGE-Net), to generate multiple gaze maps as complimentary candidates simultaneously supervised by the ground-truth gaze map. To achieve robust estimation, we adaptively fuse the gaze directions predicted on these candidate gaze maps by a regression network through a simple attention mechanism. Experiments on three benchmarks, that is, MPIIGaze, EYEDIAP, and Columbia, demonstrate that our VGE-Net outperforms state-of-the-art gaze estimation methods, especially on challenging cases. Comprehensive ablation studies also validate the effectiveness of our contributions. The code will be publicly released.
Guanhe Huang, Jingyue Shi, Jun Xu 0019, Jing Li 0027, Shengyong Chen, Yingjun Du, Xiantong Zhen, Honghai Liu 0001
IEEE Trans. Cybern.7
2024 Weakly-Supervised RGBD Video Object Segmentation
abstract
Depth information opens up new opportunities for video object segmentation (VOS) to be more accurate and robust in complex scenes. However, the RGBD VOS task is largely unexplored due to the expensive collection of RGBD data and time-consuming annotation of segmentation. In this work, we first introduce a new benchmark for RGBD VOS, named DepthVOS, which contains 350 videos (over 55k frames in total) annotated with masks and bounding boxes. We futher propose a novel, strong baseline model - Fused Color-Depth Network (FusedCDNet), which can be trained solely under the supervision of bounding boxes, while being used to generate masks with a bounding box guideline only in the first frame. Thereby, the model possesses three major advantages: a weakly-supervised training strategy to overcome the high-cost annotation, a cross-modal fusion module to handle complex scenes, and weakly-supervised inference to promote ease of use. Extensive experiments demonstrate that our proposed method performs on par with top fully-supervised algorithms. We will open-source our project on https://github.com/yjybuaa/depthvos/ to facilitate the development of RGBD VOS.
Mingqi Gao 0003, Feng Zheng 0001, Xiantong Zhen, Rongrong Ji, Ling Shao 0001, Ales Leonardis
IEEE Trans. Image Process.4
2024 Variational Neuron Shifting for Few-Shot Image Classification Across Domains
abstract
Few-shot image classification aims to recognize unseen classes with few labeled samples. Existing meta-learning models learn the ability of learning good representation or model parameters, in order to adapt to new tasks with a few training samples. However, when there exists a domain gap between training and test tasks, the learned ability often does not generalize well across domains, resulting in degraded performance on new tasks. In this article, we propose variational neuron shifting to generate adapted feature representations for few-shot learning. To do so, we introduce a working memory module to store the shifted neurons from the support set, which will be accessed to generate adapted feature representations of query samples. Under the meta-learning paradigm, the model is learned to acquire the ability of adaptation with single sample at meta-training time so as to further adapt itself to each single test sample at meta-test time. We formulate the adaptation process as a variational Bayesian inference problem, which incorporates the test sample as the condition into the generation of the model neuron shifting. We conduct extensive experiments on both within and across domain few-shot classification tasks. The new state-of-the-art performance substantiates the effectiveness of our variational neuron shifting. The thorough ablation studies further demonstrate the benefit of each component in our model.
Liyun Zuo, Baoyan Wang, Lei Zhang 0093, Jun Xu 0019, Xiantong Zhen
IEEE Trans. Multim.5
2023 Coarse-Fine View Attention Alignment-Based GAN for CT Reconstruction from Biplanar X-Rays
abstract
For surgical planning and intra-operation imaging, CT reconstruction using X-ray images can potentially be an important alternative when CT imaging is not available or not feasible. In this paper, we aim to use biplanar X-rays to reconstruct a 3D CT image, because biplanar X-rays convey richer information than single-view X-rays and are more commonly used by surgeons. Different from previous studies in which the two X-ray views were treated indifferently when fusing the cross-view data, we propose a novel attention-informed coarse-to-fine cross-view fusion method to combine the features extracted from the orthogonal biplanar views. This method consists of a view attention alignment sub-module and a fine-distillation sub-module that are designed to work together to highlight the unique or complementary information from each of the views. Experiments have demonstrated the superiority of our proposed method over the SOTA methods.
Hanqiang Ouyang, Dongheng Chu, Huishu Yuan, Xiantong Zhen, Pei Dong
BIBM5
2023 SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail
abstract
Modern image classifiers perform well on populated classes, while degrading considerably on tail classes with only a few instances. Humans, by contrast, effortlessly handle the long-tailed recognition challenge, since they can learn the tail representation based on different levels of semantic abstraction, making the learned tail features more discriminative. This phenomenon motivated us to propose SuperDisco, an algorithm that discovers super-class representations for long-tailed recognition using a graph model. We learn to construct the super-class graph to guide the representation learning to deal with long-tailed distributions. Through message passing on the super-class graph, image representations are rectified and refined by attending to the most relevant entities based on the semantic similarity among their super-classes. Moreover, we propose to meta-learn the super-class graph under the supervision of a prototype graph constructed from a small amount of imbalanced data. By doing so, we obtain a more robust super-class graph that further improves the long-tailed recognition performance. The consistent state-of-the-art experiments on the long-tailed CIFAR-100, ImageNet, Places and iNaturalist demonstrate the benefit of the discovered super-class graph for dealing with long-tailed distributions.
Yingjun Du, Xiantong Zhen, Cees Snoek
CVPR3
2023 Implicit Diffusion Models for Continuous Super-Resolution
abstract
Image super-resolution (SR) has attracted increasing attention due to its widespread applications. However, current SR methods generally suffer from over-smoothing and artifacts, and most work only with fixed magnifications. This paper introduces an Implicit Diffusion Model (IDM) for high-fidelity continuous image super-resolution. IDM integrates an implicit neural representation and a denoising diffusion model in a unified end-to-end framework, where the implicit neural representation is adopted in the decoding process to learn continuous-resolution representation. Furthermore, we design a scale-adaptive conditioning mechanism that consists of a low-resolution (LR) conditioning network and a scaling factor. The scaling factor regulates the resolution and accordingly modulates the proportion of the LR information and generated features in the final output, which enables the model to accommodate the continuous-resolution requirement. Extensive experiments validate the effectiveness of our IDM and demonstrate its superior performance over prior arts. The source code will be available at https://github.com/Ree1s/IDM.
Sicheng Gao, Xuhui Liu, Bohan Zeng, Sheng Xu 0007, Yanjing Li, Xiaoyan Luo, Jianzhuang Liu, Xiantong Zhen, Baochang Zhang 0001
CVPR8
2023 Order-preserving Consistency Regularization for Domain Adaptation and Generalization
abstract
Deep learning models fail on cross-domain challenges if the model is oversensitive to domain-specific attributes, e.g., lightning, background, camera angle, etc. To alleviate this problem, data augmentation coupled with consistency regularization are commonly adopted to make the model less sensitive to domain-specific attributes. Consistency regularization enforces the model to output the same representation or prediction for two views of one image. These constraints, however, are either too strict or not order-preserving for the classification probabilities. In this work, we propose the Order-preserving Consistency Regularization (OCR) for cross-domain tasks. The order-preserving property for the prediction makes the model robust to task-irrelevant transformations. As a result, the model becomes less sensitive to the domain-specific attributes. The comprehensive experiments show that our method achieves clear advantages on five different cross-domain tasks.
Mengmeng Jing, Xiantong Zhen, Jingjing Li 0001, Cees Snoek
ICCV2
2023 Knowledge-Aware Prompt Tuning for Generalizable Vision-Language Models
abstract
Pre-trained vision-language models, e.g., CLIP, working with manually designed prompts have demonstrated great capacity of transfer learning. Recently, learnable prompts achieve state-of-the-art performance, which however are prone to overfit to seen classes, failing to generalize to unseen classes. In this paper, we propose a Knowledge-Aware Prompt Tuning (KAPT) framework for vision-language models. Our approach takes the inspiration from human intelligence in which external knowledge is usually incorporated into recognizing novel categories of objects. Specifically, we design two complementary types of knowledge-aware prompts for the text encoder to leverage the distinctive characteristics of category-related external knowledge. The discrete prompt extracts the key information from descriptions of an object category, and the learned continuous prompt captures overall contexts. We further design an adaptation head for the visual encoder to aggregate salient attentive visual cues, which establishes discriminative and task-aware visual representations. We conduct extensive experiments on 11 widely-used benchmark datasets and the results verify the effectiveness in few-shot image classification, especially in generalizing to unseen categories. Compared with the state-of-the-art CoCoOp method, KAPT exhibits favorable performance and achieves an absolute gain of 3.22% on new classes and 2.57% in terms of harmonic mean.
Baoshuo Kan, Teng Wang 0007, Wenpeng Lu, Xiantong Zhen, Weili Guan, Feng Zheng 0001
ICCV4
2023 Learning Cross-Modal Affinity for Referring Video Object Segmentation Targeting Limited Samples
abstract
Referring video object segmentation (RVOS), as a supervised learning task, relies on sufficient annotated data for a given scene. However, in more realistic scenarios, only minimal annotations are available for a new scene, which poses significant challenges to existing RVOS methods. With this in mind, we propose a simple yet effective model with a newly designed cross-modal affinity (CMA) module based on a Transformer architecture. The CMA module builds multimodal affinity with a few samples, thus quickly learning new semantic information, and enabling the model to adapt to different scenarios. Since the proposed method targets limited samples for new scenes, we generalize the problem as - few-shot referring video object segmentation (FS-RVOS). To foster research in this direction, we build up a new FS-RVOS benchmark based on currently available datasets. The benchmark covers a wide range and includes multiple situations, which can maximally simulate real-world scenarios. Extensive experiments show that our model adapts well to different scenarios with only a few samples, reaching state-of-the-art performance on the benchmark. On Mini-Ref-YouTube-VOS, our model achieves an average performance of 53.1 ${\mathcal{J}}$ and 54.8 ${\mathcal{F}}$, which are 10% better than the baselines. Furthermore, we show impressive results of 77.7 ${\mathcal{J}}$ and 74.8 ${\mathcal{F}}$ on Mini-Ref-SAIL-VOS, which are significantly better than the baselines. Code is publicly available at https://github.com/hengliusky/Few_shot_RVOS.
Mingqi Gao 0003, Heng Liu 0002, Xiantong Zhen, Feng Zheng 0001
ICCV4
2023 Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot Learning
Ivona Najdenkoska, Xiantong Zhen, Marcel Worring
ICLR2
2023 Energy-Based Test Sample Adaptation for Domain Generalization
Zehao Xiao, Xiantong Zhen, Shengcai Liao, Cees Snoek
ICLR2
2023 MetaModulation: Learning Variational Feature Hierarchies for Few-Shot Learning with Fewer Tasks
abstract
Meta-learning algorithms are able to learn a new task using previously learned knowledge, but they often require a large number of meta-training tasks which may not be readily available. To address this issue, we propose a method for few-shot learning with fewer tasks, which we call MetaModulation. The key idea is to use a neural network to increase the density of the meta-training tasks by modulating batch normalization parameters during meta-training. Additionally, we modify parameters at various neural network levels, rather than just a single layer, to increase task diversity. To account for the uncertainty caused by the reduced number of training tasks, we propose a variational MetaModulation where the modulation parameters are treated as latent variables. We also introduce learning variational feature hierarchies by the variational MetaModulation, which modulates features at all layers and can take into account task uncertainty and generate more diverse tasks. The ablation studies illustrate the advantages of utilizing a learnable task modulation at different levels and demonstrate the benefit of incorporating probabilistic variants in few-task meta-learning. Our MetaModulation and its variational variants consistently outperform state-of-the-art alternatives on four few-task meta-learning benchmarks.
Wenfang Sun, Yingjun Du, Xiantong Zhen, Cees Snoek
ICML3
2023 Episodic Multi-Task Learning with Heterogeneous Neural Processes
abstract
This paper focuses on the data-insufficiency problem in multi-task learning within an episodic training setup. Specifically, we explore the potential of heterogeneous information across tasks and meta-knowledge among episodes to effectively tackle each task with limited data. Existing meta-learning methods often fail to take advantage of crucial heterogeneous information in a single episode, while multi-task learning models neglect reusing experience from earlier episodes. To address the problem of insufficient data, we develop Heterogeneous Neural Processes (HNPs) for the episodic multi-task setup. Within the framework of hierarchical Bayes, HNPs effectively capitalize on prior experiences as meta-knowledge and capture task-relatedness among heterogeneous tasks, mitigating data-insufficiency. Meanwhile, transformer-structured inference modules are designed to enable efficient inferences toward meta-knowledge and task-relatedness. In this way, HNPs can learn more powerful functional priors for adapting to novel heterogeneous tasks in each meta-test episode. Experimental results show the superior performance of the proposed HNPs over typical baselines, and ablation studies verify the effectiveness of the designed inference modules.
Xiantong Zhen, Qi Wang 0009, Marcel Worring
NeurIPS2
2023 Probabilistic Integration of Object Level Annotations in Chest X-ray Classification
abstract
Medical image datasets and their annotations are not growing as fast as their equivalents in the general domain. This makes translation from the newest, more data-intensive methods that have made a large impact on the vision field increasingly more difficult and less efficient. In this paper, we propose a new probabilistic latent variable model for disease classification in chest X-ray images. Specifically we consider chest X-ray datasets that contain global disease labels, and for a smaller subset contain object level expert annotations in the form of eye gaze patterns and disease bounding boxes. We propose a two-stage optimization algorithm which is able to handle these different label granularities through a single training pipeline in a two-stage manner. In our pipeline global dataset features are learned in the lower level layers of the model. The specific details and nuances in the fine-grained expert object-level annotations are learned in the final layers of the model using a knowledge distillation method inspired by conditional variational inference. Subsequently, model weights are frozen to guide this learning process and prevent overfitting on the smaller richly annotated data subsets. The proposed method yields consistent classification improvement across different back-bones on the common benchmark datasets Chest X-ray14 and MIMIC-CXR. This shows how two-stage learning of labels from coarse to fine-grained, in particular with object level annotations, is an effective method for more optimal annotation usage.
Tom van Sonsbeek, Xiantong Zhen, Dwarikanath Mahapatra, Marcel Worring
WACV2
2023 Attentional prototype inference for few-shot segmentation
Haoliang Sun, Xiankai Lu, Yilong Yin, Xiantong Zhen, Cees Snoek, Ling Shao 0001
Pattern Recognit.5
2023 Continuous cross-modal hashing
Hao Zheng 0008, Jinbao Wang 0001, Xiantong Zhen, Jingkuan Song, Feng Zheng 0001, Ke Lu 0002, Guo-Jun Qi
Pattern Recognit.3
2023 Learning to Learn With Variational Inference for Cross-Domain Image Classification
abstract
Learning models that can generalize to previously unseen domains to which we have no access is a fundamental yet challenging problem in machine learning. In this paper, we propose meta variational inference (MetaVI), a variational Bayesian framework of meta-learning for cross domain image classification. Within the meta learning setting, MetaVI is derived to learn a probabilistic latent variable model by maximizing a meta evidence lower bound (Meta ELBO) for knowledge transfer across domains. To enhance the discriminative ability of the model, we further introduce a Wasserstein distance based constraint to the variational objective, leading to the Wasserstein MetaVI, which largely improves classification performance. By casting into a probabilistic inference problem, MetaVI offers the first, principled variational meta-learning framework for cross domain learning. In addition, we collect a new visual recognition dataset to contribute a more challenging benchmark for cross domain learning, which will be released to the public. Extensive experimental evaluation and ablation studies on four benchmarks show that our Wasserstein MetaVI achieves new state-of-the-art performance and surpasses previous methods, demonstrating its great effectiveness.
Lei Zhang 0093, Yingjun Du, Xiantong Zhen
IEEE Trans. Multim.4
2023 Latent Domain Generation for Unsupervised Domain Adaptation Object Counting
abstract
Unsupervised cross-domain crowd counting has recently received great attention in computer vision, which generalizes the model from the source domain to the unlabeled target domain. However, it is an extremely challenging task because only unlabeled data is available from the target domain and the domain gap between two domains is implicit in crowd counting. In this paper, we propose a latent domain generation method to improve the generalization ability of unsupervised domain adaptation crowd counting by generating a latent domain. To this end, we propose a domain generator with random perturbations to learn a new latent distribution derived from the original source distribution. The latent domain generator can extract target information sampled in its stochastic latent representation, which preserves the original target information and enhances the variational ability. Meanwhile, to ensure that the generated latent domain is consistent with the source domain in counting performance, we introduce a consistency loss to encourage similar output from latent and source domains. Moreover, to enhance the adaptation ability of the generated latent domain, we apply the adversarial loss to achieve alignment between the latent and target domains. The domain generator with the adversarial loss and consistency loss ensures that the generated domain is aligned to the target while also improving the robustness of the original source domain model. The experiment indicates that our framework can effortlessly extend to scenarios with different objects (crowd, cars). The experiments also demonstrate the effectiveness of our method on unsupervised realistic-to-realistic crowd counting problems.
Yandan Yang, Jun Xu 0019, Xianbin Cao 0001, Xiantong Zhen, Ling Shao 0001
IEEE Trans. Multim.5
2022 Hierarchical Variational Memory for Few-shot Learning Across Domains
Yingjun Du, Xiantong Zhen, Ling Shao 0001, Cees Snoek
ICLR2
2022 Learning to Generalize across Domains on Single Test Samples
Zehao Xiao, Xiantong Zhen, Ling Shao 0001, Cees Snoek
ICLR2
2022 LifeLonger: A Benchmark for Continual Disease Classification
Mohammad Mahdi Derakhshani, Ivona Najdenkoska, Tom van Sonsbeek, Xiantong Zhen, Dwarikanath Mahapatra, Marcel Worring, Cees Snoek
MICCAI (2)4
2022 Variational Model Perturbation for Source-Free Domain Adaptation
abstract
We aim for source-free domain adaptation, where the task is to deploy a model pre-trained on source domains to target domains. The challenges stem from the distribution shift from the source to the target domain, coupled with the unavailability of any source data and labeled target data for optimization. Rather than fine-tuning the model by updating the parameters, we propose to perturb the source model to achieve adaptation to target domains. We introduce perturbations into the model parameters by variational Bayesian inference in a probabilistic framework. By doing so, we can effectively adapt the model to the target domain while largely preserving the discriminative ability. Importantly, we demonstrate the theoretical connection to learning Bayesian neural networks, which proves the generalizability of the perturbed model to target domains. To enable more efficient optimization, we further employ a parameter sharing strategy, which substantially reduces the learnable parameters compared to a fully Bayesian neural network. Our model perturbation provides a new probabilistic way for domain adaptation which enables efficient adaptation to target domains while maximally preserving knowledge in source models. Experiments on several source-free benchmarks under three different evaluation settings verify the effectiveness of the proposed variational model perturbation for source-free domain adaptation.
Mengmeng Jing, Xiantong Zhen, Jingjing Li 0001, Cees Snoek
NeurIPS2
2022 Association Graph Learning for Multi-Task Classification with Category Shifts
abstract
In this paper, we focus on multi-task classification, where related classification tasks share the same label space and are learned simultaneously. In particular, we tackle a new setting, which is more realistic than currently addressed in the literature, where categories shift from training to test data. Hence, individual tasks do not contain complete training data for the categories in the test set. To generalize to such test data, it is crucial for individual tasks to leverage knowledge from related tasks. To this end, we propose learning an association graph to transfer knowledge among tasks for missing classes. We construct the association graph with nodes representing tasks, classes and instances, and encode the relationships among the nodes in the edges to guide their mutual knowledge transfer. By message passing on the association graph, our model enhances the categorical information of each instance, making it more discriminative. To avoid spurious correlations between task and class nodes in the graph, we introduce an assignment entropy maximization that encourages each class node to balance its edge weights. This enables all tasks to fully utilize the categorical information from related tasks. An extensive evaluation on three general benchmarks and a medical dataset for skin lesion classification reveals that our method consistently performs better than representative baselines.
Zehao Xiao, Xiantong Zhen, Cees Snoek, Marcel Worring
NeurIPS3
2022 Attentive encoder-decoder networks for crowd counting
Xuhui Liu, Yutao Hu 0002, Baochang Zhang 0001, Xiantong Zhen, Xiaoyan Luo, Xianbin Cao 0001
Neurocomputing4
2022 Uncertainty-aware report generation for chest X-rays by variational topic inference
abstract
Automating report generation for medical imaging promises to minimize labor and aid diagnosis in clinical practice. Deep learning algorithms have recently been shown to be capable of captioning natural photos. However, doing a similar thing for medical data, is difficult due to the variety in reports written by different radiologists with fluctuating levels of knowledge and experience. Current methods for automatic report generation tend to merely copy one of the training samples in the created report. To tackle this issue, we propose variational topic inference, a probabilistic approach for automatic chest X-ray report generation. Specifically, we introduce a probabilistic latent variable model where a latent variable defines a single topic. The topics are inferred in a conditional variational inference framework by aligning vision and language modalities in a latent space, with each topic governing the generation of one sentence in the report. We further adopt a visual attention module that enables the model to attend to different locations in the image while generating the descriptions. We conduct extensive experiments on two benchmarks, namely Indiana U. Chest X-rays and MIMIC-CXR. The results demonstrate that our proposed variational topic inference method can generate reports with novel sentence structure, rather than mere copies of reports used in training, while still achieving comparable performance to state-of-the-art methods in terms of standard language generation criteria.
Ivona Najdenkoska, Xiantong Zhen, Marcel Worring, Ling Shao 0001
Medical Image Anal.2
2022 Deep learning DCE-MRI parameter estimation: Application in pancreatic cancer
abstract
Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is an MRI technique for quantifying perfusion that can be used in clinical applications for classification of tumours and other types of diseases. Conventionally, the non-linear least squares (NLLS) methods is used for tracer-kinetic modelling of DCE data. However, despite promising results, NLLS suffers from long processing times (minutes-hours) and noisy parameter maps due to the non-convexity of the cost function. In this work, we investigated physics-informed deep neural networks for estimating physiological parameters from DCE-MRI signal-curves. Three voxel-wise temporal frameworks (FCN, LSTM, GRU) and two spatio-temporal frameworks (CNN, U-Net) were investigated. The accuracy and precision of parameter estimation by the temporal frameworks were evaluated in simulations. All networks showed higher precision than the NLLS. Specifically, the GRU showed to decrease the random error on ve by a factor of 4.8 with respect to the NLLS for noise (SD) of 1/20. The accuracy was better for the prediction of the ve parameter in all networks compared to the NLLS. The GRU and LSTM worked with arbitrary acquisition lengths. The GRU was selected for in vivo evaluation and compared to the spatio-temporal frameworks in 28 patients with pancreatic cancer. All neural network approaches showed less noisy parameter maps than the NLLS. The GRU had better test-retest repeatability than the NLLS for all three parameters and was able to detect one additional patient with significant changes in DCE parameters post chemo-radiotherapy. Although the U-Net and CNN had even better test-retest characteristics than the GRU, and were able to detect even more responders, they also showed potential systematic errors in the parameter maps. Therefore, we advise using our GRU framework for analysing DCE data.
Tim Ottens, Sebastiano Barbieri, Matthew Orton, Remy Klaassen, Hanneke W. M. van Laarhoven, Johannes Crezee, Aart J. Nederveen, Xiantong Zhen, Oliver J. Gurney-Champion
Medical Image Anal.8
2022 Spherical Zero-Shot Learning
abstract
Zero-shot Learning (ZSL) is a highly non-trivial task to generalize from seen to unseen classes. In this paper, we propose spherical zero-shot learning (SZSL) to address the major challenges in ZSL. By decoupling the similarity metric in the spherical embedding space into radius and angle, our SZSL can map classes to hyperspherical surfaces of different radiuses, which greatly increases its flexibility. Specifically, we introduce the spherical alignment on angles to spread classes as uniformly as possible to alleviate the hubness problem and simultaneously preserve the inter-class semantic structure to make the alignment more reasonable. We also introduce the spherical calibration with a minimum entropy based regularizer by adopting a larger radius for unseen classes than seen classes to reduce the prediction bias. Extensive experiments on five middle-scale benchmarks and large-scale ImageNet dataset demonstrate that the proposed approach consistently achieves superior performance for the traditional and generalized settings of ZSL.
Zehao Xiao, Xiantong Zhen, Lei Zhang 0093
IEEE Trans. Circuits Syst. Video Technol.3
2022 Cross-Domain Attention Network for Unsupervised Domain Adaptation Crowd Counting
abstract
Unsupervised domain adaptation crowd counting (UDACC) has been studied with practical research utility by getting rid of the labeling burden on large-scale dense crowds in the target domain. Current methods generalize well within the specific domain gap by directly aligning domain distributions or translating synthetic data to realistic images. However, it is difficult to define domain gaps among complex real-world datasets, in which the images vary greatly in style, density level and/or content. To tackle this problem, in this paper, we propose a Cross-Domain Attention Network (CDANet), which can effectively generalize the model to the unlabeled domain on both unsupervised synthetic-to-realistic and realistic-to-realistic crowd counting. Specifically, we propose a Cross-Domain Attention Module (CDAM) to learn domain-related information between the source and target domain, which extracts relations in cross-domain attentive information, thus enhancing crowd-informative features. Moreover, to make our CDAM invariant to domain shifts, we introduce a consistency penalty to ensure that the attention maps are consistent before and after the domain shifting. Thus our CDANet can pay attention to the shared counting information across domains, while remaining its invariant ability during domain adaptation. Extensive experiments on several common benchmarks for UDACC demonstrate that our CDANet gets competitive results on both unsupervised synthetic-to-realistic and realistic-to-realistic UDACC tasks.
Jun Xu 0019, Xiaoyan Luo, Xianbin Cao 0001, Xiantong Zhen
IEEE Trans. Circuits Syst. Video Technol.5
2022 Variational Hyperparameter Inference for Few-Shot Learning Across Domains
abstract
The focus of few shot learning research has been on the development of meta-learning recently, where a meta-learner is trained on a variety of tasks in hopes of being generalizable to new tasks. Tasks in meta training and meta test are usually assumed to be from the same domain, which would not necessarily hold in real world scenarios. In this paper, we propose variational hyperparameter inference for few-shot learning across domains. Based on an especially successful algorithm named model agnostic meta learning, the proposed variational hyperparameter inference integrates meta learning and variational inference into the optimization of hyperparameters, which enables the meta-learner with adaptivity for generalization across domains. In particular, we choose to learn adaptive hyperparameters including the learning rate and weight decay to avoid the failure in the face of few labeled examples across domain. Moreover, we model hyperparameters as distributions instead of fixed values, which will further enhance the generalization ability by capturing the uncertainty. Extensive experiments are conducted on two benchmark datasets including few shot learning dataset within-domain and across-domain. The results demonstrate that our methods outperforms previous approaches consistently, and comprehensive ablation studies further validate its effectiveness on few shot learning both within domains and across domains.
Lei Zhang 0093, Liyun Zuo, Baoyan Wang, Xin Li 0100, Xiantong Zhen
IEEE Trans. Circuits Syst. Video Technol.5
2022 Variational Abnormal Behavior Detection With Motion Consistency
abstract
Abnormal crowd behavior detection has recently attracted increasing attention due to its wide applications in computer vision research areas. However, it is still an extremely challenging task due to the great variability of abnormal behavior coupled with huge ambiguity and uncertainty of video contents. To tackle these challenges, we propose a new probabilistic framework named variational abnormal behavior detection (VABD), which can detect abnormal crowd behavior in video sequences. We make three major contributions: (1) We develop a new probabilistic latent variable model that combines the strengths of the U-Net and conditional variational auto-encoder, which also are the backbone of our model; (2) We propose a motion loss based on an optical flow network to impose the motion consistency of generated video frames and input video frames; (3) We embed a Wasserstein generative adversarial network at the end of the backbone network to enhance the framework performance. VABD can accurately discriminate abnormal video frames from video sequences. Experimental results on UCSD, CUHK Avenue, IITB-Corridor, and ShanghaiTech datasets show that VABD outperforms the state-of-the-art algorithms on abnormal crowd behavior detection. Without data augmentation, our VABD achieves 72.24% in terms of AUC on IITB-Corridor, which surpasses the state-of-the-art methods by nearly 5%.
Jing Li 0027, Qingwang Huang, Yingjun Du, Xiantong Zhen, Shengyong Chen, Ling Shao 0001
IEEE Trans. Image Process.4
2022 Memory Attention Networks for Skeleton-Based Action Recognition
abstract
Skeleton-based action recognition has been extensively studied, but it remains an unsolved problem because of the complex variations of skeleton joints in 3-D spatiotemporal space. To handle this issue, we propose a newly temporal-then-spatial recalibration method named memory attention networks (MANs) and deploy MANs using the temporal attention recalibration module (TARM) and spatiotemporal convolution module (STCM). In the TARM, a novel temporal attention mechanism is built based on residual learning to recalibrate frames of skeleton data temporally. In the STCM, the recalibrated sequence is transformed or encoded as the input of CNNs to further model the spatiotemporal information of skeleton sequence. Based on MANs, a new collaborative memory fusion module (CMFM) is proposed to further improve the efficiency, leading to the collaborative MANs (C-MANs), trained with two streams of base MANs. TARM, STCM, and CMFM form a single network seamlessly and enable the whole network to be trained in an end-to-end fashion. Comparing with the state-of-the-art methods, MANs and C-MANs improve the performance significantly and achieve the best results on six data sets for action recognition. The source code has been made publicly available at https://github.com/memory-attention-networks.
Ce Li 0002, Chunyu Xie, Baochang Zhang 0001, Jungong Han, Xiantong Zhen, Jie Chen 0001
IEEE Trans. Neural Networks Learn. Syst.5
2021 Meta-Learning with Variational Semantic Memory for Word Sense Disambiguation
abstract
Yingjun Du, Nithin Holla, Xiantong Zhen, Cees Snoek, Ekaterina Shutova. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yingjun Du, Nithin Holla, Xiantong Zhen, Cees Snoek, Ekaterina Shutova
ACL/IJCNLP (1)3
2021 Seminar Learning for Click-Level Weakly Supervised Semantic Segmentation
abstract
Annotation burden has become one of the biggest barriers to semantic segmentation. Approaches based on click-level annotations have therefore attracted increasing attention due to their superior trade-off between supervision and annotation cost. In this paper, we propose seminar learning, a new learning paradigm for semantic segmentation with click-level supervision. The fundamental rationale of seminar learning is to leverage the knowledge from different networks to compensate for insufficient information provided in click-level annotations. Mimicking a seminar, our seminar learning involves a teacher-student and a student-student module, where a student can learn from both skillful teachers and other students. The teacher-student module uses a teacher network based on the exponential moving average to guide the training of the student network. In the student-student module, heterogeneous pseudo-labels are proposed to bridge the transfer of knowledge among students to enhance each other’s performance. Experimental results demonstrate the effectiveness of seminar learning, which achieves the new state-of-the-art performance of 72.51% (mIOU), surpassing previous methods by a large margin of up to 16.88% on the Pascal VOC 2012 dataset.
Jinbao Wang 0001, Hong Cai Chen, Xiantong Zhen, Feng Zheng 0001, Rongrong Ji, Ling Shao 0001
ICCV4
2021 MetaNorm: Learning to Normalize Few-Shot Batches Across Domains
Yingjun Du, Xiantong Zhen, Ling Shao 0001, Cees Snoek
ICLR2
2021 Kernel Continual Learning
abstract
This paper introduces kernel continual learning, a simple but effective variant of continual learning that leverages the non-parametric nature of kernel methods to tackle catastrophic forgetting. We deploy an episodic memory unit that stores a subset of samples for each task to learn task-specific classifiers based on kernel ridge regression. This does not require memory replay and systematically avoids task interference in the classifiers. We further introduce variational random features to learn a data-driven kernel for each task. To do so, we formulate kernel continual learning as a variational inference problem, where a random Fourier basis is incorporated as the latent variable. The variational posterior distribution over the random Fourier basis is inferred from the coreset of each task. In this way, we are able to generate more informative kernels specific to each task, and, more importantly, the coreset size can be reduced to achieve more compact memory, resulting in more efficient continual learning based on episodic memory. Extensive evaluation on four benchmarks demonstrates the effectiveness and promise of kernels for continual learning.
Mohammad Mahdi Derakhshani, Xiantong Zhen, Ling Shao 0001, Cees Snoek
ICML2
2021 A Bit More Bayesian: Domain-Invariant Learning with Uncertainty
abstract
Domain generalization is challenging due to the domain shift and the uncertainty caused by the inaccessibility of target domain data. In this paper, we address both challenges with a probabilistic framework based on variational Bayesian inference, by incorporating uncertainty into neural network weights. We couple domain invariance in a probabilistic formula with the variational Bayesian inference. This enables us to explore domain-invariant learning in a principled way. Specifically, we derive domain-invariant representations and classifiers, which are jointly established in a two-layer Bayesian neural network. We empirically demonstrate the effectiveness of our proposal on four widely used cross-domain visual recognition benchmarks. Ablation studies validate the synergistic benefits of our Bayesian treatment when jointly learning domain-invariant representations and classifiers for domain generalization. Further, our method consistently delivers state-of-the-art mean accuracy on all benchmarks.
Zehao Xiao, Xiantong Zhen, Ling Shao 0001, Cees Snoek
ICML3
2021 Variational Topic Inference for Chest X-Ray Report Generation
Ivona Najdenkoska, Xiantong Zhen, Marcel Worring, Ling Shao 0001
MICCAI (3)2
2021 Learning Hierarchical Embedding for Video Instance Segmentation
abstract
In this paper, we address video instance segmentation using a new generative model that learns effective representations of the target and background appearance. We propose to exploit hierarchical structural embedding over spatio-temporal space, which is compact, powerful, and flexible in contrast to current tracking-by-detection methods. Specifically, our model segments and tracks instances across space and time in a single forward pass, which is formulated as hierarchical embedding learning. The model is trained to locate the pixels belonging to specific instances over a video clip. We firstly take advantage of a novel mixing function to better fuse spatio-temporal embeddings. Moreover, we introduce normalizing flows to further improve the robustness of the learned appearance embedding, which theoretically extends conventional generative flows to a factorized conditional scheme. Comprehensive experiments on the video instance segmentation benchmark, i.e., YouTube-VIS, demonstrate the effectiveness of the proposed approach. Furthermore, we evaluate our method on an unsupervised video object segmentation dataset to demonstrate its generalizability.
Zheyun Qin, Xiankai Lu, Xiushan Nie, Xiantong Zhen, Yilong Yin
ACM Multimedia4
2021 Variational Multi-Task Learning with Gumbel-Softmax Priors
abstract
Multi-task learning aims to explore task relatedness to improve individual tasks, which is of particular significance in the challenging scenario that only limited data is available for each task. To tackle this challenge, we propose variational multi-task learning (VMTL), a general probabilistic inference framework for learning multiple related tasks. We cast multi-task learning as a variational Bayesian inference problem, in which task relatedness is explored in a unified manner by specifying priors. To incorporate shared knowledge into each task, we design the prior of a task to be a learnable mixture of the variational posteriors of other related tasks, which is learned by the Gumbel-Softmax technique. In contrast to previous methods, our VMTL can exploit task relatedness for both representations and classifiers in a principled way by jointly inferring their posteriors. This enables individual tasks to fully leverage inductive biases provided by related tasks, therefore improving the overall performance of all tasks. Experimental results demonstrate that the proposed VMTL is able to effectively tackle a variety of challenging multi-task learning settings with limited training data for both classification and regression. Our method consistently surpasses previous methods, including strong Bayesian approaches, and achieves state-of-the-art performance on five benchmark datasets.
Xiantong Zhen, Marcel Worring, Ling Shao 0001
NeurIPS2
2021 Learning to Learn Dense Gaussian Processes for Few-Shot Learning
abstract
Gaussian processes with deep neural networks demonstrate to be a strong learner for few-shot learning since they combine the strength of deep learning and kernels while being able to well capture uncertainty. However, it remains an open problem to leverage the shared knowledge provided by related tasks. In this paper, we propose to learn Gaussian processes with dense inducing variables by meta-learning for few-shot learning. In contrast to sparse Gaussian processes, we define a set of dense inducing variables to be of a much larger size than the support set in each task, which collects prior knowledge from experienced tasks. The dense inducing variables specify a shared Gaussian process prior over prediction functions of all tasks, which are learned in a variational inference framework and offer a strong inductive bias for learning new tasks. To achieve task-specific prediction functions, we propose to adapt the inducing variables to each task by efficient gradient descent. We conduct extensive experiments on common benchmark datasets for a variety of few-shot learning tasks. Our dense Gaussian processes present significant improvements over vanilla Gaussian processes and comparable or even better performance with state-of-the-art methods.
Ze Wang 0008, Zichen Miao, Xiantong Zhen, Qiang Qiu 0001
NeurIPS3
2021 Variational Prototype Inference for Few-Shot Semantic Segmentation
abstract
In this paper, we propose variational prototype inference to address few-shot semantic segmentation in a probabilistic framework. A probabilistic latent variable model infers the distribution of the prototype that is treated as the latent variable. We formulate the optimization as a variational inference problem, which is established with an amortized inference network based on an auto-encoder architecture. The probabilistic modeling of the prototype enhances its generalization ability to handle the inherent uncertainty caused by limited data and the huge intra-class variations of objects. Moreover, it offers a principled way to incorporate the prototype extracted from support images into the prediction of the segmentation maps for query images. We conduct extensive experimental evaluations on three benchmark datasets. Ablation studies show the effectiveness of variational prototype inference for few-shot semantic segmentation by probabilistic modeling. On all three benchmarks, our proposal achieves high segmentation accuracy and surpasses previous methods by considerable margins.
Yandan Yang, Xianbin Cao 0001, Xiantong Zhen, Cees Snoek, Ling Shao 0001
WACV4
2021 Deep 3D human pose estimation: A review
abstract
Three-dimensional (3D) human pose estimation involves estimating the articulated 3D joint locations of a human body from an image or video. Due to its widespread applications in a great variety of areas, such as human motion analysis, human–computer interaction, robots, 3D human pose estimation has recently attracted increasing attention in the computer vision community, however, it is a challenging task due to depth ambiguities and the lack of in-the-wild datasets. A large number of approaches, with many based on deep learning, have been developed over the past decade, largely advancing the performance on existing benchmarks. To guide future development, a comprehensive literature review is highly desired in this area. However, existing surveys on 3D human pose estimation mainly focus on traditional methods and a comprehensive review on deep learning based methods remains lacking in the literature. In this paper, we provide a thorough review of existing deep learning based works for 3D pose estimation, summarize the advantages and disadvantages of these methods and provide an in-depth understanding of this area. Furthermore, we also explore the commonly-used benchmark datasets on which we conduct a comprehensive study for comparison and analysis. Our study sheds light on the state of research development in 3D human pose estimation and provides insights that can facilitate the future design of models and algorithms.
Jinbao Wang 0001, Shujie Tan, Xiantong Zhen, Feng Zheng 0001, Zhenyu He 0001, Ling Shao 0001
Comput. Vis. Image Underst.3
2021 Attentional Kernel Encoding Networks for Fine-Grained Visual Categorization
abstract
Fine-grained visual categorization aims to recognize objects from different sub-ordinate categories, which is a challenging task due to subtle visual differences between images. It is highly desired to identify discriminative regions while achieving highly non-linear compact representation for fine-grained visual categorization. However, existing methods either rely on manually defined part-based annotations to indicate the distinctive regions or operate on longitudinal vectors to capture the non-linear information, which may lose important spatial layout information. In this paper, we propose the Attentional Kernel Encoding Networks (AKEN) for fine-grained visual categorization. Specifically, the AKEN aggregates feature maps from the last convolutional layer of ConvNets to obtain a holistic feature representation. By Fourier embedding, it encodes features from both the longitudinal and transverse directions, which largely retains the spatial layout information. Moreover, we incorporate a Cascaded Attention (Cas-Attention) module to highlight local regions that distinguish among subordinate categories, enabling the AKEN to extract the most discriminative features. Working in conjunction with the attention mechanism, the proposed AKEN combines the strengths of ConvNets and kernels for non-linear feature learning, which can establish discriminative and descriptive feature representations for fine-grained image categorization. Experiments on three benchmark datasets show that the proposed AKEN delivers highly competitive performance, surpassing most existed methods and achieving state-of-the-art results.
Yutao Hu 0002, Yandan Yang, Jun Zhang 0007, Xianbin Cao 0001, Xiantong Zhen
IEEE Trans. Circuits Syst. Video Technol.5
2021 Learning to Adapt With Memory for Probabilistic Few-Shot Learning
abstract
Few-shot learning has recently generated increasing popularity in machine learning, which addresses the fundamental yet challenging problem of learning to adapt to new tasks with the limited data. In this paper, we propose a new probabilistic framework that learns to fast adapt with external memory. We model the classifier parameters as distributions that are inferred from the support set and directly applied to the query set for prediction. The model is optimized by formulating as a variational inference problem. The probabilistic modeling enables better handling prediction uncertainty due to the limited data. We impose a discriminative constraint on the feature representations by exploring the class structure, which can improve the classification performance. We further introduce a memory unit to store task-specific information extracted from the support set and used for the query set to achieve explicit adaption to individual tasks. By episodic training, the model learns to acquire the capability of adapting to specific tasks, which guarantees its performance on new related tasks. We conduct extensive experiments on widely-used benchmarks for few-shot recognition. Our method achieves new state-of-the-art performance and largely surpassing previous methods by large margins. The ablation study further demonstrates the effectiveness of the proposed discriminative learning and memory unit.
Lei Zhang 0093, Liyun Zuo, Yingjun Du, Xiantong Zhen
IEEE Trans. Circuits Syst. Video Technol.4
2021 Pixel-Level Non-local Image Smoothing With Objective Evaluation
abstract
Recently, imagesmoothing has gained increasing attention due to its prerequisite role in other image processing tasks, e.g., image enhancement and editing. However, the evaluation of image smoothing algorithms is usually performed by subjective observation on images without corresponding ground truths. To promote the development of image smoothing algorithms, in this paper, we construct a novel Nankai Smoothing (NKS) dataset containing 200 images blended by versatile structure images and natural textures. The structure images are inherently smooth and naturally taken as ground truths. On our NKS dataset, we comprehensively evaluate 14 popular image smoothing algorithms. Moreover, we propose a Pixel-level Non-Local Smoothing (PNLS) method to well preserve the structure of the smoothed images, by exploiting the pixel-level non-local self-similarity prior of natural images. Extensive experiments on several benchmark datasets demonstrate that our PNLS outperforms previous algorithms on the image smoothing task. Ablation studies also reveal the work mechanism of our PNLS on image smoothing. To further show its effectiveness, we apply our PNLS on several applications such as semantic region smoothing, detail/edge enhancement, and image abstraction. The dataset and code are available athttps://github.com/zal0302/PNLS.
Jun Xu 0019, Yingkun Hou, Xiantong Zhen, Ling Shao 0001, Ming-Ming Cheng
IEEE Trans. Multim.4
2020 Learning to Learn with Variational Information Bottleneck for Domain Generalization
Yingjun Du, Jun Xu 0019, Huan Xiong, Qiang Qiu 0001, Xiantong Zhen, Cees Snoek, Ling Shao 0001
ECCV (10)5
2020 Few-Shot Semantic Segmentation with Democratic Attention Networks
Yutao Hu 0002, Yandan Yang, Xianbin Cao 0001, Xiantong Zhen
ECCV (13)6
2020 You Only Need The Image: Unsupervised Few-Shot Semantic Segmentation With Co-Guidance Network
abstract
Few-shot semantic segmentation has recently attracted attention for its ability to segment unseen-class images with only a few annotated support samples. Yet existing methods not only need to be trained with a large scale of pixel-level annotations on certain seen classes, but also require a few annotated support image-mask pairs for the guidance of segmentation on each unseen class. In this paper, we propose the Co-guidance Network (CGNet) for unsupervised few-shot segmentation, which eliminates requirements of annotation on both seen and unseen classes. Specifically, CGNet segments unseen-class images with only unlabeled support images by the newly designed co-guidance mechanism. Moreover, CGNet is trained on seen classes by a novel co-existence recognition loss, which further removes the need of pixel-level annotations. Extensive experiments on the PASCAL -5idataset show that the unsupervised CGNet performs comparably with the state-of-the-art fully-supervised few-shot methods, while largely alleviating annotation requirement.
Yandan Yang, Xianbin Cao 0001, Xiantong Zhen
ICIP5
2020 Learning to Learn Kernels with Variational Random Features
abstract
We introduce kernels with random Fourier features in the meta-learning framework for few-shot learning. We propose meta variational random features (MetaVRF) to learn adaptive kernels for the base-learner, which is developed in a latent variable model by treating the random feature basis as the latent variable. We formulate the optimization of MetaVRF as a variational inference problem by deriving an evidence lower bound under the meta-learning framework. To incorporate shared knowledge from related tasks, we propose a context inference of the posterior, which is established by an LSTM architecture. The LSTM-based inference network can effectively integrate the context information of previous tasks with task-specific information, generating informative and adaptive features. The learned MetaVRF can produce kernels of high representational power with a relatively low spectral sampling rate and also enables fast adaptation to new tasks. Experimental results on a variety of few-shot regression and classification tasks demonstrate that MetaVRF delivers much better, or at least competitive, performance compared to existing meta-learning alternatives.
Xiantong Zhen, Haoliang Sun, Yingjun Du, Jun Xu 0019, Yilong Yin, Ling Shao 0001, Cees Snoek
ICML1
2020 Transductive Relation-Propagation Network for Few-shot Learning
abstract
Few-shot learning, aiming to learn novel concepts from few labeled examples, is an interesting and very challenging problem with many practical advantages. To accomplish this task, one should concentrate on revealing the accurate relations of the support-query pairs. We propose a transductive relation-propagation graph neural network (TRPN) to explicitly model and propagate such relations across support-query pairs. Our TRPN treats the relation of each support-query pair as a graph node, named relational node, and resorts to the known relations between support samples, including both intra-class commonality and inter-class uniqueness, to guide the relation propagation in the graph, generating the discriminative relation embeddings for support-query pairs. A pseudo relational node is further introduced to propagate the query characteristics, and a fast, yet effective transductive learning strategy is devised to fully exploit the relation information among different queries. To the best of our knowledge, this is the first work that explicitly takes the relations of support-query pairs into consideration in few-shot learning, which might offer a new way to solve the few-shot learning problem. Extensive experiments conducted on several benchmark datasets demonstrate that our method can significantly outperform a variety of state-of-the-art few-shot learning methods.
Yuqing Ma, Shihao Bai, Shan An, Wei Liu 0005, Aishan Liu, Xiantong Zhen, Xianglong Liu 0001
IJCAI6
2020 Few-Shot Ensemble Learning for Video Classification with SlowFast Memory Networks
abstract
In the era of big data, few-shot learning has recently received much attention in multimedia analysis and computer vision due to its appealing ability of learning from scarce labeled data. However, it has been largely underdeveloped in the video domain, which is even more challenging due to the huge spatial-temporal variability of video data. In this paper, we address few-shot video classification by learning an ensemble of SlowFast networks augmented with memory units. Specifically, we introduce a family of few-shot learners based on SlowFast networks which are used to extract informative features at multiple rates, and we incorporate a memory unit into each network to enable encoding and retrieving crucial information instantly. Furthermore, we propose a choice controller network to leverage the diversity of few-shot learners by learning to adaptively assign a confidence score to each SlowFast memory network, leading to a strong classifier for enhanced prediction. Experimental results on two widely-adopted video datasets demonstrate the effectiveness of the proposed method, as well as its superior performance over the state-of-the-art approaches.
Mengshi Qi, Jie Qin 0004, Xiantong Zhen, Di Huang 0001, Yi Yang 0001, Jiebo Luo 0001
ACM Multimedia3
2020 Learning to Learn Variational Semantic Memory
abstract
In this paper, we introduce variational semantic memory into meta-learning to acquire long-term knowledge for few-shot learning. The variational semantic memory accrues and stores semantic information for the probabilistic inference of class prototypes in a hierarchical Bayesian framework. The semantic memory is grown from scratch and gradually consolidated by absorbing information from tasks it experiences. By doing so, it is able to accumulate long-term, general knowledge that enables it to learn new concepts of objects. We formulate memory recall as the variational inference of a latent memory variable from addressed contents, which offers a principled way to adapt the knowledge to individual tasks. Our variational semantic memory, as a new long-term memory module, confers principled recall and update mechanisms that enable semantic information to be efficiently accrued and adapted for few-shot learning. Experiments demonstrate that the probabilistic modelling of prototypes achieves a more informative representation of object classes compared to deterministic vectors. The consistent new state-of-the-art performance on four benchmarks shows the benefit of variational semantic memory in boosting few-shot recognition.
Xiantong Zhen, Yingjun Du, Huan Xiong, Qiang Qiu 0001, Cees Snoek, Ling Shao 0001
NeurIPS1
2020 Variational Image Deraining
abstract
Images captured in severe weather such as rain and snow significantly degrade the accuracy of vision systems, e.g., for outdoor video surveillance or autonomous driving. Image deraining is a critical yet highly challenging task, due to the fact that rain density varies across spatial locations, while the distribution patterns simultaneously vary across color channels. In this paper, we propose a variational image deraining (VID) method by formulating image deraining in a conditional variational auto-encoder framework. To achieve adaptive deraining to spatial rain density, we generate a density estimation map for each color channel, which can largely avoid over and under deraining. In addition, to address cross-channel variations, we conduct channel-wise deraining, motivated by our observation that bright pixels do not tend to remain bright after deraining unless their color channels are handled separately. Experimental results show that the proposed deraining method achieves superior performance on both synthesized and real rainy images, surpassing previous state-of-the-art methods by large margins.
Yingjun Du, Jun Xu 0019, Qiang Qiu 0001, Xiantong Zhen, Lei Zhang 0093
WACV4
2020 Model-Agnostic Metric for Zero-Shot Learning
abstract
Zero-shot Learning (ZSL) aims to learn a classifier to recognize unseen categories without training samples. Most ZSL works based on embedding models handle the visual space and the semantic space through a common metric space and then apply a simple nearest neighbor search which directly leads to the hubness problem, one of the main challenges of ZSL. Contrary to recent works, whose conclusions about hubs are drawn based on Euclidean and specific models like ridge regression, we adopt cosine metric and for the first time prove cosine is model-agnostic to alleviate the hubness problem in ZSL. Assuming that the normalized mapped semantic vectors follow a uniform distribution, we provide theoretical analysis which demonstrates that hubs can be better reduced with a higher-dimensional cosine metric space. Moreover, we introduce a diversity-based regularizer with the cosine metric which underpins the assumption about the uniform distribution and further improves the model's discriminative ability. Extensive experiments on five benchmarks and large-scale Imagenet dataset show that our method can improve the performance, surpassing previous embedding methods by large margins.
Qiang Qiu 0001, Xiantong Zhen, Xianbin Cao 0001
WACV5
2020 Heterogenous output regression network for direct face alignment
Xiantong Zhen, Mengyang Yu, Zehao Xiao, Lei Zhang 0093, Ling Shao 0001
Pattern Recognit.1
2020 Calibrated Multivariate Regression Networks
abstract
In this paper, we propose a new multi-layer learning architecture, the calibrated multivariate regression network (CMRN). Compared to previous multivariate models, the CMRN is able to simultaneously handle major challenges in multivariate regression including highly nonlinear input-output relationships, underlying inter-output correlations and calibration of multiple outputs within one single framework. The CMRN is comprised of a nonlinear module with cosine activations and a linear module with the low-rank expansion, which establishes a compact multivariate regression network. By seamlessly working with the ℓ2,1loss, the CMRN automatically calibrates multiple outputs with distinct noise levels to achieve improved performance. Being succinctly formulated but theoretically well-founded, the CMRN offers a compact multi-layer learning architecture that can be efficiently trained to scale up with massive datasets. We conduct extensive experimental evaluation on two representative large multivariate regression tasks for both machine learning and computer vision. The proposed CMRN can produce high performance on all tasks, which is better or competitive to state-of-the-art models. Extensive ablation studies offer deep insights into the effectiveness of the proposed CMRN.
Lei Zhang 0093, Yingjun Du, Xin Li 0100, Xiantong Zhen
IEEE Trans. Circuits Syst. Video Technol.4
2020 Conditional Variational Image Deraining
abstract
Image deraining is an important yet challenging image processing task. Though deterministic image deraining methods are developed with encouraging performance, they are infeasible to learn flexible representations for probabilistic inference and diverse predictions. Besides, rain intensity varies both in spatial locations and across color channels, making this task more difficult. In this paper, we propose a Conditional Variational Image Deraining (CVID) network for better deraining performance, leveraging the exclusive generative ability of Conditional Variational Auto-Encoder (CVAE) on providing diverse predictions for the rainy image. To perform spatially adaptive deraining, we propose a spatial density estimation (SDE) module to estimate a rain density map for each image. Since rain density varies across different color channels, we also propose a channel-wise (CW) deraining scheme. Experiments on synthesized and real-world datasets show that the proposed CVID network achieves much better performance than previous deterministic methods on image deraining. Extensive ablation studies validate the effectiveness of the proposed SDE module and CW scheme in our CVID network. The code is available at https://github.com/Yingjun-Du/VID.
Yingjun Du, Jun Xu 0019, Xiantong Zhen, Ming-Ming Cheng, Ling Shao 0001
IEEE Trans. Image Process.3
2020 The Structure Transfer Machine Theory and Applications
abstract
Representation learning is a fundamental but challenging problem, especially when the distribution of data is unknown. In this paper, we propose a new representation learning method, named Structure Transfer Machine (STM), which enables feature learning process to converge at the representation expectation in a probabilistic way. We theoretically show that such an expected value of the representation (mean) is achievable if the manifold structure can be transferred from the data space to the feature space. The resulting structure regularization term, named manifold loss, is incorporated into the loss function of the typical deep learning pipeline. The STM architecture is constructed to enforce the learned deep representation to satisfy the intrinsic manifold structure from the data, which results in robust features that suit various application scenarios, such as digit recognition, image classification and object tracking. Compared with state-of-the-art CNN architectures, we achieve better results on several commonly used public benchmarks.
Baochang Zhang 0001, Wankou Yang, Ze Wang 0008, Lian Zhuo, Jungong Han, Xiantong Zhen
IEEE Trans. Image Process.6
2019 Attentive Temporal Pyramid Network for Dynamic Scene Classification
abstract
Dynamic scene classification is an important yet challenging problem especially with the presence of defected or irrelevant frames due to unconstrained imaging conditions such as illumination, camera motion and irrelevant background. In this paper, we propose the attentive temporal pyramid network (ATP-Net) to establish effective representations of dynamic scenes by extracting and aggregating the most informative and discriminative features. The proposed ATP-Net detects informative features of frames that contain the most relevant information to scenes by a temporal pyramid structure with the incorporated attention mechanism. These frame features are effectively fused by a newly designed kernel aggregation layer based on kernel approximation into a discriminative holistic representations of dynamic scenes. The proposed ATP-Net leverages the strength of attention mechanism to select the most relevant frame features and the ability of kernels to achieve optimal feature fusion for discriminative representations of dynamic scenes. Extensive experiments and comparisons are conducted on three benchmark datasets and the results show our superiority over the state-of-the-art methods on all these three benchmark datasets.
Yuanjun Huang, Xianbin Cao 0001, Xiantong Zhen, Jungong Han
AAAI3
2019 Crowd Counting and Density Estimation by Trellis Encoder-Decoder Networks
abstract
Crowd counting has recently attracted increasing interest in computer vision but remains a challenging problem. In this paper, we propose a trellis encoder-decoder network (TEDnet) for crowd counting, which focuses on generating high-quality density estimation maps. The major contributions are four-fold. First, we develop a new trellis architecture that incorporates multiple decoding paths to hierarchically aggregate features at different encoding stages, which improves the representative capability of convolutional features for large variations in objects. Second, we employ dense skip connections interleaved across paths to facilitate sufficient multi-scale feature fusions, which also helps TEDnet to absorb the supervision information. Third, we propose a new combinatorial loss to enforce similarities in local coherence and spatial correlation between maps. By distributedly imposing this combinatorial loss on intermediate outputs, TEDnet can improve the back-propagation process and alleviate the gradient vanishing problem. Finally, on four widely-used benchmarks, our TEDnet achieves the best overall performance in terms of both density map quality and counting accuracy, with an improvement up to 14% in MAE metric. These results validate the effectiveness of TEDnet for crowd counting.
Zehao Xiao, Baochang Zhang 0001, Xiantong Zhen, Xianbin Cao 0001, David S. Doermann, Ling Shao 0001
CVPR4
2019 Relational Attention Network for Crowd Counting
abstract
Crowd counting is receiving rapidly growing research interests due to its potential application value in numerous real-world scenarios. However, due to various challenges such as occlusion, insufficient resolution and dynamic backgrounds, crowd counting remains an unsolved problem in computer vision. Density estimation is a popular strategy for crowd counting, where conventional density estimation methods perform pixel-wise regression without explicitly accounting the interdependence of pixels. As a result, independent pixel-wise predictions can be noisy and inconsistent. In order to address such an issue, we propose a Relational Attention Network (RANet) with a self-attention mechanism for capturing interdependence of pixels. The RANet enhances the self-attention mechanism by accounting both short-range and long-range interdependence of pixels, where we respectively denote these implementations as local self-attention (LSA) and global self-attention (GSA). We further introduce a relation module to fuse LSA and GSA to achieve more informative aggregated feature representations. We conduct extensive experiments on four public datasets, including ShanghaiTech A, ShanghaiTech B, UCF-CC-50 and UCF-QNRF. Experimental results on all datasets suggest RANet consistently reduces estimation errors and surpasses the state-of-the-art approaches by large margins.
Zehao Xiao, Fan Zhu 0001, Xiantong Zhen, Xianbin Cao 0001, Ling Shao 0001
ICCV5
2019 Attentional Neural Fields for Crowd Counting
abstract
Crowd counting has recently generated huge popularity in computer vision, and is extremely challenging due to the huge scale variations of objects. In this paper, we propose the Attentional Neural Field (ANF) for crowd counting via density estimation. Within the encoder-decoder network, we introduce conditional random fields (CRFs) to aggregate multi-scale features, which can build more informative representations. To better model pair-wise potentials in CRFs, we incorperate non-local attention mechanism implemented as inter- and intra-layer attentions to expand the receptive field to the entire image respectively within the same layer and across different layers, which captures long-range dependencies to conquer huge scale variations. The CRFs coupled with the attention mechanism are seamlessly integrated into the encoder-decoder network, establishing an ANF that can be optimized end-to-end by back propagation. We conduct extensive experiments on four public datasets, including ShanghaiTech, WorldEXPO 10, UCF-CC-50 and UCF-QNRF. The results show that our ANF achieves high counting performance, surpassing most previous methods.
Fan Zhu 0001, Xiantong Zhen, Xianbin Cao 0001, Ling Shao 0001
ICCV5
2019 Two-Stream Multi-Task Network for Fashion Recognition
abstract
In this paper, we present a two-stream multi-task network for fashion recognition. This task is challenging as fashion clothing always contain multiple attributes, which need to be predicted simultaneously for real-time industrial systems. To handle these challenges, we formulate fashion recognition into a multi-task learning problem, including landmark detection, category and attribute classifications, and solve it with the proposed deep convolutional neural network. We design two knowledge sharing strategies which enable information transfer between tasks and improve the overall performance. The proposed model achieves state-of-the-art results on large-scale fashion dataset comparing to the existing methods, which demonstrates its great effectiveness and superiority for fashion recognition.
Peizhao Li, Yanjing Li, Xiantong Zhen
ICIP4
2019 Learning the Set Graphs: Image-Set Classification Using Sparse Graph Convolutional Networks
abstract
Image-set classification has recently made great progress in computer vision. Compared with traditional image classification tasks, set-based classification exhibits great challenges due to huge intra-class variability and high inter-class ambiguity. In this paper, we propose to model the image set as a graph and formulate image set classification as the graph matching task. Without relying on the strong structure assumption, we build the first end-to-end graph convolutional network, the Deep SetNet, to learn the graph structure of an image set. Specifically, the SetNet consists of one convolutional network (CNN) to sufficiently extract the discriminative vertex, one graph convolutional Network (GCN) to faithfully learn the substructure in set graphs and graph pooling layers to aggregate the vertex features from the GCN. Moreover, we propose imposing the ℓ1,2-norm based sparsity constraint to select vertex features, which largely improves the model generalization capability. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art methods, showing its great effectiveness in set-based image classification.
Haoliang Sun, Xiantong Zhen, Yilong Yin
ICIP2
2019 Model-Free Tracking With Deep Appearance and Motion Features Integration
abstract
Being able to track an anonymous object, a model-free tracker is comprehensively applicable regardless of the target type. However, designing such a generalized framework is challenged by the lack of object-oriented prior information. As one solution, a real-time model-free object tracking approach is designed in this work relying on Convolutional Neural Networks (CNNs). To overcome the object-centric information scarcity, both appearance and motion features are deeply integrated by the proposed AMNet, which is an end-to-end offline trained two-stream network. Between the two parallel streams, the ANet investigates appearance features with a multi-scale Siamese atrous CNN, enabling the tracking-by-matching strategy. The MNet achieves deep motion detection to localize anonymous moving objects by processing generic motion features. The final tracking result at each frame is generated by fusing the output response maps from both sub-networks. The proposed AMNet reports leading performance on both OTB and VOT benchmark datasets with favorable real-time processing speed.
Peizhao Li, Xiantong Zhen, Xianbin Cao 0001
WACV3
2019 Multi-Scale Aggregation Network for Direct Face Alignment
abstract
Face alignment has been extensively researched in computer vision while remaining a challenging task. Direct face alignment based on convolutional neural networks (CNN) without relying on cascaded regression has recently emerged and achieved promising performance. In this paper, we propose a multi-scale aggregation network (MAN) for direct face alignment by aggregating features from intermediate layers of a CNN. Specifically, MAN adopts a new convolutional architecture to aggregate features at all scales in different semantic levels, which establishes highly informative facial representations for accurate alignment. Moreover, we introduce the attention mechanism into the network, which drives it to focus on the spatial regions closely related to facial landmarks for further improved performance. Our MAN achieves a general end-to-end learning architecture for multi-scale feature aggregation, which, coupled with spatial attention mechanism, is well-suited for direct face alignment. Extensive experiments conducted on four benchmark datasets, including AFLW, 300W, CelebA and 300VW, show that MAN consistently produces high performance and surpasses several state-of-the-art methods in most cases.
Peizhao Li, Xiantong Zhen, Xianbin Cao 0001
WACV4
2019 Starts Better and Ends Better: A Target Adaptive Image Signature Tracker
abstract
Correlation filter (CF) trackers have achieved outstanding performance in visual object tracking tasks, in which the cosine mask plays an essential role in alleviating boundary effects caused by the circular assumption. However, the cosine mask imposes a larger weight on its center position, which greatly affects CF trackers, that is, their performance will drop significantly if a bad starting point happens to occur. To address the above issue, we propose a target adaptive image signature (TaiS) model to refine the starting point in each frame for CF trackers. Specifically, we incorporate the target priori into the image signature to build a target-specific saliency map, and iteratively refine the starting point with a closed-form solution during the tracking process. As a result, our TaiS is able to find a better starting point close to the center of targets; more importantly, it is independent of specific CF trackers and can efficiently improve their performance. Experiments on two benchmark datasets, i.e., OTB100 and UAV123, demonstrate that our TaiS consistently achieves high performance and updates the state of the arts in visual tracking. The source code of our approach will be made publicly available.
Xingchao Liu, Ce Li 0002, Hongren Wang 0001, Xiantong Zhen, Baochang Zhang 0001, Qixiang Ye
WACV4
2019 Attentional Information Fusion Networks for Cross-Scene Power Line Detection
abstract
The power line is one of the most hazardous obstacles for low-altitude aircrafts. As aircrafts usually encounter scenes like never before during the flight, cross-scene power line detection is the key for their flight safety. However, compared to regular object detection tasks, cross-scene power line detection is extremely challenging due to its weak visual appearance and widespread existence. In this letter, we propose a cross-scene power line detection method based on attentional information fusion networks. Specifically, we construct a fully convolutional network with attention and information fusion mechanism for cross-scene detection. The two main modules make full use of the semantic and location information, which enables the model to focus more on power lines rather than the unexpected scenes. To the best of author knowledge, our method establishes the first end-to-end convolutional architecture for pixelwise power line detection. Experimental results have shown that our method outperforms previous methods by large margins for cross-scene power line detection.
Yan Li 0054, Zehao Xiao, Xiantong Zhen, Xianbin Cao 0001
IEEE Geosci. Remote. Sens. Lett.3
2019 Long-Short-Term Features for Dynamic Scene Classification
abstract
Dynamic scene classification has been extensively studied in computer vision due to its widespread applications. The key to dynamic scene classification lies in jointly characterizing spatial appearance and temporal dynamics to achieve informative representation, which remains an outstanding task in the literature. In this paper, we propose a unified framework to extract spatial and temporal features for dynamic scene representation. More specifically, we deploy two variants of deep convolutional neural networks to encode spatial appearance and short-term dynamics into short-term deep features (STDF). Based on STDF, we propose using the autoregressive moving average model to extract long-term frequency features (LTFF). By combining STDF and LTFF, we establish the long-short-term feature (LSTF) representations of dynamic scenes. The LSTF characterizes both spatial and temporal patterns of dynamic scenes for comprehensive and information representation that enables more accurate classification. Extensive experiments on three-dynamic scene classification benchmarks have shown that the proposed LSTF achieves high performance and substantially surpasses the state-of-the-art methods.
Yuanjun Huang, Xianbin Cao 0001, Qi Wang 0009, Baochang Zhang 0001, Xiantong Zhen, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2019 Glance and Stare: Trapping Flying Birds in Aerial Videos by Adaptive Deep Spatio-Temporal Features
abstract
Flying bird detection has recently attracted increasing attention in computer vision, which becomes an urgent task with the opening up of the low-altitude airspace. However, compared to conventional object detection tasks, it is much more challenging to trap flying birds in aerial videos due to small target sizes, complex backgrounds of great variations and disturbances of bird-like objects. In this paper, we propose a unified framework termed glance-and-stare detection (GSD) to trap flying birds in aerial videos. The GSD is inspired by the fact that human beings first glance at the whole image and then stare at the areas where the suspected object is most likely to appear until the confirmation is obtained. Specifically, we propose the zooming-in algorithm to generate region proposals for accurate localization of flying birds; to represent region proposal sequences of different lengths, we propose adaptive deep spatio-temporal features by leveraging the strength of 3D convolutional neural networks, based on which classification is conducted to achieve final detection. In contrast to conventional methods, the GSD enables localization and classification to be conducted jointly in an alternating iterative way, which mutually enhances each other to improve their performance. In order to validate the proposed GSD algorithm, we build flying bird data sets including images and videos, which provide new benchmarks for evaluation of flying bird detection systems. Experiments on the data sets demonstrate that the GSD can achieve high detection accuracy and largely outperform the state-of-the-art detection methods.
Shuman Tian, Xianbin Cao 0001, Yan Li 0054, Xiantong Zhen, Baochang Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2019 Learning Match Kernels on Grassmann Manifolds for Action Recognition
abstract
Action recognition has been extensively researched in computer vision due to its potential applications in a broad range of areas. The key to action recognition lies in modeling actions and measuring their similarity, which however poses great challenges. In this paper, we propose learning match kernels between actions on Grassmann manifold for action recognition. Specifically, we propose modeling actions as a linear subspace on the Grassmann manifold; the subspace is a set of convolutional neural network (CNN) feature vectors pooled temporally over frames in semantic video clips, which simultaneously captures local discriminant patterns and temporal dynamics of motion. To measure the similarity between actions, we propose Grassmann match kernels (GMK) based on canonical correlations of linear subspaces to directly match videos for action recognition; GMK is learned in a supervised way via kernel target alignment, which is endowed with a great discriminative ability to distinguish actions from different classes. The proposed approach leverages the strengths of CNNs for feature extraction and kernels for measuring similarity, which accomplishes a general learning framework of match kernels for action recognition. We have conducted extensive experiments on five challenging realistic data sets including Youtube, UCF50, UCF101, Penn action, and HMDB51. The proposed approach achieves high performance and substantially surpasses the state-of-the-art algorithms by large margins, which demonstrates the great effectiveness of proposed approach for action recognition.
Lei Zhang 0093, Xiantong Zhen, Ling Shao 0001, Jingkuan Song
IEEE Trans. Image Process.2
2019 Deep Ensemble Machine for Video Classification
abstract
Video classification has been extensively researched in computer vision due to its wide spread applications. However, it remains an outstanding task because of the great challenges in effective spatial-temporal feature extraction and efficient classification with high-dimensional video representations. To address these challenges, in this paper, we propose an end-to-end learning framework called deep ensemble machine (DEM) for video classification. Specifically, to establish effective spatio-temporal features, we propose using two deep convolutional neural networks (CNNs), i.e., vision and graphics group and C3-D to extract heterogeneous spatial and temporal features for complementary representations. To achieve efficient classification, we propose ensemble learning based on random projections aiming to transform high-dimensional features into a set of lower dimensional compact features in subspaces; an ensemble of classifiers is trained on the subspaces and combined with a weighting layer during the backpropagation. To further enhance the performance, we introduce rectified linear encoding (RLE) inspired from error-correcting output coding to encode the initial outputs of classifiers, followed by a softmax layer to produce the final classification results. DEM combines the strengths of deep CNNs and ensemble learning, which establishes a new end-to-end learning architecture for more accurate and efficient video classification. We show the great effectiveness of DEM by extensive experiments on four data sets for diverse video classification tasks including action recognition and dynamic scene classification. Results have shown that DEM achieves high performance on all tasks with an improvement of up to 13% on CIFAR10 data set over the baseline model.
Jiewan Zheng, Xianbin Cao 0001, Baochang Zhang 0001, Xiantong Zhen, Xiangbo Su
IEEE Trans. Neural Networks Learn. Syst.4
2018 Deep Collaborative Tracking Networks
Xiantong Zhen, Baochang Zhang 0001, Xianbin Cao 0001
BMVC2
2018 In Defense of Single-column Networks for Crowd Counting
Ze Wang 0008, Zehao Xiao, Qiang Qiu 0001, Xiantong Zhen, Xianbin Cao 0001
BMVC5
2018 Attentional Alignment Networks
Baochang Zhang 0001, Xiantong Zhen, Xianbin Cao 0001
BMVC5
2018 Direct Shape Regression Networks for End-to-End Face Alignment
abstract
Face alignment has been extensively studied in computer vision community due to its fundamental role in facial analysis, but it remains an unsolved problem. The major challenges lie in the highly nonlinear relationship between face images and associated facial shapes, which is coupled by underlying correlation of landmarks. Existing methods mainly rely on cascaded regression, suffering from intrinsic shortcomings, e.g., strong dependency on initialization and failure to exploit landmark correlations. In this paper, we propose the direct shape regression network (DSRN) for end-to-end face alignment by jointly handling the aforementioned challenges in a unified framework. Specifically, by deploying doubly convolutional layer and by using the Fourier feature pooling layer proposed in this paper, DSRN efficiently constructs strong representations to disentangle highly nonlinear relationships between images and shapes; by incorporating a linear layer of low-rank learning, DSRN effectively encodes correlations of landmarks to improve performance. DSRN leverages the strengths of kernels for nonlinear feature extraction and neural networks for structured prediction, and provides the first end-to-end learning architecture for direct face alignment. Its effectiveness and generality are validated by extensive experiments on five benchmark datasets, including AFLW, 300W, CelebA, MAFL, and 300VW. All empirical results demonstrate that DSRN consistently produces high performance and in most cases surpasses state-of-the-art.
Xiantong Zhen, Xianglong Liu 0001, Cheng Deng 0002, Vassilis Athitsos, Heng Huang 0001
CVPR2
2018 Spatial Ensemble Kernel Learning for Scene Classification
abstract
Scene recognition is one of the most important tasks in computer vision. Apart from appearance, spatial layout carries the crucial cue for discriminative representation. In this paper, we propose spatial ensemble kernel (SEK) learning, which enables fusion of multi-scale spatial information to achieve compact while discriminative representation of scenes. Based on the spatial pyramid, SEK combines the CNN features in each level of the pyramid in an ensemble and fuse them by kernels. By kernel approximation, we achieve Fourier feature embedding of CNN features in each scale, which establishes a nonlinear layer of the neural network with a cosine activation function. The parameters of the nonlinear layer can be learned jointly in one single optimization framework by supervised learning, which enables compact and discriminative feature representations. We show the effectiveness of the proposed SEK on two recent scene benchmark datasets, i.e., MIT indoor and SUN 397. The propose SEK produces high performance on two datasets which are competitive to state-of-the-art algorithms.
Lei Zhang 0093, Xiantong Zhen, Qiujing Zhang
ICASSP2
2018 Deep appearance and motion learning for egocentric activity recognition
Xuanhan Wang, Lianli Gao, Jingkuan Song, Xiantong Zhen, Nicu Sebe, Heng Tao Shen
Neurocomputing4
2018 Multi-Target Regression via Robust Low-Rank Learning
abstract
Multi-target regression has recently regained great popularity due to its capability of simultaneously learning multiple relevant regression tasks and its wide applications in data mining, computer vision and medical image analysis, while great challenges arise from jointly handling inter-target correlations and input-output relationships. In this paper, we propose Multi-layer Multi-target Regression (MMR) which enables simultaneously modeling intrinsic inter-target correlations and nonlinear input-output relationships in a general framework via robust low-rank learning. Specifically, the MMR can explicitly encode inter-target correlations in a structure matrix by matrix elastic nets (MEN); the MMR can work in conjunction with the kernel trick to effectively disentangle highly complex nonlinear input-output relationships; the MMR can be efficiently solved by a new alternating optimization algorithm with guaranteed convergence. The MMR leverages the strength of kernel methods for nonlinear feature learning and the structural advantage of multi-layer learning architectures for inter-target correlation modeling. More importantly, it offers a new multi-layer learning paradigm for multi-target regression which is endowed with high generality, flexibility and expressive ability. Extensive experimental evaluation on 18 diverse real-world datasets demonstrates that our MMR can achieve consistently high performance and outperforms representative state-of-the-art algorithms, which shows its great effectiveness and generality for multivariate prediction.
Xiantong Zhen, Mengyang Yu, Xiaofei He 0001, Shuo Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Multitarget Sparse Latent Regression
abstract
Multitarget regression has recently generated intensive popularity due to its ability to simultaneously solve multiple regression tasks with improved performance, while great challenges stem from jointly exploring inter-target correlations and input-output relationships. In this paper, we propose multitarget sparse latent regression (MSLR) to simultaneously model intrinsic intertarget correlations and complex nonlinear input-output relationships in one single framework. By deploying a structure matrix, the MSLR accomplishes a latent variable model which is able to explicitly encode intertarget correlations via -norm-based sparse learning; the MSLR naturally admits a representer theorem for kernel extension, which enables it to flexibly handle highly complex nonlinear input-output relationships; the MSLR can be solved efficiently by an alternating optimization algorithm with guaranteed convergence, which ensures efficient multitarget regression. Extensive experimental evaluation on both synthetic data and six greatly diverse real-world data sets shows that the proposed MSLR consistently outperforms the state-of-the-art algorithms, which demonstrates its great effectiveness for multivariate prediction.
Xiantong Zhen, Mengyang Yu, Feng Zheng 0001, Ilanit Ben Nachum, Mousumi Bhaduri, David T. Laidley, Shuo Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2017 Learning Deep Match Kernels for Image-Set Classification
abstract
Image-set classification has recently generated great popularity due to its widespread applications in computer vision. The great challenges arise from effectively and efficiently measuring the similarity between image sets with high inter-class ambiguity and huge intra-class variability. In this paper, we propose deep match kernels (DMK) to directly measure the similarity between image sets in the match kernel framework. Specifically, we build deep local match kernels between images upon arc-cosine kernels, which can faithfully characterize the similarity between images by mimicking deep neural networks, we introduce anchors to aggregate those deep local match kernels into a global match kernel between image sets, which is learned in a supervised way by kernel alignment and therefore more discriminative. The DMK provides the first match kernel framework for image-set classification, which removes specific assumptions usually required in previous approaches and is computationally more efficient. We conduct extensive experiments on four datasets for three diverse image-set classification tasks. The DMK achieves high performance and consistently surpasses state-of-the-art methods, showing its great effectiveness for image-set classification.
Haoliang Sun, Xiantong Zhen, Yuanjie Zheng, Gongping Yang 0001, Yilong Yin, Shuo Li 0001
CVPR2
2017 Realistic human action recognition: When CNNS meet LDS
abstract
In this paper, we proposed new framework for human action representation, which leverages the strengths of convolutional neural networks (CNNs) and the linear dynamical system (LDS) to represent both spatial and temporal structures of actions in videos. We make two principal contributions: first, we incorporate image-trained CNNs to detect action clip concepts, which takes advantage of different levels of information by combining the two layers in CNNs trained from images; Second, we further propose adopting a linear dynamical system (LDS) to model the relationships between these clip concepts, which captures temporal structures of actions. We have applied the proposed method on two challenging realistic benchmark datasets, and our method achieves high performance up to 86.16% on the YouTube and 82.76% UCF50 datasets, which largely outperforms most of the state-of-the-art algorithms with more sophisticated techniques.
Lei Zhang 0093, Yangyang Feng, Xuezhi Xiang, Xiantong Zhen
ICASSP4
2017 Learning discriminant grassmann kernels for image-set classification
abstract
Image-set classification has recently generated great popularity due to widespread application to challenging tasks in computer vision. The great challenges arise from measuring the similarity between image sets which usually exhibit huge inter-class ambiguity and intra-class variation. In this paper, based on the assumption that each image set as a linear subspace can be treated as a point on a Grassmann manifold, we propose discriminant Grassmann kernels (DGK) of principal angles between subspaces. To tackle the ambiguity and variation, we propose learning the DGK via kernel target alignment, which achieves kernels of great discrimination by maximizing correlations with class labels. The proposed DGK has been evaluated on two challenging datasets including the ETH-80 and UCSD datasets for object recognition and video-based traffic congestion recognition, respectively. Extensive experiments have shown that the proposed DGKs achieves state-of-the-art performance and surpasses most of previous methods, which demonstrates the great effectiveness of the DGKs for image-set classification.
Lei Zhang 0093, Xuezhi Xiang, Xiantong Zhen
ICIP5
2017 Direct and simultaneous estimation of cardiac four chamber volumes by multioutput sparse regression
Xiantong Zhen, Heye Zhang, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001
Medical Image Anal.1
2017 Real-time visual tracking based on improved perceptual hashing
Mengjuan Fei, Zhaojie Ju, Xiantong Zhen, Jing Li 0027
Multim. Tools Appl.3
2017 Supervised Local Descriptor Learning for Human Action Recognition
abstract
Local features have been widely used in computer vision tasks, e.g., human action recognition, but it tends to be an extremely challenging task to deal with large-scale local features of high dimensionality with redundant information. In this paper, we propose a novel fully supervised local descriptor learning algorithm called discriminative embedding method based on the image-to-class distance (I2CDDE) to learn compact but highly discriminative local feature descriptors for more accurate and efficient action recognition. By leveraging the advantages of the I2C distance, the proposed I2CDDE incorporates class labels to enable fully supervised learning of local feature descriptors, which achieves highly discriminative but compact local descriptors. The objective of our I2CDDE is to minimize the I2C distances from samples to their corresponding classes while maximizing the I2C distances to the other classes in the low-dimensional space. To further improve the performance, we propose incorporating a manifold regularization based on the graph Laplacian into the objective function, which can enhance the smoothness of the embedding by extracting the local intrinsic geometrical structure. The proposed I2CDDE for the first time achieves fully supervised learning of local feature descriptors. It significantly improves the performance of I2C-based methods by increasing the discriminative ability of local features while greatly reducing the computational burden by dimensionality reduction to handle large-scale data. We apply the proposed I2CDDE algorithm to human action recognition on four widely used benchmark datasets. The results have shown that I2CDDE can significantly improve I2C-based classifiers and achieves state-of-the-art performance.
Xiantong Zhen, Feng Zheng 0001, Ling Shao 0001, Xianbin Cao 0001, Dan Xu 0002
IEEE Trans. Multim.1
2017 Descriptor Learning via Supervised Manifold Regularization for Multioutput Regression
abstract
Multioutput regression has recently shown great ability to solve challenging problems in both computer vision and medical image analysis. However, due to the huge image variability and ambiguity, it is fundamentally challenging to handle the highly complex input-target relationship of multioutput regression, especially with indiscriminate high-dimensional representations. In this paper, we propose a novel supervised descriptor learning (SDL) algorithm for multioutput regression, which can establish discriminative and compact feature representations to improve the multivariate estimation performance. The SDL is formulated as generalized low-rank approximations of matrices with a supervised manifold regularization. The SDL is able to simultaneously extract discriminative features closely related to multivariate targets and remove irrelevant and redundant information by transforming raw features into a new low-dimensional space aligned to targets. The achieved discriminative while compact descriptor largely reduces the variability and ambiguity for multioutput regression, which enables more accurate and efficient multivariate estimation. We conduct extensive evaluation of the proposed SDL on both synthetic data and real-world multioutput regression tasks for both computer vision and medical image analysis. Experimental results have shown that the proposed SDL can achieve high multivariate estimation accuracy on all tasks and largely outperforms the algorithms in the state of the arts. Our method establishes a novel SDL framework for multioutput regression, which can be widely used to boost the performance in different applications.
Xiantong Zhen, Mengyang Yu, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2016 Realistic human action recognition: When deep learning meets VLAD
abstract
Human action recognition from realistic scenarios is extremely challenging due to large intra-class variation and complex background clutters. In this paper, by leveraging the strength of deep learning and vector of locally aggregated descriptors (VLAD), we propose a new methods for human action recognition from realistic datsets. We adopt stack convolu-tional independent subspace analysis (ISA) networks to learn 3D cuboid representation directly from spatio-temporal video data; we propose an improved VLAD by incorporating the spatio-temporal geometrical information to encode the deep learned local features. On two challenging realistic datasets: the YouTube action and HMDB51 datasets, the proposed method achieves state-of-the-art performance with an efficient linear SVM classifier, which is competitive with and even better than existing sophisticated algorithms.
Lei Zhang 0093, Yangyang Feng, Jiqing Han 0001, Xiantong Zhen
ICASSP4
2016 Towards optimal vlad for human action recognition from still images
abstract
Human action recognition from still image has recently drawn increasing attention in human behavior analysis vision and also poses great challenges due to the huge inter ambiguity and intra variability. Vector of locally aggregated descriptors (VLAD) has achieved state-of-the-art performance in many image classification tasks based on local features. The great success of VLAD is largely due to its high descriptive ability and computational efficiency. In this paper, towards optimal VLAD representations for human action recognition from still images, we improve VLAD by tackling two important issues in VLAD including empty cavity and assignment ambiguity. The empty cavity issue severely compromises the performance of VLAD and has long been overlooked. We investigate the empty cavity and provide an effective solution to deal with it, which largely improves the performance of VLAD; we propose middle level assignments to conquer the assignment ambiguity, which are more reliable and can provide more useful information for realistic activity. We have conducted extensive experiments on two widely-used benchmarks to validate the proposed method for human action recognition from still images. Our method produces competitive performance with state-of-the-art algorithms.
Lei Zhang 0093, Xiantong Zhen, Jiqing Han 0001
ICASSP2
2016 Multi-task Shape Regression for Medical Image Segmentation
abstract
In this paper, we propose a general segmentation framework of Multi-Task Shape Regression (MTSR) which formulates segmentation as multi-task learning to leverage its strength of jointly solving multiple tasks enhanced by capturing task correlations. The MTSR entirely estimates coordinates of all points on shape contours by multi-task regression, where estimation of each coordinate corresponds to a regression task; the MTSR can jointly handle nonlinear relationships between image appearance and shapes while capturing holistic shape information by encoding coordinate correlations, which enables estimation of highly variable shapes, even with vague edge or region inhomogeneity. The MTSR achieves a long-desired general framework without relying on any specific assumptions or initialization, which enables flexible and fully automatic segmentation of multiple objects simultaneously, for different applications irrespective of modalities. The MTSR is validated on six representative applications of diverse images, achieves consistently high performance with dice similarity coefficient (DSC) up to 0.93 and largely outperforms state of the arts in each application, which demonstrates its effectiveness and generality for medical image segmentation.
Xiantong Zhen, Yilong Yin, Mousumi Bhaduri, Ilanit Ben Nachum, David T. Laidley, Shuo Li 0001
MICCAI (3)1
2016 Spatial and temporal scoring for egocentric video summarization
Zhao Guo, Lianli Gao, Xiantong Zhen, Fuhao Zou, Fumin Shen, Kai Zheng 0001
Neurocomputing3
2016 Action recognition via spatio-temporal local features: A comprehensive study
Xiantong Zhen, Ling Shao 0001
Image Vis. Comput.1
2016 Handcrafted vs. learned representations for human action recognition
Xiantong Zhen, Ling Shao 0001, Stephen J. Maybank, Rama Chellappa
Image Vis. Comput.1
2016 Multi-scale deep networks and regression forests for direct bi-ventricular volume estimation
Xiantong Zhen, Zhijie Wang 0003, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001
Medical Image Anal.1
2016 Local Feature Discriminant Projection
abstract
In this paper, we propose a novel subspace learning algorithm called Local Feature Discriminant Projection (LFDP) for supervised dimensionality reduction of local features. LFDP is able to efficiently seek a subspace to improve the discriminability of local features for classification. We make three novel contributions. First, the proposed LFDP is a general supervised subspace learning algorithm which provides an efficient way for dimensionality reduction of large-scale local feature descriptors. Second, we introduce the Differential Scatter Discriminant Criterion (DSDC) to the subspace learning of local feature descriptors which avoids the matrix singularity problem. Third, we propose a generalized orthogonalization method to impose on projections, leading to a more compact and less redundant subspace. Extensive experimental validation on three benchmark datasets including UIUC-Sports, Scene-15 and MIT Indoor demonstrates that the proposed LFDP outperforms other dimensionality reduction methods and achieves state-of-the-art performance for image classification.
Mengyang Yu, Ling Shao 0001, Xiantong Zhen, Xiaofei He 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Supervised descriptor learning for multi-output regression
abstract
Descriptor learning has recently drawn increasing attention in computer vision, Existing algorithms are mainly developed for classification rather than for regression which however has recently emerged as a powerful tool to solve a broad range of problems, e.g., head pose estimation. In this paper, we propose a novel supervised descriptor learning (SDL) algorithm to establish a discriminative and compact feature representation for multi-output regression. By formulating as generalized low-rank approximations of matrices with a supervised manifold regularization (SMR), the SDL removes irrelevant and redundant information from raw features by transforming into a low-dimensional space under the supervision of multivariate targets. The obtained discriminative while compact descriptor largely reduces the variability and ambiguity in multi-output regression, and therefore enables more accurate and efficient multivariate estimation. We demonstrate the effectiveness of the proposed SDL algorithm on a representative multi-output regression task: head pose estimation using the benchmark Pointing'04 dataset. Experimental results show that the SDL can achieve high pose estimation accuracy and significantly outperforms state-of-the-art algorithms by an error reduction up to 27.5%. The proposed SDL algorithm provides a general descriptor learning framework in a supervised way for multi-output regression which can largely boost the performance of existing multi-output regression tasks.
Xiantong Zhen, Zhijie Wang 0003, Mengyang Yu, Shuo Li 0001
CVPR1
2015 Dimensionality reduction by supervised locality analysis
abstract
High-dimensional feature representations have recently been widely used for image classification, which not only induce large storage requirement and high computational complexity, but also tend to be lack of discrimination due to redundant and noisy features. In this paper, we propose a novel algorithm named supervised locality analysis (SLA) for dimensionality reduction. In contrast to conventional dimensionality reduction methods, the proposed SLA incorporates supervision into locality analysis by fully exploring multi-class distributions, which can handle the non-linear data structure while preserving intrinsic discriminative information. The obtained compact and highly discriminative features by the SLA is enables more accurate and efficient classification. Moreover, the SLA can be used for supervised dimensionality reduction of both handcrafted and deep learning based features. We have conduced experiments to evaluate the proposed SLA on three datasets for image classification. The SLA has produced state-of-the-art performance and largely outperformed widely-used dimensionality reduction methods.
Lei Zhang 0093, Peipei Peng, Xuezhi Xiang, Xiantong Zhen
ICIP4
2015 Direct and Simultaneous Four-Chamber Volume Estimation by Multi-Output Regression
Xiantong Zhen, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001
MICCAI (1)1
2015 Regression Segmentation for M3 Spinal Images
abstract
Clinical routine often requires to analyze spinal images of multiple anatomic structures in multiple anatomic planes from multiple imaging modalities (M(3)). Unfortunately, existing methods for segmenting spinal images are still limited to one specific structure, in one specific plane or from one specific modality (S(3)). In this paper, we propose a novel approach, Regression Segmentation, that is for the first time able to segment M(3) spinal images in one single unified framework. This approach formulates the segmentation task innovatively as a boundary regression problem: modeling a highly nonlinear mapping function from substantially diverse M(3) images directly to desired object boundaries. Leveraging the advancement of sparse kernel machines, regression segmentation is fulfilled by a multi-dimensional support vector regressor (MSVR) which operates in an implicit, high dimensional feature space where M(3) diversity and specificity can be systematically categorized, extracted, and handled. The proposed regression segmentation approach was thoroughly tested on images from 113 clinical subjects including both disc and vertebral structures, in both sagittal and axial planes, and from both MRI and CT modalities. The overall result reaches a high dice similarity index (DSI) 0.912 and a low boundary distance (BD) 0.928 mm. With our unified and expendable framework, an efficient clinical tool for M(3) spinal image segmentation can be easily achieved, and will substantially benefit the diagnosis and treatment of spinal diseases.
Zhijie Wang 0003, Xiantong Zhen, KengYeow Tay, Said Osman, Walter Romano, Shuo Li 0001
IEEE Trans. Medical Imaging2
2015 Correction to "Regression Segmentation for M3 Spinal Images"
Zhijie Wang 0003, Xiantong Zhen, KengYeow Tay, Said Osman, Walter Romano, Shuo Li 0001
IEEE Trans. Medical Imaging2
2014 Discriminative Embedding via Image-to-Class Distances
Xiantong Zhen, Ling Shao 0001, Feng Zheng 0001
BMVC1
2014 Learning semantic kernels for scene classification
abstract
In this paper we propose to learn semantic kernels for scene classification. We first decompose the Object Bank representation into subspaces associated with each object, Anchor Objects are then created by clustering for each scene class separately. The Anchor Distances are computed to measure the distance between objects to scene classes. In order to take the advantage of the discriminative information from different scene classes, we propose semantic kernels based on the anchor distances to different classes for scene classification. Through extensive experiments on two benchmark datasets: UIUC-Sports dataset and 15-Scene dataset, we prove that the proposed Semantic Kernels can significantly improve the original Object Bank and achieve state-of-the-art performance.
Lei Zhang 0093, Xiantong Zhen, Jiqing Han 0001, Xuezhi Xiang
ICASSP2
2014 A Performance Evaluation on Action Recognition with Local Features
abstract
Local features have played an important role in visual recognition. Methods based on local features, e.g., the bag-of-words (BoW) model and sparse coding, have shown their effectiveness in image and object recognition in the past decades. Recently, many new techniques, including the improvements of BoW and sparse coding as well as the non-parametric naive bayes nearest neighbor (NBNN) classifier, have been proposed and advanced the state-of-the-art in the image domain. However, in the video domain, the BoW model still dominates the action recognition field. It is unclear how effective the state-of-the-art techniques widely used in the image domain would perform on action recognition. To fill this gap, we aim to implement and provide a systematic study of these techniques on action recognition, and compare their performance under a unified evaluation framework. Other techniques such as match kernels and random forest, which have also demonstrated their potential in handling local features, are also included for a comprehensive evaluation. Extensive experiments have been conducted on three benchmarks including the KTH, the UCF-YouTube and the HMDB51 datasets, and results and findings are analyzed and discussed.
Xiantong Zhen, Ling Shao 0001
ICPR1
2014 Direct Estimation of Cardiac Bi-ventricular Volumes with Regression Forests
Xiantong Zhen, Zhijie Wang 0003, Ali Islam, Mousumi Bhaduri, Ian Chan, Shuo Li 0001
MICCAI (2)1
2014 Action recognition by spatio-temporal oriented energies
Xiantong Zhen, Ling Shao 0001, Xuelong Li 0001
Inf. Sci.1
2014 Robust point pattern matching based on spectral context
Jun Tang 0007, Ling Shao 0001, Xiantong Zhen
Pattern Recognit.3
2014 Spatio-Temporal Laplacian Pyramid Coding for Action Recognition
abstract
We present a novel descriptor, called spatio-temporal Laplacian pyramid coding (STLPC), for holistic representation of human actions. In contrast to sparse representations based on detected local interest points, STLPC regards a video sequence as a whole with spatio-temporal features directly extracted from it, which prevents the loss of information in sparse representations. Through decomposing each sequence into a set of band-pass-filtered components, the proposed pyramid model localizes features residing at different scales, and therefore is able to effectively encode the motion information of actions. To make features further invariant and resistant to distortions as well as noise, a bank of 3-D Gabor filters is applied to each level of the Laplacian pyramid, followed by max pooling within filter bands and over spatio-temporal neighborhoods. Since the convolving and pooling are performed spatio-temporally, the coding model can capture structural and motion information simultaneously and provide an informative representation of actions. The proposed method achieves superb recognition rates on the KTH, the multiview IXMAS, the challenging UCF Sports, and the newly released HMDB51 datasets. It outperforms state of the art methods showing its great potential on action recognition.
Ling Shao 0001, Xiantong Zhen, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Cybern.2
2014 Learning Object-to-Class Kernels for Scene Classification
abstract
High-level image representations have drawn increasing attention in visual recognition, e.g., scene classification, since the invention of the object bank. The object bank represents an image as a response map of a large number of pretrained object detectors and has achieved superior performance for visual recognition. In this paper, based on the object bank representation, we propose the object-to-class (O2C) distances to model scene images. In particular, four variants of O2C distances are presented, and with the O2C distances, we can represent the images using the object bank by lower-dimensional but more discriminative spaces, called distance spaces, which are spanned by the O2C distances. Due to the explicit computation of O2C distances based on the object bank, the obtained representations can possess more semantic meanings. To combine the discriminant ability of the O2C distances to all scene classes, we further propose to kernalize the distance representation for the final classification. We have conducted extensive experiments on four benchmark data sets, UIUC-Sports, Scene-15, MIT Indoor, and Caltech-101, which demonstrate that the proposed approaches can significantly improve the original object bank approach and achieve the state-of-the-art performance.
Lei Zhang 0093, Xiantong Zhen, Ling Shao 0001
IEEE Trans. Image Process.2
2013 Human Action Retrieval via efficient feature matching
abstract
As a large proportion of the available video media concerns humans, human action retrieval is posed as a new topic in the domain of content-based video retrieval. For retrieving complex human actions, measuring the similarity between two videos represented by local features is a critical issue. In this paper, a fast and explicit feature correspondence approach is presented to compute the match cost serving as the similarity metric. Then the proposed similarity metric is embedded into the framework of manifold ranking for action retrieval. In contrast to the Bag-of-Words model and its variants, our method yields an encouraging improvement of accuracy on the KTH and the UCF YouTube datasets with reasonably efficient computation.
Jun Tang 0007, Ling Shao 0001, Xiantong Zhen
AVSS3
2013 Towards optimal object bank for scene classification
abstract
High-level image representations have drawn increasing attention in visual recognition, e.g., scene classification, since the invention of the object bank (OB). The object bank represents an image as a response map of a large number of pre-trained object detectors and has achieved superior performances for visual recognition. However, the object bank representation can be further improved by considering the distributions of the object across categories and the discriminative contributions to the image representation. In this paper, we propose an optimal object bank (OOB) by imposing weights on the detectors according to their discriminative abilities. Through extensive experiments on two benchmark datasets: UIUC-Sports dataset and 15-Scene dataset, we prove that the proposed OOB can significantly improve the original object bank and achieves state-of-the-art performances.
Lei Zhang 0093, Shouzhi Xie, Xiantong Zhen
ICASSP3
2013 Recognizing actions via sparse coding on structure projection
abstract
In this paper, we propose a novel method for human action recognition based on sparse coding with a pyramid matching. Spatio-temporal interest points (STIPs) are firstly detected by a newly developed detector named spatio-temporal steerable detector (STSD). To effectively capture the distribution of STIPs in the video sequence, we propose to project the STIPs onto the three orthogonal planes (TOP), and we employ a sparse coding algorithm combined with the spatial pyramid matching to encode the layout of STIPs. Therefore the structure of an action are sufficiently encoded, obtaining a informative holistic descriptor for action representation. Extensive experiments have been conducted on KTH and HMDB51 datasets. Our method achieves the state-of-the-art performance for action recognition showing the effectiveness of the proposed methods for human action representation.
Lei Zhang 0093, Xiantong Zhen
ICIP3
2013 Discriminative high-level representations for scene classification
abstract
High-level image representations, e.g, Object Bank, have drawn increasing attention in visual recognition. In this paper, we propose a discriminative high-level representation based on object bank for scene classification. By projecting the high-level features from the object bank into discriminative subspaces, which are obtained by clustering the features in a supervised way, the final representations are more compact and discriminative. We have conducted extensive experiments on two benchmark datasets: UIUC-Sports dataset and 15-Scene dataset, which demonstrates that the proposed approach can significantly improve the original object bank and achieves state-of-the-art performances.
Lei Zhang 0093, Shouzhi Xie, Xiantong Zhen
ICIP3
2013 Combining appearance and structural features for human action recognition
Ling Shao 0001, Xiantong Zhen, Yan Liu 0004
Neurocomputing3
2013 A local descriptor based on Laplacian pyramid coding for action recognition
Xiantong Zhen, Ling Shao 0001
Pattern Recognit. Lett.1
2013 Embedding Motion and Structure Features for Action Recognition
abstract
We propose a novel method to model human actions by explicitly coding motion and structure features that are separately extracted from video sequences. Firstly, the motion template (one feature map) is applied to encode the motion information and image planes (five feature maps) are extracted from the volume of differences of frames to capture the structure information. The Gaussian pyramid and center-surround operations are performed on each of the six obtained feature maps, decomposing each feature map into a set of subband maps. Biologically inspired features are then extracted by successively applying Gabor filtering and max pooling on each subband map. To make a compact representation, discriminative locality alignment is employed to embed the high-dimensional features into a low-dimensional manifold space. In contrast to sparse representations based on detected interest points, which suffer from the loss of structure information, the proposed model takes into account the motion and structure information simultaneously and integrates them in a unified framework; it therefore provides an informative and compact representation of human actions. The proposed method is evaluated on the KTH, the multiview IXMAS, and the challenging UCF sports datasets and outperforms state-of-the-art techniques on action recognition.
Xiantong Zhen, Ling Shao 0001, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2013 Learning Discriminative Key Poses for Action Recognition
abstract
In this paper, we present a new approach for human action recognition based on key-pose selection and representation. Poses in video frames are described by the proposed extensive pyramidal features (EPFs), which include the Gabor, Gaussian, and wavelet pyramids. These features are able to encode the orientation, intensity, and contour information and therefore provide an informative representation of human poses. Due to the fact that not all poses in a sequence are discriminative and representative, we further utilize the AdaBoost algorithm to learn a subset of discriminative poses. Given the boosted poses for each video sequence, a new classifier named weighted local naive Bayes nearest neighbor is proposed for the final action classification, which is demonstrated to be more accurate and robust than other classifiers, e.g., support vector machine (SVM) and naive Bayes nearest neighbor. The proposed method is systematically evaluated on the KTH data set, the Weizmann data set, the multiview IXMAS data set, and the challenging HMDB51 data set. Experimental results manifest that our method outperforms the state-of-the-art techniques in terms of recognition rate.
Li Liu 0004, Ling Shao 0001, Xiantong Zhen, Xuelong Li 0001
IEEE Trans. Cybern.3
2012 High order co-occurrence of visualwords for action recognition
abstract
This paper exploits the high order co-occurrence information for human action representation. Based on the bag-of-words (BoW) model, visual words are mapped into a co-occurrence space through latent semantic analysis (LSA). High order co-occurrence of the visual words is well captured and therefore the representation of actions in the co-occurrence space becomes more informative and compact. Since the representation is effective and efficient, and is less affected by the sizes of the codebook, it can be easily integrated into models based on BoW. Evaluations on the benchmark KTH dataset and the realistic HMDB51 dataset demonstrates that the proposed approach significantly improves the baseline BoW model and therefore is promising for human action recognition.
Lei Zhang 0093, Xiantong Zhen, Ling Shao 0001
ICIP2