Yi Zhou 0007

dblp:01/1901-7 · DBLP profile ↗
← Back
52ranked-venue papers
16as first author
36since 2021 · last 2026
0000-0003-3021-3229ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 11 first-author · 19 since 2021Artificial intelligence and machine learning · 23 · 7 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 6 first-author · 16 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Bidirectional Channel-selective Semantic Interaction for Semi-Supervised Medical Segmentation
abstract
Semi-supervised medical image segmentation is an effective method for addressing scenarios with limited labeled data. Existing methods mainly rely on frameworks such as mean teacher and dual-stream consistency learning. These approaches often face issues like error accumulation and model structural complexity, while also neglecting the interaction between labeled and unlabeled data streams. To overcome these challenges, we propose a Bidirectional Channel-selective Semantic Interaction (BCSI) framework for semi-supervised medical image segmentation. First, we propose a Semantic-Spatial Perturbation (SSP) mechanism, which disturbs the data using two strong augmentation operations and leverages unsupervised learning with pseudo-labels from weak augmentations. Additionally, we employ consistency on the predictions from the two strong augmentations to further improve model stability and robustness. Second, to reduce noise during the interaction between labeled and unlabeled data, we propose a Channel-selective Router (CR) component, which dynamically selects the most relevant channels for information exchange. This mechanism ensures that only highly relevant features are activated, minimizing unnecessary interference. Finally, the Bidirectional Channel-wise Interaction (BCI) strategy is employed to supplement additional semantic information and enhance the representation of important channels. Experimental results on multiple benchmarking 3D medical datasets demonstrate that the proposed method outperforms existing semi-supervised approaches.
Kaiwen Huang 0002, Yizhe Zhang 0001, Yi Zhou 0007, Tianyang Xu 0001, Tao Zhou 0002
AAAI3
2026 Structured prompt-guided knowledge injection for medical image segmentation
Kelei He, Yizhe Zhang 0001, Yi Zhou 0007, Tao Zhou 0002, Dong Liang 0001
Pattern Recognit.4
2026 Uncertainty-Guided Prototype Reliability Enhancement Network for Few-Shot Medical Image Segmentation
abstract
Few-Shot Learning (FSL) has garnered increasing attention for data-scarce scenarios, particularly in medical segmentation tasks where only a few labeled data points are available. Existing few-shot segmentation methods typically learn prototypes from support images and employ nearest-neighbor searching to segment query images. Despite notable progress, effectively learning prototypes for each class remains a challenging task to achieve promising results. In this paper, we propose an Uncertainty-guided Prototype Reliability Enhancement Network (UPRE-Net) for few-shot medical image segmentation. Specifically, we present a dual-support branch to maximize the extraction of information from support images through augmentation techniques. To enhance the reliability of prototypes, we propose an Uncertainty-guided Prototype Generation (UPG) module. Within the UPG module, we first extract both global and local prototypes for each class and then apply uncertainty measures to select the most informative prototypes. Additionally, to effectively combine the prediction results from the dual-support branch, we present a Reliable Dynamic Fusion (RDF) module. This module dynamically integrates the two prediction results to generate a more reliable output. Furthermore, we present an Uncertainty-induced Weighted Loss (UWL) to ensure that the model pays more attention to these regions with high uncertainty. Experiments on four benchmark medical image datasets demonstrate that our proposed model significantly outperforms state-of-the-art methods. The code will be released at https://github.com/taozh2017/UPRENet.
Tao Zhou 0002, Kaiwen Huang 0002, Yi Zhou 0007, Haofeng Zhang 0001, Boqiang Fan, Huazhu Fu
IEEE Trans. Medical Imaging4
2025 Text-Driven Multiplanar Visual Interaction for Semi-supervised Medical Image Segmentation
Kaiwen Huang 0002, Yi Zhou 0007, Huazhu Fu, Yizhe Zhang 0001, Chen Gong 0002, Tao Zhou 0002
MICCAI (5)2
2025 Continual Retinal Vision-Language Pre-training upon Incremental Imaging Modalities
Yuang Yao, Yi Zhou 0007, Tao Zhou 0002
MICCAI (5)3
2025 FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait Synthesis
abstract
Achieving high-fidelity and accurate lip-speech synchronization in audio-driven talking portrait synthesis is particularly challenging. Some studies utilize multi-stage pipelines or diffusion models for high-quality talking portraits; however, they suffer from excessive computational costs. Some approaches achieve remarkable progress on specific individuals with low resource requirements, yet still exhibit mismatched lip movements. The aforementioned methods are modeled in the pixel domain. We observed that there are noticeable discrepancies in the frequency domain between the synthesized talking videos and natural videos. Currently, no research on talking portrait synthesis has considered this aspect. To address this, we propose a FREquency-modulated, high-fidelity, and real-time Audio-driven talKing portrait synthesis framework, named FREAK, which models talking portraits from the frequency domain perspective, enhancing the fidelity and naturalness of the synthesized portraits. FREAK introduces two novel frequency-based modules: 1) the Visual Encoding Frequency Modulator (VEFM) to couple multi-scale visual features in the frequency domain, better preserving visual frequency information and reducing the gap in the frequency spectrum between synthesized and natural frames. and 2) the Audio Visual Frequency Modulator (AVFM) to help the model learn the talking pattern in the frequency domain and improve audio-visual synchronization. Additionally, we optimize the model in both pixel domain and frequency domain jointly. Furthermore, FREAK supports seamless switching between one-shot and video dubbing settings, offering enhanced flexibility. Due to its superior performance, it can simultaneously support high-resolution video results and real-time inference. Extensive experiments demonstrate that our method synthesizes high-fidelity talking portraits with detailed facial textures and precise lip synchronization in real-time, outperforming state-of-the-art methods.
Ziqi Ni, Ao Fu, Yi Zhou 0007
ICMR3
2025 WeakPolyp-SAM: Segment Anything Model-driven weakly-supervised polyp segmentation
Tao Zhou 0002, Yunqi Gu, Yi Zhou 0007, Yizhe Zhang 0001, Ye Wu 0001, Huazhu Fu
Knowl. Based Syst.4
2025 Dual-scale enhanced and cross-generative consistency learning for semi-supervised medical image segmentation
Yunqi Gu, Tao Zhou 0002, Yizhe Zhang 0001, Yi Zhou 0007, Kelei He, Chen Gong 0002, Huazhu Fu
Pattern Recognit.4
2025 Uncertainty-Aware Cross-Training for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised learning has gained considerable popularity in medical image segmentation tasks due to its capability to reduce reliance on expert-examined annotations. Several mean-teacher (MT) based semi-supervised methods utilize consistency regularization to effectively leverage valuable information from unlabeled data. However, these methods often heavily rely on the student model and overlook the potential impact of cognitive biases within the model. Furthermore, some methods employ co-training using pseudo-labels derived from different inputs, yet generating high-confidence pseudo-labels from perturbed inputs during training remains a significant challenge. In this paper, we propose an Uncertainty-aware Cross-training framework for semi-supervised medical image Segmentation (UC-Seg). Our UC-Seg framework incorporates two distinct subnets to effectively explore and leverage the correlation between them, thereby mitigating cognitive biases within the model. Specifically, we present a Cross-subnet Consistency Preservation (CCP) strategy to enhance feature representation capability and ensure feature consistency across the two subnets. This strategy enables each subnet to correct its own biases and learn shared semantics from both labeled and unlabeled data. Additionally, we propose an Uncertainty-aware Pseudo-label Generation (UPG) component that leverages segmentation results and corresponding uncertainty maps from both subnets to generate high-confidence pseudo-labels. We extensively evaluate the proposed UC-Seg on various medical image segmentation tasks involving different modality images, such as MRI, CT, ultrasound, colonoscopy, and so on. The results demonstrate that our method achieves superior segmentation accuracy and generalization performance compared to other state-of-the-art semi-supervised methods. Our code and segmentation maps will be released at https://github.com/taozh2017/UCSeg.
Kaiwen Huang 0002, Tao Zhou 0002, Huazhu Fu, Yizhe Zhang 0001, Yi Zhou 0007, Xiaojun Wu 0001
IEEE Trans. Image Process.5
2025 Learnable Prompting SAM-Induced Knowledge Distillation for Semi-Supervised Medical Image Segmentation
abstract
The limited availability of labeled data has driven advancements in semi-supervised learning for medical image segmentation. Modern large-scale models tailored for general segmentation, such as the Segment Anything Model (SAM), have revealed robust generalization capabilities. However, applying these models directly to medical image segmentation still exposes performance degradation. In this paper, we propose a learnable prompting SAM-induced Knowledge distillation framework (KnowSAM) for semi-supervised medical image segmentation. Firstly, we propose a Multi-view Co-training (MC) strategy that employs two distinct sub-networks to employ a co-teaching paradigm, resulting in more robust outcomes. Secondly, we present a Learnable Prompt Strategy (LPS) to dynamically produce dense prompts and integrate an adapter to fine-tune SAM specifically for medical image segmentation tasks. Moreover, we propose SAM-induced Knowledge Distillation (SKD) to transfer useful knowledge from SAM to two sub-networks, enabling them to learn from SAM's predictions and alleviate the effects of incorrect pseudo-labels during training. Notably, the predictions generated by our subnets are used to produce mask prompts for SAM, facilitating effective inter-module information exchange. Extensive experimental results on various medical segmentation tasks demonstrate that our model outperforms the state-of-the-art semi-supervised segmentation approaches. Crucially, our SAM distillation framework can be seamlessly integrated into other semi-supervised segmentation methods to enhance performance. The code will be released upon acceptance of this manuscript at https://github.com/taozh2017/KnowSAM.
Kaiwen Huang 0002, Tao Zhou 0002, Huazhu Fu, Yizhe Zhang 0001, Yi Zhou 0007, Chen Gong 0002, Dong Liang 0001
IEEE Trans. Medical Imaging5
2025 Domain-Interactive Contrastive Learning and Prototype-Guided Self-Training for Cross-Domain Polyp Segmentation
abstract
Accurate polyp segmentation plays a critical role in the diagnosis and treatment of colorectal cancer from colonoscopy images. While deep learning-based polyp segmentation models have made significant progress, they often suffer from performance degradation when applied to unseen target domain datasets collected from different imaging devices. To address this challenge, unsupervised domain adaptation (UDA) methods have gained attention by leveraging labeled source data and unlabeled target data to reduce the domain gap. However, existing UDA methods primarily focus on capturing class-wise representations, neglecting domain-wise representations. Additionally, uncertainty in pseudo-labels could hinder the segmentation performance. To tackle these issues, we propose a novel Domain-interactive Contrastive Learning and Prototype-guided Self-training (DCL-PS) framework for cross-domain polyp segmentation. Specifically, domain-interactive contrastive learning (DCL) with a domain-mixed prototype updating strategy is proposed to discriminate class-wise feature representations across domains. Then, to enhance the feature extraction ability of the encoder, we present a contrastive learning-based cross-consistency training (CL-CCT) strategy, which is imposed on both the prototypes obtained by the outputs of the main decoder and perturbed auxiliary outputs. Furthermore, we propose a prototype-guided self-training (PS) strategy, which dynamically assigns a weight for each pixel during self-training, filtering out unreliable pixels and improving the quality of pseudo-labels. Experimental results demonstrate the superiority of DCL-PS in improving polyp segmentation performance in the target domain. The code is released at https://github.com/taozh2017/DCLPS.
Ziru Lu, Yizhe Zhang 0001, Yi Zhou 0007, Ye Wu 0001, Tao Zhou 0002
IEEE Trans. Medical Imaging3
2024 MedSegViG: Medical Image Segmentation with a Vision Graph Neural Network
abstract
Medical image segmentation is a crucial step toward automatic clinical diagnosis, which has received growing interest. Although some existing methods based on convolutional neural networks or transformers have achieved remarkable success in this task, they still show limitations in effectively modeling the relationships among different objects in images. In this paper, we propose a novel deep learning based model to address this issue by leveraging a vision graph neural network (ViG). Our model, MedSegViG, mainly consists of a hierarchical ViG encoder and a lightweight convolutional decoder. The hierarchical encoder extracts multi-level features from the image and captures the object relationships with graph neural networks. The lightweight decoder then fuses these features and generates the corresponding segmentation map. Extensive experiments are conducted on seven datasets for three typical medical image segmentation tasks: polyp segmentation, skin lesion segmentation, and retinal vessel segmentation. The results demonstrate the superiority of our MedSegViG over state-of-the-art models across various tasks and datasets. The code is released on https://github.com/Xinhong-Li/MedSegViG.
Geng Chen 0001, Yuanfeng Wu, Junqing Yang, Tao Zhou 0002, Yi Zhou 0007, Wentao Zhu 0002
BIBM6
2024 Memory-Assisted Sub-Prototype Mining for Universal Domain Adaptation
abstract
Universal domain adaptation aims to align the classes and reduce the feature gap between the same category of the source and target domains. The target private category is set as the unknown class during the adaptation process, as it is not included in the source domain. However, most existing methods overlook the intra-class structure within a category, especially in cases where there exists significant concept shift between the samples belonging to the same category. When samples with large concept shift are forced to be pushed together, it may negatively affect the adaptation performance. Moreover, from the interpretability aspect, it is unreasonable to align visual features with significant differences, such as fighter jets and civil aircraft, into the same category. Unfortunately, due to such semantic ambiguity and annotation cost, categories are not always classified in detail, making it difficult for the model to perform precise adaptation. To address these issues, we propose a novel Memory-Assisted Sub-Prototype Mining (MemSPM) method that can learn the differences between samples belonging to the same category and mine sub-classes when there exists significant concept shift between them. By doing so, our model learns a more reasonable feature space that enhances the transferability and reflects the inherent differences among samples annotated as the same category. We evaluate the effectiveness of our MemSPM method over multiple scenarios, including UniDA, OSDA, and PDA. Our method achieves state-of-the-art performance on four benchmarks in most cases.
Yuxiang Lai, Yi Zhou 0007, Xinghong Liu, Tao Zhou 0002
ICLR2
2024 MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise
Chenran Zhang, Jianle Zhang, Yi Zhou 0007, Tao Zhou 0002, Huazhu Fu
MICCAI (1)4
2024 SimTxtSeg: Weakly-Supervised Medical Image Segmentation with Simple Text Cues
Tao Zhou 0002, Yi Zhou 0007, Geng Chen 0001
MICCAI (8)3
2024 TextPolyp: Point-Supervised Polyp Segmentation with Text Cues
Yi Zhou 0007, Yizhe Zhang 0001, Ye Wu 0001, Tao Zhou 0002
MICCAI (11)2
2024 Uncertainty-Aware Hierarchical Aggregation Network for Medical Image Segmentation
abstract
Medical image segmentation is an essential process to assist clinics with computer-aided diagnosis and treatment. Recently, a large amount of convolutional neural network (CNN)-based methods have been rapidly developed and achieved remarkable performances in several different medical image segmentation tasks. However, the same type of infected region or lesions often has a diversity of scales, making it a challenging task to achieve accurate medical image segmentation. In this paper, we present a novel Uncertainty-aware Hierarchical Aggregation Network, namely UHA-Net, for medical image segmentation, which can fully make utilization of cross-level and multi-scale features to handle scale variations. Specifically, we propose a hierarchical feature fusion (HFF) module to aggregate high-level features, which is used to produce a global map for the coarse localization of the segmented target. Then, we propose an uncertainty-induced cross-level fusion (UCF) module to fully fuse features from the adjacent levels, which can learn knowledge guidance to capture the contextual information from adjacent resolutions. Further, a scale aggregation module (SAM) is presented to learn multi-scale features by using different convolution kernels, to effectively deal with scale variations. At last, we formulate a unified framework to simultaneously fuse inter-layer convolutional features and learn the discriminability of multi-scale representations from the intra-layer features, leading to accurate segmentation results. We carry out experiments on three different medical image segmentation tasks, and the results demonstrate that our UHA-Net outperforms state-of-the-art segmentation methods. Our implementation code and segmentation maps will be publicly at https://github.com/taozh2017/UHANet.
Tao Zhou 0002, Yi Zhou 0007, Geng Chen 0001, Jianbing Shen
IEEE Trans. Circuits Syst. Video Technol.2
2023 Specificity-preserving RGB-D saliency detection
abstract
RGB-D saliency detection has attracted increasing attention, due to its effectiveness and the fact that depth cues can now be conveniently captured. Existing works often focus on learning a shared representation through various fusion strategies, with few methods explicitly considering how to preserve modality-specific characteristics. In this paper, taking a new perspective, we propose a specificity-preserving network (SP-Net) for RGB-D saliency detection, which benefits saliency detection performance by exploring both the shared information and modality-specific properties (e.g., specificity). Specifically, two modality-specific networks and a shared learning network are adopted to generate individual and shared saliency maps. A cross-enhanced integration module (CIM) is proposed to fuse cross-modal features in the shared learning network, which are then propagated to the next layer for integrating cross-level information. Besides, we propose a multi-modal feature aggregation (MFA) module to integrate the modality-specific features from each individual decoder into the shared decoder, which can provide rich complementary multi-modal information to boost the saliency detection performance. Further, a skip connection is used to combine hierarchical features between the encoder and decoder layers. Experiments on six benchmark datasets demonstrate that our SP-Net outperforms other state-of-the-art methods. Code is available at: https://github.com/taozh2017/SPNet.
Tao Zhou 0002, Deng-Ping Fan, Geng Chen 0001, Yi Zhou 0007, Huazhu Fu
Comput. Vis. Media4
2023 Combating medical noisy labels by disentangled distribution learning and consistency regularization
Yi Zhou 0007, Lei Huang 0015, Tao Zhou 0002, Hanshi Sun
Future Gener. Comput. Syst.1
2023 Normalization Techniques in Training DNNs: Methodology, Analysis and Application
abstract
Normalization techniques are essential for accelerating the training and improving the generalization of deep neural networks (DNNs), and have successfully been used in various applications. This paper reviews and comments on the past, present and future of normalization methods in the context of DNN training. We provide a unified picture of the main motivation behind different approaches from the perspective of optimization, and present a taxonomy for understanding the similarities and differences between them. Specifically, we decompose the pipeline of the most representative normalizing activation methods into three components: the normalization area partitioning, normalization operation and normalization representation recovery. In doing so, we provide insight for designing new normalization technique. Finally, we discuss the current progress in understanding normalization methods, and provide a comprehensive review of the applications of normalization for particular tasks, in which it can effectively solve the key issues.
Lei Huang 0015, Jie Qin 0004, Yi Zhou 0007, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Cross-level Feature Aggregation Network for Polyp Segmentation
Tao Zhou 0002, Yi Zhou 0007, Kelei He, Chen Gong 0002, Jian Yang 0003, Huazhu Fu, Dinggang Shen
Pattern Recognit.2
2023 Flexible Fusion Network for Multi-Modal Brain Tumor Segmentation
abstract
Automated brain tumor segmentation is crucial for aiding brain disease diagnosis and evaluating disease progress. Currently, magnetic resonance imaging (MRI) is a routinely adopted approach in the field of brain tumor segmentation that can provide different modality images. It is critical to leverage multi-modal images to boost brain tumor segmentation performance. Existing works commonly concentrate on generating a shared representation by fusing multi-modal data, while few methods take into account modality-specific characteristics. Besides, how to efficiently fuse arbitrary numbers of modalities is still a difficult task. In this study, we present a flexible fusion network (termed F$^{2}$Net) for multi-modal brain tumor segmentation, which can flexibly fuse arbitrary numbers of multi-modal information to explore complementary information while maintaining the specific characteristics of each modality. Our F$^{2}$Net is based on the encoder-decoder structure, which utilizes two Transformer-based feature learning streams and a cross-modal shared learning network to extract individual and shared feature representations. To effectively integrate the knowledge from the multi-modality data, we propose a cross-modal feature-enhanced module (CFM) and a multi-modal collaboration module (MCM), which aims at fusing the multi-modal features into the shared learning network and incorporating the features from encoders into the shared decoder, respectively. Extensive experimental results on multiple benchmark datasets demonstrate the effectiveness of our F$^{2}$Net over other state-of-the-art segmentation methods.
Hengyi Yang, Tao Zhou 0002, Yi Zhou 0007, Yizhe Zhang 0001, Huazhu Fu
IEEE J. Biomed. Health Informatics3
2023 Blind Super-Resolution of 3D MRI via Unsupervised Domain Transformation
abstract
High-resolution medical images can be effectively used for clinical diagnosis. However, the acquisition of high-resolution images is difficult and often limited by medical instruments. Super-resolution (SR) methods provide a solution, where high-resolution (HR) images can be reconstructed from low-resolution (LR) ones. Most of existing deep neural networks for 3D SR medical images trained in a non-blind process, where LR images are directly degraded from HR data via a pre-determined downscale method. Such approaches rely heavily on the assumed degradation model, resulting in inevitable deviations in real clinical practice. Blind super-resolution, as a more attractive research line for this field, aims to generate HR images from LR inputs containing unknown degradation. Towards generalizing SR models for diverse types of degradation, we propose a robust blind SR of 3D medical images in an unsupervised manner with domain correction and upscaling treatment. First, a CycleGAN-based architecture is implemented to generate the LR data from the source domain to the target one for domain correction. Then, an upscaling network is learned via pre-determined HR-LR couples for reconstruction. The proposed framework is able to automatically learn noisy and blurry correction kernels for unpaired 3D SR magnetic resonance images (MRI). Our method achieves better and more robust performances in reconstruction of HR images from LR MRI with multiple unknown degradation processes, and show its superiority to other state-of-the-art supervised models and cycle-consistency based methods, especially in severe distortion cases.
Hexiang Zhou, Yawen Huang, Yuexiang Li, Yi Zhou 0007, Yefeng Zheng 0001
IEEE J. Biomed. Health Informatics4
2023 Multi-Scale Transformer Network With Edge-Aware Pre-Training for Cross-Modality MR Image Synthesis
abstract
Cross-modality magnetic resonance (MR) image synthesis can be used to generate missing modalities from given ones. Existing (supervised learning) methods often require a large number of paired multi-modal data to train an effective synthesis model. However, it is often challenging to obtain sufficient paired data for supervised training. In reality, we often have a small number of paired data while a large number of unpaired data. To take advantage of both paired and unpaired data, in this paper, we propose a Multi-scale Transformer Network (MT-Net) with edge-aware pre-training for cross-modality MR image synthesis. Specifically, an Edge-preserving Masked AutoEncoder (Edge-MAE) is first pre-trained in a self-supervised manner to simultaneously perform 1) image imputation for randomly masked patches in each image and 2) whole edge map estimation, which effectively learns both contextual and structural information. Besides, a novel patch-wise loss is proposed to enhance the performance of Edge-MAE by treating different masked patches differently according to the difficulties of their respective imputations. Based on this proposed pre-training, in the subsequent fine-tuning stage, a Dual-scale Selective Fusion (DSF) module is designed (in our MT-Net) to synthesize missing-modality images by integrating multi-scale features extracted from the encoder of the pre-trained Edge-MAE. Furthermore, this pre-trained encoder is also employed to extract high-level features from the synthesized image and corresponding ground-truth image, which are required to be similar (consistent) in the training. Experimental results show that our MT-Net achieves comparable performance to the competing methods even using 70% of all available paired data. Our code will be released at https://github.com/lyhkevin/MT-Net.
Yonghao Li, Tao Zhou 0002, Kelei He, Yi Zhou 0007, Dinggang Shen
IEEE Trans. Medical Imaging4
2022 Delving into the Estimation Shift of Batch Normalization in a Network
abstract
Batch normalization (BN) is a milestone technique in deep learning. It normalizes the activation using mini-batch statistics during training but the estimated population statistics during inference. This paper focuses on investigating the estimation of population statistics. We define the estimation shift magnitude of BN to quantitatively measure the difference between its estimated population statistics and expected ones. Our primary observation is that the estimation shift can be accumulated due to the stack of BN in a network, which has detriment effects for the test performance. We further find a batch-free normalization (BFN) can block such an accumulation of estimation shift. These observations motivate our design of XBNBlock that replace one BN with BFN in the bottleneck block of residual-style networks. Experiments on the ImageNet and COCO benchmarks show that XBNBlock consistently improves the performance of different architectures, including ResNet and ResNeXt, by a significant margin and seems to be more robust to distribution shift.
Lei Huang 0015, Yi Zhou 0007, Tian Wang 0002, Jie Luo 0004, Xianglong Liu 0001
CVPR2
2022 Delving into Local Features for Open-Set Domain Adaptation in Fundus Image Analysis
Yi Zhou 0007, Shaochen Bai, Tao Zhou 0002, Yu Zhang 0009, Huazhu Fu
MICCAI (8)1
2022 Feature Aggregation and Propagation Network for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) aims to detect/segment camouflaged objects embedded in the environment, which has attracted increasing attention over the past decades. Although several COD methods have been developed, they still suffer from unsatisfactory performance due to the intrinsic similarities between the foreground objects and background surroundings. In this paper, we propose a novel Feature Aggregation and Propagation Network (FAP-Net) for camouflaged object detection. Specifically, we propose a Boundary Guidance Module (BGM) to explicitly model the boundary characteristic, which can provide boundary-enhanced features to boost the COD performance. To capture the scale variations of the camouflaged objects, we propose a Multi-scale Feature Aggregation Module (MFAM) to characterize the multi-scale information from each layer and obtain the aggregated feature representations. Furthermore, we propose a Cross-level Fusion and Propagation Module (CFPM). In the CFPM, the feature fusion part can effectively integrate the features from adjacent layers to exploit the cross-level correlations, and the feature propagation part can transmit valuable context information from the encoder to the decoder network via a gate unit. Finally, we formulate a unified and end-to-end trainable framework where cross-level features can be effectively fused and propagated for capturing rich context information. Extensive experiments on three benchmark camouflaged datasets demonstrate that our FAP-Net outperforms other state-of-the-art COD models. Moreover, our model can be extended to the polyp segmentation task, and the comparison results further validate the effectiveness of the proposed model in segmenting polyps. The source code and results will be released at https://github.com/taozh2017/FAPNet.
Tao Zhou 0002, Yi Zhou 0007, Chen Gong 0002, Jian Yang 0003, Yu Zhang 0009
IEEE Trans. Image Process.2
2022 DR-GAN: Conditional Generative Adversarial Network for Fine-Grained Lesion Synthesis on Diabetic Retinopathy Images
abstract
Diabetic retinopathy (DR) is a complication of diabetes that severely affects eyes. It can be graded into five levels of severity according to international protocol. However, optimizing a grading model to have strong generalizability requires a large amount of balanced training data, which is difficult to collect, particularly for the high severity levels. Typical data augmentation methods, including random flipping and rotation, cannot generate data with high diversity. In this paper, we propose a diabetic retinopathy generative adversarial network (DR-GAN) to synthesize high-resolution fundus images which can be manipulated with arbitrary grading and lesion information. Thus, large-scale generated data can be used for more meaningful augmentation to train a DR grading and lesion segmentation model. The proposed retina generator is conditioned on the structural and lesion masks, as well as adaptive grading vectors sampled from the latent grading space, which can be adopted to control the synthesized grading severity. Moreover, a multi-scale spatial and channel attention module is devised to improve the generation ability to synthesize small details. Multi-scale discriminators are designed to operate from large to small receptive fields, and joint adversarial losses are adopted to optimize the whole network in an end-to-end manner. With extensive experiments evaluated on the EyePACS dataset connected to Kaggle, as well as the FGADR dataset, we validate the effectiveness of our method, which can both synthesize highly realistic ( 1280 ×1280) controllable fundus images and contribute to the DR grading task.
Yi Zhou 0007, Xiaodong He 0004, Shanshan Cui, Ling Shao 0001
IEEE J. Biomed. Health Informatics1
2021 Group-Wise Semantic Mining for Weakly Supervised Semantic Segmentation
abstract
Acquiring sufficient ground-truth supervision to train deep vi- sual models has been a bottleneck over the years due to the data-hungry nature of deep learning. This is exacerbated in some structured prediction tasks, such as semantic segmen- tation, which requires pixel-level annotations. This work ad- dresses weakly supervised semantic segmentation (WSSS), with the goal of bridging the gap between image-level anno- tations and pixel-level segmentation. We formulate WSSS as a novel group-wise learning task that explicitly models se- mantic dependencies in a group of images to estimate more reliable pseudo ground-truths, which can be used for training more accurate segmentation models. In particular, we devise a graph neural network (GNN) for group-wise semantic min- ing, wherein input images are represented as graph nodes, and the underlying relations between a pair of images are char- acterized by an efficient co-attention mechanism. Moreover, in order to prevent the model from paying excessive atten- tion to common semantics only, we further propose a graph dropout layer, encouraging the model to learn more accurate and complete object responses. The whole network is end-to- end trainable by iterative message passing, which propagates interaction cues over the images to progressively improve the performance. We conduct experiments on the popular PAS- CAL VOC 2012 and COCO benchmarks, and our model yields state-of-the-art performance. Our code is available at: https://github.com/Lixy1997/Group-WSSS.
Xueyi Li 0006, Tianfei Zhou, Jianwu Li, Yi Zhou 0007, Zhaoxiang Zhang 0001
AAAI4
2021 Many-to-One Distribution Learning and K-Nearest Neighbor Smoothing for Thoracic Disease Identification
abstract
Chest X-rays are an important and accessible clinical imaging tool for the detection of many thoracic diseases. Over the past decade, deep learning, with a focus on the convolutional neural network (CNN), has become the most powerful computer-aided diagnosis technology for improving disease identification performance. However, training an effective and robust deep CNN usually requires a large amount of data with high annotation quality. For chest X-ray imaging, annotating large-scale data requires professional domain knowledge and is time-consuming. Thus, existing public chest X-ray datasets usually adopt language pattern based methods to automatically mine labels from reports. However, this results in label uncertainty and inconsistency. In this paper, we propose many-to-one distribution learning (MODL) and K-nearest neighbor smoothing (KNNS) methods from two perspectives to improve a single model's disease identification performance, rather than focusing on an ensemble of models. MODL integrates multiple models to obtain a soft label distribution for optimizing the single target model, which can reduce the effects of original label uncertainty. Moreover, KNNS aims to enhance the robustness of the target model to provide consistent predictions on images with similar medical findings. Extensive experiments on the public NIH Chest X-ray and CheXpert datasets show that our model achieves consistent improvements over the state-of-the-art methods.
Yi Zhou 0007, Lei Huang 0015, Tianfei Zhou, Ling Shao 0001
AAAI1
2021 Group Whitening: Balancing Learning Efficiency and Representational Capacity
abstract
Batch normalization (BN) is an important technique commonly incorporated into deep learning models to perform standardization within mini-batches. The merits of BN in improving a model’s learning efficiency can be further amplified by applying whitening, while its drawbacks in estimating population statistics for inference can be avoided through group normalization (GN). This paper proposes group whitening (GW), which exploits the advantages of the whitening operation and avoids the disadvantages of normalization within mini-batches. In addition, we analyze the constraints imposed on features by normalization, and show how the batch size (group number) affects the performance of batch (group) normalized networks, from the perspective of model’s representational capacity. This analysis provides theoretical guidance for applying GW in practice. Finally, we apply the proposed GW to ResNet and ResNeXt architectures and conduct experiments on the ImageNet and COCO benchmarks. Results show that GW consistently improves the performance of different architectures, with absolute gains of 1.02% ∼ 1.49% in top-1 accuracy on ImageNet and 1.82% ∼ 3.21% in bounding box AP on COCO.
Lei Huang 0015, Yi Zhou 0007, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001
CVPR2
2021 Specificity-preserving RGB-D Saliency Detection
Tao Zhou 0002, Huazhu Fu, Geng Chen 0001, Yi Zhou 0007, Deng-Ping Fan, Ling Shao 0001
ICCV4
2021 CCT-Net: Category-Invariant Cross-Domain Transfer for Medical Single-to-Multiple Disease Diagnosis
abstract
A medical imaging model is usually explored for the diagnosis of a single disease. However, with the expanding demand for multi-disease diagnosis in clinical applications, multi-function solutions need to be investigated. Previous works proposed to either exploit different disease labels to conduct transfer learning through fine-tuning, or transfer knowledge across different domains with similar diseases. However, these methods still cannot address the real clinical challenge - a multi-disease model is required but annotations for each disease are not always available. In this paper, we introduce the task of transferring knowledge from single-disease diagnosis (source domain) to enhance multi-disease diagnosis (target domain). A category-invariant cross-domain transfer (CCT) method is proposed to address this single-to-multiple extension. First, for domain-specific task learning, we present a confidence weighted pooling (CWP) to obtain coarse heatmaps for different disease categories. Then, conditioned on these heatmaps, category-invariant feature refinement (CIFR) blocks are proposed to better localize discriminative semantic regions related to the corresponding diseases. The category-invariant characteristic enables transferability from the source domain to the target domain. We validate our method in two popular areas: extending diabetic retinopathy to identifying multiple ocular diseases, and extending glioma identification to the diagnosis of other brain tumors.
Yi Zhou 0007, Lei Huang 0015, Tao Zhou 0002, Ling Shao 0001
ICCV1
2021 Visual-Textual Attentive Semantic Consistency for Medical Report Generation
abstract
Automatic report generation on medical radiographs have recently gained interest. However, identifying diseases as well as correctly predicting their corresponding sizes, locations and other medical description patterns, which is essential for generating high-quality reports, is challenging. Although previous methods focused on producing readable reports, how to accurately detect and describe findings that match with the query X-Ray has not been successfully addressed. In this paper, we propose a multi-modality semantic attention model to integrate visual features, predicted key finding embeddings, as well as clinical features, and progressively decode reports with visual-textual semantic consistency. First, multi-modality features are extracted and attended with the hidden states from the sentence de-coder, to encode enriched context vectors for better decoding a report. These modalities include regional visual features of scans, semantic word embeddings of the top-K findings predicted with high probabilities, and clinical features of indications. Second, the progressive report decoder consists of a sentence decoder and a word decoder, where we propose image-sentence matching and description accuracy losses to constrain the visual-textual semantic consistency. Extensive experiments on the public MIMIC-CXR and IU X-Ray datasets show that our model achieves consistent improvements over the state-of-the-art methods.
Yi Zhou 0007, Lei Huang 0015, Tao Zhou 0002, Huazhu Fu, Ling Shao 0001
ICCV1
2021 A Benchmark for Studying Diabetic Retinopathy: Segmentation, Grading, and Transferability
abstract
People with diabetes are at risk of developing an eye disease called diabetic retinopathy (DR). This disease occurs when high blood glucose levels cause damage to blood vessels in the retina. Computer-aided DR diagnosis has become a promising tool for the early detection and severity grading of DR, due to the great success of deep learning. However, most current DR diagnosis systems do not achieve satisfactory performance or interpretability for ophthalmologists, due to the lack of training data with consistent and fine-grained annotations. To address this problem, we construct a large fine-grained annotated DR dataset containing 2,842 images (FGADR). Specifically, this dataset has 1,842 images with pixel-level DR-related lesion annotations, and 1,000 images with image-level labels graded by six board-certified ophthalmologists with intra-rater consistency. The proposed dataset will enable extensive studies on DR diagnosis. Further, we establish three benchmark tasks for evaluation: 1. DR lesion segmentation; 2. DR grading by joint classification and segmentation; 3. Transfer learning for ocular multi-disease identification. Moreover, a novel inductive transfer learning method is introduced for the third task. Extensive experiments using different state-of-the-art methods are conducted on our FGADR dataset, which can serve as baselines for future research. Our dataset will be released in https://csyizhou.github.io/FGADR/.
Yi Zhou 0007, Lei Huang 0015, Shanshan Cui, Ling Shao 0001
IEEE Trans. Medical Imaging1
2021 Contrast-Attentive Thoracic Disease Recognition With Dual-Weighting Graph Reasoning
abstract
Automatic thoracic disease diagnosis is a rising research topic in the medical imaging community, with many potential applications. However, the inconsistent appearances and high complexities of various lesions in chest X-rays currently hinder the development of a reliable and robust intelligent diagnosis system. Attending to the high-probability abnormal regions and exploiting the priori of a related knowledge graph offers one promising route to addressing these issues. As such, in this paper, we propose two contrastive abnormal attention models and a dual-weighting graph convolution to improve the performance of thoracic multi-disease recognition. First, a left-right lung contrastive network is designed to learn intra-attentive abnormal features to better identify the most common thoracic diseases, whose lesions rarely appear in both sides symmetrically. Moreover, an inter-contrastive abnormal attention model aims to compare the query scan with multiple anchor scans without lesions to compute the abnormal attention map. Once the intra- and inter-contrastive attentions are weighted over the features, in addition to the basic visual spatial convolution, a chest radiology graph is constructed for dual-weighting graph reasoning. Extensive experiments on the public NIH ChestX-ray and CheXpert datasets show that our model achieves consistent improvements over the state-of-the-art methods both on thoracic disease identification and localization.
Yi Zhou 0007, Tianfei Zhou, Tao Zhou 0002, Huazhu Fu, Jiacheng Liu 0011, Ling Shao 0001
IEEE Trans. Medical Imaging1
2020 Motion-Attentive Transition for Zero-Shot Video Object Segmentation
abstract
In this paper, we present a novel Motion-Attentive Transition Network (MATNet) for zero-shot video object segmentation, which provides a new way of leveraging motion information to reinforce spatio-temporal object representation. An asymmetric attention block, called Motion-Attentive Transition (MAT), is designed within a two-stream encoder, which transforms appearance features into motion-attentive representations at each convolutional stage. In this way, the encoder becomes deeply interleaved, allowing for closely hierarchical interactions between object motion and appearance. This is superior to the typical two-stream architecture, which treats motion and appearance separately in each stream and often suffers from overfitting to appearance information. Additionally, a bridge network is proposed to obtain a compact, discriminative and scale-sensitive representation for multi-level encoder features, which is further fed into a decoder to achieve segmentation results. Extensive experiments on three challenging public benchmarks (i.e., DAVIS-16, FBMS and Youtube-Objects) show that our model achieves compelling performance against the state-of-the-arts. Code is available at: https://github.com/tfzhou/MATNet.
Tianfei Zhou, Shunzhou Wang, Yi Zhou 0007, Yazhou Yao, Jianwu Li, Ling Shao 0001
AAAI3
2020 An Investigation Into the Stochasticity of Batch Whitening
abstract
Batch Normalization (BN) is extensively employed in various network architectures by performing standardization within mini-batches. A full understanding of the process has been a central target in the deep learning communities. Unlike existing works, which usually only analyze the standardization operation, this paper investigates the more general Batch Whitening (BW). Our work originates from the observation that while various whitening transformations equivalently improve the conditioning, they show significantly different behaviors in discriminative scenarios and training Generative Adversarial Networks (GANs). We attribute this phenomenon to the stochasticity that BW introduces. We quantitatively investigate the stochasticity of different whitening transformations and show that it correlates well with the optimization behaviors during training. We also investigate how stochasticity relates to the estimation of population statistics during inference. Based on our analysis, we provide a framework for designing and comparing BW algorithms in different scenarios. Our proposed BW algorithm improves the residual networks by a significant margin on ImageNet classification. Besides, we show that the stochasticity of BW can improve the GAN's performance with, however, the sacrifice of the training stability.
Lei Huang 0015, Yi Zhou 0007, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001
CVPR3
2020 Inf-Net: Automatic COVID-19 Lung Infection Segmentation From CT Images
abstract
Coronavirus Disease 2019 (COVID-19) spread globally in early 2020, causing the world to face an existential health crisis. Automated detection of lung infections from computed tomography (CT) images offers a great potential to augment the traditional healthcare strategy for tackling COVID-19. However, segmenting infected regions from CT slices faces several challenges, including high variation in infection characteristics, and low intensity contrast between infections and normal tissues. Further, collecting a large amount of data is impractical within a short time period, inhibiting the training of a deep model. To address these challenges, a novel COVID-19 Lung Infection Segmentation Deep Network (Inf-Net) is proposed to automatically identify infected regions from chest CT slices. In our Inf-Net, a parallel partial decoder is used to aggregate the high-level features and generate a global map. Then, the implicit reverse attention and explicit edge-attention are utilized to model the boundaries and enhance the representations. Moreover, to alleviate the shortage of labeled data, we present a semi-supervised segmentation framework based on a randomly selected propagation strategy, which only requires a few labeled images and leverages primarily unlabeled data. Our semi-supervised framework can improve the learning ability and achieve a higher performance. Extensive experiments on our COVID-SemiSeg and real CT volumes demonstrate that the proposed Inf-Net outperforms most cutting-edge segmentation models and advances the state-of-the-art performance.
Deng-Ping Fan, Tao Zhou 0002, Ge-Peng Ji, Yi Zhou 0007, Geng Chen 0001, Huazhu Fu, Jianbing Shen, Ling Shao 0001
IEEE Trans. Medical Imaging4
2019 Iterative Normalization: Beyond Standardization Towards Efficient Whitening
abstract
Batch Normalization (BN) is ubiquitously employed for accelerating neural network training and improving the generalization capability by performing standardization within mini-batches. Decorrelated Batch Normalization (DBN) further boosts the above effectiveness by whitening. However, DBN relies heavily on either a large batch size, or eigen-decomposition that suffers from poor efficiency on GPUs. We propose Iterative Normalization (IterNorm), which employs Newton’s iterations for much more efficient whitening, while simultaneously avoiding the eigen-decomposition. Furthermore, we develop a comprehensive study to show IterNorm has better trade-off between optimization and generalization, with theoretical and experimental support. To this end, we exclusively introduce Stochastic Normalization Disturbance (SND), which measures the inherent stochastic uncertainty of samples when applied to normalization operations. With the support of SND, we provide natural explanations to several phenomena from the perspective of optimization, e.g., why group-wise whitening of DBN generally outperforms full-whitening and why the accuracy of BN degenerates with reduced batch sizes. We demonstrate the consistently improved performance of IterNorm with extensive experiments on CIFAR-10 and ImageNet over BN and DBN.
Lei Huang 0015, Yi Zhou 0007, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001
CVPR2
2019 Building Detail-Sensitive Semantic Segmentation Networks With Polynomial Pooling
abstract
Semantic segmentation is an important computer vision task, which aims to allocate a semantic label to each pixel in an image. When training a segmentation model, it is common to fine-tune a classification network pre-trained on a large-scale dataset. However, as an intrinsic property of the classification model, invariance to spatial perturbation resulting from the lose of detail-sensitivity prevents segmentation networks from achieving high performance. The use of standard poolings is one of the key factors for this invariance. The most common standard poolings are max and average pooling. Max pooling can increase both the invariance to spatial perturbations and the non-linearity of the networks. Average pooling, on the other hand, is sensitive to spatial perturbations, but is a linear function. For semantic segmentation, we prefer both the preservation of detailed cues within a local feature region and non-linearity that increases a network's functional complexity. In this work, we propose a polynomial pooling (P-pooling) function that finds an intermediate form between max and average pooling to provide an optimally balanced and self-adjusted pooling strategy for semantic segmentation. The P-pooling is differentiable and can be applied into a variety of pre-trained networks. Extensive studies on the PASCAL VOC, Cityscapes and ADE20k datasets demonstrate the superiority of P-pooling over other poolings. Experiments on various network architectures and state-of-the-art training strategies also show that models with P-pooling layers consistently outperform those directly fine-tuned using pre-trained classification models.
Zhen Wei 0001, Jingyi Zhang 0005, Li Liu 0004, Fan Zhu 0001, Fumin Shen, Yi Zhou 0007, Si Liu 0001, Yao Sun 0004, Ling Shao 0001
CVPR6
2019 Collaborative Learning of Semi-Supervised Segmentation and Classification for Medical Images
abstract
Medical image analysis has two important research areas: disease grading and fine-grained lesion segmentation. Although the former problem often relies on the latter, the two are usually studied separately. Disease severity grading can be treated as a classification problem, which only requires image-level annotations, while the lesion segmentation requires stronger pixel-level annotations. However, pixel-wise data annotation for medical images is highly time-consuming and requires domain experts. In this paper, we propose a collaborative learning method to jointly improve the performance of disease grading and lesion segmentation by semi-supervised learning with an attention mechanism. Given a small set of pixel-level annotated data, a multi-lesion mask generation model first performs the traditional semantic segmentation task. Then, based on initially predicted lesion maps for large quantities of image-level annotated data, a lesion attentive disease grading model is designed to improve the severity classification accuracy. Meanwhile, the lesion attention model can refine the lesion maps using class-specific information to fine-tune the segmentation model in a semi-supervised manner. An adversarial architecture is also integrated for training. With extensive experiments on a representative medical problem called diabetic retinopathy (DR), we validate the effectiveness of our method and achieve consistent improvements over state-of-the-art methods on three public datasets.
Yi Zhou 0007, Xiaodong He 0004, Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Shanshan Cui, Ling Shao 0001
CVPR1
2019 DME-Net: Diabetic Macular Edema Grading by Auxiliary Task Learning
Xiaodong He 0004, Yi Zhou 0007, Shanshan Cui, Ling Shao 0001
MICCAI (1)2
2019 High-Resolution Diabetic Retinopathy Image Synthesis Manipulated by Grading and Lesions
Yi Zhou 0007, Xiaodong He 0004, Shanshan Cui, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001
MICCAI (1)1
2018 Viewpoint-Aware Attentive Multi-View Inference for Vehicle Re-Identification
abstract
Vehicle re-identification (re-ID) has the huge potential to contribute to the intelligent video surveillance. However, it suffers from challenges that different vehicle identities with a similar appearance have little inter-instance discrepancy while one vehicle usually has large intra-instance differences under viewpoint and illumination variations. Previous methods address vehicle re-ID by simply using visual features from originally captured views and usually exploit the spatial-temporal information of the vehicles to refine the results. In this paper, we propose a Viewpoint-aware Attentive Multi-view Inference (VAMI) model that only requires visual information to solve the multi-view vehicle reID problem. Given vehicle images of arbitrary viewpoints, the VAMI extracts the single-view feature for each input image and aims to transform the features into a global multiview feature representation so that pairwise distance metric learning can be better optimized in such a viewpointinvariant feature space. The VAMI adopts a viewpoint-aware attention model to select core regions at different viewpoints and implement effective multi-view feature inference by an adversarial training architecture. Extensive experiments validate the effectiveness of each proposed component and illustrate that our approach achieves consistent improvements over state-of-the-art vehicle re-ID methods on two public datasets: VeRi and VehicleID.
Yi Zhou 0007, Ling Shao 0001
CVPR1
2018 Vehicle Re-Identification by Adversarial Bi-Directional LSTM Network
abstract
Vehicle re-identification (re-ID) is an area that has received far less attention in the computer vision community than the prevalent person re-ID. Possible reasons for this slow progress are the lack of appropriate research data and the special 3D structure of a vehicle. Previous works have generally focused on limited views (e.g. front and rear), but these methods are less effective in realistic scenarios where vehicles usually appear in arbitrary views to cameras. In this paper, we focus on the uncertainty of vehicle viewpoint in re-ID, proposing an Adversarial Bi-directional LSTM Network (ABLN). Our model exploits the great advantages of the Long Short-Term Memory (LSTM) to model transformations across continuous view variations of a vehicle and adopts the adversarial architecture to enhance training. Thus, a global vehicle representation containing all views' information can be inferred from only one visible view, and then used for learning to measure the distance between two vehicles with arbitrary views. To verify our model, we evaluate the proposed method on the public VehicleID and VeRi datasets. Experimental results illustrate that our approach achieves consistent improvements over state-of-the-art vehicle re-ID methods.
Yi Zhou 0007, Ling Shao 0001
WACV1
2018 Deep Action Parsing in Videos With Large-Scale Synthesized Data
abstract
Action parsing in videos with complex scenes is an interesting but challenging task in computer vision. In this paper, we propose a generic 3D convolutional neural network in a multi-task learning manner for effective Deep Action Parsing (DAP3D-Net) in videos. Particularly, in the training phase, action localization, classification, and attributes learning can be jointly optimized on our appearance-motion data via DAP3D-Net. For an upcoming test video, we can describe each individual action in the video simultaneously as: Where the action occurs, What the action is, and How the action is performed. To well demonstrate the effectiveness of the proposed DAP3D-Net, we also contribute a new Numerous-category Aligned Synthetic Action data set, i.e., NASA, which consists of 200 000 action clips of over 300 categories and with 33 pre-defined action attributes in two hierarchical levels (i.e., low-level attributes of basic body part movements and high-level attributes related to action motion). We learn DAP3D-Net using the NASA data set and then evaluate it on our collected Human Action Understanding data set and the public THUMOS data set. Experimental results show that our approach can accurately localize, categorize, and describe multiple actions in realistic videos.
Li Liu 0004, Yi Zhou 0007, Ling Shao 0001
IEEE Trans. Image Process.2
2018 Vehicle Re-Identification by Deep Hidden Multi-View Inference
abstract
Vehicle re-identification (re-ID) is an area that has received far less attention in the computer vision community than the prevalent person re-ID. Possible reasons for this slow progress are the lack of appropriate research data and the special 3D structure of a vehicle. Previous works have generally focused on some specific views (e.g., front); but, these methods are less effective in realistic scenarios, where vehicles usually appear in arbitrary views to cameras. In this paper, we focus on the uncertainty of vehicle viewpoint in re-ID, proposing two end-to-end deep architectures: the Spatially Concatenated ConvNet and convolutional neural network (CNN)-LSTM bi-directional loop. Our models exploit the great advantages of the CNN and long short-term memory (LSTM) to learn transformations across different viewpoints of vehicles. Thus, a multi-view vehicle representation containing all viewpoints' information can be inferred from the only one input view, and then used for learning to measure distance. To verify our models, we also introduce a Toy Car RE-ID data set with images from multiple viewpoints of 200 vehicles. We evaluate our proposed methods on the Toy Car RE-ID data set and the public Multi-View Car, VehicleID, and VeRi data sets. Experimental results illustrate that our models achieve consistent improvements over the state-of-the-art vehicle re-ID approaches.
Yi Zhou 0007, Li Liu 0004, Ling Shao 0001
IEEE Trans. Image Process.1
2018 Fast Automatic Vehicle Annotation for Urban Traffic Surveillance
abstract
Automatic vehicle detection and annotation for streaming video data with complex scenes is an interesting but challenging task for intelligent transportation systems. In this paper, we present a fast algorithm: detection and annotation for vehicles (DAVE), which effectively combines vehicle detection and attributes annotation into a unified framework. DAVE consists of two convolutional neural networks: a shallow fully convolutional fast vehicle proposal network (FVPN) for extracting all vehicles' positions, and a deep attributes learning network (ALN), which aims to verify each detection candidate and infer each vehicle's pose, color, and type information simultaneously. These two nets are jointly optimized so that abundant latent knowledge learned from the deep empirical ALN can be exploited to guide training the much simpler FVPN. Once the system is trained, DAVE can achieve efficient vehicle detection and attributes annotation for real-world traffic surveillance data, while the FVPN can be independently adopted as a real-time high-performance vehicle detector as well. We evaluate the DAVE on a new self-collected urban traffic surveillance data set and the public PASCAL VOC2007 car and LISA 2010 data sets, with consistent improvements over existing algorithms.
Yi Zhou 0007, Li Liu 0004, Ling Shao 0001, Matt Mellor
IEEE Trans. Intell. Transp. Syst.1
2017 Cross-View GAN Based Vehicle Generation for Re-identification
Yi Zhou 0007, Ling Shao 0001
BMVC1
2017 DAP3D-Net: Where, what and how actions occur in videos?
abstract
Action parsing in videos with complex scenes is an interesting but challenging task in computer vision. In this paper, we propose a novel deep model based on 3D CNN (convolutional neural network) and LSTM (long short-term memory) module with a multi-task learning manner for effective Deep Action Parsing (DAP3D-Net) in videos. Particularly in the training phase, each action clip, sliced to several short consecutive segments, is fed into 3D CNN followed by LSTM to model the whole action dynamic information, so that action localization, classification and attributes learning can be jointly optimized via our deep model. Once the DAP3D-Net is trained, for an upcoming test video, we can describe each individual action in the video simultaneously as: Where the action occurs; What the action is and How the action is performed. To well demonstrate the effectiveness of the proposed DAP3D-Net, we also contribute a new Numerous-category Aligned Synthetic Action dataset, i.e., NASA, which consists of 200,000 action clips of 300 categories and with 33 pre-defined action attributes in two hierarchical levels (i.e., low-level attributes of basic body part movements and high-level attributes related to action motion). We learn DAP3D-Net using the NASA dataset and then evaluate it on our collected Human Action Understanding (HAU) dataset and the public THUMOS dataset. Experimental results show that our approach can accurately localize, categorize and describe multiple actions in realistic videos.
Li Liu 0004, Yi Zhou 0007, Ling Shao 0001
ICRA2
2016 DAVE: A Unified Framework for Fast Vehicle Detection and Annotation
Yi Zhou 0007, Li Liu 0004, Ling Shao 0001, Matt Mellor
ECCV (2)1