EDBT 2026 Demo / reviewers in the wild / expert
Luping Zhou
dblp:45/933
· DBLP profile ↗
145ranked-venue papers
13as first author
94since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 82 · 10 first-author · 54 since 2021Artificial intelligence and machine learning · 72 · 6 first-author · 47 since 2021Applied, interdisciplinary, general and emerging computing · 49 · 6 first-author · 31 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EndoIR: Degradation-Agnostic All-in-One Endoscopic Image Restoration via Noise-Aware Routing DiffusionabstractEndoscopic images often suffer from diverse and co-occurring degradations such as low lighting, smoke, and bleeding, which obscure critical clinical details. Existing restoration methods are typically task-specific and often require prior knowledge of the degradation type, limiting their robustness in real-world clinical use. We propose EndoIR, an all-in-one, degradation-agnostic diffusion-based framework that restores multiple degradation types using a single model. EndoIR introduces a Dual-Domain Prompter that extracts joint spatial–frequency features, coupled with an adaptive embedding that encodes both shared and task-specific cues as conditioning for denoising. To mitigate feature confusion in conventional concatenation-based conditioning, we design a Dual-Stream Diffusion architecture that processes clean and degraded inputs separately, with a Rectified Fusion Block integrating them in a structured, degradation-aware manner. Furthermore, Noise-Aware Routing Block improves efficiency by dynamically selecting only noise-relevant features during denoising. Experiments on SegSTRONG-C and CEC datasets demonstrate that EndoIR achieves state-of-the-art performance across multiple degradation scenarios while using fewer parameters than strong baselines, and downstream segmentation experiments confirm its clinical utility. Tong Chen 0011, Long Bai 0008, Luping Zhou |
AAAI | 6 |
| 2026 | ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report GenerationabstractAutomated radiology report generation (R2Gen) has advanced significantly, yet evaluation remains challenging due to the complexity of assessing report quality. Traditional metrics often misalign with human judgments, failing to identify specific deficiencies. To address this, we introduce ReFINE, a framework for training an Evaluation Model using a novel margin-based reward enforcement loss. This approach decomposes report quality into fine-grained sub-scores across user-defined criteria, improving interpretability. Leveraging GPT-4, we generate diverse training data with paired accepted and rejected reports to train our model under a reward-based system. The trained ReFINE Score provides both granular sub-scores and an aggregated quality assessment, enabling criterion-specific evaluation. Experimental results demonstrate ReFINE's superior alignment with human judgments, outperforming traditional metrics in model selection. Its robustness is validated across three expert-annotated datasets—including chest X-rays and multimodal reports covering 9 imaging modalities—and under two distinct scoring systems. Yunyi Liu, Yingshu Li 0001, Zhanyu Wang, Lingqiao Liu, Lei Wang 0001, Luping Zhou |
AAAI | 7 |
| 2026 | A renaissance of explicit motion information mining from transformers for action recognition
Peiqin Zhuang, Lei Bai 0001, Yichao Wu, Ding Liang, Luping Zhou, Yali Wang 0001, Wanli Ouyang |
Pattern Recognit. | 5 |
| 2026 | Prototype-Based Multi-Dimension Intensity Mapping Density Sampling Network for Corrosion SegmentationabstractCorrosion semantic segmentation (CSS) is essential for early and accurate detection and positioning of corrosion in complex real-life scenarios. However, the unique characteristics of corrosion patterns, including the diverse forms, blurred boundaries, and intra-class heterogeneity, pose significant challenges in CSS. To address these challenges, we propose a Prototype-based Multi-dimension Sample-Adaptive Intensity Mapping with Density Sampling network (PMSAD) for CSS. PMSAD leverages nonparametric nearest prototype retrieving to enhance intra-class cohesion and inter-class separation, thereby handling the challenge of diverse forms. In PMSAD, prototypes are equally assigned to each class during training to mitigate class imbalance and capture intra-class variations. In addition, we elaborately design and implement three core components in PMSAD, including Multi-Scale Dual Attention (MSDA), Multi-dimension Sample-adaptive Intensity Mapping (MSAIM), and Density Sampling (DS). The MSDA enhances feature discrimination, facilitating robust representation learning. The end-to-end MSAIM adaptively adjusts RGB channel intensity contrasts of the input corrosion image to enhance feature robustness, counteracting the effects of uneven natural illumination. The DS is proposed for training refinement to tackle fuzzy boundaries and internal interference between corrosion classes. It focuses on high-density, high-error regions, offering refined guidance to correct intra-cluster centers and reduce inter-cluster similarity. Extensive evaluations on real-world datasets, including coarse and relabeled fine-grained dataset, validate the superior performance and generalization ability of PMSAD, achieving the new state-of-the-art performance in precise boundary delineation and accurate corrosion classification. The code is available at: https://github.com/c1oTTpD/PMSAD. Bohao Zhao, Gaoyang Pang, Luping Zhou, Yonghui Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | OnUVS: An Online Motion Transfer Framework With Content-Texture Decoupling for High-Fidelity Ultrasound Video SynthesisabstractUltrasound (US) imaging plays a crucial role in diagnosing heart and pelvic diseases, where sonographers tend to evaluate dynamic motion and structure. However, the scarcity of US videos for rare cases limitstraining opportunities for novice sonographers and deep learning models, hindering detection rates and clinical diagnostic applications. US video synthesis is a promising solution to this issue. Nevertheless, accurately imitating the intricate motion of the anatomy while preserving image fidelity presents asignificant challenge. In this work, we propose OnUVS, a novel online feature-decoupling framework for high-fidelity US video synthesis. First, to simulate realistic motion, we incorporate keypoints into anatomical learning through a weakly supervised training approach, which enhances motion representation and minimizes the need for fully annotated data. Second, we implement a dual-decoder generator that effectively balances content and textural features of generated frames, significantly enhancing the image fidelity of US videos. Third, a multi-scale discriminator further refines the sharpness and fine details, ensuring high-fidelity video synthesis. Fourth, an online learning strategy is designed to smooth coherence between frames by constraining the keypoint trajectories during inference. Validation on echocardiographic and pelvic floor US datasets demonstrates that OnUVS outperforms existing methods, achieving a 22.08% improvement in motion consistency (FVD) and 25.04% in image fidelity (FID). Rusi Chen, Xin Yang 0009, Ao Chang, Junxuan Yu, Yuhao Huang 0001, Ruobing Huang, Luping Zhou, Jiamin Liang, Haoran Dou, Yongsong Zhou, Mengyun Qiao, Deng-Ping Fan, Hongkui Yu, Dong Ni 0001, Zhongshan Gou |
IEEE J. Biomed. Health Informatics | 9 |
| 2026 | Medical Referring Image Segmentation via Next-Token Mask Prediction
Gaoyang Pang, Jiafu Hao, Chentao Yue, Luping Zhou, Yonghui Li 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2026 | Diversity-Enhanced Collaborative Mamba for Semi-Supervised Medical Image SegmentationabstractAcquiring high-quality annotated data for medical image segmentation is tedious and costly. Semi-supervised segmentation techniques alleviate this burden by leveraging unlabeled data to generate pseudo labels. Recently, advanced state space models, represented by Mamba, have shown efficient handling of long-range dependencies. This drives us to explore their potential in semi-supervised medical image segmentation. In this paper, we propose a novel Diversity-enhanced Collaborative Mamba framework (namely DCMamba) for semi-supervised medical image segmentation, which explores and utilizes the diversity from data, network, and feature perspectives. Firstly, from the data perspective, we develop patch-level weak-strong mixing augmentation with Mamba's scanning modeling characteristics. Moreover, from the network perspective, we introduce a diverse-scan collaboration module, which could benefit from the prediction discrepancies arising from different scanning directions. Furthermore, from the feature perspective, we adopt an uncertainty-weighted contrastive learning mechanism to enhance the diversity of feature representation. Experiments demonstrate that our DCMamba significantly outperforms other semi-supervised medical image segmentation methods, e.g., yielding the latest SSM-based method by 6.69% on the Synapse dataset with 20% labeled data. The code is available at https://github.com/ShumengLI/DCMamba. Shumeng Li, Jian Zhang 0090, Lei Qi 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | TB-HSU: Hierarchical 3D Scene Understanding with Contextual AffordancesabstractThe concept of function and affordance is a critical aspect of 3D scene understanding and supports task-oriented objectives. In this work, we develop a model that learns to structure and vary functional affordance across a 3D hierarchical scene graph representing the spatial organization of a scene. The varying functional affordance is designed to integrate with the varying spatial context of the graph. More specifically, we develop an algorithm that learns to construct a 3D hierarchical scene graph (3DHSG) that captures the spatial organization of the scene. Starting from segmented object point clouds and object semantic labels, we develop a 3DHSG with a top node that identifies the room label, child nodes that define local spatial regions inside the room with region-specific affordances, and grand-child nodes indicating object locations and object-specific affordances. To support this work, we create a custom 3DHSG dataset that provides ground truth data for local spatial regions with region-specific affordances and also object-specific affordances for each object. We employ a Transformer Based Hierarchical Scene Understanding (TB-HSU) model to learn the 3DHSG. We use a multi-task learning framework that learns both room classification and learns to define spatial regions within the room with region-specific affordances. Our work improves on the performance of state-of-the-art baseline models and shows one approach for applying transformer models to 3D scene understanding and the generation of 3DHSGs that capture the spatial organization of a room. The code and dataset are publicly available. Wenting Xu, Viorela Ila, Luping Zhou, Craig T. Jin |
AAAI | 3 |
| 2025 | DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy LabelsabstractMedical Visual Question Answering (Med-VQA) systems benefit the interpretation of medical images containing critical clinical information. However, the challenge of noisy labels and limited high-quality datasets remains underexplored. To address this, we establish the first benchmark for noisy labels in Med-VQA by simulating human mislabeling with semantically designed noise types. More importantly, we introduce the DiN framework, which leverages a diffusion model to handle noisy labels in Med-VQA. Unlike the dominant classification-based VQA approaches that directly predict answers, our Answer Diffuser (AD) module employs a coarse-to-fine process, refining answer candidates with a diffusion model for improved accuracy. The Answer Condition Generator (ACG) further enhances this process by generating task-specific conditional information via integrating answer embeddings with fused image-question features. To address label noise, our Noisy Label Refinement(NLR) module introduces a robust loss function and dynamic answer adjustment to further boost the performance of the AD module. Our DiN framework consistently outperforms existing methods across multiple benchmarks with varying noise levels1. Erjian Guo, Zhen Zhao 0001, Zicheng Wang 0012, Tong Chen 0011, Yunyi Liu, Luping Zhou |
CVPR | 6 |
| 2025 | InfGen: A Resolution-Agnostic Paradigm for Scalable Image SynthesisabstractArbitrary resolution image generation provides a consistent visual experience across devices, having extensive applications for producers and consumers. Current diffusion models increase computational demand quadratically with resolution, causing 4K image generation delays over 100 seconds. To solve this, we explore the second generation upon the latent diffusion models, where the fixed latent generated by diffusion models is regarded as the content representation and we propose to decode arbitrary resolution images with a compact generated latent using a one-step generator. Thus, we present the \textbf{InfGen}, replacing the VAE decoder with the new generator, for generating images at any resolution from a fixed-size latent without retraining the diffusion models, which simplifies the process, reducing computational complexity and can be applied to any model using the same latent space. Experiments show InfGen is capable of improving many models into the arbitrary high-resolution era while cutting 4K image generation time to under 10 seconds. Tao Han 0002, Wanghan Xu, Junchao Gong, Xiaoyu Yue, Song Guo 0001, Luping Zhou, Lei Bai 0001 |
ICCV | 6 |
| 2025 | Optimizing Efficiency and Visual-Textual Alignment for LLM-Based Radiology Report GenerationabstractLLM-based radiology report generation (R2Gen) systems have demonstrated promising performance but face significant challenges in bridging the gap between the visual encoder and the LLM. Specifically, two issues hinder progress: (1) parameter-heavy visual projector that increases complexity and degrades performance, and (2) insufficient alignment between visual and textual modalities, limiting system efficacy. To address these, we propose R2Gen-EVA, a novel framework emphasizing Efficiency and Visual-Textual Alignment (VTA), which introduces two key innovations: (1) a parameter-free visual projector that enhances model efficiency while improving performance, and (2) an LLM-adapted VTA module that enhances the alignment of visual features with LLM’s textual embeddings. Our design significantly improves model efficacy without adding extra parameters, achieving both streamlined complexity and higher computational efficiency during inference. Extensive experiments demonstrate that R2Gen-EVA enhances the fluency and clinical accuracy of generated reports, establishing it as a more effective and efficient solution for LLM-based R2Gen. The code is available at https://github.com/zailongchen/R2Gen-EVA. Zailong Chen, Yujian Lee, Johan Barthelemy, Luping Zhou, Lei Wang 0001 |
ICME | 5 |
| 2025 | SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
Tong Chen 0011, Shuya Yang, Long Bai 0008, Hongliang Ren 0001, Luping Zhou |
MICCAI (10) | 6 |
| 2025 | HiLa: Hierarchical Vision-Language Collaboration for Cancer Survival Prediction
Lu Wen, Yuchen Fei, Bo Liu 0113, Luping Zhou, Dinggang Shen, Yan Wang 0015 |
MICCAI (5) | 5 |
| 2025 | WiD-PET: PET Image Reconstruction from Low-Dose Data Using a Wavelet-Informed Diffusion Model with Fast Inference
Qingcheng Lyu, Tong Chen 0011, Erjian Guo, Luping Zhou |
MICCAI (16) | 5 |
| 2025 | MAK-GAN: Multi-level Adaptive Convolutional Kernels for Asymmetric Multi-modal PET Reconstruction
Xinyi Zeng, Pinxian Zeng, Yan Wang 0015, Luping Zhou, Caiwen Jiang, Han Zhang 0002, Dinggang Shen |
MICCAI (2) | 5 |
| 2025 | Unifying Generative Self-Supervised Paradigms with Diffusion ModelsabstractLearning semantic representations with generative self-supervised pre-training has demonstrated immense potential in the field of Natural Language Processing (NLP). However, this success has not been fully replicated in computer vision, primarily due to two key challenges: (1) visual generation tasks are difficult to formalize into a unified task, and (2) different downstream tasks require visual representations at varying semantic levels. To address these challenges, we propose a novel Generative framework with a UNified Self-supervised training paradigm (GUNS) to learn semantic information at different granularities and unify multiple image generation tasks within a denoising diffusion model. Specifically, GUNS employs various data augmentations to construct training objectives at different visual levels and unifies these training objectives into a generative paradigm. To facilitate this multi-task pre-training, the denoising diffusion model is introduced into GUNS as the image decoder, thereby making GUNS also an image-conditional generative model. Unlike other visual self-supervised methods that can only be applied to downstream tasks via transfer learning, GUNS can be directly applied to image generation tasks such as colorization and out-painting. We evaluate GUNS on downstream recognition tasks and image generation tasks, and experimental results demonstrate that GUNS can achieve competitive performance on both tasks simultaneously. Luping Zhou, Xiaoyu Yue |
MMAsia | 1 |
| 2025 | Understand Before You Generate: Self-Guided Training for Autoregressive Image GenerationabstractRecent studies have demonstrated the importance of high-quality visual representations in image generation and have highlighted the limitations of generative models in image understanding. As a generative paradigm originally designed for natural language, autoregressive models face similar challenges. In this work, we present the first systematic investigation into the mechanisms of applying the next-token prediction paradigm to the visual domain. We identify three key properties that hinder the learning of high-level visual semantics: local and conditional dependence, inter-step semantic inconsistency, and spatial invariance deficiency. We show that these issues can be effectively addressed by introducing self-supervised objectives during training, leading to a novel training framework, Self-guided Training for AutoRegressive models (ST-AR). Without relying on pre-trained representation models, ST-AR significantly enhances the image understanding ability of autoregressive models and leads to improved generation quality. Specifically, ST-AR brings approximately 42% FID improvement for LlamaGen-L and 49% FID improvement for LlamaGen-XL, while maintaining the same sampling strategy. Xiaoyu Yue, Zidong Wang 0004, Xihui Liu, Wanli Ouyang, Lei Bai 0001, Luping Zhou |
NeurIPS | 8 |
| 2025 | Enriching Category Representations with LLMs Towards Robust Zero-Shot OOD Detection
Dian Chao, Luping Zhou |
ECML/PKDD (1) | 3 |
| 2025 | Diagnostic Captioning by Cooperative Task Interactions and Sample-Graph ConsistencyabstractRadiographic images are similar to each other, making it challenging for diagnostic captioning to narrate fine-grained visual differences of clinical importance. In this paper, we propose a self-boosting framework integrating two novel strategies to learn tightly correlated image and text features for diagnostic captioning. The first strategy explicitly aligns image and text features through training an auxiliary task of image-text matching (ITM) jointly with the main task of report generation (RG) as two branches of a network model. The ITM branch explicitly learns image-text alignment and provides highly correlated visual and textual features for the RG branch to generate high-quality reports. The high-quality reports generated by RG branch, in turn, are utilized as additional harder negative samples to push the ITM branch to evolve towards better image-text alignment. These two branches help improve each other progressively, so that the whole model is self-boosted without requiring external resources. The second strategy aligns image-sample space and report-sample space to achieve consistent image and text feature embeddings. To achieve this, the sample graph of the embedded ground-truth reports is built and used as the target to train the sample graph of the embedded images so that the fine discrepancy in the ground-truth reports could be captured by the learned visual feature embeddings. Our proposed framework demonstrates its superiority on two medical report generation benchmarks, including the largest dataset MIMIC-CXR. Zhanyu Wang, Lei Wang 0001, Xiu Li 0001, Luping Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Neural Vector Fields: Generalizing Distance Vector Fields by Codebooks and Zero-Curl RegularizationabstractRecent neural networks based surface reconstruction can be roughly divided into two categories, one warping templates explicitly and the other representing 3D surfaces implicitly. To enjoy the advantages of both, we propose a novel 3D representation, Neural Vector Fields (NVF), which adopts the explicit learning process to manipulate meshes and implicit unsigned distance function (UDF) representation to break the barriers in resolution and topology. This is achieved by directly predicting the displacements from surface queries and modeling shapes as Vector Fields, rather than relying on network differentiation to obtain direction fields as most existing UDF-based methods do. In this way, our approach is capable of encoding both the distance and the direction fields so that the calculation of direction fields is differentiation-free, circumventing the non-trivial surface extraction step. Furthermore, building upon NVFs, we propose to incorporate two types of shape codebooks, i.e., NVFs (Lite or Ultra), to promote cross-category reconstruction through encoding cross-object priors. Moreover, we propose a new regularization based on analyzing the zero-curl property of NVFs, and implement this through the fully differentiable framework of our NVF (ultra). We evaluate both NVFs on four surface reconstruction scenarios, including watertight vs non-watertight shapes, category-agnostic reconstruction vs category-unseen reconstruction, category-specific, and cross-domain reconstruction. Xianghui Yang, Guosheng Lin, Luping Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Multi-Modal Long-Short Distance Attention-Based Transformer-GAN for PET Reconstruction With Auxiliary MRIabstractTo obtain high-quality PET scans while minimizing potential radiation hazards for patients, various GAN-based methods have been developed to reconstruct high-quality standard-count PET (SPET) images from low-count PET (LPET) ones. While recent efforts try to integrate MRI or CT to enhance reconstruction in a multi-modal way, current architectures mainly face two limitations: 1) CNN backbones or simple Transformer bottleneck layers are insufficient for robust semantic understanding; and 2) the identical strategies for multi-modal feature extraction and fusion overlook each modality’s respective importance for the reconstruction task. In this work, we propose the Multi-modal Long-Short Distance Attention-based Transformer-GAN (MLSDA-GAN), a novel network combining 3D transformer and CNN architecture for PET image reconstruction. Specifically, to extract fine-grained features with a small number of parameters, our MLSDA-GAN integrates multi-scale convolution into the embedding part of the transformer. As for our multi-modal design, given the strong correlation between LPET and SPET in structural characteristics, we treat MRI as an auxiliary modality to LPET and achieve effective multi-modal extraction and fusion strategies. These strategies include 1) a PET-specific Self-attention Extraction (PSE) block for comprehensive feature extraction of the primary LPET and 2) a Multi-modality Cross-attention Fusion (MCF) block for effective multi-modal interaction and fusion, enabling us to more efficiently model both long- and short-range relationships in the corresponding feature extraction and fusion processes. Experiments demonstrate superiority of our method quantitatively and qualitatively. Code is available athttps://github.com/Aru321/MLSDA-GAN. Pinxian Zeng, Xinyi Zeng, Yan Wang 0015, Luping Zhou, Chen Zu, Xi Wu 0004, Jiliu Zhou, Dinggang Shen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Enhancing Radiology Report Generation via Multi-Phased SupervisionabstractRadiology report generation using large language models has recently produced reports with more realistic styles and better language fluency. However, their clinical accuracy remains inadequate. Considering the significant imbalance between clinical phrases and general descriptions in a report, we argue that using an entire report for supervision is problematic as it fails to emphasize the crucial clinical phrases, which require focused learning. To address this issue, we propose a multi-phased supervision method, inspired by the spirit of curriculum learning where models are trained by gradually increasing task complexity. Our approach organizes the learning process into structured phases at different levels of semantical granularity, each building on the previous one to enhance the model. During the first phase, disease labels are used to supervise the model, equipping it with the ability to identify underlying diseases. The second phase progresses to use entity-relation triples to guide the model to describe associated clinical findings. Finally, in the third phase, we introduce conventional whole-report-based supervision to quickly adapt the model for report generation. Throughout the phased training, the model remains the same and consistently operates in the generation mode. As experimentally demonstrated, this proposed change in the way of supervision enhances report generation, achieving state-of-the-art performance in both language fluency and clinical accuracy. Our work underscores the importance of training process design in radiology report generation. Our code is available on https://github.com/zailongchen/MultiP-R2Gen. Zailong Chen, Yingshu Li 0002, Zhanyu Wang, Johan Barthelemy, Luping Zhou, Lei Wang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Mamba-Sea: A Mamba-Based Framework With Global-to-Local Sequence Augmentation for Generalizable Medical Image SegmentationabstractTo segment medical images with distribution shifts, domain generalization (DG) has emerged as a promising setting to train models on source domains that can generalize to unseen target domains. Existing DG methods are mainly based on CNN or ViT architectures. Recently, advanced state space models, represented by Mamba, have shown promising results in various supervised medical image segmentation. The success of Mamba is primarily owing to its ability to capture long-range dependencies while keeping linear complexity with input sequence length, making it a promising alternative to CNNs and ViTs. Inspired by the success, in the paper, we explore the potential of the Mamba architecture to address distribution shifts in DG for medical image segmentation. Specifically, we propose a novel Mamba-based framework, Mamba-Sea, incorporating global-to-local sequence augmentation to improve the model's generalizability under domain shift issues. Our Mamba-Sea introduces a global augmentation mechanism designed to simulate potential variations in appearance across different sites, aiming to suppress the model's learning of domain-specific information. At the local level, we propose a sequence-wise augmentation along input sequences, which perturbs the style of tokens within random continuous sub-sequences by modeling and resampling style statistics associated with domain shifts. To our best knowledge, Mamba-Sea is the first work to explore the generalization of Mamba for medical image segmentation, providing an advanced and promising Mamba-based architecture with strong robustness to domain shifts. Remarkably, our proposed method is the first to surpass a Dice coefficient of 90% on the Prostate dataset, which exceeds previous SOTA of 88.61%. The code is available at https://github.com/orange-czh/Mamba-Sea. Zihan Cheng 0001, Jintao Guo, Jian Zhang 0090, Lei Qi 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Imbalanced Medical Image Segmentation With Pixel-Dependent Noisy LabelsabstractAccurate medical image segmentation is often hindered by noisy labels in training data, due to the challenges of annotating medical images. Prior research works addressing noisy labels tend to make class-dependent assumptions, overlooking the pixel-dependent nature of most noisy labels. Furthermore, existing methods typically apply fixed thresholds to filter out noisy labels, risking the removal of minority classes and consequently degrading segmentation performance. To bridge these gaps, our proposed framework, Collaborative Learning with Curriculum Selection (CLCS), addresses pixel-dependent noisy labels with class imbalance. CLCS advances the existing works by i) treating noisy labels as pixel-dependent and addressing them through a collaborative learning framework, and ii) employing a curriculum dynamic thresholding approach adapting to model learning progress to select clean data samples to mitigate the class imbalance issue, and iii) applying a noise balance loss to noisy data samples to improve data utilization instead of discarding them outright. Specifically, our CLCS contains two modules: Curriculum Noisy Label Sample Selection (CNS) and Noise Balance Loss (NBL). In the CNS module, we designed a two-branch network with discrepancy loss for collaborative learning so that different feature representations of the same instance could be extracted from distinct views and used to vote the class probabilities of pixels. Besides, a curriculum dynamic threshold is adopted to select clean-label samples through probability voting. In the NBL module, instead of directly dropping the suspiciously noisy labels, we further adopt a robust loss to leverage such instances to boost the performance. We verify our CLCS on two benchmarks with different types of segmentation noise. Our method can obtain new state-of-the-art performance in different settings, yielding more than 3% Dice and mIoU improvements. Our code is available at https://github.com/Erjian96/CLCS.git. Erjian Guo, Zicheng Wang 0012, Zhen Zhao 0001, Luping Zhou |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Improving CXR Bone Suppression by Exploiting Domain-Level and Instance-Level InformationabstractFor chest X-ray image (CXR) analysis, effective bone structure suppression is essential for uncovering lung abnormalities and facilitating accurate clinical diagnoses. While recent deep generative models, to some extent, improve the reconstruction quality of bone-suppressed CXRs, they often fall short in delivering substantial improvements in downstream diagnosis tasks. This limitation is attributed to a narrow focus on instance-specific details, neglecting broader domain-level knowledge, which hampers bone-suppression effectiveness. In response to these challenges, our proposed framework adopts a novel approach that integrates both instance-level and domain-level information. To capture instance information, our model employs a hybrid approach using both cross-covariance attention blocks (CABs) to underscore relevant image information and a followed Vision Transformers (ViTs) encoder for image feature embedding. To capture domain information, we introduce multi-head codebook attention (MCA) which leverages codebook structure with multi-head attention mechanism to capture global, domain-level information specific to the bone-suppressed CXR domain, thereby refining the synthesis process. During optimization, our two-stage training scheme involves a MCA learning stage that encapsulates the domain of bone-suppressed CXRs in MCA through a ViT-based GAN model, and a synthesis stage that employs the learned codebook to generate bone-suppressed CXRs from the original ones, enhancing instance synthesis through domain insights. Moreover, the incorporation of CABs further refines pixel-level instance information. Extensive experiments demonstrate the superior performance of our approach, improving PSNR by 8.36% and SSIM by 2.7% for bone suppression while boosting lung disease classification by 2.8% and 4.2% on two datasets and segmentation by 1.5%. Kaisiyuan Wang, Luping Zhou |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Asynchronous Functional Brain Network Construction With Spatiotemporal Transformer for MCI ClassificationabstractConstruction and analysis of functional brain networks (FBNs) with resting-state functional magnetic resonance imaging (rs-fMRI) is a promising method to diagnose functional brain diseases. Nevertheless, the existing methods suffer from several limitations. First, the functional connectivities (FCs) of the FBN are usually measured by the temporal co-activation level between rs-fMRI time series from regions of interest (ROIs). While enjoying simplicity, the existing approach implicitly assumes simultaneous co-activation of all the ROIs, and models only their synchronous dependencies. However, the FCs are not necessarily always synchronous due to the time lag of information flow and cross-time interactions between ROIs. Therefore, it is desirable to model asynchronous FCs. Second, the traditional methods usually construct FBNs at individual level, leading to large variability and degraded diagnosis accuracy when modeling asynchronous FBN. Third, the FBN construction and analysis are conducted in two independent steps without joint alignment for the target diagnosis task. To address the first limitation, this paper proposes an effective sliding-window-based method to model spatiotemporal FCs in Transformer. Regarding the second limitation, we propose to learn common and individual FBNs adaptively with the common FBN as prior knowledge, thus alleviating the variability and enabling the network to focus on the individual disease-specific asynchronous FCs. To address the third limitation, the common and individual asynchronous FBNs are built and analyzed by an integrated network, enabling end-to-end training and improving the flexibility and discriminability. The effectiveness of the proposed method is consistently demonstrated on three data sets for mild cognitive impairment (MCI) diagnosis. Jianjia Zhang, Xiaotong Wu, Xiang Tang, Luping Zhou, Lei Wang 0001, Weiwen Wu, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Roll with the Punches: Expansion and Shrinkage of Soft Label Selection for Semi-supervised Fine-Grained LearningabstractWhile semi-supervised learning (SSL) has yielded promising results, the more realistic SSL scenario remains to be explored, in which the unlabeled data exhibits extremely high recognition difficulty, e.g., fine-grained visual classification in the context of SSL (SS-FGVC). The increased recognition difficulty on fine-grained unlabeled data spells disaster for pseudo-labeling accuracy, resulting in poor performance of the SSL model. To tackle this challenge, we propose Soft Label Selection with Confidence-Aware Clustering based on Class Transition Tracking (SoC) by reconstructing the pseudo-label selection process by jointly optimizing Expansion Objective and Shrinkage Objective, which is based on a soft label manner. Respectively, the former objective encourages soft labels to absorb more candidate classes to ensure the attendance of ground-truth class, while the latter encourages soft labels to reject more noisy classes, which is theoretically proved to be equivalent to entropy minimization. In comparisons with various state-of-the-art methods, our approach demonstrates its superior performance in SS-FGVC. Checkpoints and source code are available at https://github.com/NJUyued/SoC4SS-FGVC. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
AAAI | 4 |
| 2024 | Noise-Aware Image Captioning with Progressively Exploring Mismatched WordsabstractImage captioning aims to automatically generate captions for images by learning a cross-modal generator from vision to language. The large amount of image-text pairs required for training is usually sourced from the internet due to the manual cost, which brings the noise with mismatched relevance that affects the learning process. Unlike traditional noisy label learning, the key challenge in processing noisy image-text pairs is to finely identify the mismatched words to make the most use of trustworthy information in the text, rather than coarsely weighing the entire examples. To tackle this challenge, we propose a Noise-aware Image Captioning method (NIC) to adaptively mitigate the erroneous guidance from noise by progressively exploring mismatched words. Specifically, NIC first identifies mismatched words by quantifying word-label reliability from two aspects: 1) inter-modal representativeness, which measures the significance of the current word by assessing cross-modal correlation via prediction certainty; 2) intra-modal informativeness, which amplifies the effect of current prediction by combining the quality of subsequent word generation. During optimization, NIC constructs the pseudo-word-labels considering the reliability of the origin word-labels and model convergence to periodically coordinate mismatched words. As a result, NIC can effectively exploit both clean and noisy image-text pairs to learn a more robust mapping function. Extensive experiments conducted on the MS-COCO and Conceptual Caption datasets validate the effectiveness of our method in various noisy scenarios. Zhongtian Fu, Kefei Song, Luping Zhou |
AAAI | 3 |
| 2024 | UFDA: Universal Federated Domain Adaptation with Practical AssumptionsabstractConventional Federated Domain Adaptation (FDA) approaches usually demand an abundance of assumptions, which makes them significantly less feasible for real-world situations and introduces security hazards. This paper relaxes the assumptions from previous FDAs and studies a more practical scenario named Universal Federated Domain Adaptation (UFDA). It only requires the black-box model and the label set information of each source domain, while the label sets of different source domains could be inconsistent, and the target-domain label set is totally blind. Towards a more effective solution for our newly proposed UFDA scenario, we propose a corresponding methodology called Hot-Learning with Contrastive Label Disambiguation (HCLD). It particularly tackles UFDA's domain shifts and category gaps problems by using one-hot outputs from the black-box models of various source domains. Moreover, to better distinguish the shared and unknown classes, we further present a cluster-level strategy named Mutual-Voting Decision (MVD) to extract robust consensus knowledge across peer classes from both source and target domains. Extensive experiments on three benchmark datasets demonstrate that our method achieves comparable performance for our UFDA scenario with much fewer assumptions, compared to previous methodologies with comprehensive additional assumptions. Luping Zhou, Dong Xu 0001, Wei Xi 0003, Gairui Bai, Yihan Zhao, Jizhong Zhao |
AAAI | 3 |
| 2024 | Progressive Classifier and Feature Extractor Adaptation for Unsupervised Domain Adaptation on Point Clouds
Zicheng Wang 0012, Zhen Zhao 0001, Yiming Wu 0005, Luping Zhou, Dong Xu 0001 |
ECCV (28) | 4 |
| 2024 | Alternate Diverse Teaching for Semi-supervised Medical Image Segmentation
Zhen Zhao 0001, Zicheng Wang 0012, Longyue Wang, Dian Yu 0001, Yixuan Yuan, Luping Zhou |
ECCV (5) | 6 |
| 2024 | Towards Speech Classification from Acoustic and Vocal Tract data in Real-time MRI
Yaoyao Yue, Michael Proctor, Luping Zhou, Rijul Gupta, Tharinda Piyadasa, Amelia Gully, Kirrie J. Ballard, Craig T. Jin |
INTERSPEECH | 3 |
| 2024 | LighTDiff: Surgical Endoscopic Image Low-Light Enhancement with T-Diffusion
Tong Chen 0011, Qingcheng Lyu, Long Bai 0008, Erjian Guo, Huxin Gao, Xiaoxiao Yang, Hongliang Ren 0001, Luping Zhou |
MICCAI (6) | 8 |
| 2024 | KARGEN: Knowledge-Enhanced Automated Radiology Report Generation Using Large Language Models
Yingshu Li 0002, Zhanyu Wang, Yunyi Liu, Lei Wang 0001, Lingqiao Liu, Luping Zhou |
MICCAI (5) | 6 |
| 2024 | MRScore: Evaluating Medical Report with LLM-Based Reward System
Yunyi Liu, Zhanyu Wang, Yingshu Li 0002, Lingqiao Liu, Lei Wang 0001, Luping Zhou |
MICCAI (3) | 7 |
| 2024 | Group-aware Parameter-efficient Updating for Content-Adaptive Neural Video CompressionabstractContent-adaptive compression is crucial for enhancing the adaptability of the pre-trained neural codec for various contents. Though, its application in neural video compression (NVC) is still limited due to two main aspects: 1), video compression relies heavily on temporal redundancy, therefore updating just one or a few frames can lead to significant errors accumulating over time; 2), NVC frameworks are generally complex, with many comprehensive components that are not trivial to update quickly during the encoding procedure. To address these challenges, we have developed a content-adaptive NVC technique called Group-aware Parameter-efficient Updating (GPU). Initially, to minimize error accumulation, we adopt a group-aware approach for updating encoder parameters. This involves adopting a patch-based Group of Pictures (GoP) updating strategy to segment a video into patch-based GoPs, which will be updated to facilitate a globally optimized domain-transferable solution. Subsequently, we introduce a parameter-efficient delta-tuning strategy, which is achieved by integrating several light-weight adapters into each encoding component by using both serial and parallel configuration. Such architecture-agnostic modules stimulate the components with large parameters, thereby reducing the updating cost during the encoding stage. We incorporate our GPU into the latest NVC framework and conduct extensive experiments, whose results showcase outstanding video compression efficiency across six video compression benchmarks and the adaptability of one medical volumetric image compression benchmark. Luping Zhou, Dong Xu 0001 |
ACM Multimedia | 2 |
| 2024 | GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation
Zhanyu Wang, Longyue Wang, Zhen Zhao 0001, Minghao Wu, Chenyang Lyu, Deng Cai 0002, Luping Zhou, Shuming Shi 0001, Zhaopeng Tu |
ACM Multimedia | 8 |
| 2024 | DSANet: Dual-path segmentation-guided attention network for radiotherapy dose prediction from CT images only
Lu Wen, Zhengyang Jiao, Jianghong Xiao, Luping Zhou, Yanmei Luo, Jiliu Zhou, Xingchen Peng, Yan Wang 0015 |
Knowl. Based Syst. | 5 |
| 2024 | 3D multi-modality Transformer-GAN for high-quality PET reconstruction
Yan Wang 0015, Yanmei Luo, Chen Zu, Bo Zhan, Zhengyang Jiao, Xi Wu 0004, Jiliu Zhou, Dinggang Shen, Luping Zhou |
Medical Image Anal. | 9 |
| 2024 | Constructing hierarchical attentive functional brain networks for early AD diagnosis
Jianjia Zhang, Yunan Guo, Luping Zhou, Lei Wang 0001, Weiwen Wu, Dinggang Shen |
Medical Image Anal. | 3 |
| 2024 | A Survey on Efficient Vision Transformers: Algorithms, Techniques, and Performance BenchmarkingabstractVision Transformer (ViT) architectures are becoming increasingly popular and widely employed to tackle computer vision applications. Their main feature is the capacity to extract global information through the self-attention mechanism, outperforming earlier convolutional neural networks. However, ViT deployment and performance have grown steadily with their size, number of trainable parameters, and operations. Furthermore, self-attention's computational and memory cost quadratically increases with the image resolution. Generally speaking, it is challenging to employ these architectures in real-world applications due to many hardware and environmental restrictions, such as processing and computational capabilities. Therefore, this survey investigates the most efficient methodologies to ensure sub-optimal estimation performances. More in detail, four efficient categories will be analyzed: compact architecture, pruning, knowledge distillation, and quantization strategies. Moreover, a new metric called Efficient Error Rate has been introduced in order to normalize and compare models' features that affect hardware devices at inference time, such as the number of parameters, bits, FLOPs, and model size. Summarizing, this paper first mathematically defines the strategies used to make Vision Transformer efficient, describes and discusses state-of-the-art methodologies, and analyzes their performances over different application scenarios. Toward the end of this paper, we also discuss open challenges and promising research directions. Lorenzo Papa, Paolo Russo 0001, Irene Amerini, Luping Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Semi-supervised medical image segmentation via hard positives oriented contrastive learning
Cheng Tang 0003, Xinyi Zeng, Luping Zhou, Qizheng Zhou, Xi Wu 0004, Hongping Ren, Jiliu Zhou, Yan Wang 0015 |
Pattern Recognit. | 3 |
| 2024 | CL-TransFER: Collaborative learning based transformer for facial expression recognition with masked reconstruction
Chen Zu, Jianjia Zhang, Jiliu Zhou, Luping Zhou, Yan Wang 0015 |
Pattern Recognit. | 8 |
| 2024 | 3D Point-Based Multi-Modal Context Clusters GAN for Low-Dose PET Image DenoisingabstractTo obtain high-quality Positron emission tomography (PET) images while minimizing radiation hazards, various methods have been developed to acquire standard-dose PET (SPET) images from low-dose PET (LPET) images. Recent efforts mainly focus on improving the denoising quality by utilizing multi-modal inputs. However, these methods exhibit certain limitations. First, they neglect the varied significance of each modality in denoising. Second, they rely on inflexible voxel-based representations, failing to explicitly preserve intricate structures and contexts in images. To alleviate these problems, we propose a 3D Point-based Multi-modal Context Clusters GAN, namely PMC2-GAN, for obtaining high-quality SPET images from LPET and magnetic resonance imaging (MRI) images. Specifically, we transform the 3D image into unorganized points to flexibly and precisely express its complex structure. Moreover, a self-context clusters (Self-CC) block is devised to explore fine-grained contextual relationships of the image from the perspective of points. Additionally, considering the diverse importance of different modalities, we introduce a cross-context clusters (Cross-CC) block, which prioritizes PET as the primary modality while regarding MRI as the auxiliary one, to effectively integrate the knowledge from the two modalities. Overall, built on the smart integration of Self- and Cross-CC blocks, our PMC2-GAN follows GAN architecture. Extensive experiments validate our superiority. Yan Wang 0015, Luping Zhou, Yuchen Fei, Jiliu Zhou, Dinggang Shen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Rebalanced Vision-Language Retrieval Considering Structure-Aware DistillationabstractVision-language retrieval aims to search for similar instances in one modality based on queries from another modality. The primary objective is to learn cross-modal matching representations in a latent common space. Actually, the assumption underlying cross-modal matching is modal balance, where each modality contains sufficient information to represent the others. However, noise interference and modality insufficiency often lead to modal imbalance, making it a common phenomenon in practice. The impact of imbalance on retrieval performance remains an open question. In this paper, we first demonstrate that ultimate cross-modal matching is generally sub-optimal for cross-modal retrieval when imbalanced modalities exist. The structure of instances in the common space is inherently influenced when facing imbalanced modalities, posing a challenge to cross-modal similarity measurement. To address this issue, we emphasize the importance of meaningful structure-preserved matching. Accordingly, we propose a simple yet effective method to rebalance cross-modal matching by learning structure-preserved matching representations. Specifically, we design a novel multi-granularity cross-modal matching that incorporates structure-aware distillation alongside the cross-modal matching loss. While the cross-modal matching loss constraints instance-level matching, the structure-aware distillation further regularizes the geometric consistency between learned matching representations and intra-modal representations through the developed relational matching. Extensive experiments on different datasets affirm the superior cross-modal retrieval performance of our approach, simultaneously enhancing single-modal retrieval capabilities compared to the baseline models. Yang Yang 0074, Wenjuan Xi, Luping Zhou, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | MutexMatch: Semi-Supervised Learning With Mutex-Based Consistency RegularizationabstractThe core issue in semi-supervised learning (SSL) lies in how to effectively leverage unlabeled data, whereas most existing methods tend to put a great emphasis on the utilization of high-confidence samples yet seldom fully explore the usage of low-confidence samples. In this article, we aim to utilize low-confidence samples in a novel way with our proposed mutex-based consistency regularization, namely MutexMatch. Specifically, the high-confidence samples are required to exactly predict "what it is" by the conventional true-positive classifier (TPC), while low-confidence samples are employed to achieve a simpler goal-to predict with ease "what it is not" by the true-negative classifier (TNC). In this sense, we not only mitigate the pseudo-labeling errors but also make full use of the low-confidence unlabeled data by the consistency of dissimilarity degree. MutexMatch achieves superior performance on multiple benchmark datasets, i.e., Canadian Institute for Advanced Research (CIFAR)-10, CIFAR-100, street view house numbers (SVHN), self-taught learning 10 (STL-10), and mini-ImageNet. More importantly, our method further shows superiority when the amount of labeled data is scarce, e.g., 92.23% accuracy with only 20 labeled data on CIFAR-10. Code has been released at https://github.com/NJUyued/MutexMatch4SSL. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Lei Wang 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Learning Partial Correlation based Deep Visual Representation for Image ClassificationabstractVisual representation based on covariance matrix has demonstrates its efficacy for image classification by characterising the pairwise correlation of different channels in convolutional feature maps. However, pairwise correlation will become misleading once there is another channel correlating with both channels of interest, resulting in the “confounding” effect. For this case, “partial correlation” which removes the confounding effect shall be estimated instead. Nevertheless, reliably estimating partial correlation requires to solve a symmetric positive definite matrix optimisation, known as sparse inverse covariance estimation (SICE). How to incorporate this process into CNN remains an open issue. In this work, we formulate SICE as a novel structured layer of CNN. To ensure end-to-end trainability, we develop an iterative method to solve the above matrix optimisation during forward and backward propagation steps. Our work obtains a partial correlation based deep visual representation and mitigates the small sample problem often encountered by covariance matrix estimation in CNN. Computationally, our model can be effectively trained with GPU and works well with a large number of channels of advanced CNNs. Experiments show the efficacy and superior classification performance of our deep visual representation compared to covariance matrix based counterparts. Saimunur Rahman, Piotr Koniusz, Lei Wang 0001, Luping Zhou, Peyman Moghadam, Changming Sun |
CVPR | 4 |
| 2023 | METransformer: Radiology Report Generation by Transformer with Multiple Learnable Expert TokensabstractIn clinical scenarios, multi-specialist consultation could significantly benefit the diagnosis, especially for intricate cases. This inspires us to explore a “multi-expert joint diagnosis” mechanism to upgrade the existing “single expert” framework commonly seen in the current literature. To this end, we propose METransformer, a method to realize this idea with a transformer-based backbone. The key design of our method is the introduction of multiple learnable “expert” tokens into both the transformer encoder and decoder. In the encoder, each expert token interacts with both vision tokens and other expert tokens to learn to attend different image regions for image representation. These expert tokens are encouraged to capture complementary information by an orthogonal loss that minimizes their over-lap. In the decoder, each attended expert token guides the cross-attention between input words and visual tokens, thus influencing the generated report. A metrics-based expert voting strategy is further developed to generate the final report. By the multi-experts concept, our model enjoys the merits of an ensemble-based approach but through a manner that is computationally more efficient and supports more sophisticated interactions among experts. Experimental results demonstrate the promising performance of our proposed model on two widely used benchmarks. Last but not least, the framework-level innovation makes our work ready to incorporate advances on existing “single-expert” models to further improve its performance. Zhanyu Wang, Lingqiao Liu, Lei Wang 0001, Luping Zhou |
CVPR | 4 |
| 2023 | Conflict-Based Cross-View Consistency for Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation (SSS) has recently gained increasing research interest as it can reduce the requirement for large-scale fully-annotated training data. The current methods often suffer from the confirmation bias from the pseudo-labelling process, which can be alleviated by the co-training framework. The current co-training-based SSS methods rely on hand-crafted perturbations to prevent the different sub-nets from collapsing into each other, but these artificial perturbations cannot lead to the optimal solution. In this work, we propose a new conflict-based cross-view consistency (CCVC) method based on a two-branch co-training framework which aims at enforcing the two sub-nets to learn informative features from irrelevant views. In particular, we first propose a new cross-view consistency (CVC) strategy that encourages the two sub-nets to learn distinct features from the same input by introducing a feature discrepancy loss, while these distinct features are expected to generate consistent prediction scores of the input. The CVC strategy helps to prevent the two sub-nets from stepping into the collapse. In addition, we further propose a conflict-based pseudo-labelling (CPL) method to guarantee the model will learn more useful information from conflicting predictions, which will lead to a stable training process. We validate our new CCVC approach on the SSS benchmark datasets where our method achieves new state-of-the-art performance. Our code is available at https://github.com/xiaoyao3302/CCVC. Zicheng Wang 0012, Zhen Zhao 0001, Xiaoxia Xing, Dong Xu 0001, Luping Zhou |
CVPR | 6 |
| 2023 | Neural Vector Fields: Implicit Representation by Explicit LearningabstractDeep neural networks (DNNs) are widely applied for nowadays 3D surface reconstruction tasks and such methods can be further divided into two categories, which respectively warp templates explicitly by moving vertices or represent 3D surfaces implicitly as signed or unsigned distance functions. Taking advantage of both advanced explicit learning process and powerful representation ability of implicit functions, we propose a novel 3D representation method, Neural Vector Fields (NVF). It not only adopts the explicit learning process to manipulate meshes directly, but also leverages the implicit representation of unsigned distance functions (UDFs) to break the barriers in resolution and topology. Specifically, our method first predicts the displacements from queries towards the surface and models the shapes as Vector Fields. Rather than relying on network differentiation to obtain direction fields as most existing UDF-based methods, the produced vector fields encode the distance and direction fields both and mitigate the ambiguity at “ridge” points, such that the calculation of direction fields is straightforward and differentiation-free. The differentiation-free characteristic enables us to further learn a shape codebook via Vector Quantization, which encodes the cross-object priors, accelerates the training procedure, and boosts model generalization on cross-category reconstruction. The extensive experiments on surface reconstruction benchmarks indicate that our method outperforms those state-of-the-art methods in different evaluation scenarios including watertight vs non-watertight shapes, category-specific vs category-agnostic reconstruction, category-unseen reconstruction, and cross-domain reconstruction. Our code is released at https://github.com/Wi-sc/NVF. Xianghui Yang, Guosheng Lin, Luping Zhou |
CVPR | 4 |
| 2023 | Instance-Specific and Model-Adaptive Supervision for Semi-Supervised Semantic SegmentationabstractRecently, semi-supervised semantic segmentation has achieved promising performance with a small fraction of labeled data. However, most existing studies treat all unlabeled data equally and barely consider the differences and training difficulties among unlabeled instances. Differentiating unlabeled instances can promote instance-specific supervision to adapt to the model's evolution dynamically. In this paper, we emphasize the cruciality of instance differences and propose an instance-specific and model-adaptive supervision for semi-supervised semantic segmentation, named iMAS. Relying on the model's performance, iMAS employs a class-weighted symmetric intersection-over-union to evaluate quantitative hardness of each unlabeled instance and supervises the training on unlabeled data in a model-adaptive manner. Specifically, iMAS learns from unlabeled instances progressively by weighing their corresponding consistency losses based on the evaluated hardness. Besides, iMAS dynamically adjusts the augmentation for each instance such that the distortion degree of augmented instances is adapted to the model's generalization capability across the training course. Not integrating additional losses and training procedures, iMAS can obtain remarkable performance gains against current state-of-the-art approaches on segmentation benchmarks under different semi-supervised partition protocols11Code and logs: https://github.com/zhenzhao/iMAS. Zhen Zhao 0001, Sifan Long 0001, Jimin Pi, Jingdong Wang 0001, Luping Zhou |
CVPR | 5 |
| 2023 | Augmentation Matters: A Simple-Yet-Effective Approach to Semi-Supervised Semantic SegmentationabstractRecent studies on semi-supervised semantic segmentation (SSS) have seen fast progress. Despite their promising performance, current state-of-the-art methods tend to increasingly complex designs at the cost of introducing more network components and additional training procedures. Differently, in this work, we follow a standard teacher-student framework and propose AugSeg, a simple and clean approach that focuses mainly on data perturbations to boost the SSS performance. We argue that various data augmentations should be adjusted to better adapt to the semi-supervised scenarios instead of directly applying these techniques from supervised learning. Specifically, we adopt a simplified intensity-based augmentation that selects a random number of data transformations with uniformly sampling distortion strengths from a continuous space. Based on the estimated confidence of the model on different unlabeled samples, we also randomly inject labelled information to augment the unlabeled samples in an adaptive manner. Without bells and whistles, our simple AugSeg can readily achieve new state-of-the-art performance on SSS benchmarks under different partition protocols11Code and logs: https://github.com/zhenzhao/AugSeg.. Zhen Zhao 0001, Lihe Yang, Sifan Long 0001, Jimin Pi, Luping Zhou, Jingdong Wang 0001 |
CVPR | 5 |
| 2023 | Automatic Radiology Report Generation by Learning with Increasingly Hard NegativesabstractAutomatic radiology report generation is challenging as medical images or reports are usually similar to each other due to the common content of anatomy. This makes a model hard to capture the uniqueness of individual images and is prone to producing undesired generic or mismatched reports. This situation calls for learning more discriminative features that could capture even fine-grained mismatches between images and reports. To achieve this, this paper proposes a novel framework to learn discriminative image and report features by distinguishing them from their closest peers, i.e., hard negatives. Especially, to attain more discriminative features, we gradually raise the difficulty of such a learning task by creating increasingly hard negative reports for each image in the feature space during training, respectively. By treating the increasingly hard negatives as auxiliary variables, we formulate this process as a min-max alternating optimisation problem. At each iteration, conditioned on a given set of hard negative reports, image and report features are learned as usual by minimising the loss functions related to report generation. After that, a new set of harder negative reports will be created by maximising a loss reflecting image-report alignment. By solving this optimisation, we attain a model that can generate more specific and accurate reports. It is noteworthy that our framework enhances discriminative feature learning without introducing extra network weights. Also, in contrast to the existing way of generating hard negatives, our framework extends beyond the granularity of the dataset by generating harder samples out of the training set. Experimental study on benchmark datasets verifies the efficacy of our framework and shows that it can serve as a plug-in to readily improve existing medical report generation models. The code is publicly available at https://github.com/Bhanu068/ITHN. Bhanu Prakash Voutharoja, Lei Wang 0001, Luping Zhou |
ECAI | 3 |
| 2023 | Towards Semi-supervised Learning with Non-random Missing LabelsabstractSemi-supervised learning (SSL) tackles the label missing problem by enabling the effective usage of unlabeled data. While existing SSL methods focus on the traditional setting, a practical and challenging scenario called label Missing Not At Random (MNAR) is usually ignored. In MNAR, the labeled and unlabeled data fall into different class distributions resulting in biased label imputation, which deteriorates the performance of SSL models. In this work, class transition tracking based Pseudo-Rectifying Guidance (PRG) is devised for MNAR. We explore the class-level guidance information obtained by the Markov random walk, which is modeled on a dynamically created graph built over the class tracking matrix. PRG unifies the historical information of class distribution and class transitions caused by the pseudo-rectifying procedure to maintain the model’s unbiased enthusiasm towards assigning pseudo-labels to all classes, so as the quality of pseudo-labels on both popular classes and rare classes in MNAR could be improved. Finally, we show the superior performance of PRG across a variety of MNAR scenarios, outperforming the latest SSL approaches combining bias removal solutions by a large margin. Code and model weights are available at https://github.com/NJUyued/PRG4SSL-MNAR. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
ICCV | 4 |
| 2023 | Enhancing Sample Utilization through Sample Adaptive Augmentation in Semi-Supervised LearningabstractIn semi-supervised learning, unlabeled samples can be utilized through augmentation and consistency regularization. However, we observed certain samples, even undergoing strong augmentation, are still correctly classified with high confidence, resulting in a loss close to zero. It indicates that these samples have been already learned well and do not provide any additional optimization benefits to the model. We refer to these samples as "naive samples". Unfortunately, existing SSL models overlook the characteristics of naive samples, and they just apply the same learning strategy to all samples. To further optimize the SSL model, we emphasize the importance of giving attention to naive samples and augmenting them in a more diverse manner. Sample adaptive augmentation (SAA) is proposed for this stated purpose and consists of two modules: 1) sample selection module; 2) sample augmentation module. Specifically, the sample selection module picks out naive samples based on historical training information at each epoch, then the naive samples will be augmented in a more diverse manner in the sample augmentation module. Thanks to the extreme ease of implementation of the above modules, SAA is advantageous for being simple and lightweight. We add SAA on top of FixMatch and FlexMatch respectively, and experiments demonstrate SAA can significantly improve the models. For example, SAA helped improve the accuracy of FixMatch from 92.50% to 94.76% and that of FlexMatch from 95.01% to 95.31% on CIFAR-10 with 40 labels. The code is available at https://github.com/GuanGui-nju/SAA. Guan Gui 0002, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
ICCV | 4 |
| 2023 | Task-Oriented Multi-Modal Mutual Learning for Vision-Language ModelsabstractPrompt learning has become one of the most efficient paradigms for adapting large pre-trained vision-language models to downstream tasks. Current state-of-the-art methods, like CoOp and ProDA, tend to adopt soft prompts to learn an appropriate prompt for each specific task. Recent CoCoOp further boosts the base-to-new generalization performance via an image-conditional prompt. However, it directly fuses identical image semantics to prompts of different labels and significantly weakens the discrimination among different classes as shown in our experiments. Motivated by this observation, we first propose a class-aware text prompt (CTP) to enrich generated prompts with label-related image information. Unlike CoCoOp, CTP can effectively involve image semantics and avoid introducing extra ambiguities into different prompts. On the other hand, instead of reserving the complete image representations, we propose text-guided feature tuning (TFT) to make the image branch attend to class-related representation. A contrastive loss is employed to align such augmented text and image representations on downstream tasks. In this way, the image-to-text CTP and text-to-image TFT can be mutually promoted to enhance the adaptation of VLMs for downstream tasks. Extensive experiments demonstrate that our method outperforms the existing methods by a significant margin. Especially, compared to CoCoOp, we achieve an average improvement of 4.03% on new classes and 3.19% on harmonic-mean over eleven classification benchmarks. Sifan Long 0001, Zhen Zhao 0001, Junkun Yuan, Zichang Tan, Jiangjiang Liu 0006, Luping Zhou, Sheng-Sheng Wang 0001, Jingdong Wang 0001 |
ICCV | 6 |
| 2023 | Learning Spatial-context-aware Global Visual Feature Representation for Instance Image RetrievalabstractIn instance image retrieval, considering local spatial information within an image has proven effective to boost retrieval performance, as demonstrated by local visual descriptor based geometric verification. Nevertheless, it will be highly valuable to make ordinary global image representations spatial-context-aware because global representation based image retrieval is appealing thanks to its algorithmic simplicity, low memory cost, and being friendly to sophisticated data structures. To this end, we propose a novel feature learning framework for instance image retrieval, which embeds local spatial context information into the learned global feature representations. Specifically, in parallel to the visual feature branch in a CNN backbone, we design a spatial context branch that consists of two modules called online token learning and distance encoding. For each local descriptor learned in CNN, the former module is used to indicate the types of its surrounding descriptors, while their spatial distribution information is captured by the latter module. After that, the visual feature branch and the spatial context branch are fused to produce a single global feature representation per image. As experimentally demonstrated, with the spatial-context-aware characteristic, we can well improve the performance of global representation based image retrieval while maintaining all of its appealing properties. Our code is available at https://github.com/Zy-Zhang/SpCa. Zhongyan Zhang, Lei Wang 0001, Luping Zhou, Piotr Koniusz |
ICCV | 3 |
| 2023 | Contrastive Diffusion Model with Auxiliary Guidance for Coarse-to-Fine PET Reconstruction
Zeyu Han, Luping Zhou, Binyu Yan, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
MICCAI (10) | 3 |
| 2023 | Neural Video Compression with Spatio-Temporal Cross-Covariance TransformersabstractAlthough existing neural video compression~(NVC) methods have achieved significant success, most of them focus on improving either temporal or spatial information separately. They generally use simple operations such as concatenation or subtraction to utilize this information, while such operations only partially exploit spatio-temporal redundancies. This work aims to effectively and jointly leverage robust temporal and spatial information by proposing a new 3D-based transformer module: Spatio-Temporal Cross-Covariance Transformer (ST-XCT). The ST-XCT module combines two individual extracted features into a joint spatio-temporal feature, followed by 3D convolutional operations and a novel spatio-temporal-aware cross-covariance attention mechanism. Unlike conventional transformers, the cross-covariance attention mechanism is applied across the feature channels without breaking down the spatio-temporal features into local tokens. Such design allows for modeling global cross-channel correlations of the spatio-temporal context while lowering the computational requirement. Based on ST-XCT, we introduce a novel transformer-based end-to-end optimized NVC framework. ST-XCT-based modules are integrated into various key coding components of NVC, such as feature extraction, frame reconstruction, and entropy modeling, demonstrating its generalizability. Extensive experiments show that our ST-XCT-based NVC proposal achieves state-of-the-art compression performances on various standard video benchmark datasets. Lucas Relic, Roberto Azevedo, Yang Zhang 0003, Markus Gross 0001, Dong Xu 0001, Luping Zhou, Christopher Schroers |
ACM Multimedia | 7 |
| 2023 | Entropy-based Optimization on Individual and Global Predictions for Semi-Supervised LearningabstractPseudo-labelling-based semi-supervised learning (SSL) has demonstrated remarkable success in enhancing model performance by effectively leveraging a large amount of unlabeled data. However, existing studies focus mainly on rectifying individual predictions (i.e., pseudo-labels) on each unlabeled instance but ignore the overall prediction statistics from a global perspective. Such neglect may lead to model collapse and performance degradation in SSL, especially in label-scarce scenarios. In this paper, we emphasize the cruciality of global prediction constraints and propose a new SSL method that employs Entropy-based optimization on both Individual and Global predictions of unlabeled instances, dubbed EntInG. Specifically, we propose two criteria for leveraging unlabeled data in SSL: individual prediction entropy minimization (IPEM) and global distribution entropy maximization (GDEM). On the one hand, we show that current dominant SSL methods can be viewed as an implicit form of IPEM improved by recent augmentation techniques. On the other hand, we construct a new distribution loss to encourage GDEM, which greatly benefits producing better pseudo-labels for unlabeled data. Theoretical analysis also demonstrates that our proposed criteria can be derived by enforcing mutual information maximization on unlabeled instances. Despite its simplicity, our proposed method can achieve significant accuracy gains on popular SSL classification benchmarks. Zhen Zhao 0001, Meng Zhao 0003, Ye Liu 0013, Luping Zhou |
ACM Multimedia | 5 |
| 2023 | Dataset-Driven Unsupervised Object Discovery for Region-Based Instance Image RetrievalabstractInstance image retrieval could greatly benefit from discovering objects in the image dataset. This not only helps produce more reliable feature representation but also better informs users by delineating query-matched object regions. However, object classes are usually not predefined in a retrieval dataset and class label information is generally unavailable in image retrieval. This situation makes object discovery a challenging task. To address this, we propose a novel dataset-driven unsupervised object discovery framework. By utilizing deep feature representation and weakly-supervised object detection, we explore supervisory information from within an image dataset, construct class-wise object detectors, and assign multiple detectors to each image for detection. To efficiently construct object detectors for large image datasets, we propose a novel "base-detector repository" and derive a fast way to generate the base detectors. In addition, the whole framework is designed to work in a self-boosting manner to iteratively refine object discovery. Compared with existing unsupervised object detection methods, our framework produces more accurate object discovery results. Different from supervised detection, we need neither manual annotation nor auxiliary datasets to train object detectors. Experimental study demonstrates the effectiveness of the proposed framework and the improved performance for region-based instance image retrieval. Zhongyan Zhang, Lei Wang 0001, Yang Wang 0002, Luping Zhou, Jianjia Zhang, Fang Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Kernel-based feature aggregation framework in point cloud networks
Jianjia Zhang, Lei Wang 0001, Luping Zhou, Xiaocai Zhang, Weiwen Wu |
Pattern Recognit. | 4 |
| 2023 | Stay in Grid: Improving Video Captioning via Fully Grid-Level RepresentationabstractVideo captioning is a challenging task of automatically generating natural and meaningful textual descriptions given some context videos. The state-of-the-art methods aggregate the spatial-wise information in the video encoder at the early stage, which has two drawbacks: 1) Early aggregation in the encoder can cause considerable spatial details missing, which may consequently lead to incorrect word choices in the following text encoder. 2) The spatial attention learned in the video encoder may not be compelling enough without text guidance. To solve these problems, we propose a Stay-in-Grid video CAPtioning method SGCAP, which makes full use of the grid-level spatial features and consists of a Bilinear Sequential Attention Encoder (BSAE) and a Cross-modal Sequential Attention Decoder (CSAD). The former explores and retains fully grid-level discriminative representations in the video encoder, while the latter performs the late spatial aggregation in the decoder to attend to the most relevant regions with the supervision of the input words. Experimental results demonstrate the effectiveness of our method on three public datasets, showing its superior performance over multiple state-of-the-art video captioning models. Source codes and the pre-trained models will be made available to the public. Mingkang Tang, Zhanyu Wang, Zhaoyang Zeng, Xiu Li 0001, Luping Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Single-View 3D Mesh Reconstruction for Seen and Unseen CategoriesabstractSingle-view 3D object reconstruction is a fundamental and challenging computer vision task that aims at recovering 3D shapes from single-view RGB images. Most existing deep learning based reconstruction methods are trained and evaluated on the same categories, and they cannot work well when handling objects from novel categories that are not seen during training. Focusing on this issue, this paper tackles Single-view 3D Mesh Reconstruction, to study the model generalization on unseen categories and encourage models to reconstruct objects literally. Specifically, we propose an end-to-end two-stage network, GenMesh, to break the category boundaries in reconstruction. Firstly, we factorize the complicated image-to-mesh mapping into two simpler mappings, i.e., image-to-point mapping and point-to-mesh mapping, while the latter is mainly a geometric problem and less dependent on object categories. Secondly, we devise a local feature sampling strategy in 2D and 3D feature spaces to capture the local geometry shared across objects to enhance model generalization. Thirdly, apart from the traditional point-to-point supervision, we introduce a multi-view silhouette loss to supervise the surface generation process, which provides additional regularization and further relieves the overfitting problem. The experimental results show that our method significantly outperforms the existing works on the ShapeNet and Pix3D under different scenarios and various metrics, especially for novel objects. Xianghui Yang, Guosheng Lin, Luping Zhou |
IEEE Trans. Image Process. | 3 |
| 2023 | Bridging Synthetic and Real Images: A Transferable and Multiple Consistency Aided Fundus Image Enhancement FrameworkabstractDeep learning based image enhancement models have largely improved the readability of fundus images in order to decrease the uncertainty of clinical observations and the risk of misdiagnosis. However, due to the difficulty of acquiring paired real fundus images at different qualities, most existing methods have to adopt synthetic image pairs as training data. The domain shift between the synthetic and the real images inevitably hinders the generalization of such models on clinical data. In this work, we propose an end-to-end optimized teacher-student framework to simultaneously conduct image enhancement and domain adaptation. The student network uses synthetic pairs for supervised enhancement, and regularizes the enhancement model to reduce domain-shift by enforcing teacher-student prediction consistency on the real fundus images without relying on enhanced ground-truth. Moreover, we also propose a novel multi-stage multi-attention guided enhancement network (MAGE-Net) as the backbones of our teacher and student network. Our MAGE-Net utilizes multi-stage enhancement module and retinal structure preservation module to progressively integrate the multi-scale features and simultaneously preserve the retinal structures for better fundus image quality enhancement. Comprehensive experiments on both real and synthetic datasets demonstrate that our framework outperforms the baseline approaches. Moreover, our method also benefits the downstream clinical tasks. Erjian Guo, Huazhu Fu, Luping Zhou, Dong Xu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Breast Fibroglandular Tissue Segmentation for Automated BPE Quantification With Iterative Cycle-Consistent Semi-Supervised LearningabstractBackground Parenchymal Enhancement (BPE) quantification in Dynamic Contrast-Enhanced Magnetic Resonance Imaging (DCE-MRI) plays a pivotal role in clinical breast cancer diagnosis and prognosis. However, the emerging deep learning-based breast fibroglandular tissue segmentation, a crucial step in automated BPE quantification, often suffers from limited training samples with accurate annotations. To address this challenge, we propose a novel iterative cycle-consistent semi-supervised framework to leverage segmentation performance by using a large amount of paired pre-/post-contrast images without annotations. Specifically, we design the reconstruction network, cascaded with the segmentation network, to learn a mapping from the pre-contrast images and segmentation predictions to the post-contrast images. Thus, we can implicitly use the reconstruction task to explore the inter-relationship between these two-phase images, which in return guides the segmentation task. Moreover, the reconstructed post-contrast images across multiple auto-context modeling-based iterations can be viewed as new augmentations, facilitating cycle-consistent constraints across each segmentation output. Extensive experiments on two datasets with various data distributions show great segmentation and BPE quantification accuracy compared with other state-of-the-art semi-supervised methods. Importantly, our method achieves 11.80 times of quantification accuracy improvement along with 10 times faster, compared with clinical physicians, demonstrating its potential for automated BPE quantification. The code is available at https://github.com/ZhangJD-ong/Iterative-Cycle-consistent-Semi-supervised-Learning-for-fibroglandular-tissue-segmentation. Zhiming Cui 0001, Luping Zhou, Yiqun Sun, Zhenhui Li, Zaiyi Liu, Dinggang Shen |
IEEE Trans. Medical Imaging | 3 |
| 2022 | LaSSL: Label-Guided Self-Training for Semi-supervised LearningabstractThe key to semi-supervised learning (SSL) is to explore adequate information to leverage the unlabeled data. Current dominant approaches aim to generate pseudo-labels on weakly augmented instances and train models on their corresponding strongly augmented variants with high-confidence results. However, such methods are limited in excluding samples with low-confidence pseudo-labels and under-utilization of the label information. In this paper, we emphasize the cruciality of the label information and propose a Label-guided Self-training approach to Semi-supervised Learning (LaSSL), which improves pseudo-label generations from two mutually boosted strategies. First, with the ground-truth labels and iteratively-polished pseudo-labels, we explore instance relations among all samples and then minimize a class-aware contrastive loss to learn discriminative feature representations that make same-class samples gathered and different-class samples scattered. Second, on top of improved feature representations, we propagate the label information to the unlabeled samples across the potential data manifold at the feature-embedding level, which can further improve the labelling of samples with reference to their neighbours. These two strategies are seamlessly integrated and mutually promoted across the whole training process. We evaluate LaSSL on several classification benchmarks under partially labeled settings and demonstrate its superiority over the state-of-the-art approaches. Zhen Zhao 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001 |
AAAI | 2 |
| 2022 | DC-SSL: Addressing Mismatched Class Distribution in Semi-supervised LearningabstractConsistency-based Semi-supervised learning (SSL) has achieved promising performance recently. However, the success largely depends on the assumption that the labeled and unlabeled data share an identical class distribution, which is hard to meet in real practice. The distribution mismatch between the labeled and unlabeled sets can cause severe bias in the pseudo-labels of SSL, resulting in significant performance degradation. To bridge this gap, we put forward a new SSL learning framework, named Distribution Consistency SSL (DC-SSL), which rectifies the pseudolabels from a distribution perspective. The basic idea is to directly estimate a reference class distribution (RCD), which is regarded as a surrogate of the ground truth class distribution about the unlabeled data, and then improve the pseudo-labels by encouraging the predicted class distribution (PCD) of the unlabeled data to approach RCD gradually. To this end, this paper revisits the Exponentially Moving Average (EMA) model and utilizes it to estimate RCD in an iteratively improved manner, which is achieved with a momentum-update scheme throughout the training procedure. On top of this, two strategies are proposed for RCD to rectify the pseudo-label prediction, respectively. They correspond to an efficient training-free scheme and a training-based alternative that generates more accurate and reliable predictions. DC-SSL is evaluated on multiple SSL benchmarks and demonstrates remarkable performance improvement over competitive methods under matched- and mismatched-distribution scenarios. Zhen Zhao 0001, Luping Zhou, Yue Duan, Lei Wang 0001, Lei Qi 0001, Yinghuan Shi |
CVPR | 2 |
| 2022 | RDA: Reciprocal Distribution Alignment for Robust Semi-supervised Learning
Yue Duan, Lei Qi 0001, Lei Wang 0001, Luping Zhou, Yinghuan Shi |
ECCV (30) | 4 |
| 2022 | A Medical Semantic-Assisted Transformer for Radiographic Report Generation
Zhanyu Wang, Mingkang Tang, Lei Wang 0001, Xiu Li 0001, Luping Zhou |
MICCAI (3) | 5 |
| 2022 | 3D CVT-GAN: A 3D Convolutional Vision Transformer-GAN for PET Reconstruction
Pinxian Zeng, Luping Zhou, Chen Zu, Xinyi Zeng, Zhengyang Jiao, Xi Wu 0004, Jiliu Zhou, Dinggang Shen, Yan Wang 0015 |
MICCAI (6) | 2 |
| 2022 | Improving Barely Supervised Learning by Discriminating Unlabeled Samples with Super-ClassabstractIn semi-supervised learning (SSL), a common practice is to learn consistent information from unlabeled data and discriminative information from labeled data to ensure both the immutability and the separability of the classification model. Existing SSL methods suffer from failures in barely-supervised learning (BSL), where only one or two labels per class are available, as the insufficient labels cause the discriminative information being difficult or even infeasible to learn. To bridge this gap, we investigate a simple yet effective way to leverage unlabeled samples for discriminative learning, and propose a novel discriminative information learning module to benefit model training. Specifically, we formulate the learning objective of discriminative information at the super-class level and dynamically assign different classes into different super-classes based on model performance improvement. On top of this on-the-fly process, we further propose a distribution-based loss to learn discriminative information by utilizing the similarity relationship between samples and super-classes. It encourages the unlabeled samples to stay closer to the distribution of their corresponding super-class than those of others. Such a constraint is softer than the direct assignment of pseudo labels, while the latter could be very noisy in BSL. We compare our method with state-of-the-art SSL and BSL methods through extensive experiments on standard SSL benchmarks. Our method can achieve superior results, \eg, an average accuracy of 76.76\% on CIFAR-10 with merely 1 label per class. Guan Gui 0002, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
NeurIPS | 4 |
| 2022 | An Efficient Semi-Supervised Framework with Multi-Task and Curriculum Learning for Medical Image SegmentationabstractA practical problem in supervised deep learning for medical image segmentation is the lack of labeled data which is expensive and time-consuming to acquire. In contrast, there is a considerable amount of unlabeled data available in the clinic. To make better use of the unlabeled data and improve the generalization on limited labeled data, in this paper, a novel semi-supervised segmentation method via multi-task curriculum learning is presented. Here, curriculum learning means that when training the network, simpler knowledge is preferentially learned to assist the learning of more difficult knowledge. Concretely, our framework consists of a main segmentation task and two auxiliary tasks, i.e. the feature regression task and target detection task. The two auxiliary tasks predict some relatively simpler image-level attributes and bounding boxes as the pseudo labels for the main segmentation task, enforcing the pixel-level segmentation result to match the distribution of these pseudo labels. In addition, to solve the problem of class imbalance in the images, a bounding-box-based attention (BBA) module is embedded, enabling the segmentation network to concern more about the target region rather than the background. Furthermore, to alleviate the adverse effects caused by the possible deviation of pseudo labels, error tolerance mechanisms are also adopted in the auxiliary tasks, including inequality constraint and bounding-box amplification. Our method is validated on ACDC2017 and PROMISE12 datasets. Experimental results demonstrate that compared with the full supervision method and state-of-the-art semi-supervised methods, our method yields a much better segmentation performance on a small labeled dataset. Code is available at https://github.com/DeepMedLab/MTCL. Kaiping Wang, Yan Wang 0015, Bo Zhan, Chen Zu, Xi Wu 0004, Jiliu Zhou, Dong Nie, Luping Zhou |
Int. J. Neural Syst. | 9 |
| 2022 | D2FE-GAN: Decoupled dual feature extraction based GAN for MRI image synthesis
Bo Zhan, Luping Zhou, Xi Wu 0004, Yi-Fei Pu, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
Knowl. Based Syst. | 2 |
| 2022 | Adaptive rectification based adversarial network with spectrum constraint for high-quality PET image synthesis
Yanmei Luo, Luping Zhou, Bo Zhan, Fei-Yue Wang 0001, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
Medical Image Anal. | 2 |
| 2022 | Semi-supervised medical image segmentation via a tripled-uncertainty guided mean teacher model with contrastive learning
Kaiping Wang, Bo Zhan, Chen Zu, Xi Wu 0004, Jiliu Zhou, Luping Zhou, Yan Wang 0015 |
Medical Image Anal. | 6 |
| 2022 | FOD-Net: A deep learning method for fiber orientation distribution angular super resolution
Jinglei Lv, He Wang 0016, Luping Zhou, Michael Barnett 0006, Fernando Calamante, Chenyu Wang 0001 |
Medical Image Anal. | 4 |
| 2022 | ASMFS: Adaptive-similarity-based multi-modality feature selection for classification of Alzheimer's disease
Yuang Shi, Chen Zu, Luping Zhou, Lei Wang 0001, Xi Wu 0004, Jiliu Zhou, Daoqiang Zhang, Yan Wang 0015 |
Pattern Recognit. | 4 |
| 2022 | Action Recognition With Motion Diversification and Dynamic SelectionabstractMotion modeling is crucial in modern action recognition methods. As motion dynamics like moving tempos and action amplitude may vary a lot in different video clips, it poses great challenge on adaptively covering proper motion information. To address this issue, we introduce a Motion Diversification and Selection (MoDS) module to generate diversified spatio-temporal motion features and then select the suitable motion representation dynamically for categorizing the input video. To be specific, we first propose a spatio-temporal motion generation (StMG) module to construct a bank of diversified motion features with varying spatial neighborhood and time range. Then, a dynamic motion selection (DMS) module is leveraged to choose the most discriminative motion feature both spatially and temporally from the feature bank. As a result, our proposed method can make full use of the diversified spatio-temporal motion information, while maintaining computational efficiency at the inference stage. Extensive experiments on five widely-used benchmarks, demonstrate the effectiveness of the method and we achieve state-of-the-art performance on Something-Something V1 & V2 that are of large motion variation. Peiqin Zhuang, Luping Zhou, Lei Bai 0001, Ding Liang, Zhiyong Wang 0001, Yali Wang 0001, Wanli Ouyang |
IEEE Trans. Image Process. | 4 |
| 2022 | The Bounds of Improvements Toward Real-Time Forecast of Multi-Scenario Train DelaysabstractDifferent from the existing train delay studies that had strived to explore sophisticated algorithms, this paper focuses on finding the bound of improvements on predicting multi-scenario train delays with different machine learning methods. Motivated by the observation of deep learning methods failing to improve the prediction performance if the delay occurs rarely, we present a novel augmented machine learning approach to improve the overall prediction accuracy further. Our solution proposes a rule-driven automation (RDA) method, including a delay status labeling (DSL) algorithm, and the resilience of section (RSE) and resilience of station (RST) indicators to generate the forecast for train delays. The experiment results demonstrate that the Random Forest based implementation of our RDA method (RF-RDA) can significantly improve the generalization ability of multivariate multi-step forecast models for multi-scenario train delay prediction. The proposed solution surpasses state-of-art baselines based on real-world traffic datasets, which treat various real-time delays differently. Even when the predictability of conventional deep learning methods decreases, the performance of our method is still acceptable for practical use to provide accurate forecasts. Jianqing Wu 0002, Yihui Wang 0001, Bo Du 0004, Qiang Wu 0010, Yanlong Zhai, Jun Shen 0001, Luping Zhou, Wei Wei 0006, Qingguo Zhou |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2022 | Automated Radiographic Report Generation Purely on Transformer: A Multicriteria Supervised ApproachabstractAutomated radiographic report generation is challenging in at least two aspects. First, medical images are very similar to each other and the visual differences of clinic importance are often fine-grained. Second, the disease-related words may be submerged by many similar sentences describing the common content of the images, causing the abnormal to be misinterpreted as the normal in the worst case. To tackle these challenges, this paper proposes a pure transformer-based framework to jointly enforce better visual-textual alignment, multi-label diagnostic classification, and word importance weighting, to facilitate report generation. To the best of our knowledge, this is the first pure transformer-based framework for medical report generation, which enjoys the capacity of transformer in learning long range dependencies for both image regions and sentence words. Specifically, for the first challenge, we design a novel mechanism to embed an auxiliary image-text matching objective into the transformer's encoder-decoder structure, so that better correlated image and text features could be learned to help a report to discriminate similar images. For the second challenge, we integrate an additional multi-label classification task into our framework to guide the model in making correct diagnostic predictions. Also, a term-weighting scheme is proposed to reflect the importance of words for training so that our model would not miss key discriminative information. Our work achieves promising performance over the state-of-the-arts on two benchmark datasets, including the largest dataset MIMIC-CXR. Zhanyu Wang, Lei Wang 0001, Xiu Li 0001, Luping Zhou |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Diffusion Kernel Attention Network for Brain Disorder ClassificationabstractConstructing and analyzing functional brain networks (FBN) has become a promising approach to brain disorder classification. However, the conventional successive construct-and-analyze process would limit the performance due to the lack of interactions and adaptivity among the subtasks in the process. Recently, Transformer has demonstrated remarkable performance in various tasks, attributing to its effective attention mechanism in modeling complex feature relationships. In this paper, for the first time, we develop Transformer for integrated FBN modeling, analysis and brain disorder classification with rs-fMRI data by proposing a Diffusion Kernel Attention Network to address the specific challenges. Specifically, directly applying Transformer does not necessarily admit optimal performance in this task due to its extensive parameters in the attention module against the limited training samples usually available. Looking into this issue, we propose to use kernel attention to replace the original dot-product attention module in Transformer. This significantly reduces the number of parameters to train and thus alleviates the issue of small sample while introducing a non-linear attention mechanism to model complex functional connections. Another limit of Transformer for FBN applications is that it only considers pair-wise interactions between directly connected brain regions but ignores the important indirect connections. Therefore, we further explore diffusion process over the kernel attention to incorporate wider interactions among indirectly connected brain regions. Extensive experimental study is conducted on ADHD-200 data set for ADHD classification and on ADNI data set for Alzheimer's disease classification, and the results demonstrate the superior performance of the proposed method over the competing methods. Jianjia Zhang, Luping Zhou, Lei Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2021 | A Self-Boosting Framework for Automated Radiographic Report GenerationabstractAutomated radiographic report generation is a challenging task since it requires to generate paragraphs describing fine-grained visual differences of cases, especially for those between the diseased and the healthy. Existing image captioning methods commonly target at generic images, and lack mechanism to meet this requirement. To bridge this gap, in this paper, we propose a self-boosting framework that improves radiographic report generation based on the cooperation of the main task of report generation and an auxiliary task of image-text matching. The two tasks are built as the two branches of a network model and influence each other in a cooperative way. On one hand, the image-text matching branch helps to learn highly text-correlated visual features for the report generation branch to output high quality reports. On the other hand, the improved reports produced by the report generation branch provide additional harder samples for the image-text matching branch and enforce the latter to improve itself by learning better visual and text feature representations. This, in turn, helps improve the report generation branch again. These two branches are jointly trained to help improve each other iteratively and progressively, so that the whole model is self-boosted without requiring external resources. Experimental results demonstrate the effectiveness of our method on two public datasets, showing its superior performance over multiple state-of-the-art image captioning and medical report generation methods. Zhanyu Wang, Luping Zhou, Lei Wang 0001, Xiu Li 0001 |
CVPR | 2 |
| 2021 | 3D Transformer-GAN for High-Quality PET Reconstruction
Yanmei Luo, Yan Wang 0015, Chen Zu, Bo Zhan, Xi Wu 0004, Jiliu Zhou, Dinggang Shen, Luping Zhou |
MICCAI (6) | 8 |
| 2021 | Tripled-Uncertainty Guided Mean Teacher Model for Semi-supervised Medical Image Segmentation
Kaiping Wang, Bo Zhan, Chen Zu, Xi Wu 0004, Jiliu Zhou, Luping Zhou, Yan Wang 0015 |
MICCAI (2) | 6 |
| 2021 | Few-shot Unsupervised Domain Adaptation with Image-to-Class Sparse Similarity EncodingabstractThis paper investigates a valuable setting called few-shot unsupervised domain adaptation (FS-UDA), which has not been sufficiently studied in the literature. In this setting, the source domain data are labelled, but with few-shot per category, while the target domain data are unlabelled. To address the FS-UDA setting, we develop a general UDA model to solve the following two key issues: the few-shot labeled data per category and the domain adaptation between support and query sets. Our model is general in that once trained it will be able to be applied to various FS-UDA tasks from the same source and target domains. Inspired by the recent local descriptor based few-shot learning (FSL), our general UDA model is fully built upon local descriptors (LDs) for image classification and domain adaptation. By proposing a novel concept called similarity patterns (SPs), our model not only effectively considers the spatial relationship of LDs that was ignored in previous FSL methods, but also makes the learned image similarity better serve the required domain alignment. Specifically, we propose a novel IMage-to-class sparse Similarity Encoding (IMSE) method. It learns SPs to extract the local discriminative information for classification and meanwhile aligns the covariance matrix of the SPs for domain adaptation. Also, domain adversarial training and multi-scale local feature matching are performed upon LDs. Extensive experiments conducted on a multi-domain benchmark dataset DomainNet demonstrates the state-of-the-art performance of our IMSE for the novel setting of FS-UDA. In addition, for FSL, our IMSE can also show better performance than most of recent FSL methods on miniImageNet. Shengqi Huang, Wanqi Yang, Lei Wang 0001, Luping Zhou, Ming Yang 0014 |
ACM Multimedia | 4 |
| 2021 | Interactive medical image segmentation via a point-based interaction
Jian Zhang 0090, Yinghuan Shi, Jinquan Sun, Lei Wang 0001, Luping Zhou, Yang Gao 0001, Dinggang Shen |
Artif. Intell. Medicine | 5 |
| 2021 | Beyond Covariance: SICE and Kernel Based Visual Feature Representation
Jianjia Zhang, Lei Wang 0001, Luping Zhou, Wanqing Li 0001 |
Int. J. Comput. Vis. | 3 |
| 2021 | DA-DSUnet: Dual Attention-based Dense SU-net for automatic head-and-neck tumor segmentation in MRI images
Pin Tang, Chen Zu, Xingchen Peng, Jianghong Xiao, Xi Wu 0004, Jiliu Zhou, Luping Zhou, Yan Wang 0015 |
Neurocomputing | 9 |
| 2021 | Unsupervised brain tumor segmentation using a symmetric-driven adversarial network
Xinheng Wu, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Luping Zhou, Jinman Kim |
Neurocomputing | 5 |
| 2021 | Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action LocalizationabstractSpatio-temporal action localization consists of three levels of tasks: spatial localization, action classification, and temporal localization. In this work, we propose a new progressive cross-stream cooperation (PCSC) framework that improves all three tasks above. The basic idea is to utilize both spatial region (resp., temporal segment proposals) and features from one stream (i.e., the Flow/RGB stream) to help another stream (i.e., the RGB/Flow stream) to iteratively generate better bounding boxes in the spatial domain (resp., temporal segments in the temporal domain). In this way, not only the actions could be more accurately localized both spatially and temporally, but also the action classes could be predicted more precisely. Specifically, we first combine the latest region proposals (for spatial detection) or segment proposals (for temporal localization) from both streams to form a larger set of labelled training samples to help learn better action detection or segment detection models. Second, to learn better representations, we also propose a new message passing approach to pass information from one stream to another stream, which also leads to better action detection and segment detection models. By first using our newly proposed PCSC framework for spatial localization at the frame-level and then applying our temporal PCSC framework for temporal localization at the tube-level, the action localization results are progressively improved at both the frame level and the video level. Comprehensive experiments on two benchmark datasets UCF-101-24 and J-HMDB demonstrate the effectiveness of our newly proposed approaches for spatio-temporal action localization in realistic scenarios. Dong Xu 0001, Luping Zhou, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Improving Weakly Supervised Temporal Action Localization by Exploiting Multi-Resolution Information in Temporal DomainabstractWeakly supervised temporal action localization is a challenging task as only the video-level annotation is available during the training process. To address this problem, we propose a two-stage approach to generate high-quality frame-level pseudo labels by fully exploiting multi-resolution information in the temporal domain and complementary information between the appearance (i.e., RGB) and motion (i.e., optical flow) streams. In the first stage, we propose an Initial Label Generation (ILG) module to generate reliable initial frame-level pseudo labels. Specifically, in this newly proposed module, we exploit temporal multi-resolution consistency and cross-stream consistency to generate high quality class activation sequences (CASs), which consist of a number of sequences with each sequence measuring how likely each video frame belongs to one specific action class. In the second stage, we propose a Progressive Temporal Label Refinement (PTLR) framework to iteratively refine the pseudo labels, in which we use a set of selected frames with highly confident pseudo labels to progressively train two networks and better predict action class scores at each frame. Specifically, in our newly proposed PTLR framework, two networks called Network-OTS and Network-RTS, which are respectively used to generate CASs for the original temporal scale and the reduced temporal scales, are used as two streams (i.e., the OTS stream and the RTS stream) to refine the pseudo labels in turn. By this way, multi-resolution information in the temporal domain is exchanged at the pseudo label level, and our work can help improve each network/stream by exploiting the refined pseudo labels from another network/stream. Comprehensive experiments on two benchmark datasets THUMOS14 and ActivityNet v1.3 demonstrate the effectiveness of our newly proposed method for weakly supervised temporal action localization. Dong Xu 0001, Luping Zhou, Wanli Ouyang |
IEEE Trans. Image Process. | 3 |
| 2021 | SA-LuT-Nets: Learning Sample-Adaptive Intensity Lookup Tables for Brain Tumor SegmentationabstractIn clinics, the information about the appearance and location of brain tumors is essential to assist doctors in diagnosis and treatment. Automatic brain tumor segmentation on the images acquired by magnetic resonance imaging (MRI) is a common way to attain this information. However, MR images are not quantitative and can exhibit significant variation in signal depending on a range of factors, which increases the difficulty of training an automatic segmentation network and applying it to new MR images. To deal with this issue, this paper proposes to learn a sample-adaptive intensity lookup table (LuT) that dynamically transforms the intensity contrast of each input MR image to adapt to the following segmentation task. Specifically, the proposed deep SA-LuT-Net framework consists of a LuT module and a segmentation module, trained in an end-to-end manner: the LuT module learns a sample-specific nonlinear intensity mapping function through communication with the segmentation module, aiming at improving the final segmentation performance. In order to make the LuT learning sample-adaptive, we parameterize the intensity mapping function by exploring two families of non-linear functions (i.e., piece-wise linear and power functions) and predict the function parameters for each given sample. These sample-specific parameters make the intensity mapping adaptive to samples. We develop our SA-LuT-Nets separately based on two backbone networks for segmentation, i.e., DMFNet and the modified 3D Unet, and validate them on BRATS2018 and BRATS2019 datasets for brain tumor segmentation. Our experimental results clearly demonstrate the superior performance of the proposed SA-LuT-Nets using either single or multiple MR modalities. It not only significantly improves the two baselines (DMFNet and the modified 3D Unet), but also wins a set of state-of-the-art segmentation methods. Moreover, we show that, the LuTs learnt using one segmentation model could also be applied to improving the performance of another segmentation model, indicating the general segmentation information captured by LuTs. Biting Yu, Luping Zhou, Lei Wang 0001, Wanqi Yang, Ming Yang 0014, Pierrick Bourgeat, Jurgen Fripp |
IEEE Trans. Medical Imaging | 2 |
| 2021 | Dense Video Captioning Using Graph-Based Sentence SummarizationabstractRecently, dense video captioning has made attractive progress in detecting and captioning all events in a long untrimmed video. Despite promising results were achieved, most existing methods do not sufficiently explore the scene evolution within an event temporal proposal for captioning, and therefore perform less satisfactorily when the scenes and objects change over a relatively long proposal. To address this problem, we propose a graph-based partition-and-summarization (GPaS) framework for dense video captioning within two stages. For the “partition” stage, a whole event proposal is split into short video segments for captioning at a finer level. For the “summarization” stage, the generated sentences carrying rich description information for each segment are summarized into one sentence to describe the whole event. We particularly focus on the “summarization” stage, and propose a framework that effectively exploits the relationship between semantic words for summarization. We achieve this goal by treating semantic words as the nodes in a graph and learning their interactions by coupling Graph Convolutional Network (GCN) and Long Short Term Memory (LSTM), with the aid of visual cues. Two schemes of GCN-LSTM Interaction (GLI) modules are proposed for seamless integration of GCN and LSTM. The effectiveness of our approach is demonstrated via an extensive comparison with the state-of-the-arts methods on the two benchmarks ActivityNet Captions dataset and YouCook II dataset. Zhiwang Zhang, Dong Xu 0001, Wanli Ouyang, Luping Zhou |
IEEE Trans. Multim. | 4 |
| 2020 | BriNet: Towards Bridging the Intra-class and Inter-class Gaps in One-Shot Segmentation
Xianghui Yang, Bairun Wang, Xinchi Zhou, Kaige Chen, Shuai Yi, Wanli Ouyang, Luping Zhou |
BMVC | 7 |
| 2020 | ReDro: Efficiently Learning Large-Sized SPD Visual Representation
Saimunur Rahman, Lei Wang 0001, Changming Sun, Luping Zhou |
ECCV (15) | 4 |
| 2020 | Learning Sample-Adaptive Intensity Lookup Table for Brain Tumor Segmentation
Biting Yu, Luping Zhou, Lei Wang 0001, Wanqi Yang, Ming Yang 0014, Pierrick Bourgeat, Jurgen Fripp |
MICCAI (4) | 2 |
| 2020 | Improving Auto-Augment via Augmentation-Wise Weight SharingabstractThe recent progress on automatically searching augmentation policies has boosted the performance substantially for various tasks. A key component of automatic augmentation search is the evaluation process for a particular augmentation policy, which is utilized to return reward and usually runs thousands of times. A plain evaluation process, which includes full model training and validation, would be time-consuming. To achieve efficiency, many choose to sacrifice evaluation reliability for speed. In this paper, we dive into the dynamics of augmented training of the model. This inspires us to design a powerful and efficient proxy task based on the Augmentation-Wise Weight Sharing (AWS) to form a fast yet accurate evaluation process in an elegant way. Comprehensive analysis verifies the superiority of this approach in terms of effectiveness and efficiency. The augmentation policies found by our method achieve superior accuracies compared with existing auto-augmentation search methods. On CIFAR-10, we achieve a top-1 error rate of 1.24%, which is currently the best performing single model without extra training data. On ImageNet, we get a top-1 error rate of 20.36% for ResNet-50, which leads to 3.34% absolute error rate reduction over the baseline augmentation. Keyu Tian, Chen Lin 0003, Ming Sun 0008, Luping Zhou, Wanli Ouyang |
NeurIPS | 4 |
| 2020 | Deep learning based HEp-2 image classification: A comprehensive review
Saimunur Rahman, Lei Wang 0001, Changming Sun, Luping Zhou |
Medical Image Anal. | 4 |
| 2020 | Coherent Pattern in Multi-Layer Brain Networks: Application to Epilepsy IdentificationabstractCurrently, how to conjointly fuse structural connectivity (SC) and functional connectivity (FC) for identifying brain diseases is a hot topic in the area of brain network analysis. Most of the existing works combine two types of connectivity in decision level, thus ignoring the underlying relationship between SC and FC. To solve this problem, in this paper, we model the brain network as the multi-layer network formed by the SC and FC, and then propose a coherent pattern to represent structural information of the multi-layer network for the brain disease identification. The proposed coherent pattern consists of a paired-subgraph extracted from the FC and SC within the same node-set. Compared with the previous methods, this coherent pattern not only describes the connectivity information of both SC and FC by subgraphs at each layer, but also reflects their intrinsic relationship by the co-occurrence pattern of the paired-subgraph. Based on this coherent pattern, we further develop a framework for identifying brain diseases. Specifically, we first construct multi-layer networks by using SC and FC for each subject and then mine coherent patterns that frequently appear in each group. Next, we select the discriminative coherent pattern from these frequent coherent patterns according to their frequency of occurrence. Finally, we construct a feature matrix for each subject based on the binary indicator vector and then use the support vector machine (SVM) as its classifier. Experimental results on real epilepsy datasets demonstrate that our method outperforms several state-of-the-art approaches in the tasks of brain disease classification. Jiashuang Huang, Qi Zhu 0001, Luping Zhou, Daoqiang Zhang |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Epileptic Seizure Classification With Symmetric and Hybrid Bilinear ModelsabstractEpilepsy affects nearly [Formula: see text] of the global population, of which two thirds can be treated by anti-epileptic drugs and a much lower percentage by surgery. Diagnostic procedures for epilepsy and monitoring are highly specialized and labour-intensive. The accuracy of the diagnosis is also complicated by overlapping medical symptoms, varying levels of experience and inter-observer variability among clinical professions. This paper proposes a novel hybrid bilinear deep learning network with an application in the clinical procedures of epilepsy classification diagnosis, where the use of surface electroencephalogram (sEEG) and audiovisual monitoring is standard practice. Hybrid bilinear models based on two types of feature extractors, namely Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs), are trained using Short-Time Fourier Transform (STFT) of one-second sEEG. In the proposed hybrid models, CNNs extract spatio-temporal patterns, while RNNs focus on the characteristics of temporal dynamics in relatively longer intervals given the same input data. Second-order features, based on interactions between these spatio-temporal features are further explored by bilinear pooling and used for epilepsy classification. Our proposed methods obtain an F1-score of [Formula: see text] on the Temple University Hospital Seizure Corpus and [Formula: see text] on the EPILEPSIAE dataset, comparing favourably to existing benchmarks for sEEG-based seizure type classification. The open-source implementation of this study is available at https://github.com/NeuroSyd/Epileptic-Seizure-Classification. Tennison Liu, Nhan Duy Truong, Armin Nikpour, Luping Zhou, Omid Kavehei |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Attention-Diffusion-Bilinear Neural Network for Brain Network AnalysisabstractBrain network provides essential insights in diagnosing many brain disorders. Integrative analysis of multiple types of connectivity, e.g, functional connectivity (FC) and structural connectivity (SC), can take advantage of their complementary information and therefore may help to identify patients. However, traditional brain network methods usually focus on either FC or SC for describing node interactions and only consider the interaction between paired network nodes. To tackle this problem, in this paper, we propose an Attention-Diffusion-Bilinear Neural Network (ADB-NN) framework for brain network analysis, which is trained in an end-to-end manner. The proposed network seamlessly couples FC and SC to learn wider node interactions and generates a joint representation of FC and SC for diagnosis. Specifically, a brain network (graph) is first defined, where each node corresponding to a brain region is governed by the features of brain activities (i.e., FC) extracted from functional magnetic resonance imaging (fMRI), and the presence of edges is determined by neural fiber physical connections (i.e., SC) extracted from Diffusion Tensor Imaging (DTI). Based on this graph, we train two Attention-Diffusion-Bilinear (ADB) modules jointly. In each module, an attention model is utilized to automatically learn the strength of node interactions. This information further guides a diffusion process that generates new node representations by considering the influence from other nodes as well. After that, the second-order statistics of these node representations are extracted by bilinear pooling to form connectivity-based features for disease prediction. The two ADB modules correspond to the one-step and two-step diffusion, respectively. Experiments on a real epilepsy dataset demonstrate the effectiveness and advantages of our proposed method. Jiashuang Huang, Luping Zhou, Lei Wang 0001, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Sample-Adaptive GANs: Linking Global and Local Mappings for Cross-Modality MR Image SynthesisabstractGenerative adversarial network (GAN) has been widely explored for cross-modality medical image synthesis. The existing GAN models usually adversarially learn a global sample space mapping from the source-modality to the target-modality and then indiscriminately apply this mapping to all samples in the whole space for prediction. However, due to the scarcity of training samples in contrast to the complicated nature of medical image synthesis, learning a single global sample space mapping that is "optimal" to all samples is very challenging, if not intractable. To address this issue, this paper proposes sample-adaptive GAN models, which not only cater for the global sample space mapping between the source- and the target-modalities but also explore the local space around each given sample to extract its unique characteristic. Specifically, the proposed sample-adaptive GANs decompose the entire learning model into two cooperative paths. The baseline path learns a common GAN model by fitting all the training samples as usual for the global sample space mapping. The new sample-adaptive path additionally models each sample by learning its relationship with its neighboring training samples and using the target-modality features of these training samples as auxiliary information for synthesis. Enhanced by this sample-adaptive path, the proposed sample-adaptive GANs are able to flexibly adjust themselves to different samples, and therefore optimize the synthesis performance. Our models have been verified on three cross-modality MR image synthesis tasks from two public datasets, and they significantly outperform the state-of-the-art methods in comparison. Moreover, the experiment also indicates that our sample-adaptive strategy could be utilized to improve various backbone GAN models. It complements the existing GANs models and can be readily integrated when needed. Biting Yu, Luping Zhou, Lei Wang 0001, Yinghuan Shi, Jurgen Fripp, Pierrick Bourgeat |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Improving Action Localization by Progressive Cross-Stream CooperationabstractSpatio-temporal action localization consists of three levels of tasks: spatial localization, action classification, and temporal segmentation. In this work, we propose a new Progressive Cross-stream Cooperation (PCSC) framework to iterative improve action localization results and generate better bounding boxes for one stream (i.e., Flow/RGB) by leveraging both region proposals and features from another stream (i.e., RGB/Flow) in an iterative fashion. Specifically, we first generate a larger set of region proposals by combining the latest region proposals from both streams, from which we can readily obtain a larger set of labelled training samples to help learn better action detection models. Second, we also propose a new message passing approach to pass information from one stream to another stream in order to learn better representations, which also leads to better action detection models. As a result, our iterative framework progressively improves action localization results at the frame level. To improve action localization results at the video level, we additionally propose a new strategy to train class-specific actionness detectors for better temporal segmentation, which can be readily learnt by using the training samples around temporal boundaries. Comprehensive experiments on two benchmark datasets UCF-101-24 and J-HMDB demonstrate the effectiveness of our newly proposed approaches for spatio-temporal action localization in realistic scenarios. Wanli Ouyang, Luping Zhou, Dong Xu 0001 |
CVPR | 3 |
| 2019 | A Novel Unsupervised Camera-Aware Domain Adaptation Framework for Person Re-IdentificationabstractUnsupervised cross-domain person re-identification (Re-ID) faces two key issues. One is the data distribution discrepancy between source and target domains, and the other is the lack of discriminative information in target domain. From the perspective of representation learning, this paper proposes a novel end-to-end deep domain adaptation framework to address them. For the first issue, we highlight the presence of camera-level sub-domains as a unique characteristic in person Re-ID, and develop a “camera-aware” domain adaptation method via adversarial learning. With this method, the learned representation reduces distribution discrepancy not only between source and target domains but also across all cameras. For the second issue, we exploit the temporal continuity in each camera of target domain to create discriminative information. This is implemented by dynamically generating online triplets within each batch, in order to maximally take advantage of the steadily improved representation in training process. Together, the above two methods give rise to a new unsupervised domain adaptation framework for person Re-ID. Extensive experiments and ablation studies conducted on benchmark datasets demonstrate its superiority and interesting properties. Lei Qi 0001, Lei Wang 0001, Jing Huo, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
ICCV | 4 |
| 2019 | Miss Detection vs. False Alarm: Adversarial Learning for Small Object Segmentation in Infrared ImagesabstractA key challenge of infrared small object segmentation (ISOS) is to balance miss detection (MD) and false alarm (FA). This usually needs ``opposite'' strategies to suppress the two terms, and has not been well resolved in the literature. In this paper, we propose a deep adversarial learning framework to improve this situation. Departing from the tradition of jointly reducing MD and FA via a single objective, we decompose this difficult task into two sub-tasks handled by two models trained adversarially, with each focusing on reducing either MD or FA. Such a new design brings forth at least three advantages. First, as each model focuses on a relatively simpler sub-task, the overall difficulty of ISOS is somehow decreased. Second, the adversarial training of the two models naturally produces a delicate balance of MD and FA, and low rates for both MD and FA could be achieved at Nash equilibrium. Third, this MD-FA detachment gives us more flexibility to develop specific models dedicated to each sub-task. To realize the above design, we propose a conditional Generative Adversarial Network comprising of two generators and one discriminator. Each generator strives for one sub-task, while the discriminator differentiates the three segmentation results from the two generators and the ground truth. Moreover, in order to better serve the sub-tasks, the two generators, based on context aggregation networks, utilzse different size of receptive fields, providing both local and global views of objects for segmentation. As verified on multiple infrared image data sets, our method consistently achieves better segmentation than many state-of-the-art ISOS methods. Luping Zhou, Lei Wang 0001 |
ICCV | 2 |
| 2019 | Integrating Functional and Structural Connectivities via Diffusion-Convolution-Bilinear Neural Network
Jiashuang Huang, Luping Zhou, Lei Wang 0001, Daoqiang Zhang |
MICCAI (3) | 2 |
| 2019 | A Probabilistic Approach to Cross-Region Matching-Based Image RetrievalabstractWith deep convolutional features, cross-region matching (CRM) has recently shown superior performance on image retrieval. It evaluates image similarity by comparing image regions at different locations and scales, and is, therefore, more robust to geometric variance of objects. This paper first scrutinizes CRM-based image retrieval to provide a rigorous probabilistic interpretation by following the probability ranking principle. In addition to manifesting the assumptions implicitly taken by CRM, our interpretation highlights a fundamental issue hindering the performance of CRM-when comparing two image regions, CRM ignores modeling the distribution of the visual concept class associated with an image region, making the similarity comparison less precise. Taking advantage of the unprecedented representation capability of deep convolutional features, this paper proposes one approach to tackle that issue. It treats locally clustered image regions as a pseudo-labeled class sharing the same visual concept and utilizes them to model the distribution of the visual concept class associated with an image region. Both non-parametric and parametric methods are developed for this purpose, with careful probabilistic justification. Extensive experimental study on multiple benchmark data sets demonstrates the superior performance of the proposed pseudo-label approach to CRM and other comparable methods, with the maximum improvement of more than 10 percentage points over CRM. Zhimin Gao, Lei Wang 0001, Luping Zhou |
IEEE Trans. Image Process. | 3 |
| 2019 | 3D Auto-Context-Based Locality Adaptive Multi-Modality GANs for PET SynthesisabstractPositron emission tomography (PET) has been substantially used recently. To minimize the potential health risk caused by the tracer radiation inherent to PET scans, it is of great interest to synthesize the high-quality PET image from the low-dose one to reduce the radiation exposure. In this paper, we propose a 3D auto-context-based locality adaptive multi-modality generative adversarial networks model (LA-GANs) to synthesize the high-quality FDG PET image from the low-dose one with the accompanying MRI images that provide anatomical information. Our work has four contributions. First, different from the traditional methods that treat each image modality as an input channel and apply the same kernel to convolve the whole image, we argue that the contributions of different modalities could vary at different image locations, and therefore a unified kernel for a whole image is not optimal. To address this issue, we propose a locality adaptive strategy for multi-modality fusion. Second, we utilize 1 ×1 ×1 kernel to learn this locality adaptive fusion so that the number of additional parameters incurred by our method is kept minimum. Third, the proposed locality adaptive fusion mechanism is learned jointly with the PET image synthesis in a 3D conditional GANs model, which generates high-quality PET images by employing large-sized image patches and hierarchical features. Fourth, we apply the auto-context strategy to our scheme and propose an auto-context LA-GANs model to further refine the quality of synthesized images. Experimental results show that our method outperforms the traditional multi-modality fusion methods used in deep networks, as well as the state-of-the-art PET estimation approaches. Yan Wang 0015, Luping Zhou, Biting Yu, Lei Wang 0001, Chen Zu, David S. Lalush, Weili Lin, Xi Wu 0004, Jiliu Zhou, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Ea-GANs: Edge-Aware Generative Adversarial Networks for Cross-Modality MR Image SynthesisabstractMagnetic resonance (MR) imaging is a widely used medical imaging protocol that can be configured to provide different contrasts between the tissues in human body. By setting different scanning parameters, each MR imaging modality reflects the unique visual characteristic of scanned body part, benefiting the subsequent analysis from multiple perspectives. To utilize the complementary information from multiple imaging modalities, cross-modality MR image synthesis has aroused increasing research interest recently. However, most existing methods only focus on minimizing pixel/voxel-wise intensity difference but ignore the textural details of image content structure, which affects the quality of synthesized images. In this paper, we propose edge-aware generative adversarial networks (Ea-GANs) for cross-modality MR image synthesis. Specifically, we integrate edge information, which reflects the textural structure of image content and depicts the boundaries of different objects in images, to reduce this gap. Corresponding to different learning strategies, two frameworks are proposed, i.e., a generator-induced Ea-GAN (gEa-GAN) and a discriminator-induced Ea-GAN (dEa-GAN). The gEa-GAN incorporates the edge information via its generator, while the dEa-GAN further does this from both the generator and the discriminator so that the edge similarity is also adversarially learned. In addition, the proposed Ea-GANs are 3D-based and utilize hierarchical features to capture contextual information. The experimental results demonstrate that the proposed Ea-GANs, especially the dEa-GAN, outperform multiple state-of-the-art methods for cross-modality MR image synthesis in both qualitative and quantitative measures. Moreover, the dEa-GAN also shows excellent generality to generic image synthesis tasks on benchmark datasets about facades, maps, and cityscapes. Biting Yu, Luping Zhou, Lei Wang 0001, Yinghuan Shi, Jurgen Fripp, Pierrick Bourgeat |
IEEE Trans. Medical Imaging | 2 |
| 2018 | Modelling Diffusion Process by Deep Neural Networks for Image Retrieval
Yan Zhao 0019, Lei Wang 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
BMVC | 3 |
| 2018 | Data Fusion for MaaS: Opportunities and ChallengesabstractComputer Supported Cooperative Work (CSCW) in design is an essential facilitator for the development and implementation of smart cities, where modern cooperative transportation and integrated mobility are highly demanded. Owing to greater availability of different data sources, data fusion problem in intelligent transportation systems (ITS) has been very challenging, where machine learning modelling and approaches are promising to offer an important yet comprehensive solution. In this paper, we provide an overview of the recent advances in data fusion for Mobility as a Service (MaaS), including the basics of data fusion theory and the related machine learning methods. We also highlight the opportunities and challenges on MaaS, and discuss potential future directions of research on the integrated mobility modelling. Jianqing Wu 0002, Luping Zhou, Jun Shen 0001, Sim Kim Lau, Jianming Yong |
CSCWD | 2 |
| 2018 | DeepKSPD: Learning Kernel-Matrix-Based SPD Representation For Fine-Grained Image Recognition
Melih Engin, Lei Wang 0001, Luping Zhou, Xinwang Liu 0002 |
ECCV (2) | 3 |
| 2018 | A Novel Image-Specific Transfer Approach for Prostate Segmentation in MR ImagesabstractProstate segmentation in Magnetic Resonance (MR) Images is a significant yet challenging task for prostate cancer treatment. Most of the existing works attempted to design a global classifier for all MR images, which neglect the discrepancy of images across different patients. To this end, we propose a novel transfer approach for prostate segmentation in MR images. Firstly, an image-specific classifier is built for each training image. Secondly, a pair of dictionaries and a mapping matrix are jointly obtained by a novel Semi-Coupled Dictionary Transfer Learning (SCDTL). Finally, the classifiers on the source domain could be selectively transferred to the target domain (i.e. testing images) by the dictionaries and the mapping matrix. The evaluation demonstrates that our approach has a competitive performance compared with the state-of-the-art transfer learning methods. Moreover, the proposed transfer approach outperforms the conventional deep neural network based method. Pinzhuo Tian, Lei Qi 0001, Yinghuan Shi, Luping Zhou, Yang Gao 0001, Dinggang Sheri |
ICASSP | 4 |
| 2018 | Locality Adaptive Multi-modality GANs for High-Quality PET Image Synthesis
Yan Wang 0015, Luping Zhou, Lei Wang 0001, Biting Yu, Chen Zu, David S. Lalush, Weili Lin, Xi Wu 0004, Jiliu Zhou, Dinggang Shen |
MICCAI (1) | 2 |
| 2018 | Instance Image Retrieval by Aggregating Sample-based Discriminative CharacteristicsabstractIdentifying the discriminative characteristic of a query is important for image retrieval. For retrieval without human interaction, such characteristic is usually obtained by average query expansion (AQE) or its discriminative variant (DQE) learned from pseudo-examples online, among others. In this paper, we propose a new query expansion method to further improve the above ones. The key idea is to learn a "unique'' discriminative characteristic for each database image, in an offline manner. During retrieval, the characteristic of a query is obtained by aggregating the unique characteristics of the query-relevant images collected from an initial retrieval result. Compared with AQE which works in the original feature space, our method works in the space of the unique characteristics of database images, significantly enhancing the discriminative power of the characteristic identified for a query. Compared with DQE, our method needs neither pseudo-labeled negatives nor the online learning process, leading to more efficient retrieval and even better performance. The experimental study conducted on seven benchmark datasets verifies the considerable improvement achieved by the proposed method, and also demonstrates its application to the state-of-the-art diffusion-based image retrieval. Zhongyan Zhang, Lei Wang 0001, Yang Wang 0002, Luping Zhou, Jianjia Zhang, Fang Chen 0001 |
ICMR | 4 |
| 2018 | OPML: A one-pass closed-form solution for online metric learning
Wenbin Li 0006, Yang Gao 0001, Lei Wang 0001, Luping Zhou, Jing Huo, Yinghuan Shi |
Pattern Recognit. | 4 |
| 2017 | Revisiting Metric Learning for SPD Matrix Based Visual RepresentationabstractThe success of many visual recognition tasks largely depends on a good similarity measure, and distance metric learning plays an important role in this regard. Meanwhile, Symmetric Positive Definite (SPD) matrix is receiving increased attention for feature representation in multiple computer vision applications. However, distance metric learning on SPD matrices has not been sufficiently researched. A few existing works approached this by learning either d2× p or d × k transformation matrix for d× d SPD matrices. Different from these methods, this paper proposes a new member to the family of distance metric learning for SPD matrices. It learns only d parameters to adjust the eigenvalues of the SPD matrices through an efficient optimisation scheme. Also, it is shown that the proposed method can be interpreted as learning a sample-specific transformation matrix, instead of the fixed transformation matrix learned for all the samples in the existing works. The optimised d parameters can be used to massage the SPD matrices for better discrimination while still keeping them in the original space. From this perspective, the proposed method complements, rather than competes with, the existing linear-transformation-based methods, as the latter can always be applied to the output of the former to perform distance metric learning in further. The proposed method has been tested on multiple SPD-based visual representation data sets used in the literature, and the results demonstrate its interesting properties and attractive performance. Luping Zhou, Lei Wang 0001, Jianjia Zhang, Yinghuan Shi, Yang Gao 0001 |
CVPR | 1 |
| 2017 | Infomax principle based pooling of deep convolutional activations for image retrievalabstractNeural activations produced by deep convolutional networks have recently become state-of-the-art representation for image retrieval. To obtain a global image representation, sum-pooling has been frequently used to aggregate activations of convolutional feature maps. This work first presents an understanding on the effectiveness of sum-pooling via probabilistic interpretation, by proving that sum-pooling is an upper bound of the probability that a visual pattern is present in an image. To further answer the optimality of sum-pooling, a quantitative analysis based on the Infomax principle in neural networks is provided. It shows that sum-pooling aligns well with the leading eigenvector of principal component analysis (PCA) applied to the activations of a feature map. Moreover, considering the 2D matrix structure of feature maps, a two-directional 2DPCA-based pooling scheme is proposed to aggregate the convolutional activations. Experiments on multiple benchmark image retrieval datasets demonstrate the above analysis and the superiority of the proposed pooling scheme. Zhimin Gao, Lei Wang 0001, Luping Zhou, Ming Yang 0014 |
ICME | 3 |
| 2017 | Machine learning in medical imaging
Kenji Suzuki 0001, Luping Zhou, Qian Wang 0001 |
Pattern Recognit. | 2 |
| 2017 | Subject-adaptive Integration of Multiple SICE Brain Networks with Different Sparsity
Jianjia Zhang, Luping Zhou, Lei Wang 0001 |
Pattern Recognit. | 2 |
| 2017 | HEp-2 Cell Image Classification With Deep Convolutional Neural NetworksabstractEfficient Human Epithelial-2 cell image classification can facilitate the diagnosis of many autoimmune diseases. This paper proposes an automatic framework for this classification task, by utilizing the deep convolutional neural networks (CNNs) which have recently attracted intensive attention in visual recognition. In addition to describing the proposed classification framework, this paper elaborates several interesting observations and findings obtained by our investigation. They include the important factors that impact network design and training, the role of rotation-based data augmentation for cell images, the effectiveness of cell image masks for classification, and the adaptability of the CNN-based classification system across different datasets. Extensive experimental study is conducted to verify the above findings and compares the proposed framework with the well-established image classification models in the literature. The results on benchmark datasets demonstrate that 1) the proposed framework can effectively outperform existing models by properly applying data augmentation, 2) our CNN-based framework has excellent adaptability across different datasets, which is highly desirable for cell image classification under varying laboratory settings. Our system is ranked high in the cell image classification competition hosted by ICPR 2014. Zhimin Gao, Lei Wang 0001, Luping Zhou, Jianjia Zhang |
IEEE J. Biomed. Health Informatics | 3 |
| 2017 | A Graph-Embedding Approach to Hierarchical Visual Word MergenceabstractAppropriately merging visual words are an effective dimension reduction method for the bag-of-visual-words model in image classification. The approach of hierarchically merging visual words has been extensively employed, because it gives a fully determined merging hierarchy. Existing supervised hierarchical merging methods take different approaches and realize the merging process with various formulations. In this paper, we propose a unified hierarchical merging approach built upon the graph-embedding framework. Our approach is able to merge visual words for any scenario, where a preferred structure and an undesired structure are defined, and, therefore, can effectively attend to all kinds of requirements for the word-merging process. In terms of computational efficiency, we show that our algorithm can seamlessly integrate a fast search strategy developed in our previous work and, thus, well maintain the state-of-the-art merging speed. To the best of our survey, the proposed approach is the first one that addresses the hierarchical visual word mergence in such a flexible and unified manner. As demonstrated, it can maintain excellent image classification performance even after a significant dimension reduction, and outperform all the existing comparable visual word-merging methods. In a broad sense, our work provides an open platform for applying, evaluating, and developing new criteria for hierarchical word-merging tasks. Lei Wang 0001, Lingqiao Liu, Luping Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | A Hidden Markov Model based efficient next best viewpoint planning for autonomous explorationabstractA next best viewpoint planning algorithm based on Hidden Markov Model (HMM) was presented for autonomous exploration. The grid map and probabilistic assessment model based on HMM was established to evaluate the next best exploration perspective. Leveraging the relative information entropy, the maximum information gain of candidate viewpoints were assessed. After the establishment of the actual scene graph rule, the state transfer matrix and observation occupied state transition matrix were derived. Used Bayesian filtering to update the probabilistic grid diagram, the path planning was selected from starting point to the maximum information return point. By simulation experiment scene of different complexity, the efficiency and its good robustness was confirmed. Bingen Li, Luping Zhou |
CSCWD | 5 |
| 2016 | Learning Discriminative Bayesian Networks from High-Dimensional Continuous Neuroimaging DataabstractDue to its causal semantics, Bayesian networks (BN) have been widely employed to discover the underlying data relationship in exploratory studies, such as brain research. Despite its success in modeling the probability distribution of variables, BN is naturally a generative model, which is not necessarily discriminative. This may cause the ignorance of subtle but critical network changes that are of investigation values across populations. In this paper, we propose to improve the discriminative power of BN models for continuous variables from two different perspectives. This brings two general discriminative learning frameworks for Gaussian Bayesian networks (GBN). In the first framework, we employ Fisher kernel to bridge the generative models of GBN and the discriminative classifiers of SVMs, and convert the GBN parameter learning to Fisher kernel learning via minimizing a generalization error bound of SVMs. In the second framework, we employ the max-margin criterion and build it directly upon GBN models to explicitly optimize the classification performance of the GBNs. The advantages and disadvantages of the two frameworks are discussed and experimentally compared. Both of them demonstrate strong power in learning discriminative parameters of GBNs for neuroimaging based brain network analysis, as well as maintaining reasonable representation capacity. The contributions of this paper also include a new Directed Acyclic Graph (DAG) constraint with theoretical guarantee to ensure the graph validity of GBN. Luping Zhou, Lei Wang 0001, Lingqiao Liu, Philip Ogunbona, Dinggang Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Learning Discriminative Stein Kernel for SPD Matrices and Its ApplicationsabstractStein kernel (SK) has recently shown promising performance on classifying images represented by symmetric positive definite (SPD) matrices. It evaluates the similarity between two SPD matrices through their eigenvalues. In this paper, we argue that directly using the original eigenvalues may be problematic because: 1) eigenvalue estimation becomes biased when the number of samples is inadequate, which may lead to unreliable kernel evaluation, and 2) more importantly, eigenvalues reflect only the property of an individual SPD matrix. They are not necessarily optimal for computing SK when the goal is to discriminate different classes of SPD matrices. To address the two issues, we propose a discriminative SK (DSK), in which an extra parameter vector is defined to adjust the eigenvalues of input SPD matrices. The optimal parameter values are sought by optimizing a proxy of classification performance. To show the generality of the proposed method, three kernel learning criteria that are commonly used in the literature are employed as a proxy. A comprehensive experimental study is conducted on a variety of image classification tasks to compare the proposed DSK with the original SK and other methods for evaluating the similarity between SPD matrices. The results demonstrate that the DSK can attain greater discrimination and better align with classification tasks by altering the eigenvalues. This makes it produce higher classification performance than the original SK and other commonly used methods. Jianjia Zhang, Lei Wang 0001, Luping Zhou, Wanqing Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Beyond Covariance: Feature Representation with Nonlinear Kernel MatricesabstractCovariance matrix has recently received increasing attention in computer vision by leveraging Riemannian geometry of symmetric positive-definite (SPD) matrices. Originally proposed as a region descriptor, it has now been used as a generic representation in various recognition tasks. However, covariance matrix has shortcomings such as being prone to be singular, limited capability in modeling complicated feature relationship, and having a fixed form of representation. This paper argues that more appropriate SPD-matrix-based representations shall be explored to achieve better recognition. It proposes an open framework to use the kernel matrix over feature dimensions as a generic representation and discusses its properties and advantages. The proposed framework significantly elevates covariance representation to the unlimited opportunities provided by this new representation. Experimental study shows that this representation consistently outperforms its covariance counterpart on various visual recognition tasks. In particular, it achieves significant improvement on skeleton-based human action recognition, demonstrating the state-of-the-art performance over both the covariance and the existing non-covariance representations. Lei Wang 0001, Jianjia Zhang, Luping Zhou, Chang Tang, Wanqing Li 0001 |
ICCV | 3 |
| 2015 | An efficient radius-incorporated MKL algorithm for Alzheimer's disease prediction
Xinwang Liu 0002, Luping Zhou, Lei Wang 0001, Jian Zhang 0002, Jianping Yin, Dinggang Shen |
Pattern Recognit. | 2 |
| 2014 | Discriminative Sparse Inverse Covariance Matrix: Application in Brain Functional Network ClassificationabstractRecent studies show that mental disorders change the functional organization of the brain, which could be investigated via various imaging techniques. Analyzing such changes is becoming critical as it could provide new biomarkers for diagnosing and monitoring the progression of the diseases. Functional connectivity analysis studies the covary activity of neuronal populations in different brain regions. The sparse inverse covariance estimation (SICE), also known as graphical LASSO, is one of the most important tools for functional connectivity analysis, which estimates the interregional partial correlations of the brain. Although being increasingly used for predicting mental disorders, SICE is basically a generative method that may not necessarily perform well on classifying neuroimaging data. In this paper, we propose a learning framework to effectively improve the discriminative power of SICEs by taking advantage of the samples in the opposite class. We formulate our objective as convex optimization problems for both one-class and two-class classifications. By analyzing these optimization problems, we not only solve them efficiently in their dual form, but also gain insights into this new learning framework. The proposed framework is applied to analyzing the brain metabolic covariant networks built upon FDG-PET images for the prediction of the Alzheimer's disease, and shows significant improvement of classification performance for both one-class and two-class scenarios. Moreover, as SICE is a general method for learning undirected Gaussian graphical models, this paper has broader meanings beyond the scope of brain research. Luping Zhou, Lei Wang 0001, Philip Ogunbona |
CVPR | 1 |
| 2014 | Max-Margin Based Learning for Discriminative Bayesian Network from Neuroimaging Data
Luping Zhou, Lei Wang 0001, Lingqiao Liu, Philip Ogunbona, Dinggang Shen |
MICCAI (3) | 1 |
| 2014 | A Hierarchical Word-Merging Algorithm with Class Separability MeasureabstractIn image recognition with the bag-of-features model, a small-sized visual codebook is usually preferred to obtain a low-dimensional histogram representation and high computational efficiency. Such a visual codebook has to be discriminative enough to achieve excellent recognition performance. To create a compact and discriminative codebook, in this paper we propose to merge the visual words in a large-sized initial codebook by maximally preserving class separability. We first show that this results in a difficult optimization problem. To deal with this situation, we devise a suboptimal but very efficient hierarchical word-merging algorithm, which optimally merges two words at each level of the hierarchy. By exploiting the characteristics of the class separability measure and designing a novel indexing structure, the proposed algorithm can hierarchically merge 10,000 visual words down to two words in merely 90 seconds. Also, to show the properties of the proposed algorithm and reveal its advantages, we conduct detailed theoretical analysis to compare it with another hierarchical word-merging algorithm that maximally preserves mutual information, obtaining interesting findings. Experimental studies are conducted to verify the effectiveness of the proposed algorithm on multiple benchmark data sets. As shown, it can efficiently produce more compact and discriminative codebooks than the state-of-the-art hierarchical word-merging algorithms, especially when the size of the codebook is significantly reduced. Lei Wang 0001, Luping Zhou, Chunhua Shen, Lingqiao Liu, Huan Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Multiple Kernel Learning in the Primal for Multimodal Alzheimer's Disease ClassificationabstractTo achieve effective and efficient detection of Alzheimer's disease (AD), many machine learning methods have been introduced into this realm. However, the general case of limited training samples, as well as different feature representations typically makes this problem challenging. In this paper, we propose a novel multiple kernel-learning framework to combine multimodal features for AD classification, which is scalable and easy to implement. Contrary to the usual way of solving the problem in the dual, we look at the optimization from a new perspective. By conducting Fourier transform on the Gaussian kernel, we explicitly compute the mapping function, which leads to a more straightforward solution of the problem in the primal. Furthermore, we impose the mixed L21 norm constraint on the kernel weights, known as the group lasso regularization, to enforce group sparsity among different feature modalities. This actually acts as a role of feature modality selection, while at the same time exploiting complementary information among different kernels. Therefore, it is able to extract the most discriminative features for classification. Experiments on the ADNI dataset demonstrate the effectiveness of the proposed method. Fayao Liu, Luping Zhou, Chunhua Shen, Jianping Yin |
IEEE J. Biomed. Health Informatics | 2 |
| 2013 | A Fast Approximate AIB Algorithm for Distributional Word ClusteringabstractDistributional word clustering merges the words having similar probability distributions to attain reliable parameter estimation, compact classification models and even better classification performance. Agglomerative Information Bottleneck (AIB) is one of the typical word clustering algorithms and has been applied to both traditional text classification and recent image recognition. Although enjoying theoretical elegance, AIB has one main issue on its computational efficiency, especially when clustering a large number of words. Different from existing solutions to this issue, we analyze the characteristics of its objective function-the loss of mutual information, and show that by merely using the ratio of word-class joint probabilities of each word, good candidate word pairs for merging can be easily identified. Based on this finding, we propose a fast approximate AIB algorithm and show that it can significantly improve the computational efficiency of AIB while well maintaining or even slightly increasing its classification performance. Experimental study on both text and image classification benchmark data sets shows that our algorithm can achieve more than 100 times speedup on large real data sets over the state-of-the-art method. Lei Wang 0001, Jianjia Zhang, Luping Zhou, Wanqing Li 0001 |
CVPR | 3 |
| 2013 | Discriminative Brain Effective Connectivity Analysis for Alzheimer's Disease: A Kernel Learning Approach upon Sparse Gaussian Bayesian NetworkabstractAnalyzing brain networks from neuroimages is becoming a promising approach in identifying novel connectivity-based biomarkers for the Alzheimer's disease (AD). In this regard, brain ``effective connectivity" analysis, which studies the causal relationship among brain regions, is highly challenging and of many research opportunities. Most of the existing works in this field use generative methods. Despite their success in data representation and other important merits, generative methods are not necessarily discriminative, which may cause the ignorance of subtle but critical disease-induced changes. In this paper, we propose a learning-based approach that integrates the benefits of generative and discriminative methods to recover effective connectivity. In particular, we employ Fisher kernel to bridge the generative models of sparse Bayesian networks (SBN) and the discriminative classifiers of SVMs, and convert the SBN parameter learning to Fisher kernel learning via minimizing a generalization error bound of SVMs. Our method is able to simultaneously boost the discriminative power of both the generative SBN models and the SBN-induced SVM classifiers via Fisher kernel. The proposed method is tested on analyzing brain effective connectivity for AD from ADNI data, and demonstrates significant improvements over the state-of-the-art work. Luping Zhou, Lei Wang 0001, Lingqiao Liu, Philip Ogunbona, Dinggang Shen |
CVPR | 1 |
| 2012 | MR-Less Surface-Based Amyloid Estimation by Subject-Specific Atlas Selection and Bayesian Fusion
Luping Zhou, Olivier Salvado, Vincent Doré, Pierrick Bourgeat, Parnesh Raniga, Victor Villemagne, Christopher Rowe, Jurgen Fripp |
MICCAI (2) | 1 |
| 2011 | Hierarchical anatomical brain networks for MCI prediction by partial least square analysisabstractOwning to its clinical accessibility, T1-weighted MRI has been extensively studied for the prediction of mild cognitive impairment (MCI) and Alzheimer's disease (AD). The tissue volumes of GM, WM and CSF are the most commonly used measures for MCI and AD prediction. We note that disease-induced structural changes may not happen at isolated spots, but in several inter-related regions. Therefore, in this paper we propose to directly extract the inter-region connectivity based features for MCI prediction. This involves constructing a brain network for each subject, with each node representing an ROI and each edge representing regional interactions. This network is also built hierarchically to improve the robustness of classification. Compared with conventional methods, our approach produces a significant larger pool of features, which if improperly dealt with, will result in intractability when used for classifier training. Therefore based on the characteristics of the network features, we employ Partial Least Square analysis to efficiently reduce the feature dimensionality to a manageable level while at the same time preserving discriminative information as much as possible. Our experiment demonstrates that without requiring any new information in addition to T1-weighted images, the prediction accuracy of MCI is statistically improved. Luping Zhou, Yang Li 0010, Pew-Thian Yap, Dinggang Shen |
CVPR | 1 |
| 2010 | Hippocampal Shape Classification Using Redundancy Constrained Feature Selection
Luping Zhou, Lei Wang 0001, Chunhua Shen, Nick Barnes |
MICCAI (2) | 1 |
| 2010 | Feature selection with redundancy-constrained class separabilityabstractScatter-matrix-based class separability is a simple and efficient feature selection criterion in the literature. However, the conventional trace-based formulation does not take feature redundancy into account and is prone to selecting a set of discriminative but mutually redundant features. In this brief, we first theoretically prove that in the context of this trace-based criterion the existence of sufficiently correlated features can always prevent selecting the optimal feature set. Then, on top of this criterion, we propose the redundancy-constrained feature selection (RCFS). To ensure the algorithm's efficiency and scalability, we study the characteristic of the constraints with which the resulted constrained 0-1 optimization can be efficiently and globally solved. By using the totally unimodular (TUM) concept in integer programming, a necessary condition for such constraints is derived. This condition reveals an interesting special case in which qualified redundancy constraints can be conveniently generated via a clustering of features. We study this special case and develop an efficient feature selection approach based on Dinkelbach's algorithm. Experiments on benchmark data sets demonstrate the superior performance of our approach to those without redundancy constraints. Luping Zhou, Lei Wang 0001, Chunhua Shen |
IEEE Trans. Neural Networks | 1 |
| 2009 | Discriminative Maximum Margin Image Object Categorization with Exact InferenceabstractCategorizing multiple objects in images is essentially a structured prediction problem: the label of an object is in general dependent on the labels of other objects in the image. We explicitly model object dependencies in a sparse graphical topology induced by the adjacency of objects in the image, which benefits inference, and then use maximum margin principle to learn the model discriminatively. Moreover, we propose a novel exact inference method, which is used in training to find the most violated constraint required by cutting plane method. A slightly modified inference method is used in testing when the target labels are unseen. Experiment results on both synthetic and real datasets demonstrate the improvement of the proposed approach over the state-of-the-art methods. Qinfeng Shi, Luping Zhou, Li Cheng 0001, Dale Schuurmans |
ICIG | 2 |
| 2009 | Identifying Anatomical Shape Difference by Regularized Discriminative DirectionabstractIdentifying the shape difference between two groups of anatomical objects is important for medical image analysis and computer-aided diagnosis. A method called "discriminative direction" in the literature has been proposed to solve this problem. In that method, the shape difference between groups is identified by deforming a shape along the discriminative direction. This paper conducts a thorough study about inferring this discriminative direction in an efficient and accurate way. First, finding the discriminative direction is reformulated as a preimage problem in kernel-based learning. This provides a complementary but conceptually simpler solution than the previous method. More importantly, we find that a shape deforming along the original discriminative direction cannot faithfully maintain its anatomical correctness. This unnecessarily introduces spurious shape differences and leads to inaccurate analysis. To overcome this problem, this paper further proposes a regularized discriminative direction by requiring a shape to conform to its underlying distribution when it deforms. Two different approaches are developed to impose the regularization, one from the perspective of probability distributions and the other from a geometric point of view, and their relationship is discussed. After verifying their superior performance through controlled experiments, we apply the proposed methods to detecting and localizing the hippocampal shape difference between sexes. We get results consistent with other independent research, providing a more compact representation of the shape difference compared with the established discriminative direction method. Luping Zhou, Richard I. Hartley, Lei Wang 0001, Paulette Lieby, Nick Barnes |
IEEE Trans. Medical Imaging | 1 |
| 2008 | A Fast Algorithm for Creating a Compact and Discriminative Visual Codebook
Lei Wang 0001, Luping Zhou, Chunhua Shen |
ECCV (4) | 2 |
| 2008 | Regularized Discriminative Direction for Shape Difference Analysis
Luping Zhou, Richard I. Hartley, Lei Wang 0001, Paulette Lieby, Nick Barnes |
MICCAI (1) | 1 |
| 2008 | A Kernel-Induced Space Selection Approach to Model Selection in KLDAabstractModel selection in kernel linear discriminant analysis (KLDA) refers to the selection of appropriate parameters of a kernel function and the regularizer. By following the principle of maximum information preservation, this paper formulates the model selection problem as a problem of selecting an optimal kernel-induced space in which different classes are maximally separated from each other. A scatter-matrix-based criterion is developed to measure the "goodness" of a kernel-induced space, and the kernel parameters are tuned by maximizing this criterion. This criterion is computationally efficient and is differentiable with respect to the kernel parameters. Compared with the leave-one-out (LOO) or k-fold cross validation (CV), the proposed approach can achieve a faster model selection, especially when the number of training samples is large or when many kernel parameters need to be tuned. To tune the regularization parameter in the KLDA, our criterion is used together with the method proposed by Saadi (2004). Experiments on benchmark data sets verify the effectiveness of this model selection approach. Lei Wang 0001, Kap Luk Chan, Ping Xue 0001, Luping Zhou |
IEEE Trans. Neural Networks | 4 |
| 2007 | A Study of Hippocampal Shape Difference Between Genders by Efficient Hypothesis Test and Discriminative Deformation
Luping Zhou, Richard I. Hartley, Paulette Lieby, Nick Barnes, Kaarin Anstey, Nicolas Cherbuin, Perminder S. Sachdev |
MICCAI (1) | 1 |
| 2001 | Polygonizing Non-Uniformly Distributed 3D Points by Advancing Mesh Frontiersabstract3D digitization devices produce very large sets of 3D points sampled from the surfaces of the objects being scanned. A mesh construction procedure needs to be applied to derive polygon mesh from the 3D point sets. As the 3D points derived from digitization devices based on digital imaging technologies are inherently non-uniformly distributed over regions that may contain surface discontinuities, existing methods are not suitable for polygonizing them. This paper describes a novel polygonization algorithm for constructing triangle mesh from unorganized 3D points. In contrast to existing methods, this algorithm begins the mesh construction process from 3D points lying on smooth surfaces, and advances the mesh frontier towards 3D points lying near surface discontinuities. If 3D points along the edges and at the corners are sampled, then the algorithm will form an edge where two advancing frontiers meet, and a corner where three or more frontiers meet. Otherwise, the algorithm constructs approximations of the edges and corners. It can be shown that this frontier advancing algorithm performs 2D Delaunay triangulation of 3D points lying on a plane in 3D space. Indriyati Atmosukarto, Luping Zhou, Wee Kheng Leow |
Computer Graphics International | 2 |