Jun Zhang 0018

dblp:29/4190-18 · DBLP profile ↗
← Back
81ranked-venue papers
16as first author
46since 2021 · last 2026
0000-0001-5579-7094ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 33 · 4 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 32 · 7 first-author · 14 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DST-Net: A closed-loop dual-stream transformer with identity-guided video matting for visible-infrared person re-identification
Yanze Zhu, Rongbo Fan, Jun Zhang 0018, Jianhua Yang 0005
Neurocomputing4
2025 Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models
abstract
Recent research showcases the considerable potential of conditional diffusion models for generating consistent stories. However, current methods, which primarily generate stories in a caption-dependent manner, often overlook the importance of contextual consistency and the relevance of frames during sequential generation. To address this, we propose a novel Rich-contextual Conditional Diffusion Models (RCDMs), a two-stage approach designed to enhance story generation's semantic consistency and temporal consistency. Specifically, in the first stage, the frame-prior transformer diffusion model is presented to predict the frame semantic embedding of the unknown clip by aligning the semantic correlations between the captions and frames of the known clip. The second stage establishes a robust model with rich contextual conditions, including reference images of the known clip, the predicted frame semantic embedding of the unknown clip, and text embeddings of all captions. By jointly injecting these rich contextual conditions at the image and feature levels, RCDMs can generate semantic and temporal consistency stories. Moreover, RCDMs can generate consistent stories with a single forward inference compared to autoregressive models. Our qualitative and quantitative results demonstrate that our proposed RCDMs outperform in challenging scenarios.
Fei Shen 0004, Hu Ye, Sibo Liu, Jun Zhang 0018, Cong Wang 0034, Xiao Han 0011
AAAI4
2025 Ensembling Diffusion Models via Adaptive Feature Aggregation
abstract
The success of the text-guided diffusion model has inspired the development and release of numerous powerful diffusion models within the open-source community. These models are typically fine-tuned on various expert datasets, showcasing diverse denoising capabilities. Leveraging multiple high-quality models to produce stronger generation ability is valuable, but has not been extensively studied. Existing methods primarily adopt parameter merging strategies to produce a new static model. However, they overlook the fact that the divergent denoising capabilities of the models may dynamically change across different states, such as when experiencing different prompts, initial noises, denoising steps, and spatial locations. In this paper, we propose a novel ensembling method, Adaptive Feature Aggregation (AFA), which dynamically adjusts the contributions of multiple models at the feature level according to various states (i.e., prompts, initial noises, denoising steps, and spatial locations), thereby keeping the advantages of multiple diffusion models, while suppressing their disadvantages. Specifically, we design a lightweight Spatial-Aware Block-Wise (SABW) feature aggregator that adaptive aggregates the block-wise intermediate features from multiple U-Net denoisers into a unified one. The core idea lies in dynamically producing an individual attention map for each model's features by comprehensively considering various states. It is worth noting that only SABW is trainable with about 50 million parameters, while other models are frozen. Both the quantitative and qualitative experiments demonstrate the effectiveness of our proposed method.
Cong Wang 0034, Kuan Tian, Yonghang Guan, Fei Shen 0004, Zhiwei Jiang 0001, Qing Gu 0001, Jun Zhang 0018
ICLR7
2025 LLM-Primitives: Large Language Model for 3D Reconstruction with Primitives
abstract
We present LLM-Primitives: Large Language Model for 3D Reconstruction with Primitives, a novel approach to shape abstraction. By incorporating multi-modal conditional inputs, our method enables LLMs to reconstruct high-quality 3D primitives using only a modest amount of training data (tens of thousands of samples). This work marks a significant milestone in applying large language models to 3D primitive-based reconstruction, demonstrating both their feasibility and effectiveness in this domain. Specifically, we leverage the point clouds of existing 3D models as conditional inputs to the LLM via a multi-modal connector. Instead of directly estimating primitive parameters, we introduce a center-to-surface vector representation, ensuring deterministic outputs and avoiding the ambiguity often associated with primitive parameterization. Experimental results show that LLM-Primitives surpass state-of-the-art 3D primitive methods across various quantitative metrics. Notably, the substantial improvements in visual quality further confirm that LLM-Primitives can reconstruct high-quality, practical 3D primitives. (Project page: https://llm-primitives.github.io/LLM-Primitives/)
Kuan Tian, Yonghang Guan, Jun Zhang 0018
SIGGRAPH Asia4
2025 NAPG: Neighborhood-Assisted Multiprototype Group Model for Cross-Domain Semantic Segmentation of Remote Sensing Images
abstract
Unsupervised domain adaptation (UDA) is crucial for semantic segmentation of remote sensing images (RS-SS), particularly when data distributions differ between source and target domains. Existing prototype-based UDA methods struggle with complex land cover class distributions and spatial information capture. To address these limitations, the Neighborhood-Assisted Multi-Prototype Group (NAPG) model is proposed. This model enhances cross-domain adaptability and spatial context richness by dynamically determining the number of prototype features and integrating neighborhood similarity gradients. Specifically, NAPG employs the Cross-Domain Representation of Multi-Prototype Group (CDR-MPG) module to generate multi-prototype group, capturing land cover complexity more effectively. Additionally, the Gradient Neighborhood Consistency Estimation (GNCE) module improves spatial representation by reducing intra-class variance and alleviating feature inconsistency. Experiments demonstrate that the proposed NAPG model outperforms state-of-the-art UDA methods across multiple datasets, achieving an mean intersection over union (mIoU) improvement of 3%. The source code is publicly available at https://github.com/Fanrongbo/NAPG-UDA-RS-SS.
Rongbo Fan, Jialin Xie, Junmin Liu, Jun Zhang 0018, Yan Zhang 0109, Hong Hou, Jianhua Yang 0005
IEEE Trans. Geosci. Remote. Sens.4
2024 Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models
abstract
Recent work has showcased the significant potential of diffusion models in pose-guided person image synthesis. However, owing to the inconsistency in pose between the source and target images, synthesizing an image with a distinct pose, relying exclusively on the source image and target pose information, remains a formidable challenge. This paper presents Progressive Conditional Diffusion Models (PCDMs) that incrementally bridge the gap between person images under the target and source poses through three stages. Specifically, in the first stage, we design a simple prior conditional diffusion model that predicts the global features of the target image by mining the global alignment relationship between pose coordinates and image appearance. Then, the second stage establishes a dense correspondence between the source and target images using the global features from the previous stage, and an inpainting conditional diffusion model is proposed to further align and enhance the contextual features, generating a coarse-grained person image. In the third stage, we propose a refining conditional diffusion model to utilize the coarsely generated image from the previous stage as a condition, achieving texture restoration and enhancing fine-detail consistency. The three-stage PCDMs work progressively to generate the final high-quality and high-fidelity synthesized image. Both qualitative and quantitative results demonstrate the consistency and photorealism of our proposed PCDMs under challenging scenarios. The code and model will be available at https://github.com/tencent-ailab/PCDMs.
Fei Shen 0004, Hu Ye, Jun Zhang 0018, Cong Wang 0034, Xiao Han 0011
ICLR3
2024 VersVideo: Leveraging Enhanced Temporal Diffusion Models for Versatile Video Generation
abstract
Creating stable, controllable videos is a complex task due to the need for significant variation in temporal dynamics and cross-frame temporal consistency. To address this, we enhance the spatial-temporal capability and introduce a versatile video generation model, VersVideo, which leverages textual, visual, and stylistic conditions. Current video diffusion models typically extend image diffusion architectures by supplementing 2D operations (such as convolutions and attentions) with temporal operations. While this approach is efficient, it often restricts spatial-temporal performance due to the oversimplification of standard 3D operations. To counter this, we incorporate two key elements: (1) multi-excitation paths for spatial-temporal convolutions with dimension pooling across different axes, and (2) multi-expert spatial-temporal attention blocks. These enhancements boost the model's spatial-temporal performance without significantly escalating training and inference costs. We also tackle the issue of information loss that arises when a variational autoencoder is used to transform pixel space into latent features and then back into pixel frames. To mitigate this, we incorporate temporal modules into the decoder to maintain inter-frame consistency. Lastly, by utilizing the innovative denoising UNet and decoder, we develop a unified ControlNet model suitable for various conditions, including image, Canny, HED, depth, and style. Examples of the videos generated by our model can be found at https://jinxixiang.github.io/versvideo/.
Jinxi Xiang, Ricong Huang, Jun Zhang 0018, Guanbin Li, Xiao Han 0011
ICLR3
2024 Semi-Open 3D Object Retrieval via Hierarchical Equilibrium on Hypergraph
abstract
Existing open-set learning methods consider only the single-layer labels of objects and strictly assume no overlap between the training and testing sets, leading to contradictory optimization for superposed categories. In this paper, we introduce a more practical Semi-Open Environment setting for open-set 3D object retrieval with hierarchical labels, in which the training and testing set share a partial label space for coarse categories but are completely disjoint from fine categories. We propose the Hypergraph-Based Hierarchical Equilibrium Representation (HERT) framework for this task. Specifically, we propose the Hierarchical Retrace Embedding (HRE) module to overcome the global disequilibrium of unseen categories by fully leveraging the multi-level category information. Besides, tackling the feature overlap and class confusion problem, we perform the Structured Equilibrium Tuning (SET) module to utilize more equilibrial correlations among objects and generalize to unseen categories, by constructing a superposed hypergraph based on the local coherent and global entangled correlations. Furthermore, we generate four semi-open 3DOR datasets with multi-level labels for benchmarking. Results demonstrate that the proposed method can effectively generate the hierarchical embeddings of 3D objects and generalize them towards semi-open environments.
Yang Xu 0064, Yifan Feng 0001, Jun Zhang 0018, Jun-Hai Yong, Yue Gao 0002
NeurIPS3
2024 Assembly Fuzzy Representation on Hypergraph for Open-Set 3D Object Retrieval
abstract
The lack of object-level labels presents a significant challenge for 3D object retrieval in the open-set environment. However, part-level shapes of objects often share commonalities across categories but remain underexploited in existing retrieval methods. In this paper, we introduce the Hypergraph-Based Assembly Fuzzy Representation (HARF) framework, which navigates the intricacies of open-set 3D object retrieval through a bottom-up lens of Part Assembly. To tackle the challenge of assembly isomorphism and unification, we propose the Hypergraph Isomorphism Convolution (HIConv) for smoothing and adopt the Isomorphic Assembly Embedding (IAE) module to generate assembly embeddings with geometric-semantic consistency. To address the challenge of open-set category generalization, our method employs high-order correlations and fuzzy representation to mitigate distribution skew through the Structure Fuzzy Reconstruction (SFR) module, by constructing a leveraged hypergraph based on local certainty and global uncertainty correlations. We construct three open-set retrieval datasets for 3D objects with part-level annotations: OP-SHNP, OP-INTRA, and OP-COSEG. Extensive experiments and ablation studies on these three benchmarks show our method outperforms current state-of-the-art methods.
Yang Xu 0064, Yifan Feng 0001, Jun Zhang 0018, Jun-Hai Yong, Yue Gao 0002
NeurIPS3
2024 CoNIC Challenge: Pushing the frontiers of nuclear detection, segmentation, classification and counting
abstract
Nuclear detection, segmentation and morphometric profiling are essential in helping us further understand the relationship between histology and patient outcome. To drive innovation in this area, we setup a community-wide challenge using the largest available dataset of its kind to assess nuclear segmentation and cellular composition. Our challenge, named CoNIC, stimulated the development of reproducible algorithms for cellular recognition with real-time result inspection on public leaderboards. We conducted an extensive post-challenge analysis based on the top-performing models using 1,658 whole-slide images of colon tissue. With around 700 million detected nuclei per model, associated features were used for dysplasia grading and survival analysis, where we demonstrated that the challenge's improvement over the previous state-of-the-art led to significant boosts in downstream performance. Our findings also suggest that eosinophils and neutrophils play an important role in the tumour microevironment. We release challenge models and WSI-level results to foster the development of further methods for biomarker discovery.
Simon Graham, Quoc Dang Vu, Mostafa Jahanifar, Martin Weigert 0001, Jun Zhang 0018, Sen Yang 0006, Jinxi Xiang, Josef Lorenz Rumberger, Elias Baumann, Peter Hirsch 0001, Chenyang Hong, Angelica I. Avilés-Rivero, Ayushi Jain, Heeyoung Ahn, Yiyu Hong, Hussam Azzuni, Min Xu 0009, Mohammad Yaqub, Marie-Claire Blache, Benoît Piégu, Bertrand Vernay, Tim Scherr, Moritz Böhland, Katharina Löffler, Weiqin Ying, Chixin Wang, David R. J. Snead, Shan E Ahmed Raza, Fayyaz ul Amir Afsar Minhas, Nasir M. Rajpoot
Medical Image Anal.7
2023 RLogist: Fast Observation Strategy on Whole-Slide Images with Deep Reinforcement Learning
abstract
Whole-slide images (WSI) in computational pathology have high resolution with gigapixel size, but are generally with sparse regions of interest, which leads to weak diagnostic relevance and data inefficiency for each area in the slide. Most of the existing methods rely on a multiple instance learning framework that requires densely sampling local patches at high magnification. The limitation is evident in the application stage as the heavy computation for extracting patch-level features is inevitable. In this paper, we develop RLogist, a benchmarking deep reinforcement learning (DRL) method for fast observation strategy on WSIs. Imitating the diagnostic logic of human pathologists, our RL agent learns how to find regions of observation value and obtain representative features across multiple resolution levels, without having to analyze each part of the WSI at the high magnification. We benchmark our method on two whole-slide level classification tasks, including detection of metastases in WSIs of lymph node sections, and subtyping of lung cancer. Experimental results demonstrate that RLogist achieves competitive classification performance compared to typical multiple instance learning algorithms, while having a significantly short observation path. In addition, the observation path given by RLogist provides good decision-making interpretability, and its ability of reading path navigation can potentially be used by pathologists for educational/assistive purposes. Our code is available at: https://github.com/tencent-ailab/RLogist.
Boxuan Zhao, Jun Zhang 0018, Deheng Ye, Jian Cao 0001, Xiao Han 0011, Qiang Fu 0016, Wei Yang 0032
AAAI2
2023 Exploring Low-Rank Property in Multiple Instance Learning for Whole Slide Image Classification
Jinxi Xiang, Jun Zhang 0018
ICLR2
2023 MIMT: Masked Image Modeling Transformer for Video Compression
Jinxi Xiang, Kuan Tian, Jun Zhang 0018
ICLR3
2023 Forensic Histopathological Recognition via a Context-Aware MIL Network Powered by Self-supervised Contrastive Learning
Jun Zhang 0018, Xinggong Liang, Zeyi Hao, Kehan Li 0005, Fan Wang 0023, Zhenyuan Wang, Chunfeng Lian
MICCAI (6)2
2023 Dynamic Low-Rank Instance Adaptation for Universal Neural Image Compression
Yue Lv, Jinxi Xiang, Jun Zhang 0018, Wenming Yang, Xiao Han 0011, Wei Yang 0032
ACM Multimedia3
2023 Towards Real-Time Neural Video Codec for Cross-Platform Application Using Calibration Information
abstract
The state-of-the-art neural video codecs have outperformed the most sophisticated traditional codecs in terms of rate-distortion (RD) performance in certain cases. However, utilizing them for practical applications is still challenging for two major reasons. 1) Cross-platform computational errors resulting from floating point operations can lead to inaccurate decoding of the bitstream. 2) The high computational complexity of the encoding and decoding process poses a challenge in achieving real-time performance. In this paper, we propose a real-time cross-platform neural video codec, which is capable of efficiently decoding (25FPS) of 720P video bitstream from other encoding platforms on a consumer-grade GPU (e.g., NVIDIA RTX 2080). First, to solve the problem of inconsistency of codec caused by the uncertainty of floating point calculations across platforms, we design a calibration transmitting system to guarantee the consistent quantization of entropy parameters between the encoding and decoding stages. The parameters that may have transboundary quantization between encoding and decoding are identified in the encoding stage, and their coordinates will be delivered by auxiliary transmitted bitstream. By doing so, these inconsistent parameters can be processed properly in the decoding stage. Furthermore, to reduce the bitrate of the auxiliary bitstream, we rectify the distribution of entropy parameters using a piecewise Gaussian constraint. Second, to match the computational limitations on the decoding side for real-time video codec, we design a lightweight model. A series of efficiency techniques, such as model pruning, motion downsampling, and arithmetic coding skipping, enable our model to achieve 25 FPS decoding speed on NVIDIA RTX 2080 GPU. Experimental results demonstrate that our model can achieve real-time decoding of 720P videos while encoding on another platform. Furthermore, the real-time model brings up to a maximum of 24.2% BD-rate improvement from the perspective of PSNR with the anchor H.265 (medium).
Kuan Tian, Yonghang Guan, Jinxi Xiang, Jun Zhang 0018, Xiao Han 0011, Wei Yang 0032
ACM Multimedia4
2023 Graph-Based Self-Learning for Robust Person Re-identification
abstract
Existing deep learning approaches for person re-identification (Re-ID) mostly rely on large-scale and well-annotated training data. However, human-annotated labels are prone to label noise in real-world applications. Previous person Re-ID works mainly focus on random label noise, which doesn’t properly reflect the characteristic of label noise in practical human-annotated process. In this work, we find the visual ambiguity noise is more common and reasonable noise assumption in annotation of person Re-ID. To handle the kind of noise, we propose a simple and effective robust person Re-ID framework, namely Graph-Based Self-Learning (GBSL), to iteratively learn discriminative representation and rectify noisy labels with limited annotated samples for each identity. Meanwhile, considering the practical annotation process in person Re-ID, we further extend the visual ambiguity noise assumption and propose a type of more practical label noise in person Re-ID, namely the tracklet-level label noise (TLN). Without modifying network architecture or loss function, our approach significantly improves the robustness against label noise of the Re-ID system. Our model obtains competitive performance with training data corrupted by various types of label noise and outperforms the existing methods for robust Re-ID on public benchmarks.
Yuqiao Xian, Jinrui Yang, Fufu Yu, Jun Zhang 0018, Xing Sun 0001
WACV4
2023 CLC-Net: Contextual and local collaborative network for lesion segmentation in diabetic retinopathy images
Yuqi Fang, Sen Yang 0006, Delong Zhu 0001, Jing Zhang 0051, Jun Zhang 0018, Jun Cheng 0003, Raymond Kai-Yu Tong, Xiao Han 0011
Neurocomputing7
2023 RetCCL: Clustering-guided contrastive learning for whole-slide image retrieval
Yuexi Du, Sen Yang 0006, Jun Zhang 0018, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011
Medical Image Anal.4
2023 A generalizable and robust deep learning algorithm for mitosis detection in multicenter breast histopathological images
Jun Zhang 0018, Sen Yang 0006, Jingxi Xiang, Feng Luo 0003, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011
Medical Image Anal.2
2023 Merging nucleus datasets by correlation-based cross-training
Jun Zhang 0018, Sen Yang 0006, Junzhou Huang, Wei Yang 0032, Xiao Han 0011
Medical Image Anal.2
2023 High-Order Correlation-Guided Slide-Level Histology Retrieval With Self-Supervised Hashing
abstract
Histopathological Whole Slide Images (WSIs) play a crucial role in cancer diagnosis. It is of significant importance for pathologists to search for images sharing similar content with the query WSI, especially in the case-based diagnosis. While slide-level retrieval could be more intuitive and practical in clinical applications, most methods are designed for patch-level retrieval. A few recently unsupervised slide-level methods only focus on integrating patch features directly, without perceiving slide-level information, and thus severely limits the performance of WSI retrieval. To tackle the issue, we propose a High-Order Correlation-Guided Self-Supervised Hashing-Encoding Retrieval (HSHR) method. Specifically, we train an attention-based hash encoder with slide-level representation in a self-supervised manner, enabling it to generate more representative slide-level hash codes of cluster centers and assign weights for each. These optimized and weighted codes are leveraged to establish a similarity-based hypergraph, in which a hypergraph-guided retrieval module is adopted to explore high-order correlations in the multi-pairwise manifold to conduct WSI retrieval. Extensive experiments on multiple TCGA datasets with over 24,000 WSIs spanning 30 cancer subtypes demonstrate that HSHR achieves state-of-the-art performance compared with other unsupervised histology WSI retrieval methods.
Shengrui Li, Jun Zhang 0018, Ting Yu 0004, Ji Zhang 0001, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Markov chain-based frequency correlation processing algorithm for wideband DOA estimation
Jun Zhang 0018, Ming Bao, Zhifei Chen, Hong Hou, Jianhua Yang 0005
Signal Process.1
2022 Node-aligned Graph Convolutional Network for Whole-slide Image Representation and Classification
abstract
The large-scale whole-slide images (WSIs) facilitate the learning-based computational pathology methods. However, the gigapixel size of WSIs makes it hard to train a conventional model directly. Current approaches typically adopt multiple-instance learning (MIL) to tackle this problem. Among them, MIL combined with graph convolutional network (GCN) is a significant branch, where the sampled patches are regarded as the graph nodes to further discover their correlations. However, it is difficult to build correspondence across patches from different WSIs. Therefore, most methods have to perform non-ordered node pooling to generate the bag-level representation. Direct non-ordered pooling will lose much structural and contextual information, such as patch distribution and heterogeneous patterns, which is critical for WSI representation. In this paper, we propose a hierarchical global-to-local clustering strategy to build a Node-Aligned GCN (NAGCN) to represent WSI with rich local structural information as well as global distribution. We first deploy a global clustering operation based on the instance features in the dataset to build the correspondence across different WSIs. Then, we perform a local clustering-based sampling strategy to select typical instances belonging to each cluster within the WSI. Finally, we employ the graph convolution to obtain the representation. Since our graph construction strategy ensures the alignment among different WSIs, WSI-level representation can be easily generated and used for the subsequent classification. The experiment results on two cancer subtype classification datasets demonstrate our method achieves better performance compared with the state-of-the-art methods.
Yonghang Guan, Jun Zhang 0018, Kuan Tian, Sen Yang 0006, Pei Dong, Jinxi Xiang, Wei Yang 0032, Junzhou Huang, Yuyao Zhang 0005, Xiao Han 0011
CVPR2
2022 PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining
abstract
Large-scale vision-language pre-training has achieved promising results on downstream tasks. Existing methods highly rely on the assumption that the image-text pairs crawled from the Internet are in perfect one-to-one correspondence. However, in real scenarios, this assumption can be difficult to hold: the text description, obtained by crawling the affiliated metadata of the image, often suffers from the semantic mismatch and the mutual compatibility. To address these issues, we introduce PyramidCLIP, which constructs an input pyramid with different semantic levels for each modality, and aligns visual elements and linguistic elements in the form of hierarchy via peer-level semantics alignment and cross-level relation alignment. Furthermore, we soften the loss of negative samples (unpaired samples) so as to weaken the strict constraint during the pre-training stage, thus mitigating the risk of forcing the model to distinguish compatible negative pairs. Experiments on five downstream tasks demonstrate the effectiveness of the proposed PyramidCLIP. In particular, with the same amount of 15 million pre-training image-text pairs, PyramidCLIP exceeds CLIP on ImageNet zero-shot classification top-1 accuracy by 10.6%/13.2%/10.0% with ResNet50/ViT-B32/ViT-B16 based image encoder respectively. When scaling to larger datasets, PyramidCLIP achieves the state-of-the-art results on several downstream tasks. In particular, the results of PyramidCLIP-ResNet50 trained on 143M image-text pairs surpass that of CLIP using 400M data on ImageNet zero-shot classification task, significantly improving the data efficiency of CLIP.
Jinfeng Liu 0007, Jun Zhang 0018, Ke Li 0015, Rongrong Ji, Chunhua Shen
NeurIPS4
2022 Multi-dataset Training of Transformers for Robust Action Recognition
abstract
We study the task of robust feature representations, aiming to generalize well on multiple datasets for action recognition. We build our method on Transformers for its efficacy. Although we have witnessed great progress for video action recognition in the past decade, it remains challenging yet valuable how to train a single model that can perform well across multiple datasets. Here, we propose a novel multi-dataset training paradigm, MultiTrain, with the design of two new loss terms, namely informative loss and projection loss, aiming tolearn robust representations for action recognition. In particular, the informative loss maximizes the expressiveness of the feature embedding while the projection loss for each dataset mines the intrinsic relations between classes across datasets. We verify the effectiveness of our method on five challenging datasets, Kinetics-400, Kinetics-700, Moments-in-Time, Activitynet and Something-something-v2 datasets. Extensive experimental results show that our method can consistently improve state-of-the-art performance. Code and models are released.
Junwei Liang 0001, Enwei Zhang, Jun Zhang 0018, Chunhua Shen
NeurIPS3
2022 Text-Adaptive Multiple Visual Prototype Matching for Video-Text Retrieval
abstract
Cross-modal retrieval between videos and texts has gained increasing interest because of the rapid emergence of videos on the web. Generally, a video contains rich instance and event information and the query text only describes a part of the information. Thus, a video can have multiple different text descriptions and queries. We call it the Video-Text Correspondence Ambiguity problem. Current techniques mostly concentrate on mining local or multi-level alignment between contents of video and text (e.g., object to entity and action to verb). It is difficult for these methods to alleviate video-text correspondence ambiguity by describing a video using only one feature, which is required to be matched with multiple different text features at the same time. To address this problem, we propose a Text-Adaptive Multiple Visual Prototype Matching Model. It automatically captures multiple prototypes to describe a video by adaptive aggregation on video token features. Given a query text, the similarity is determined by the most similar prototype to find correspondence in the video, which is called text-adaptive matching. To learn diverse prototypes for representing the rich information in videos, we propose a variance loss to encourage different prototypes to attend to different contents of the video. Our method outperforms the state-of-the-art methods on four public video retrieval datasets.
Chengzhi Lin, Ancong Wu, Junwei Liang 0001, Jun Zhang 0018, Wenhang Ge, Wei-Shi Zheng 0001, Chunhua Shen
NeurIPS4
2022 SCL-WC: Cross-Slide Contrastive Learning for Weakly-Supervised Whole-Slide Image Classification
abstract
Weakly-supervised whole-slide image (WSI) classification (WSWC) is a challenging task where a large number of unlabeled patches (instances) exist within each WSI (bag) while only a slide label is given. Despite recent progress for the multiple instance learning (MIL)-based WSI analysis, the major limitation is that it usually focuses on the easy-to-distinguish diagnosis-positive regions while ignoring positives that occupy a small ratio in the entire WSI. To obtain more discriminative features, we propose a novel weakly-supervised classification method based on cross-slide contrastive learning (called SCL-WC), which depends on task-agnostic self-supervised feature pre-extraction and task-specific weakly-supervised feature refinement and aggregation for WSI-level prediction. To enable both intra-WSI and inter-WSI information interaction, we propose a positive-negative-aware module (PNM) and a weakly-supervised cross-slide contrastive learning (WSCL) module, respectively. The WSCL aims to pull WSIs with the same disease types closer and push different WSIs away. The PNM aims to facilitate the separation of tumor-like patches and normal ones within each WSI. Extensive experiments demonstrate state-of-the-art performance of our method in three different classification tasks (e.g., over 2% of AUC in Camelyon16, 5% of F1 score in BRACS, and 3% of AUC in DiagSet). Our method also shows superior flexibility and scalability in weakly-supervised localization and semi-supervised classification experiments (e.g., first place in the BRIGHT challenge). Our code will be available at https://github.com/Xiyue-Wang/SCL-WC.
Jinxi Xiang, Jun Zhang 0018, Sen Yang 0006, Zhongyi Yang, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011
NeurIPS3
2022 Multi-level attention graph neural network based on co-expression gene modules for disease diagnosis and prognosis
abstract
MOTIVATION: Advanced deep learning techniques have been widely applied in disease diagnosis and prognosis with clinical omics, especially gene expression data. In the regulation of biological processes and disease progression, genes often work interactively rather than individually. Therefore, investigating gene association information and co-functional gene modules can facilitate disease state prediction. RESULTS: To explore the gene modules and inter-gene relational information contained in the omics data, we propose a novel multi-level attention graph neural network (MLA-GNN) for disease diagnosis and prognosis. Specifically, we format omics data into co-expression graphs via weighted correlation network analysis, and then construct multi-level graph features, finally fuse them through a well-designed multi-level graph feature fully fusion module to conduct predictions. For model interpretation, a novel full-gradient graph saliency mechanism is developed to identify the disease-relevant genes. MLA-GNN achieves state-of-the-art performance on transcriptomic data from TCGA-LGG/TCGA-GBM and proteomic data from coronavirus disease 2019 (COVID-19)/non-COVID-19 patient sera. More importantly, the relevant genes selected by our model are interpretable and are consistent with the clinical understanding. AVAILABILITYAND IMPLEMENTATION: The codes are available at https://github.com/TencentAILabHealthcare/MLA-GNN.
Xiaohan Xing, Fan Yang 0081, Jun Zhang 0018, Yu Zhao 0009, Mingxuan Gao, Junzhou Huang, Jianhua Yao 0001
Bioinform.4
2022 Adaptive context- and scale-aware aggregation with feature alignment for one-shot object detection
Chengdong Dong, Jun Zhang 0018, Hangguan Shan, Eryun Liu
Neurocomputing3
2022 Multitarget Tracking Based on Dynamic Bayesian Network With Reparameterized Approximate Variational Inference
abstract
Multitarget tracking (MTT) is an important component of situation-awareness based on the Internet of Things (IoT). Existing algorithms mainly focus on tracking based on conventional measurements, e.g., bearings or ranges. However, measurement parameter estimations are considered in isolation, limiting the accuracy and resolution of MTT, and the related data association is an NP-hard multidimensional assignment problem. In this article, we develop a new one-step MTT algorithm based on a novel dynamic Bayesian network (DBN), i.e., DBNMTT. The new MTT algorithm directly infers target states from the raw measurement data by fusing the array signal model, the signal propagation model, and the motion model. In this new DBNMTT framework, we treat target states and conventional measurements, such as bearings and target energies as hidden random variables. The posterior joint probability optimization problem is translated into the problem of graphical model learning. In this way, we can improve the accuracy and resolution of MTT and convert the NP-hard data association problem to a hidden variable learning problem. For nonconjugate models in the DBNMTT, we develop a novel reparameterized approximation variational inference (ReAVI) approach to solve the learning problem. The ReAVI converts nonconjugate models to conjugate models with new parameters and reuses the mean-field algorithm. The performance of our proposed new MTT method, namely, DBNMTT based on ReAVI (DBNMTT-ReAVI), is analyzed on extensive simulations in challenging scenarios. The simulation results show that the DBNMTT-ReAVI algorithm is superior to conventional measurement-based MTT algorithms in several aspects, including the success probability, convergence, resolution, and accuracy.
Wenqiong Zhang, Jun Zhang 0018, Ming Bao, Xiao-Ping Zhang 0002, Xiaodong Li 0002
IEEE Internet Things J.2
2022 Transformer-based unsupervised contrastive learning for histopathological image classification
Sen Yang 0006, Jun Zhang 0018, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011
Medical Image Anal.3
2022 PFVNet: A Partial Fingerprint Verification Network Learned From Large Fingerprint Matching
abstract
With the decreasing size of fingerprint scanners in portable devices, e.g., mobile phone and smart watch, partial fingerprint recognition has become a challenging and urgently needed technique due to the limited features contained in small area as well as the large rotation and translation between query and reference images. Deep learning as a powerful modeling method has advanced the research progress of fingerprint recognition, but it still suffers from the lack of labeled data in the scenario of partial fingerprint matching. In this paper, we propose a novel partial fingerprint verification network (PFVNet) based on spatial transformer network (STN) and the local self-attention mechanism. Our model can be trained end-to-end and learn multi-level fingerprint features automatically. To alleviate the data annotation work, the model is trained in a self-supervision and domain adaptation manner with data generated from large fingerprint image matching. The experimental results compared with other methods on FVC2006 DB1 dataset and in-house datasets (i.e., ZJUPartial database) show that our method achieves state-of-the-art performance, and also robust to different types of scanners.
Jun Zhang 0018, Liaojun Pang, Eryun Liu
IEEE Trans. Inf. Forensics Secur.2
2022 Big-Hypergraph Factorization Neural Network for Survival Prediction From Whole Slide Image
abstract
Survival prediction for patients based on histopa- thological whole-slide images (WSIs) has attracted increasing attention in recent years. Due to the massive pixel data in a single WSI, fully exploiting cell-level structural information (e.g., stromal/tumor microenvironment) from the gigapixel WSI is challenging. Most of the current studies resolve the problem by sampling limited image patches to construct a graph-based model (e.g., hypergraph). However, the sampling scale is a critical bottleneck since it is a fundamental obstacle of broadening samples for transductive learning. To overcome the limitation of the sampling scale for constructing a big hypergraph model, we propose a factorization neural network that embeds the correlation among large-scale vertices and hyperedges into two low-dimensional latent semantic spaces separately, empowering the dense sampling. Thanks to the compressed low-dimensional correlation embedding, the hypergraph convolutional layers generate the high-order global representation for each WSI. To minimize the effect of the uncertainty data as well as to achieve the metric-driven learning, we also propose a multi-level ranking supervision to enable the network learning by a queue of patients on the global horizon. Extensive experiments are conducted on three public carcinoma datasets (i.e., LUSC, GBM, and NLST), and the quantitative results demonstrate the proposed method outperforms state-of-the-art methods across-the-board.
Donglin Di, Jun Zhang 0018, Fuqiang Lei, Qi Tian 0001, Yue Gao 0002
IEEE Trans. Image Process.2
2022 Conditional Feature Learning Based Transformer for Text-Based Person Search
abstract
Text-based person search aims at retrieving the target person in an image gallery using a descriptive sentence of that person. The core of this task is to calculate a similarity score between the pedestrian image and description, which requires inferring the complex latent correspondence between image sub-regions and textual phrases at different scales. Transformer is an intuitive way to model the complex alignment by its self-attention mechanism. Most previous Transformer-based methods simply concatenate image region features and text features as input and learn a cross-modal representation in a brute force manner. Such weakly supervised learning approaches fail to explicitly build alignment between image region features and text features, causing an inferior feature distribution. In this paper, we present CFLT, Conditional Feature Learning based Transformer. It maps the sub-regions and phrases into a unified latent space and explicitly aligns them by constructing conditional embeddings where the feature of data from one modality is dynamically adjusted based on the data from the other modality. The output of our CFLT is a set of similarity scores for each sub-region or phrase rather than a cross-modal representation. Furthermore, we propose a simple and effective multi-modal re-ranking method named Re-ranking scheme by Visual Conditional Feature (RVCF). Benefit from the visual conditional feature and better feature distribution in our CFLT, the proposed RVCF achieves significant performance improvement. Experimental results show that our CFLT outperforms the state-of-the-art methods by 7.03% in terms of top-1 accuracy and 5.01% in terms of top-5 accuracy on the text-based person search dataset.
Chenyang Gao, Guanyu Cai, Xinyang Jiang, Feng Zheng 0001, Jun Zhang 0018, Yifei Gong, Fangzhou Lin, Xing Sun 0001, Xiang Bai
IEEE Trans. Image Process.5
2022 Knowledge-Based Representation Learning for Nucleus Instance Classification From Histopathological Images
abstract
The classification of nuclei in H&E-stained histopathological images is a fundamental step in the quantitative analysis of digital pathology. Most existing methods employ multi-class classification on the detected nucleus instances, while the annotation scale greatly limits their performance. Moreover, they often downplay the contextual information surrounding nucleus instances that is critical for classification. To explicitly provide contextual information to the classification model, we design a new structured input consisting of a content-rich image patch and a target instance mask. The image patch provides rich contextual information, while the target instance mask indicates the location of the instance to be classified and emphasizes its shape. Benefiting from our structured input format, we propose Structured Triplet for representation learning, a triplet learning framework on unlabelled nucleus instances with customized positive and negative sampling strategies. We pre-train a feature extraction model based on this framework with a large-scale unlabeled dataset, making it possible to train an effective classification model with limited annotated data. We also add two auxiliary branches, namely the attribute learning branch and the conventional self-supervised learning branch, to further improve its performance. As part of this work, we will release a new dataset of H&E-stained pathology images with nucleus instance masks, containing 20,187 patches of size 1024 ×1024 , where each patch comes from a different whole-slide image. The model pre-trained on this dataset with our framework significantly reduces the burden of extensive labeling. We show a substantial improvement in nucleus classification accuracy compared with the state-of-the-art methods.
Jun Zhang 0018, Sen Yang 0006, Wei Yang 0032, Junzhou Huang, Xiao Han 0011
IEEE Trans. Medical Imaging2
2021 Diagnose Like A Pathologist: Weakly-Supervised Pathologist-Tree Network for Slide-Level Immunohistochemical Scoring
abstract
The immunohistochemistry (IHC) test of biopsy tissue is crucial to develop targeted treatment and evaluate prognosis for cancer patients. The IHC staining slide is usually digitized into the whole-slide image (WSI) with gigapixels for quantitative image analysis. To perform a whole image prediction (e.g., IHC scoring, survival prediction, and cancer grading) from this kind of high-dimensional image, algorithms are often developed based on multi-instance learning (MIL) framework. However, the multi-scale information of WSI and the associations among instances are not well explored in existing MIL based studies. Inspired by the fact that pathologists jointly analyze visual fields at multiple powers of objective for diagnostic predictions, we propose a Pathologist-Tree Network (PTree-Net) to sparsely model the WSI efficiently in multi-scale manner. Specifically, we propose a Focal-Aware Module (FAM) that can approximately estimate diagnosis-related regions with an extractor trained using the thumbnail of WSI. With the initial diagnosis-related regions, we hierarchically model the multi-scale patches in a tree structure, where both the global and local information can be captured. To explore this tree structure in an end-to-end network, we propose a patch Relevance-enhanced Graph Convolutional Network (RGCN) to explicitly model the correlations of adjacent parent-child nodes, accompanied by patch relevance to exploit the implicit contextual information among distant nodes. In addition, tree-based self-supervision is devised to improve representation learning and suppress irrelevant instances adaptively. Extensive experiments are performed on a large-scale IHC HER2 dataset. The ablation study confirms the effectiveness of our design, and our approach outperforms state-of-the-art by a large margin.
Zhen Chen 0013, Jun Zhang 0018, Shuanlong Che, Junzhou Huang, Xiao Han 0011, Yixuan Yuan
AAAI2
2021 Minimizing Labeling Cost for Nuclei Instance Segmentation and Classification with Cross-domain Images and Weak Labels
abstract
Nucleus instance segmentation and classification in histopathological images is an essential prerequisite in pathology diagnosis/prognosis. However, nucleus annotations (e.g., segmentation and labeling) require domain experts, and annotating nuclei at pixel-level is time-consuming and labor-intensive. Moreover, nuclei from different cancer types vary in shapes and appearances. These inter-cancer variations require careful annotations for specific cancer types. Therefore, to minimize the labeling cost, we propose a novel application that considers each cancer type as an individual domain and apply domain adaptation techniques to improve the segmentation/classification performance among different cancer types. Unlike the previous studies that focus on unsupervised or weakly-supervised domain adaptation independently, we would like to discover what kinds of labeling can achieve the most cost-effective domain adaptation performance in nucleus instance segmentation and classification. Specifically, we propose a unified framework that is applicable to different level annotations: no annotations, image-level, and point-level annotations. Cyclic adaptation with pseudo labels and adversarial discriminator are utilized for unsupervised domain alignment. Image-level or point-level annotations are additionally adopted to supervise the nucleus classification and refine the pseudo labels. Experiments demonstrate the effectiveness and efficacy of the proposed framework (jointly using unsupervised and weakly supervised learning) on adapting the segmentation and classification model from one cancer type to 18 other cancer types.
Siqi Yang 0001, Jun Zhang 0018, Junzhou Huang, Brian C. Lovell, Xiao Han 0011
AAAI2
2021 An Interpretable Multi-Level Enhanced Graph Attention Network for Disease Diagnosis with Gene Expression Data
abstract
Clinical omics, especially gene expression data, have been widely studied and successfully applied for disease diagnosis using machine learning techniques. As genes often work interactively rather than individually, investigating co-functional gene modules can improve our understanding of disease mechanisms and facilitate disease state prediction. To this end, we in this paper propose a novel Multi-Level Enhanced Graph ATtention (MLE-GAT) network to explore the gene modules and intergene relational information contained in the omics data. In specific, we first format the omics data of each patient into co-expression graphs using weighted correlation network analysis (WGCNA) and then feed them to a well-designed multi-level graph feature fully fusion (MGFFF) module for disease diagnosis. For model interpretation, we develop a novel full-gradient graph saliency (FGS) mechanism to identify the disease-relevant genes. Comprehensive experiments show that our proposed MLE-GAT achieves state-of-the-art performance on transcriptomics data from TCGA-LGG/TCGA-GBM and proteomics data from COVID-19/non-COVID-19 patient sera.
Xiaohan Xing, Fan Yang 0081, Jun Zhang 0018, Yu Zhao 0009, Mingxuan Gao, Junzhou Huang, Jianhua Yao 0001
BIBM4
2021 Learning 3D Shape Feature for Texture-Insensitive Person Re-Identification
abstract
It is well acknowledged that person re-identification (person ReID) highly relies on visual texture information like clothing. Despite significant progress has been made in recent years, texture-confusing situations like clothing changing and persons wearing the same clothes receive little attention from most existing ReID methods. In this paper, rather than relying on texture based information, we propose to improve the robustness of person ReID against clothing texture by exploiting the information of a person’s 3D shape. Existing shape learning schemas for person ReID either ignore the 3D information of a person, or require extra physical devices to collect 3D source data. Differently, we propose a novel ReID learning framework that directly extracts a texture-insensitive 3D shape embedding from a 2D image by adding 3D body reconstruction as an auxiliary task and regularization, called 3D Shape Learning (3DSL). The 3D reconstruction based regularization forces the ReID model to decouple the 3D shape information from the visual texture, and acquire discriminative 3D shape ReID features. To solve the problem of lacking 3D ground truth, we design an adversarial self-supervised projection (ASSP) model, performing 3D reconstruction without ground truth. Extensive experiments on common ReID datasets and texture-confusing datasets validate the effectiveness of our model.
Xinyang Jiang, Fudong Wang 0001, Jun Zhang 0018, Feng Zheng 0001, Xing Sun 0001, Wei-Shi Zheng 0001
CVPR4
2021 Ask&Confirm: Active Detail Enriching for Cross-Modal Retrieval with Partial Query
abstract
Text-based image retrieval has seen considerable progress in recent years. However, the performance of existing methods suffers in real life since the user is likely to provide an incomplete description of an image, which often leads to results filled with false positives that fit the incomplete description. In this work, we introduce the partial-query problem and extensively analyze its influence on text-based image retrieval. Previous interactive methods tackle the problem by passively receiving users’ feedback to supplement the incomplete query iteratively, which is time-consuming and requires heavy user effort. Instead, we propose a novel retrieval framework that conducts the interactive process in an Ask-and-Confirm fashion, where AI actively searches for discriminative details missing in the current query, and users only need to confirm AI’s proposal. Specifically, we propose an object-based interaction to make the interactive retrieval more user-friendly and present a reinforcement-learning-based policy to search for discriminative objects. Furthermore, since fully-supervised training is often infeasible due to the difficulty of obtaining human-machine dialog data, we present a weakly-supervised training strategy that needs no human-annotated dialogs other than a text-image dataset. Experiments show that our framework significantly improves the performance of text-based image retrieval. Code is available at https://github.com/CuthbertCai/Ask-Confirm.
Guanyu Cai, Jun Zhang 0018, Xinyang Jiang, Yifei Gong, Lianghua He, Fufu Yu, Feiyue Huang, Xing Sun 0001
ICCV2
2021 Multi-modal Multi-instance Learning Using Weakly Correlated Histopathological Images and Tabular Clinical Information
Fan Yang 0081, Xiaohan Xing, Yu Zhao 0009, Jun Zhang 0018, Yueping Liu, Mengxue Han, Junzhou Huang, Liansheng Wang 0002, Jianhua Yao 0001
MICCAI (8)5
2021 DT-MIL: Deformable Transformer for Multi-instance Learning on Histopathological Image
Fan Yang 0081, Yu Zhao 0009, Xiaohan Xing, Jun Zhang 0018, Mingxuan Gao, Junzhou Huang, Liansheng Wang 0002, Jianhua Yao 0001
MICCAI (8)5
2021 TransPath: Transformer-Based Self-supervised Learning for Histopathological Image Classification
Sen Yang 0006, Jun Zhang 0018, Jing Zhang 0051, Junzhou Huang, Wei Yang 0032, Xiao Han 0011
MICCAI (8)3
2021 Joint fully convolutional and graph convolutional networks for weakly-supervised segmentation of pathology images
Jun Zhang 0018, Zhiyuan Hua, Kezhou Yan, Kuan Tian, Jianhua Yao 0001, Eryun Liu, Mingxia Liu 0001, Xiao Han 0011
Medical Image Anal.1
2021 Group-Wise Learning for Aurora Image Classification With Multiple Representations
abstract
In conventional aurora image classification methods, it is general to employ only one single feature representation to capture the morphological characteristics of aurora images, which is difficult to describe the complicated morphologies of different aurora categories. Although several studies have proposed to use multiple feature representations, the inherent correlation among these representations are usually neglected. To address this problem, we propose a group-wise learning (GWL) method for the automatic aurora image classification using multiple representations. Specifically, we first extract the multiple feature representations for aurora images, and then construct a graph in each of multiple feature spaces. To model the correlation among different representations, we partition multiple graphs into several groups via a clustering algorithm. We further propose a GWL model to automatically estimate class labels for aurora images and optimal weights for the multiple representations in a data-driven manner. Finally, we develop a label fusion approach to make a final classification decision for new testing samples. The proposed GWL method focuses on the diverse properties of multiple feature representations, by clustering the correlated representations into the same group. We evaluate our method on an aurora image data set that contains 12 682 aurora images from 19 days. The experimental results demonstrate that the proposed GWL method achieves approximately 6% improvement in terms of classification accuracy, compared to the methods using a single feature representation.
Jun Zhang 0018, Mingxia Liu 0001, Ke Lu 0002, Yue Gao 0002
IEEE Trans. Cybern.1
2020 Predicting Lymph Node Metastasis Using Histopathological Images Based on Multiple Instance Learning With Deep Graph Convolution
abstract
Multiple instance learning (MIL) is a typical weakly-supervised learning method where the label is associated with a bag of instances instead of a single instance. Despite extensive research over past years, effectively deploying MIL remains an open and challenging problem, especially when the commonly assumed standard multiple instance (SMI) assumption is not satisfied. In this paper, we propose a multiple instance learning method based on deep graph convolutional network and feature selection (FS-GCN-MIL) for histopathological image classification. The proposed method consists of three components, including instance-level feature extraction, instance-level feature selection, and bag-level classification. We develop a self-supervised learning mechanism to train the feature extractor based on a combination model of variational autoencoder and generative adversarial network (VAE-GAN). Additionally, we propose a novel instance-level feature selection method to select the discriminative instance features. Furthermore, we employ a graph convolutional network (GCN) for learning the bag-level representation and then performing the classification. We apply the proposed method in the prediction of lymph node metastasis using histopathological images of colorectal cancer. Experimental results demonstrate that the proposed method achieves superior performance compared to the state-of-the-art methods.
Yu Zhao 0009, Fan Yang 0081, Yuqi Fang, Hailing Liu, Niyun Zhou, Jun Zhang 0018, Sen Yang 0006, Bjoern Menze, Xinjuan Fan, Jianhua Yao 0001
CVPR6
2020 GATCluster: Self-supervised Gaussian-Attention Network for Image Clustering
Chuang Niu, Jun Zhang 0018, Ge Wang 0001, Jimin Liang
ECCV (25)2
2020 Do Not Disturb Me: Person Re-identification Under the Interference of Other Pedestrians
Shizhen Zhao, Changxin Gao, Jun Zhang 0018, Hao Cheng 0012, Chuchu Han, Xinyang Jiang, Wei-Shi Zheng 0001, Nong Sang, Xing Sun 0001
ECCV (6)3
2020 Ranking-Based Survival Prediction on Histopathological Whole-Slide Images
Donglin Di, Shengrui Li, Jun Zhang 0018, Yue Gao 0002
MICCAI (5)3
2020 Deep Active Learning for Breast Cancer Segmentation on Immunohistochemistry Images
Haocheng Shen, Kuan Tian, Pei Dong, Jun Zhang 0018, Kezhou Yan, Shannon Che, Jianhua Yao 0001, Pifu Luo, Xiao Han 0011
MICCAI (5)4
2020 Weakly-Supervised Nucleus Segmentation Based on Point Annotations: A Coarse-to-Fine Self-Stimulated Learning Strategy
Kuan Tian, Jun Zhang 0018, Haocheng Shen, Kezhou Yan, Pei Dong, Jianhua Yao 0001, Shannon Che, Pifu Luo, Xiao Han 0011
MICCAI (5)2
2020 Context-guided fully convolutional networks for joint craniomaxillofacial bone segmentation and landmark digitization
Jun Zhang 0018, Mingxia Liu 0001, Li Wang 0026, Peng Yuan 0001, Jianfu Li, Steve G. Shen, Ken-Chung Chen, James J. Xia, Dinggang Shen
Medical Image Anal.1
2020 Hierarchical Fully Convolutional Network for Joint Atrophy Localization and Alzheimer's Disease Diagnosis Using Structural MRI
abstract
Structural magnetic resonance imaging (sMRI) has been widely used for computer-aided diagnosis of neurodegenerative disorders, e.g., Alzheimer's disease (AD), due to its sensitivity to morphological changes caused by brain atrophy. Recently, a few deep learning methods (e.g., convolutional neural networks, CNNs) have been proposed to learn task-oriented features from sMRI for AD diagnosis, and achieved superior performance than the conventional learning-based methods using hand-crafted features. However, these existing CNN-based methods still require the pre-determination of informative locations in sMRI. That is, the stage of discriminative atrophy localization is isolated to the latter stages of feature extraction and classifier construction. In this paper, we propose a hierarchical fully convolutional network (H-FCN) to automatically identify discriminative local patches and regions in the whole brain sMRI, upon which multi-scale feature representations are then jointly learned and fused to construct hierarchical classification models for AD diagnosis. Our proposed H-FCN method was evaluated on a large cohort of subjects from two independent datasets (i.e., ADNI-1 and ADNI-2), demonstrating good performance on joint discriminative atrophy localization and brain disease diagnosis.
Chunfeng Lian, Mingxia Liu 0001, Jun Zhang 0018, Dinggang Shen
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Weakly Supervised Deep Learning for Brain Disease Prognosis Using MRI and Incomplete Clinical Scores
abstract
As a hot topic in brain disease prognosis, predicting clinical measures of subjects based on brain magnetic resonance imaging (MRI) data helps to assess the stage of pathology and predict future development of the disease. Due to incomplete clinical labels/scores, previous learning-based studies often simply discard subjects without ground-truth scores. This would result in limited training data for learning reliable and robust models. Also, existing methods focus only on using hand-crafted features (e.g., image intensity or tissue volume) of MRI data, and these features may not be well coordinated with prediction models. In this paper, we propose a weakly supervised densely connected neural network (wiseDNN) for brain disease prognosis using baseline MRI data and incomplete clinical scores. Specifically, we first extract multiscale image patches (located by anatomical landmarks) from MRI to capture local-to-global structural information of images, and then develop a weakly supervised densely connected network for task-oriented extraction of imaging features and joint prediction of multiple clinical measures. A weighted loss function is further employed to make full use of all available subjects (even those without ground-truth scores at certain time-points) for network training. The experimental results on 1469 subjects from both ADNI-1 and ADNI-2 datasets demonstrate that our proposed method can efficiently predict future clinical measures of subjects.
Mingxia Liu 0001, Jun Zhang 0018, Chunfeng Lian, Dinggang Shen
IEEE Trans. Cybern.2
2019 Particle Swarm Loss for Lightweight Object Detection
abstract
Currently in object detection, deep learning based detectors are gaining their momentum. However, the supervision involved in the widely-used anchor paradigm within the detection pipeline is inadequate. Traditional object detectors opt for densely picking anchors to increase the training samples for faster convergence and better detection quality. However, dense anchor scheme requires extra computational budget which renders it infeasible for lightweight detectors. To address the problem, inspired by the cognitive consistency, we propose a novel Particle Swarm Loss for lightweight object detection. Experiments upon the MS-COCO challenge show that detectors compensated by PS loss can not only converge faster but also acquire better detection quality than their vanilla versions (YOLOv3 and SSD improves 2.0% and 2.5% on the harsh AP50respectively) without extra computational overhead. In addition, we propose a dapper backbone with high cost-efficiency for the resource-limited scenarios.
Peizhen Zhang, Feng Zheng 0001, Junlong Du, Jun Zhang 0018, Wei-Shi Zheng 0001
ICME4
2019 Multimedia analysis for medical applications
Jun Zhang 0018, Mingxia Liu 0001, Yi Zhen
Multim. Syst.1
2019 Hierarchical Convolutional Neural Networks for Segmentation of Breast Tumors in MRI With Application to Radiogenomics
abstract
Breast tumor segmentation based on dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is a challenging problem and an active area of research. Particular challenges, similarly as in other segmentation problems, include the class-imbalance problem as well as confounding background in DCE-MR images. To address these issues, we propose a mask-guided hierarchical learning (MHL) framework for breast tumor segmentation via fully convolutional networks (FCN). Specifically, we first develop an FCN model to generate a 3D breast mask as the region of interest (ROI) for each image, to remove confounding information from input DCE-MR images. We then design a two-stage FCN model to perform coarse-to-fine segmentation for breast tumors. Particularly, we propose a Dice-Sensitivity-like loss function and a reinforcement sampling strategy to handle the class-imbalance problem. To precisely identify locations of tumors that underwent a biopsy, we further propose an FCN model to detect two landmarks located at two nipples. We finally selected the biopsied tumor based on both identified landmarks and segmentations. We validate our MHL method on 272 patients, achieving a mean Dice similarity coefficient (DSC) of 0.72 which is comparable to mutual DSC between expert radiologists. Using the segmented biopsied tumors, we also demonstrate that the automatically generated masks can be applied to radiogenomics and can identify luminal A subtype from other molecular subtypes with the similar accuracy with the analysis based on semi-manual tumor segmentation.
Jun Zhang 0018, Ashirbani Saha, Zhe Zhu, Maciej A. Mazurowski
IEEE Trans. Medical Imaging1
2018 Multi-channel multi-scale fully convolutional network for 3D perivascular spaces segmentation in 7T MR images
Chunfeng Lian, Jun Zhang 0018, Mingxia Liu 0001, Xiaopeng Zong, Sheng-Che Hung, Weili Lin, Dinggang Shen
Medical Image Anal.2
2018 Landmark-based deep multi-instance learning for brain disease diagnosis
Mingxia Liu 0001, Jun Zhang 0018, Ehsan Adeli-Mosabbeb, Dinggang Shen
Medical Image Anal.2
2018 Multi-task neural networks for joint hippocampus segmentation and clinical score regression
Jifeng Zheng, Feng Yin 0001, Jun Zhang 0018
Multim. Tools Appl.7
2018 Guest Editorial: Large-scale 3D Multimedia Analysis and Applications
Sicheng Zhao, Jun Zhang 0018, Yi Zhen
Multim. Tools Appl.2
2018 Weakly Supervised Semantic Segmentation for Joint Key Local Structure Localization and Classification of Aurora Image
abstract
In this paper, we propose a novel weakly supervised semantic segmentation (WSSS) method that uses image tags as supervision to achieve joint pixel-level localization of the key local structure (KLS) and image-level classification of the aurora images captured by the ground-based optical all-sky imager. First, a patch-scale model (PSM) based on the small-scale structure of aurora is designed to identify the type-specific regions for each training image. Second, a region-scale model is trained with the identified type-specific regions to coarsely localize the KLS from multiple sizes of field of view, based on which the aurora image is classified. Finally, given the predicted image type, the PSM further refines the KLS in a pixel level. By localizing KLS from coarse to fine, the proposed method captures both overall shape with a bottom-up processing and local structure details of aurora in a top-down manner. Extensive experiments on the expert labeled data sets have demonstrated the efficacy of the proposed method in benchmarking with the state-of-the-art WSSS methods.
Chuang Niu, Jun Zhang 0018, Qian Wang 0019, Jimin Liang
IEEE Trans. Geosci. Remote. Sens.2
2018 Anatomical Landmark Based Deep Feature Representation for MR Images in Brain Disease Diagnosis
abstract
Most automated techniques for brain disease diagnosis utilize hand-crafted (e.g., voxel-based or region-based) biomarkers from structural magnetic resonance (MR) images as feature representations. However, these hand-crafted features are usually high-dimensional or require regions-of-interest defined by experts. Also, because of possibly heterogeneous property between the hand-crafted features and the subsequent model, existing methods may lead to sub-optimal performances in brain disease diagnosis. In this paper, we propose a landmark-based deep feature learning (LDFL) framework to automatically extract patch-based representation from MRI for automatic diagnosis of Alzheimer's disease. We first identify discriminative anatomical landmarks from MR images in a data-driven manner, and then propose a convolutional neural network for patch-based deep feature learning. We have evaluated the proposed method on subjects from three public datasets, including the Alzheimer's disease neuroimaging initiative (ADNI-1), ADNI-2, and the minimal interval resonance imaging in alzheimer's disease (MIRIAD) dataset. Experimental results of both tasks of brain disease classification and MR image retrieval demonstrate that the proposed LDFL method improves the performance of disease classification and MR image retrieval.
Mingxia Liu 0001, Jun Zhang 0018, Dong Nie, Pew-Thian Yap, Dinggang Shen
IEEE J. Biomed. Health Informatics2
2017 Deformable Image Registration Based on Similarity-Steered CNN Regression
Xiaohuan Cao, Jianhua Yang 0005, Jun Zhang 0018, Dong Nie, Minjeong Kim 0001, Qian Wang 0001, Dinggang Shen
MICCAI (1)3
2017 Deep Multi-task Multi-channel Learning for Joint Classification and Regression of Brain Status
Mingxia Liu 0001, Jun Zhang 0018, Ehsan Adeli-Mosabbeb, Dinggang Shen
MICCAI (3)2
2017 Joint Craniomaxillofacial Bone Segmentation and Landmark Digitization by Context-Guided Fully Convolutional Networks
Jun Zhang 0018, Mingxia Liu 0001, Li Wang 0026, Peng Yuan 0001, Jianfu Li, Steve G. Shen, Ken-Chung Chen, James J. Xia, Dinggang Shen
MICCAI (2)1
2017 Hypergraph regularized sparse feature learning
Mingxia Liu 0001, Jun Zhang 0018, Xiaochun Guo, Liujuan Cao
Neurocomputing2
2017 Auroral event representation based on the n-ary fusion of multiple oriented energies
Jun Zhang 0018, Qian Wang 0019, Zejun Hu, Mingxia Liu 0001
Neurocomputing1
2017 View-aligned hypergraph learning for Alzheimer's disease diagnosis with incomplete multi-modality data
Mingxia Liu 0001, Jun Zhang 0018, Pew-Thian Yap, Dinggang Shen
Medical Image Anal.2
2017 Multi-view texture classification using hierarchical synthetic images
Jun Zhang 0018, Jimin Liang, Haihong Hu
Multim. Tools Appl.1
2017 Detecting Anatomical Landmarks From Limited Medical Imaging Data Using Two-Stage Task-Oriented Deep Neural Networks
abstract
One of the major challenges in anatomical landmark detection, based on deep neural networks, is the limited availability of medical imaging data for network learning. To address this problem, we present a two-stage task-oriented deep learning method to detect large-scale anatomical landmarks simultaneously in real time, using limited training data. Specifically, our method consists of two deep convolutional neural networks (CNN), with each focusing on one specific task. Specifically, to alleviate the problem of limited training data, in the first stage, we propose a CNN based regression model using millions of image patches as input, aiming to learn inherent associations between local image patches and target anatomical landmarks. To further model the correlations among image patches, in the second stage, we develop another CNN model, which includes a) a fully convolutional network that shares the same architecture and network weights as the CNN used in the first stage and also b) several extra layers to jointly predict coordinates of multiple anatomical landmarks. Importantly, our method can jointly detect large-scale (e.g., thousands of) landmarks in real time. We have conducted various experiments for detecting 1200 brain landmarks from the 3D T1-weighted magnetic resonance images of 700 subjects, and also 7 prostate landmarks from the 3D computed tomography images of 73 subjects. The experimental results show the effectiveness of our method regarding both accuracy and efficiency in the anatomical landmark detection.
Jun Zhang 0018, Mingxia Liu 0001, Dinggang Shen
IEEE Trans. Image Process.1
2017 Alzheimer's Disease Diagnosis Using Landmark-Based Features From Longitudinal Structural MR Images
abstract
Structural magnetic resonance imaging (MRI) has been proven to be an effective tool for Alzheimer's disease (AD) diagnosis. While conventional MRI-based AD diagnosis typically uses images acquired at a single time point, a longitudinal study is more sensitive in detecting early pathological changes of AD, making it more favorable for accurate diagnosis. In general, there are two challenges faced in MRI-based diagnosis. First, extracting features from structural MR images requires time-consuming nonlinear registration and tissue segmentation, whereas the longitudinal study with involvement of more scans further exacerbates the computational costs. Moreover, the inconsistent longitudinal scans (i.e., different scanning time points and also the total number of scans) hinder extraction of unified feature representations in longitudinal studies. In this paper, we propose a landmark-based feature extraction method for AD diagnosis using longitudinal structural MR images, which does not require nonlinear registration or tissue segmentation in the application stage and is also robust to inconsistencies among longitudinal scans. Specifically, first, the discriminative landmarks are automatically discovered from the whole brain using training images, and then efficiently localized using a fast landmark detection method for testing images, without the involvement of any nonlinear registration and tissue segmentation; and second, high-level statistical spatial features and contextual longitudinal features are further extracted based on those detected landmarks, which can characterize spatial structural abnormalities and longitudinal landmark variations. Using these spatial and longitudinal features, a linear support vector machine is finally adopted to distinguish AD subjects or mild cognitive impairment (MCI) subjects from healthy controls (HCs). Experimental results on the Alzheimer's Disease Neuroimaging Initiative database demonstrate the superior performance and efficiency of the proposed method, with classification accuracies of 88.30% for AD versus HC and 79.02% for MCI versus HC, respectively.
Jun Zhang 0018, Mingxia Liu 0001, Yaozong Gao, Dinggang Shen
IEEE J. Biomed. Health Informatics1
2016 Semi-supervised Hierarchical Multimodal Feature and Sample Selection for Alzheimer's Disease Diagnosis
Ehsan Adeli-Mosabbeb, Mingxia Liu 0001, Jun Zhang 0018, Dinggang Shen
MICCAI (2)4
2016 Diagnosis of Alzheimer's Disease Using View-Aligned Hypergraph Learning with Incomplete Multi-modality Data
Mingxia Liu 0001, Jun Zhang 0018, Pew-Thian Yap, Dinggang Shen
MICCAI (1)2
2016 Detecting Anatomical Landmarks for Fast Alzheimer's Disease Diagnosis
abstract
Structural magnetic resonance imaging (MRI) is a very popular and effective technique used to diagnose Alzheimer's disease (AD). The success of computer-aided diagnosis methods using structural MRI data is largely dependent on the two time-consuming steps: 1) nonlinear registration across subjects, and 2) brain tissue segmentation. To overcome this limitation, we propose a landmark-based feature extraction method that does not require nonlinear registration and tissue segmentation. In the training stage, in order to distinguish AD subjects from healthy controls (HCs), group comparisons, based on local morphological features, are first performed to identify brain regions that have significant group differences. In general, the centers of the identified regions become landmark locations (or AD landmarks for short) capable of differentiating AD subjects from HCs. In the testing stage, using the learned AD landmarks, the corresponding landmarks are detected in a testing image using an efficient technique based on a shape-constrained regression-forest algorithm. To improve detection accuracy, an additional set of salient and consistent landmarks are also identified to guide the AD landmark detection. Based on the identified AD landmarks, morphological features are extracted to train a support vector machine (SVM) classifier that is capable of predicting the AD condition. In the experiments, our method is evaluated on landmark detection and AD classification sequentially. Specifically, the landmark detection error (manually annotated versus automatically detected) of the proposed landmark detector is 2.41 mm , and our landmark-based AD classification accuracy is 83.7%. Lastly, the AD classification performance of our method is comparable to, or even better than, that achieved by existing region-based and voxel-based methods, while the proposed method is approximately 50 times faster.
Jun Zhang 0018, Yue Gao 0002, Yaozong Gao, Brent C. Munsell, Dinggang Shen
IEEE Trans. Medical Imaging1
2015 Automatic Craniomaxillofacial Landmark Digitization via Segmentation-Guided Partially-Joint Regression Forest Model
Jun Zhang 0018, Yaozong Gao, Li Wang 0026, James J. Xia, Dinggang Shen
MICCAI (3)1
2015 Scale invariant texture representation based on frequency decomposition and gradient orientation
Jun Zhang 0018, Jimin Liang, Heng Zhao 0001
Pattern Recognit. Lett.1
2015 A new shape prior model with rotation invariance
Jimin Liang, Jun Zhang 0018, Heng Zhao 0001
Pattern Recognit. Lett.3
2013 Continuous rotation invariant local descriptors for texton dictionary-based texture classification
Jun Zhang 0018, Heng Zhao 0001, Jimin Liang
Comput. Vis. Image Underst.1
2013 Local Energy Pattern for Texture Classification Using Self-Adaptive Quantization Thresholds
abstract
Local energy pattern, a statistical histogram-based representation, is proposed for texture classification. First, we use normalized local-oriented energies to generate local feature vectors, which describe the local structures distinctively and are less sensitive to imaging conditions. Then, each local feature vector is quantized by self-adaptive quantization thresholds determined in the learning stage using histogram specification, and the quantized local feature vector is transformed to a number by N-nary coding, which helps to preserve more structure information during vector quantization. Finally, the frequency histogram is used as the representation feature. The performance is benchmarked by material categorization on KTH-TIPS and KTH-TIPS2-a databases. Our method is compared with typical statistical approaches, such as basic image features, local binary pattern (LBP), local ternary pattern, completed LBP, Weber local descriptor, and VZ algorithms (VZ-MR8 and VZ-Joint). The results show that our method is superior to other methods on the KTH-TIPS2-a database, and achieving competitive performance on the KTH-TIPS database. Furthermore, we extend the representation from static image to dynamic texture, and achieve favorable recognition results on the University of California at Los Angeles (UCLA) dynamic texture database.
Jun Zhang 0018, Jimin Liang, Heng Zhao 0001
IEEE Trans. Image Process.1