Junchao Zhu

dblp:32/10449 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A comprehensive survey of computer vision methods for spatial transcriptomics
abstract
Spatial transcriptomics (ST) enables the simultaneous measurement of gene expression and spatial localization within tissue sections, providing unprecedented opportunities to dissect tissue architecture and functional organization. As a relatively new omics technology, bioinformatics has driven much of the innovation in ST. However, within these frameworks, spatial information is often reduced to locations and relationships between molecular profiles, without fully leveraging the wealth of sub-micron morphological detail and histological knowledge available. Advances in computer vision-based artificial intelligence (AI) are opening exciting new avenues beyond conventional bioinformatics approaches by modeling complex histological patterns and linking morphology to molecular states. More excitingly, they bring fresh perspectives to potentially address key limitations of ST, including its high cost, limited clinical applicability, and reliance on 2D analysis of inherently 3D tissues. For instance, models that predict ST directly from histology images enable virtual sequencing, drastically reducing costs while integrating morphological insights from pathology with molecular biomarkers, thus accelerating clinical translation. Moreover, computer vision techniques can reconstruct pixel-aligned 3D tissue models, overcoming the technical barriers of 2D acquisition and advancing 3D spatial omics analytics. In this paper, we present the first systematic survey of computer vision AI models for ST analytics, categorizing approaches across architectures, learning paradigms, tasks, and datasets, and tracing their technological evolution. We highlight key challenges and future directions, offering a panoramic perspective on vision-driven ST and its potential to transform both basic research and clinical practice. The curated collection of vision-driven ST papers is available at https://github.com/hrlblab/computer_vision_spatial_omics.
Junchao Zhu, Ruining Deng, Junlin Guo, Tianyuan Yao, Siqi Lu, Chongyu Qu, Juming Xiong, Yanfan Zhu, Zhengyi Lu, Yuechen Yang, Marilyn Lionts, Yucheng Tang, Daguang Xu, Shilin Zhao, Haichun Yang, Yuankai Huo
Briefings Bioinform.1
2026 CytoAL: Toward Label-Efficient Cytology Diagnosis via Cellularity-Guided Active Learning
abstract
Examining thyroid fine needle aspiration (FNA) can grade cancer risks, derive prognostic information, and guide follow-up care or surgery decision-making. However, thyroid cytology's diagnostic cues are more dispersed compared with pathology images in other disciplines, making standard annotation strategy for AI diagnosis labor-intensive. Inspired by how cytologists diagnose under the microscope, we propose an innovative cellularity-based active learning framework, namely Cyto-AL, to correlate cellularity with diagnostic categories for the active learning query. We also improve the Whole Slide Image (WSI) category of The Bethesda System for Reporting Thyroid Cytology (TBSRTC) prediction by proposing severe-stage pinpointed Multiple Instance Learning (MIL). Additionally, we introduce a lightweight score model to optimize the query in human in the loop (HITL) annotation strategy. Given scarce public thyroid cytology datasets, we release our collected and labeled images as benchmarks. The benchmark comprises 138 WSIs (27,496 valid image patches) collected from 2021-2023 across six classes, annotated by three pathologists using TBSRTC. At patch-level verification, Cyto-AL achieves a 2.2% average classification accuracy improvement over state-of-the-art methods with an equally labeled dataset, and its lightweight ranking-aware model reduces training time by around 65%. Moreover, the WSI-level MIL approach improves average accuracy by 10.7% and Macro-F1 score by 3.5%, outperforming standard sampling methods such as Monte Carlo sampling. The source code and dataset are available at https://github.com/Junchao-Zhu/Cyto-AL.
Junchao Zhu, Yiqing Shen 0003, Rui Fei Du, Arcot Sowmya, Caifeng Wan, Jing Ke
IEEE Trans. Image Process.1
2025 ASIGN: An Anatomy-aware Spatial Imputation Graphic Network for 3D Spatial Transcriptomics
abstract
Spatial transcriptomics (ST) is an emerging technology that enables medical computer vision scientists to automatically interpret the molecular profiles underlying morphological features. Currently, however, most deep learning-based ST analyses are limited to two-dimensional (2D) sections, which can introduce diagnostic errors due to the heterogeneity of pathological tissues across 3D sections. Expanding ST to three-dimensional (3D) volumes is challenging due to the prohibitive costs; a 2D ST acquisition already costs over 50 times more than whole slide imaging (WSI), and a full 3D volume with 10 sections can be an order of magnitude more expensive. To reduce costs, scientists have attempted to predict ST data directly from WSI without performing actual ST acquisition. However, these methods typically yield unsatisfying results. To address this, we introduce a novel problem setting: 3D ST imputation using 3D WSI histology sections combined with a single 2D ST slide. To do so, we present the Anatomy-aware Spatial Imputation Graph Network (ASIGN) for more precise, yet affordable, 3D ST modeling. The ASIGN architecture extends existing 2D spatial relationships into 3D by leveraging cross-layer overlap and similarity-based expansion. Moreover, a multi-level spatial attention graph network integrates features comprehensively across different data sources. We evaluated ASIGN on three public spatial transcriptomics datasets, with experimental results demonstrating that ASIGN achieves state-of-the-art performance on both 2D and 3D scenarios. The code for this paper is publicly available1.
Junchao Zhu, Ruining Deng, Tianyuan Yao, Juming Xiong, Chongyu Qu, Junlin Guo, Siqi Lu, Mengmeng Yin, Shilin Zhao, Haichun Yang, Yuankai Huo
CVPR1
2025 Dual Teacher with Dempster-Shafer Guidance for Decision Making in Semi-Supervised Small Object Detection
abstract
Small-scale object detection remains a major challenge in semi-supervised object detection (SSOD), particularly in medical image analysis. Conventional teacher models often struggle to accurately capture the features of low-contrast small lesions, leading to noisy pseudo-labels in both localization and classification, which introduces severe uncertainty and degrades detection performance. To address this issue, we propose Dual Teacher, a novel multimodal semi-supervised detection framework designed to enhance pseudo-label reliability and improve small-scale lesion detection. Specifically, we introduce two complementary teacher models: Hybrid-Scale Teacher, which exploits downsampled views to strengthen multi-scale feature learning, and Entropy-Based Multi-Modal Teacher, which leverages entropy maps to refine the quality of small-scale pseudo-labels. To effectively fuse predictions from both teachers and resolve conflicts, we propose a Dempster-Shafer-based Dual-Teacher pseudo-label fusion strategy that explicitly models uncertainty and optimizes classification confidence. Additionally, we introduce a class-adaptive threshold mechanism that dynamically adjusts pseudo-label selection based on dual-teacher predictions, further boosting the recall of small-scale lesions. Extensive experiments on the Dental Disease Dataset, ChestX-Det and M3FD demonstrate that our method consistently surpasses state-of-the-art SSOD approaches. Code is available at: https://github.com/z316910/Dual-Teacher.git.
Nan Gao 0001, Junchao Zhu, Yilong Zhang 0001, Ronghua Liang, Guodao Sun, Peng Chen 0008
ACM Multimedia2
2024 ACPNet: Enhancing Small-Scale Dieases Detection in Panoramic X-rays
abstract
Deep learning-based disease detection can automatically identify dental diseases in panoramic X-rays and improve the accuracy and efficiency of doctors’ diagnoses. However, due to the complex data distribution of panoramic oral X-rays, significant scale differences among lesions, and the presence of many small-scale diseases, automated disease detection in panoramic oral X-rays faces considerable challenges. To alleviate the aforementioned issues, we propose ACPNet, which introduces a novel two-stage approach for detecting small-scale dental diseases in panoramic X-rays using the Contextual Attention Alignment Network (CAAN) and the Point-to-Patch Module (PTPM). To the best of our knowledge, we are the first to explore the detection of small-scale dental diseases in panoramic X-rays under limited sample conditions. Specifically, CAAN integrates deformable convolution with the global attention mechanism of transformer attention, enabling the model to more accurately extract small target foreground features in the complex background of panoramic X-rays. PTPM employs key point detection and cascade dynamic patches to adjust the bounding boxes of lesions, ensuring that small-scale diseases have sufficient high-quality proposals, thereby enhancing detector performance. Additionally, we collected a dataset containing 1157 instances of dental diseases to validate the effectiveness of our algorithm. Extensive experiments demonstrate that ACPNet achieves state-of-the-art performance, highlighting its superiority over baseline and other detection methods.
Nan Gao 0001, Junchao Zhu, Peng Chen 0008, Jijun Tang, Ronghua Liang
BIBM2
2024 Revealing cell-cell communication pathways with their spatially coupled gene programs
abstract
Inference of cell-cell communication (CCC) provides valuable information in understanding the mechanisms of many important life processes. With the rise of spatial transcriptomics in recent years, many methods have emerged to predict CCCs using spatial information of cells. However, most existing methods only describe CCCs based on ligand-receptor interactions, but lack the exploration of their upstream/downstream pathways. In this paper, we proposed a new method to infer CCCs, called Intercellular Gene Association Network (IGAN). Specifically, it is for the first time that we can estimate the gene associations/network between two specific single spatially adjacent cells. By using the IGAN method, we can not only infer CCCs in an accurate manner, but also explore the upstream/downstream pathways of ligands/receptors from the network perspective, which are actually exhibited as a new panoramic cell-interaction-pathway graph, and thus provide extensive information for the regulatory mechanisms behind CCCs. In addition, IGAN can measure the CCC activity at single cell/spot resolution, and help to discover the CCC spatial heterogeneity. Interestingly, we found that CCC patterns from IGAN are highly consistent with the spatial microenvironment patterns for each cell type, which further indicated the accuracy of our method. Analyses on several public datasets validated the advantages of IGAN.
Junchao Zhu, Luonan Chen
Briefings Bioinform.1
2023 An Anti-biased TBSRTC-Category Aware Nuclei Segmentation Framework with a Multi-label Thyroid Cytology Benchmark
Junchao Zhu, Yiqing Shen 0003, Jing Ke
MICCAI (6)1
2023 ClusterSeg: A crowd cluster pinpointed nucleus segmentation framework with cross-modality datasets
Jing Ke, Yizhou Lu, Yiqing Shen 0003, Junchao Zhu, Yijin Zhou, Jinghan Huang 0002, Jieteng Yao, Xiaoyao Liang, Yi Guo 0001, Zhonghua Wei, Fusong Jiang, Dinggang Shen
Medical Image Anal.4
2023 Learning Adaptive Node Embeddings Across Graphs
abstract
Recently, learning embeddings of nodes in graphs has attracted increasing research attention. There are two main kinds of graph embedding methods, i.e., transductive embedding methods and inductive embedding methods. The former focuses on directly optimizing the embedding vectors, and the latter tries to learn a mapping function for the given nodes and features. However, little work has focused on applying the learned model from one graph to another, which is a pervasive idea in Computer Vision or Natural Language Processing. Although some of the graph neural networks (GNNs) present a similar motivation, none of them considers graph biases between graphs. In this paper, we present a novel graph embedding problem called Adaptive Task (AT), and propose a unified framework for the adaptive task, which introduces two types of alignment to learn adaptive node embeddings across graphs. Then, based on the proposed framework, a novel Graph Adaptive Embedding network (GraphAE) is designed to address the adaptive task. Furthermore, we extend GraphAE to a multi-graph version to consider a more complex adaptive situation. The extensive experimental results demonstrate that our model significantly outperforms the state-of-the-art methods, and also show that our framework can make a great improvement over a number of existing GNNs.
Gaoyang Guo, Chaokun Wang, Bencheng Yan, Yunkai Lou, Hao Feng 0007, Junchao Zhu, Jun Chen 0004, Fei He 0001, Philip S. Yu
IEEE Trans. Knowl. Data Eng.6
2019 Forbidden Nodes Aware Community Search
abstract
Community search is an important problem in network analysis, which has attracted much attention in recent years. It starts with some given nodes, pays more attention to local network structures, and gets personalized resultant communities quickly. In this paper, we argue that there are many real scenarios where some nodes are not allowed to appear in the community. Then, we introduce a new concept called forbidden nodes and present a new problem of forbidden nodes aware community search to describe these scenarios.To address the above problem, three methods are proposed, i.e., k-core based FORTE (Forbidden nOdes awaRe communiTy sEarch), k-truss based FORTE and CW based FORTE, where the effects of both forbidden nodes and query nodes are thoroughly considered for each node in the resultant community. The former two methods are able to make use of popular community structures, while the latter is based on a new metric called weighted conductance. The extensive experiments conducted on real data sets demonstrate the effectiveness of the proposed methods.
Chaokun Wang, Junchao Zhu
AAAI2
2018 Super Wide Regression Network for Unsupervised Cross-Database Facial Expression Recognition
abstract
Unsupervised cross-database facial expression recognition (FER) is a challenging problem, in which the training and testing samples belong to different facial expression databases. For this reason, the training (source) and testing (target) facial expression samples would have different feature distributions and hence the performance of lots of existing FER methods may decrease. To solve this problem, in this paper we propose a novel super wide regression network (SWiRN) model, which serves as the regression parameter to bridge the original feature space and the label space and herein in each layer the maximum mean discrepancy (MMD) criterion is used to enforce the source and target facial expression samples to share the same or similar feature distributions. Consequently, the learned SWiRN is able to predict the expression categories of the target samples although we have no access to any label information of target samples. We conduct extensive cross-database FER experiments on CK+, eNTERFACE, and Oulu-CASIA VIS facial expression databases to evaluate the proposed SWiRN. Experimental results show that our SWiRN model achieves more promising performance than recent proposed cross-database emotion recognition methods.
Baofeng Zhang, Yuan Zong, Li Liu 0002, Jie Chen 0001, Guoying Zhao 0001, Junchao Zhu
ICASSP7
2018 Unsupervised Cross-Corpus Speech Emotion Recognition Using Domain-Adaptive Subspace Learning
abstract
In this paper, we investigate an interesting problem, i.e., unsupervised cross-corpus speech emotion recognition (SER), in which the training and testing speech signals come from two different speech emotion corpora. Meanwhile, the training speech signals are labeled, while the label information of the testing speech signals is entirely unknown. Due to this setting, the training (source) and testing (target) speech signals may have different feature distributions and therefore lots of existing SER methods would not work. To deal with this problem, we propose a domain-adaptive subspace learning (DoSL) method for learning a projection matrix with which we can transform the source and target speech signals from the original feature space to the label space. The transformed source and target speech signals in the label space would have similar feature distributions. Consequently, the classifier learned on the labeled source speech signals can effectively predict the emotional states of the unlabeled target speech signals. To evaluate the performance of the proposed DoSL method, we carry out extensive cross-corpus SER experiments on three speech emotion corpora including EmoDB, eNTERFACE, and AFEW 4.0. Compared with recent state-of-the-art cross-corpus SER methods, the proposed DoSL can achieve more satisfactory overall results.
Yuan Zong, Baofeng Zhang, Li Liu 0002, Jie Chen 0001, Guoying Zhao 0001, Junchao Zhu
ICASSP7