Junhan Zhao

dblp:225/1522 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-0316-8365ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2026 NeuroSonic: Instant Neural Impact Sound Synthesis With Learned Acoustic Transfer
abstract
ABSTRACT Physically based modal sound synthesis generates realistic audio by modeling structural vibration and acoustic radiation from object geometry and material properties. However, traditional approaches rely on computationally expensive eigenvalue analysis and acoustic transfer simulations, making them impractical for real‐time virtual reality (VR), especially for dynamically generated or fractured objects. We present NeuroSonic, a unified neural framework for real‐time physically based sound synthesis directly from point cloud geometry. Our key idea is to amortize modal analysis and acoustic transfer using neural networks, replacing costly numerical solvers with efficient feed‐forward inference. The framework predicts modal eigenvalues, eigenvectors, and acoustic transfer functions, forming a physically grounded representation for sound synthesis. By operating on point clouds, our method avoids voxelization artifacts and handles geometrically complex and thin structures. Once trained, NeuroSonic computes complete modal and acoustic information in under 0.02 s, enabling real‐time performance. Experiments show that our approach achieves accuracy comparable to physically based solvers while providing orders‐of‐magnitude speedup, making it a practical solution for interactive VR environments.
Junhan Zhao, Xutong Jin, Meng Gai, Sheng Li 0008
Comput. Animat. Virtual Worlds1
2026 Toward General-Purpose Video Reconstruction Through Synergy of Grid-Splicing Diffusion and Large Language Models
abstract
Various forms of degradation, including noise, blur, and adverse weather conditions (e.g., rain, snow, and fog), significantly compromise video quality and system reliability across critical domains ranging from surveillance and medical imaging to entertainment. Previous research mainly focuses on network models tailored to specific degradation types, while recent unified frameworks and foundation models still face critical challenges in temporal consistency, automated degradation recognition, and detail preservation. Despite recent advances in foundation models, current approaches rely heavily on predefined degradation labels and remain focused on image-level operations, limiting their generalization to real-world scenarios and struggling with preserving fine-grained details. To address these challenges, we propose Grid Splicing Diffusion Model (GSDiff), a general framework for video reconstruction that leverages a novel grid splicing execution alongside instruction-tuned Large Language Model (LLM). GSDiff introduces three key innovative modules: (1) a LLM-driven degradation recognition module that enables automatic and fine-grained restoration guidance through zero-shot degradation analysis, (2) a Grid Splicing Module that organizes multiple frames into a unified grid structure to facilitate spatiotemporal feature processing, and (3) a Detail Preservation Module integrated with a Tail Refine Network to enhance fine-grained details during diffusion and post-processing. Extensive experiments demonstrate that GSDiff delivers state-of-the-art performance across a wide range of reconstruction tasks, including deraining, desnowing, denoising, and deblurring, propelling advancements in medical diagnostics and smart city applications.
Sen Yang 0006, Jinxi Xiang, Jieqiong Zhao, Zongxin Yang, Junhan Zhao
IEEE Trans. Circuits Syst. Video Technol.8
2025 Prompt-based multimodal representation learning for drug repurposing
abstract
Drug repurposing significantly reduces development costs and shortens research cycles, making it a critical strategy in drug discovery. An emerging class of drug repurposing approaches applies deep learning to structural data. However, these methods often depend on static representations of molecular and protein structures, which may not fully capture the dynamic character of compound-protein interactions. To address these challenges and enhance the accuracy of compound-protein interaction predictions, we introduce an innovative prompt-based multimodal representation learning framework that dynamically encodes task-specific contextual information for drug repurposing. Specifically, the framework includes a dynamic prompt generation module that adaptively creates receptor-specific prompts and a prompt calibration module for effective multimodal feature integration and optimization. When applied to identifying FDA-approved drug candidates targeting G-protein-coupled receptors, our method achieved a 7.4% improvement in mean absolute error compared with state-of-the-art methods, with up to a 25.1% improvement for specific target-of-interest. By demonstrating potential in repurposing non-opioid treatments without the risk of addiction for safe pain management, our method has the capacity to advance drug discovery and meet a wide range of therapeutic needs.
Kaicheng U, Dhruv Rana, Sophia Meixuan Zhang, Sen Yang 0006, Zongxin Yang, Hongping Tang, Junhan Zhao
Briefings Bioinform.11
2025 Rethinking Copy-Paste for Consistency Learning in Medical Image Segmentation
abstract
Semi-supervised learning based on consistency learning offers significant promise for enhancing medical image segmentation. Current approaches use copy-paste as an effective data perturbation technique to facilitate weak-to-strong consistency learning. However, these techniques often lead to a decrease in the accuracy of synthetic labels corresponding to the synthetic data and introduce excessive perturbations to the distribution of the training data. Such over-perturbation causes the data distribution to stray from its true distribution, thereby impairing the model's generalization capabilities as it learns the decision boundaries. We propose a weak-to-strong consistency learning framework that integrally addresses these issues with two primary designs: 1) it emphasizes the use of highly reliable data to enhance the quality of labels in synthetic datasets through cross-copy-pasting between labeled and unlabeled datasets; 2) it employs uncertainty estimation and foreground region constraints to meticulously filter the regions for copy-pasting, thus the copy-paste technique implemented introduces a beneficial perturbation to the training data distribution. Our framework expands the copy-paste method by addressing its inherent limitations, and amplifying the potential of data perturbations for consistency learning. We extensively validated our model using six publicly available medical image segmentation datasets across different diagnostic tasks, including the segmentation of cardiac structures, prostate structures, brain structures, skin lesions, and gastrointestinal polyps. The results demonstrate that our method significantly outperforms state-of-the-art models. For instance, on the PROMISE12 dataset for the prostate structure segmentation task, using only 10% labeled data, our method achieves a 15.31% higher Dice score compared to the baseline models. Our experimental code will be made publicly available at https://github.com/slhuang24/RCP4CL.
Senlong Huang, Yongxin Ge, Dongfang Liu, Mingjian Hong, Junhan Zhao, Alexander C. Loui
IEEE Trans. Image Process.5
2025 Counterfactual Bidirectional Co-Attention Transformer for Integrative Histology-Genomic Cancer Risk Stratification
abstract
Applying deep learning to predict patient prognostic survival outcomes using histological whole-slide images (WSIs) and genomic data is challenging due to the morphological and transcriptomic heterogeneity present in the tumor microenvironment. Existing deep learning-enabled methods often exhibit learning biases, primarily because the genomic knowledge used to guide directional feature extraction from WSIs may be irrelevant or incomplete. This results in a suboptimal and sometimes myopic understanding of the overall pathological landscape, potentially overlooking crucial histological insights. To tackle these challenges, we propose the CounterFactual Bidirectional Co-Attention Transformer framework. By integrating a bidirectional co-attention layer, our framework fosters effective feature interactions between the genomic and histology modalities and ensures consistent identification of prognostic features from WSIs. Using counterfactual reasoning, our model utilizes causality to model unimodal and multimodal knowledge for cancer risk stratification. This approach directly addresses and reduces bias, enables the exploration of 'what-if' scenarios, and offers a deeper understanding of how different features influence survival outcomes. Our framework, validated across eight diverse cancer benchmark datasets from The Cancer Genome Atlas (TCGA), represents a major improvement over current histology-genomic model learning methods. It shows an average 2.5% improvement in c-index performance over 18 state-of-the-art models in predicting patient prognoses across eight cancer types.
Zheyi Ji, Yongxin Ge, Chijioke Chukwudi, Kaicheng U, Sophia Meixuan Zhang, Yulong Peng, Junyou Zhu, Hossam Zaki, Xueling Zhang, Sen Yang 0006, Junhan Zhao
IEEE J. Biomed. Health Informatics13
2024 Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation
abstract
Recently, text-to-image diffusion models have emerged as a powerful tool for image-to-image translation (I2I), allowing flexible image translation via user-provided text prompts. This paper proposes frequency-controlled diffusion model (FCDiffusion), an end-to-end diffusion-based framework contributing a novel solution to text-guided I2I from a frequency-domain perspective. At the heart of our framework is a feature-space frequency-domain filtering module based on Discrete Cosine Transform, which extracts image features carrying different DCT spectral bands to control the text-to-image generation process of the Latent Diffusion Model, realizing versatile I2I applications including style-guided content creation, image semantic manipulation, image scene translation, and image style translation. Different from related methods, FCDiffusion establishes a unified text-driven I2I framework suiting diverse I2I application scenarios simply by switching among different frequency control branches. The effectiveness and superiority of our method for text-guided I2I are demonstrated with extensive experiments both qualitatively and quantitatively. Our project is publicly available at: https://xianggao1102.github.io/FCDiffusion/.
Xiang Gao 0014, Zhengbo Xu, Junhan Zhao, Jiaying Liu 0001
AAAI3
2024 Radiance Field Learners As UAV First-Person Viewers
Liqi Yan, Qifan Wang 0001, Junhan Zhao, Qiang Guan, Dongfang Liu
ECCV (61)3
2024 M²PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning
abstract
Taowen Wang, Yiyang Liu, James Chenhao Liang, Junhan Zhao, Yiming Cui, Yuning Mao, Shaoliang Nie, Jiahao Liu, Fuli Feng, Zenglin Xu, Cheng Han, Lifu Huang, Qifan Wang, Dongfang Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Taowen Wang, Yiyang Liu 0003, James Liang, Junhan Zhao, Yiming Cui 0002, Yuning Mao, Shaoliang Nie, Fuli Feng, Zenglin Xu, Cheng Han 0001, Lifu Huang, Qifan Wang 0001, Dongfang Liu
EMNLP4
2024 A systematic evaluation of computational methods for cell segmentation
abstract
Cell segmentation is a fundamental task in analyzing biomedical images. Many computational methods have been developed for cell segmentation and instance segmentation, but their performances are not well understood in various scenarios. We systematically evaluated the performance of 18 segmentation methods to perform cell nuclei and whole cell segmentation using light microscopy and fluorescence staining images. We found that general-purpose methods incorporating the attention mechanism exhibit the best overall performance. We identified various factors influencing segmentation performances, including image channels, choice of training data, and cell morphology, and evaluated the generalizability of methods across image modalities. We also provide guidelines for choosing the optimal segmentation methods in various real application scenarios. We developed Seggal, an online resource for downloading segmentation models already pre-trained with various tissue and cell types, substantially reducing the time and effort for training cell segmentation models.
Junhan Zhao, Hongye Xu, Cheng Han 0001, Zhiqiang Tao, Tong Geng, Dongfang Liu
Briefings Bioinform.2
2024 Predicting physical functioning status in older adults: insights from wrist accelerometer sensors and derived digital biomarkers of physical activity
abstract
OBJECTIVE: Conventional physical activity (PA) metrics derived from wearable sensors may not capture the cumulative, transitions from sedentary to active, and multidimensional patterns of PA, limiting the ability to predict physical function impairment (PFI) in older adults. This study aims to identify unique temporal patterns and develop novel digital biomarkers from wrist accelerometer data for predicting PFI and its subtypes using explainable artificial intelligence techniques. MATERIALS AND METHODS: Wrist accelerometer streaming data from 747 participants in the National Health and Aging Trends Study (NHATS) were used to calculate 231 PA features through time-series analysis techniques-Tsfresh. Predictive models for PFI and its subtypes (walking, balance, and extremity strength) were developed using 6 machine learning (ML) algorithms with hyperparameter optimization. The SHapley Additive exPlanations method was employed to interpret the ML models and rank the importance of input features. RESULTS: Temporal analysis revealed peak PA differences between PFI and healthy controls from 9:00 to 11:00 am. The best-performing model (Gradient boosting Tree) achieved an area under the curve score of 85.93%, accuracy of 81.52%, sensitivity of 77.03%, and specificity of 87.50% when combining wrist accelerometer streaming data (WAPAS) features with demographic data. DISCUSSION: The novel digital biomarkers, including change quantiles, Fourier transform (FFT) coefficients, and Aggregated (AGG) Linear Trend, outperformed traditional PA metrics in predicting PFI. These findings highlight the importance of capturing the multidimensional nature of PA patterns for PFI. CONCLUSION: This study investigates the potential of wrist accelerometer digital biomarkers in predicting PFI and its subtypes in older adults. Integrated PFI monitoring systems with digital biomarkers would improve the current state of remote PFI surveillance.
Lingjie Fan, Junhan Zhao, Mengyi Wu, Tao Lin 0022
J. Am. Medical Informatics Assoc.2
2024 Error-Robust and Label-Efficient Deep Learning for Understanding Tumor Microenvironment From Spatial Transcriptomics
abstract
Spatial transcriptomics (ST) has become an important methodology in the analysis of the tumor microenvironment (TME) due to its ability to provide gene expression information with spatial resolution, enabling the identification and characterization of TME gene markers. Deep learning methods are proposed for analyzing spatial transcriptomic data for clustering the spatial regions of the TME based on gene expression. However, deep learning methods are often imposed by errors, which can impact the accuracy of gene expression quantification and TME gene identification. To address this issue, we propose a label-efficient method that utilizes curriculum learning and confidence learning to identify errors in graph deep learning when analyzing ST data. Our method explicitly incorporates the effect of noise in the learning process and employs probabilistic models or uncertainty estimates to represent the uncertainty in the data. Validated on human breast cancer ST data, we studied spatial gene expression in HER2-positive breast tumors using our method. The evaluation results suggest that the error quantification helps identify the noisy samples and subset the samples that results in more accurate gene expression quantification and TME gene identification. Additionally, there are biological insights obtained from the new subset formed by error samples. This error-robust deep learning method offers promising avenues for the analysis of spatial transcriptomic data, enabling accurate and label-efficient quantification of gene expression and identification of TME gene markers.
Jiake Leng, Yiming Cui 0002, Junhan Zhao, Yongxin Ge
IEEE Trans. Circuits Syst. Video Technol.5
2024 SAC-Net: Enhancing Spatiotemporal Aggregation in Cervical Histological Image Classification via Label-Efficient Weakly Supervised Learning
abstract
Cervical cancer is the fourth most common cancer in women and its subtyping requires examining histopathological slides or digital images, such as whole slide images (WSIs). However, manually inspecting WSIs with gigapixel sizes can be laborious and prone to errors for pathologists. To address this issue, computer-aided approaches based on weakly-supervised learning techniques have been proposed. These methods can predict disease types directly from WSIs and highlight diagnosis-relevant regions, which can help pathologists achieve faster and more accurate diagnoses. WSIs are divided into overlapping patches using a sliding window approach, and these patches are subsequently screened in a sequential zig-zag pattern to identify spatiotemporal dependencies. These dependencies are further analyzed to generate predictions at the WSI level. Therefore, effective patch feature learning and spatiotemporal aggregation are two key issues in the weakly-supervised WSI classification (WSWC) task. In this paper, we present a label-efficient WSWC method called spatiotemporal aggregation for cervical WSIs (SAC-Net), which jointly performs online feature extraction and feature aggregation to infer the WSI-level prediction in an end-to-end manner. The online feature extractor helps to learn cervical-cancer-specific features and obtain more accurate patch representations. The feature aggregator uses an online instance clustering method to learn proper weight parameters for each cluster, which generates the WSI embedding with enhanced spatiotemporal aggregation. SAC-Net is developed and evaluated on a public cervical WSI dataset (TissueNet) containing 1015 WSIs, which are also externally tested on three independent cervical WSI datasets. Our results demonstrate that SAC-Net achieves state-of-the-art classification performance and is robust. SAC-Net has the potential to be a useful tool for clinical cervical cancer detection.
De Cai, Sen Yang 0006, Yiming Cui 0002, Junyou Zhu, Kanran Wang, Junhan Zhao
IEEE Trans. Circuits Syst. Video Technol.7
2024 HiCervix: An Extensive Hierarchical Dataset and Benchmark for Cervical Cytology Classification
abstract
Cervical cytology is a critical screening strategy for early detection of pre-cancerous and cancerous cervical lesions. The challenge lies in accurately classifying various cervical cytology cell types. Existing automated cervical cytology methods are primarily trained on databases covering a narrow range of coarse-grained cell types, which fail to provide a comprehensive and detailed performance analysis that accurately represents real-world cytopathology conditions. To overcome these limitations, we introduce HiCervix, the most extensive, multi-center cervical cytology dataset currently available to the public. HiCervix includes 40,229 cervical cells from 4,496 whole slide images, categorized into 29 annotated classes. These classes are organized within a three-level hierarchical tree to capture fine-grained subtype information. To exploit the semantic correlation inherent in this hierarchical tree, we propose HierSwin, a hierarchical vision transformer-based classification network. HierSwin serves as a benchmark for detailed feature learning in both coarse-level and fine-level cervical cancer classification tasks. In our comprehensive experiments, HierSwin demonstrated remarkable performance, achieving 92.08% accuracy for coarse-level classification and 82.93% accuracy averaged across all three levels. When compared to board-certified cytopathologists, HierSwin achieved high classification performance (0.8293 versus 0.7359 averaged accuracy), highlighting its potential for clinical applications. This newly released HiCervix dataset, along with our benchmark HierSwin method, is poised to make a substantial impact on the advancement of deep learning algorithms for rapid cervical cancer screening and greatly improve cancer prevention and patient outcomes in real-world clinical settings.
De Cai, Jie Chen 0081, Junhan Zhao, Yuan Xue 0002, Sen Yang 0006, Wei Yuan 0015, Min Feng 0012, Haiyan Weng, Yulong Peng, Junyou Zhu, Kanran Wang, Christopher Jackson, Hongping Tang, Junzhou Huang
IEEE Trans. Medical Imaging3
2022 Memorizing All for Implicit Discourse Relation Recognition
abstract
Implicit discourse relation recognition is a challenging task due to the absence of the necessary informative clues from explicit connectives. An implicit discourse relation recognizer has to carefully tackle the semantic similarity of sentence pairs and the severe data sparsity issue. In this article, we learn token embeddings to encode the structure of a sentence from a dependency point of view in their representations and use them to initialize a baseline model to make it really strong. Then, we propose a novel memory component to tackle the data sparsity issue by allowing the model to master the entire training set, which helps in achieving further performance improvement. The memory mechanism adequately memorizes information by pairing representations and discourse relations of all training instances, thus filling the slot of the data-hungry issue in the current implicit discourse relation recognizer. The proposed memory component, if attached with any suitable baseline, can help in performance enhancement. The experiments show that our full model with memorizing the entire training data provides excellent results on PDTB and CDTB datasets, outperforming the baselines by a fair margin.
Kashif Munir, Hongxiao Bai, Hai Zhao 0001, Junhan Zhao
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2021 Phoenixmap: An Abstract Approach to Visualize 2D Spatial Distributions
abstract
The multidimensional nature of spatial data poses a challenge for visualization. In this paper, we introduce Phoenixmap, a simple abstract visualization method to address the issue of visualizing multiple spatial distributions at once. The Phoenixmap approach starts by identifying the enclosed outline of the point collection, then assigns different widths to outline segments according to the segments' corresponding inside regions. Thus, one 2D distribution is represented as an outline with varied thicknesses. Phoenixmap is capable of overlaying multiple outlines and comparing them across categories of objects in a 2D space. We chose heatmap as a benchmark spatial visualization method and conducted user studies to compare performances among Phoenixmap, heatmap, and dot distribution map. Based on the analysis and participant feedback, we demonstrate that Phoenixmap 1) allows users to perceive and compare spatial distribution data efficiently; 2) frees up graphics space with a concise form that can provide visualization design possibilities like overlapping; and 3) provides a good quantitative perceptual estimating capability given the proper legends. Finally, we discuss several possible applications of Phoenixmap and present one visualization of multiple species of birds' active regions in a nature preserve.
Junhan Zhao, Xiang Liu 0016, Cheryl Z. Qian, Victor Y. Chen
IEEE Trans. Vis. Comput. Graph.1