EDBT 2026 Demo / reviewers in the wild / expert
Jiaxin Zhuang
dblp:243/9261 · also Jia-Xin Zhuang
· DBLP profile ↗
20ranked-venue papers
8as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DCAC: Dynamic Class-Aware Cache Creates Stronger Out-of-Distribution DetectorsabstractOut-of-distribution (OOD) detection remains a fundamental challenge for deep neural networks, particularly due to overconfident predictions on unseen OOD samples during testing. We reveal a key insight: OOD samples predicted as the same class, or given high probabilities for it, are visually more similar to each other than to the true in-distribution (ID) samples. Motivated by this class-specific observation, we propose DCAC (Dynamic Class-Aware Cache), a training-free, test-time calibration module that maintains separate caches for each ID class to collect high-entropy samples and calibrate the raw predictions of input samples. DCAC leverages cached visual features and predicted probabilities through a lightweight two-layer module to mitigate overconfident predictions on OOD samples. This module can be seamlessly integrated with various existing OOD detection methods across both unimodal and vision-language models while introducing minimal computational overhead. Extensive experiments on multiple OOD benchmarks demonstrate that DCAC significantly enhances existing methods, achieving substantial improvements, i.e., reducing FPR95 by 6.55% when integrated with ASH-S on ImageNet OOD benchmark. Yanqi Wu, Qichao Chen, Runhe Lai, Xinhua Lu, Jiaxin Zhuang, Zhi-Lin Zhao 0001, Wei-Shi Zheng 0001 |
AAAI | 5 |
| 2026 | MG-3D: Multi-grained knowledge-enhanced vision-language pre-training for 3D medical image analysis
Xuefeng Ni, Linshan Wu, Jiaxin Zhuang, Qiong Wang 0001, Mingxiang Wu, Varut Vardhanabhuti, Lihai Zhang, Hanyu Gao, Hao Chen 0011 |
Medical Image Anal. | 3 |
| 2026 | Large-Scale 3D Medical Image Pre-Training With Geometric Context PriorsabstractThe scarcity of annotations poses a significant challenge in medical image analysis, which demands extensive efforts from radiologists, especially for high-dimension 3D medical images. Large-scale pre-training has emerged as a promising label-efficient solution, owing to the utilization of large-scale data, large models, and advanced pre-training techniques. However, its development in medical images remains underexplored. The primary challenge lies in harnessing large-scale unlabeled data and learning high-level semantics without annotations. We observe that 3D medical images exhibit consistent geometric context, i.e., consistent geometric relations between different organs, which leads to a promising way for learning consistent representations. Motivated by this, we introduce a simple-yet-effective Volume Contrast (VoCo) framework to leverage geometric context priors for self-supervision. Given an input volume, we extract base crops from different regions to construct positive and negative pairs for contrastive learning. Then we predict the contextual position of a random crop by contrasting its similarity to the base crops. In this way, VoCo implicitly encodes the inherent geometric context into model representations, facilitating high-level semantic learning without annotations. To assess effectiveness, we (1) introduce PreCT-160 K, the largest medical image pre-training dataset to date, which comprises 160 K Computed Tomography (CT) volumes covering diverse anatomic structures; (2) investigate scaling laws and propose guidelines for tailoring different model sizes to various medical tasks; (3) build a comprehensive benchmark encompassing 51 medical tasks, including segmentation, classification, registration, and vision-language. Extensive experiments highlight the superiority of VoCo, showcasing promising transferability to unseen modalities and datasets. VoCo notably enhances performance on datasets with limited labeled cases and significantly expedites fine-tuning convergence. Linshan Wu, Jiaxin Zhuang, Hao Chen 0011 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Beyond H&E: Unlocking Pathological Insights with Polarization ImagingabstractHistopathology image analysis is fundamental to digital pathology, with hematoxylin and eosin (H&E) staining as the gold standard for diagnostic and prognostic assessments. While H&E imaging effectively highlights cellular and tissue structures, it lacks sensitivity to birefringence and tissue anisotropy, which are crucial for assessing collagen organization, fiber alignment, and microstructural alterations-key indicators of tumor progression, fibrosis, and other pathological conditions. To bridge this gap, we construct a polarization imaging system and curate a new dataset of over 13,000 paired Polar-H&E images. Visualizations of polarization properties reveal distinctive optical signatures in pathological tissues, underscoring its diagnostic value. Building on this dataset, we propose PolarHE, a dual-modality fusion framework that integrates H&E with polarization imaging, leveraging the latter's ability to enhance tissue characterization. Our approach employs a feature decomposition strategy to disentangle common and modality-specific features, ensuring effective multimodal representation learning. Through comprehensive validation, our approach significantly outperforms previous methods, achieving an accuracy of 86.70 % on the Chaoyang dataset and 89.06 % on the MHIST dataset. These results demonstrate that polarization imaging is a powerful and underutilized modality in computational pathology, enriching feature representation and improving diagnostic accuracy. PolarHE establishes a promising direction for multimodal learning, paving the way for more interpretable and generalizable pathology models. Jiaxin Zhuang, Jing Cong, Limei Guo, Xiaomeng Li 0001 |
BIBM | 2 |
| 2025 | Normal Distribution Priority Path Planning Method for Unmanned SystemabstractNormal Distribution Priority (NDP) algorithm is a path-planning strategy for unmanned systems navigating complex environments. NDP efficiently circumvents obstacles and navigates through constricted spaces while optimizing energy use. The algorithm consists of two stages: waypoint expansion and energy optimization. First, it selects optimal waypoints using a normal distribution-based weighting mechanism to create a secure route. Then, it fine-tunes the path to minimize energy consumption. Dynamic programming optimizes the process to ensure efficiency. Comparative analyses with RRT* and PSO algorithms demonstrate NDP’s advantages in obstacle negotiation, time efficiency, and path refinement, particularly in static environments. Experimental validations confirm the algorithm’s adaptability and effectiveness for complex path planning. Jiaxin Zhuang, Xiaolin Mou |
IECON | 2 |
| 2025 | Diffusion-Based Virtual Staining from Polarimetric Mueller Matrix Imaging
Jiaxin Zhuang, Jing Cong, Limei Guo, Hao Chen 0011 |
MICCAI (1) | 3 |
| 2025 | Bio2Vol: Adapting 2D Biomedical Foundation Models for Volumetric Medical Image Segmentation
Jiaxin Zhuang, Linshan Wu, Xuefeng Ni, Xi Wang 0013, Liansheng Wang 0002, Hao Chen 0011 |
MICCAI (6) | 1 |
| 2025 | Boundary-Guided Contrastive Learning for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning methods, compared to fully supervised learning, offer significant potential to alleviate the burden of manual annotations on clinicians. By leveraging unlabeled data, these methods can aid in the development of medical image segmentation systems for improving efficiency. Boundary segmentation is crucial in medical image analysis. However, accurate segmentation of boundary regions is under-explored in existing methods since boundary pixels constitute only a small fraction of the overall image, resulting in suboptimal segmentation performance for boundary regions. In this paper, we introduce boundary-guided contrastive learning for semi-supervised medical image segmentation (BoCLIS). Specifically, we first propose conservative-to-radical teacher networks with an uncertainty-weighted aggregation strategy to generate higher quality pseudo-labels, enabling more efficient utilization of unlabeled data. To further improve the performance of segmentation in boundary regions, we propose a boundary-guided patch sampling strategy to guide the framework in learning discriminative representations for these regions. Lastly, the patch-based contrastive learning is proposed to simultaneously compute the (dis)similarities of the discriminative representations across intra- and inter-images. Extensive experiments on three public datasets show that our method consistently outperforms existing methods, especially in the boundary region, with DSC improvements of 20.47%, 16.75%, and 17.18%, respectively. A comprehensive analysis is further performed to demonstrate the effectiveness of our approach. Our code is released publicly at https://github.com/youngyzzZ/BoCLIS. Yang Yang 0002, Jiaxin Zhuang, Guoying Sun, Jingyong Su |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Advancing Volumetric Medical Image Segmentation via Global-Local Masked AutoencodersabstractMasked Autoencoder (MAE) is a self-supervised pre-training technique that holds promise in improving the representation learning of neural networks. However, the current application of MAE directly to volumetric medical images poses two challenges: (i) insufficient global information for clinical context understanding of the holistic data, and (ii) the absence of any assurance of stabilizing the representations learned from randomly masked inputs. To conquer these limitations, we propose the Global-Local Masked AutoEncoders (GL-MAE), a simple yet effective self-supervised pre-training strategy. GL-MAE acquires robust anatomical structure features by incorporating multi-level reconstruction from fine-grained local details to high-level global semantics. Furthermore, a complete global view serves as an anchor to direct anatomical semantic alignment and stabilize the learning process through global-to-global consistency learning and global-to-local consistency learning. Our fine-tuning results on eight mainstream public datasets demonstrate the superiority of our method over other state-of-the-art self-supervised algorithms, highlighting its effectiveness on versatile volumetric medical image segmentation and classification tasks. We will release codes upon acceptance at https://github.com/JiaxinZhuang/GL-MAE. Jiaxin Zhuang, Luyang Luo, Qiong Wang 0001, Mingxiang Wu, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | MiM: Mask in Mask Self-Supervised Pre-Training for 3D Medical Image AnalysisabstractThe Vision Transformer (ViT) has demonstrated remarkable performance in Self-Supervised Learning (SSL) for 3D medical image analysis. Masked AutoEncoder (MAE) for feature pre-training can further unleash the potential of ViT on various medical vision tasks. However, due to large spatial sizes with much higher dimensions of 3D medical images, the lack of hierarchical design for MAE may hinder the performance of downstream tasks. In this paper, we propose a novel Mask in Mask (MiM) pre-training framework for 3D medical images, which aims to advance MAE by learning discriminative representation from hierarchical visual tokens across varying scales. We introduce multiple levels of granularity for masked inputs from the volume, which are then reconstructed simultaneously ranging at both fine and coarse levels. Additionally, a cross-level alignment mechanism is applied to adjacent level volumes to enforce anatomical similarity hierarchically. Furthermore, we adopt a hybrid backbone to enhance the hierarchical representation learning efficiently during the pre-training. MiM was pre-trained on a large scale of available 3D volumetric images, i.e., Computed Tomography (CT) images containing various body parts. Extensive experiments on twelve public datasets demonstrate the superiority of MiM over other SSL methods in organ/tumor segmentation and disease classification. We further scale up the MiM to large pre-training datasets with more than 10k volumes, showing that large-scale pre-training can further enhance the performance of downstream tasks. Code is available at https://github.com/JiaxinZhuang/MiM. Jiaxin Zhuang, Linshan Wu, Qiong Wang 0001, Peng Fei, Varut Vardhanabhuti, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 1 |
| 2024 | VoCo: A Simple-Yet-Effective Volume Contrastive Learning Framework for 3D Medical Image AnalysisabstractSelf-Supervised Learning (SSL) has demonstrated promising results in 3D medical image analysis. However, the lack of high-level semantics in pre-training still heavily hinders the performance of downstream tasks. We ob-serve that 3D medical images contain relatively consistent contextual position information, i.e., consistent geometric relations between different organs, which leads to a potential way for us to learn consistent semantic representations in pre-training. In this paper, we propose a simple-yet-effective Volume Contrast (VoCo) framework to leverage the contextual position priors for pre-training. Specif-ically, we first generate a group of base crops from different regions while enforcing feature discrepancy among them, where we employ them as class assignments of dif-ferent regions. Then, we randomly crop sub-volumes and predict them belonging to which class (located at which re-gion) by contrasting their similarity to different base crops, which can be seen as predicting contextual positions of different sub-volumes. Through this pretext task, VoCo implic-itly encodes the contextual position priors into model rep-resentations without the guidance of annotations, enabling us to effectively improve the performance of downstream tasks that require high-level semantics. Extensive exper-imental results on six downstream tasks demonstrate the superior effectiveness of VoCo. Code will be available at httpsu/github.com/luffytls/vo'Co. Linshan Wu, Jiaxin Zhuang, Hao Chen 0011 |
CVPR | 2 |
| 2024 | Graph Link Prediction via Decay Coefficient based Proportional Aggregation and Hybrid ConcatenationabstractAs one of the popular topics in social network analysis, link prediction aims to predict the likely but unobserved links between two nodes. It has a wide field of applications such as knowledge graph completion and recommender systems. Currently, graph neural network (GNN) is the most common method with excellent performance, which contains two main steps, i.e., node aggregation and edge concatenation. However, there are some drawbacks with the existing methods. Firstly, traditional node aggregation methods usually iteratively aggregate all or a fixed number of neighbors, which is inflexible and inefficient. Secondly, most of the existing edge concatenation methods only apply a single concatenation approach to obtain edge embeddings, which cannot fully guarantee the embedding quality. To tackle the two problems, this paper proposes a graph embedding approach for link prediction via decay coefficient based Proportional Aggregation and Hybrid Concatenation (PAHC). On one hand, PAHC directly aggregates only proportional neighbors from different orders by correlating the decaying phenomenon of information propagation in social networks with the decaying phenomenon of temperature in Newton’s cooling theorem. On the other hand, PAHC proposes a hybrid concatenation approach to obtain the final edge embedding by mixing weighted summation and weighted direct concatenation. Experiments on five datasets show that through decay proportional aggregation and hybrid concatenation, the proposed PAHC can better predict links in social networks, outperforming state-of-the-art methods. Jiaxin Zhuang, Yahui Chai, Xiaobin Rui |
IJCNN | 1 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 45 |
| 2024 | Unleashing the Denoising Capability of Diffusion Prior for Solving Inverse ProblemsabstractThe recent emergence of diffusion models has significantly advanced the precision of learnable priors, presenting innovative avenues for addressing inverse problems. Previous works have endeavored to integrate diffusion priors into the maximum a posteriori estimation (MAP) framework and design optimization methods to solve the inverse problem. However, prevailing optimization-based rithms primarily exploit the prior information within the diffusion models while neglecting their denoising capability. To bridge this gap, this work leverages the diffusion process to reframe noisy inverse problems as a two-variable constrained optimization task by introducing an auxiliary optimization variable that represents a 'noisy' sample at an equivalent denoising step. The projection gradient descent method is efficiently utilized to solve the corresponding optimization problem by truncating the gradient through the $\mu$-predictor. The proposed algorithm, termed ProjDiff, effectively harnesses the prior information and the denoising capability of a pre-trained diffusion model within the optimization framework. Extensive experiments on the image restoration tasks and source separation and partial generation tasks demonstrate that ProjDiff exhibits superior performance across various linear and nonlinear inverse problems, highlighting its potential for practical applications. Code is available at https://github.com/weigerzan/ProjDiff/. Jiaxin Zhuang, Yuantao Gu |
NeurIPS | 2 |
| 2023 | Adaptive Stanley Control Method Based on Dynamic Window ApproachabstractStanley algorithm is a classical algorithm in automatic driving path tracking algorithm, which has the advantages of low computational complexity and good tracking effect. However, the tracking accuracy of the traditional Stanley algorithm is limited by the gain parameters and the gain coefficients cannot be dynamically adjusted according to the road conditions. To improve the tracking accuracy and driving stability, an adaptive Stanley control algorithm based on DWA (DWA-Stanley) is proposed, which is used to sampling the velocity vector space and selecting the evaluation function, the vehicle motion planning is based on the best set of velocity values. In this paper, the DWA-Stanley algorithm is validated based on three different driving speeds. Simulation results show that an racecar using the DWA-Stanley algorithm has higher tracking accuracy and smoothness than a conventional Stanley. Jiaxin Zhuang, Guanrong Huang, Yizhen Wu, Qiang Hua, Bian Gong, Xiaolin Mou |
IECON | 1 |
| 2023 | Class attention to regions of lesion for imbalanced medical image recognition
Jiaxin Zhuang, Jiabin Cai, Jianguo Zhang 0001, Wei-Shi Zheng 0001 |
Neurocomputing | 1 |
| 2022 | View Dialogue in 2D: A Two-stream Model in Time-speaker Perspective for Dialogue Summarization and beyondabstractExisting works on dialogue summarization often follow the common practice in document summarization and view the dialogue, which comprises utterances of different speakers, as a single utterance stream ordered by time. However, this single-stream approach without specific attention to the speaker-centered points has limitations in fully understanding the dialogue. To better capture the dialogue information, we propose a 2D view of dialogue based on a time-speaker perspective, where the time and speaker streams of dialogue can be obtained as strengthened input. Based on this 2D view, we present an effective two-stream model called ATM to combine the two streams. Extensive experiments on various summarization datasets demonstrate that ATM significantly surpasses other models regarding diverse metrics and beats the state-of-the-art models on the QMSum dataset in ROUGE scores. Besides, ATM achieves great improvements in summary faithfulness and human evaluation. Moreover, results on machine reading comprehension datasets show the generalization ability of the proposed methods and shed light on other dialogue-based tasks. Our code will be publicly available online. Keli Xie, Dongchen He, Jiaxin Zhuang, Siyuan Lu 0002, Zhongfeng Wang 0001 |
COLING | 3 |
| 2022 | DisCo: Remedying Self-supervised Learning on Lightweight Models with Distilled Contrastive Learning
Jiaxin Zhuang, Shaohui Lin, Hao Cheng 0012, Xing Sun 0001, Ke Li 0015, Chunhua Shen |
ECCV (26) | 2 |
| 2022 | OpenMedIA: Open-Source Medical Image Analysis Toolbox and Benchmark Under Heterogeneous AI Computing Platforms
Jiaxin Zhuang, Xiansong Huang, Yang Yang 0002, Jiancong Chen, Yue Yu 0001, Wei Gao 0003, Ge Li 0002, Jie Chen 0001, Tong Zhang 0017 |
PRCV (1) | 1 |
| 2020 | Deep kNN for Medical Image Classification
Jiaxin Zhuang, Jiabin Cai, Jianguo Zhang 0001, Wei-Shi Zheng 0001 |
MICCAI (1) | 1 |