Zhongzhen Huang

dblp:320/0462 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Segmentation and scene understanding · 54% Vision and language · 24% Question answering and dialogue systems · 11%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
medical image segmentation
1.522024
CAT: Coordinating Anatomical-Textual Prompts for Multi-Organ and Tumor Segmentation · NeurIPS 2024
ZePT: Zero-Shot Pan-Tumor Segmentation via Query-Disentangling and Self-Prompting · CVPR 2024
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
medical question answering
1.012026
MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration · ACL (1) 2026
Natural language and speech › Language models and text generation › agentic language model
tool-augmented language models
1.012026
MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration · ACL (1) 2026
Computer vision › Segmentation and scene understanding › medical image segmentation
multi-organ and tumor segmentation
0.812024
CAT: Coordinating Anatomical-Textual Prompts for Multi-Organ and Tumor Segmentation · NeurIPS 2024
Computer vision › Segmentation and scene understanding
prompt-based segmentation
0.812024
CAT: Coordinating Anatomical-Textual Prompts for Multi-Organ and Tumor Segmentation · NeurIPS 2024
Computer vision › Segmentation and scene understanding › medical image segmentation
tumor segmentation
0.812024
ZePT: Zero-Shot Pan-Tumor Segmentation via Query-Disentangling and Self-Prompting · CVPR 2024
Computer vision › Vision and language
vision-language model
0.812024
CAT: Coordinating Anatomical-Textual Prompts for Multi-Organ and Tumor Segmentation · NeurIPS 2024
Computer vision › Segmentation and scene understanding › open-world segmentation
zero-shot segmentation
0.812024
ZePT: Zero-Shot Pan-Tumor Segmentation via Query-Disentangling and Self-Prompting · CVPR 2024
Computer vision › Vision and language
image captioning
0.712023
KiUT: Knowledge-injected U-Transformer for Radiology Report Generation · CVPR 2023
Computer vision › Vision and language › image captioning
medical image captioning
0.712023
KiUT: Knowledge-injected U-Transformer for Radiology Report Generation · CVPR 2023
Medical and health informatics › medical report generation
radiology report generation
0.712023
KiUT: Knowledge-injected U-Transformer for Radiology Report Generation · CVPR 2023

Methods — techniques the papers use, named apart from their topics

u-transformer · 1.3symptom graph · 1.3knowledge distillation · 1.3model context protocol · 1.0sharerefiner · 0.8self-prompting · 0.8query-knowledge alignment · 0.8query-based segmentation · 0.8query disentangling · 0.8dual-prompt schema · 0.8
YearPublicationVenuePosition
2026 MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration
abstract
Yakun Zhu, Yutong Huang, Shengqian Qin, Zhongzhen Huang, Shaoting Zhang, Xiaofan Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yakun Zhu, Yutong Huang, Shengqian Qin, Zhongzhen Huang, Shaoting Zhang 0001, Xiaofan Zhang 0002
ACL (1)4
2026 PathFound: An agentic multimodal model activating evidence-seeking pathological diagnosis
Shengyi Hua, Tianle Shen 0001, Kangzhe Hu, Zhongzhen Huang, Shujuan Ni
Medical Image Anal.5
2025 Surgical Action Planning with Large Language Models
Mengya Xu, Zhongzhen Huang, Xiaofan Zhang 0002, Qi Dou 0001
MICCAI (9)2
2024 ZePT: Zero-Shot Pan-Tumor Segmentation via Query-Disentangling and Self-Prompting
abstract
The long-tailed distribution problem in medical image analysis reflects a high prevalence of common conditions and a low prevalence of rare ones, which poses a significant challenge in developing a unified model capable of identifying rare or novel tumor categories not encountered during training. In this paper, we propose a new Zero-shot Pan-Tumor segmentation framework (ZePT) based on query-disentangling and self-prompting to segment unseen tumor categories beyond the training set. ZePT disentangles the object queries into two subsets and trains them in two stages. Initially, it learns a set of fundamental queries for organ segmentation through an object-aware feature grouping strategy, which gathers organ-level visual features. Subsequently, it refines the other set of advanced queries that focus on the auto-generated visual prompts for unseen tumor segmentation. Moreover, we introduce query-knowledge alignment at the feature level to enhance each query's discriminative representation and generalizability. Extensive experiments on various tumor segmentation tasks demonstrate the performance superiority of ZePT, which surpasses the previous counterparts and evidences the promising ability for zero-shot tumor segmentation in real-world settings.
Yankai Jiang 0003, Zhongzhen Huang, Rongzhao Zhang, Xiaofan Zhang 0002, Shaoting Zhang 0001
CVPR2
2024 Modality-Aware and Shift Mixer for Multi-Modal Brain Tumor Segmentation
abstract
Combining images from multi-modalities is beneficial for exploring various information in computer vision, especially in the medical domain. As an essential part of clinical diagnosis, multi-modal brain tumor segmentation presents a set of distinct challenges for accurately delineating both the normal anatomy and the pathologic deviations caused by the tumor. In this paper, we aim to fuse information on different imaging modalities with the medical domain knowledge to segment tumors. We present MASM, a novel Modality Aware and Shift Mixer that integrates intra-modality and inter-modality dependencies of multi-modal images for effective and robust brain tumor segmentation. Specifically, we introduce a Modality-Aware (MA) module according to neuroimaging studies for modeling the specific modality pair relationships at low levels, and a Modality-Shift (MS) module with specific mosaic patterns is developed to explore the complex relationships that are not addressed by the MA module across modalities efficiently. Experimentally, we outperform previous state-of-the-art approaches on the public Brain Tumor Segmentation dataset. Further qualitative experiments demonstrate the effectiveness and robustness of MASM.
Zhongzhen Huang, Linda Wei, Shaoting Zhang 0001, Xiaofan Zhang 0002
ECAI1
2024 CAT: Coordinating Anatomical-Textual Prompts for Multi-Organ and Tumor Segmentation
abstract
Existing promptable segmentation methods in the medical imaging field primarily consider either textual or visual prompts to segment relevant objects, yet they often fall short when addressing anomalies in medical images, like tumors, which may vary greatly in shape, size, and appearance. Recognizing the complexity of medical scenarios and the limitations of textual or visual prompts, we propose a novel dual-prompt schema that leverages the complementary strengths of visual and textual prompts for segmenting various organs and tumors. Specifically, we introduce $\textbf{\textit{CAT}}$, an innovative model that $\textbf{C}$oordinates $\textbf{A}$natomical prompts derived from 3D cropped images with $\textbf{T}$extual prompts enriched by medical domain knowledge. The model architecture adopts a general query-based design, where prompt queries facilitate segmentation queries for mask prediction. To synergize two types of prompts within a unified framework, we implement a ShareRefiner, which refines both segmentation and prompt queries while disentangling the two types of prompts. Trained on a consortium of 10 public CT datasets, $\textbf{\textit{CAT}}$ demonstrates superior performance in multiple segmentation tasks. Further validation on a specialized in-house dataset reveals the remarkable capacity of segmenting tumors across multiple cancer stages. This approach confirms that coordinating multimodal prompts is a promising avenue for addressing complex scenarios in the medical domain.
Zhongzhen Huang, Yankai Jiang 0003, Rongzhao Zhang, Shaoting Zhang 0001, Xiaofan Zhang 0002
NeurIPS1
2023 KiUT: Knowledge-injected U-Transformer for Radiology Report Generation
abstract
Radiology report generation aims to automatically generate a clinically accurate and coherent paragraph from the X-ray image, which could relieve radiologists from the heavy burden of report writing. Although various image caption methods have shown remarkable performance in the natural image field, generating accurate reports for medical images requires knowledge of multiple modalities, including vision, language, and medical terminology. We propose a Knowledge-injected U-Transformer (KiUT) to learn multi-level visual representation and adaptively distill the information with contextual and clinical knowledge for word prediction. In detail, a U-connection schema between the encoder and decoder is designed to model interactions between different modalities. And a symptom graph and an injected knowledge distiller are developed to assist the report generation. Experimentally, we outperform state-of-the-art methods on two widely used benchmark datasets: IU-Xray and MIMIC-CXR. Further experimental results prove the advantages of our architecture and the complementary benefits of the injected knowledge.
Zhongzhen Huang, Xiaofan Zhang 0002, Shaoting Zhang 0001
CVPR1
2023 One for all: One-stage referring expression comprehension with dynamic reasoning
Zhimin Wei, Zhongzhen Huang, Rui Niu
Neurocomputing3