VLDB 2026 Research / reviewers in the wild / expert
Xiaofan Zhang 0002
dblp:28/9804-2
· DBLP profile ↗
39ranked-venue papers
6as first author
32since 2021 · last 2026
0000-0003-3999-1449ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GuideGen: A Text-Guided Framework for Paired Full-torso Anatomy and CT Volume GenerationabstractThe recently emerging conditional diffusion models seem promising for mitigating the labor and expenses in building large 3D medical imaging datasets. However, previous studies on 3D CT generation primarily focus on specific organs characterized by a local structure and fixed contrast and have yet to fully capitalize on the benefits of both semantic and textual conditions. In this paper, we present GuideGen, a controllable framework based on easily-acquired text prompts to generate anatomical masks and corresponding CT volumes for the entire torso—from chest to pelvis. Our approach includes three core components: a text-conditional semantic synthesizer for creating realistic full-torso anatomies; an anatomy-aware high-dynamic-range (HDR) autoencoder for high-fidelity feature extraction across varying intensity levels; and a latent feature generator that ensures alignment between CT images, anatomical semantics and input prompts. Combined, these components enable data synthesis for segmentation tasks from only textual instructions. To train and evaluate GuideGen, we compile a multi-modality cancer imaging dataset with paired CT and clinical descriptions from 12 public TCIA datasets and one private real-world dataset. Comprehensive evaluations across generation quality, cross-modality alignment, and data usability on multi-organ and tumor segmentation tasks demonstrate GuideGen's superiority over existing CT generation methods. Linrui Dai, Rongzhao Zhang, Yongrui Yu, Xiaofan Zhang 0002 |
AAAI | 4 |
| 2026 | Concepts from Representations: Post-hoc Concept Bottleneck Models via Sparse Decomposition of Visual RepresentationsabstractDeep learning has achieved remarkable success in image recognition, yet their inherent opacity poses challenges for deployment in critical domains. Concept-based interpretations aim to address this by explaining model reasoning through human-understandable concepts. However, existing post-hoc methods and ante-hoc concept bottleneck models (CBMs), suffer from limitations such as unreliable concept relevance, non-visual or labor-intensive concept definitions, and model/data-agnostic assumptions. This paper introduces Post-hoc Concept Bottleneck Model via Representation Decomposition (PCBM-ReD), a novel pipeline that retrofits interpretability onto pretrained opaque models. PCBM-ReD automatically extracts visual concepts from a pre-trained encoder, employs multimodal large language models (MLLMs) to label and filter concepts based on visual identifiability and task relevance, and selects an independent subset via reconstruction-guided optimization. Leveraging CLIP’s visual-text alignment, it decomposes image representations into linear combination of concept embeddings to fit into the CBMs abstraction. Extensive experiments across 11 image classification tasks show PCBM-ReD achieves state-of-the-art accuracy, narrows the performance gap with end-to-end models, and exhibits better interpretability. Shizhan Gong, Xiaofan Zhang 0002, Qi Dou 0001 |
AAAI | 2 |
| 2026 | MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP IntegrationabstractYakun Zhu, Yutong Huang, Shengqian Qin, Zhongzhen Huang, Shaoting Zhang, Xiaofan Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yakun Zhu, Yutong Huang, Shengqian Qin, Zhongzhen Huang, Shaoting Zhang 0001, Xiaofan Zhang 0002 |
ACL (1) | 6 |
| 2025 | Meta-Tool: Unleash Open-World Function Calling Capabilities of General-Purpose Large Language ModelsabstractLarge language models (LLMs) have showcased remarkable capabilities as autonomous agents when augmented with external tools.Equipped with fixed tool sets, LLMs struggle with addressing diverse user inquiries in open-world tasks.To evaluate and boost the performance of LLMs in dealing with complex demands in the real-world, we propose open-world function calling, where LLMs need to retrieve suitable tools from a pre-defined external tool library and use retrieved tools to resolve the user's problem.We introduce Meta-Tool, a versatile and plug-and-play tool retrieval system as the access of LLMs to external tool library.Drawing inspiration from the myriad of enhanced approaches associated with Retrieval-Augmented Generation (RAG), Meta-Tool employs a hypothesizeretrieve-invoke framework.We further propose Meta-Bench, a comprehensive benchmark for evaluating LLMs in open-world function calling and associated tasks.Meta-Bench encompasses 2, 800 dialogues and 7, 361 tools, spanning ten distinct scenarios to provide robust and diverse test categories.In conjunction, we present MT-LLaMA, a finetuned version of LLaMA-3.1, which exhibits remarkable performance improvements.Our empirical experiments reveal that Meta-Tool significantly enhances the ability of advanced LLMs to retrieve and leverage the most suitable tools compared to previous tool retrieval methods.Moreover, our fine-tuning enables even smallersized LLMs to achieve comparable even exceeding results to GPT-4o.Both the benchmark and the model are made publicly available at https://github.com/qinshengqian/Meta-Tool to foster further research and development in the field. Shengqian Qin, Yakun Zhu, Linjie Mu, Shaoting Zhang 0001, Xiaofan Zhang 0002 |
ACL (1) | 5 |
| 2025 | Mesenteric Vasculature Guided Segmentation of Metastatic Lymph Nodes in Colorectal CancerabstractAccurate segmentation of lymph node metastases in colorectal cancer is crucial for disease staging, prognosis evaluation, and treatment planning. However, the identification of metastatic lymph nodes in colorectal cancer from CT images remains extremely challenging due to their small size, indistinguishable appearance, and the potential for distal spread from the primary tumor, making them difficult to differentiate within the complex abdominal cavity. Based on the metastatic patterns of tumors, where lymph node metastases in colorectal cancer often accompany the mesenteric vessels, we propose to leverage the mesenteric vasculature as guidance for the identification of metastatic lymph nodes. Specifically, we utilize vessel segmentation maps of the superior and inferior mesenteric vessels as mesenteric vascular guidance. To explicitly model the spatial relationships between metastatic lymph nodes and mesenteric vessels, we further employ explicit distance relationship modeling to represent the distance from each voxel in the CT image to the nearest mesenteric vessel. In addition, to incorporate mesenteric vascular guidance and explicit distance relationship modeling into the segmentation process of metastatic lymph nodes in colorectal cancer, we integrate a guidance encoder and a guidance signal fusion module into the U-Net segmentation network. Inference-time vascular localization is also employed to assist in localizing the region of interest and filtering segmentation results. Experimental results on the colorectal cancer lymph node metastasis dataset demonstrate the effectiveness of our proposed method. The code will be released at https://github.com/YuriYu12/VesselGuidedSegmentation. Yongrui Yu, Linrui Dai, Shaoting Zhang 0001, Xiaofan Zhang 0002 |
BIBM | 5 |
| 2025 | An LLM-based Framework for Biomedical Terminology Normalization in Social Media via Multi-Agent CollaborationabstractBiomedical Terminology Normalization aims to identify the standard term in a specified termbase for non-standardized mentions from social media or clinical texts, employing the mainstream “Recall and Re-rank” framework. Instead of the traditional pretraining-finetuning paradigm, we would like to explore the possibility of accomplishing this task through a tuning-free paradigm using powerful Large Language Models (LLMs), hoping to address the costs of re-training due to discrepancies of both standard termbases and annotation protocols. Another major obstacle in this task is that both mentions and terms are short texts. Short texts contain an insufficient amount of information that can introduce ambiguity, especially in a biomedical context. Therefore, besides using the advanced embedding model, we implement a Retrieval-Augmented Generation (RAG) based knowledge card generation module. This module introduces an LLM agent that expands the short texts into accurate, harmonized, and more informative descriptions using a search engine and a domain knowledge base. Furthermore, we present an innovative tuning-free agent collaboration framework for the biomedical terminology normalization task in social media. By leveraging the internal knowledge and the reasoning capabilities of LLM, our framework conducts more sophisticated recall, ranking and re-ranking processes with the collaboration of different LLM agents. Experimental results across multiple datasets indicate that our approach exhibits competitive performance. We release our code and data on the github repository JOHNNY-fans/RankNorm. Yongqi Fan, Kui Xue, Zelin Li 0004, Xiaofan Zhang 0002, Tong Ruan |
COLING | 4 |
| 2025 | Interactive Evaluation for Medical LLMs via Task-oriented Dialogue SystemabstractThis study focuses on evaluating proactive communication and diagnostic capabilities of medical Large Language Models (LLMs), which directly impact their effectiveness in patient consultations. In typical medical scenarios, doctors often ask a set of questions to gain a comprehensive understanding of patients’ conditions. We argue that single-turn question-answering tasks such as MultiMedQA are insufficient for evaluating LLMs’ medical consultation abilities. To address this limitation, we developed an evaluation benchmark called Multi-turn Medical Dialogue Evaluation (MMD-Eval), specifically designed to evaluate the proactive communication and diagnostic capabilities of medical LLMs during consultations. Considering the high cost and potential for hallucinations in LLMs, we innovatively trained a task-oriented dialogue system to simulate patients engaging in dialogues with the medical LLMs using our structured medical records dataset. This approach enabled us to generate multi-turn dialogue data. Subsequently, we evaluate the communication skills and medical expertise of the medical LLMs. All resources associated with this study will be made publicly available. Ruoyu Liu, Kui Xue, Xiaofan Zhang 0002, Shaoting Zhang 0001 |
COLING | 3 |
| 2025 | Unleashing the Potential of Vision-Language Pre-Training for 3D Zero-Shot Lesion Segmentation via Mask-Attribute AlignmentabstractRecent advancements in medical vision-language pre-training models have driven significant progress in zero-shot disease recognition. However, transferring image-level knowledge to pixel-level tasks, such as lesion segmentation in 3D CT scans, remains a critical challenge. Due to the complexity and variability of pathological visual characteristics, existing methods struggle to align fine-grained lesion features not encountered during training with disease-related textual representations. In this paper, we present Malenia, a novel multi-scale lesion-level mask-attribute alignment framework, specifically designed for 3D zero-shot lesion segmentation. Malenia improves the compatibility between mask representations and their associated elemental attributes, explicitly linking the visual features of unseen lesions with the extensible knowledge learned from previously seen ones. Furthermore, we design a Cross-Modal Knowledge Injection module to enhance both visual and textual features with mutually beneficial information, effectively guiding the generation of segmentation results. Comprehensive experiments across three datasets and 12 lesion categories validate the superior performance of Malenia. Yankai Jiang 0003, Wenhui Lei, Xiaofan Zhang 0002, Shaoting Zhang 0001 |
ICLR | 3 |
| 2025 | Interactive Segmentation and Report Generation for CT Images
Yannian Gu, Wenhui Lei, Shaoting Zhang 0001, Xiaofan Zhang 0002 |
MICCAI (5) | 5 |
| 2025 | LesionDiffusion: Towards Text-Controlled General Lesion Synthesis
Wenhui Lei, Hengrui Tian 0001, Linrui Dai, Xiaofan Zhang 0002 |
MICCAI (5) | 5 |
| 2025 | Surgical Action Planning with Large Language Models
Mengya Xu, Zhongzhen Huang, Xiaofan Zhang 0002, Qi Dou 0001 |
MICCAI (9) | 4 |
| 2025 | MeNTi: Bridging Medical Calculator and LLM Agent with Nested Tool CallingabstractYakun Zhu, Shaohang Wei, Xu Wang, Kui Xue, Shaoting Zhang, Xiaofan Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yakun Zhu, Shaohang Wei, Kui Xue, Shaoting Zhang 0001, Xiaofan Zhang 0002 |
NAACL (Long Papers) | 6 |
| 2025 | Maxillofacial bone movements-aware dual graph convolution approach for postoperative facial appearance prediction
Xinrui Huang, Dongming He, Zhenming Li, Xiaofan Zhang 0002 |
Medical Image Anal. | 4 |
| 2025 | MedLSAM: Localize and segment anything model for 3D CT images
Wenhui Lei, Wei Xu 0046, Kang Li 0004, Xiaofan Zhang 0002, Shaoting Zhang 0001 |
Medical Image Anal. | 4 |
| 2024 | GECSum: Generative Evaluation-Driven Sequence Level Contrastive Learning for Abstractive SummarizationabstractWhile dominant in abstractive summarization, transformer-based language models with the standard maximum likelihood estimation (MLE) training remain challenged by two discrepancies: the misalignment between token-level training and sequence-level evaluation, and the divergence between teacher-forcing training manner and auto-regressive generation behavior. Recent studies have shown that sequence-level contrastive learning, which utilizes the quality differences between multiple summaries as prior information, can effectively mitigate these issues. However, as certain evaluation metrics often determine the contrastive signals in existing methods, this leads to the model performance aligning with the preferences of these metrics being limited by the evaluation capabilities of these metrics. Inspired by prior works that treat the evaluation of generated text as a text generation problem, we propose a generative evaluation-driven contrastive learning framework, which leverages the semantic understanding capabilities of the abstractive model itself to evaluate summary in reference-based settings. In this way, our method establishes a connection between the model’s reference-based evaluation and reference-free generation scenarios, allowing them to share the benefits of model capability enhancements. Extensive experiments on four summarization datasets demonstrate that our method outperforms the previous state-of-the-art regarding comprehensive performance. Various empirical analyses further substantiate the effectiveness of our method. Jiawen Xie, Shaoting Zhang 0001, Xiaofan Zhang 0002 |
LREC/COLING | 3 |
| 2024 | ZePT: Zero-Shot Pan-Tumor Segmentation via Query-Disentangling and Self-PromptingabstractThe long-tailed distribution problem in medical image analysis reflects a high prevalence of common conditions and a low prevalence of rare ones, which poses a significant challenge in developing a unified model capable of identifying rare or novel tumor categories not encountered during training. In this paper, we propose a new Zero-shot Pan-Tumor segmentation framework (ZePT) based on query-disentangling and self-prompting to segment unseen tumor categories beyond the training set. ZePT disentangles the object queries into two subsets and trains them in two stages. Initially, it learns a set of fundamental queries for organ segmentation through an object-aware feature grouping strategy, which gathers organ-level visual features. Subsequently, it refines the other set of advanced queries that focus on the auto-generated visual prompts for unseen tumor segmentation. Moreover, we introduce query-knowledge alignment at the feature level to enhance each query's discriminative representation and generalizability. Extensive experiments on various tumor segmentation tasks demonstrate the performance superiority of ZePT, which surpasses the previous counterparts and evidences the promising ability for zero-shot tumor segmentation in real-world settings. Yankai Jiang 0003, Zhongzhen Huang, Rongzhao Zhang, Xiaofan Zhang 0002, Shaoting Zhang 0001 |
CVPR | 4 |
| 2024 | Modality-Aware and Shift Mixer for Multi-Modal Brain Tumor SegmentationabstractCombining images from multi-modalities is beneficial for exploring various information in computer vision, especially in the medical domain. As an essential part of clinical diagnosis, multi-modal brain tumor segmentation presents a set of distinct challenges for accurately delineating both the normal anatomy and the pathologic deviations caused by the tumor. In this paper, we aim to fuse information on different imaging modalities with the medical domain knowledge to segment tumors. We present MASM, a novel Modality Aware and Shift Mixer that integrates intra-modality and inter-modality dependencies of multi-modal images for effective and robust brain tumor segmentation. Specifically, we introduce a Modality-Aware (MA) module according to neuroimaging studies for modeling the specific modality pair relationships at low levels, and a Modality-Shift (MS) module with specific mosaic patterns is developed to explore the complex relationships that are not addressed by the MA module across modalities efficiently. Experimentally, we outperform previous state-of-the-art approaches on the public Brain Tumor Segmentation dataset. Further qualitative experiments demonstrate the effectiveness and robustness of MASM. Zhongzhen Huang, Linda Wei, Shaoting Zhang 0001, Xiaofan Zhang 0002 |
ECAI | 4 |
| 2024 | PathoTune: Adapting Visual Foundation Model to Pathological Specialists
Jiaxuan Lu, Fang Yan 0002, Xiaofan Zhang 0002, Yue Gao 0002, Shaoting Zhang 0001 |
MICCAI (4) | 3 |
| 2024 | Incorporating Clinical Guidelines Through Adapting Multi-modal Large Language Model for Prostate Cancer PI-RADS Scoring
Manxi Lin, Hongda Guo, Xiaofan Zhang 0002, Ka Fung Peter Chiu, Aasa Feragen, Qi Dou 0001 |
MICCAI (5) | 4 |
| 2024 | CAT: Coordinating Anatomical-Textual Prompts for Multi-Organ and Tumor SegmentationabstractExisting promptable segmentation methods in the medical imaging field primarily consider either textual or visual prompts to segment relevant objects, yet they often fall short when addressing anomalies in medical images, like tumors, which may vary greatly in shape, size, and appearance. Recognizing the complexity of medical scenarios and the limitations of textual or visual prompts, we propose a novel dual-prompt schema that leverages the complementary strengths of visual and textual prompts for segmenting various organs and tumors. Specifically, we introduce $\textbf{\textit{CAT}}$, an innovative model that $\textbf{C}$oordinates $\textbf{A}$natomical prompts derived from 3D cropped images with $\textbf{T}$extual prompts enriched by medical domain knowledge. The model architecture adopts a general query-based design, where prompt queries facilitate segmentation queries for mask prediction. To synergize two types of prompts within a unified framework, we implement a ShareRefiner, which refines both segmentation and prompt queries while disentangling the two types of prompts. Trained on a consortium of 10 public CT datasets, $\textbf{\textit{CAT}}$ demonstrates superior performance in multiple segmentation tasks. Further validation on a specialized in-house dataset reveals the remarkable capacity of segmenting tumors across multiple cancer stages. This approach confirms that coordinating multimodal prompts is a promising avenue for addressing complex scenarios in the medical domain. Zhongzhen Huang, Yankai Jiang 0003, Rongzhao Zhang, Shaoting Zhang 0001, Xiaofan Zhang 0002 |
NeurIPS | 5 |
| 2024 | PathoDuet: Foundation models for pathological slide analysis of H&E and IHC stains
Shengyi Hua, Fang Yan 0002, Tianle Shen 0001, Lei Ma 0006, Xiaofan Zhang 0002 |
Medical Image Anal. | 5 |
| 2024 | USFM: A universal ultrasound foundation model generalized to tasks and organs towards label efficient image analysis
Jing Jiao, Menghua Xia, Yi Huang 0018, Xiaofan Zhang 0002, Shichong Zhou, Yuanyuan Wang 0001, Yi Guo 0002 |
Medical Image Anal. | 8 |
| 2024 | Deblurring masked image modeling for ultrasound image analysis
Qingbo Kang, Qicheng Lao, Jingyan Liu, Huahui Yi, Buyun Ma, Xiaofan Zhang 0002, Kang Li 0004 |
Medical Image Anal. | 7 |
| 2024 | One-Shot Weakly-Supervised Segmentation in 3D Medical ImagesabstractDeep neural networks typically require accurate and a large number of annotations to achieve outstanding performance in medical image segmentation. One-shot and weakly-supervised learning are promising research directions that reduce labeling effort by learning a new class from only one annotated image and using coarse labels instead, respectively. In this work, we present an innovative framework for 3D medical image segmentation with one-shot and weakly-supervised settings. Firstly a propagation-reconstruction network is proposed to propagate scribbles from one annotated volume to unlabeled 3D images based on the assumption that anatomical patterns in different human bodies are similar. Then a multi-level similarity denoising module is designed to refine the scribbles based on embeddings from anatomical- to pixel-level. After expanding the scribbles to pseudo masks, we observe the miss-classified voxels mainly occur at the border region and propose to extract self-support prototypes for the specific refinement. Based on these weakly-supervised segmentation results, we further train a segmentation model for the new class with the noisy label training strategy. Experiments on three CT and one MRI datasets show the proposed method obtains significant improvement over the state-of-the-art methods and performs robustly even under severe class imbalance and low contrast. Code is publicly available at https://github.com/LWHYC/OneShot_WeaklySeg. Wenhui Lei, Ran Gu, Xinglong Liu, Guotai Wang, Xiaofan Zhang 0002, Shaoting Zhang 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2023 | MidMed: Towards Mixed-Type Dialogues for Medical ConsultationabstractXiaoming Shi, Zeming Liu, Chuan Wang, Haitao Leng, Kui Xue, Xiaofan Zhang, Shaoting Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zeming Liu, Chuan Wang 0002, Haitao Leng, Kui Xue, Xiaofan Zhang 0002, Shaoting Zhang 0001 |
ACL (1) | 6 |
| 2023 | KiUT: Knowledge-injected U-Transformer for Radiology Report GenerationabstractRadiology report generation aims to automatically generate a clinically accurate and coherent paragraph from the X-ray image, which could relieve radiologists from the heavy burden of report writing. Although various image caption methods have shown remarkable performance in the natural image field, generating accurate reports for medical images requires knowledge of multiple modalities, including vision, language, and medical terminology. We propose a Knowledge-injected U-Transformer (KiUT) to learn multi-level visual representation and adaptively distill the information with contextual and clinical knowledge for word prediction. In detail, a U-connection schema between the encoder and decoder is designed to model interactions between different modalities. And a symptom graph and an injected knowledge distiller are developed to assist the report generation. Experimentally, we outperform state-of-the-art methods on two widely used benchmark datasets: IU-Xray and MIMIC-CXR. Further experimental results prove the advantages of our architecture and the complementary benefits of the injected knowledge. Zhongzhen Huang, Xiaofan Zhang 0002, Shaoting Zhang 0001 |
CVPR | 2 |
| 2023 | Efficient Subclass Segmentation in Medical Images
Linrui Dai, Wenhui Lei, Xiaofan Zhang 0002 |
MICCAI (2) | 3 |
| 2023 | Contrastive Semi-Supervised Learning for Domain Adaptive Segmentation Across Similar Anatomical StructuresabstractConvolutional Neural Networks (CNNs) have achieved state-of-the-art performance for medical image segmentation, yet need plenty of manual annotations for training. Semi-Supervised Learning (SSL) methods are promising to reduce the requirement of annotations, but their performance is still limited when the dataset size and the number of annotated images are small. Leveraging existing annotated datasets with similar anatomical structures to assist training has a potential for improving the model's performance. However, it is further challenged by the cross-anatomy domain shift due to the image modalities and even different organs in the target domain. To solve this problem, we propose Contrastive Semi-supervised learning for Cross Anatomy Domain Adaptation (CS-CADA) that adapts a model to segment similar structures in a target domain, which requires only limited annotations in the target domain by leveraging a set of existing annotated images of similar structures in a source domain. We use Domain-Specific Batch Normalization (DSBN) to individually normalize feature maps for the two anatomical domains, and propose a cross-domain contrastive learning strategy to encourage extracting domain invariant features. They are integrated into a Self-Ensembling Mean-Teacher (SE-MT) framework to exploit unlabeled target domain images with a prediction consistency constraint. Extensive experiments show that our CS-CADA is able to solve the challenging cross-anatomy domain shift problem, achieving accurate segmentation of coronary arteries in X-ray images with the help of retinal vessel images and cardiac MR images with the help of fundus images, respectively, given only a small number of annotations in the target domain. Our code is available at https://github.com/HiLab-git/DAG4MIA. Ran Gu, Jingyang Zhang, Guotai Wang, Wenhui Lei, Tao Song 0002, Xiaofan Zhang 0002, Kang Li 0004, Shaoting Zhang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2022 | DigestPath: A benchmark dataset with challenge review for the pathological detection and segmentation of digestive-system
Qian Da, Zhongyu Li 0002, Yanfei Zuo, Chenbin Zhang, Jingxin Liu 0005, Wen Chen 0001, Jiahui Li 0005, Dou Xu, Hongmei Yi, Zhe Wang 0043, Li Zhang 0040, Xianying He, Xiaofan Zhang 0002, Ke Mei, Chuang Zhu, Weizeng Lu, LinLin Shen, Jun Shi 0006, Jun Li 0106, Sreehari S, Ganapathy Krishnamurthi, Jiangcheng Yang, Tiancheng Lin 0001, Qingyu Song 0004, Xuechen Liu 0004, Simon Graham, Raja Muhammad Saad Bashir, Canqian Yang, Shaofei Qin, Xinmei Tian 0001, Jie Zhao 0014, Dimitris N. Metaxas, Hongsheng Li 0001, Chaofu Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 17 |
| 2022 | Segmentation only uses sparse annotations: Unified weakly and semi-supervised learning in medical images
Feng Gao 0023, Minhao Hu, Min-Er Zhong, Shixiang Feng, Xuwei Tian, Xiaochun Meng, Mayidili Nijiati, Zeping Huang, Minyi Lv, Tao Song 0002, Xiaofan Zhang 0002, Xiaoguang Zou, Xiaojian Wu |
Medical Image Anal. | 11 |
| 2022 | WORD: A large scale dataset, benchmark and clinical applicable study for abdominal organ segmentation from CT image
Xiangde Luo, Wenjun Liao, Jianghong Xiao, Jieneng Chen, Tao Song 0002, Xiaofan Zhang 0002, Kang Li 0004, Dimitris N. Metaxas, Guotai Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 6 |
| 2021 | Large-scale gastric cancer screening and localization using multi-task deep neural network
Xiaofan Zhang 0002, Lingjun Song, Liren Jiang, Wen Chen 0022, Chenbin Zhang, Jiahui Li 0005, Jiji Yang, Qi Duan, Wanyuan Chen, Xianglei He, Jinshuang Fan, Weihai Jiang, Li Zhang 0040, Chengmin Qiu, Minmin Gu, Yangqiong Zhang, Guangyin Peng, Weiwei Shen, Guohui Fu |
Neurocomputing | 2 |
| 2018 | Large-scale retrieval for medical image analytics: A comprehensive review
Zhongyu Li 0002, Xiaofan Zhang 0002, Henning Müller, Shaoting Zhang 0001 |
Medical Image Anal. | 2 |
| 2016 | Embedding Label Structures for Fine-Grained Feature RepresentationabstractRecent algorithms in convolutional neural networks (CNN) considerably advance the fine-grained image classification, which aims to differentiate subtle differences among subordinate classes. However, previous studies have rarely focused on learning a fined-grained and structured feature representation that is able to locate similar images at different levels of relevance, e.g., discovering cars from the same make or the same model, both of which require high precision. In this paper, we propose two main contributions to tackle this problem. 1) A multitask learning framework is designed to effectively learn fine-grained feature representations by jointly optimizing both classification and similarity constraints. 2) To model the multi-level relevance, label structures such as hierarchy or shared attributes are seamlessly embedded into the framework by generalizing the triplet loss. Extensive and thorough experiments have been conducted on three finegrained datasets, i.e., the Stanford car, the Car-333, and the food datasets, which contain either hierarchical labels or shared attributes. Our proposed method has achieved very competitive performance, i.e., among state-of-the-art classification accuracy when not using parts. More importantly, it significantly outperforms previous fine-grained feature representations for image retrieval at different levels of relevance. Xiaofan Zhang 0002, Feng Zhou 0002, Yuanqing Lin, Shaoting Zhang 0001 |
CVPR | 1 |
| 2016 | Fusing Heterogeneous Features From Stacked Sparse Autoencoder for Histopathological Image AnalysisabstractIn the analysis of histopathological images, both holistic (e.g., architecture features) and local appearance features demonstrate excellent performance, while their accuracy may vary dramatically when providing different inputs. This motivates us to investigate how to fuse results from these features to enhance the accuracy. Particularly, we employ content-based image retrieval approaches to discover morphologically relevant images for image-guided diagnosis, using holistic and local features, both of which are generated from the cell detection results by a stacked sparse autoencoder. Because of the dramatically different characteristics and representations of these heterogeneous features (i.e., holistic and local), their results may not agree with each other, causing difficulties for traditional fusion methods. In this paper, we employ a graph-based query-specific fusion approach where multiple retrieval results (i.e., rank lists) are integrated and reordered based on a fused graph. The proposed method is capable of combining the strengths of local or holistic features adaptively for different inputs. We evaluate our method on a challenging clinical problem, i.e., histopathological image-guided diagnosis of intraductal breast lesions, and it achieves 91.67% classification accuracy on 120 breast tissue images from 40 patients. Xiaofan Zhang 0002, Hang Dou, Tao Ju 0001, Jun Xu 0005, Shaoting Zhang 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2015 | Fine-grained histopathological image analysis via robust segmentation and large-scale retrievalabstractComputer-aided diagnosis of medical images requires thorough analysis of image details. For example, examining all cells enables fine-grained categorization of histopathological images. Traditional computational methods may have efficiency issues when performing such detailed analysis. In this paper, we propose a robust and scalable solution to achieve this. Specifically, a robust segmentation method is developed to delineate region-of-interests (e.g., cells) accurately, using hierarchical voting and repulsive active contour. A hashing-based large-scale retrieval approach is also designed to examine and classify them by comparing with a massive training database. We evaluate this proposed framework on a challenging and important clinical use case, i.e., differentiation of two types of lung cancers (the adenocarcinoma and the squamous carcinoma), using thousands of histopathological images extracted from hundreds of patients. Our method has achieved promising performance, i.e., 87.3% accuracy and 1.68 seconds by searching among half-million cells. Xiaofan Zhang 0002, Hai Su, Lin Yang 0002, Shaoting Zhang 0001 |
CVPR | 1 |
| 2015 | High-throughput histopathological image analysis via robust cell segmentation and hashing
Xiaofan Zhang 0002, Fuyong Xing, Hai Su, Lin Yang 0002, Shaoting Zhang 0001 |
Medical Image Anal. | 1 |
| 2015 | Towards Large-Scale Histopathological Image Analysis: Hashing-Based Image RetrievalabstractAutomatic analysis of histopathological images has been widely utilized leveraging computational image-processing methods and modern machine learning techniques. Both computer-aided diagnosis (CAD) and content-based image-retrieval (CBIR) systems have been successfully developed for diagnosis, disease detection, and decision support in this area. Recently, with the ever-increasing amount of annotated medical data, large-scale and data-driven methods have emerged to offer a promise of bridging the semantic gap between images and diagnostic information. In this paper, we focus on developing scalable image-retrieval techniques to cope intelligently with massive histopathological images. Specifically, we present a supervised kernel hashing technique which leverages a small amount of supervised information in learning to compress a 10 000-dimensional image feature vector into only tens of binary bits with the informative signatures preserved. These binary codes are then indexed into a hash table that enables real-time retrieval of images in a large database. Critically, the supervised information is employed to bridge the semantic gap between low-level image features and high-level diagnostic information. We build a scalable image-retrieval framework based on the supervised hashing technique and validate its performance on several thousand histopathological images acquired from breast microscopic tissues. Extensive evaluations are carried out in terms of image classification (i.e., benign versus actionable categorization) and retrieval tests. Our framework achieves about 88.1% classification accuracy as well as promising time efficiency. For example, the framework can execute around 800 queries in only 0.01 s, comparing favorably with other commonly used dimensionality reduction and feature selection methods. Xiaofan Zhang 0002, Wei Liu 0005, Murat Dundar, Sunil Badve, Shaoting Zhang 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2014 | Mining Histopathological Images via Composite Hashing and Online Learning
Xiaofan Zhang 0002, Lin Yang 0002, Wei Liu 0005, Hai Su, Shaoting Zhang 0001 |
MICCAI (2) | 1 |