EDBT 2026 Demo / reviewers in the wild / expert
Junjun He
dblp:128/7027
· DBLP profile ↗
74ranked-venue papers
3as first author
63since 2021 · last 2026
0000-0002-1813-1784ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 3 first-author · 31 since 2021Artificial intelligence and machine learning · 37 · 2 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 30 · 1 first-author · 29 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AIabstractDespite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annotations into high-quality image-text pairs. This dataset offers comprehensive task coverage, diverse modalities, and rich image-text data. Building upon this dataset, we develop GMAI-VL, a 7B-parameter general medical vision-language model, with a three-stage training strategy that enhances the integration of visual and textual information. This approach significantly improves the model's ability to process multimodal data, supporting accurate diagnoses and clinical decision-making. Experiments show that GMAI-VL achieves state-of-the-art performance across various multimodal medical tasks, including visual question answering and medical image diagnosis. Tianbin Li, Yanzhou Su, Wei Li 0320, Zhe Chen 0017, Ziyan Huang, Guoan Wang, Chenglong Ma 0002, Yanjun Li 0007, Shixiang Tang, Xiaowei Hu 0001, Zhongying Deng, Yuanfeng Ji, Jin Ye 0002, Yu Qiao 0001, Junjun He |
AAAI | 19 |
| 2026 | S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything Without SupervisionabstractRecent self-supervised image segmentation models have achieved promising performance on semantic segmentation and class-agnostic instance segmentation. However, their pretraining schedule is multi-stage, requiring a time-consuming pseudo-masks generation process between each training epoch. This time-consuming offline process not only makes it difficult to scale with training dataset size, but also leads to sub-optimal solutions due to its discontinuous optimization routine. To solve these, we first present a novel pseudo-mask algorithm, Fast Universal Agglomerative Pooling (UniAP). Each layer of UniAP can identify groups of similar nodes in parallel, allowing to generate both semantic-level and instance-level and multi-granular pseudo-masks within ens of milliseconds for one image. Based on the fast UniAP, we propose the Scalable Self-Supervised Universal Segmentation (S2-UniSeg), which employs a student and a momentum teacher for continuous pretraining. A novel segmentation-oriented pretext task, Query-wise Self-Distillation (QuerySD), is proposed to pretrain S2-UniSeg to learn the local-to-global correspondences. Under the same setting, S2-UniSeg outperforms the SOTA UnSAM model, achieving notable improvements of AP+6.9 on COCO, AR+11.1 on UVO, PixelAcc+4.5 on COCOStuff-27, RQ+8.0 on Cityscapes. After scaling up to a larger 2M-image subset of SA-1B, S2-UniSeg further achieves performance gains on all four benchmarks. Jin Ye 0002, Hongqiu Wang, Changkai Ji, Jiashi Lin, Ziyan Huang, Chenglong Ma 0002, Tianbin Li, Junjun He, Lei Zhu 0003 |
AAAI | 12 |
| 2026 | Multimodal Medical Image Binding via Shared Text EmbeddingsabstractMedical image analysis increasingly relies on the integration of multiple imaging modalities to capture complementary anatomical and functional information, enabling more accurate diagnosis and treatment planning. Achieving aligned feature representations across these diverse modalities is therefore important for effective multimodal analysis. While contrastive language-image pre-training (CLIP) and its variant have enabled image-text alignments, they require explicitly paired data between arbitrary two modalities, which is difficult to acquire in medical contexts. To address the gap, we present Multimodal Medical Image Binding with Text (M3Bind), a novel pre-training framework that enables seamless alignment of multiple medical imaging modalities through a shared text representation space without requiring explicit paired data between any two medical image modalities. Specifically, based on the insight that different images can naturally bind with text, M3Bind first fine-tunes pre-trained CLIP-like image-text models, which are derived from different medical modalities, to align their modality-specific text embedding space while preserving their original image-text alignments. Subsequently, we distill these modality-specific text encoders into a unified model, creating a shared text embedding space. Notably, M3Bind is a flexible framework in which the selection of CLIP-like models is not fixed and can be adapted according to the requirements of the task. Experiments on X-ray, CT, retina, ECG, and pathological images on multiple downstream tasks demonstrate that M3Bind achieves competitive or even superior performance in zero-shot, few-shot classification and cross-modal retrieval tasks compared to its CLIP-like counterparts. These results validate M3Bind's effectiveness in achieving cross-image-modal alignment for medical analysis. Suyang Xi, Chicheng Jin, Chong Zhong, Junjun He, Catherine C. Liu, Yiqing Shen 0003 |
WACV | 7 |
| 2026 | Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey and Benchmark
Yi Xin 0003, Jianjiang Yang, Yuntao Du 0001, Haoxing Chen, Kangrui Cen, Yangfan He, Yuewen Cao, Junjun He, Xiaokang Yang 0001, Guangtao Zhai, Ming-Hsuan Yang 0001, Xiaohong Liu 0001 |
Int. J. Comput. Vis. | 11 |
| 2026 | MorphoNet: Morphological sub-region-based structure learning for WSI analysisabstract• Propose MorphoNet framework to capture morphological patterns and long-range tissue structures in WSIs. • Develop a Morphological Sub-Region Grouping mechanism to model WSIs as spatially coherent sub-regions. • Introduce a spatial-aware clustering approach and a sub-region aggregation strategy to derive sub-region embeddings. • Achieve superior performance across 10 public benchmarks, outperforming state-of-the-art methods in tumor subtyping and survival prediction. Representation learning of Whole slide image (WSI) is fundamental to computational pathology, enabling tasks such as tumor subtyping, survival prediction, and cancer grading. Existing methods typically tile WSIs into thousands of small patches and aggregate patch features into slide-level embeddings, but this patch-centric paradigm suffers from redundancy and suboptimal spatial modeling. Built upon these patch-level embeddings, Multiple Instance Learning (MIL) methods overfit to scattered discriminative patches, graph-based models mainly capture local neighborhoods, and prototype-based approaches often ignore spatial coherence and under-represent rare tissue patterns. To address these challenges, we propose MorphoNet, a Morph ological structure learning Net work that captures long-range spatial tissue relationships while extracting informative morphological patterns. The key idea of MorphoNet is Morphological Sub-Region Grouping (MSRG), which clusters spatially adjacent patches with similar appearance into compact sub-region embeddings, reducing redundancy and forming semantically coherent morphological units. Sub-region graph is then constructed and processed by a lightweight Graph Neural Network (GNN) to model contextual dependencies and derive slide-level representations. Importantly, MSRG is a plug-and-play module that can be integrated into MIL, graph-based, and prototype-based pipelines, consistently improving their performance. Experiments on ten public benchmarks demonstrate that MorphoNet achieves superior performance on tumor subtyping and survival prediction. Our code is available at https://github.com/fuying-wang/MorphoNet . Fuying Wang, Junjun He, Liansheng Wang 0002, Jianning Chen, Lequan Yu |
Medical Image Anal. | 4 |
| 2026 | Generative models for noise-robust training in unsupervised domain adaptation
Zhongying Deng, Da Li 0001, Junjun He, Xiaojiang Peng, Yi-Zhe Song, Tao Xiang 0002 |
Pattern Recognit. | 3 |
| 2026 | Brain foundation models with hypergraph dynamic adapter for brain disease analysisabstractBrain diseases, such as Alzheimer’s disease and brain tumors, present profound challenges due to their complexity and societal impact. Recent advancements in brain foundation models have shown significant promise in addressing a range of brain-related tasks. However, current brain foundation models are limited by task and data homogeneity, restricted generalization beyond segmentation or classification, and inefficient adaptation to diverse clinical tasks. In this work, we propose SAM-Brain3D, a brain-specific foundation model trained on over 66,000 brain image-label pairs across 14 MRI sub-modalities, and Hypergraph Dynamic Adapter (HyDA), a lightweight adapter for efficient and effective downstream adaptation. SAM-Brain3D captures detailed brain-specific anatomical and modality priors for segmenting diverse brain targets and broader downstream tasks. HyDA leverages hypergraphs to fuse complementary multi-modal data and dynamically generate patient-specific convolutional kernels for multi-scale feature fusion and personalized patient-wise adaptation. Together, our framework excels across a broad spectrum of brain disease segmentation and classification tasks. Extensive experiments demonstrate that our method consistently outperforms existing state-of-the-art approaches, offering a new paradigm for brain disease analysis through multi-modal, multi-scale, and dynamic foundation modeling. Zhongying Deng, Ziyan Huang, Lipei Zhang, Angelica I. Avilés-Rivero, Chaoyu Liu, Junjun He, Zoe Kourtzi, Carola-Bibiane Schönlieb |
Pattern Recognit. | 7 |
| 2026 | Universal pre-training for generalizable incomplete-view CT reconstruction
Chenglong Ma 0002, Zilong Li 0001, Junjun He, Junping Zhang, Yi Zhang 0018, Hongming Shan |
Pattern Recognit. | 3 |
| 2026 | DiffM4RI: A Latent Diffusion Model With Modality Inpainting for Synthesizing Missing Modalities in MRI AnalysisabstractFoundation Models (FMs) have shown great promise for multimodal medical image analysis such as Magnetic Resonance Imaging (MRI). However, certain MRI sequences may be unavailable due to various constraints, such as limited scanning time, patient discomfort, or scanner limitations. The absence of certain modalities can hinder the performance of FMs in clinical applications, making effective missing modality imputation crucial for ensuring their applicability. Previous approaches, including generative adversarial networks (GANs), have been employed to synthesize missing modalities in either a one-to-one or many-to-one manner. However, these methods have limitations, as they require training a new model for different missing scenarios and are prone to mode collapse, generating limited diversity in the synthesized images. To address these challenges, we propose DiffM4RI, a diffusion model for many-to-many missing modality imputation in MRI. DiffM4RI innovatively formulates the missing modality imputation as a modality-level inpainting task, enabling it to handle arbitrary missing modality situations without the need for training multiple networks. Experiments on the BraTs datasets demonstrate DiffM4RI can achieve an average SSIM improvement of 0.15 over MustGAN, 0.1 over SynDiff, and 0.02 over VQ-VAE-2. These results highlight the potential of DiffM4RI in enhancing the reliability of FMs in clinical applications. The code is available athttps://github.com/27yw/DiffM4RI. Zhetao Guo, Yuxiang Ren, Yushi Shen, Junjun He, Jing Ke, Yiqing Shen 0003 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World ConversationabstractRecent multimodal large language models (MLLMs) have demonstrated significant potential in open-ended conversation, generating more accurate and personalized responses. However, their abilities to memorize, recall, and reason in sustained interactions within real-world scenarios remain underexplored. This paper introduces MMRC, a Multi-Modal Real-world Conversation benchmark for evaluating six core open-ended abilities of MLLMs: information extraction, multi-turn reasoning, information update, image management, memory recall, and answer refusal. With data collected from real-world scenarios, MMRC comprises 5,120 conversations and 28,720 corresponding manually labeled questions, posing a significant challenge to existing MLLMs. Evaluations on 20 MLLMs in MMRC indicate an accuracy drop during open-ended interactions. We identify four common failure patterns: long-term memory degradation, inadequacies in updating factual knowledge, accumulated assumption of error propagation, and reluctance to “say no.” To mitigate these issues, we propose a simple yet effective NOTE-TAKING strategy, which can record key information from the conversation and remind the model during its responses, enhancing conversational capabilities. Experiments across six MLLMs demonstrate significant performance improvements. Haochen Xue, Yexin Liu, Qidong Huang, Yulong Li 0002, Zhongxing Xu, Chong Zhang 0006, Yutong Xie 0001, Muhammad Imran Razzak, ZongYuan Ge, Jionglong Su, Junjun He, Yu Qiao 0001 |
ACL (1) | 15 |
| 2025 | SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image UnderstandingabstractDespite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essential contextual information at the whole-slide level. The lack of large-scale instruction datasets and the gigapixel scale of whole slide images (WSIs) pose significant developmental challenges. In this paper, we present SlideChat, the first vision-language assistant capable of understanding gigapixel whole-slide images, exhibiting excellent multimodal conversational capability and response complex instruction across diverse pathology scenarios. To support its development, we created SlideInstruction, the largest instruction-following dataset for WSIs consisting of 4.2K WSI captions and 176K VQA pairs with multiple categories. Furthermore, we propose SlideBench, a multimodal benchmark that incorporates captioning and VQA tasks to assess SlideChat’s capabilities in various settings such as microscopy, diagnosis and clinical. Compared to both general and specialized MLLMs, SlideChat exhibits exceptional capabilities, achieving state-of-the-art performance on 18 of 22 tasks. For example, it achieved an overall accuracy of 81.17% on SlideBench-VQA (TCGA), and 54.15% on SlideBench-VQA (BCNB). Our code, data, and model is publicly accessible at https://uni-medical.github.io/SlideChat.github.io. Guoan Wang, Yuanfeng Ji, Yanjun Li 0007, Jin Ye 0002, Tianbin Li, Rongshan Yu, Yu Qiao 0001, Junjun He |
CVPR | 10 |
| 2025 | Interactive Medical Image Segmentation: A Benchmark Dataset and BaselineabstractInteractive Medical Image Segmentation (IMIS) has long been constrained by the limited availability of large-scale, diverse, and densely annotated datasets, which hinders model generalization and consistent evaluation across different models. In this paper, we introduce the IMed-361M benchmark dataset, a significant advancement in general IMIS research. First, we collect and standardize over 6.4 million medical images and their corresponding ground truth masks from multiple data sources. Then, leveraging the strong object recognition capabilities of a vision foundational model, we automatically generated dense interactive masks for each image and ensured their quality through rigorous quality control and granularity management. Unlike previous datasets, which are limited by specific modalities or sparse annotations, IMed-361M spans 14 modalities and 204 segmentation targets, totaling 361 million masks—an average of 56 masks per image. Finally, we developed an IMIS baseline network on this dataset that supports high-quality mask generation through interactive inputs, including clicks, bounding boxes, text prompts, and their combinations. We evaluate its performance on medical image segmentation tasks from multiple perspectives, demonstrating superior accuracy and scalability compared to existing interactive segmentation models. To facilitate research on foundational models in medical computer vision, we release the IMed-361M and model at https://github.com/uni-medical/IMIS-Bench. Junlong Cheng, Jin Ye 0002, Guoan Wang, Tianbin Li, Haoyu Wang 0010, He Yao, Yanzhou Su, Min Zhu 0005, Junjun He |
CVPR | 13 |
| 2025 | Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
Qi Chen 0014, Xinze Zhou, Hao Chen 0011, Zekun Jiang, Ziyan Huang, Dexin Yu, Junjun He, Yefeng Zheng 0001, Ling Shao 0001, Alan L. Yuille, Zongwei Zhou |
ICCV | 10 |
| 2025 | Fontanimate: High Quality Few-Shot Font Generation Via Animating Font Transfer Process
Kainan Yan, Shitian Zhao, Jie Wen 0001, Junjun He, Peng Gao 0007 |
ICCV | 7 |
| 2025 | OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language PretrainingabstractSurgical practice involves complex visual interpretation, procedural skills, and advanced medical knowledge, making surgical vision-language pretraining (VLP) particularly challenging due to this complexity and the limited availability of annotated data. To address the gap, we propose OphCLIP, a hierarchical retrieval-augmented vision-language pretraining framework specifically designed for ophthalmic surgical workflow understanding. OphCLIP leverages the OphVL dataset we constructed, a large-scale and comprehensive collection of over 375K hierarchically structured video-text pairs with tens of thousands of different combinations of attributes (surgeries, phases/operations/actions, instruments, medications, as well as more advanced aspects like the causes of eye diseases, surgical objectives, and postoperative recovery recommendations, etc). These hierarchical video-text correspondences enable OphCLIP to learn both fine-grained and long-term visual representations by aligning short video clips with detailed narrative descriptions and full videos with structured titles, capturing intricate surgical details and high-level procedural insights, respectively. Our OphCLIP also designs a retrieval-augmented pretraining framework to leverage the underexplored large-scale silent surgical procedure videos, automatically retrieving semantically relevant content to enhance the representation learning of narrative videos. Evaluation across 11 datasets for phase recognition and multi-instrument identification shows OphCLIP's robust generalization and superior performance. Kun Yuan 0004, Yaling Shen, Xiaohao Xu, Wei Li 0320, Zhongxing Xu, Zelin Peng, Siyuan Yan, Vinkle Srivastav, Diping Song, Tianbin Li, Danli Shi, Jin Ye 0002, Nicolas Padoy, Nassir Navab, Junjun He, ZongYuan Ge |
ICCV | 19 |
| 2025 | Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution GenerationabstractSora unveils the potential of scaling Diffusion Transformer (DiT) for generating photorealistic images and videos at arbitrary resolutions, aspect ratios, and durations, yet it still lacks sufficient implementation details. In this paper, we introduce the Lumina-T2X family -- a series of Flow-based Large Diffusion Transformers (Flag-DiT) equipped with zero-initialized attention, as a simple and scalable generative framework that can be adapted to various modalities, e.g., transforming noise into images, videos, multi-view 3D objects, or audio clips conditioned on text instructions. By tokenizing the latent spatial-temporal space and incorporating learnable placeholders such as |[nextline]| and |[nextframe]| tokens, Lumina-T2X seamlessly unifies the representations of different modalities across various spatial-temporal resolutions. Advanced techniques like RoPE, KQ-Norm, and flow matching enhance the stability, flexibility, and scalability of Flag-DiT, enabling models of Lumina-T2X to scale up to 7 billion parameters and extend the context window to 128K tokens. This is particularly beneficial for creating ultra-high-definition images with our Lumina-T2I model and long 720p videos with our Lumina-T2V model. Remarkably, Lumina-T2I, powered by a 5-billion-parameter Flag-DiT, requires only 35% of the training computational costs of a 600-million-parameter naive DiT (PixArt-alpha), indicating that increasing the number of parameters significantly accelerates convergence of generative models without compromising visual quality. Our further comprehensive analysis underscores Lumina-T2X's preliminary capability in resolution extrapolation, high-resolution editing, generating consistent 3D views, and synthesizing videos with seamless transitions. All code and checkpoints of Lumina-T2X are released at https://github.com/Alpha-VLLM/Lumina-T2X to further foster creativity, transparency, and diversity in the generative AI community. Peng Gao 0007, Le Zhuo, Ruoyi Du, Longtian Qiu, Rongjie Huang 0001, Shijie Geng, Renrui Zhang, Junlin Xie, Wenqi Shao, Zhengkai Jiang 0001, Tianshuo Yang, Weicai Ye, Tong He 0001, Jingwen He, Junjun He, Yu Qiao 0001, Hongsheng Li 0001 |
ICLR | 18 |
| 2025 | HyperPath: Knowledge-Guided Hyperbolic Semantic Hierarchy Modeling for WSI Analysis
Peixiang Huang, Yanyan Huang, Weiqin Zhao, Junjun He, Lequan Yu |
MICCAI (5) | 4 |
| 2025 | Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model
Wei Li 0320, Guoan Wang, Kaijing Zhou, Junzhi Ning, ZongYuan Ge, Lixu Gu, Junjun He |
MICCAI (9) | 10 |
| 2025 | Multi-modal MRI Translation via Evidential Regression and Distribution Calibration
Jiyao Liu, Shangqi Gao, Zhaohu Xing, Junzhi Ning, Yanzhou Su, Xiao-Yong Zhang, Junjun He, Ningsheng Xu, Xiahai Zhuang |
MICCAI (8) | 10 |
| 2025 | Towards Interpretable Counterfactual Generation via Multimodal Autoregression
Chenglong Ma 0002, Yuanfeng Ji, Jin Ye 0002, Lu Zhang 0060, Tianbin Li, Mingjie Li 0006, Junjun He, Hongming Shan |
MICCAI (2) | 8 |
| 2025 | RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions
Junzhi Ning, Cheng Tang 0003, Kaijing Zhou, Diping Song, Wei Li 0320, Yanzhou Su, Tianbin Li, Jiyao Liu, Jin Ye 0002, Yuanfeng Ji, Junjun He |
MICCAI (16) | 15 |
| 2025 | Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation
Xinkun Wang, Yifang Wang 0013, Senwei Liang, Junjun He, ZongYuan Ge, Muhammad Imran Razzak |
MICCAI (8) | 8 |
| 2025 | MSWAL: 3D Multi-class Segmentation of Whole Abdominal Lesions Dataset
Zhaodong Wu, Qiaochu Zhao, Yulong Li 0002, Haochen Xue, Zhengyong Jiang, Angelos Stefanidis, Muhammad Imran Razzak, ZongYuan Ge, Junjun He, Yu Qiao 0001, Kang Dang, Jionglong Su |
MICCAI (2) | 11 |
| 2025 | MedGround-R1: Advancing Medical Image Grounding via Spatial-Semantic Rewarded Group Relative Policy Optimization
Yuanpeng Nie, Hualiang Wang, Wei Li 0320, Junzhi Ning, Hongqiu Wang, Jiyao Liu, Junjun He |
MICCAI (5) | 12 |
| 2025 | FPN-in-FPN: A Nested Multi-scale Aggregation Network for Polyp Segmentation
Jin Ye 0002, Yanzhou Su, Yicheng Wu 0001, Junjun He, Bohan Zhuang, Zhaolin Chen, Jianfei Cai 0001 |
MICCAI (11) | 4 |
| 2025 | Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic SurgeryabstractAccurate 3D reconstruction of hands and instruments is critical for vision-based analysis of ophthalmic microsurgery, yet progress has been hampered by the lack of realistic, large-scale datasets and reliable annotation tools. In this work, we introduce OphNet-3D, the first extensive RGB-D dynamic 3D reconstruction dataset for ophthalmic surgery, comprising 41 sequences from 40 surgeons and totaling 7.1 million frames, with fine-grained annotations of 12 surgical phases, 10 instrument categories, dense MANO hand meshes, and full 6-DoF instrument poses. To scalably produce high-fidelity labels, we design a multi-stage automatic annotation pipeline that integrates multi-view data observation, data-driven motion prior with cross-view geometric consistency and biomechanical constraints, along with a combination of collision-aware interaction constraints for instrument interactions. Building upon OphNet-3D, we establish two challenging benchmarks—bimanual hand pose estimation and hand–instrument interaction reconstruction—and propose two dedicated architectures: H-Net for dual-hand mesh recovery and OH-Net for joint reconstruction of two-hand–two-instrument interactions. These models leverage a novel spatial reasoning module with weak-perspective camera modeling and collision-aware center-based representation. Both architectures outperform existing methods by substantial margins, achieving improvements of over 2mm in Mean Per Joint Position Error (MPJPE) and up to 23\% in ADD-S metrics for hand and instrument reconstruction, respectively. Zhengdi Yu, Yulong Li 0002, Muhammad Imran Razzak, Junjun He, Tolga Birdal, Kaijing Zhou, ZongYuan Ge |
NeurIPS | 7 |
| 2025 | AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray InterpretationabstractChest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Medical Large Multimodal Models (MLMMs) have enabled automated CXR interpretation, improving diagnostic accuracy and efficiency. However, despite their strong visual understanding, current MLMMs still face two major challenges: (1) insufficient region-level understanding and interaction, and (2) limited accuracy and interpretability due to single-step prediction. In this paper, we address these challenges by empowering MLMMs with anatomy-centric reasoning capabilities to enhance their interactivity and explainability. Specifically, we propose an Anatomical Ontology-Guided Reasoning (AOR) framework that accommodates both textual and optional visual prompts, centered on region-level information to enable multimodal multi-step reasoning. We also develop AOR-Instruction, a large instruction dataset for MLMMs training, under the guidance of expert physicians. Our experiments demonstrate AOR's superior performance in both Visual Question Answering (VQA) and report generation tasks. Code and data are available at: https://github.com/Liqq1/AOR. Qingqiu Li, Zihang Cui, Seongsu Bae, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Quanli Shen, Shang Gao 0003, Junjun He |
NeurIPS | 11 |
| 2025 | Reliable Lifelong Multimodal Editing: Conflict-Aware Retrieval Meets Multi-Level GuidanceabstractThe dynamic nature of real-world information demands efficient knowledge editing in multimodal large language models (MLLMs) to ensure continuous knowledge updates. However, existing methods often struggle with precise matching in large-scale knowledge retrieval and lack multi-level guidance for coordinated editing, leading to less reliable outcomes. To tackle these challenges, we propose CARML, a novel retrieval-augmented editing framework that integrates conflict-aware dynamic retrieval with multi-level implicit and explicit guidance for reliable lifelong multimodal editing. Specifically, CARML introduces intra-modal uncertainty and inter-modal conflict quantification to dynamically integrate multi-channel retrieval results, so as to pinpoint the most relevant knowledge to the incoming edit samples. Afterwards, an edit scope classifier discerns whether the edit sample semantically aligns with the edit scope of the retrieved knowledge. If deemed in-scope, CARML refines the retrieved knowledge into information-rich continuous prompt prefixes, serving as the implicit knowledge guide. These prefixes not only include static knowledge prompt that capture key textual semantics but also incorporate token-level, context-aware dynamic prompt to explore fine-grained cross-modal associations between the edit sample and retrieved knowledge. To further enhance reliability, CARML incorporates a "hard correction" mechanism, leveraging explicit label knowledge to adjust the model’s output logits. Extensive experiments across multiple MLLMs and datasets indicate the superior performance of CARML in lifelong multimodal editing scenarios. Qiang Zhang 0051, Fanrui Zhang, Jiawei Liu 0001, Junjun He, Zhengjun Zha |
NeurIPS | 5 |
| 2025 | A survey for large language models in biomedicine
Chong Wang 0027, Junjun He, Zhongruo Wang, Erfan Darzi, Jin Ye 0002, Tianbin Li, Yanzhou Su, Jing Ke, Kaili Qu, Pietro Liò, Tianyun Wang, Yu Guang Wang 0001, Yiqing Shen 0003 |
Artif. Intell. Medicine | 3 |
| 2025 | Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai |
Sci. China Inf. Sci. | 33 |
| 2025 | FCN+: Global receptive convolution makes FCN great again
Zhongying Deng, Jin Ye 0002, Junjun He |
Neurocomputing | 4 |
| 2025 | PitVis-2023 challenge: Workflow recognition in videos of endoscopic pituitary surgeryabstractThe field of computer vision applied to videos of minimally invasive surgery is ever-growing. Workflow recognition pertains to the automated recognition of various aspects of a surgery, including: which surgical steps are performed; and which surgical instruments are used. This information can later be used to assist clinicians when learning the surgery or during live surgery. The Pituitary Vision (PitVis) 2023 Challenge tasks the community to step and instrument recognition in videos of endoscopic pituitary surgery. This is a particularly challenging task when compared to other minimally invasive surgeries due to: the smaller working space, which limits and distorts vision; and higher frequency of instrument and step switching, which requires more precise model predictions. Participants were provided with 25-videos, with results presented at the MICCAI-2023 conference as part of the Endoscopic Vision 2023 Challenge in Vancouver, Canada, on 08-Oct-2023. There were 18-submissions from 9-teams across 6-countries, using a variety of deep learning models. The top performing model for step recognition utilised a transformer based architecture, uniquely using an autoregressive decoder with a positional encoding input. The top performing model for instrument recognition utilised a spatial encoder followed by a temporal encoder, which uniquely used a 2-layer temporal architecture. In both cases, these models outperformed purely spatial based models, illustrating the importance of sequential and temporal information. This PitVis-2023 therefore demonstrates state-of-the-art computer vision models in minimally invasive surgery are transferable to a new dataset. Benchmark results are provided in the paper, and the dataset is publicly available at: https://doi.org/10.5522/04/26531686. Adrito Das, Danyal Z. Khan, Dimitris Psychogyios, John G. Hanrahan, Francisco Vasconcelos 0001, You Pang, Zhen Chen 0018, Jinlin Wu, Xiaoyang Zou, Guoyan Zheng, Abdul Qayyum 0002, Moona Mazher, Muhammad Imran Razzak, Tianbin Li, Jin Ye 0002, Junjun He, Szymon Plotka, Joanna Kaleta, Amine Yamlahi, Antoine Jund, Patrick Godau, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Dominik Rivoir, Stefanie Speidel, Alejandra Pérez, Santiago Rodríguez, Pablo Andrés Arbeláez, Danail Stoyanov, Hani J. Marcus, Sophia Bano |
Medical Image Anal. | 17 |
| 2025 | A-Eval: A benchmark for cross-dataset and cross-modality evaluation of abdominal multi-organ segmentation
Ziyan Huang, Zhongying Deng, Jin Ye 0002, Haoyu Wang 0010, Yanzhou Su, Tianbin Li, Junlong Cheng, Jianpin Chen, Junjun He, Yun Gu, Shaoting Zhang 0001, Lixu Gu, Yu Qiao 0001 |
Medical Image Anal. | 10 |
| 2025 | SegRap2023: A benchmark of organs-at-risk and gross tumor volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma
Xiangde Luo, Yunxin Zhong, Shuolin Liu, Mehdi Astaraki, Simone Bendazzoli, Iuliana Toma-Dasu, Yiwen Ye, Ziyang Chen 0003, Yong Xia 0001, Yanzhou Su, Jin Ye 0002, Junjun He, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Kaixiang Yang 0004, Zhiwei Wang 0002, Chan Woong Lee, Sang Joon Park, Jaehee Chun, Constantin Ulrich, Klaus H. Maier-Hein, Nchongmaje Ndipenoch, Alina Dana Miron, Yongmin Li 0001, Chengyang An, Lisheng Wang, Kaiwen Huang 0002, Yunqi Gu, Tao Zhou 0002, Mu Zhou, Shichuan Zhang, Wenjun Liao, Guotai Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 14 |
| 2025 | RFMiD: Retinal Image Analysis for multi-Disease Detection challenge
Samiksha Pachade, Prasanna Porwal, Manesh Kokare, Girish Deshmukh, Vivek Sahasrabuddhe, Zhengbo Luo, Zitang Sun, Li Qihan, Edward Ho, Asaanth Sivajohan, Saerom Youn, Kevin Lane, Jin Chun, Yunchao Gu, Sixu Lu, Young-tack Oh, Hyunjin Park, Chia-Yen Lee, Hung Yeh, Kai-Wen Cheng, Haoyu Wang 0010, Jin Ye 0002, Junjun He, Lixu Gu, Dominik Müller, Iñaki Soto Rey, Frank Kramer 0001, Hidehisa Arai, Yuma Ochi, Takami Okada, Luca Giancardo, Gwenolé Quellec, Fabrice Mériaudeau |
Medical Image Anal. | 27 |
| 2025 | SegAnyPath: A Foundation Model for Multi- Resolution Stain-Variant and Multi-Task Pathology Image SegmentationabstractFoundation models like the Segment Anything Model (SAM) have shown promising performance in general image segmentation tasks. However, their effectiveness is limited when applied to pathology images due to the inherent multi-scale structural complexity and staining heterogeneity. To address these challenges, we introduce SegAnyPath, a foundational model specifically designed for pathology image segmentation. SegAnyPath is trained on an extensive public pathology dataset comprising over 1.5 million images and 3.5 million masks. We propose a multi-scale proxy task to handle the diverse resolutions in pathology images, complementing the reconstruction objective in the supervised learning stage. To enhance segmentation performance across stain variations, we introduce a novel self-distillation scheme based on stain augmentations. Furthermore, we propose an innovative task-guided Mixture of Experts (MoE) architecture in the decoder of SegAnyPath for efficient management of distinct pathology segmentation tasks, including cell, tissue, and tumor segmentation. Experimental results demonstrate SegAnyPath's zero-shot generalization capability, achieving a Dice score of 0.6797 across multiple datasets and organs while maintaining consistent performance across varying staining styles and resolutions. In comparison, the fine-tuned SAM achieves a Dice score of only 0.5258 on the same external test sets, indicating a substantial 29.27% improvement by SegAnyPath. SegAnyPath has the potential to advance the field of pathology analysis and improve diagnostic accuracy in clinical settings. The code is available at https://github.com/wagnchogn/SegAnyPath. Chong Wang 0027, Yajie Wan, Kaili Qu, Xuezhi Zhou, Junjun He, Jing Ke, Tianyun Wang, Yiqing Shen 0003 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Learning With Explicit Shape Priors for Medical Image SegmentationabstractMedical image segmentation is a fundamental task for medical image analysis and surgical planning. In recent years, UNet-based networks have prevailed in the field of medical image segmentation. However, convolutional neural networks (CNNs) suffer from limited receptive fields, which fail to model the long-range dependency of organs or tumors. Besides, these models are heavily dependent on the training of the final segmentation head. And existing methods can not well address aforementioned limitations simultaneously. Hence, in our work, we proposed a novel shape prior module (SPM), which can explicitly introduce shape priors to promote the segmentation performance of UNet-based models. The explicit shape priors consist of global and local shape priors. The former with coarse shape representations provides networks with capabilities to model global contexts. The latter with finer shape information serves as additional guidance to relieve the heavy dependence on the learnable prototype in the segmentation head. To evaluate the effectiveness of SPM, we conduct experiments on three challenging public datasets. And our proposed model achieves state-of-the-art performance. Furthermore, SPM can serve as a plug-and-play structure into classic CNNs and Transformer-based backbones, facilitating the segmentation task on different datasets. Source codes are available at https://github.com/AlexYouXin/Explicit-Shape-Priors. Xin You 0002, Junjun He, Jie Yang 0002, Yun Gu |
IEEE Trans. Medical Imaging | 2 |
| 2025 | SAM-Med3D: A Vision Foundation Model for General-Purpose Segmentation on Volumetric Medical ImagesabstractExisting volumetric medical image segmentation models are typically task-specific, excelling at specific targets but struggling to generalize across anatomical structures or modalities. This limitation restricts their broader clinical use. In this article, we introduce segment anything model (SAM)-Med3D, a vision foundation model (VFM) for general-purpose segmentation on volumetric medical images. Given only a few 3-D prompt points, SAM-Med3D can accurately segment diverse anatomical structures and lesions across various modalities. To achieve this, we gather and preprocess a large-scale 3-D medical image segmentation dataset, SA-Med3D-140K, from 70 public datasets and 8K licensed private cases from hospitals. This dataset includes 22K 3-D images and 143K corresponding masks. SAM-Med3D, a promptable segmentation model characterized by its fully learnable 3-D structure, is trained on this dataset using a two-stage procedure and exhibits impressive performance on both seen and unseen segmentation targets. We comprehensively evaluate SAM-Med3D on 16 datasets covering diverse medical scenarios, including different anatomical structures, modalities, targets, and zero-shot transferability to new/unseen tasks. The evaluation demonstrates the efficiency and efficacy of SAM-Med3D, as well as its promising application to diverse downstream tasks as a pretrained model. Our approach illustrates that substantial medical resources can be harnessed to develop a general-purpose medical AI for various potential applications. Our dataset, code, and models are available at: https://github.com/uni-medical/SAM-Med3D. Haoyu Wang 0010, Sizheng Guo, Jin Ye 0002, Zhongying Deng, Junlong Cheng, Tianbin Li, Jianpin Chen, Yanzhou Su, Ziyan Huang, Yiqing Shen 0003, Shaoting Zhang 0001, Junjun He |
IEEE Trans. Neural Networks Learn. Syst. | 13 |
| 2025 | Toward the unification of generative and discriminative visual foundation model: a survey
Chong Wang 0027, Yuanxin Wang 0001, Qinjingwen Cao, Weizhi Du, Yonghuan Yang, Junjun He, Yu Qiao 0001, Yiqing Shen 0003 |
Vis. Comput. | 9 |
| 2024 | Histology Image Artifact Restoration with Lightweight Transformer Based Diffusion Model
Chong Wang 0027, Zhenqi He, Junjun He, Jin Ye 0002, Yiqing Shen 0003 |
AIME (2) | 3 |
| 2024 | A Fine-tuning Dataset and Benchmark for Large Language Models for Protein UnderstandingabstractThe high similarities between protein sequences and natural language, particularly in their sequential data structures, have driven parallel advancements in deep learning models for both domains. In natural language processing (NLP), large language models (LLMs) have achieved remarkable success in tasks such as text generation, translation, and conversational agents, owing to their extensive training on diverse datasets that enable them to capture complex language patterns and generate human-like text. Inspired by these advancements, researchers have attempted to adapt LLMs for protein understanding by integrating a protein sequence encoder with a pre-trained LLM, following designs like LLaVa. However, this adaptation raises a fundamental question: "Can LLMs, originally designed for NLP, effectively comprehend protein sequences as a form of language?" Current datasets fall short in addressing this question due to the lack of a direct correlation between protein sequences and corresponding text descriptions, limiting the ability to train and evaluate LLMs for protein understanding effectively. To bridge this gap, we introduce ProteinLMDataset, a dataset specifically designed for further self-supervised pretraining and supervised fine-tuning (SFT) of LLMs to enhance their capability for protein sequence comprehension. Specifically, ProteinLMDataset includes 17.46 billion tokens for pretraining and 893K instructions for SFT. Additionally, we present ProteinLMBench, the first benchmark dataset consisting of 944 manually verified multiple-choice questions for assessing the protein understanding capabilities of LLMs. ProteinLMBench incorporates protein-related details and sequences in multiple languages, establishing a new standard for evaluating LLMs’ abilities in protein comprehension. The large language model InternLM2-7B, pretrained and fine-tuned on the ProteinLMDataset, outperforms GPT-4 on ProteinLMBench, achieving the highest accuracy score. The dataset and the benchmark are available at https://huggingface. co/datasets/tsynbio/ProteinLMDataset/ and https://huggingface.co/datasets/tsynbio/ProteinLMBench. The code is available at https://github.com/tsynbio/ProteinLMDataset/. Yiqing Shen 0003, Michail Mamalakis, Luhan He, Tianbin Li, Yanzhou Su, Junjun He, Yu Guang Wang 0001 |
BIBM | 8 |
| 2024 | TourSynbio: A Multi-Modal Large Model and Agent Framework to Bridge Text and Protein Sequences for Protein EngineeringabstractThe structural similarities between protein sequences and natural languages have led to parallel advancements in deep learning across both domains. While large language models (LLMs) have achieved much progress in the domain of natural language processing, their potential in protein engineering remains largely unexplored. Previous approaches have equipped LLMs with protein understanding capabilities by incorporating external protein encoders, but this fails to fully leverage the inherent similarities between protein sequences and natural languages, resulting in sub-optimal performance and increased model complexity. To address this gap, we present TourSynbio-7B, the first multi-modal large model specifically designed for protein engineering tasks without external protein encoders. TourSynbio-7B demonstrates that LLMs can inherently learn to understand proteins as language. The model is post-trained and instruction fine-tuned on InternLM2-7B using ProteinLM-Dataset, a dataset comprising 17.46 billion tokens of text and protein sequence for self-supervised pretraining and 893K instructions for supervised fine-tuning. TourSynbio7B outperforms GPT-4 on the ProteinLMBench, a benchmark of 944 manually verified multiple-choice questions, with 62.18% accuracy. Leveraging TourSynbio-7B’s enhanced protein sequence understanding capability, we introduce TourSynbioAgent, an innovative framework capable of performing various protein engineering tasks, including mutation analysis, inverse folding, protein folding, and visualization. TourSynbio-Agent integrates previously disconnected deep learning models in the protein engineering domain, offering a unified conversational user interface for improved usability. Finally, we demonstrate the efficacy of TourSynbio-7B and TourSynbio-Agent through two wet lab case studies on vanilla key enzyme modification and steroid compound catalysis. Our results show that this combination facilitates protein engineering tasks in wet labs, leading to higher positive rates, improved mutations, shorter delivery times, and increased automation. The model weights are available at https://huggingface.co/tsynbio/Toursynbio and codes at https://github.com/tsynbio/TourSynbio. Yiqing Shen 0003, Michail Mamalakis, Yungeng Liu, Tianbin Li, Yanzhou Su, Junjun He, Pietro Liò, Yu Guang Wang 0001 |
BIBM | 7 |
| 2024 | A Tiny Efficient U-Net with Gated Linear Attention for Medical Image SegmentationabstractMedical image segmentation is crucial for diagnosis and treatment planning. While recent advancements in deep learning, particularly UNet variants, have improved segmentation performance, they often result in increased model complexity, which limits their real-time applicability on resource-constrained devices in clinical settings. To address this challenge, we present the Tiny Efficient U-Net (TE-UNet), a novel lightweight model balancing efficiency and accuracy. TE-UNet uses a U-shaped encoder-decoder framework with a gated linear attention mechanism to process low-level and high-level features, preserving details and reducing complexity. It employs depth-wise separable convolutions for higher-level processing, enhancing efficiency without losing performance. Additionally, skip connections improve multi-scale feature extraction and information flow. Experiments on two public datasets across different modalities demonstrate that TE-UNet outperforms ten state-of-the-art methods, maintaining a parameter size under 40KB and low computational cost. TE-UNet makes real-time segmentation more accessible for various clinical applications. Sibo Ju, Zhaozhen Chen, Xiangwen Liao, Yiqing Shen 0003, Junjun He, Yanzhou Su |
BIBM | 5 |
| 2024 | Generate Like Experts: Multi-Stage Font Generation by Incorporating Font Transfer Process into Diffusion ModelsabstractFew-shot font generation (FFG) produces stylized font images with a limited number of reference samples, which can significantly reduce labor costs in manual font designs. Most existing FFG methods follow the style-content dis-entanglement paradigm and employ the Generative Adver-sarial Network (GAN) to generate target fonts by combining the decoupled content and style representations. The complicated structure and detailed style are simultaneously generated in those methods, which may be the sub-optimal solutions for FFG task. Inspired by most manual font design processes of expert designers, in this paper, we model font generation as a multi-stage generative process. Specifically, as the injected noise and the data distribution in diffusion models can be well-separated into different sub-spaces, we are able to incorporate the font transfer process into these models. Based on this observation, we generalize diffusion methods to modelfont generative process by separating the reverse diffusion process into three stages with different functions: The structure construction stage first generates the structure information for the target character based on the source image, and the font transfer stage subsequently transforms the source font to the target font. Finally, the font refinement stage enhances the appearances and local details of the target font images. Based on the above multi-stage generative process, we construct our font generation framework. named MSD-Font, with a dual-network approach to generate font images. The superior performance demonstrates the effectiveness of our model. The code is available at: https://github.com/fubinfbIMSD-Font. Fanghua Yu, Jie Wen 0001, Junjun He, Yu Qiao 0001 |
CVPR | 6 |
| 2024 | OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLMabstractLarge Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in various multimodal tasks. However, their potential in the medical domain re-mains largely unexplored. A significant challenge arises from the scarcity of diverse medical images spanning various modalities and anatomical regions, which is essential in real-world medical applications. To solve this problem, in this paper, we introduce OmniMedVQA, a novel comprehensive medical Visual Question Answering (VQA) benchmark. This benchmark is collected from 73 different medical datasets, including 12 different modalities and covering more than 20 distinct anatomical regions. Importantly, all images in this benchmark are sourced from authentic medical scenarios, ensuring alignment with the requirements of the medical field and suitability for evaluating LVLMs. Through our extensive experiments, we have found that existing LVLMs struggle to address these medical VQA problems effectively. Moreover, what surprises us is that medical-specialized LVLMs even exhibit inferior performance to those general-domain models, calling for a more versatile and robust LVLM in the biomedical field. The evaluation results not only reveal the current limitations of LVLM in understanding real medical images but also highlight our dataset's significance. Our code with dataset are available at https://github.com/OpenGVLab/ Multi Modality-Arena. Yutao Hu 0002, Tianbin Li, Quanfeng Lu, Wenqi Shao, Junjun He, Yu Qiao 0001, Ping Luo 0002 |
CVPR | 5 |
| 2024 | SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language ModelsabstractWe propose SPHINX-X, an extensive Multi-modality Large Language Model (MLLM) series developed upon SPHINX. To improve the architecture and training efficiency, we modify the SPHINX framework by removing redundant visual encoders, bypassing fully-padded sub-images with skip tokens, and simplifying multi-stage training into a one-stage all-in-one paradigm. To fully unleash the potential of MLLMs, we assemble a comprehensive multi-domain and multi-modal dataset covering publicly available resources in language, vision, and vision-language tasks. We further enrich this collection with our curated OCR intensive and Set-of-Mark datasets, extending the diversity and generality. By training over different base LLMs including TinyLlama-1.1B, InternLM2-7B, LLaMA2-13B, and Mixtral-8$\times$7B, we obtain a spectrum of MLLMs that vary in parameter size and multilingual capabilities. Comprehensive benchmarking reveals a strong correlation between the multi-modal performance with the data and parameter scales. Code and models are released at https://github.com/Alpha-VLLM/LLaMA2-Accessory. Renrui Zhang, Longtian Qiu, Siyuan Huang 0004, Weifeng Lin, Shitian Zhao, Shijie Geng, Kaipeng Zhang, Wenqi Shao, Conghui He, Junjun He, Hao Shao, Pan Lu, Yu Qiao 0001, Hongsheng Li 0001, Peng Gao 0007 |
ICML | 14 |
| 2024 | SAM-Med3D-MoE: Towards a Non-Forgetting Segment Anything Model via Mixture of Experts for 3D Medical Image Segmentation
Guoan Wang, Jin Ye 0002, Junlong Cheng, Tianbin Li, Zhaolin Chen, Jianfei Cai 0001, Junjun He, Bohan Zhuang |
MICCAI (9) | 7 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 12 |
| 2024 | GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AIabstractLarge Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial assistance for diagnosis and treatment. Before that, it is crucial to develop benchmarks to evaluate LVLMs' effectiveness in various medical applications. Current benchmarks are often built upon specific academic literature, mainly focusing on a single domain, and lacking varying perceptual granularities. Thus, they face specific challenges, including limited clinical relevance, incomplete evaluations, and insufficient guidance for interactive LVLMs. To address these limitations, we developed the GMAI-MMBench, the most comprehensive general medical AI benchmark with well-categorized data structure and multi-perceptual granularity to date. It is constructed from 284 datasets across 38 medical image modalities, 18 clinical-related tasks, 18 departments, and 4 perceptual granularities in a Visual Question Answering (VQA) format. Additionally, we implemented a lexical tree structure that allows users to customize evaluation tasks, accommodating various assessment needs and substantially supporting medical AI research and applications. We evaluated 50 LVLMs, and the results show that even the advanced GPT-4o only achieves an accuracy of 53.96\%, indicating significant room for improvement. Moreover, we identified five key insufficiencies in current cutting-edge LVLMs that need to be addressed to advance the development of better medical applications. We believe that GMAI-MMBench will stimulate the community to build the next generation of LVLMs toward GMAI. Jin Ye 0002, Guoan Wang, Yanjun Li 0007, Zhongying Deng, Wei Li 0320, Tianbin Li, Haodong Duan, Ziyan Huang, Yanzhou Su, Benyou Wang, Shaoting Zhang 0001, Jianfei Cai 0001, Bohan Zhuang, Eric J. Seibel, Junjun He, Yu Qiao 0001 |
NeurIPS | 17 |
| 2023 | Neural Transformation Fields for Arbitrary-Styled Font GenerationabstractFew-shot font generation (FFG), aiming at generating font images with a few samples, is an emerging topic in recent years due to the academic and commercial values. Typically, the FFG approaches follow the style-content disentanglement paradigm, which transfers the target font styles to characters by combining the content representations of source characters and the style codes of reference samples. Most existing methods attempt to increase font generation ability via exploring powerful style representations, which may be a sub-optimal solution for the FFG task due to the lack of modeling spatial transformation in transferring font styles. In this paper, we model font generation as a continuous transformation process from the source character image to the target font image via the creation and dissipation of font pixels, and embed the corresponding transformations into a neural transformation field. With the estimated transformation path, the neural transformation field generates a set of intermediate transformation results via the sampling process, and a font rendering formula is developed to accumulate them into the target font image. Extensive experiments show that our method achieves state-of-the-art performance on few-shot font generation task, which demonstrates the effectiveness of our proposed model. Our implementation is available at: https://github.com/fubinfb/NTF. Junjun He, Yu Qiao 0001 |
CVPR | 2 |
| 2023 | Vision Transformer Adapter for Dense Predictions
Zhe Chen 0017, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu 0002, Jifeng Dai, Yu Qiao 0001 |
ICLR | 4 |
| 2023 | Artifact Restoration in Histology Images with Diffusion Probabilistic Models
Zhenqi He, Junjun He, Jin Ye 0002, Yiqing Shen 0003 |
MICCAI (6) | 2 |
| 2023 | Revisiting Feature Propagation and Aggregation in Polyp Segmentation
Yanzhou Su, Yiqing Shen 0003, Jin Ye 0002, Junjun He, Jian Cheng 0003 |
MICCAI (5) | 4 |
| 2023 | Pick the Best Pre-trained Model: Towards Transferability Estimation for Medical Image Segmentation
Yuncheng Yang, Junjun He, Jin Ye 0002, Yun Gu |
MICCAI (1) | 3 |
| 2023 | Accurate polyp segmentation through enhancing feature fusion and boosting boundary performance
Yanzhou Su, Jian Cheng 0003, Chuqiao Zhong, Chengzhi Jiang, Jin Ye 0002, Junjun He |
Neurocomputing | 6 |
| 2023 | A Holistically-Guided Decoder for Deep Representation Learning With Applications to Semantic Segmentation and Object DetectionabstractBoth high-level and high-resolution feature representations are of great importance in various visual understanding tasks. To acquire high-resolution feature maps with high-level semantic information, one common strategy is to adopt dilated convolutions in the backbone networks to extract high-resolution feature maps, such as the dilatedFCN-based methods for semantic segmentation. However, due to many convolution operations are conducted on the high-resolution feature maps, such methods have large computational complexity and memory consumption. To balance the performance and efficiency, there also exist encoder-decoder structures that gradually recover the spatial information by combining multi-level feature maps from a feature encoder, such as the FPN architecture for object detection and the U-Net for semantic segmentation. Although being more efficient, the performances of existing encoder-decoder methods for semantic segmentation are far from comparable with the dilatedFCN-based methods. In this paper, we propose one novel holistically-guided decoder which is introduced to obtain the high-resolution semantic-rich feature maps via the multi-scale features from the encoder. The decoding is achieved via novel holistic codeword generation and codeword assembly operations, which take advantages of both the high-level and low-level features from the encoder features. With the proposed holistically-guided decoder, we implement the EfficientFCN architecture for semantic segmentation and HGD-FPN for object detection and instance segmentation. The EfficientFCN achieves comparable or even better performance than state-of-the-art methods with only 1/3 of their computational costs for semantic segmentation on PASCAL Context, PASCAL VOC, ADE20K datasets. Meanwhile, the proposed HGD-FPN achieves higher mean Average Precision (mAP) when integrated into several object detection frameworks with ResNet-50 encoding backbones. Junjun He, Yuanjie Zheng, Shuai Yi, Xiaogang Wang 0001, Hongsheng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | StructToken: Rethinking Semantic Segmentation With Structural PriorabstractIn previous deep-learning-based methods, semantic segmentation has been regarded as a static or dynamic per-pixel classification task, i.e., classify each pixel representation to a specific category. However, these methods only focus on learning better pixel representations or classification kernels while ignoring the structural information of objects, which is critical to human decision-making mechanism. In this paper, we present a new paradigm for semantic segmentation, named structure-aware extraction. Specifically, it generates the segmentation results via the interactions between a set of learned structure tokens and the image feature, which aims to progressively extract the structural information of each category from the feature. Extensive experiments show that our StructToken outperforms the state-of-the-art on three widely-used benchmarks, including ADE20K, Cityscapes, and COCO-Stuff-10K. Fangjian Lin, Sitong Wu, Junjun He, Shengwei Tian |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Dual Relation Network for Scene Text RecognitionabstractLocal visual and long-range contextual features yield two complementary cues for human reading text in natural scene. Existing scene text recognition methods mainly extract local features at a low level and then model long-range dependencies at a high level, this sequential pipeline may be sub-optimal to construct complete and effective representation. Except for high-level features, long-range contextual relation is of importance in low-level features as well since it can help separate different characters based on the intervals between characters and thus enhance the character features. To address this issue, we develop a dual relation module to extract complementary features in a parallel manner for scene text recognition, which consists of a local visual branch and a long-range contextual branch. The local visual branch employs a topological-aware operation to model intra-character characteristic and extract discriminative features of different characters. Meanwhile, the long-range contextual branch utilizes a simple but effective strategy to incorporate inter-character relations into feature maps. Our dual relation module is a plug-and-play block which can be easily incorporated into modern deep architectures. Experimental results demonstrate that our methods achieved top performance on several standard benchmarks. Code and models will become publicly available in the future. Ming Li 0010, Junjun He, Yu Qiao 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Region-Aware Arbitrary-Shaped Text Detection With Progressive FusionabstractSegmentation-based text detectors are flexible to capture arbitrary-shaped text regions. Due to large geometry variance, it is necessary to construct effective and robust representations to identify text regions with various shapes and scales. In this paper, we focus on designing effective multi-scale contextual features for locating text instances. Specially, we develop a Region Context Module (RCM) to summarize the semantic response and adaptively extract text-region-aware information in a limited local area. To construct complementary multi-scale contextual representations, multiple RCM branches with different scales are employed and integrated via Progressive Fusion Module (PFM). Our proposed RCM and PFM serve as the plug-and-play modules which can be incorporated into existing scene text detection platforms to further boost detection performance. Extensive experiments show that our methods achieve state-of-the-art performances on Total-Text, SCUT-CTW1500 and MSRA-TD500 datasets. The code with models will become publicly available athttps://github.com/wqtwjt1996/RP-Text. Qitong Wang 0001, Ming Li 0010, Junjun He, Xi Peng 0005, Yu Qiao 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Dynamic Instance Domain AdaptationabstractMost existing studies on unsupervised domain adaptation (UDA) assume that each domain's training samples come with domain labels (e.g., painting, photo). Samples from each domain are assumed to follow the same distribution and the domain labels are exploited to learn domain-invariant features via feature alignment. However, such an assumption often does not hold true-there often exist numerous finer-grained domains (e.g., dozens of modern painting styles have been developed, each differing dramatically from those of the classic styles). Therefore, forcing feature distribution alignment across each artificially-defined and coarse-grained domain can be ineffective. In this paper, we address both single-source and multi-source UDA from a completely different perspective, which is to view each instance as a fine domain. Feature alignment across domains is thus redundant. Instead, we propose to perform dynamic instance domain adaptation (DIDA). Concretely, a dynamic neural network with adaptive convolutional kernels is developed to generate instance-adaptive residuals to adapt domain-agnostic deep features to each individual instance. This enables a shared classifier to be applied to both source and target domain data without relying on any domain annotation. Further, instead of imposing intricate feature alignment losses, we adopt a simple semi-supervised learning paradigm using only a cross-entropy loss for both labeled source and pseudo labeled target data. Our model, dubbed DIDA-Net, achieves state-of-the-art performance on several commonly used single-source and multi-source UDA datasets including Digits, Office-Home, DomainNet, Digit-Five, and PACS. Zhongying Deng, Kaiyang Zhou, Da Li 0001, Junjun He, Yi-Zhe Song, Tao Xiang 0002 |
IEEE Trans. Image Process. | 4 |
| 2021 | A Novel Hybrid Convolutional Neural Network for Accurate Organ Segmentation in 3D Head and Neck CT Images
Cheng Li 0008, Junjun He, Jin Ye 0002, Diping Song, Shanshan Wang 0002, Lixu Gu, Yu Qiao 0001 |
MICCAI (1) | 3 |
| 2021 | Group Shift Pointwise Convolution for Volumetric Medical Image Segmentation
Junjun He, Jin Ye 0002, Cheng Li 0008, Diping Song, Shanshan Wang 0002, Lixu Gu, Yu Qiao 0001 |
MICCAI (3) | 1 |
| 2021 | Deep Relation Transformer for Diagnosing Glaucoma With Optical Coherence Tomography and Visual Field FunctionabstractGlaucoma is the leading reason for irreversible blindness. Early detection and timely treatment of glaucoma are essential for preventing visual field loss or even blindness. In clinical practice, Optical Coherence Tomography (OCT) and Visual Field (VF) exams are two widely-used and complementary techniques for diagnosing glaucoma. OCT provides quantitative measurements of the optic nerve head (ONH) structure, while VF test is the functional assessment of peripheral vision. In this paper, we propose a Deep Relation Transformer (DRT) to perform glaucoma diagnosis with OCT and VF information combined. A novel deep reasoning mechanism is proposed to explore implicit pairwise relations between OCT and VF information in global and regional manners. With the pairwise relations, a carefully-designed deep transformer mechanism is developed to enhance the representation with complementary information for each modal. Based on reasoning and transformer mechanisms, three successive modules are designed to extract and collect valuable information for glaucoma diagnosis, the global relation module, the guided regional relation module, and the interaction transformer module, namely. Moreover, we build a large dataset, namely ZOC-OCT&VF dataset, which includes 1395 OCT-VF pairs for developing and evaluating our DRT. We conduct extensive experiments to validate the effectiveness of the proposed method. Experimental results show that our method achieves 88.3% accuracy and outperforms the existing single-modal approaches with a large margin. The codes and dataset will be publicly available in the future. Diping Song, Junjun He, Xiulan Zhang, Yu Qiao 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Dynamic Sampling Network for Semantic SegmentationabstractSampling is a basic operation of modern convolutional neural networks (CNN) since down-sampling operators are employed to enlarge the receptive field while up-sampling operators are adopted to increase resolution. Most existing deep segmentation networks employ regular grid sampling operators, which can be suboptimal for semantic segmentation task due to large shape and scale variance. To address this problem, this paper proposes a Context Guided Dynamic Sampling (CGDS) module to obtain an effective representation with rich shape and scale information by adaptively sampling useful segmentation information in spatial space. Moreover, we utilize the multi-scale contextual representations to guide the sampling process. Therefore, our CGDS can adaptively capture shape and scale information according to not only the input feature map but also the multi-scale semantic context. CGDS provides a plug-and-play module which can be easily incorporated in deep segmentation networks. We incorporate our proposed CGDS module into Dynamic Sampling Network (DSNet) and perform extensive experiments on segmentation datasets. Experimental results show that our CGDS significantly improves semantic segmentation performance and achieves state-of-the-art performance on PASCAL VOC 2012 and ADE20K datasets. Our model achieves 85.2% mIOU on PASCAL VOC 2012 test set without MS COCO dataset pre-trained and 46.4% on ADE20K validation set. The codes will become publicly available after publication. Junjun He, Zhengfu Zhang, Yu Qiao 0001 |
AAAI | 2 |
| 2020 | Tensor Low-Rank Reconstruction for Semantic Segmentation
Xinge Zhu, Ruoqi Sun, Junjun He, Ruiyu Li, Xiaoyong Shen, Bei Yu 0001 |
ECCV (17) | 4 |
| 2020 | Learning to Predict Context-Adaptive Convolution for Semantic Segmentation
Junjun He, Yu Qiao 0001, Jimmy S. J. Ren, Hongsheng Li 0001 |
ECCV (25) | 2 |
| 2020 | EfficientFCN: Holistically-Guided Decoding for Semantic Segmentation
Junjun He, Jiawei Zhang 0002, Jimmy S. J. Ren, Hongsheng Li 0001 |
ECCV (26) | 2 |
| 2020 | Attention-Driven Dynamic Graph Convolutional Network for Multi-label Image Recognition
Jin Ye 0002, Junjun He, Xiaojiang Peng, Yu Qiao 0001 |
ECCV (21) | 2 |
| 2020 | MIA-Prognosis: A Deep Learning Framework to Predict Therapy Response
Jiancheng Yang, Kaiming Kuang, Tiancheng Lin 0001, Junjun He, Bingbing Ni |
MICCAI (2) | 5 |
| 2019 | Adaptive Pyramid Context Network for Semantic SegmentationabstractRecent studies witnessed that context features can significantly improve the performance of deep semantic segmentation networks. Current context based segmentation methods differ with each other in how to construct context features and perform differently in practice. This paper firstly introduces three desirable properties of context features in segmentation task. Specially, we find that Global-guided Local Affinity (GLA) can play a vital role in constructing effective context features, while this property has been largely ignored in previous works. Based on this analysis, this paper proposes Adaptive Pyramid Context Network (APCNet) for semantic segmentation. APCNet adaptively constructs multi-scale contextual representations with multiple well-designed Adaptive Context Modules (ACMs). Specifically, each ACM leverages a global image representation as a guidance to estimate the local affinity coefficients for each sub-region, and then calculates a context vector with these affinities. We empirically evaluate our APCNet on three semantic segmentation and scene parsing datasets, including PASCAL VOC 2012, Pascal-Context, and ADE20K dataset. Experimental results show that APCNet achieves state-of-the-art performance on all three benchmarks, and obtains a new record 84.2% on PASCAL VOC 2012 test set without MS COCO pre-trained and any post-processing. Junjun He, Zhongying Deng, Lei Zhou 0003, Yali Wang 0001, Yu Qiao 0001 |
CVPR | 1 |
| 2019 | Dynamic Multi-Scale Filters for Semantic SegmentationabstractMulti-scale representation provides an effective way to address scale variation of objects and stuff in semantic segmentation. Previous works construct multi-scale representation by utilizing different filter sizes, expanding filter sizes with dilated filters or pooling grids, and the parameters of these filters are fixed after training. These methods often suffer from heavy computational cost or have more parameters, and are not adaptive to the input image during inference. To address these problems, this paper proposes a Dynamic Multi-scale Network (DMNet) to adaptively capture multi-scale contents for predicting pixel-level semantic labels. DMNet is composed of multiple Dynamic Convolutional Modules (DCMs) arranged in parallel, each of which exploits context-aware filters to estimate semantic representation for a specific scale. The outputs of multiple DCMs are further integrated for final segmentation. We conduct extensive experiments to evaluate our DMNet on three challenging semantic segmentation and scene parsing datasets, PASCAL VOC 2012, Pascal-Context, and ADE20K. DMNet achieves a new record 84.4% mIoU on PASCAL VOC 2012 test set without MS COCO pre-trained and post-processing, and also obtains state-of-the-art performance on Pascal-Context and ADE20K. Junjun He, Zhongying Deng, Yu Qiao 0001 |
ICCV | 1 |
| 2019 | Prostate Segmentation using 2D Bridged U-netabstractIn this paper, we focus on three problems in deep learning based medical image segmentation. Firstly, U-net, as a popular model for medical image segmentation, is difficult to train when convolutional layers increase even though a deeper network usually has a better generalization ability because of more learnable parameters. Secondly, the exponential ReLU (ELU), as an alternative of ReLU, is not much different from ReLU when the network of interest gets deep. Thirdly, the Dice loss, as one of the pervasive loss functions for medical image segmentation, is not effective when the prediction is close to ground truth and will cause oscillation during training. To address the aforementioned three problems, we propose and validate a deeper network that can fit medical image datasets that are usually small in the sample size. Meanwhile, we propose a new loss function to accelerate the learning process and a combination of different activation functions to improve the network performance. Our experimental results suggest that our network is comparable or superior to state-of-the-art methods. Yue Zhang 0033, Junjun He, Yu Qiao 0001, Yifan Chen 0001, Hongjian Shi, Ed X. Wu, Xiaoying Tang 0001 |
IJCNN | 3 |
| 2019 | Dual-supervised attention network for deep cross-modal hashing
Hanyu Peng, Junjun He, Shifeng Chen, Yali Wang 0001, Yu Qiao 0001 |
Pattern Recognit. Lett. | 2 |
| 2018 | StripNet: Towards Topology Consistent Strip Structure SegmentationabstractIn this work, we propose to study a special semantic segmentation problem where the targets are long and continuous strip patterns. Strip patterns widely exist in medical images and natural photos, such as retinal layers in OCT images and lanes on the roads, and segmentation of them has practical significance. Traditional pixel-level segmentation methods largely ignore the structure prior of strip patterns and thus easily suffer from the topological inconformity problem, such as holes and isolated islands in segmentation results. To tackle this problem, we design a novel deep framework, StripNet, that leverages the strong end-to-end learning ability of CNNs to predict the structured outputs as a sequence of boundary locations of the target strips. Specifically, StripNet decomposes the original segmentation problem into more easily solved local boundary-regression problems, and takes account of the topological constraints on the predicted boundaries. Moreover, our framework adopts a coarse-to-fine strategy and uses carefully designed heatmaps for training the boundary localization network. We examine StripNet on two challenging strip pattern segmentation tasks, retinal layer segmentation and lane detection. Extensive experiments demonstrate that StripNet achieves excellent results and outperforms state-of-the-art methods in both tasks. Guoxiang Qu, Zhe Wang 0006, Xing Dai, Jianping Shi, Junjun He, Xiulan Zhang, Yu Qiao 0001 |
ACM Multimedia | 6 |