EDBT 2026 Demo / reviewers in the wild / expert
Yuanfeng Ji
dblp:227/4488
· DBLP profile ↗
18ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0001-6799-5493ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AIabstractDespite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annotations into high-quality image-text pairs. This dataset offers comprehensive task coverage, diverse modalities, and rich image-text data. Building upon this dataset, we develop GMAI-VL, a 7B-parameter general medical vision-language model, with a three-stage training strategy that enhances the integration of visual and textual information. This approach significantly improves the model's ability to process multimodal data, supporting accurate diagnoses and clinical decision-making. Experiments show that GMAI-VL achieves state-of-the-art performance across various multimodal medical tasks, including visual question answering and medical image diagnosis. Tianbin Li, Yanzhou Su, Wei Li 0320, Zhe Chen 0017, Ziyan Huang, Guoan Wang, Chenglong Ma 0002, Yanjun Li 0007, Shixiang Tang, Xiaowei Hu 0001, Zhongying Deng, Yuanfeng Ji, Jin Ye 0002, Yu Qiao 0001, Junjun He |
AAAI | 16 |
| 2025 | SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image UnderstandingabstractDespite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essential contextual information at the whole-slide level. The lack of large-scale instruction datasets and the gigapixel scale of whole slide images (WSIs) pose significant developmental challenges. In this paper, we present SlideChat, the first vision-language assistant capable of understanding gigapixel whole-slide images, exhibiting excellent multimodal conversational capability and response complex instruction across diverse pathology scenarios. To support its development, we created SlideInstruction, the largest instruction-following dataset for WSIs consisting of 4.2K WSI captions and 176K VQA pairs with multiple categories. Furthermore, we propose SlideBench, a multimodal benchmark that incorporates captioning and VQA tasks to assess SlideChat’s capabilities in various settings such as microscopy, diagnosis and clinical. Compared to both general and specialized MLLMs, SlideChat exhibits exceptional capabilities, achieving state-of-the-art performance on 18 of 22 tasks. For example, it achieved an overall accuracy of 81.17% on SlideBench-VQA (TCGA), and 54.15% on SlideBench-VQA (BCNB). Our code, data, and model is publicly accessible at https://uni-medical.github.io/SlideChat.github.io. Guoan Wang, Yuanfeng Ji, Yanjun Li 0007, Jin Ye 0002, Tianbin Li, Rongshan Yu, Yu Qiao 0001, Junjun He |
CVPR | 3 |
| 2025 | CompGS: Unleashing 2D Compositionality for Compositional Text-to-3D via Dynamically Optimizing 3D GaussiansabstractRecent breakthroughs in text-guided image generation have significantly advanced the field of 3D generation. While generating a single high-quality 3D object is now feasible, generating multiple objects with reasonable interactions within a 3D space, a.k.a. compositional 3D generation, presents substantial challenges. This paper introduces COMPGS, a novel generative framework that employs 3D Gaussian Splatting (GS) for efficient, compositional text-to-3D content generation. To achieve this goal, two core designs are proposed: (1) 3D Gaussians Initialization with 2D compositionality: We transfer the well-established 2D compositionality to initialize the Gaussian parameters on an entity-by-entity basis, ensuring both consistent 3D priors for each entity and reasonable interactions among multiple entities; (2) Dynamic Optimization: We propose a dynamic strategy to optimize 3D Gaussians using Score Distillation Sampling (SDS) loss. COMPGS first automatically decomposes 3D Gaussians into distinct entity parts, enabling optimization at both the entity and composition levels. Additionally, COMPGS optimizes across objects of varying scales by dynamically adjusting the spatial parameters of each entity, enhancing the generation of fine-grained details, particularly in smaller entities. Qualitative comparisons and quantitative evaluations on T3Bench demonstrate the effectiveness of COMPGS in generating compositional 3D objects with superior image quality and semantic alignment over existing methods. COMPGS can also be easily extended to progressively adding 3D objects, facilitating complex scene generation. We hope COMPGS will provide new insights to the compositional 3D generation. Chongjian Ge, Chenfeng Xu, Yuanfeng Ji, Chensheng Peng, Masayoshi Tomizuka, Ping Luo 0002, Mingyu Ding, Varun Jampani |
CVPR | 3 |
| 2025 | Towards Interpretable Counterfactual Generation via Multimodal Autoregression
Chenglong Ma 0002, Yuanfeng Ji, Jin Ye 0002, Lu Zhang 0060, Tianbin Li, Mingjie Li 0006, Junjun He, Hongming Shan |
MICCAI (2) | 2 |
| 2025 | RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions
Junzhi Ning, Cheng Tang 0003, Kaijing Zhou, Diping Song, Wei Li 0320, Yanzhou Su, Tianbin Li, Jiyao Liu, Jin Ye 0002, Yuanfeng Ji, Junjun He |
MICCAI (16) | 14 |
| 2024 | Large Language Models as Automated Aligners for benchmarking Vision-Language ModelsabstractWith the advancements in Large Language Models (LLMs), Vision-Language Models (VLMs) have reached a new level of sophistication, showing notable competence in executing intricate cognition and reasoning tasks. However, existing evaluation benchmarks, primarily relying on rigid, hand-crafted datasets to measure task-specific performance, face significant limitations in assessing the alignment of these increasingly anthropomorphic models with human intelligence. In this work, we address the limitations via Auto-Bench, which delves into exploring LLMs as proficient aligners, measuring the alignment between VLMs and human intelligence and value through automatic data curation and assessment. Specifically, for data curation, Auto-Bench utilizes LLMs (e.g., GPT-4) to automatically generate a vast set of question-answer-reasoning triplets via prompting on visual symbolic representations (e.g., captions, object locations, instance relationships, and etc. The curated data closely matches human intent, owing to the extensive world knowledge embedded in LLMs. Through this pipeline, a total of 28.5K human-verified and 3,504K unfiltered question-answer-reasoning triplets have been curated, covering 4 primary abilities and 16 sub-abilities. We subsequently engage LLMs like GPT-3.5 to serve as judges, implementing the quantitative and qualitative automated assessments to facilitate a comprehensive evaluation of VLMs. Our validation results reveal that LLMs are proficient in both evaluation data curation and model assessment, achieving an average agreement rate of 85%. We envision Auto-Bench as a flexible, scalable, and comprehensive benchmark for evaluating the evolving sophisticated VLMs. Yuanfeng Ji, Chongjian Ge, Weikai Kong, Enze Xie, Zhengying Liu, Zhenguo Li, Ping Luo 0002 |
ICLR | 1 |
| 2023 | DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery - a Focus on Affinity Prediction Problems with Noise AnnotationsabstractAI-aided drug discovery (AIDD) is gaining popularity due to its potential to make the search for new pharmaceuticals faster, less expensive, and more effective. Despite its extensive use in numerous fields (e.g., ADMET prediction, virtual screening), little research has been conducted on the out-of-distribution (OOD) learning problem with noise. We present DrugOOD, a systematic OOD dataset curator and benchmark for AIDD. Particularly, we focus on the drug-target binding affinity prediction problem, which involves both macromolecule (protein target) and small-molecule (drug compound). DrugOOD offers an automated dataset curator with user-friendly customization scripts, rich domain annotations aligned with biochemistry knowledge, realistic noise level annotations, and rigorous benchmarking of SOTA OOD algorithms, as opposed to only providing fixed datasets. Since the molecular data is often modeled as irregular graphs using graph neural network (GNN) backbones, DrugOOD also serves as a valuable testbed for graph OOD learning problems. Extensive empirical studies have revealed a significant performance gap between in-distribution and out-of-distribution experiments, emphasizing the need for the development of more effective schemes that permit OOD generalization under noise for AIDD. Yuanfeng Ji, Lu Zhang 0060, Jiaxiang Wu 0001, Bingzhe Wu, Lanqing Li, Long-Kai Huang, Tingyang Xu, Yu Rong 0001, Ding Xue, Houtim Lai, Wei Liu 0005, Junzhou Huang, Shuigeng Zhou, Ping Luo 0002, Peilin Zhao, Yatao Bian |
AAAI | 1 |
| 2023 | DDP: Diffusion Model for Dense Visual PredictionabstractWe propose a simple, efficient, yet powerful framework for dense visual predictions based on the conditional diffusion pipeline. Our approach follows a "noise-to-map" generative paradigm for prediction by progressively removing noise from a random Gaussian distribution, guided by the image. The method, called DDP, efficiently extends the denoising diffusion process into the modern perception pipeline. Without task-specific design and architecture customization, DDP is easy to generalize to most dense prediction tasks, e.g., semantic segmentation and depth estimation. In addition, DDP shows attractive properties such as dynamic inference and uncertainty awareness, in contrast to previous single-step discriminative methods. We show top results on three representative tasks with six diverse benchmarks, without tricks, DDP achieves state-of-the-art or competitive performance on each task compared to the specialist counterparts. For example, semantic segmentation (83.9 mIoU on Cityscapes), BEV map segmentation (70.6 mIoU on nuScenes), and depth estimation (0.05 REL on KITTI). We hope that our approach will serve as a solid baseline and facilitate future research. Yuanfeng Ji, Zhe Chen 0017, Enze Xie, Lanqing Hong, Xihui Liu, Zhaoqiang Liu, Tong Lu 0002, Zhenguo Li, Ping Luo 0002 |
ICCV | 1 |
| 2023 | TAGNet: Learning Configurable Context Pathways for Semantic SegmentationabstractState-of-the-art semantic segmentation methods capture the relationship between pixels to facilitate contextual information exchange. Advanced methods utilize fixed pathways for context exchange, lacking the flexibility to harness the most relevant context for each pixel. In this paper, we present Configurable Context Pathways (CCPs), a novel model for establishing pathways for augmenting contextual information. In contrast to previous pathway models, CCPs are learned, leveraging configurable regions to form information flows between pairs of pixels. We propose TAGNet to adaptively configure the regions, which span over the entire image space, driven by the relationships between the remote pixels. Subsequently, the information flows along the pathways are updated gradually by the information provided by sequences of configurable regions, forming more powerful contextual information. We extensively evaluate the traveling, adaption, and gathering (TAG) stages of our network on the public benchmarks, demonstrating that all of the stages successfully improve the segmentation accuracy and help to surpass the state-of-the-art results. The code package is available at: https://github.com/dilincv/TAGNet. Di Lin 0002, Dingguo Shen, Yuanfeng Ji, Siting Shen, Mingrui Xie, Wei Feng 0005, Hui Huang 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | AMOS: A Large-Scale Abdominal Multi-Organ Benchmark for Versatile Medical Image SegmentationabstractDespite the considerable progress in automatic abdominal multi-organ segmentation from CT/MRI scans in recent years, a comprehensive evaluation of the models' capabilities is hampered by the lack of a large-scale benchmark from diverse clinical scenarios. Constraint by the high cost of collecting and labeling 3D medical data, most of the deep learning models to date are driven by datasets with a limited number of organs of interest or samples, which still limits the power of modern deep models and makes it difficult to provide a fully comprehensive and fair estimate of various methods. To mitigate the limitations, we present AMOS, a large-scale, diverse, clinical dataset for abdominal organ segmentation. AMOS provides 500 CT and 100 MRI scans collected from multi-center, multi-vendor, multi-modality, multi-phase, multi-disease patients, each with voxel-level annotations of 15 abdominal organs, providing challenging examples and test-bed for studying robust segmentation algorithms under diverse targets and scenarios. We further benchmark several state-of-the-art medical segmentation models to evaluate the status of the existing methods on this new challenging dataset. We have made our datasets, benchmark servers, and baselines publicly available, and hope to inspire future research. Information can be found at https://amos22.grand-challenge.org. Yuanfeng Ji, Haotian Bai, Chongjian Ge, Ruimao Zhang, Zhen Li 0026, Wanling Ma, Ping Luo 0002 |
NeurIPS | 1 |
| 2021 | Multi-compound Transformer for Accurate Biomedical Image Segmentation
Yuanfeng Ji, Ruimao Zhang, Huijie Wang, Zhen Li 0026, Lingyun Wu, Shaoting Zhang 0001, Ping Luo 0002 |
MICCAI (1) | 1 |
| 2021 | Multi-frame Collaboration for Effective Endoscopic Video Polyp Detection via Spatial-Temporal Feature Transformation
Lingyun Wu, Yuanfeng Ji, Ping Luo 0002, Shaoting Zhang 0001 |
MICCAI (5) | 3 |
| 2020 | UXNet: Searching Multi-level Feature Aggregation for 3D Medical Image Segmentation
Yuanfeng Ji, Ruimao Zhang, Zhen Li 0026, Jiamin Ren, Shaoting Zhang 0001, Ping Luo 0002 |
MICCAI (1) | 1 |
| 2020 | RANet: Region Attention Network for Semantic SegmentationabstractRecent semantic segmentation methods model the relationship between pixels to construct the contextual representations. In this paper, we introduce the \emph{Region Attention Network} (RANet), a novel attention network for modeling the relationship between object regions. RANet divides the image into object regions, where we select representative information. In contrast to the previous methods, RANet configures the information pathways between the pixels in different regions, enabling the region interaction to exchange the regional context for enhancing all of the pixels in the image. We train the construction of object regions, the selection of the representative regional contents, the configuration of information pathways and the context exchange between pixels, jointly, to improve the segmentation accuracy. We extensively evaluate our method on the challenging segmentation benchmarks, demonstrating that RANet effectively helps to achieve the state-of-the-art results. Dingguo Shen, Yuanfeng Ji, Ping Li 0016, Yi Wang 0031, Di Lin 0002 |
NeurIPS | 2 |
| 2020 | SCN: Switchable Context Network for Semantic Segmentation of RGB-D ImagesabstractContext representations have been widely used to profit semantic image segmentation. The emergence of depth data provides additional information to construct more discriminating context representations. Depth data preserves the geometric relationship of objects in a scene, which is generally hard to be inferred from RGB images. While deep convolutional neural networks (CNNs) have been successful in solving semantic segmentation, we encounter the problem of optimizing CNN training for the informative context using depth data to enhance the segmentation accuracy. In this paper, we present a novel switchable context network (SCN) to facilitate semantic segmentation of RGB-D images. Depth data is used to identify objects existing in multiple image regions. The network analyzes the information in the image regions to identify different characteristics, which are then used selectively through switching network branches. With the content extracted from the inherent image structure, we are able to generate effective context representations that are aware of both image structures and object relationships, leading to a more coherent learning of semantic segmentation network. We demonstrate that our SCN outperforms state-of-the-art methods on two public datasets. Di Lin 0002, Ruimao Zhang, Yuanfeng Ji, Ping Li 0016, Hui Huang 0004 |
IEEE Trans. Cybern. | 3 |
| 2019 | ZigZagNet: Fusing Top-Down and Bottom-Up Context for Object SegmentationabstractMulti-scale context information has proven to be essential for object segmentation tasks. Recent works construct the multi-scale context by aggregating convolutional feature maps extracted by different levels of a deep neural network. This is typically done by propagating and fusing features in a one-directional, top-down and bottom-up, manner. In this work, we introduce ZigZagNet, which aggregates a richer multi-context feature map by using not only dense top-down and bottom-up propagation, but also by introducing pathways crossing between different levels of the top-down and the bottom-up hierarchies, in a zig-zag fashion. Furthermore, the context information is exchanged and aggregated over multiple stages, where the fused feature maps from one stage are fed into the next one, yielding a more comprehensive context for improved segmentation performance. Our extensive evaluation on the public benchmarks demonstrates that ZigZagNet surpasses the state-of-the-art accuracy for both semantic segmentation and instance segmentation tasks. Di Lin 0002, Dingguo Shen, Siting Shen, Yuanfeng Ji, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
CVPR | 4 |
| 2019 | PRSNet: Part Relation and Selection Network for Bone Age Assessment
Yuanfeng Ji, Hao Chen 0011, Dan Lin 0009, Di Lin 0002 |
MICCAI (6) | 1 |
| 2018 | Multi-scale Context Intertwining for Semantic Segmentation
Di Lin 0002, Yuanfeng Ji, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
ECCV (3) | 2 |