Guoan Wang

dblp:59/9431 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AI
abstract
Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annotations into high-quality image-text pairs. This dataset offers comprehensive task coverage, diverse modalities, and rich image-text data. Building upon this dataset, we develop GMAI-VL, a 7B-parameter general medical vision-language model, with a three-stage training strategy that enhances the integration of visual and textual information. This approach significantly improves the model's ability to process multimodal data, supporting accurate diagnoses and clinical decision-making. Experiments show that GMAI-VL achieves state-of-the-art performance across various multimodal medical tasks, including visual question answering and medical image diagnosis.
Tianbin Li, Yanzhou Su, Wei Li 0320, Zhe Chen 0017, Ziyan Huang, Guoan Wang, Chenglong Ma 0002, Yanjun Li 0007, Shixiang Tang, Xiaowei Hu 0001, Zhongying Deng, Yuanfeng Ji, Jin Ye 0002, Yu Qiao 0001, Junjun He
AAAI7
2025 SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding
abstract
Despite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essential contextual information at the whole-slide level. The lack of large-scale instruction datasets and the gigapixel scale of whole slide images (WSIs) pose significant developmental challenges. In this paper, we present SlideChat, the first vision-language assistant capable of understanding gigapixel whole-slide images, exhibiting excellent multimodal conversational capability and response complex instruction across diverse pathology scenarios. To support its development, we created SlideInstruction, the largest instruction-following dataset for WSIs consisting of 4.2K WSI captions and 176K VQA pairs with multiple categories. Furthermore, we propose SlideBench, a multimodal benchmark that incorporates captioning and VQA tasks to assess SlideChat’s capabilities in various settings such as microscopy, diagnosis and clinical. Compared to both general and specialized MLLMs, SlideChat exhibits exceptional capabilities, achieving state-of-the-art performance on 18 of 22 tasks. For example, it achieved an overall accuracy of 81.17% on SlideBench-VQA (TCGA), and 54.15% on SlideBench-VQA (BCNB). Our code, data, and model is publicly accessible at https://uni-medical.github.io/SlideChat.github.io.
Guoan Wang, Yuanfeng Ji, Yanjun Li 0007, Jin Ye 0002, Tianbin Li, Rongshan Yu, Yu Qiao 0001, Junjun He
CVPR2
2025 Interactive Medical Image Segmentation: A Benchmark Dataset and Baseline
abstract
Interactive Medical Image Segmentation (IMIS) has long been constrained by the limited availability of large-scale, diverse, and densely annotated datasets, which hinders model generalization and consistent evaluation across different models. In this paper, we introduce the IMed-361M benchmark dataset, a significant advancement in general IMIS research. First, we collect and standardize over 6.4 million medical images and their corresponding ground truth masks from multiple data sources. Then, leveraging the strong object recognition capabilities of a vision foundational model, we automatically generated dense interactive masks for each image and ensured their quality through rigorous quality control and granularity management. Unlike previous datasets, which are limited by specific modalities or sparse annotations, IMed-361M spans 14 modalities and 204 segmentation targets, totaling 361 million masks—an average of 56 masks per image. Finally, we developed an IMIS baseline network on this dataset that supports high-quality mask generation through interactive inputs, including clicks, bounding boxes, text prompts, and their combinations. We evaluate its performance on medical image segmentation tasks from multiple perspectives, demonstrating superior accuracy and scalability compared to existing interactive segmentation models. To facilitate research on foundational models in medical computer vision, we release the IMed-361M and model at https://github.com/uni-medical/IMIS-Bench.
Junlong Cheng, Jin Ye 0002, Guoan Wang, Tianbin Li, Haoyu Wang 0010, He Yao, Yanzhou Su, Min Zhu 0005, Junjun He
CVPR4
2025 MCTrack: A Unified 3D Multi-Object Tracking Framework for Autonomous Driving
abstract
This paper introduces MCTrack, a new 3D multi-object tracking method that achieves performance across KITTI, nuScenes, and Waymo datasets. Addressing the gap in existing tracking paradigms, which often perform well on specific datasets but lack generalizability, MCTrack offers a unified solution. Additionally, we have standardized the format of perceptual results across various datasets, termed BaseVersion, facilitating researchers in the field of MOT) to concentrate on the core algorithmic development without the undue burden of data preprocessing. Finally, recognizing the limitations of current evaluation metrics, we introduce a novel set of metrics designed to evaluate the output of motion information, including velocity and acceleration, which are essential for subsequent tasks. The source codes of the proposed method are available at this link: https://github.com/megvii-research/MCTrack
Xiyang Wang 0002, Shouzheng Qi, Jieyou Zhao, Hangning Zhou, Siyu Zhang 0002, Guoan Wang, Kai Tu, Songlin Guo, Jianbo Zhao 0001, Hailong Qin, Mu Yang
IROS6
2025 Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model
Wei Li 0320, Guoan Wang, Kaijing Zhou, Junzhi Ning, ZongYuan Ge, Lixu Gu, Junjun He
MICCAI (9)3
2024 StreamMOTP: Streaming and Unified Framework for Joint 3D Multi-Object Tracking and Trajectory Prediction
Jiaheng Zhuang, Guoan Wang, Siyu Zhang 0002, Xiyang Wang 0002, Hangning Zhou, Ziyao Xu 0003, Chi Zhang 0067, Zhiheng Li 0001
ACCV (2)2
2024 SAM-Med3D-MoE: Towards a Non-Forgetting Segment Anything Model via Mixture of Experts for 3D Medical Image Segmentation
Guoan Wang, Jin Ye 0002, Junlong Cheng, Tianbin Li, Zhaolin Chen, Jianfei Cai 0001, Junjun He, Bohan Zhuang
MICCAI (9)1
2024 GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI
abstract
Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial assistance for diagnosis and treatment. Before that, it is crucial to develop benchmarks to evaluate LVLMs' effectiveness in various medical applications. Current benchmarks are often built upon specific academic literature, mainly focusing on a single domain, and lacking varying perceptual granularities. Thus, they face specific challenges, including limited clinical relevance, incomplete evaluations, and insufficient guidance for interactive LVLMs. To address these limitations, we developed the GMAI-MMBench, the most comprehensive general medical AI benchmark with well-categorized data structure and multi-perceptual granularity to date. It is constructed from 284 datasets across 38 medical image modalities, 18 clinical-related tasks, 18 departments, and 4 perceptual granularities in a Visual Question Answering (VQA) format. Additionally, we implemented a lexical tree structure that allows users to customize evaluation tasks, accommodating various assessment needs and substantially supporting medical AI research and applications. We evaluated 50 LVLMs, and the results show that even the advanced GPT-4o only achieves an accuracy of 53.96\%, indicating significant room for improvement. Moreover, we identified five key insufficiencies in current cutting-edge LVLMs that need to be addressed to advance the development of better medical applications. We believe that GMAI-MMBench will stimulate the community to build the next generation of LVLMs toward GMAI.
Jin Ye 0002, Guoan Wang, Yanjun Li 0007, Zhongying Deng, Wei Li 0320, Tianbin Li, Haodong Duan, Ziyan Huang, Yanzhou Su, Benyou Wang, Shaoting Zhang 0001, Jianfei Cai 0001, Bohan Zhuang, Eric J. Seibel, Junjun He, Yu Qiao 0001
NeurIPS3
2023 SQT: Debiased Visual Question Answering via Shuffling Question Types
abstract
Visual Question Answering (VQA) aims to obtain answers through image-question pairs. Nowadays, the VQA model tends to get answers only through questions, ignoring the information in the images. This phenomenon is caused by bias. As indicated by previous studies, the bias in VQA mainly comes from text modality. Our analysis of bias suggests that the question type is a crucial factor in bias formation. To interrupt the shortcut from question type to answer for de-biasing, we propose a self-supervised method for Shuffling Question Types (SQT) to reduce bias from text modality, which overcomes the prior language problem by mitigating the question-to-answer bias without introducing external annotations. Moreover, we propose a new objective function for negative samples. Experimental results show that our approach can achieve 61.76% accuracy on the VQA-CP v2 dataset, which outperforms the state-of-the-art in both self-supervised and supervised methods.
Tianyu Huai, Junhang Zhang, Guoan Wang, Xinru Yu, Tianlong Ma, Liang He 0001
ICME4
2023 PRFNet: Progressive Region Focusing Network for Polyp Segmentation
Jilong Chen, Junlong Cheng, Pengyu Yin, Guoan Wang
PRCV (5)5
2022 FlexVAA: A Flexible, Passive van Atta Retroreflector for Roadside Infrastructure Tagging and Identification
abstract
We propose FlexVAA, a system for identifying roadside infrastructure using flexible, passive wireless retroreflectors. FlexVAA can be easily attached to any surface, allowing roadside infrastructure to be upgraded for autonomous systems without impairing existing operations. Preliminary results show that attaching FlexVAA to a surface reflects more power than without FlexVAA attached. In the future, we plan to arrange multiple FlexVAA elements in order to embed data in the passively reflected signal, allowing identification for multiple unique tags.
Nicholas Junker, Jinqun Ge, Guoan Wang, Sanjib Sur 0001
SenSys3
2020 Design and Optimization Methodology of Coplanar Waveguide Test Structures for Dielectric Characterization of Thin Films
Jinqun Ge, Tian Xia 0005, Guoan Wang
J. Electron. Test.3
2018 A Simplified Calibration Methodology for On-Chip Couplers
Wei Jiang 0017, Guoan Wang
J. Electron. Test.2
2015 On-Wafer Calibration Technique for High Frequency Measurement with Simultaneous Voltage and Current Tuning
B. M. Farid Rahman, Yujia Peng, TengXing Wang, Tian Xia 0005, Guoan Wang
J. Electron. Test.5
2014 Characterization of a Passive Telemetric System for ISM Band Pressure Sensors
Yujia Peng, B. M. Farid Rahman, TengXing Wang, Guoan Wang, Xinchuan Liu, Xuejun Wen
J. Electron. Test.4
2010 An adaptive body-bias low voltage low power LC VCO
abstract
A novel LC VCO employing an adaptive body-bias is proposed for low power and low voltage high performance signal generation. By using the adaptive body-bias technique, the threshold voltage of the cross-coupled pair is reduced, and meanwhile the effective source-to-drain voltage and oscillation amplitude are increased, which improves the phase noise. In addition, the back-gate capacitances from the cross-coupled pair are used for frequency tuning which avoids quality factor deteriorations resulted from using extra tuning capacitance. In a standard 65nm CMOS process with a 0.4-V power supply, the proposed VCO achieves FOM of 189.8 with a phase noise of −122.5 dBc/Hz at 10MHz offset when operating at 24.25 GHz. The power consumption of the VCO is 0.94 mW.
Pinping Sun, Guoan Wang, Wayne H. Woods, Ya Jun Yu
ISCAS2