VLDB 2026 Research / reviewers in the wild / expert
Jin Ye 0002
dblp:33/2058-2
· DBLP profile ↗
30ranked-venue papers
4as first author
28since 2021 · last 2026
0000-0003-0667-9889ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 17 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AIabstractDespite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annotations into high-quality image-text pairs. This dataset offers comprehensive task coverage, diverse modalities, and rich image-text data. Building upon this dataset, we develop GMAI-VL, a 7B-parameter general medical vision-language model, with a three-stage training strategy that enhances the integration of visual and textual information. This approach significantly improves the model's ability to process multimodal data, supporting accurate diagnoses and clinical decision-making. Experiments show that GMAI-VL achieves state-of-the-art performance across various multimodal medical tasks, including visual question answering and medical image diagnosis. Tianbin Li, Yanzhou Su, Wei Li 0320, Zhe Chen 0017, Ziyan Huang, Guoan Wang, Chenglong Ma 0002, Yanjun Li 0007, Shixiang Tang, Xiaowei Hu 0001, Zhongying Deng, Yuanfeng Ji, Jin Ye 0002, Yu Qiao 0001, Junjun He |
AAAI | 17 |
| 2026 | S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything Without SupervisionabstractRecent self-supervised image segmentation models have achieved promising performance on semantic segmentation and class-agnostic instance segmentation. However, their pretraining schedule is multi-stage, requiring a time-consuming pseudo-masks generation process between each training epoch. This time-consuming offline process not only makes it difficult to scale with training dataset size, but also leads to sub-optimal solutions due to its discontinuous optimization routine. To solve these, we first present a novel pseudo-mask algorithm, Fast Universal Agglomerative Pooling (UniAP). Each layer of UniAP can identify groups of similar nodes in parallel, allowing to generate both semantic-level and instance-level and multi-granular pseudo-masks within ens of milliseconds for one image. Based on the fast UniAP, we propose the Scalable Self-Supervised Universal Segmentation (S2-UniSeg), which employs a student and a momentum teacher for continuous pretraining. A novel segmentation-oriented pretext task, Query-wise Self-Distillation (QuerySD), is proposed to pretrain S2-UniSeg to learn the local-to-global correspondences. Under the same setting, S2-UniSeg outperforms the SOTA UnSAM model, achieving notable improvements of AP+6.9 on COCO, AR+11.1 on UVO, PixelAcc+4.5 on COCOStuff-27, RQ+8.0 on Cityscapes. After scaling up to a larger 2M-image subset of SA-1B, S2-UniSeg further achieves performance gains on all four benchmarks. Jin Ye 0002, Hongqiu Wang, Changkai Ji, Jiashi Lin, Ziyan Huang, Chenglong Ma 0002, Tianbin Li, Junjun He, Lei Zhu 0003 |
AAAI | 2 |
| 2026 | SegRap2025: A benchmark of gross tumor volume and lymph node clinical target volume Segmentation for Radiotherapy Planning of nasopharyngeal carcinoma
Litingyu Wang, Chenyuan Bian, Zijun Gao, Chunbin Gu, Xin Weng, Jianghao Wu 0001, Yicheng Wu 0001, Jin Ye 0002, Linhao Li, Yiwen Ye, Yong Xia 0001, Elias Tappeiner, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, Junqiang Chen, Chuanyi Huang, Lisheng Wang, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Shichuan Zhang, Shaoting Zhang 0001, Wenjun Liao, Guotai Wang |
Medical Image Anal. | 12 |
| 2025 | SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image UnderstandingabstractDespite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essential contextual information at the whole-slide level. The lack of large-scale instruction datasets and the gigapixel scale of whole slide images (WSIs) pose significant developmental challenges. In this paper, we present SlideChat, the first vision-language assistant capable of understanding gigapixel whole-slide images, exhibiting excellent multimodal conversational capability and response complex instruction across diverse pathology scenarios. To support its development, we created SlideInstruction, the largest instruction-following dataset for WSIs consisting of 4.2K WSI captions and 176K VQA pairs with multiple categories. Furthermore, we propose SlideBench, a multimodal benchmark that incorporates captioning and VQA tasks to assess SlideChat’s capabilities in various settings such as microscopy, diagnosis and clinical. Compared to both general and specialized MLLMs, SlideChat exhibits exceptional capabilities, achieving state-of-the-art performance on 18 of 22 tasks. For example, it achieved an overall accuracy of 81.17% on SlideBench-VQA (TCGA), and 54.15% on SlideBench-VQA (BCNB). Our code, data, and model is publicly accessible at https://uni-medical.github.io/SlideChat.github.io. Guoan Wang, Yuanfeng Ji, Yanjun Li 0007, Jin Ye 0002, Tianbin Li, Rongshan Yu, Yu Qiao 0001, Junjun He |
CVPR | 5 |
| 2025 | Interactive Medical Image Segmentation: A Benchmark Dataset and BaselineabstractInteractive Medical Image Segmentation (IMIS) has long been constrained by the limited availability of large-scale, diverse, and densely annotated datasets, which hinders model generalization and consistent evaluation across different models. In this paper, we introduce the IMed-361M benchmark dataset, a significant advancement in general IMIS research. First, we collect and standardize over 6.4 million medical images and their corresponding ground truth masks from multiple data sources. Then, leveraging the strong object recognition capabilities of a vision foundational model, we automatically generated dense interactive masks for each image and ensured their quality through rigorous quality control and granularity management. Unlike previous datasets, which are limited by specific modalities or sparse annotations, IMed-361M spans 14 modalities and 204 segmentation targets, totaling 361 million masks—an average of 56 masks per image. Finally, we developed an IMIS baseline network on this dataset that supports high-quality mask generation through interactive inputs, including clicks, bounding boxes, text prompts, and their combinations. We evaluate its performance on medical image segmentation tasks from multiple perspectives, demonstrating superior accuracy and scalability compared to existing interactive segmentation models. To facilitate research on foundational models in medical computer vision, we release the IMed-361M and model at https://github.com/uni-medical/IMIS-Bench. Junlong Cheng, Jin Ye 0002, Guoan Wang, Tianbin Li, Haoyu Wang 0010, He Yao, Yanzhou Su, Min Zhu 0005, Junjun He |
CVPR | 3 |
| 2025 | OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language PretrainingabstractSurgical practice involves complex visual interpretation, procedural skills, and advanced medical knowledge, making surgical vision-language pretraining (VLP) particularly challenging due to this complexity and the limited availability of annotated data. To address the gap, we propose OphCLIP, a hierarchical retrieval-augmented vision-language pretraining framework specifically designed for ophthalmic surgical workflow understanding. OphCLIP leverages the OphVL dataset we constructed, a large-scale and comprehensive collection of over 375K hierarchically structured video-text pairs with tens of thousands of different combinations of attributes (surgeries, phases/operations/actions, instruments, medications, as well as more advanced aspects like the causes of eye diseases, surgical objectives, and postoperative recovery recommendations, etc). These hierarchical video-text correspondences enable OphCLIP to learn both fine-grained and long-term visual representations by aligning short video clips with detailed narrative descriptions and full videos with structured titles, capturing intricate surgical details and high-level procedural insights, respectively. Our OphCLIP also designs a retrieval-augmented pretraining framework to leverage the underexplored large-scale silent surgical procedure videos, automatically retrieving semantically relevant content to enhance the representation learning of narrative videos. Evaluation across 11 datasets for phase recognition and multi-instrument identification shows OphCLIP's robust generalization and superior performance. Kun Yuan 0004, Yaling Shen, Xiaohao Xu, Wei Li 0320, Zhongxing Xu, Zelin Peng, Siyuan Yan, Vinkle Srivastav, Diping Song, Tianbin Li, Danli Shi, Jin Ye 0002, Nicolas Padoy, Nassir Navab, Junjun He, ZongYuan Ge |
ICCV | 16 |
| 2025 | Towards Interpretable Counterfactual Generation via Multimodal Autoregression
Chenglong Ma 0002, Yuanfeng Ji, Jin Ye 0002, Lu Zhang 0060, Tianbin Li, Mingjie Li 0006, Junjun He, Hongming Shan |
MICCAI (2) | 3 |
| 2025 | RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions
Junzhi Ning, Cheng Tang 0003, Kaijing Zhou, Diping Song, Wei Li 0320, Yanzhou Su, Tianbin Li, Jiyao Liu, Jin Ye 0002, Yuanfeng Ji, Junjun He |
MICCAI (16) | 12 |
| 2025 | New Multiple Sclerosis Lesion Segmentation via Calibrated Inter-patch Blending
Jin Ye 0002, Son Duy Dao, Yicheng Wu 0001, Yasmeen M. George, Thanh Nguyen-Duc, Daniel F. Schmidt, Hengcan Shi, Winston Chong, Jianfei Cai 0001 |
MICCAI (16) | 1 |
| 2025 | FPN-in-FPN: A Nested Multi-scale Aggregation Network for Polyp Segmentation
Jin Ye 0002, Yanzhou Su, Yicheng Wu 0001, Junjun He, Bohan Zhuang, Zhaolin Chen, Jianfei Cai 0001 |
MICCAI (11) | 1 |
| 2025 | A survey for large language models in biomedicine
Chong Wang 0027, Junjun He, Zhongruo Wang, Erfan Darzi, Jin Ye 0002, Tianbin Li, Yanzhou Su, Jing Ke, Kaili Qu, Pietro Liò, Tianyun Wang, Yu Guang Wang 0001, Yiqing Shen 0003 |
Artif. Intell. Medicine | 7 |
| 2025 | FCN+: Global receptive convolution makes FCN great again
Zhongying Deng, Jin Ye 0002, Junjun He |
Neurocomputing | 3 |
| 2025 | PitVis-2023 challenge: Workflow recognition in videos of endoscopic pituitary surgeryabstractThe field of computer vision applied to videos of minimally invasive surgery is ever-growing. Workflow recognition pertains to the automated recognition of various aspects of a surgery, including: which surgical steps are performed; and which surgical instruments are used. This information can later be used to assist clinicians when learning the surgery or during live surgery. The Pituitary Vision (PitVis) 2023 Challenge tasks the community to step and instrument recognition in videos of endoscopic pituitary surgery. This is a particularly challenging task when compared to other minimally invasive surgeries due to: the smaller working space, which limits and distorts vision; and higher frequency of instrument and step switching, which requires more precise model predictions. Participants were provided with 25-videos, with results presented at the MICCAI-2023 conference as part of the Endoscopic Vision 2023 Challenge in Vancouver, Canada, on 08-Oct-2023. There were 18-submissions from 9-teams across 6-countries, using a variety of deep learning models. The top performing model for step recognition utilised a transformer based architecture, uniquely using an autoregressive decoder with a positional encoding input. The top performing model for instrument recognition utilised a spatial encoder followed by a temporal encoder, which uniquely used a 2-layer temporal architecture. In both cases, these models outperformed purely spatial based models, illustrating the importance of sequential and temporal information. This PitVis-2023 therefore demonstrates state-of-the-art computer vision models in minimally invasive surgery are transferable to a new dataset. Benchmark results are provided in the paper, and the dataset is publicly available at: https://doi.org/10.5522/04/26531686. Adrito Das, Danyal Z. Khan, Dimitris Psychogyios, John G. Hanrahan, Francisco Vasconcelos 0001, You Pang, Zhen Chen 0018, Jinlin Wu, Xiaoyang Zou, Guoyan Zheng, Abdul Qayyum 0002, Moona Mazher, Muhammad Imran Razzak, Tianbin Li, Jin Ye 0002, Junjun He, Szymon Plotka, Joanna Kaleta, Amine Yamlahi, Antoine Jund, Patrick Godau, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Dominik Rivoir, Stefanie Speidel, Alejandra Pérez, Santiago Rodríguez, Pablo Andrés Arbeláez, Danail Stoyanov, Hani J. Marcus, Sophia Bano |
Medical Image Anal. | 16 |
| 2025 | A-Eval: A benchmark for cross-dataset and cross-modality evaluation of abdominal multi-organ segmentation
Ziyan Huang, Zhongying Deng, Jin Ye 0002, Haoyu Wang 0010, Yanzhou Su, Tianbin Li, Junlong Cheng, Jianpin Chen, Junjun He, Yun Gu, Shaoting Zhang 0001, Lixu Gu, Yu Qiao 0001 |
Medical Image Anal. | 3 |
| 2025 | SegRap2023: A benchmark of organs-at-risk and gross tumor volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma
Xiangde Luo, Yunxin Zhong, Shuolin Liu, Mehdi Astaraki, Simone Bendazzoli, Iuliana Toma-Dasu, Yiwen Ye, Ziyang Chen 0003, Yong Xia 0001, Yanzhou Su, Jin Ye 0002, Junjun He, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Kaixiang Yang 0004, Zhiwei Wang 0002, Chan Woong Lee, Sang Joon Park, Jaehee Chun, Constantin Ulrich, Klaus H. Maier-Hein, Nchongmaje Ndipenoch, Alina Dana Miron, Yongmin Li 0001, Chengyang An, Lisheng Wang, Kaiwen Huang 0002, Yunqi Gu, Tao Zhou 0002, Mu Zhou, Shichuan Zhang, Wenjun Liao, Guotai Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 13 |
| 2025 | RFMiD: Retinal Image Analysis for multi-Disease Detection challenge
Samiksha Pachade, Prasanna Porwal, Manesh Kokare, Girish Deshmukh, Vivek Sahasrabuddhe, Zhengbo Luo, Zitang Sun, Li Qihan, Edward Ho, Asaanth Sivajohan, Saerom Youn, Kevin Lane, Jin Chun, Yunchao Gu, Sixu Lu, Young-tack Oh, Hyunjin Park, Chia-Yen Lee, Hung Yeh, Kai-Wen Cheng, Haoyu Wang 0010, Jin Ye 0002, Junjun He, Lixu Gu, Dominik Müller, Iñaki Soto Rey, Frank Kramer 0001, Hidehisa Arai, Yuma Ochi, Takami Okada, Luca Giancardo, Gwenolé Quellec, Fabrice Mériaudeau |
Medical Image Anal. | 26 |
| 2025 | Multi-Center Fetal Brain Tissue Annotation (FeTA) Challenge 2022 ResultsabstractSegmentation is a critical step in analyzing the developing human fetal brain. There have been vast improvements in automatic segmentation methods in the past several years, and the Fetal Brain Tissue Annotation (FeTA) Challenge 2021 helped to establish an excellent standard of fetal brain segmentation. However, FeTA 2021 was a single center study, limiting real-world clinical applicability and acceptance. The multi-center FeTA Challenge 2022 focused on advancing the generalizability of fetal brain segmentation algorithms for magnetic resonance imaging (MRI). In FeTA 2022, the training dataset contained images and corresponding manually annotated multi-class labels from two imaging centers, and the testing data contained images from these two centers as well as two additional unseen centers. The multi-center data included different MR scanners, imaging parameters, and fetal brain super-resolution algorithms applied. 16 teams participated and 17 algorithms were evaluated. Here, the challenge results are presented, focusing on the generalizability of the submissions. Both in- and out-of-domain, the white matter and ventricles were segmented with the highest accuracy (Top Dice scores: 0.89, 0.87 respectively), while the most challenging structure remains the grey matter (Top Dice score: 0.75) due to anatomical complexity. The top 5 average Dices scores ranged from 0.81-0.82, the top 5 average percentile Hausdorff distance values ranged from 2.3-2.5mm, and the top 5 volumetric similarity scores ranged from 0.90-0.92. The FeTA Challenge 2022 was able to successfully evaluate and advance generalizability of multi-class fetal brain tissue segmentation algorithms for MRI and it continues to benchmark new algorithms. Kelly Payette, Céline Steger, Roxane Licandro, Priscille de Dumast, Hongwei Li 0004, Matthew J. Barkovich, Liu Li 0001, Maik Dannecker, Chen Chen 0042, Cheng Ouyang, Niccolò McConnell, Alina Dana Miron, Yongmin Li 0001, Alena Uus, Irina Grigorescu, Paula Ramirez Gilliland, Md Mahfuzur Rahman Siddiquee, Daguang Xu, Andriy Myronenko, Haoyu Wang 0010, Ziyan Huang, Jin Ye 0002, Mireia Alenyà, Valentin Comte, Oscar Camara 0001, Jean-Baptiste Masson, Astrid Nilsson, Charlotte Godard, Moona Mazher, Abdul Qayyum 0002, Yibo Gao, Hangqi Zhou, Shangqi Gao, Guiming Dong, Guotai Wang, ZunHyan Rieu, HyeonSik Yang, Szymon Plotka, Michal K. Grzeszczyk, Arkadiusz Sitek, Luisa Vargas Daza, Santiago Usma, Pablo Andrés Arbeláez, Wenying Lu, Romain Valabrègue, Anand A. Joshi, Krishna N. Nayak, Richard M. Leahy, Luca Wilhelmi, Aline Dändliker, Antonio G. Gennari, Anton Jakovcic, Melita Klaic, Ana Adzic, Pavel Markovic, Gracia Grabaric, Gregor Kasprian, Gregor Dovjak, Milan Rados, Lana Vasung, Meritxell Bach Cuadra, András Jakab |
IEEE Trans. Medical Imaging | 22 |
| 2025 | SAM-Med3D: A Vision Foundation Model for General-Purpose Segmentation on Volumetric Medical ImagesabstractExisting volumetric medical image segmentation models are typically task-specific, excelling at specific targets but struggling to generalize across anatomical structures or modalities. This limitation restricts their broader clinical use. In this article, we introduce segment anything model (SAM)-Med3D, a vision foundation model (VFM) for general-purpose segmentation on volumetric medical images. Given only a few 3-D prompt points, SAM-Med3D can accurately segment diverse anatomical structures and lesions across various modalities. To achieve this, we gather and preprocess a large-scale 3-D medical image segmentation dataset, SA-Med3D-140K, from 70 public datasets and 8K licensed private cases from hospitals. This dataset includes 22K 3-D images and 143K corresponding masks. SAM-Med3D, a promptable segmentation model characterized by its fully learnable 3-D structure, is trained on this dataset using a two-stage procedure and exhibits impressive performance on both seen and unseen segmentation targets. We comprehensively evaluate SAM-Med3D on 16 datasets covering diverse medical scenarios, including different anatomical structures, modalities, targets, and zero-shot transferability to new/unseen tasks. The evaluation demonstrates the efficiency and efficacy of SAM-Med3D, as well as its promising application to diverse downstream tasks as a pretrained model. Our approach illustrates that substantial medical resources can be harnessed to develop a general-purpose medical AI for various potential applications. Our dataset, code, and models are available at: https://github.com/uni-medical/SAM-Med3D. Haoyu Wang 0010, Sizheng Guo, Jin Ye 0002, Zhongying Deng, Junlong Cheng, Tianbin Li, Jianpin Chen, Yanzhou Su, Ziyan Huang, Yiqing Shen 0003, Shaoting Zhang 0001, Junjun He |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Histology Image Artifact Restoration with Lightweight Transformer Based Diffusion Model
Chong Wang 0027, Zhenqi He, Junjun He, Jin Ye 0002, Yiqing Shen 0003 |
AIME (2) | 4 |
| 2024 | SAM-Med3D-MoE: Towards a Non-Forgetting Segment Anything Model via Mixture of Experts for 3D Medical Image Segmentation
Guoan Wang, Jin Ye 0002, Junlong Cheng, Tianbin Li, Zhaolin Chen, Jianfei Cai 0001, Junjun He, Bohan Zhuang |
MICCAI (9) | 2 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 11 |
| 2024 | GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AIabstractLarge Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial assistance for diagnosis and treatment. Before that, it is crucial to develop benchmarks to evaluate LVLMs' effectiveness in various medical applications. Current benchmarks are often built upon specific academic literature, mainly focusing on a single domain, and lacking varying perceptual granularities. Thus, they face specific challenges, including limited clinical relevance, incomplete evaluations, and insufficient guidance for interactive LVLMs. To address these limitations, we developed the GMAI-MMBench, the most comprehensive general medical AI benchmark with well-categorized data structure and multi-perceptual granularity to date. It is constructed from 284 datasets across 38 medical image modalities, 18 clinical-related tasks, 18 departments, and 4 perceptual granularities in a Visual Question Answering (VQA) format. Additionally, we implemented a lexical tree structure that allows users to customize evaluation tasks, accommodating various assessment needs and substantially supporting medical AI research and applications. We evaluated 50 LVLMs, and the results show that even the advanced GPT-4o only achieves an accuracy of 53.96\%, indicating significant room for improvement. Moreover, we identified five key insufficiencies in current cutting-edge LVLMs that need to be addressed to advance the development of better medical applications. We believe that GMAI-MMBench will stimulate the community to build the next generation of LVLMs toward GMAI. Jin Ye 0002, Guoan Wang, Yanjun Li 0007, Zhongying Deng, Wei Li 0320, Tianbin Li, Haodong Duan, Ziyan Huang, Yanzhou Su, Benyou Wang, Shaoting Zhang 0001, Jianfei Cai 0001, Bohan Zhuang, Eric J. Seibel, Junjun He, Yu Qiao 0001 |
NeurIPS | 2 |
| 2023 | Artifact Restoration in Histology Images with Diffusion Probabilistic Models
Zhenqi He, Junjun He, Jin Ye 0002, Yiqing Shen 0003 |
MICCAI (6) | 3 |
| 2023 | Revisiting Feature Propagation and Aggregation in Polyp Segmentation
Yanzhou Su, Yiqing Shen 0003, Jin Ye 0002, Junjun He, Jian Cheng 0003 |
MICCAI (5) | 3 |
| 2023 | Pick the Best Pre-trained Model: Towards Transferability Estimation for Medical Image Segmentation
Yuncheng Yang, Junjun He, Jin Ye 0002, Yun Gu |
MICCAI (1) | 5 |
| 2023 | Accurate polyp segmentation through enhancing feature fusion and boosting boundary performance
Yanzhou Su, Jian Cheng 0003, Chuqiao Zhong, Chengzhi Jiang, Jin Ye 0002, Junjun He |
Neurocomputing | 5 |
| 2021 | A Novel Hybrid Convolutional Neural Network for Accurate Organ Segmentation in 3D Head and Neck CT Images
Cheng Li 0008, Junjun He, Jin Ye 0002, Diping Song, Shanshan Wang 0002, Lixu Gu, Yu Qiao 0001 |
MICCAI (1) | 4 |
| 2021 | Group Shift Pointwise Convolution for Volumetric Medical Image Segmentation
Junjun He, Jin Ye 0002, Cheng Li 0008, Diping Song, Shanshan Wang 0002, Lixu Gu, Yu Qiao 0001 |
MICCAI (3) | 2 |
| 2020 | Attention-Driven Dynamic Graph Convolutional Network for Multi-label Image Recognition
Jin Ye 0002, Junjun He, Xiaojiang Peng, Yu Qiao 0001 |
ECCV (21) | 1 |
| 2019 | Visual-Textual Sentiment Analysis in Product ReviewsabstractSentiment analysis has attracted increasing attention recently due to its potential wide applications in opinion analysis, recommendation system, etc. Visual-textual sentiment analysis aims to improve the performance of sentiment analysis by leveraging both visual and textual signals. In this paper, we address the visual-textual sentiment analysis in product reviews. Our main contributions are two-fold. First, instead of crawling data from Flickr or Twitter with positive and negative labels in existing works, we introduce a new dataset for visual-textual sentiment analysis, termed as Product Reviews150K (PR-150K), which is collected from the product reviews of online shopping websites. Second, we propose a deep Tucker fusion method for visual-textual sentiment analysis, which efficiently combines visual and textual deep representations based on the Tucker decomposition and a bilinear pooling operation. Extensive experiments on our PR-150K, MVSO, and VSO datasets show that our method outperforms several state-of-the-art methods. Jin Ye 0002, Xiaojiang Peng, Yu Qiao 0001, Rongrong Ji |
ICIP | 1 |