VLDB 2026 Research / reviewers in the wild / expert
Yuxuan Sun 0002
dblp:194/3079-2
· DBLP profile ↗
26ranked-venue papers
9as first author
22since 2021 · last 2026
0000-0002-1277-4316ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIRA: Evaluating Multimodal AI on Complex Clinical Reasoning in Interventional RadiologyabstractWe present MIRA (Multimodal Interventional RAdiology evaluation), a comprehensive benchmark for evaluating large multimodal models in expert-level interventional radiology tasks requiring specialized domain knowledge and advanced visual reasoning capabilities. Unlike existing medical benchmarks that primarily provide binary labels without contextual depth, MIRA offers diverse question formats, including open-ended, closed-ended, single-choice, and multiple-choice categories, each accompanied by detailed expert-validated explanations. The benchmark incorporates approximately 184K high-quality medical images spanning multiple imaging modalities with 1.2M meticulously generated question-answer pairs across various anatomical regions. These pairs were created through a sophisticated cascade methodology involving expert interventional radiologists at both the data collection and validation stages. Our comprehensive evaluation, encompassing zero-shot testing and fine-tuning experiments of large multimodal models, revealing significant performance gaps between AI systems and human specialists. Fine-tuning experiments demonstrate substantial improvements, with models achieving up to 0.80 accuracy on single-choice questions. MIRA establishes a challenging benchmark that suggests promising directions for developing specialized clinical AI systems for interventional radiology. Jingxiong Li, Chenglu Zhu, Sunyi Zheng, Yuxuan Sun 0002, Yixuan Si, Lin Yang 0002, Liang Xiao 0001 |
AAAI | 4 |
| 2026 | Towards Effective and Efficient Context-aware Nucleus Detection in Histopathology Whole Slide ImagesabstractNucleus detection in histopathology whole slide images (WSIs) is crucial for a broad spectrum of clinical applications. The gigapixel size of WSIs necessitates the use of sliding window methodology for nucleus detection. However, mainstream methods process each sliding window independently, which overlooks broader contextual information and easily leads to inaccurate predictions. To address this limitation, recent studies additionally crop a large Filed-of-View (LFoV) patch centered on each sliding window to extract contextual features. However, such methods substantially increase whole-slide inference latency. In this work, we propose an effective and efficient context-aware nucleus detection approach. Specifically, instead of using lFoV patches, we aggregate contextual clues from off-the-shelf features of historically visited sliding windows, which greatly enhances the inference efficiency. Moreover, compared to lFoV patches used in previous works, the sliding window patches have higher magnification and provide finer-grained tissue details, thereby enhancing the classification accuracy. To develop the proposed context-aware model, we utilize annotated patches along with their surrounding unlabeled patches for training. Beyond exploiting high-level tissue context from these surrounding regions, we design a post-training strategy that leverages abundant unlabeled nucleus samples within them to enhance the model's context adaptability. Extensive experimental results on three challenging benchmarks demonstrate the superiority of our method. Zhongyi Shui, Honglin Li 0001, Yuxuan Sun 0002, Yiwen Ye, Pingyi Chen, Ruizhe Guo, Lei Cui 0004, Chenglu Zhu, Lin Yang 0002 |
AAAI | 4 |
| 2025 | MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkabstractXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, Graham Neubig. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang 0019, Kai Zhang 0033, Shengbang Tong, Yuxuan Sun 0002, Botao Yu, Ge Zhang 0009, Huan Sun 0001, Yu Su 0001, Wenhu Chen, Graham Neubig |
ACL (1) | 7 |
| 2025 | CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational PathologyabstractThe emergence of large multimodal models (LMMs) has brought significant advancements to pathology. Previous research has primarily focused on separately training patch-level and whole-slide image (WSI)-level models, limiting the integration of learned knowledge across patches and WSIs and resulting in redundant models. In this work, we introduce CPath-Omni, the first 15B parameter LMM that unifies patch and WSI analysis, consolidating a variety of tasks at both levels, including classification, visual question answering, captioning, and visual referring prompting. Extensive experiments demonstrate that CPath-Omni achieves state-of-the-art (SOTA) performance across seven diverse tasks on 39 out of 42 datasets, outperforming or matching task-specific models trained for individual tasks. Additionally, we develop a specialized pathology CLIP-based visual processor for CPath-Omni, CPath-CLIP, which, for the first time, integrates different vision models and incorporates a large language model as a text encoder to build a more powerful CLIP model, which achieves SOTA performance on nine zero-shot and four few-shot datasets. Our findings highlight CPath-Omni’s ability to unify diverse pathology tasks, demonstrating its potential to streamline and advance the field of foundation model in pathology. The code and model are available at CPath-Omni. Yuxuan Sun 0002, Yixuan Si, Chenglu Zhu, Kai Zhang 0033, Pingyi Chen, Zhongyi Shui, Tao Lin 0004, Lin Yang 0002 |
CVPR | 1 |
| 2025 | Stable Test-Time Training for Semantic Segmentation with Output Contrastive LossabstractDeep learning-based models have achieved impressive performance on public segmentation benchmarks, yet generalizing to unseen environments remains challenging. Test-time training (TTT) addresses this by adapting source-pretrained models during evaluation. While existing TTT methods have shown promise in image classification, they often exhibit instability with small test batches and class imbalance—challenges that intensify in semantic segmentation tasks. To tackle this issue, we present Output Contrastive Loss (OCL) to improve the stability of contrastive loss when applied to TTT for segmentation. OCL applies contrastive loss directly to the output space, avoiding the need for extra regularization, and employs a high temperature to prevent model collapse. To further stabilize the TTT process, we integrate BN statistics Modulation and Stochastic Restoration techniques. Extensive experiments across diverse datasets, settings, architectures, and pretrained methods demonstrate consistent performance improvements, achieving a 7.5 mIoU gain on the GTA→CS benchmark and showing effectiveness even with domain adaptation pretraining. Code is available at https://github.com/dazhangyul23/OCL. Zhongyi Shui, Honglin Li 0001, Yuxuan Sun 0002, Chenglu Zhu, Lin Yang 0002 |
ICASSP | 4 |
| 2025 | PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent CollaborationabstractVision Language Models (VLMs) like CLIP have attracted substantial attention in pathology, serving as backbones for applications such as zero-shot image classification and Whole Slide Image (WSI) analysis. Additionally, they can function as vision encoders when combined with large language models (LLMs) to support broader capabilities. Current efforts to train pathology VLMs rely on pathology image-text pairs from platforms like PubMed, YouTube, and Twitter, which provide limited, unscalable data with generally suboptimal image quality. In this work, we leverage large-scale WSI datasets like TCGA to extract numerous high-quality image patches. We then train a large multimodal model (LMM) to generate captions for extracted images, creating PathGen-1.6M, a dataset containing 1.6 million high-quality image-caption pairs. Our approach involves multiple agent models collaborating to extract representative WSI patches, generating and refining captions to obtain high-quality image-text pairs. Extensive experiments show that integrating these generated pairs with existing datasets to train a pathology-specific CLIP model, PathGen-CLIP, significantly enhances its ability to analyze pathological images, with substantial improvements across nine pathology-related zero-shot image classification tasks and three whole-slide image tasks. Furthermore, we construct 200K instruction-tuning data based on PathGen-1.6M and integrate PathGen-CLIP with the Vicuna LLM to create more powerful multimodal models through instruction tuning. Overall, we provide a scalable pathway for high-quality data generation in pathology, paving the way for next-generation general pathology models. Our dataset, code, and model are open-access at https://github.com/PathFoundation/PathGen-1.6M. Yuxuan Sun 0002, Yixuan Si, Chenglu Zhu, Kai Zhang 0033, Zhongyi Shui, Jingxiong Li, Xinheng Lyu, Tao Lin 0004, Lin Yang 0002 |
ICLR | 1 |
| 2025 | AAAR-1.0: Assessing AI's Potential to Assist ResearchabstractNumerous studies have assessed the proficiency of AI systems, particularly large language models (LLMs), in facilitating everyday tasks such as email writing, question answering, and creative content generation. However, researchers face unique challenges and opportunities in leveraging LLMs for their own work, such as brainstorming research ideas, designing experiments, and writing or reviewing papers. In this study, we introduce AAAR-1.0, a benchmark dataset designed to evaluate LLM performance in three fundamental, expertise-intensive research tasks: (i) EquationInference, assessing the correctness of equations based on the contextual information in paper submissions; (ii) ExperimentDesign, designing experiments to validate research ideas and solutions; and (iii) PaperWeakness, identifying weaknesses in paper submissions. AAAR-1.0 differs from prior benchmarks in two key ways: first, it is explicitly research-oriented, with tasks requiring deep domain expertise; second, it is researcher-oriented, mirroring the primary activities that researchers engage in on a daily basis. An evaluation of both open-source and proprietary LLMs reveals their potential as well as limitations in conducting sophisticated research tasks. We will release the AAAR-1.0 and keep iterating it to new versions. Renze Lou, Hanzi Xu, Jiangshu Du, Ryo Kamoi, Xiaoxin Lu, Yuxuan Sun 0002, Yusen Zhang 0001, Jihyun Janice Ahn, Hongchao Fang, Zhuoyang Zou, Kai Zhang 0033, Congying Xia, Lifu Huang, Wenpeng Yin 0001 |
ICML | 8 |
| 2025 | AEM: Attention Entropy Maximization for Multiple Instance Learning Based Whole Slide Image Classification
Honglin Li 0001, Yuxuan Sun 0002, Zhongyi Shui, Jingxiong Li, Chenglu Zhu, Lin Yang 0002 |
MICCAI (7) | 3 |
| 2025 | CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic LogicabstractRecent advances in computational pathology have led to the emergence of numerous foundation models. These models typically rely on general-purpose encoders with multi-instance learning for whole slide image (WSI) classification or apply multimodal approaches to generate reports directly from images. However, these models cannot emulate the diagnostic approach of pathologists, who systematically examine slides at low magnification to obtain an overview before progressively zooming in on suspicious regions to formulate comprehensive diagnoses. Instead, existing models directly output final diagnoses without revealing the underlying reasoning process.
To address this gap, we introduce CPathAgent, an innovative agent-based approach that mimics pathologists' diagnostic workflow by autonomously navigating across WSI through zoom-in/out and move operations based on observed visual features, thereby generating substantially more transparent and interpretable diagnostic summaries. To achieve this, we develop a multi-stage training strategy that unifies patch-level, region-level, and WSI-level capabilities within a single model, which is essential for replicating how pathologists understand and reason across diverse image scales.
Additionally, we construct PathMMU-HR², the first expert-validated benchmark for large region analysis. This represents a critical intermediate scale between patches and whole slides, reflecting a key clinical reality where pathologists typically examine several key large regions rather than entire slides at once. Extensive experiments demonstrate that CPathAgent consistently outperforms existing approaches across benchmarks at three different image scales, validating the effectiveness of our agent-based diagnostic approach and highlighting a promising direction for computational pathology. Yuxuan Sun 0002, Yixuan Si, Chenglu Zhu, Kai Zhang 0033, Zhongyi Shui, Tao Lin 0004, Lin Yang 0002 |
NeurIPS | 1 |
| 2025 | ToPoFM: Topology-Guided Pathology Foundation Model for High-Resolution Pathology Image Synthesis With Cellular-Level ControlabstractSynthetic data generation emerges as a strategy to mitigate data scarcity in digital pathology, where complicated tissue and cellular features are correlated with cancer diagnosis. The synthesis of such visuals, however, suffers from limited inter class diversity and scarcity of cellular annotations. Current methodologies struggle with capturing the broad spectrum of pathology features, causing unpredictable objects and defected fidelity. Moreover, discrepancies in image resolution across developmental and operational phases can amplify the distribution shifts, undermining the precision of diagnosis. To address these challenges, we introduce TOpology guided PathOlogy Foundation Model (ToPoFM), a visual foundation model designed for the synthesis of high-resolution pathology images with cellular-level control. Our approach integrates a topology-informed cell arrangement generator to steer large language models for crafting synthetic cell arrangements. We correlate cell arrangement guidance with diffusion model for pathology content generation, then further implement a random sliding inference strategy, merging discrete low-resolution samplings into single high-resolution representation. Our model requires only small patches for training. The efficacy of ToPoFM is demonstrated through extensive experiments, complemented by expert validations, showing high fidelity on data synthesis. Additionally, we underscore the utility of our generated imagery as an augmentation tool, enhancing the performance of downstream tasks, including cancer subtype classification and segmentation. Jingxiong Li, Chenglu Zhu, Sunyi Zheng, Pingyi Chen, Yuxuan Sun 0002, Honglin Li 0001, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | PathBench: Advancing the Benchmark of Large Multimodal Models for Pathology Image Understanding at Patch and Whole Slide LevelabstractRapid advancements in large multimodal models (LMMs) have significantly enhanced their applications in pathology, particularly in image classification, pathology image description, and whole slide image (WSI) classification. In pathology, WSIs represent gigapixel-scale images composed of thousands of image patches. Therefore, both patch-level and WSI-level evaluations are essential and inherently interconnected for assessing LMM capabilities. In this work, we propose PathBench, which comprises three subsets at both patch and WSI levels, to refine and enhance the validation of LMMs. At the patch-level, evaluations using existing multi-choice Q&A datasets reveal that some LMMs can predict answers without genuine image analysis. To address this, we introduce PatchVQA, a large-scale visual question answering (VQA) dataset containing 5,382 images and 6,335 multiple-choice questions designed with distractor options to prevent shortcut learning. These new questions are rigorously validated by professional pathologists to ensure reliable model assessments. At the WSI-level, current efforts primarily focus on image classification tasks and lack diverse validation datasets for multimodal models. To address this, we generate a detailed WSI report dataset through an innovative approach that integrates detailed patch descriptions generated by foundational models into comprehensive WSI reports. These are then combined with physician-written reports corresponding to TCGA WSIs, resulting in WSICap, a detailed report dataset containing 7,000 samples. Based on WSICap, we further develop a WSI-level VQA dataset, WSIVQA, to serve as a validation set for WSI LMMs. Using these PathBench subsets, we conduct extensive experiments to benchmark the performance of state-of-the-art LMMs at both the patch and WSI levels. The proposed dataset is available at https://github.com/superjamessyx/PathBench. Yuxuan Sun 0002, Hao Wu 0072, Chenglu Zhu, Yixuan Si, Qizi Chen, Kai Zhang 0033, Jingxiong Li, Jiatong Cai, Lin Sun 0006, Tao Lin 0004, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 1 |
| 2024 | PathAsst: A Generative Foundation AI Assistant towards Artificial General Intelligence of PathologyabstractAs advances in large language models (LLMs) and multimodal techniques continue to mature, the development of general-purpose multimodal large language models (MLLMs) has surged, offering significant applications in interpreting natural images. However, the field of pathology has largely remained untapped, particularly in gathering high-quality data and designing comprehensive model frameworks. To bridge the gap in pathology MLLMs, we present PathAsst, a multimodal generative foundation AI assistant to revolutionize diagnostic and predictive analytics in pathology. The development of PathAsst involves three pivotal steps: data acquisition, CLIP model adaptation, and the training of PathAsst's multimodal generative capabilities. Firstly, we collect over 207K high-quality pathology image-text pairs from authoritative sources. Leveraging the advanced power of ChatGPT, we generate over 180K instruction-following samples. Furthermore, we devise additional instruction-following data specifically tailored for invoking eight pathology-specific sub-models we prepared, allowing the PathAsst to effectively collaborate with these models, enhancing its diagnostic ability. Secondly, by leveraging the collected data, we construct PathCLIP, a pathology-dedicated CLIP, to enhance PathAsst's capabilities in interpreting pathology images. Finally, we integrate PathCLIP with the Vicuna-13b and utilize pathology-specific instruction-tuning data to enhance the multimodal generation capacity of PathAsst and bolster its synergistic interactions with sub-models. The experimental results of PathAsst show the potential of harnessing AI-powered generative foundation model to improve pathology diagnosis and treatment processes. We open-source our dataset, as well as a comprehensive toolkit for extensive pathology data collection and preprocessing at https://github.com/superjamessyx/Generative-Foundation-AI-Assistant-for-Pathology. Yuxuan Sun 0002, Chenglu Zhu, Sunyi Zheng, Kai Zhang 0033, Lin Sun 0006, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002 |
AAAI | 1 |
| 2024 | MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIabstractWe introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and text-books, covering six core disciplines: Art & Design, Busi-ness, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly het-erogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. The evaluation of 28 open-source LMMs as well as the propri-etary GPT-4V(ision) and Gemini highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V and Gemini Ultra only achieve accuracies of 56% and 59% respectively, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence. Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 0033, Ruoqi Liu, Ge Zhang 0009, Samuel Stevens 0001, Dongfu Jiang, Weiming Ren, Yuxuan Sun 0002, Cong Wei 0001, Botao Yu, Ruibin Yuan, Renliang Sun, Boyuan Zheng 0001, Zhenzhu Yang, Wenhao Huang 0001, Huan Sun 0001, Yu Su 0001, Wenhu Chen |
CVPR | 10 |
| 2024 | Unleashing the Power of Prompt-Driven Nucleus Instance Segmentation
Zhongyi Shui, Chenglu Zhu, Sunyi Zheng, Jingxiong Li, Honglin Li 0001, Yuxuan Sun 0002, Ruizhe Guo, Lin Yang 0002 |
ECCV (27) | 8 |
| 2024 | PathMMU: A Massive Multimodal Expert-Level Benchmark for Understanding and Reasoning in Pathology
Yuxuan Sun 0002, Hao Wu 0072, Chenglu Zhu, Sunyi Zheng, Qizi Chen, Kai Zhang 0033, Dan Wan, Xiaoxiao Lan, Mengyue Zheng, Jingxiong Li, Xinheng Lyu, Tao Lin 0004, Lin Yang 0002 |
ECCV (62) | 1 |
| 2024 | MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction FollowingabstractIn the realm of large language models (LLMs), enhancing instruction-following capability often involves curating expansive training data. This is achieved through two primary schemes: i) Scaling-Inputs: Amplifying (input, output) pairs per task instruction, aiming for better instruction adherence. ii) Scaling Input-Free Tasks: Enlarging tasks, each composed of an (instruction, output) pair (without requiring a separate input anymore). However, LLMs under Scaling-Inputs tend to be overly sensitive to inputs, leading to misinterpretation or non-compliance with instructions. Conversely, Scaling Input-Free Tasks demands a substantial number of tasks but is less effective in instruction following when dealing with instances in Scaling-Inputs. This work introduces MUFFIN, a new scheme of instruction-following dataset curation. Specifically, we automatically Scale Tasks per Input by diversifying these tasks with various input facets. Experimental results across four zero-shot benchmarks, spanning both Scaling-Inputs and Scaling Input-Free Tasks schemes, reveal that LLMs, at various scales, trained on MUFFIN generally demonstrate superior instruction-following capabilities compared to those trained on the two aforementioned schemes. Renze Lou, Kai Zhang 0033, Yuxuan Sun 0002, Jihyun Janice Ahn, Hanzi Xu, Yu Su 0001, Wenpeng Yin 0001 |
ICLR | 4 |
| 2024 | Context-Aware Text-Assisted Multimodal Framework for Cervical Cytology Cell Diagnosis and ChattingabstractRecent advancements underscore the potential of deep learning-based Computer-Assisted Diagnosis (CAD) systems for cervical cytology image analysis. However, traditional methods focusing solely on single-view of cells fall short in performance due to the lack of contextual information. Moreover, the unclear reasoning behind model’s classification hinders their interpretability. To overcome these issues, we present Cervi-CAT, a context-aware, text-assisted multimodal framework for cervical cytology cell classification. CerviCAT captures visual cell representations from both global and local perspectives and subsequently generates textual descriptions based on the visual representation. A multimodal transformer then integrates these descriptions with visual features for interpretable and accurate cell classification. Additionally, we introduce Cyto-Vicuna, a cytology-specific large language model fine-tuned based on Vicuna-7b using collected cytology-specific data. When integrated into CerviCAT, it produces more detailed diagnostic reports while simultaneously fostering interaction between the model and cytologists, promoting collaborative diagnosis. Our results demonstrate that CerviCAT not only surpasses traditional CAD methods in performance but also provides interpretable diagnosis. Yuxuan Sun 0002, Chenglu Zhu, Sunyi Zheng, Honglin Li 0001, Lin Yang 0002 |
ICME | 1 |
| 2024 | PathUp: Patch-wise Timestep Tracking for Multi-class Large Pathology Image Synthesising Diffusion ModelabstractIn digital pathology, cancer lesions are identified by analyzing the spatial context within pathology images. Synthesizing such complex spatial context is challenging as pathology whole slide images typically exhibit high resolution, low inter-class variety, and are sparsely labeled. To address these challenges, we propose PathUp, a novel diffusion model tailored for the synthesis of multi-class high-resolution pathology images. Our approach includes a latent space patch-wise timestep tracking, which helps to generate high-quality images without tiling artifacts. Pathology knowledge is integrated through our patho-align. The robust generation of lesion subtypes and scale information is ensured by introducing a feature entropy loss. The effectiveness of our method is evaluated through extensive experiments, supplemented by assessments from human experts, demonstrating the authenticity of the synthetic data produced. Furthermore, we highlight the potential utility of our generated images as an augmentation method, thereby enhancing the performance of downstream tasks such as cancer subtype classification. Jingxiong Li, Sunyi Zheng, Chenglu Zhu, Yuxuan Sun 0002, Pingyi Chen, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002 |
ACM Multimedia | 4 |
| 2024 | Masked Conditional Variational Autoencoders for Chromosome StraighteningabstractKaryotyping is of importance for detecting chromosomal aberrations in human disease. However, chromosomes easily appear curved in microscopic images, which prevents cytogeneticists from analyzing chromosome types. To address this issue, we propose a framework for chromosome straightening, which comprises a preliminary processing algorithm and a generative model called masked conditional variational autoencoders (MC-VAE). The processing method utilizes patch rearrangement to address the difficulty in erasing low degrees of curvature, providing reasonable preliminary results for the MC-VAE. The MC-VAE further straightens the results by leveraging chromosome patches conditioned on their curvatures to learn the mapping between banding patterns and conditions. During model training, we apply a masking strategy with a high masking ratio to train the MC-VAE with eliminated redundancy. This yields a non-trivial reconstruction task, allowing the model to effectively preserve chromosome banding patterns and structure details in the reconstructed results. Extensive experiments on three public datasets with two stain styles show that our framework surpasses the performance of state-of-the-art methods in retaining banding patterns and structure details. Compared to using real-world bent chromosomes, the use of high-quality straightened chromosomes generated by our proposed method can improve the performance of various deep learning models for chromosome classification by a large margin. Such a straightening approach has the potential to be combined with other karyotyping systems to assist cytogeneticists in chromosome analysis. Jingxiong Li, Sunyi Zheng, Zhongyi Shui, Shichuan Zhang, Linyi Yang, Yuxuan Sun 0002, Honglin Li 0001, Yuanxin Ye, Peter M. A. van Ooijen, Kang Li 0004, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 6 |
| 2023 | Task-Specific Fine-Tuning via Variational Information Bottleneck for Weakly-Supervised Pathology Whole Slide Image ClassificationabstractWhile Multiple Instance Learning (MIL) has shown promising results in digital Pathology Whole Slide Image (WSI) analysis, such a paradigm still faces performance and generalization problems due to high computational costs and limited supervision of Gigapixel WSIs. To deal with the computation problem, previous methods utilize a frozen model pretrained from ImageNet to obtain representations, however, it may lose key information owing to the large domain gap and hinder the generalization ability without image-level training-time augmentation. Though Self-supervised Learning (SSL) proposes viable representation learning schemes, the downstream task-specific features via partial label tuning are not explored. To alleviate this problem, we propose an efficient WSI fine-tuning framework motivated by the Information Bottleneck theory. The theory enables the framework to find the minimal sufficient statistics of WSI, thus supporting us to fine-tune the backbone into a task-specific representation only depending on WSI-level weak labels. The WSI-MIL problem is further analyzed to theoretically deduce our fine-tuning method. We evaluate the method on five pathological WSI datasets on various WSI heads. The experimental results show significant improvements in both accuracy and generalization compared with previous works. Source code will be available at https://github.com/invoker-LL/WSI-finetuning. Honglin Li 0001, Chenglu Zhu, Yuxuan Sun 0002, Zhongyi Shui, Wenwei Kuang, Sunyi Zheng, Lin Yang 0002 |
CVPR | 4 |
| 2023 | Assessing the Robustness of Deep Learning-Assisted Pathological Image Analysis Under Practical Variables of Imaging SystemabstractWith the advancement of deep learning, computer-assisted clinical diagnosis, such as liquid-based cervical cytology, has attracted more attention. However, the fragile robustness of deep learning models has a non-negligible impact on their classification accuracy and reliability. To be more specific, various scanner parameters will be used depending on the pathologist’s preferences during the clinical diagnosis process (e.g., field source brightness, contrast, saturation, etc.), and this variation will lead to the unstable performance of the model. In this paper, we construct an evaluation pathway to assess the stability and consistency of deep learning models under various customized scanner parameters. Specifically, a multi-scanned dataset consists of 4200 whole slide images (WSIs) is generated by scanning 200 stained slices using various scanner parameters. Moreover, we conducted a large number of experiments to investigate the robustness of numerous models, including convolution-based and transformer-based models concerning various scanner parameter settings. Furthermore, we introduce several indicators to analyze the prediction accuracy, consistency and robustness of the model on the constructed dataset. The experimental results indicate that the deep learning models are sensitive to luminance-related scanner parameters. In addition, transformer-based models have better robustness than traditional convolutional neural networks. Our code has been made available1. Yuxuan Sun 0002, Chenglu Zhu, Honglin Li 0001, Pingyi Chen, Lin Yang 0002 |
ICASSP | 1 |
| 2022 | Benchmarking the Robustness of Deep Neural Networks to Common Corruptions in Digital Pathology
Yuxuan Sun 0002, Honglin Li 0001, Sunyi Zheng, Chenglu Zhu, Lin Yang 0002 |
MICCAI (2) | 2 |
| 2020 | RIVA: A Pre-trained Tweet Multimodal Model Based on Text-image Relation for Multimodal NERabstractMultimodal named entity recognition (MNER) for tweets has received increasing attention recently.Most of the multimodal methods used attention mechanisms to capture the text-related visual information.However, unrelated or weakly related text-image pairs account for a large proportion in tweets.Visual clues unrelated to the text would incur uncertain or even negative effects for multimodal model learning.In this paper, we propose a novel pre-trained multimodal model based on Relationship Inference and Visual Attention (RIVA) for tweets.The RIVA model controls the attention-based visual clues with a gate regarding the role of image to the semantics of text.We use a teacher-student semi-supervised paradigm to leverage a large unlabeled multimodal tweet corpus with a labeled data set for text-image relation classification.In the multimodal NER task, the experimental results show the significance of text-related visual features for the visual-linguistic model and our approach achieves SOTA performance on the MNER datasets. Lin Sun 0006, Jiquan Wang, Yindu Su, Fangsheng Weng, Yuxuan Sun 0002, Zengwei Zheng |
COLING | 5 |
| 2020 | Attention-based Deep Learning Model for Text Readability EvaluationabstractText readability is useful to measure the comprehensibility of a piece of writing and plays an important role in the field of education. The classic evaluation methods are linear functions with two or three parameters. However, it needs laborious human tests to find the appropriate parameters. This paper presents a recurrent neural network to quickly build a text readability evaluation model. The model consists of bi-directional gated recurrent unit (bi-GRU) and attention layer. It can predict the readability level of the sentences or paragraphs. The weight of the attention layer can finely locate the distribution of difficulty in the comprehension of the sentences. We train the model on rough leveled reading materials without tremendous reading comprehension tests. Our experiments are performed on various text materials. Compared with the popular readability formulas, the results show good performance on readability measurement and visualization. Yuxuan Sun 0002, Keying Chen, Lin Sun 0006, Chenlu Hu |
IJCNN | 1 |
| 2020 | ETIP: a lengthy nested NER problem for Chinese insurance policy analysis
Lin Sun 0006, Kai Zhang 0033, Yuxuan Sun 0002, Fangsheng Weng |
Pattern Anal. Appl. | 3 |
| 2020 | Joint Learning of Token Context and Span Feature for Span-Based Nested NERabstractNested named entity recognition (NER) is a linguistic phenomenon that has received increasing attention in the NLP community. In this article, we propose a novel joint learning network of the token context, and span feature (TCSF) for nested NER. Our model is a combination of token context network (TCN) for token learning, and deep residual convolutional neural network (CNN), and span relation network (SRN) for span learning. Span features are represented at both the token level, and span relation level. The span relation representation of SRN is trained on the similarity of span, and its positional features by attention weights. The IoU metric is employed for span filtering, and analysis of the overlapping adjacent spans. Moreover, we propose a novel head-inner-tail (HIP T) operation to extend the inner features in spans for a fixed-length representation. TCSF is a fully end-to-end entity recognition model. We perform a set of comprehensive experiments, including an ablation study on the ACE, and GENIA datasets, and show state-of-the-art performance without, and with the pretrained language models. Lin Sun 0006, Yuxuan Sun 0002, Fule Ji |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |