Miao Li 0003

dblp:39/9-3 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
12since 2021 · last 2025
0009-0009-5672-7448ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Enhancing Elusive Clues in Knowledge Learning by Contrasting Attention of Language Models
abstract
Causal language models acquire vast amount of knowledge from general text corpus during pretraining, but the efficiency of knowledge learning is known to be unsatisfactory, especially when learning from knowledge-dense and small-sized corpora. The deficiency can come from long-distance dependencies which are hard to capture by language models, and overfitting to co-occurrence patterns and distracting clues in the training text. To address these issues, the paper proposes a method to enhance knowledge learning during language model pretraining, by enhancing elusive but important clues in text discovered by the language model themselves. We found that larger language models pay more attention to non-obvious but important clues, which are often overlooked by smaller language models. Therefore, we can identify these clues by contrasting the attention weights of large and small language models. We use the identified clues as a guide to perform token-dropout data augmentation on the training text, and observed a significant boost in both small and large models' performance in fact memorization. This shows that the behavior contrast between more and less-performant language models contains important clues for knowledge learning, and it can be "amplified" for a straight-forward improvement in knowledge learning efficiency.
Xiao Zhang 0001, Miao Li 0003, Ji Wu 0002
AAAI3
2025 Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model
abstract
3D medical image analysis is essential for modern healthcare, yet traditional task-specific models are inadequate due to limited generalizability across diverse clinical scenarios. Multimodal large language models (MLLMs) offer a promising solution to these challenges. However, existing MLLMs have limitations in fully leveraging the rich, hierarchical information embedded in 3D medical images. Inspired by clinical practice, where radiologists focus on both 3D spatial structure and 2D planar content, we propose Med-2E3, a 3D medical MLLM that integrates a dual 3D-2D encoder architecture. To aggregate 2D features effectively, we design a Text-Guided Inter-Slice (TG-IS) scoring module, which scores the attention of each 2D slice based on slice contents and task instructions. To the best of our knowledge, Med-2E3 is the first MLLM to integrate both 3D and 2D features for 3D medical image analysis. Experiments on large-scale, open-source 3D medical multimodal datasets demonstrate that TG- IS exhibits task-specific attention distribution and sig-nificantly outperforms current state-of-the-art models. The code is available at: https://github.com/MSIIPlMed-2E3
Yiming Shi, Chenyi Guo, Miao Li 0003, Ji Wu 0002
BIBM6
2025 3D-HSPA: Integrating 3D Spatial Information with Hierarchical Slice-Patch Attention for Knee MRI Analysis
abstract
Magnetic Resonance Imaging (MRI) is a crucial modality for diagnosing knee joint diseases. However, accurately extracting disease-relevant features from complex multi-slice, multi-sequence MRI scans remains a considerable challenge. To address this, we propose 3D-HSPA, a novel diagnostic framework for multi-slice, multi-sequence knee MRI, which integrates disease-specific information at both slice and patch levels and establishes intrinsic spatial connections among different sequences. Specifically, we introduce a patch-level and slice-level label attention mechanism, guiding the model to automatically learn a precise alignment between image regions and disease labels. Furthermore, by mapping 2D images from various sequences into a unified 3D spatial coordinate system, we enhance the spatial consistency and robustness of the attention distributions. We validated 3D-HSPA on a large-scale MRI dataset comprising 50 fine-grained types of knee joint diseases. The experimental results demonstrate that 3D-HSPA not only achieves superior diagnostic performance but also exhibits strong model interpretability.
Jingzhi Yang, Yiming Shi, Ji Wu 0002, Huishu Yuan, Miao Li 0003, Xiangling Fu
BIBM8
2025 Connector-S: A Survey of Connectors in Multi-modal Large Language Models
abstract
With the rapid advancements in multi-modal large language models (MLLMs), connectors play a pivotal role in bridging diverse modalities and enhancing model performance. However, the design and evolution of connectors have not been comprehensively analyzed, leaving gaps in understanding how these components function and hindering the development of more powerful connectors. In this survey, we systematically review the current progress of connectors in MLLMs and present a structured taxonomy that categorizes connectors into atomic operations (mapping, compression, mixture of experts) and holistic designs (multi-layer, multi-encoder, multi-modal scenarios), highlighting their technical contributions and advancements. Furthermore, we discuss several promising research frontiers and challenges, including high-resolution input, dynamic compression, guide information selection, combination strategy, and interpretability. This survey is intended to serve as a foundational reference and a clear roadmap for researchers, providing valuable insights into the design and optimization of next-generation connectors to enhance the performance and adaptability of MLLMs.
Xi Chen 0009, Yiming Shi, Miao Li 0003, Ji Wu 0002
IJCAI5
2025 Medical Contrastive Learning of Positive and Negative Mentions
WeiLong Wu, Jingzhi Yang, Xiao Zhang 0001, ZiYu Liu, Miao Li 0003, Ji Wu 0002
MICCAI (11)6
2025 MIPS: A Multimodal Infinite Polymer Sequence Pre-training Framework for Polymer Property Prediction
abstract
Polymers, composed of repeating structural units called monomers, are fundamental materials with a wide range of applications in daily life and industry. Accurate property prediction for polymers is essential for their design, development, and application. However, existing modeling approaches, which typically represent polymers by the constituent monomers, struggle to capture the whole properties of polymer, since the properties change during the polymerization process. In this study, we propose a Multimodal Infinite Polymer Sequence (MIPS) pre-training framework, which represents polymers as infinite sequences of monomers and integrates both topological and spatial information for comprehensive modeling. From the topological perspective, we generalize message passing mechanism (MPM) and graph attention mechanism (GAM) to infinite polymer sequences. For MPM, we demonstrate that applying MPM to infinite polymer sequences is equivalent to applying MPM on the induced star-linking graph of monomers. For GAM, we propose to further replace global graph attention with localized graph attention (LGA). Moreover, we show the robustness of the ''star linking'' strategy through an adversarial evaluation method named Repeat and Shift Invariance Test (RSIT). Despite its robustness, ''star linking'' strategy exhibits limitations when monomer side chains contain ring structures, a common characteristic of polymers, as it fails the Weisfeiler-Lehman (WL) test. To overcome this issue, we propose backbone embedding to enhance the capability of MPM and LGA on infinite polymer sequences. From the spatial perspective, we extract 3D descriptors of repeating monomers to capture spatial information. Finally, we design a cross-modal fusion mechanism to unify the topological and spatial information. Experimental validation across eight diverse polymer property prediction tasks reveals that MIPS achieves state-of-the-art performance. Ablation studies further comfirm the efficacy of our infinite polymer sequence modeling approach and multimodal pre-training framework.
Yaosen Min, Miao Li 0003, Ji Wu 0002
ACM Multimedia4
2025 Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation Data
abstract
The emergence of medical generalist foundation models has revolutionized conventional task-specific model development paradigms, aiming to better handle multiple tasks through joint training on large-scale medical datasets. However, recent advances prioritize simple data scaling or architectural component enhancement, while neglecting to re-examine multi-task learning from a data-centric perspective. Critically, simply aggregating existing data resources leads to decentralized image-task alignment, which fails to cultivate comprehensive image understanding or align with clinical needs for multi-dimensional image interpretation. In this paper, we introduce the image-centric multi-annotation X-ray dataset (IMAX), the first attempt to enhance the multi-task learning capabilities of medical multi-modal large language models (MLLMs) from the data construction level. To be specific, IMAX is featured from the following attributes: 1) High-quality data curation. A comprehensive collection of more than 354K entries applicable to seven different medical tasks. 2) Image-centric dense annotation. Each X-ray image is associated with an average of 4.10 tasks and 7.46 training entries, ensuring multi-task representation richness per image. Compared to the general decentralized multi-annotation X-ray dataset (DMAX), IMAX consistently demonstrates significant multi-task average performance gains ranging from 3.20% to 21.05% across seven open-source state-of-the-art medical MLLMs. Moreover, we investigate differences in statistical patterns exhibited by IMAX and DMAX training processes, exploring potential correlations between optimization dynamics and multi-task performance. Finally, leveraging the core concept of IMAX data construction, we propose an optimized DMAX-based training strategy to alleviate the dilemma of obtaining high-quality IMAX data in practical scenarios. Related resources are available at https://github.com/MSIIP/IMAX.
Fanbin Mo, Yiming Shi, Ming Wu 0001, Miao Li 0003, Ji Wu 0002
ACM Multimedia8
2024 Slice-Level Label Attention with Global-Guided Attention Regularization for Multi-Label Classification in Knee MRI Sequences
abstract
Magnetic Resonance Imaging (MRI) is crucial for diagnosing various knee-related diseases, and developing automatic diagnostic models based on knee MRI data is highly valuable. However, this task presents significant challenges due to the need to manage MRI data with multiple sequences and numerous images, where different diseases are often associated with specific images within certain sequences. To address these challenges, we propose a multi-label classification framework designed to effectively process MRI data and handle a large-scale label space encompassing hundreds of disease categories. Our approach introduces a Slice-Level Label Attention mechanism, which enables the model to learn the alignment between labels and images within sequences, thereby enhancing both performance and interpretability. Additionally, we present a Global-Guided Attention Regularization mechanism that further improves the consistency and robustness of the Slice-Level Label Attention results. We validate our framework on a large-scale MRI dataset involving multi-label classification across hundreds of fine-grained disease categories. Experimental results demonstrate that our method not only achieves superior performance but also provides more robust and consistent interpretability.
Jingzhi Yang, Weilong Wu, Ji Wu 0002, Huishu Yuan, Xiangling Fu, Miao Li 0003
IEEE Big Data9
2024 Conditional Language Learning with Context
abstract
Language models can learn sophisticated language understanding skills from fitting raw text. They also unselectively learn useless corpus statistics and biases, especially during finetuning on domain-specific corpora. In this paper, we propose a simple modification to causal language modeling called conditional finetuning, which performs language modeling conditioned on a context. We show that a context can "explain away" certain corpus statistics and make the model avoid learning them. In this fashion, conditional finetuning achieves selective learning from a corpus, learning knowledge useful for downstream tasks while avoiding learning useless corpus statistics like topic biases. This selective learning effect leads to less forgetting and better stability-plasticity tradeoff in domain finetuning, potentially benefitting lifelong learning with language models.
Xiao Zhang 0001, Miao Li 0003, Ji Wu 0002
ICML2
2024 Co-occurrence is not Factual Association in Language Models
abstract
Pretrained language models can encode a large amount of knowledge and utilize it for various reasoning tasks, yet they can still struggle to learn novel factual knowledge effectively from finetuning on limited textual demonstrations. In this work, we show that the reason for this deficiency is that language models are biased to learn word co-occurrence statistics instead of true factual associations. We identify the differences between two forms of knowledge representation in language models: knowledge in the form of co-occurrence statistics is encoded in the middle layers of the transformer model and does not generalize well to reasoning scenarios beyond simple question answering, while true factual associations are encoded in the lower layers and can be freely utilized in various reasoning tasks. Based on these observations, we propose two strategies to improve the learning of factual associations in language models. We show that training on text with implicit rather than explicit factual associations can force the model to learn factual associations instead of co-occurrence statistics, significantly improving the generalization of newly learned knowledge. We also propose a simple training method to actively forget the learned co-occurrence statistics, which unblocks and enhances the learning of factual associations when training on plain narrative text. On both synthetic and real-world corpora, the two proposed strategies improve the generalization of the knowledge learned during finetuning to reasoning scenarios such as indirect and multi-hop question answering.
Xiao Zhang 0001, Miao Li 0003, Ji Wu 0002
NeurIPS2
2024 Uni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE
abstract
Multi-modal large language models (MLLMs) have shown impressive capabilities as a general-purpose interface for various visual and linguistic tasks. However, building a unified MLLM for multi-task learning in the medical field remains a thorny challenge. To mitigate the tug-of-war problem of multi-modal multi-task optimization in MLLMs, recent advances primarily focus on improving the LLM components, while neglecting the connector that bridges the gap between modalities. In this paper, we introduce Uni-Med, a novel medical generalist foundation model which consists of a universal visual feature extraction module, a connector mixture-of-experts (CMoE) module, and an LLM. Benefiting from the proposed CMoE that leverages a well-designed router with a mixture of projection experts at the connector, Uni-Med achieves efficient solution to the tug-of-war problem and can perform six different medical tasks including question answering, visual question answering, report generation, referring expression comprehension, referring expression generation and image classification. To the best of our knowledge, Uni-Med is the first effort to tackle multi-task interference at the connector in MLLMs. Extensive ablation experiments validate the effectiveness of introducing CMoE under any configuration, with up to an average 8% performance gains. We further provide interpretation analysis of the tug-of-war problem from the perspective of gradient optimization and parameter statistics. Compared to previous state-of-the-art medical MLLMs, Uni-Med achieves competitive or superior evaluation metrics on diverse tasks. Code and resources are available at https://github.com/MSIIP/Uni-Med.
Fanbin Mo, Miao Li 0003, Ji Wu 0002
NeurIPS4
2023 Learning to Generate Radiology Findings from Impressions Based on Large Language Model
abstract
Medical imaging plays a pivotal role in clinical diagnosis, and the textual reports associated with these images are of paramount importance in aiding image comprehension and supporting treatment decisions. Automated report generation serves to alleviate the burden on radiologists and has garnered significant attention in the field of medical artificial intelligence. Previous research in text-based report generation primarily focused on generating impressions statements from radiology findings. However, the benefits in terms of reducing the workload on radiologists were not particularly evident. In this article, we propose a novel task of generating findings from radiology impressions. Leveraging advanced large language models, we trained a set of report generation models using a real dataset of knee MRI reports. Additionally, we incorporated various strategies, including data augmentation and efficient parameter fine-tuning. Objective experiments affirm the effectiveness of the methods we introduced. Furthermore, we conducted subjective assessments by radiologists, and the results demonstrate that our trained large language models significantly outperform professional radiologists in terms of overall report quality and content consistency.
Weilong Wu, Miao Li 0003, Ji Wu 0002, Huishu Yuan
IEEE Big Data2
2016 Improving the Probabilistic Framework for Representing Dialogue Systems with User Response Model
Miao Li 0003, Ji Wu 0002
INTERSPEECH1
2016 Target-Based State and Tracking Algorithm for Spoken Dialogue System
Miao Li 0003, Zhiyang He, Ji Wu 0002
INTERSPEECH1
2016 The MSIIP system for dialog state tracking challenge 5
abstract
We present our work in Dialog State Tracking Challenge 5, the main task of which is to track dialog state on human-human conversations cross language. Firstly a probabilistic enhanced framework is used to represent sub-dialog, which consists of three parts, the input model for extracting features, the enhanced model for updating dialog state and the output model to give the tracking frame. Meanwhile, parallel language systems are proposed to overcome inaccuracy caused by machine translation for cross language testing. We also introduce a new iterative alignment method extended from our work in DSTC4. Furthermore, a slot-based score averaging method is introduced to build an ensemble by combining different trackers. Results of our DSTC5 system show that our method significantly improves tracking performance compared with baseline method.
Miao Li 0003, Ji Wu 0002
SLT2
2015 An entropy minimization framework for goal-driven dialogue management
Ji Wu 0002, Miao Li 0003, Chin-Hui Lee 0001
INTERSPEECH2
2015 A Probabilistic Framework for Representing Dialog Systems and Entropy-Based Dialog Management Through Dynamic Stochastic State Evolution
abstract
In this paper, we present a probabilistic framework for goal-driven spoken dialog systems. A new dynamic stochastic state (DS-state) is then defined to characterize the goal set of a dialog state at different stages of the dialog process. Furthermore, an entropy minimization dialog management (EMDM) strategy is also proposed to combine with the DS-states to facilitate a robust and efficient solution in reaching a user's goals. A song-on-demand task, with a total of 38 117 songs and 12 attributes corresponding to each song, is used to test the performance of the proposed approach. In an ideal simulation, assuming no errors, the EMDM strategy is the most efficient goal-seeking method among all tested approaches, returning the correct song within 3.3 dialog turns on average. Furthermore, in a practical scenario, with top five candidates to handle the unavoidable automatic speech recognition (ASR) and natural language understanding (NLU) errors, the results show that only 61.7% of the dialog goals can be successfully obtained in 6.23 dialog turns on average when random questions are asked by the system, whereas if the proposed DS-states are updated with the top five candidates from the SLU output using the proposed EMDM strategy executed at every DS-state, then a 86.7% dialog success rate can be accomplished effectively within 5.17 dialog turns on average. We also demonstrate that entropy-based DM strategies are more efficient than non-entropy based DM. Moreover, using the goal set distributions in EMDM, the results are better than those without them, such as in sate-of-the-art database summary DM.
Ji Wu 0002, Miao Li 0003, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.2