EDBT 2026 Demo / reviewers in the wild / expert
Xiang Li 0001
dblp:40/1491-1
· DBLP profile ↗
71ranked-venue papers
7as first author
48since 2021 · last 2026
0000-0002-9851-6376ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 42 · 1 first-author · 22 since 2021Artificial intelligence and machine learning · 28 · 5 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task ProjectionabstractRecent advances in parameter-efficient transfer learning have demonstrated the utility of composing LoRA adapters from libraries of pretrained modules. However, most existing approaches rely on simple retrieval heuristics or uniform averaging, which overlook the latent structure of task relationships in representation space. We propose a new framework for adapter reuse that moves beyond retrieval, formulating adapter composition as a geometry-aware sparse reconstruction problem. Specifically, we represent each task by a latent prototype vector derived from the base model’s encoder and aim to approximate the target task prototype as a sparse linear combination of retrieved reference prototypes, under an L1-regularized optimization objective. The resulting combination weights are then used to blend the corresponding LoRA adapters, yielding a composite adapter tailored to the target task. This formulation not only preserves the local geometric structure of the task representation manifold, but also promotes interpretability and efficient reuse by selecting a minimal set of relevant adapters. We demonstrate the effectiveness of our approach across multiple domains—including medical image segmentation, medical report generation and image synthesis. Our results highlight the benefit of coupling retrieval with latent geometry-aware optimization for improved zero-shot generalization. Pengfei Jin, Peng Shu, Sifan Song, Sekeun Kim, Qing Xiao 0003, Cheng Chen 0013, Tianming Liu 0001, Xiang Li 0001, Quanzheng Li |
AAAI | 8 |
| 2026 | GAGM: Geometry-aware graph matching framework for weakly supervised gyral hinge correspondence
Wuyang Li, Tianming Liu 0001, Xiang Li 0001, Junwei Han 0001, Yixuan Yuan |
Medical Image Anal. | 4 |
| 2025 | SearchRAG: Can Search Engines Be Helpful for LLM-Based Medical Question Answering?abstractLarge Language Models (LLMs) have shown remarkable capabilities in general domains but often struggle with tasks requiring specialized knowledge. Conventional Retrieval-Augmented Generation (RAG) techniques typically retrieve external information from static knowledge bases, which can be outdated or incomplete, missing fine-grained clinical details essential for accurate medical question answering. In this work, we propose SearchRAG, a novel framework that overcomes these limitations by leveraging real-time search engines. Our method employs synthetic query generation to convert complex medical questions into search-engine-friendly queries and utilizes uncertainty-based knowledge selection to filter and incorporate the most relevant and informative medical knowledge into the LLM's input. Experimental results demonstrate that our method significantly improves response accuracy in medical question answering tasks, particularly for complex questions requiring detailed and up-to-date knowledge. We provide our code here11https://github.com/sycny/SearchRAG. Tianze Yang, Canyu Chen, Quanzheng Li, Tianming Liu 0001, Xiang Li 0001, Ninghao Liu 0001 |
BIBM | 6 |
| 2025 | Impromptu Cybercrime Euphemism DetectionabstractDetecting euphemisms is essential for content security on various social media platforms, but existing methods designed for detecting euphemisms are ineffective in impromptu euphemisms. In this work, we make a first attempt to an exploration of impromptu euphemism detection and introduce the Impromptu Cybercrime Euphemisms Detection (ICED) dataset. Moreover, we propose a detection framework tailored to this problem, which employs context augmentation modeling and multi-round iterative training. Our detection framework mainly consists of a coarse-grained and a fine-grained classification model. The coarse-grained classification model removes most of the harmless content in the corpus to be detected. The fine-grained model, impromptu euphemisms detector, integrates context augmentation and multi-round iterations training to better predicts the actual meaning of a masked token. In addition, we leverage ChatGPT to evaluate the mode’s capability. Experimental results demonstrate that our approach achieves a remarkable 76-fold improvement compared to the previous state-of-the-art euphemism detector. Xiang Li 0001, Yucheng Zhou 0001, Laiping Zhao, Jing Li 0034, Fangming Liu |
COLING | 1 |
| 2025 | ECHOPulse: ECG Controlled Echocardio-gram Video GenerationabstractEchocardiography (ECHO) is essential for cardiac assessments, but its video quality and interpretation heavily relies on manual expertise, leading to inconsistent results from clinical and portable devices. ECHO video generation offers a solution by improving automated monitoring through synthetic data and generating high-quality videos from routine health data. However, existing models often face high computational costs, slow inference, and rely on complex conditional prompts that require experts' annotations. To address these challenges, we propose ECHOPulse, an ECG-conditioned ECHO video generation model. ECHOPulse introduces two key advancements: (1) it accelerates ECHO video generation by leveraging VQ-VAE tokenization and masked visual token modeling for fast decoding, and (2) it conditions on readily accessible ECG signals, which are highly coherent with ECHO videos, bypassing complex conditional prompts. To the best of our knowledge, this is the first work to use time-series prompts like ECG signals for ECHO video generation. ECHOPulse not only enables controllable synthetic ECHO data generation but also provides updated cardiac function information for disease monitoring and prediction beyond ECG alone. Evaluations on three public and private datasets demonstrate state-of-the-art performance in ECHO video generation across both qualitative and quantitative measures. Additionally, ECHOPulse can be easily generalized to other modality generation tasks, such as cardiac MRI, fMRI, and 3D CT generation. We will make the synthetic ECHO dataset, along with the code and model, publicly available upon acceptance. Yiwei Li 0002, Sekeun Kim, Zihao Wu 0001, Hanqi Jiang, Yi Pan 0001, Pengfei Jin, Sifan Song, Xiaowei Yu 0001, Tianze Yang, Tianming Liu 0001, Quanzheng Li, Xiang Li 0001 |
ICLR | 13 |
| 2025 | Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized DataabstractLarge Multimodal Models (LMMs), or Vision-Language Models (VLMs), have shown impressive capabilities in a wide range of visual tasks. However, they often struggle with fine-grained visual reasoning, failing to identify domain-specific objectives and provide justifiable explanations for their predictions. To address the above challenge, we propose a novel visual rejection sampling framework to improve the cognition and explainability of LMMs using self-synthesized data. Specifically, visual fine-tuning requires images, queries, and target answers. Our approach begins by synthesizing interpretable answers that include human-verifiable visual features. These features are based on expert-defined concepts, and carefully selected based on their alignment with the image content. After each round of fine-tuning, we apply a reward model-free filtering mechanism to select the highest-quality interpretable answers for the next round of tuning. This iterative process of synthetic data generation and fine-tuning progressively improves the model's ability to generate accurate and reasonable explanations. Experimental results demonstrate the effectiveness of our method in improving both the accuracy and explainability of specialized visual classification tasks. Quanzheng Li, Jin Sun 0011, Xiang Li 0001, Ninghao Liu 0001 |
ICLR | 4 |
| 2025 | Distribution-aware Fairness Learning in Medical Image Segmentation From A Control-Theoretic PerspectiveabstractEnsuring fairness in medical image segmentation is critical due to biases in imbalanced clinical data acquisition caused by demographic attributes (e.g., age, sex, race) and clinical factors (e.g., disease severity). To address these challenges, we introduce Distribution-aware Mixture of Experts (dMoE), inspired by optimal control theory. We provide a comprehensive analysis of its underlying mechanisms and clarify dMoE's role in adapting to heterogeneous distributions in medical image segmentation. Furthermore, we integrate dMoE into multiple network architectures, demonstrating its broad applicability across diverse medical image analysis tasks. By incorporating demographic and clinical factors, dMoE achieves state-of-the-art performance on two 2D benchmark datasets and a 3D in-house dataset. Our results highlight the effectiveness of dMoE in mitigating biases from imbalanced distributions, offering a promising approach to bridging control theory and medical image segmentation within fairness learning paradigms. The source code is available at https://github.com/tvseg/dMoE. Yujin Oh, Pengfei Jin, Sangjoon Park, Sekeun Kim, Siyeop Yoon, Kyung Sang Kim, Xiang Li 0001, Quanzheng Li |
ICML | 8 |
| 2025 | Not All Layers of LLMs Are Necessary During InferenceabstractDue to the large number of parameters, the inference phase of Large Language Models (LLMs) is resource-intensive. However, not all requests posed to LLMs are equally difficult to handle. Through analysis, we show that for some tasks, LLMs can achieve results comparable to the final output at some intermediate layers. That is, not all layers of LLMs are necessary during inference. If we can predict at which layer the inferred results match the final results (produced by evaluating all layers), we could significantly reduce the inference cost. To this end, we propose a simple yet effective algorithm named AdaInfer to adaptively terminate the inference process for an input instance. AdaInfer relies on easily obtainable statistical features and classic classifiers like SVM. Experiments on well-known LLMs like the Llama2 series and OPT, show that AdaInfer can achieve an average of 17.8% pruning ratio, and up to 43% on sentiment tasks, with nearly no performance drop (<1%). Because AdaInfer does not alter LLM parameters, the LLMs incorporated with AdaInfer maintain generalizability across tasks. Siqi Fan 0001, Xin Jiang 0005, Xiang Li 0001, Xuying Meng, Peng Han 0005, Shuo Shang, Aixin Sun, Yequan Wang |
IJCAI | 3 |
| 2025 | DRARL: Disengagement-Reason-Augmented Reinforcement Learning for Efficient Improvement of Autonomous Driving PolicyabstractWith the increasing presence of automated vehicles on open roads under driver supervision, disengagement cases are becoming more prevalent. While some data-driven planning systems attempt to directly utilize these disengagement cases for policy improvement, the inherent scarcity of disengagement data (often occurring as a single instance) restricts training effectiveness. Furthermore, some disengagement data should be excluded since the disengagement may not always come from the failure of driving policies, e.g. the driver may casually intervene for a while. To this end, this work proposes disengagement-reason-augmented reinforcement learning (DRARL), which enhances driving policy improvement process according to the reason of disengagement cases. Specifically, the reason of disengagement is identified by an out-of-distribution (OOD) state estimation model. When the reason doesn’t exist, the case will be identified as a casual disengagement case, which doesn’t require additional policy adjustment. Otherwise, the policy can be updated under a reason-augmented imagination environment, improving the policy performance of disengagement cases with similar reasons. The method is evaluated using real-world disengagement cases collected by autonomous driving robotaxi. Experimental results demonstrate that the method accurately identifies policy-related disengagement reasons, allowing the agent to handle both original and semantically similar cases through reason-augmented training. Furthermore, the approach prevents the agent from becoming overly conservative after policy adjustments. Overall, this work provides an efficient way to improve driving policy performance with disengagement cases. Weitao Zhou, Bo Zhang 0106, Zhong Cao 0003, Xiang Li 0001, Diange Yang |
IROS | 4 |
| 2025 | MAST-Pro: Dynamic Mixture-of-Experts for Adaptive Segmentation of Pan-Tumors with Knowledge-Driven Prompts
Runqi Meng, Sifan Song, Pengfei Jin, Yiqun Sun, Yujin Oh, Xiang Li 0001, Quanzheng Li, Dinggang Shen |
MICCAI (16) | 9 |
| 2025 | SAMed-2: Selective Memory Enhanced Medical Segment Anything Model
Zhiling Yan, Sifan Song, Dingjie Song, Yiwei Li 0002, Rong Zhou 0007, Weixiang Sun, Zhennong Chen, Sekeun Kim, Hui Ren 0001, Tianming Liu 0001, Quanzheng Li, Xiang Li 0001, Lifang He 0001, Lichao Sun 0001 |
MICCAI (13) | 12 |
| 2025 | Cascaded 3D Diffusion Models for Whole-Body 3D 18-F FDG PET/CT Synthesis from Demographics
Siyeop Yoon, Sifan Song, Pengfei Jin, Matthew Tivnan, Yujin Oh, Sekeun Kim, Dufan Wu, Xiang Li 0001, Quanzheng Li |
MICCAI (3) | 8 |
| 2025 | Learning to Plan Like the Human Brain via Visuospatial Perception and Semantic-Episodic Synergistic Decision-MakingabstractMotion planning in high-dimensional continuous spaces remains challenging due to complex environments and computational constraints. Although learning-based planners, especially graph neural network (GNN)-based, have significantly improved planning performance, they still struggle with inaccurate graph construction and limited structural reasoning, constraining search efficiency and path quality. The human brain exhibits efficient planning through a two-stage Perception-Decision model. First, egocentric spatial representations from visual and proprioceptive input are constructed, and then semantic–episodic synergy is leveraged to support decision-making in uncertainty scenarios. Inspired by this process, we propose NeuroMP, a brain-inspired planning framework that learns to plan like the human brain. NeuroMP integrates a Perceptive Segment Selector inspired by visuospatial perception to construct safer graphs, and a Global Alignment Heuristic guide search in weakly connected graphs by modeling semantic-episodic synergistic decision-making. Experimental results demonstrate that NeuroMP significantly outperforms existing planning methods in efficiency and quality while maintaining a high success rate. Tianyuan Jia, Qing Li 0027, Xiuxing Li, Xiang Li 0001, Li Yao 0002, Xia Wu 0001 |
NeurIPS | 5 |
| 2025 | REOBench: Benchmarking Robustness of Earth Observation Foundation ModelsabstractEarth observation foundation models have shown strong generalization across multiple Earth observation tasks, but their robustness under real-world perturbations remains underexplored. To bridge this gap, we introduce REOBench, the first comprehensive benchmark for evaluating the robustness of Earth observation foundation models across six tasks and twelve types of image corruptions, including both appearance-based and geometric perturbations. To ensure realistic and fine-grained evaluation, our benchmark focuses on high-resolution optical remote sensing images, which are widely used in critical applications such as urban planning and disaster response. We conduct a systematic evaluation of a broad range of models trained using masked image modeling, contrastive learning, and vision-language pre-training paradigms. Our results reveal that (1) existing Earth observation foundation models experience significant performance degradation when exposed to input corruptions. (2) The severity of degradation varies across tasks, model architectures, backbone sizes, and types of corruption, with performance drop varying from less than 1% to over 25%. (3) Vision-language models show enhanced robustness, particularly in multimodal tasks. REOBench underscores the vulnerability of current Earth observation foundation models to real-world corruptions and provides actionable insights for developing more robust and reliable models. Xiang Li 0001, Siwei Liu 0001, Zhitong Xiong, Chunbo Luo, Lu Liu 0001, Mykola Pechenizkiy, Xiao Xiang Zhu 0001, Tianjin Huang |
NeurIPS | 1 |
| 2025 | Shapley Value Estimation based on Differential MatrixabstractThe Shapley value has been extensively used in many fields as the unique metric to fairly evaluate player contributions in cooperative settings. Since the exact computation of Shapley values is \#P-hard in the task-agnostic setting, many studies have been developed to utilize the Monte Carlo method for Shapley value estimation. The existing methods estimate the Shapley values directly. In this paper, we explore a novel idea-inferring the Shapley values by estimating the differences between them. Technically, we estimate a differential matrix consisting of pairwise Shapley value differences to reduce the variance of the estimated Shapley values. We develop a least-squares optimization solution to derive the Shapley values from the differential matrix, minimizing the estimator variances. Additionally, we devise a Monte Carlo method for efficient estimation of the differential matrix and introduce two stratified Monte Carlo methods for further variance reduction. Our experimental results on real and synthetic data sets demonstrate the effectiveness and efficiency of the differential-matrix-based sampling approaches. Junyuan Pang, Jian Pei 0001, Haocheng Xia, Xiang Li 0001, Jinfei Liu |
Proc. ACM Manag. Data | 4 |
| 2025 | AugGPT: Leveraging ChatGPT for Text Data AugmentationabstractText data augmentation is an effective strategy for overcoming the challenge of limited sample sizes in many natural language processing (NLP) tasks. This challenge is especially prominent in the few-shot learning (FSL) scenario, where the data in the target domain is generally much scarcer and of lowered quality. A natural and widely used strategy to mitigate such challenges is to perform data augmentation to better capture data invariance and increase the sample size. However, current text data augmentation methods either can’t ensure the correct labeling of the generated data (lacking faithfulness), or can’t ensure sufficient diversity in the generated data (lacking compactness), or both. Inspired by the recent success of large language models (LLM), especially the development of ChatGPT, we propose a text data augmentation approach based on ChatGPT (named ”AugGPT”). AugGPT rephrases each sentence in the training samples into multiple conceptually similar but semantically different samples. The augmented samples can then be used in downstream model training. Experiment results on multiple few-shot learning text classification tasks show the superior performance of the proposed AugGPT approach over state-of-the-art text data augmentation methods in terms of testing accuracy and distribution of the augmented samples. Haixing Dai, Zhengliang Liu, Wenxiong Liao, Zihao Wu 0001, Lin Zhao 0004, Shaochen Xu, Fang Zeng, Wei Liu 0146, Ninghao Liu 0001, Sheng Li 0001, Dajiang Zhu, Hongmin Cai, Lichao Sun 0001, Quanzheng Li, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001 |
IEEE Trans. Big Data | 19 |
| 2025 | Exploring the Trade-Offs: Unified Large Language Models vs Local Fine-Tuned Models for Highly-Specific Radiology NLI TaskabstractRecently, ChatGPT and GPT-4 have emerged and gained immense global attention due to their unparalleled performance in language processing. Despite demonstrating impressive capability in various open-domain tasks, their adequacy in highly specific fields like radiology remains untested. Radiology presents unique linguistic phenomena distinct from open-domain data due to its specificity and complexity. Assessing the performance of large language models (LLMs) in such specific domains is crucial not only for a thorough evaluation of their overall performance but also for providing valuable insights into future model design directions: whether model design should be generic or domain-specific. To this end, in this study, we evaluate the performance of ChatGPT/GPT-4 on a radiology natural language inference (NLI) task and compare it to other models fine-tuned specifically on task-related data samples. We also conduct a comprehensive investigation on ChatGPT/GPT-4’s reasoning ability by introducing varying levels of inference difficulty. Our results show that 1) ChatGPT and GPT-4 outperform other LLMs in the radiology NLI task and 2) other specifically fine-tuned Bert-based models require significant amounts of data samples to achieve comparable performance to ChatGPT/GPT-4. These findings not only demonstrate the feasibility and promise of constructing a generic model capable of addressing various tasks across different domains, but also highlight several key factors crucial for developing a unified model, particularly in a medical context, paving the way for future artificial general intelligence (AGI) systems. We release our code and data to the research community. Zihao Wu 0001, Lu Zhang 0050, Xiaowei Yu 0001, Zhengliang Liu, Lin Zhao 0004, Yiwei Li 0002, Haixing Dai, Chong Ma 0004, Gang Li 0001, Wei Liu 0146, Quanzheng Li, Dinggang Shen, Xiang Li 0001, Dajiang Zhu, Tianming Liu 0001 |
IEEE Trans. Big Data | 14 |
| 2025 | MediViSTA: Medical Video Segmentation Via Temporal Fusion SAM Adaptation for EchocardiographyabstractDespite achieving impressive results in general-purpose semantic segmentation with strong generalization on natural images, the Segment Anything Model (SAM) has shown less precision and stability in medical image segmentation. In particular, the original SAM architecture is designed for 2D natural images and is therefore not support to handle three-dimensional information, which is particularly important for medical imaging modalities that are often volumetric or video data. In this paper, we introduce MediViSTA, a parameter-efficient fine-tuning method designed to adapt the vision foundation model for medical video, with a specific focus on echocardiography segmentation. To achieve spatial adaptation, we propose a frequency feature fusion technique that injects spatial frequency information from a CNN branch. For temporal adaptation, we integrate temporal adapters within the transformer blocks of the image encoder. Using a fine-tuning strategy, only a small subset of pre-trained parameters is updated, allowing efficient adaptation to echocardiography data. The effectiveness of our method has been comprehensively evaluated on three datasets, comprising two public datasets and one multi-center in-house dataset. Our method consistently outperforms various state-of-the-art approaches without using any prompts. Furthermore, our model exhibits strong generalization capabilities on unseen datasets, surpassing the second-best approach by 2.15% in Dice and 0.09 in temporal consistency. The results demonstrate the potential of MediViSTA to significantly advance echocardiography video segmentation, offering improved accuracy and robustness in cardiac assessment applications. Sekeun Kim, Pengfei Jin, Cheng Chen 0013, Kyung Sang Kim, Zhiliang Lyu, Hui Ren 0001, Zhengliang Liu, Aoxiao Zhong, Tianming Liu 0001, Xiang Li 0001, Quanzheng Li |
IEEE J. Biomed. Health Informatics | 11 |
| 2025 | Voxel-Level Brain States Prediction Using Swin TransformerabstractUnderstanding brain dynamics is important for neuroscience and mental health. Functional magnetic resonance imaging (fMRI) enables the measurement of neural activities through blood-oxygen-level-dependent (BOLD) signals, which represent brain states. In this study, we aim to predict future human resting brain states with fMRI. Due to the 3D voxel-wise spatial organization and temporal dependencies of the fMRI data, we propose a novel architecture which employs a 4D Shifted Window (Swin) Transformer as encoder to efficiently learn spatio-temporal information and a convolutional decoder to enable brain state prediction at the same spatial and temporal resolution as the input fMRI data. We used 100 unrelated subjects from the Human Connectome Project (HCP) for model training and testing. Our novel model has shown high accuracy when predicting 7.2s resting-state brain activities based on the prior 23.04s fMRI time series. The predicted brain states highly resemble BOLD contrast and dynamics. This work shows promising evidence that the spatiotemporal organization of the human brain can be learned by a Swin Transformer model, at high resolution, which provides a potential for reducing the fMRI scan time and the development of brain-computer interfaces in the future. Yifei Sun 0013, Daniel Chahine, Qinghao Wen, Tianming Liu 0001, Xiang Li 0001, Yixuan Yuan, Fernando Calamante, Jinglei Lv |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | EchoFM: Foundation Model for Generalizable Echocardiogram AnalysisabstractEchocardiography is the first-line non-invasive cardiac imaging modality, providing rich spatio-temporal information on cardiac anatomy and physiology. Recently, foundation model trained on extensive and diverse datasets has shown strong performance in various downstream tasks. However, translating foundation models into the medical imaging domain remains challenging due to domain differences between medical and natural images, the lack of diverse patient and disease datasets. In this paper, we introduce EchoFM, a general-purpose vision foundation model for echocardiography trained on a large-scale dataset of over 20 million echocardiographic images from 6,500 patients. To enable effective learning of rich spatio-temporal representations from periodic videos, we propose a novel self-supervised learning framework based on a masked autoencoder with a spatio-temporal consistent masking strategy and periodic-driven contrastive learning. The learned cardiac representations can be readily adapted and fine-tuned for a wide range of downstream tasks, serving as a strong and flexible backbone model. We validate EchoFM through experiments across key downstream tasks in the clinical echocardiography workflow, leveraging public and multi-center internal datasets. EchoFM consistently outperforms SOTA methods, demonstrating superior generalization capabilities and flexibility. The code and checkpoints are available at: https://github.com/SekeunKim/EchoFM.git. Sekeun Kim, Pengfei Jin, Sifan Song, Cheng Chen 0013, Yiwei Li 0002, Hui Ren 0001, Xiang Li 0001, Tianming Liu 0001, Quanzheng Li |
IEEE Trans. Medical Imaging | 7 |
| 2025 | LLM-Guided Decoupled Probabilistic Prompt for Continual Learning in Medical Image DiagnosisabstractDeep learning-based traditional diagnostic models typically exhibit limitations when applied to dynamic clinical environments that require handling the emergence of new diseases. Continual learning (CL) offers a promising solution, aiming to learn new knowledge while preserving previously learned knowledge. Though recent rehearsal-free CL methods employing prompt tuning (PT) have shown promise, they rely on deterministic prompts that struggle to handle diverse fine-grained knowledge. Moreover, existing PT methods utilize randomly initialized prompts that are trained under standard classification constraints, impeding expert knowledge integration and optimal performance acquisition. In this paper, we propose an LLM-guided Decoupled Probabilistic Prompt (LDPP) for Continual Learning in medical image diagnosis. Specifically, we develop an Expert Knowledge Generation (EKG) module that leverages LLM to acquire decoupled expert knowledge and comprehensive category descriptions. Then, we introduce a Decoupled Probabilistic Prompt pool (DePP) to construct a shared decoupled probabilistic prompt pool, which constructs a shared prompt pool with probabilistic prompts derived from the expert knowledge set. These prompts dynamically provide diverse and flexible descriptions for input images. Finally, We design a Steering Prompt Pool (SPP) to enhance intra-class compactness and promote model performance by learning non-shared prompts. With extensive experimental validation, LDPP consistently sets state-of-the-art performance under the challenging class-incremental setting in CL. Code is available at: https://github.com/CUHK-AIM-Group/LDPP. Yiwen Luo, Wuyang Li, Xiang Li 0001, Tianming Liu 0001, Tianye Niu, Yixuan Yuan |
IEEE Trans. Medical Imaging | 4 |
| 2024 | DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object DetectionabstractVehicle-to-Everything (V2X) collaborative perception has recently gained significant attention due to its capability to enhance scene understanding by integrating information from various agents, e.g., vehicles, and infrastructure. However, current works often treat the information from each agent equally, ignoring the inherent domain gap caused by the utilization of different LiDAR sensors of each agent, thus leading to suboptimal performance. In this paper, we propose DI-V2X, that aims to learn Domain-Invariant representations through a new distillation framework to mitigate the domain discrepancy in the context of V2X 3D object detection. DI-V2X comprises three essential components: a domain-mixing instance augmentation (DMA) module, a progressive domain-invariant distillation (PDD) module, and a domain-adaptive fusion (DAF) module. Specifically, DMA builds a domain-mixing 3D instance bank for the teacher and student models during training, resulting in aligned data representation. Next, PDD encourages the student models from different domains to gradually learn a domain-invariant feature representation towards the teacher, where the overlapping regions between agents are employed as guidance to facilitate the distillation process. Furthermore, DAF closes the domain gap between the students by incorporating calibration-aware domain-adaptive attention. Extensive experiments on the challenging DAIR-V2X and V2XSet benchmark datasets demonstrate DI-V2X achieves remarkable performance, outperforming all the previous V2X models. Code is available at https://github.com/Serenos/DI-V2X. Xiang Li 0001, Junbo Yin, Wei Li 0111, Cheng-Zhong Xu 0001, Ruigang Yang, Jianbing Shen |
AAAI | 1 |
| 2024 | Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report GenerationabstractFine-grained vision-language models (VLM) have been widely used for inter-modality local alignment between the predefined fixed patches and textual words. However, in medical analysis, lesions exhibit varying sizes and positions, and using fixed patches may cause incomplete representations of lesions. Moreover, these methods provide explainability by using heatmaps to show the general image areas potentially associated with texts rather than specific regions, making their explanations not explicit and specific enough. To address these issues, we propose a novel Adaptive patch-word Matching (AdaMatch) model to correlate chest X-ray (CXR) image regions with words in medical reports and apply it to CXR-report generation to provide explainability for the generation process. AdaMatch exploits the fine-grained relation between adaptive patches and words to provide explanations of specific image regions with corresponding words. To capture the abnormal regions of varying sizes and positions, we introduce an Adaptive Patch extraction (AdaPatch) module to acquire adaptive patches for these regions adaptively. Aiming to provide explicit explainability for the CXR-report generation task, we propose an AdaMatch-based bidirectional LLM for Cyclic CXR-report generation (AdaMatch-Cyclic). It employs AdaMatch to obtain the keywords for CXR images and 'keypatches' for medical reports as hints to guide CXR-report generation. Extensive experiments on two publicly available CXR datasets validate the effectiveness of our method and its superior performance over existing methods. © 2024 Association for Computational Linguistics. Wenting Chen, LinLin Shen, Jiebo Luo 0001, Xiang Li 0001, Yixuan Yuan |
ACL (1) | 5 |
| 2024 | KnowCoder: Coding Structured Knowledge into LLMs for Universal Information ExtractionabstractZixuan Li, Yutao Zeng, Yuxin Zuo, Weicheng Ren, Wenxuan Liu, Miao Su, Yucan Guo, Yantao Liu, Xiang Li, Zhilei Hu, Long Bai, Wei Li, Yidan Liu, Pan Yang, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zixuan Li 0001, Yutao Zeng, Yuxin Zuo, Weicheng Ren, Wenxuan Liu 0003, Miao Su, Yucan Guo, Yantao Liu, Xiang Li 0001, Zhilei Hu, Long Bai 0002, Wei Li 0176, Yidan Liu, Xiaolong Jin 0001, Jiafeng Guo, Xueqi Cheng 0001 |
ACL (1) | 9 |
| 2024 | Position: TrustLLM: Trustworthiness in Large Language ModelsabstractLarge language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for mainstream LLMs, and discussion of open challenges and future directions. Specifically, we first propose a set of principles for trustworthy LLMs that span eight different dimensions. Based on these principles, we further establish a benchmark across six dimensions including truthfulness, safety, fairness, robustness, privacy, and machine ethics. We then present a study evaluating 16 mainstream LLMs in TrustLLM, consisting of over 30 datasets. Our findings firstly show that in general trustworthiness and capability (i.e., functional effectiveness) are positively related. Secondly, our observations reveal that proprietary LLMs generally outperform most open-source counterparts in terms of trustworthiness, raising concerns about the potential risks of widely accessible open-source LLMs. However, a few open-source LLMs come very close to proprietary ones, suggesting that open-source models can achieve high levels of trustworthiness without additional mechanisms like moderator, offering valuable insights for developers in this field. Thirdly, it is important to note that some LLMs may be overly calibrated towards exhibiting trustworthiness, to the extent that they compromise their utility by mistakenly treating benign prompts as harmful and consequently not responding. Besides these observations, we’ve uncovered key insights into the multifaceted trustworthiness in LLMs. We emphasize the importance of ensuring transparency not only in the models themselves but also in the technologies that underpin trustworthiness. We advocate that the establishment of an AI alliance between industry, academia, the open-source community to foster collaboration is imperative to advance the trustworthiness of LLMs. Yue Huang 0001, Lichao Sun 0001, Haoran Wang 0005, Siyuan Wu 0001, Qihui Zhang, Chujie Gao, Wenhan Lyu, Yixuan Zhang 0001, Xiner Li, Hanchi Sun, Zhengliang Liu, Yixin Liu 0002, Yijue Wang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric P. Xing, Furong Huang, Heng Ji 0001, Hongyi Wang 0001, Huan Zhang 0001, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang 0001, Mohit Bansal, James Zou 0001, Jian Pei 0001, Jianfeng Gao 0001, Jiawei Han 0001, Jieyu Zhao 0001, Jiliang Tang, Jindong Wang 0001, Joaquin Vanschoren, John C. Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang 0001, Lifang He 0001, Lifu Huang, Michael Backes 0001, Neil Zhenqiang Gong, Philip S. Yu, Quanquan Gu, Ran Xu 0001, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen 0001, Tianming Liu 0001, Tianyi Zhou 0001, William Yang Wang, Xiang Li 0001, Xiangliang Zhang 0001, Xiao Wang 0012, Xing Xie 0001, Xuyu Wang, Yan Liu 0002, Yanfang Ye 0001, Yinzhi Cao, Yong Chen 0016, Yue Zhao 0016 |
ICML | 61 |
| 2024 | Medical Image Synthesis via Fine-Grained Image-Text Alignment and Anatomy-Pathology Prompting
Wenting Chen, Pengyu Wang 0005, Hui Ren 0001, Lichao Sun 0001, Quanzheng Li, Yixuan Yuan, Xiang Li 0001 |
MICCAI (12) | 7 |
| 2024 | F2TNet: FMRI to T1w MRI Knowledge Transfer Network for Brain Multi-phenotype Prediction
Wuyang Li, Yu Jiang 0013, Zhihao Peng 0002, Pengyu Wang 0005, Xiang Li 0001, Tianming Liu 0001, Junwei Han 0001, Yixuan Yuan |
MICCAI (11) | 6 |
| 2024 | Hallucination Index: An Image Quality Metric for Generative Reconstruction Models
Matthew Tivnan, Siyeop Yoon, Zhennong Chen, Xiang Li 0001, Dufan Wu, Quanzheng Li |
MICCAI (10) | 4 |
| 2024 | Conditional Score-Based Diffusion Model for Cortical Thickness Trajectory Prediction
Qing Xiao 0003, Siyeop Yoon, Hui Ren 0001, Matthew Tivnan, Lichao Sun 0001, Quanzheng Li, Tianming Liu 0001, Yu Zhang 0064, Xiang Li 0001 |
MICCAI (2) | 9 |
| 2024 | Volumetric Conditional Score-Based Residual Diffusion Model for PET/MR Denoising
Siyeop Yoon, Matthew Tivnan, Yuang Wang, Young-Don Son, Dufan Wu, Xiang Li 0001, Kyung Sang Kim, Quanzheng Li |
MICCAI (7) | 7 |
| 2024 | Biomedical Visual Instruction Tuning with Clinician Preference AlignmentabstractRecent advancements in multimodal foundation models have showcased impressive capabilities in understanding and reasoning with visual and textual information. Adapting these foundation models trained for general usage to specialized domains like biomedicine requires large-scale domain-specific instruction datasets. While existing works have explored curating such datasets automatically, the resultant datasets are not explicitly aligned with domain expertise. In this work, we propose a data-centric framework, Biomedical Visual Instruction Tuning with Clinician Preference Alignment (BioMed-VITAL), that incorporates clinician preferences into both stages of generating and selecting instruction data for tuning biomedical multimodal foundation models. First, during the generation stage, we prompt the GPT-4V generator with a diverse set of clinician-selected demonstrations for preference-aligned data candidate generation. Then, during the selection phase, we train a separate selection model, which explicitly distills clinician and policy-guided model preferences into a rating function to select high-quality data for medical instruction tuning. Results show that the model tuned with the instruction-following data from our method demonstrates a significant improvement in open visual chat (18.5% relatively) and medical VQA (win rate up to 81.73%). Our instruction-following data and models are available at https://BioMed-VITAL.github.io. Hejie Cui, Lingjun Mao, Jieyu Zhang 0001, Hui Ren 0001, Quanzheng Li, Xiang Li 0001, Carl Yang 0001 |
NeurIPS | 7 |
| 2024 | Eye-gaze Guided Multi-modal Alignment for Medical Representation LearningabstractIn the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly aligned from the data, without considering the explicit relationships in the medical context. This data-reliance may lead to low generalization of the learned alignment relationships. In this work, we propose the Eye-gaze Guided Multi-modal Alignment (EGMA) framework to harness eye-gaze data for better alignment of medical visual and textual features. We explore the natural auxiliary role of radiologists' eye-gaze data in aligning medical images and text, and introduce a novel approach by using eye-gaze data, collected synchronously by radiologists during diagnostic evaluations. We conduct downstream tasks of image classification and image-text retrieval on four medical datasets, where EGMA achieved state-of-the-art performance and stronger generalization across different datasets. Additionally, we explore the impact of varying amounts of eye-gaze data on model performance, highlighting the feasibility and utility of integrating this auxiliary data into multi-modal alignment framework. Chong Ma 0004, Hanqi Jiang, Wenting Chen, Yiwei Li 0002, Zihao Wu 0001, Xiaowei Yu 0001, Zhengliang Liu, Lei Guo 0002, Dajiang Zhu, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001 |
NeurIPS | 13 |
| 2024 | Retrieval-Augmented Code Generation for Universal Information Extraction
Yucan Guo, Zixuan Li 0001, Xiaolong Jin 0001, Yantao Liu, Yutao Zeng, Wenxuan Liu 0003, Xiang Li 0001, Long Bai 0002, Jiafeng Guo, Xueqi Cheng 0001 |
NLPCC (2) | 7 |
| 2024 | Mask-guided BERT for few-shot text classification
Wenxiong Liao, Zhengliang Liu, Haixing Dai, Zihao Wu 0001, Yiyang Zhang 0003, Yuzhong Chen 0002, Xi Jiang 0001, Dajiang Zhu, Sheng Li 0001, Wei Liu 0146, Tianming Liu 0001, Quanzheng Li, Hongmin Cai, Xiang Li 0001 |
Neurocomputing | 16 |
| 2024 | Zero-shot relation triplet extraction as Next-Sentence Prediction
Wenxiong Liao, Zhengliang Liu, Yiyang Zhang 0003, Ninghao Liu 0001, Tianming Liu 0001, Quanzheng Li, Xiang Li 0001, Hongmin Cai |
Knowl. Based Syst. | 8 |
| 2024 | MA-SAM: Modality-agnostic SAM adaptation for 3D medical image segmentation
Cheng Chen 0013, Juzheng Miao, Dufan Wu, Aoxiao Zhong, Zhiling Yan, Sekeun Kim, Zhengliang Liu, Lichao Sun 0001, Xiang Li 0001, Tianming Liu 0001, Pheng-Ann Heng, Quanzheng Li |
Medical Image Anal. | 10 |
| 2024 | Structure Mapping Generative Adversarial Network for Multi-View Information Mapping Pattern MiningabstractMulti-view learning is dedicated to integrating information from different views and improving the generalization performance of models. However, in most current works, learning under different views has significant independency, overlooking common information mapping patterns that exist between these views. This paper proposes a Structure Mapping Generative adversarial network (SM-GAN) framework, which utilizes the consistency and complementarity of multi-view data from the innovative perspective of information mapping. Specifically, based on network-structured multi-view data, a structural information mapping model is proposed to capture hierarchical interaction patterns among views. Subsequently, three different types of graph convolutional operations are designed in SM-GAN based on the model. Compared with regular GAN, we add a structural information mapping module between the encoder and decoder wthin the generator, completing the structural information mapping from the micro-view to the macro-view. This paper conducted sufficient validation experiments using public imaging genetics data in Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset. It is shown that SM-GAN outperforms baseline and advanced methods in multi-label classification and evolution prediction tasks. Xia-an Bi, YangJun Huang, Zicheng Yang, Zhao-Xu Xing, Luyun Xu, Xiang Li 0001, Zhengliang Liu, Tianming Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | P-Shapley: Shapley Values on Probabilistic ClassifiersabstractThe Shapley value provides a unique approach to equitably gauge each player's contribution within a coalition and has extensive applications with various utility functions. In data valuation for machine learning, particularly for classification tasks, using classification accuracy as the utility function has become a de facto standard. However, accuracy can be an imprecise metric, potentially missing finer details crucial for valuation. In this paper, we propose the probability-based Shapley (P-Shapley) value, which leverages predicted probabilities to heighten utility differentiation. Several convex calibration functions are further incorporated for probability calibration. We prove that the P-Shapley value outperforms Shapley values based on accuracy or other coarse metrics in approximation stability and the discrimination of marginal utility change can be further improved by convex calibration functions. Extensive experiments on four real-world datasets demonstrate the effectiveness of our approaches. Haocheng Xia, Xiang Li 0001, Junyuan Pang, Jinfei Liu, Kui Ren 0001, Li Xiong 0001 |
Proc. VLDB Endow. | 2 |
| 2024 | CE-GAN: Community Evolutionary Generative Adversarial Network for Alzheimer's Disease Risk PredictionabstractIn the studies of neurodegenerative diseases such as Alzheimer's Disease (AD), researchers often focus on the associations among multi-omics pathogeny based on imaging genetics data. However, current studies overlook the communities in brain networks, leading to inaccurate models of disease development. This paper explores the developmental patterns of AD from the perspective of community evolution. We first establish a mathematical model to describe functional degeneration in the brain as the community evolution driven by entropy information propagation. Next, we propose an interpretable Community Evolutionary Generative Adversarial Network (CE-GAN) to predict disease risk. In the generator of CE-GAN, community evolutionary convolutions are designed to capture the evolutionary patterns of AD. The experiments are conducted using functional magnetic resonance imaging (fMRI) data and single nucleotide polymorphism (SNP) data. CE-GAN achieves 91.67% accuracy and 91.83% area under curve (AUC) in AD risk prediction tasks, surpassing advanced methods on the same dataset. In addition, we validated the effectiveness of CE-GAN for pathogeny extraction. The source code of this work is available at https://github.com/fmri123456/CE-GAN. Xia-an Bi, Zicheng Yang, YangJun Huang, Zhao-Xu Xing, Luyun Xu, Zihao Wu 0001, Zhengliang Liu, Xiang Li 0001, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2023 | LWSIS: LiDAR-Guided Weakly Supervised Instance Segmentation for Autonomous DrivingabstractImage instance segmentation is a fundamental research topic in autonomous driving, which is crucial for scene understanding and road safety. Advanced learning-based approaches often rely on the costly 2D mask annotations for training. In this paper, we present a more artful framework, LiDAR-guided Weakly Supervised Instance Segmentation (LWSIS), which leverages the off-the-shelf 3D data, i.e., Point Cloud, together with the 3D boxes, as natural weak supervisions for training the 2D image instance segmentation models. Our LWSIS not only exploits the complementary information in multimodal data during training but also significantly reduces the annotation cost of the dense 2D masks. In detail, LWSIS consists of two crucial modules, Point Label Assignment (PLA) and Graph-based Consistency Regularization (GCR). The former module aims to automatically assign the 3D point cloud as 2D point-wise labels, while the atter further refines the predictions by enforcing geometry and appearance consistency of the multimodal data. Moreover, we conduct a secondary instance segmentation annotation on the nuScenes, named nuInsSeg, to encourage further research on multimodal perception tasks. Extensive experiments on the nuInsSeg, as well as the large-scale Waymo, show that LWSIS can substantially improve existing weakly supervised segmentation models by only involving 3D data during training. Additionally, LWSIS can also be incorporated into 3D object detectors like PointPainting to boost the 3D detection performance for free. The code and dataset are available at https://github.com/Serenos/LWSIS. Xiang Li 0001, Junbo Yin, Botian Shi, Yikang Li 0002, Ruigang Yang, Jianbing Shen |
AAAI | 1 |
| 2023 | Coarse-to-fine Knowledge Graph Domain Adaptation based on Distantly-supervised Iterative TrainingabstractThe knowledge graph (KG) is a highly needed basis to support the high-fidelity and high-interpretability modeling of various tasks in healthcare artificial intelligence. In this work, we focus on constructing an oncology knowledge graph that will be used in downstream cancer research and solution development. Modern supervised learning for knowledge graph construction requires a large amount of manually labeled data, which makes the process time-consuming and labor-intensive. Although there exists multiple research on named entity recognition and relation extraction based on distantly supervised learning, constructing a domain-specific knowledge graph from large collections of textual data without manual annotations is still an urgent problem to be solved. In response, we propose an integrated framework for adapting and re-learning knowledge graphs from a general domain (biomedical in our case) to a fine-defined domain (oncology). In this framework, we apply distant-supervision on cross-domain knowledge graph adaptation. Consequently, no manual data annotation is required to train the model. We introduce a novel iterative training strategy to facilitate the discovery of domain-specific named entities and triplets. Experimental results indicate that the proposed framework can perform domain adaptation and construction of knowledge graphs efficiently. Wenxiong Liao, Zhengliang Liu, Yiyang Zhang 0003, Fei Qi 0007, Siqi Ding, Hui Ren 0001, Zihao Wu 0001, Haixing Dai, Sheng Li 0001, Lingfei Wu 0001, Ninghao Liu 0001, Quanzheng Li, Tianming Liu 0001, Xiang Li 0001, Hongmin Cai |
BIBM | 15 |
| 2023 | Disentangled and Robust Representation Learning for Bragging Classification in Social MediaabstractResearching bragging behavior on social media arouses interest of computational (socio) linguists. However, existing bragging classification datasets suffer from a serious data imbalance issue. Because labeling a data-balance dataset is expensive, most methods introduce external knowledge to improve model learning. Nevertheless, such methods inevitably introduce noise and non-relevance information from external knowledge. To overcome the drawback, we propose a novel bragging classification method with disentangle-based representation augmentation and domain-aware adversarial strategy. Specifically, model learns to disentangle and reconstruct representation and generate augmented features via disentangle-based representation augmentation. Moreover, domain-aware adversarial strategy aims to constrain domain of augmented features to improve their robustness. Experimental results demonstrate that our method achieves state-of-the-art performance compared to other methods. Xiang Li 0001, Yucheng Zhou 0001 |
ICASSP | 1 |
| 2023 | ShapleyFL: Robust Federated Learning Based on Shapley ValueabstractFederated Learning (FL) allows clients to form a consortium to train a global model under the orchestration of a central server while keeping data on the local client without sharing it, thus mitigating data privacy issues. However, training a robust global model is challenging since the local data is invisible to the server. The local data of clients are naturally heterogeneous, while some clients can use corrupted data or send malicious updates to interfere with the training process artificially. Meanwhile, communication and computation costs are inevitable challenges in designing a practical FL algorithm. In this paper, to improve the robustness of FL, we propose a Shapley value-inspired adaptive weighting mechanism, which regards the FL training as sequential cooperative games and adjusts clients' weights according to their contributions. We also develop a client sampling strategy based on importance sampling, which can reduce the communication cost by optimizing the variance of the global updates according to the weights of clients. Furthermore, to diminish the computation cost of the server, we propose a weight calculation method by estimating differences between the Shapley value of clients. Our experimental results on several real data sets demonstrate the effectiveness of our approaches. Qiheng Sun, Xiang Li 0001, Jiayao Zhang 0006, Li Xiong 0001, Jinfei Liu, Zhan Qin, Kui Ren 0001 |
KDD | 2 |
| 2023 | Hummingbird: Dynamic Path Validation With Hidden Equal-Probability SamplingabstractPath validation has already been incrementally deployed in the Internet architecture. It secures packet forwarding by enabling end hosts to negotiate specific forwarding paths and enforcing on-path routers to prove their forwarding behaviors along these paths. Most existing path validation solutions target static paths, paying less attention to fully dynamic paths that support flexible routing. In this paper, we present Hummingbird as the first validation solution over fully dynamic paths. It features a hidden equal-probability sampling technique. Gaining efficiency via routers probabilistically sampling packets to validate, we craft the sampling probability such that each router validates a similar amount of packets given an unknown path length. We further hide the state of whether a packet has been sampled and validated using a lightweight, non-cryptographic scheme. This prevents attackers from differentiating and selectively mis-forwarding packets. We validate security and efficiency of Hummingbird through both theoretical proof and experimental evaluation. Anxiao He, Xiang Li 0001, Jiandong Fu, Kai Bu, Chenlu Miao, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | CofeNet: Context and Former-Label Enhanced Net for Complicated Quotation ExtractionabstractQuotation extraction aims to extract quotations from written text. There are three components in a quotation: source refers to the holder of the quotation, cue is the trigger word(s), and content is the main body. Existing solutions for quotation extraction mainly utilize rule-based approaches and sequence labeling models. While rule-based approaches often lead to low recalls, sequence labeling models cannot well handle quotations with complicated structures. In this paper, we propose the Context and Former-Label Enhanced Net () for quotation extraction. is able to extract complicated quotations with components of variable lengths and complicated structures. On two public datasets (and ) and one proprietary dataset (), we show that our achieves state-of-the-art performance on complicated quotation extraction. Yequan Wang, Xiang Li 0001, Aixin Sun, Xuying Meng, Huaming Liao, Jiafeng Guo |
COLING | 2 |
| 2021 | Using Keystroke Analytics to Understand Cognitive Processes during Writing
Mo Zhang, Hongwen Guo, Xiang Li 0001 |
EDM | 3 |
| 2021 | Deep metric learning-based image retrieval system for chest radiograph and its clinical applications in COVID-19
Aoxiao Zhong, Xiang Li 0001, Dufan Wu, Hui Ren 0001, Kyung Sang Kim, Young-Gon Kim, Varun Buch, Nir Neumark, Bernardo Bizzo, Won Young Tak, Soo Young Park, Yu Rim Lee, Min Kyu Kang, Jung Gil Park, Byung Seok Kim, Woo Jin Chung, Ittai Dayan, Mannudeep K. Kalra, Quanzheng Li |
Medical Image Anal. | 2 |
| 2021 | Left Ventricle Quantification Challenge: A Comprehensive Comparison and Evaluation of Segmentation and Regression for Mid-Ventricular Short-Axis Cardiac MR DataabstractAutomatic quantification of the left ventricle (LV) from cardiac magnetic resonance (CMR) images plays an important role in making the diagnosis procedure efficient, reliable, and alleviating the laborious reading work for physicians. Considerable efforts have been devoted to LV quantification using different strategies that include segmentation-based (SG) methods and the recent direct regression (DR) methods. Although both SG and DR methods have obtained great success for the task, a systematic platform to benchmark them remains absent because of differences in label information during model learning. In this paper, we conducted an unbiased evaluation and comparison of cardiac LV quantification methods that were submitted to the Left Ventricle Quantification (LVQuan) challenge, which was held in conjunction with the Statistical Atlases and Computational Modeling of the Heart (STACOM) workshop at the MICCAI 2018. The challenge was targeted at the quantification of 1) areas of LV cavity and myocardium, 2) dimensions of the LV cavity, 3) regional wall thicknesses (RWT), and 4) the cardiac phase, from mid-ventricle short-axis CMR images. First, we constructed a public quantification dataset Cardiac-DIG with ground truth labels for both the myocardium mask and these quantification targets across the entire cardiac cycle. Then, the key techniques employed by each submission were described. Next, quantitative validation of these submissions were conducted with the constructed dataset. The evaluation results revealed that both SG and DR methods can offer good LV quantification performance, even though DR methods do not require densely labeled masks for supervision. Among the 12 submissions, the DR method LDAMT offered the best performance, with a mean estimation error of 301 mm2for the two areas, 2.15 mm for the cavity dimensions, 2.03 mm for RWTs, and a 9.5% error rate for the cardiac phase classification. Three of the SG methods also delivered comparable performances. Finally, we discussed the advantages and disadvantages of SG and DR methods, as well as the unsolved problems in automatic cardiac quantification for clinical practice applications. Wufeng Xue, Jiahui Li 0005, Eric Kerfoot, James R. Clough, Ilkay Öksüz, Vicente Grau, Fumin Guo, Matthew Ng, Xiang Li 0001, Quanzheng Li, Lihong Liu, Ilias Grinias, Georgios Tziritas, Angélica Atehortúa, Mireille Garreau, Yeonggul Jang, Alejandro Debus, Enzo Ferrante, Guanyu Yang 0001, Tiancong Hua, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 11 |
| 2020 | Multi-label Detection and Classification of Red Blood Cells in Microscopic ImagesabstractCell detection and cell type classification from biomedical images play an important role for high-throughput imaging and various clinical application. While classification of single cell sample can be performed with standard computer vision and machine learning methods, analysis of multi-label samples (region containing congregating cells) is more challenging, as separation of individual cells can be difficult (e.g. touching cells) or even impossible (e.g. overlapping cells). As multi-instance images are common in analyzing Red Blood Cell (RBC) for Sickle Cell Disease (SCD) diagnosis, we develop and implement a multi-instance cell detection and classification framework to address this challenge. The framework firstly trains a region proposal model based on Region-based Convolutional Network (RCNN) to obtain bounding-boxes of regions potentially containing single or multiple cells from input microscopic images, which are extracted as image patches. High-level image features are then calculated from image patches through a pre-trained Convolutional Neural Network (CNN) with ResNet-50 structure. Using these image features inputs, six networks are then trained to make multi-label prediction of whether a given patch contains cells belonging to a specific cell type. As the six networks are trained with image patches consisting of both individual cells and touching/overlapping cells, they can effectively recognize cell types that are presented in multi-instance image samples. Finally, for the purpose of SCD testing, we train another machine learning classifier to predict whether the given image patch contains abnormal cell type based on outputs from the six networks. Testing result of the proposed framework shows that it can achieve good performance in automatic cell detection and classification. Jiaming Guo, Xiang Li 0001, Mengjia Xu, Mo Zhang, Quanzheng Li |
IEEE BigData | 3 |
| 2020 | Efficient Classification via Partial Co-Training for Virtual MetrologyabstractDeveloping accurate and cost-effective classification techniques to facilitate virtual metrology is a critical task for modern manufacturing. In this paper, we consider the scenario in which labeling data is expensive, causing a shortage of labeled data. As a consequence, conventional classification methods suffer from a high risk of overfitting. To address this issue, we develop a novel semi-supervised classification method, namely Partial Cotraining with Logistic Regression (PCT-LR). PCT-LR finds a subset of the original features to generate a partial view, and uses this partial view to provide side information to support the complete view that includes all features. Both views are cooptimized in a Bayesian inference with a Gaussian process prior and a logistic regression classifier. The proposed method is validated with two industrial examples. Experiment results suggest that the amount of required labeled data can be reduced by up to 18% without loss in accuracy. Xin Li 0001, R. D. (Shawn) Blanton, Xiang Li 0001 |
ETFA | 4 |
| 2020 | Discovering Functional Brain Networks with 3D Residual Autoencoder (ResAE)
Qinglin Dong, Ning Qiang, Jinglei Lv, Xiang Li 0001, Tianming Liu 0001, Quanzheng Li |
MICCAI (7) | 4 |
| 2020 | Spatiotemporal Attention Autoencoder (STAAE) for ADHD Classification
Qinglin Dong, Ning Qiang, Jinglei Lv, Xiang Li 0001, Tianming Liu 0001, Quanzheng Li |
MICCAI (7) | 4 |
| 2020 | Automated Semantic Segmentation of Red Blood Cells for Sickle Cell DiseaseabstractRed blood cell (RBC) segmentation and classification from microscopic images is a crucial step for the diagnosis of sickle cell disease (SCD). In this work, we adopt a deep learning based semantic segmentation framework to solve the RBC classification task. A major challenge for robust segmentation and classification is the large variations on the size, shape and viewpoint of the cells, combining with the low image quality caused by noise and artifacts. To address these challenges, we apply deformable convolution layers to the classic U-Net structure and implement the deformable U-Net (dU-Net). U-Net architecture has been shown to offer accurate localization for image semantic segmentation. Moreover, deformable convolution enables free-form deformation of the feature learning process, thus making the network more robust to various cell morphologies and image settings. dU-Net is tested on microscopic red blood cell images from patients with sickle cell disease. Results show that dU-Net can achieve highest accuracy for both binary segmentation and multi-class semantic segmentation tasks, comparing with both unsupervised and state-of-the-art deep learning based supervised segmentation methods. Through detailed investigation of the segmentation results, we further conclude that the performance improvement is mainly caused by the deformable convolution layer, which has better ability to separate the touching cells, discriminate the background noise and predict correct cell shapes without any shape priors. Mo Zhang, Xiang Li 0001, Mengjia Xu, Quanzheng Li |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Predicting Alzheimer's Disease by Hierarchical Graph Convolution from Positron Emission Tomography ImagingabstractImaging-based early diagnosis of Alzheimer Disease (AD) has become an effective approach, especially by using nuclear medicine imaging techniques such as Positron Emission Topography (PET). In various literature it has been found that PET images can be better modeled as signals (e.g. uptake of florbetapir) defined on a network (non-Euclidean) structure which is governed by its underlying graph patterns of pathological progression and metabolic connectivity. In order to effectively apply deep learning framework for PET image analysis to overcome its limitation on Euclidean grid, we develop a solution for 3D PET image representation and analysis under a generalized, graph-based CNN architecture (PETNet), which analyzes PET signals defined on a group-wise inferred graph structure. Computations in PETNet are defined in non-Euclidean, graph (network) domain, as it performs feature extraction by convolution operations on spectral-filtered signals on the graph and pooling operations based on hierarchical graph clustering. Effectiveness of the PETNet is evaluated on the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset, which shows improved performance over both deep learning and other machine learning-based methods. Jiaming Guo, Xiang Li 0001, Xuandong Zhao, Quanzheng Li |
IEEE BigData | 3 |
| 2019 | Consensus Neural Network for Medical Imaging Denoising with Only Noisy Training Samples
Dufan Wu, Kuang Gong, Kyung Sang Kim, Xiang Li 0001, Quanzheng Li |
MICCAI (4) | 4 |
| 2019 | A Distributed Computing Platform for fMRI Big Data AnalyticsabstractSince the BRAIN Initiative and Human Brain Project began, a few efforts have been made to address the computational challenges of neuroscience Big Data. The promises of these two projects were to model the complex interaction of brain and behavior and to understand and diagnose brain diseases by collecting and analyzing large quanitites of data. Archiving, analyzing, and sharing the growing neuroimaging datasets posed major challenges. New computational methods and technologies have emerged in the domain of Big Data but have not been fully adapted for use in neuroimaging. In this work, we introduce the current challenges of neuroimaging in a big data context. We review our efforts toward creating a data management system to organize the large-scale fMRI datasets, and present our novel algorithms/methods for the distributed fMRI data processing that employs Hadoop and Spark. Finally, we demonstrate the significant performance gains of our algorithms/methods to perform distributed dictionary learning. Milad Makkie, Xiang Li 0001, Shannon Quinn, Jieping Ye, Geoffrey Mon, Tianming Liu 0001 |
IEEE Trans. Big Data | 2 |
| 2018 | RBC Semantic Segmentation for Sickle Cell Disease Based on Deformable U-Net
Mo Zhang, Xiang Li 0001, Mengjia Xu, Quanzheng Li |
MICCAI (4) | 2 |
| 2018 | Modeling 4D fMRI Data via Spatio-Temporal Convolutional Neural Networks (ST-CNN)
Yu Zhao 0007, Xiang Li 0001, Wei Zhang 0090, Shijie Zhao 0001, Milad Makkie, Mo Zhang, Quanzheng Li, Tianming Liu 0001 |
MICCAI (3) | 2 |
| 2016 | Distributed rank-1 dictionary learning: Towards fast and scalable solutions for fMRI big data analyticsabstractThe use of functional brain imaging for research and diagnosis has benefitted greatly from the recent advancements in neuroimaging technologies, as well as the explosive growth in size and availability of fMRI data. While it has been shown in literature that using multiple and large scale fMRI datasets can improve reproducibility and lead to new discoveries, the computational and informatics systems supporting the analysis and visualization of such fMRI big data are extremely limited and largely under-discussed. We propose to address these shortcomings in this work, based on previous success in using dictionary learning method for functional network decomposition studies on fMRI data. We presented a distributed dictionary learning framework based on rank-1 matrix decomposition with sparseness constraint (D-r1DL framework). The framework was implemented using the Spark distributed computing engine and deployed on three different processing units: an in-house server, in-house high performance clusters, and the Amazon Elastic Compute Cloud (EC2) service. The whole analysis pipeline was integrated with our neuroinformatics system for data management, user input/output, and real-time visualization. Performance and accuracy of D-r1DL on both individual and group-wise fMRI Human Connectome Project (HCP) dataset shows that the proposed framework is highly scalable. The resulting group-wise functional network decompositions are highly accurate, and the fast processing time confirm this claim. In addition, D-r1DL can provide real-time user feedback and results visualization which are vital for large-scale data analysis. Milad Makkie, Xiang Li 0001, Tianming Liu 0001, Shannon Quinn, Jieping Ye |
IEEE BigData | 2 |
| 2016 | Implementing dictionary learning in Apache Flink, Or: How I learned to relax and love iterationsabstractThe authors evaluate the use of Apache Flink, a novel data analysis framework offering optimizations over competitors such as Apache Spark, in order to use a rank-1 dictionary learning (r1DL) algorithm to decompose fMRI data. We first expand the functionality of the Flink Python API in order to accommodate the implementation of rank-1 dictionary learning, a model for decomposing a large matrix. Iterative algorithms, aggregators, and other features are added to the incomplete Python API, and the experiences and lessons learned are described. Using these features, we port an existing implementation of r1DL from using the Python API of Apache Spark to using the Python API of Apache Flink. In preliminary testing, this implementation suggests performance boosts over Spark for large input files, meriting further research. We conclude that Flink is likely a feasible tool for the application of dictionary learning to decompose fMRI data, and we continue to evaluate and apply it. Geoffrey Mon, Milad Makkie, Xiang Li 0001, Tianming Liu 0001, Shannon Quinn |
IEEE BigData | 3 |
| 2016 | Scalable Fast Rank-1 Dictionary Learning for fMRI Big Data AnalysisabstractIt has been shown from various functional neuroimaging studies that sparsity-regularized dictionary learning could achieve superior performance in decomposing comprehensive and neuroscientifically meaningful functional networks from massive fMRI signals. However, the computational cost for solving the dictionary learning problem has been known to be very demanding, especially when dealing with large-scale data sets. Thus in this work, we propose a novel distributed rank-1 dictionary learning (D-r1DL) model and apply it for fMRI big data analysis. The model estimates one rank-1 basis vector with sparsity constraint on its loading coefficient from the input data at each learning step through alternating least squares updates. By iteratively learning the rank-1 basis and deflating the input data at each step, the model is then capable of decomposing the whole set of functional networks. We implement and parallelize the rank-1 dictionary learning algorithm using Spark engine and deployed the resilient distributed dataset (RDDs) abstracts for the data distribution and operations. Experimental results from applying the model on the Human Connectome Project (HCP) data show that the proposed D-r1DL model is efficient and scalable towards fMRI big data analytics, thus enabling data-driven neuroscientific discovery from massive fMRI big data in the future. Xiang Li 0001, Milad Makkie, Mojtaba Sedigh Fazli, Ian Davidson, Jieping Ye, Tianming Liu 0001, Shannon Quinn |
KDD | 1 |
| 2016 | Modeling Functional Dynamics of Cortical Gyri and Sulci
Xi Jiang 0001, Xiang Li 0001, Jinglei Lv, Shijie Zhao 0001, Shu Zhang 0001, Wei Zhang 0090, Tianming Liu 0001 |
MICCAI (1) | 2 |
| 2016 | Discover Mouse Gene Coexpression Landscape Using Dictionary Learning and Sparse Coding
Yujie Li 0004, Hanbo Chen, Xi Jiang 0001, Xiang Li 0001, Jinglei Lv, Hanchuan Peng, Joe Z. Tsien, Tianming Liu 0001 |
MICCAI (1) | 4 |
| 2016 | Predicting Movie Trailer Viewer's "Like/Dislike" via Learned Shot Editing PatternsabstractNowadays, there are many movie trailers publicly available on social media website such as YouTube, and many thousands of users have independently indicated whether they like or dislike those trailers. Although it is understandable that there are multiple factors that could influence viewers' like or dislike of the trailer, we aim to address a preference question in this work: Can subjective multimedia features be developed to predict the viewer's preference presented by like (by thumbs-up) or dislike (by thumbs-down) during and after watching movie trailers? We designed and implemented a computational framework that is composed of low-level multimedia feature extraction, feature screening and selection, and classification, and applied it to a collection of 725 movie trailers. Experimental results demonstrated that, among dozens of multimedia features, the single low-level multimedia feature of shot length variance is highly predictive of a viewer's “like/dislike” for a large portion of movie trailers. We interpret these findings such that variable shot lengths in a trailer tend to produce a rhythm that is likely to stimulate a viewer's positive preference. This conclusion was also proved by the repeatability experiments results using another 600 trailer videos and it was further interpreted by viewers'eye-tracking data. Shu Zhang 0001, Xi Jiang 0001, Xiang Li 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002, L. Stephen Miller, Richard Neupert, Tianming Liu 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2015 | Sparse representation of whole-brain fMRI signals for identification of functional networks
Jinglei Lv, Xi Jiang 0001, Xiang Li 0001, Dajiang Zhu, Hanbo Chen, Shu Zhang 0001, Xintao Hu, Junwei Han 0001, Heng Huang 0001, Jing Zhang 0010, Lei Guo 0002, Tianming Liu 0001 |
Medical Image Anal. | 3 |
| 2013 | Modeling Dynamic Functional Information Flows on Large-Scale Brain Networks
Peili Lv, Lei Guo 0002, Xintao Hu, Xiang Li 0001, Changfeng Jin, Junwei Han 0001, Lingjiang Li, Tianming Liu 0001 |
MICCAI (2) | 4 |
| 2013 | Sparse Representation of Group-Wise FMRI Signals
Jinglei Lv, Xiang Li 0001, Dajiang Zhu, Xi Jiang 0001, Xin Zhang 0151, Xintao Hu, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (3) | 2 |
| 2013 | Sparse Representation of Higher-Order Functional Interaction Patterns in Task-Based FMRI Data
Shu Zhang 0001, Xiang Li 0001, Jinglei Lv, Xi Jiang 0001, Dajiang Zhu, Hanbo Chen, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (3) | 2 |
| 2013 | Characterization of task-free and task-performance brain states via functional connectome patterns
Xin Zhang 0151, Lei Guo 0002, Xiang Li 0001, Dajiang Zhu, Kaiming Li, Hanbo Chen, Jinglei Lv, Changfeng Jin, Lingjiang Li, Tianming Liu 0001 |
Medical Image Anal. | 3 |
| 2012 | Characterization of Task-Free/Task-Performance Brain States
Xin Zhang 0151, Lei Guo 0002, Xiang Li 0001, Dajiang Zhu, Kaiming Li, Zhenqiang Sun, Changfeng Jin, Xintao Hu, Junwei Han 0001, Lingjiang Li, Tianming Liu 0001 |
MICCAI (2) | 3 |
| 2011 | Fiber-Centered Granger Causality Analysis
Xiang Li 0001, Kaiming Li, Lei Guo 0002, Chulwoo Lim, Tianming Liu 0001 |
MICCAI (2) | 1 |