EDBT 2026 Demo / reviewers in the wild / expert
Li Xiao 0005
dblp:14/5505-5
· DBLP profile ↗
20ranked-venue papers
4as first author
17since 2021 · last 2027
0000-0002-3063-0869ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Multimodal cancer survival prediction with optimal transport reconstruction and fisher-guided modal reweighting
Le Feng, Li Xiao 0005 |
Expert Syst. Appl. | 2 |
| 2026 | Zero-shot diverse audio captioning with diffusion models
Yonggang Zhu, Yiming Zhang 0025, Li Xiao 0005, Wenwu Wang 0001, Aidong Men |
Knowl. Based Syst. | 3 |
| 2026 | Microscale-Searching Optimization for Transfer Learning-Based Filter Fine-TuningabstractFine-tuning has emerged as a popular technique in the field of transfer learning, demonstrating remarkable achievements in various data-scarce tasks. The performance of fine-tuning in deep convolutional neural networks depends on the selection of which parameters to fine-tune and freeze. However, it is difficult to determine which parameters in the pre-trained model need to be fine-tuned for a new task. This article proposes a filter-level discrete optimization model to identify the filter subset for fine-tuning, a core step of filter selection coding optimization. Due to the huge search space of the filter fine-tuning problem, we propose a filter interactivity decomposition strategy to find a valid search subspace (a smaller search subspace containing the optimal solution) by dividing the entire filter fine-tuning problem into multiple suboptimization problems. Based on the decomposition strategy, we design a microscale-searching transfer optimization algorithm, which solves each subproblem by searching the valid search subspace instead of the original search space of the filter fine-tuning problem. To verify the validity of the proposed algorithm, extensive experiments are conducted on seven publicly available image classification datasets: Stanford Dogs, MIT Indoors, Caltech 256-30, Caltech 256-60, Aircraft, UCF-101, and Omniglot. Experimental results show that the proposed method significantly improves the fine-tuning accuracy while effectively reducing the filter fine-tuning problem scale. Moreover, the proposed algorithm outperforms the state-of-the-art fine-tuning methods on the fine-tuning problem for transfer learning. Le Feng, Fujian Feng, Li Xiao 0005, Mian Tan, Han Huang 0002 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2025 | GEMPA: Graph-Enhanced Multimodal Prototype Alignment for Class Incremental LearningabstractPre-trained models (PTMs), recognized for their generalization capabilities, have been extensively employed in Class Incremental Learning (CIL). However, existing PTM-based CIL methods encounter two critical limitations: 1) suboptimal extraction of discriminative prototypes for classifier construction, and 2) insufficient modeling of inter-class relationships, collectively leading to catastrophic forgetting of prior knowledge when incrementally learning new classes. To address these challenges, we propose Graph-Enhanced Multimodal Prototype Alignment (GEMPA) which synergizes PTM representations with class-specific semantic information. Our framework features two key innovations: First, an Embedding Modulation (EM) module rectifies the imbalanced value distributions (characterized by Leptokurtic distribution) in PTM-generated embeddings. Second, a GCN-based Multimodal Prototype Similarity Alignment (GPA) module establishes a Visual-Semantic Co-occurrence (VSC) space through CLIP-derived semantic embeddings. This VSC space aligns visual prototypes with their corresponding semantic counterparts to model incremental class relationships. A Weighted Hybrid Loss is designed to penalize similar prototypes while maintaining the intra-class distribution in GCN. We incorporate GEMPA with seven PTM-based CIL approaches and validate them across four benchmark datasets. The results demonstrate the improvements through semantic clarity of prototypes and forgetting mitigation. Code is available at: https://github.com/AI4MyBUPT/GEMPA Zhenming Zhang, Chunxia Ren, Zixuan Zhong, Li Xiao 0005 |
ECAI | 5 |
| 2025 | Generative Adversarial Networks With Noise Optimization and Pyramid Coordinate Attention for Robust Image DenoisingabstractImage denoising is a significant challenge in computer vision. While many models perform well in low‐noise environments, their denoising capabilities are relatively weak under high‐noise conditions. In addition, these models often overlook the robustness issues under adversarial attacks, leading to a marked decrease in denoising stability when facing malicious attacks. To address the challenges of achieving consistently high‐quality denoising in both high‐noise and low‐noise environments, adapting to various complex scenarios with high robustness, and enhancing the model’s resilience against attacks, we propose the NOP‐GAN, a powerful image denoising model. This model modifies the GAN architecture by integrating a U‐Net with a pyramid coordinate attention mechanism and a noise optimization algorithm into a generator of the GAN. Experimental results demonstrate that the NOP‐GAN possesses superior performance in denoising tasks and robustness against adversarial attacks. Minling Zhu, Jiahua Yuan, En Kong, Li Xiao 0005, Dong-bing Gu |
Int. J. Intell. Syst. | 5 |
| 2025 | Noise Optimization in Artificial Neural NetworksabstractArtificial neural network (ANN) has been widely used in automation. However, the vulnerability of ANN under certain attacks poses a security threat to critical automation systems. Previous research has shown that adding noise to ANNs can enhance robustness. Nonetheless, striking a balance between robustness and task performance remains challenging, as excessive noise improves robustness but hampers performance, while low noise offers minor robustness improvement. In this work, we propose to learn the distribution of optimal injected noise, which improves the robustness as well as maintains the performance. Specifically, we compute the pathwise stochastic gradient estimate with respect to the standard deviation of the Gaussian noise added to each neuron of the ANN and optimize both the noise distribution and model parameters during training with negligible additional computational cost. In numerical experiments, our proposed method can achieve significant performance improvement on the robustness of several popular ANN structures under both black box and white box attacks. We also evaluate the proposed technique on two automation tasks: the classic reinforcement learning task of the cart pole game and a fault detection problem. Our results showed that the proposed technique outperforms a conventional neural network in terms of performance, robustness, and visual explainability.Note to Practitioners—The robustness of artificial neural networks is a critical consideration in automation applications as real-world data is often subject to unforeseen perturbations from the environment, potentially causing AI systems to behave unpredictably and unstably. For example, object detection is a widely employed AI technique in automation applications. However, current object detection systems are vulnerable to noise perturbation. Even small, imperceptible noise can lead the model to malfunction. Our work focuses on improving the robustness of neural networks. We propose a novel technique that can be added to any layer of existing neural networks to enhance robustness. Extensive experiments conducted in various scenarios have verified the effectiveness of the proposed method in enhancing both performance and robustness. Li Xiao 0005, Zeliang Zhang 0001, Kuihua Huang, Jinyang Jiang 0001, Yijie Peng |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2023 | Radiology report generation with a learned knowledge base and multi-modal alignmentabstractIn clinics, a radiology report is crucial for guiding a patient's treatment. However, writing radiology reports is a heavy burden for radiologists. To this end, we present an automatic, multi-modal approach for report generation from a chest x-ray. Our approach, motivated by the observation that the descriptions in radiology reports are highly correlated with specific information of the x-ray images, features two distinct modules: (i) Learned knowledge base: To absorb the knowledge embedded in the radiology reports, we build a knowledge base that can automatically distill and restore medical knowledge from textual embedding without manual labor; (ii) Multi-modal alignment: to promote the semantic alignment among reports, disease labels, and images, we explicitly utilize textual embedding to guide the learning of the visual feature space. We evaluate the performance of the proposed model using metrics from both natural language generation and clinic efficacy on the public IU-Xray and MIMIC-CXR datasets. Our ablation study shows that each module contributes to improving the quality of generated reports. Furthermore, the assistance of both modules, our approach outperforms state-of-the-art methods over almost all the metrics. Code is available at https://github.com/LX-doctorAI1/M2KT. Shuxin Yang, Xian Wu 0001, Shen Ge, Zhuozhao Zheng, Shaohua Kevin Zhou, Li Xiao 0005 |
Medical Image Anal. | 6 |
| 2022 | DeltaNet: Conditional Medical Report Generation for COVID-19 DiagnosisabstractFast screening and diagnosis are critical in COVID-19 patient treatment. In addition to the gold standard RT-PCR, radiological imaging like X-ray and CT also works as an important means in patient screening and follow-up. However, due to the excessive number of patients, writing reports becomes a heavy burden for radiologists. To reduce the workload of radiologists, we propose DeltaNet to generate medical reports automatically. Different from typical image captioning approaches that generate reports with an encoder and a decoder, DeltaNet applies a conditional generation process. In particular, given a medical image, DeltaNet employs three steps to generate a report: 1) first retrieving related medical reports, i.e., the historical reports from the same or similar patients; 2) then comparing retrieved images and current image to find the differences; 3) finally generating a new report to accommodate identified differences based on the conditional report. We evaluate DeltaNet on a COVID-19 dataset, where DeltaNet outperforms state-of-the-art approaches. Besides COVID-19, the proposed DeltaNet can be applied to other diseases as well. We validate its generalization capabilities on the public IU-Xray and MIMIC-CXR datasets for chest-related diseases. Xian Wu 0001, Shuxin Yang, Zhaopeng Qiu, Shen Ge, Yangtian Yan, Xingwang Wu, Yefeng Zheng 0001, Shaohua Kevin Zhou, Li Xiao 0005 |
COLING | 9 |
| 2022 | Learning Incrementally to Segment Multiple Organs in a CT Image
Pengbo Liu 0004, Mengsi Fan, Hongli Pan, Minmin Yin, Xiaohong Zhu, Dandan Du, Xiaoying Zhao, Li Xiao 0005, Lian Ding, Xingwang Wu, Shaohua Kevin Zhou |
MICCAI (4) | 9 |
| 2022 | A New Likelihood Ratio Method for Training Artificial Neural NetworksabstractWe investigate a new approach to compute the gradients of artificial neural networks (ANNs), based on the so-called push-out likelihood ratio method. Unlike the widely used backpropagation (BP) method that requires continuity of the loss function and the activation function, our approach bypasses this requirement by injecting artificial noises into the signals passed along the neurons. We show how this approach has a similar computational complexity as BP, and moreover is more advantageous in terms of removing the backward recursion and eliciting transparent formulas. We also formalize the connection between BP, a pivotal technique for training ANNs, and infinitesimal perturbation analysis, a classic path-wise derivative estimation approach, so that both our new proposed methods and BP can be better understood in the context of stochastic gradient estimation. Our approach allows efficient training for ANNs with more flexibility on the loss and activation functions, and shows empirical improvements on the robustness of ANNs under adversarial attacks and corruptions of natural noises. Summary of Contribution: Stochastic gradient estimation has been studied actively in simulation for decades and becomes more important in the era of machine learning and artificial intelligence. The stochastic gradient descent is a standard technique for training the artificial neural networks (ANNs), a pivotal problem in deep learning. The most popular stochastic gradient estimation technique is the backpropagation method. We find that the backpropagation method lies in the family of infinitesimal perturbation analysis, a path-wise gradient estimation technique in simulation. Moreover, we develop a new likelihood ratio-based method, another popular family of gradient estimation technique in simulation, for training more general ANNs, and demonstrate that the new training method can improve the robustness of the ANN. Yijie Peng, Li Xiao 0005, Bernd Heidergott, L. Jeff Hong, Henry Lam |
INFORMS J. Comput. | 2 |
| 2022 | Knowledge matters: Chest radiology report generation with general and specific knowledgeabstractAutomatic chest radiology report generation is critical in clinics which can relieve experienced radiologists from the heavy workload and remind inexperienced radiologists of misdiagnosis or missed diagnose. Existing approaches mainly formulate chest radiology report generation as an image captioning task and adopt the encoder-decoder framework. However, in the medical domain, such pure data-driven approaches suffer from the following problems: 1) visual and textual bias problem; 2) lack of expert knowledge. In this paper, we propose a knowledge-enhanced radiology report generation approach introduces two types of medical knowledge: 1) General knowledge, which is input independent and provides the broad knowledge for report generation; 2) Specific knowledge, which is input dependent and provides the fine-grained knowledge for chest X-ray report generation. To fully utilize both the general and specific knowledge, we also propose a knowledge-enhanced multi-head attention mechanism. By merging the visual features of the radiology image with general knowledge and specific knowledge, the proposed model can improve the quality of generated reports. The experimental results on the publicly available IU-Xray dataset show that the proposed knowledge-enhanced approach outperforms state-of-the-art methods in almost all metrics. And the results of MIMIC-CXR dataset show that the proposed knowledge-enhanced approach is on par with state-of-the-art methods. Ablation studies also demonstrate that both general and specific knowledge can help to improve the performance of chest radiology report generation. Shuxin Yang, Xian Wu 0001, Shen Ge, Shaohua Kevin Zhou, Li Xiao 0005 |
Medical Image Anal. | 5 |
| 2021 | AMA-GCN: Adaptive Multi-layer Aggregation Graph Convolutional Network for Disease PredictionabstractRecently, Graph Convolutional Networks (GCNs) have proven to be a powerful mean for Computer Aided Diagnosis (CADx). This approach requires building a population graph to aggregate structural information, where the graph adjacency matrix represents the relationship between nodes. Until now, this adjacency matrix is usually defined manually based on phenotypic information. In this paper, we propose an encoder that automatically selects the appropriate phenotypic measures according to their spatial distribution, and uses the text similarity awareness mechanism to calculate the edge weights between nodes. The encoder can automatically construct the population graph using phenotypic measures which have a positive impact on the final results, and further realizes the fusion of multimodal information. In addition, a novel graph convolution network architecture using multi-layer aggregation mechanism is proposed. The structure can obtain deep structure information while suppressing over-smooth, and increase the similarity between the same type of nodes. Experimental results on two databases show that our method can significantly improve the diagnostic accuracy for Autism spectrum disorder and breast cancer, indicating its universality in leveraging multimodal data for disease prediction. Hao Chen 0163, Fuzhen Zhuang, Li Xiao 0005, Ling Ma 0005, Ruifang Zhang, Huiqin Jiang, Qing He 0003 |
IJCAI | 3 |
| 2021 | One-Shot Medical Landmark Detection
Qingsong Yao, Quan Quan, Li Xiao 0005, Shaohua Kevin Zhou |
MICCAI (2) | 3 |
| 2021 | You only Learn Once: Universal Anatomical Landmark Detection
Heqin Zhu, Qingsong Yao, Li Xiao 0005, Shaohua Kevin Zhou |
MICCAI (5) | 3 |
| 2021 | An Efficient Polyp Detection Framework with Suspicious Targets Assisted Training
Li Xiao 0005, Fuzhen Zhuang, Ling Ma 0005, Huiqin Jiang, Qing He 0003 |
PRCV (4) | 2 |
| 2021 | Marginal loss and exclusion loss for partially supervised multi-organ segmentation
Gonglei Shi, Li Xiao 0005, Yang Chen 0008, Shaohua Kevin Zhou |
Medical Image Anal. | 2 |
| 2021 | Label-Free Segmentation of COVID-19 Lesions in Lung CTabstractScarcity of annotated images hampers the building of automated solution for reliable COVID-19 diagnosis and evaluation from CT. To alleviate the burden of data annotation, we herein present a label-free approach for segmenting COVID-19 lesions in CT via voxel-level anomaly modeling that mines out the relevant knowledge from normal CT lung scans. Our modeling is inspired by the observation that the parts of tracheae and vessels, which lay in the high-intensity range where lesions belong to, exhibit strong patterns. To facilitate the learning of such patterns at a voxel level, we synthesize 'lesions' using a set of simple operations and insert the synthesized 'lesions' into normal CT lung scans to form training pairs, from which we learn a normalcy-recognizing network (NormNet) that recognizes normal tissues and separate them from possible COVID-19 lesions. Our experiments on three different public datasets validate the effectiveness of NormNet, which conspicuously outperforms a variety of unsupervised anomaly detection (UAD) methods. Qingsong Yao, Li Xiao 0005, Peihang Liu, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 2 |
| 2020 | DeepACEv2: Automated Chromosome Enumeration in Metaphase Cell Images Using Deep Convolutional Neural NetworksabstractChromosome enumeration is an essential but tedious procedure in karyotyping analysis. To automate the enumeration process, we develop a chromosome enumeration framework, DeepACEv2, based on the region based object detection scheme. The framework is developed following three steps. Firstly, we take the classical ResNet-101 as the backbone and attach the Feature Pyramid Network (FPN) to the backbone. The FPN takes full advantage of the multiple level features, and we only output the level of feature map that most of the chromosomes are assigned to. Secondly, we enhance the region proposal network's ability by adding a newly proposed Hard Negative Anchors Sampling to extract unapparent but essential information about highly confusing partial chromosomes. Next, to alleviate serious occlusion problems, besides the traditional detection branch, we novelly introduce an isolated Template Module branch to extract unique embeddings of each proposal by utilizing the chromosome's geometric information. The embeddings are further incorporated into the No Maximum Suppression (NMS) procedure to improve the detection of overlapping chromosomes. Finally, we design a Truncated Normalized Repulsion Loss and add it to the loss function to avoid inaccurate localization caused by occlusion. In the newly collected 1375 metaphase images that came from a clinical laboratory, a series of ablation studies validate the effectiveness of each proposed module. Combining them, the proposed DeepACEv2 outperforms all the previous methods, yielding the Whole Correct Ratio(WCR)(%) with respect to images as 71.39, and the Average Error Ratio(AER)(%) with respect to chromosomes as about 1.17. Li Xiao 0005, Chunlong Luo, Tianqi Yu, Yufan Luo, Manqing Wang, Fuhai Yu, Chan Tian, Jie Qiao |
IEEE Trans. Medical Imaging | 1 |
| 2019 | DeepACE: Automated Chromosome Enumeration in Metaphase Cell Images Using Deep Convolutional Neural Networks
Li Xiao 0005, Chunlong Luo, Yufan Luo, Tianqi Yu, Chan Tian, Jie Qiao, Yi Zhao 0013 |
MICCAI (1) | 1 |
| 2019 | Learning from Suspected Target: Bootstrapping Performance for Breast Cancer Detection in Mammography
Li Xiao 0005, Chunlong Luo, Peifang Liu, Yi Zhao 0013 |
MICCAI (6) | 1 |