EDBT 2026 Demo / reviewers in the wild / expert
Xiaoling Luo 0001
dblp:96/9539-1
· DBLP profile ↗
57ranked-venue papers
9as first author
55since 2021 · last 2027
0000-0003-3678-3185ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 8 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | TUS-DET: Open-vocabulary pretraining with standard planes for few-shot thyroid ultrasound lesion detection
Jiansong Zhang 0005, Shunlan Liu, Xiaoling Luo 0001, Guorong Lyu, LinLin Shen |
Expert Syst. Appl. | 4 |
| 2027 | RetinaFormer: Retina inspired transformer for single image dehazing
Maowei Zeng, Xiaoling Luo 0001, Ping Li 0024, Hong Qu 0002 |
Expert Syst. Appl. | 3 |
| 2026 | VPSentry: Semi-supervised Video Polyp Segmentation via Sentry-guided Long-term Prototype Fusion with Correlation Dynamic PropagationabstractAutomated polyp segmentation in colonoscopy videos is an essential computer-aided technology for early detection and removal of polyps. However, most existing video polyp segmentation methods are designed with pixel-level temporal learning mechanisms, at the cost of time-consuming frame-wise annotations. In this paper, we present VPSentry, a novel semi-supervised segmentation model with a sentry mechanism. Our model integrates a prototype memory to store the long-term spatiotemporal cues of colonoscopy videos. Moreover, we devise adaptive prototypes to capture and generalize critical representations from individual frames, enabling long-term temporal fusion across labeled and unlabeled frames. In addition, we propose a correlation dynamic propagation module that propagates information from prototypes to features while simultaneously extracting dynamic features to perceive variations in polyp details between adjacent frames. Since colonoscopy scenes may change among consecutive frames, we further employ a sentry mechanism to assess the inter-frame continuity. This mechanism guides the prototype memory updating and the correlation dynamic propagation, further facilitating robust temporal propagation and dynamic detail perception for semi-supervised learning of long-term colonoscopy video sequences. Extensive experiments on the large-scale SUN-SEG dataset demonstrate that our model achieves optimal segmentation performance with real-time inference efficiency. Guilian Chen, Xiaoling Luo 0001, Huisi Wu, Harry Qin |
AAAI | 2 |
| 2026 | Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease DiagnosisabstractMultimodal fusion of color fundus photography (CFP) and optical coherence tomography (OCT) B-scan images has demonstrated superior diagnostic potential for retinal diseases compared to single-modality approaches. However, existing fusion paradigms - whether through naive concatenation or attention mechanisms - treat cross-modal interactions indiscriminately, lacking adaptive modulation of modality-specific contributions under varying clinical scenarios. We propose an adaptive fusion framework that dynamically routes and refines multimodal signals for enhancing disease recognition. The framework comprises two key components: 1) Dynamic Cross-Modal Expert Routing (CMER), which selectively activates convolutional neural network (CNN) experts from one modality based on contextual guidance from the other, ensuring only the most relevant feature extractors contribute to fusion; and 2) Top-K Expert-Guided Wavelet Fusion (TEWF), which performs discrete wavelet transform (DWT) to decompose selected features into low- and high-frequency subbands. Cross-modal attention is then applied specifically to high-frequency components, where lesion-specific microstructures reside, enabling frequency-aware fusion. Finally, inverse DWT (IDWT) reconstructs the fused representation, weighted by CMER-derived importance scores to amplify informative modality cues while suppressing redundancy. Experimental validation on two multimodal retinal datasets demonstrates that our method achieves state-of-the-art performance, outperforming existing fusion strategies by significant margins in disease classification accuracy and robustness. Haoran Li 0024, Haoyu Cao 0002, Yongting Hu, Qihao Xu, Chengliang Liu 0003, Xiaoling Luo 0001, Zhihao Wu 0002, Yong Xu 0001, Wei Wang 0169 |
AAAI | 7 |
| 2026 | Vision-Language Models Guided Graph Concept Reasoning for Interpretable Diabetic Retinopathy DiagnosisabstractDeep neural networks (DNNs) have significantly advanced diabetic retinopathy (DR) diagnosis, yet their black-box nature limits clinical acceptance due to a lack of interpretability. Concept bottleneck model (CBM) offers a promising solution by enabling concept-level reasoning and test-time intervention, with recent DR studies modeling lesions as concepts and grades as outcomes. However, current methods often ignore relationships between lesion concepts across different DR grades and struggle when fine-grained lesion concepts are unavailable, limiting their interpretability and real-world applicability. To bridge these gaps, we propose VLM-GCR, a vision-language model guided graph concept reasoning framework for interpretable DR diagnosis. VLM-GCR emulates the diagnostic process of ophthalmologists by constructing a grading-aware lesion concept graph that explicitly models the interactions among lesions and their relationships to disease grades. In concept-free clinical scenarios, our method introduces a vision-language guided dynamic concept pseudo-labeling mechanism to mitigate the challenges of existing concept-based models in fine-grained lesion recognition. Additionally, we introduce a multi-level intervention method that supports error correction, enabling transparent and robust human-AI collaboration. Experiments on two public DR benchmarks show that VLM-GCR achieves strong performance in both lesion and grading tasks, while delivering clear and clinically meaningful reasoning steps. Qihao Xu, Xiaoling Luo 0001, Chengliang Liu 0003, Yongting Hu, Xinheng Lyu, Yong Xu 0001 |
AAAI | 2 |
| 2026 | Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language ModelsabstractShaonan Liu, Guo Yu, Xiaoling Luo, Shiyi Zheng, Jie Liu, Wenting Chen, Linlin Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaonan Liu, Xiaoling Luo 0001, Shiyi Zheng, Jie Liu 0044, Wenting Chen, LinLin Shen |
ACL (1) | 3 |
| 2026 | A bionic spiking sequence memory model with minicolumns, dendrites and oscillation for text retrieval
Xinlin Pu, Xiaoling Luo 0001, Ping Li 0024, Hong Qu 0002 |
Expert Syst. Appl. | 3 |
| 2026 | Text condition embedded regression network for automated dental implant abutment design
Mianjie Zheng, Xinquan Yang, Xiaoling Luo 0001, Xuefen Liu, He Meng 0008, LinLin Shen |
Expert Syst. Appl. | 4 |
| 2026 | FProtoSeg: Fine-grained prototype alignment for Weakly Supervised Semantic Segmentation of histopathology images
Meidan Ding, Wenting Chen, Xiaoling Luo 0001, Haiqin Zhong, LinLin Shen |
Pattern Recognit. | 3 |
| 2026 | M2GO: Multimodal protein function prediction via heterogeneous expert interaction
Xiaoling Luo 0001, Ruli Zheng, Xiaopeng Jin, Baoyi Pan, Bob Zhang |
Pattern Recognit. | 1 |
| 2026 | SDC-Net: Semi-supervised breast ultrasound lesion segmentation via semantic decoupling
Jiansong Zhang 0005, Zhuoqin Yang, Xiaoling Luo 0001, Shaozheng He, LinLin Shen |
Pattern Recognit. | 3 |
| 2026 | ACGM: Attribute-Centric Graph Modeling Network for Concurrent Missing Tabular Data Imputation and COVID-19 PrognosisabstractCOVID-19 prognosis using clinical tabular data faces significant challenges due to missing values and class imbalance issues. Existing methods often overlook the complex high-order interrelationship among clinical attributes and struggle with training stability on imbalanced datasets. We propose ACGM, an attribute-centric graph modeling network that simultaneously addresses missing data imputation and COVID-19 prognosis. ACGM consists of three key modules: an attributes preprocessing module (APM) for coarse-grained imputation initialization, a graph-enhanced attributes imputation module (GEAIM) that models high-order inter-attribute relationships through graph structures, and a graph-enhanced disease prognosis module (GEDPM) that leverages these complex attribute interactions for final prediction. GEAIM and GEDPM employ a mean-teacher strategy with attributes graph matching to preserve high-order relationships, enhance training stability, and maintain structural integrity of attribute interactions. Extensive experiments are conducted on four public COVID-19 tabular datasets, demonstrating the superiority of our ACGM over existing methods. Through comprehensive interpretability analysis, we identify that attributes such as LDH, Difficulty In Breathing, and SaO2 significantly impact COVID-19 prognosis, aligning well with clinical insights and radiologist assessments. Zhuoru Wu, Wenting Chen, Xuechen Li 0001, Filippo Ruffini, Shaonan Liu, Lorenzo Tronchin, Domenico Albano, Eliodoro Faiella, Deborah Fazzini, Domiziana Santucci, Xiaoling Luo 0001, Valerio Guarrasi, Paolo Soda, LinLin Shen |
IEEE J. Biomed. Health Informatics | 11 |
| 2026 | Multi-View Hilbert Curve-Based Hierarchical Information Aggregation for Incomplete Multimodal Alzheimer's Disease DiagnosisabstractTimely identification of Alzheimer's disease (AD) benefits from combining neuroimaging, fluid biomarkers, and cognitive assessments, yet in practice one or more modalities are often unavailable due to various factors such as cost, patient compliance, and procedural risks. Furthermore, conventional convolutional neural network (CNN) architectures and even Transformer-based models struggle to efficiently capture both local and global dependencies, especially when dealing with high-dimensional and highly heterogeneous medical data. In this study, we introduce a novel hierarchical information aggregation and dynamic fusion (HI-AD) framework for incomplete multimodal AD diagnosis. Our method couples a multi-view Hilbert curve-guided Mamba block with hierarchical spatial feature extraction to retain spatial continuity, model long-range dependencies, and integrate local context in neuroimaging data. To balance semantic alignment and modality-specific information, we propose a unified mutual information-driven learning objective with an active confidence evaluation mechanism, thereby preventing modality collapse and promoting robust representation learning. Extensive experiments on real-world datasets validate that our HI-AD framework consistently outperforms existing state-of-the-art methods across a diverse range of modality-missing scenarios, establishing an effective and generalizable solution for early-stage AD screening in heterogeneous clinical data environments. Chengliang Liu 0003, Yuanxi Que, Wai Keung Wong, Xiaoling Luo 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Thyro-LMD: A Benchmark Dataset and Sample-Driven Data Loading, Attention, and Regularization for Long-Tailed Multi-Label Thyroid Ultrasound DiagnosisabstractDeveloping robust and effective computer-aided diagnostic (CAD) methods for thyroid ultrasound (TUS) remains a key challenge in medical imaging. Prior work has largely focused on binary or multi-class lesion classification, whereas real-world diagnosis follows standardized guidelines based on combinations of lexicon-level descriptors. These combinations naturally exhibit long-tailed distributions due to epidemiological patterns, limiting the robustness and generalizability of existing methods. Motivated by this, we introduce Thyro-LMD, the first long-tailed multi-label dataset for TUS. Using histopathology as the reference, Thyro-LMD provides retrospective, fine-grained annotations aligned with ACR TI-RADS lexicons and reveals a highly imbalanced label distribution. We benchmark representative methods, including end-to-end models, general-purpose multimodal large models (e.g., GPT-4o), and pretrained foundation models. While some methods show reasonable head-class performance, they struggle with body and tail classes. We therefore propose SynTUS-Net, a purpose-built baseline comprising collaborative modules addressing long-tailed multi-label challenges across data loading, feature encoding, and prediction regularization. SynTUS-Net achieves leading performance on Thyro-LMD, outperforming conventional traditional SOTA models by 5.3 Micro-F1 and 11.83 Macro-F1, and exceeding GPT-4o by 42.76 on Tail-F1. Extensive ablation studies confirm the contribution of each module. We believe Thyro-LMD and SynTUS-Net establish a clinically grounded benchmark and a new paradigm for interpretable and generalizable AI in ultrasound. Code and data will be released here. Jiansong Zhang 0005, Shunlan Liu, Xiaoling Luo 0001, Guorong Lyu, LinLin Shen |
IEEE Trans. Medical Imaging | 3 |
| 2025 | DAMPER: A Dual-Stage Medical Report Generation Framework with Coarse-Grained MeSH Alignment and Fine-Grained Hypergraph MatchingabstractMedical report generation is crucial for clinical diagnosis and patient management, summarizing diagnoses and recommendations based on medical imaging. However, existing work often overlook the clinical pipeline involved in report writing, where physicians typically conduct an initial quick review followed by a detailed examination. Moreover, current alignment methods may lead to misaligned relationships. To address these issues, we propose DAMPER, a dual-stage framework for medical report generation that mimics the clinical pipeline of report writing in two stages. In the first stage, a MeSH-Guided Coarse-Grained Alignment (MCG) stage that aligns chest X-ray (CXR) image features with medical subject headings (MeSH) features to generate a rough keyphrase representation of the overall impression. In the second stage, a Hypergraph-Enhanced Fine-Grained Alignment (HFG) stage that constructs hypergraphs for image patches and report annotations, modeling high-order relationships within each modality and performing hypergraph matching to capture semantic correlations between image regions and textual phrases. Finally,the coarse-grained visual features, generated MeSH representations, and visual hypergraph features are fed into a report decoder to produce the final medical report. Extensive experiments on public datasets demonstrate the effectiveness of DAMPER in generating comprehensive and accurate medical reports, outperforming state-of-the-art methods across various evaluation metrics. Wenting Chen, Jie Liu 0044, Qisheng Lu, Xiaoling Luo 0001, LinLin Shen |
AAAI | 5 |
| 2025 | Deep Hierarchies and Invariant Disease-Indicative Feature Learning for Computer Aided Diagnosis of Multiple Fundus DiseasesabstractWith the advancement of computer vision, numerous models have been proposed for screening of fundus diseases. However, the recognition of multiple fundus diseases is often hampered by the simultaneous presence of multiple disease types and the confluence of lesion types in fundus images. This paper addresses these challenges by conceptualizing them as multi-level feature fusion and self-supervised disease-indicative feature learning problems. We decode fundus images at various levels of granularity to delineate scenarios wherein multiple diseases and lesions co-occur. To effectively integrate these features, we introduce a hierarchical vision transformer (HVT) that adeptly captures both inter-level and intra-level dependencies. A novel forward-attention module is proposed to enhance the integration of lower-level semantic information into higher semantic layers, thereby enriching the representation of complex features. Additionally, we introduce a novel self-supervised mask-consistent feature learner (MCFL). Unlike traditional mask-autoencoders that reconstruct original images using encoder-decoder structures, MCFL utilizes a teacher-student framework to reconstruct mask-consistent feature maps. In this setup, exponential moving averaging is employed to derive classification-guided features, serving as labels for reconstruction rather than merely reconstructing the original images. This innovative approach facilitates the extraction of disease-indicative features. Extensive experiments demonstrate that our method significantly outperforms existing state-of-the-art models. Wei Wang 0169, Xiaoling Luo 0001, Zhihao Wu 0002, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001 |
AAAI | 3 |
| 2025 | Like an Ophthalmologist: Dynamic Selection Driven Multi-View Learning for Diabetic Retinopathy GradingabstractDiabetic retinopathy (DR), with its large patient population, has become a formidable threat to human visual health. In the clinical diagnosis of DR, multi-view fundus images are considered to be more suitable for DR diagnosis because of the wide coverage of the field of view. Therefore, different from most of the previous single-view DR grading methods, we design a dynamic selection-driven multi-view DR grading method to fit clinical scenarios better. Since lesion information plays a key role in DR diagnosis, previous methods usually boost the model performance by enhancing the lesion feature. However, during the actual diagnosis, ophthalmologists not only focus on the crucial parts, but also exclude irrelevant features to ensure the accuracy of judgment. To this end, we introduce the idea of dynamic selection and design a series of selection mechanisms from fine granularity to coarse granularity. In this work, we first introduce an Ophthalmic Image Reader (OIR) agent to provide the model with pixel-level prompts of suspected lesion areas. Moreover, a Multi-View Token Selection Module (MVTSM) is designed to prune redundant feature tokens and realize dynamic selection of key information. In the final decision stage, we dynamically fuse multi-view features through the novel Multi-View Mixture of Experts Module (MVMoEM), to enhance key views and reduce the impact of conflicting views. Extensive experiments on a large multi-view fundus image dataset with 34,452 images demonstrate that our method performs favorably against state-of-the-art models. Xiaoling Luo 0001, Qihao Xu, Huisi Wu, Chengliang Liu 0003, Zhihui Lai 0001, LinLin Shen |
AAAI | 1 |
| 2025 | UBDet: An Unsupervised Breast Tumor Detection Framework with Boundary-Aware Enhancement
Xingxin Guo, Zhihui Lai 0001, Heng Kong, Xiaoling Luo 0001 |
ICIC (22) | 4 |
| 2025 | InstCNet: A Dual-Branch Network for Enhanced Tumor Diagnosis via Joint Segmentation and Classification
Zhihui Lai 0001, Xingxin Guo, Heng Kong, Israr Hussain, Xiaoling Luo 0001 |
ICIC (14) | 5 |
| 2025 | Wavelet-based Global-Local Interaction Network with Cross-Attention for Multi-View Diabetic Retinopathy DetectionabstractMulti-view diabetic retinopathy (DR) detection has recently emerged as a promising method to address the issue of incomplete lesions faced by single-view DR. However, it is still challenging due to the variable sizes and scattered locations of lesions. Furthermore, existing multi-view DR methods typically merge multiple views without considering the correlations and redundancies of lesion information across them. Therefore, we propose a novel method to overcome the challenges of difficult lesion information learning and inadequate multi-view fusion. Specifically, we introduce a two-branch network to obtain both local lesion features and their global dependencies. The high-frequency component of the wavelet transform is used to exploit lesion edge information, which is then enhanced by global semantic to facilitate difficult lesion learning. Additionally, we present a cross-view fusion module to improve multi-view fusion and reduce redundancy. Experimental results on large public datasets demonstrate the effectiveness of our method. The code is open sourced on https://github.com/HuYongting/WGLIN. Yongting Hu, Chengliang Liu 0003, Xiaoling Luo 0001, Xiaoyan Dou, Qihao Xu, Yong Xu 0001 |
ICME | 4 |
| 2025 | Mutual Learning for SAM Adaptation: A Dual Collaborative Network Framework for Source-Free Domain TransferabstractSegment Anything Model (SAM) has demonstrated remarkable zero-shot segmentation capabilities across various visual tasks. However, its performance degrades significantly when deployed in new target domains with substantial distribution shifts. While existing self-training methods based on fixed teacher-student architectures have shown improvements, they struggle to ensure that the teacher network consistently outperforms the student under severe domain shifts. To address this limitation, we propose a novel Collaborative Mutual Learning Framework for source-free SAM adaptation, leveraging dual-networks in a dynamic and cooperative manner. Unlike fixed teacher-student paradigms, our method dynamically assigns the teacher and student roles by evaluating the reliability of each collaborative network in each training iteration. Our framework incorporates a dynamic mutual learning mechanism with three key components: a direct alignment loss for knowledge transfer, a reverse distillation loss to encourage diversity, and a triplet relationship loss to refine feature representations. These components enhance the adaptation capabilities of the collaborative networks, enabling them to generalize effectively to target domains while preserving their pre-trained knowledge. Extensive experiments on diverse target domains demonstrate that our proposed framework achieves state-of-the-art adaptation performance. Wai Keung Wong, Chengliang Liu 0003, Xiaoling Luo 0001, Yong Xu 0001 |
ICML | 4 |
| 2025 | Enhancing Multimodal Protein Function Prediction Through Dual-Branch Dynamic Selection with Reconstructive Pre-TrainingabstractMultimodal protein features play a crucial role in protein function prediction. However, these features encompass a wide range of information, ranging from structural data and sequence features to protein attributes and interaction networks, making it challenging to decipher their complex interconnections. In this work, we propose a multimodal protein function prediction method (DSRPGO) by utilizing dynamic selection and reconstructive pre-training mechanisms. To acquire complex protein information, we introduce reconstructive pre-training to mine more fine-grained information with low semantic levels. Moreover, we put forward the Bidirectional Interaction Module (BInM) to facilitate interactive learning among multimodal features. Additionally, to address the difficulty of hierarchical multi-label classification in this task, a Dynamic Selection Module (DSM) is designed to select the feature representation that is most conducive to current protein function prediction. Our proposed DSRPGO model improves significantly in BPO, MFO, and CCO on human datasets, thereby outperforming other benchmark models. Xiaoling Luo 0001, Chengliang Liu 0003, Xiaopeng Jin, Jie Wen 0001 |
IJCAI | 1 |
| 2025 | CauRDG: Enhancing Domain Generalization with Causal-Driven Semantic Consistency ReasoningabstractDomain generalization (DG) plays a pivotal role in enabling models to maintain robust performance across heterogeneous environments. However, existing DG methods are fundamentally constrained by two intertwined limitations: (1) causal misalignment, which stems from undifferentiated feature encoding that entangles causal mechanisms with environmental biases; (2)semantic conflict arises when conventional adaptation methods find it challenging to balance the preservation of class discriminability with the mitigation of domain-specific distribution discrepancies. To address these challenges of DG, we propose a novel Causal-Driven Semantic Consistency Reasoning (CauRDG) method, which synergistically integrates Prototype-Guided Causal Disentanglement (PGCD) and Dual-Space Semantic Disambiguation (DSSD). Specifically, PGCD constructs a causal framework that identifies stable relationships and decouples invariant mechanisms from domain-specific variations, preserving causal consistency while adapting to contextual differences. DSSD harnesses a dual-space paradigm, enhancing local categorical clarity and maintaining global conceptual unity, thus balancing domain-specific precision with cross-domain coherence. The robustness provided by CauRDG ensures robust extraction and interpretation of essential features by preserving invariant causal structures, thereby harmonizing discriminative semantics with domain-varying contexts. Extensive experiments on multiple benchmark datasets consistently demonstrate the effectiveness and superiority of our CauRDG over state-of-the-art baselines. Zongxin Liu 0003, Yishu Liu 0001, Guangming Lu 0002, Xiaoling Luo 0001, Bingzhi Chen |
ACM Multimedia | 4 |
| 2025 | PET-GPRA: Rethinking PET with Gradient-Aware Prompting and Router-Free Adapters for Few-shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) aims to continuously learn novel concepts from limited training samples without forgetting previously encountered classes. Recent advancements have leveraged Parameter-Efficient Tuning (PET) strategies on pre-trained models to enhance FSCIL performance. However, current PET-based FSCIL approaches still suffer from the challenges posed by catastrophic collapse of general prompt and limited adaptability of specific prompt . To this end, we redefine the function of the PET paradigm with both gradient-aware prompting (GAP) and router-free adapters (RFA) to boost the performance of FSCIL, termed as "PET-GPRA". To dynamically balance the retention of previously learned general knowledge and the acquisition of novel class information across sessions, the GAP paradigm adaptively adjusts the updated gradient of the general prompt by leveraging the angular relationship between the general knowledge gradient and the novel knowledge gradient. Meanwhile, the RFA mechanism utilizes the semantic similarity between class attributes to replace the routing network, guiding the integration of adapter information, in which adapters serve as specific prompts to enhance the adaptability. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority and effectiveness of our proposed PET-GPRA framework over state-of-the-art baselines. Yishu Liu 0001, Desen Wang, Xiaoling Luo 0001, Bingzhi Chen, Guangming Lu 0002 |
ACM Multimedia | 4 |
| 2025 | Hierarchical Information Aggregation for Incomplete Multimodal Alzheimer's Disease DiagnosisabstractAlzheimer's Disease (AD) poses a significant health threat to the aging population, underscoring the critical need for early diagnosis to delay disease progression and improve patient quality of life. Recent advances in heterogeneous multimodal artificial intelligence (AI) have facilitated comprehensive joint diagnosis, yet practical clinical scenarios frequently encounter incomplete modalities due to factors like high acquisition costs or radiation risks. Moreover, traditional convolution-based architecture face inherent limitations in capturing long-range dependencies and handling heterogeneous medical data efficiently. To address these challenges, in our proposed heterogeneous multimodal diagnostic framework (HAD), we develop a multi-view Hilbert curve-based Mamba block and a hierarchical spatial feature extraction module to simultaneously capture local spatial features and global dependencies, effectively alleviating spatial discontinuities introduced by voxel serialization. Furthermore, to balance semantic consistency and modal specificity, we build a unified mutual information learning objective in the heterogeneous multimodal embedding space, which maintains effective learning of modality-specific information to avoid modality collapse caused by model preference. Extensive experiments demonstrate that our HAD significantly outperforms state-of-the-art methods in various modality-missing scenarios, providing an efficient and reliable solution for early-stage AD diagnosis. Chengliang Liu 0003, Yuanxi Que, Qihao Xu, Jie Wen 0001, Xiaoling Luo 0001 |
NeurIPS | 7 |
| 2025 | OTMamba: Ophthalmology Image Translation Using Guidance-Controllable Mamba Diffusion Model
Huijie Deng, Yamei Lu, Zhuoru Wu, Chunhui Zou, Bowei Yuan, Xiaoling Luo 0001, LinLin Shen |
PRCV (13) | 7 |
| 2025 | Saliency-Guided Selection Driven Multi-scale Network for Breast Tumor Detection
Xinfei Gu, Chengliang Liu 0003, Xiaoling Luo 0001, Qihao Xu, Zhihui Lai 0001, Heng Kong |
PRCV (13) | 3 |
| 2025 | SEGT-GO: a graph transformer method based on PPI serialization and explanatory artificial intelligence for protein function predictionabstractBACKGROUND: A massive amount of protein sequences have been obtained, but their functions remain challenging to discern. In recent research on protein function prediction, Protein-Protein Interaction (PPI) Networks have played a crucial role. Uncovering potential function relationships between distant proteins within PPI networks is essential for improving the accuracy of protein function prediction. Most current studies attempt to capture these distant relationships by stacking graph network layers, but performance gains diminish as the number of layers increases. RESULTS: To further explore the potential functional relationships between multi-hop proteins in PPI networks, this paper proposes SEGT-GO, a Graph Transformer method based on PPI multi-hop neighborhood Serialization and Explainable artificial intelligence for large-scale multispecies protein function prediction. The multi-hop neighborhood serialization maps multi-hop information in the PPI Network into serialized feature embeddings, enabling the Graph Transformer to learn deeper functional features within the PPI Network. Based on game theory, the SHAP eXplainable Artificial Intelligence (XAI) framework optimizes model input and filters out feature noise, enhancing model performance. CONCLUSIONS: Compared to the advanced network method DeepGraphGO, SEGT-GO achieves more competitive results in standard large-scale datasets and superior results on small ones, validating its ability to extract functional information from deep proteins. Furthermore, SEGT-GO achieves superior results in cross-species learning and prediction of the functions of unseen proteins, further proving the method's strong generalization. Yundong Sun, Baohui Lin, Xiaoling Luo 0001, Xiaopeng Jin, Dongjie Zhu 0001 |
BMC Bioinform. | 5 |
| 2025 | Spike-VisNet: A novel framework for visual recognition with FocusLayer-STDP learning
Ying Liu 0070, Xiaoling Luo 0001, Wei Zhang 0308, Hong Qu 0002 |
Neural Networks | 2 |
| 2025 | Multi-view diabetic retinopathy grading via cross-view spatial alignment and adaptive vessel reinforcing
Xiaoyan Dou, Xiaoling Luo 0001, Zhihao Wu 0002, Chengliang Liu 0003, Tianyi Luo, Jie Wen 0001, Bingo Wing-Kuen Ling, Yong Xu 0001, Wei Wang 0169 |
Pattern Recognit. | 3 |
| 2025 | A Lesion-Fusion Neural Network for Multi-View Diabetic Retinopathy GradingabstractAs the most common complication of diabetes, diabetic retinopathy (DR) is one of the main causes of irreversible blindness. Automatic DR grading plays a crucial role in early diagnosis and intervention, reducing the risk of vision loss in people with diabetes. In these years, various deep-learning approaches for DR grading have been proposed. Most previous DR grading models are trained using the dataset of single-field fundus images, but the entire retina cannot be fully visualized in a single field of view. There are also problems of scattered location and great differences in the appearance of lesions in fundus images. To address the limitations caused by incomplete fundus features, and the difficulty in obtaining lesion information. This work introduces a novel multi-view DR grading framework, which solves the problem of incomplete fundus features by jointly learning fundus images from multiple fields of view. Furthermore, the proposed model combines multi-view inputs such as fundus images and lesion snapshots. It utilizes heterogeneous convolution blocks (HCB) and scalable self-attention classes (SSAC), which enhance the ability of the model to obtain lesion information. The experimental results show that our proposed method performs better than the benchmark methods on the large-scale dataset. Xiaoling Luo 0001, Qihao Xu, Zhihua Wang 0002, Chao Huang 0008, Chengliang Liu 0003, Xiaopeng Jin, Jianguo Zhang 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Retina-Inspired Lightweight Spiking Convolutional Neural Network for Single-Image DehazingabstractSuspended particles in hazy medium absorb and scatter light, severely degrading imaging quality. Numerous single-image dehazing methods have been proposed to reconstruct clear images from hazy ones. However, most of them focus on increasing depth and width to improve dehazing performance, which incurs high computation and energy costs. To address this issue, we propose a lightweight spiking convolutional neural network (CNN) referred to as retina-inspired spiking CNN (RI-SCNN) for the reconstruction of hazy images. Unlike conventional dehazing techniques, first, our proposed network simulates the hierarchical structure and cellular function of the retina and devises five network modules to efficiently encode and extract image features through ON and OFF roads. Furthermore, the linear reconstruction mechanism is introduced to integrate the outputs from different roads, adaptively preserving regions with optimal details and constructing a comprehensive visual representation. Finally, by the transformed atmospheric scattering formula, our network can generate the dehazy image. Incorporating the microscale spiking mechanism of the brain, the entire network leverages discrete binary spike trains for information encoding and transmission, directly trained by spiking surrogate gradient learning on integrate-and-fire (IF) neurons. Experimental results demonstrate the superiority of the proposed RI-SCNN in terms of quantitative dehazing performance, qualitative visual effect, energy efficiency, and run speed. Considering its lightweight architecture with ultralow computation and energy costs, the network is encouraged to be deployed in the visual sensor hardware to improve overall performance. Xiaoling Luo 0001, Qian Sun 0014, Hong Qu 0002, Zhang Yi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Toward Building Human-Like Sequential Memory Using Brain-Inspired Spiking Neural ModelsabstractThe brain is able to acquire and store memories of everyday experiences in real-time. It can also selectively forget information to facilitate memory updating. However, our understanding of the underlying mechanisms and coordination of these processes within the brain remains limited. However, no existing artificial intelligence models have yet matched human-level capabilities in terms of memory storage and retrieval. This study introduces a brain-inspired spiking neural model that integrates the learning and forgetting processes of sequential memory. The proposed model closely mimics the distributed and sparse temporal coding observed in the biological neural system. It employs one-shot online learning for memory formation and uses biologically plausible mechanisms of neural oscillation and phase precession to retrieve memorized sequences reliably. In addition, an active forgetting mechanism is integrated into the spiking neural model, enabling memory removal, flexibility, and updating. The proposed memory model not only enhances our understanding of human memory processes but also provides a robust framework for addressing temporal modeling tasks. Malu Zhang, Xiaoling Luo 0001, Jibin Wu, Ammar Belatreche, Siqi Cai 0002, Yang Yang 0002, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Attention-Induced Embedding Imputation for Incomplete Multi-View Partial Multi-Label ClassificationabstractAs a combination of emerging multi-view learning methods and traditional multi-label classification tasks, multi-view multi-label classification has shown broad application prospects. The diverse semantic information contained in heterogeneous data effectively enables the further development of multi-label classification. However, the widespread incompleteness problem on multi-view features and labels greatly hinders the practical application of multi-view multi-label classification. Therefore, in this paper, we propose an attention-induced missing instances imputation technique to enhance the generalization ability of the model. Different from existing incomplete multi-view completion methods, we attempt to approximate the latent features of missing instances in embedding space according to cross-view joint attention, instead of recovering missing views in kernel space or original feature space. Accordingly, multi-view completed features are dynamically weighted by the confidence derived from joint attention in the late fusion phase. In addition, we propose a multi-view multi-label classification framework based on label-semantic feature learning, utilizing the statistical weak label correlation matrix and graph attention network to guide the learning process of label-specific features. Finally, our model is compatible with missing multi-view and partial multi-label data simultaneously and extensive experiments on five datasets confirm the advancement and effectiveness of our embedding imputation method and multi-view multi-label classification model. Chengliang Liu 0003, Jinlong Jia, Jie Wen 0001, Xiaoling Luo 0001, Chao Huang 0008, Yong Xu 0001 |
AAAI | 5 |
| 2024 | HACDR-Net: Heterogeneous-Aware Convolutional Network for Diabetic Retinopathy Multi-Lesion SegmentationabstractDiabetic Retinopathy (DR), the leading cause of blindness in diabetic patients, is diagnosed by the condition of retinal multiple lesions. As a difficult task in medical image segmentation, DR multi-lesion segmentation faces the main concerns as follows. On the one hand, retinal lesions vary in location, shape, and size. On the other hand, because some lesions occupy only a very small part of the entire fundus image, the high proportion of background leads to difficulties in lesion segmentation. To solve the above problems, we propose a heterogeneous-aware convolutional network (HACDR-Net) that composes heterogeneous cross-convolution, heterogeneous modulated deformable convolution, and optional near-far-aware convolution. Our network introduces an adaptive aggregation module to summarize the heterogeneous feature maps and get diverse lesion areas in the heterogeneous receptive field along the channels and space. In addition, to solve the problem of the highly imbalanced proportion of focal areas, we design a new medical image segmentation loss function, Noise Adjusted Loss (NALoss). NALoss balances the predictive feature distribution of background and lesion by jointing Gaussian noise and hard example mining, thus enhancing awareness of lesions. We conduct the experiments on the public datasets IDRiD and DDR, and the experimental results show that the proposed method achieves better performance than other state-of-the-art methods. The code is open-sourced on github.com/xqh180110910537/HACDR-Net. Qihao Xu, Xiaoling Luo 0001, Chao Huang 0008, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001 |
AAAI | 2 |
| 2024 | 💎 GEM: Context-Aware Gaze EstiMation with Visual Search Behavior Matching for Chest Radiograph
Shaonan Liu, Wenting Chen, Jie Liu 0044, Xiaoling Luo 0001, LinLin Shen |
MICCAI (1) | 4 |
| 2024 | Simplify Implant Depth Prediction as Video Grounding: A Texture Perceive Implant Depth Prediction Network
Xinquan Yang, Xiaoling Luo 0001, Leilei Zeng, Yudi Zhang 0005, LinLin Shen, Yongqiang Deng |
MICCAI (5) | 3 |
| 2024 | A comprehensive review and comparison of existing computational methods for protein function predictionabstractProtein function prediction is critical for understanding the cellular physiological and biochemical processes, and it opens up new possibilities for advancements in fields such as disease research and drug discovery. During the past decades, with the exponential growth of protein sequence data, many computational methods for predicting protein function have been proposed. Therefore, a systematic review and comparison of these methods are necessary. In this study, we divide these methods into four different categories, including sequence-based methods, 3D structure-based methods, PPI network-based methods and hybrid information-based methods. Furthermore, their advantages and disadvantages are discussed, and then their performance is comprehensively evaluated and compared. Finally, we discuss the challenges and opportunities present in this field. Baohui Lin, Xiaoling Luo 0001, Xiaopeng Jin |
Briefings Bioinform. | 2 |
| 2024 | AMHGCN: Adaptive multi-level hypergraph convolution network for human motion prediction
Lian Wu, Xin Wang 0160, Xiaoling Luo 0001, Yong Xu 0001 |
Neural Networks | 5 |
| 2024 | A universal ANN-to-SNN framework for achieving high accuracy and low latency deep Spiking Neural Networks
Malu Zhang, Xiaoling Luo 0001, Hong Qu 0002 |
Neural Networks | 4 |
| 2024 | Information Recovery-Driven Deep Incomplete Multiview Clustering NetworkabstractIncomplete multiview clustering (IMC) is a hot and emerging topic. It is well known that unavoidable data incompleteness greatly weakens the effective information of multiview data. To date, existing IMC methods usually bypass unavailable views according to prior missing information, which is considered a second-best scheme based on evasion. Other methods that attempt to recover missing information are mostly applicable to specific two-view datasets. To handle these problems, in this article, we propose an information-recovery-driven-deep IMC network, termed as RecFormer. Concretely, a two-stage autoencoder network with self-attention structure is built to synchronously extract high-level semantic representations of multiple views and recover the missing data. Besides, we develop a recurrent graph reconstruction mechanism that cleverly leverages the restored views to promote representation learning and further data reconstruction. Visualization of recovery results are given and sufficient experimental results confirm that our RecFormer has obvious advantages over other top methods. Chengliang Liu 0003, Jie Wen 0001, Zhihao Wu 0002, Xiaoling Luo 0001, Chao Huang 0008, Yong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Minicolumn-Based Episodic Memory Model With Spiking Neurons, Dendrites and DelaysabstractEpisodic memory is fundamental to the brain's cognitive function, but how neuronal activity is temporally organized during its encoding and retrieval is still unknown. In this article, combining hippocampus structure with a spiking neural network (SNN), a new bionic spiking temporal memory (BSTM) model is proposed to explore the encoding, formation, and retrieval of episodic memory. For encoding episodic memory, the spike-timing-dependent-plasticity (STDP) learning algorithm and a proposed minicolumn selection algorithm are used to encode each input item into several active minicolumns. For the formation of episodic memory, a sequential memory algorithm is proposed to store the contexts between items. For retrieval of episodic memory, the local retrieval algorithm and the global retrieval algorithm are proposed to retrieve sequence information, achieving multisentence prediction and multitime step prediction. All functions of BSTM are based on bionic spiking neurons, which have biological characteristics including columnar and dendritic structures, firing and receiving spikes, and delaying transmission. To test the performance of the BSTM model, the Children's Book Test (CBT) data set was used to conduct a series of experiments under different settings, including changing the number of minicolumns, neurons and sequences, modifying sequence items, etc. Compared to other sequence memory algorithms, the experimental results show that the proposed BSTM achieves higher accuracy and better robustness. Yi Chen 0034, Jilun Zhang, Xiaoling Luo 0001, Malu Zhang, Hong Qu 0002, Zhang Yi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | DICNet: Deep Instance-Level Contrastive Network for Double Incomplete Multi-View Multi-Label ClassificationabstractIn recent years, multi-view multi-label learning has aroused extensive research enthusiasm. However, multi-view multi-label data in the real world is commonly incomplete due to the uncertain factors of data collection and manual annotation, which means that not only multi-view features are often missing, and label completeness is also difficult to be satisfied. To deal with the double incomplete multi-view multi-label classification problem, we propose a deep instance-level contrastive network, namely DICNet. Different from conventional methods, our DICNet focuses on leveraging deep neural network to exploit the high-level semantic representations of samples rather than shallow-level features. First, we utilize the stacked autoencoders to build an end-to-end multi-view feature extraction framework to learn the view-specific representations of samples. Furthermore, in order to improve the consensus representation ability, we introduce an incomplete instance-level contrastive learning scheme to guide the encoders to better extract the consensus information of multiple views and use a multi-view weighted fusion module to enhance the discrimination of semantic features. Overall, our DICNet is adept in capturing consistent discriminative representations of multi-view multi-label data and avoiding the negative effects of missing views and missing labels. Extensive experiments performed on five datasets validate that our method outperforms other state-of-the-art methods. Chengliang Liu 0003, Jie Wen 0001, Xiaoling Luo 0001, Chao Huang 0008, Zhihao Wu 0002, Yong Xu 0001 |
AAAI | 3 |
| 2023 | Incomplete Multi-View Multi-Label Learning via Label-Guided Masked View- and Category-Aware TransformersabstractAs we all know, multi-view data is more expressive than single-view data and multi-label annotation enjoys richer supervision information than single-label, which makes multi-view multi-label learning widely applicable for various pattern recognition tasks. In this complex representation learning problem, three main challenges can be characterized as follows: i) How to learn consistent representations of samples across all views? ii) How to exploit and utilize category correlations of multi-label to guide inference? iii) How to avoid the negative impact resulting from the incompleteness of views or labels? To cope with these problems, we propose a general multi-view multi-label learning framework named label-guided masked view- and category-aware transformers in this paper. First, we design two transformer-style based modules for cross-view features aggregation and multi-label classification, respectively. The former aggregates information from different views in the process of extracting view-specific features, and the latter learns subcategory embedding to improve classification performance. Second, considering the imbalance of expressive power among views, an adaptively weighted view fusion module is proposed to obtain view-consistent embedding features. Third, we impose a label manifold constraint in sample-level representation learning to maximize the utilization of supervised information. Last but not least, all the modules are designed under the premise of incomplete views and labels, which makes our method adaptable to arbitrary multi-view and multi-label data. Extensive experiments on five datasets confirm that our method has clear advantages over other state-of-the-art methods. Chengliang Liu 0003, Jie Wen 0001, Xiaoling Luo 0001, Yong Xu 0001 |
AAAI | 3 |
| 2023 | MVCINN: Multi-View Diabetic Retinopathy Detection Using a Deep Cross-Interaction Neural NetworkabstractDiabetic retinopathy (DR) is the main cause of irreversible blindness for working-age adults. The previous models for DR detection have difficulties in clinical application. The main reason is that most of the previous methods only use single-view data, and the single field of view (FOV) only accounts for about 13% of the FOV of the retina, resulting in the loss of most lesion features. To alleviate this problem, we propose a multi-view model for DR detection, which takes full advantage of multi-view images covering almost all of the retinal field. To be specific, we design a Cross-Interaction Self-Attention based Module (CISAM) that interfuses local features extracted from convolutional blocks with long-range global features learned from transformer blocks. Furthermore, considering the pathological association in different views, we use the feature jigsaw to assemble and learn the features of multiple views. Extensive experiments on the latest public multi-view MFIDDR dataset with 34,452 images demonstrate the superiority of our method, which performs favorably against state-of-the-art models. To the best of our knowledge, this work is the first study on the public large-scale multi-view fundus images dataset for DR detection. Xiaoling Luo 0001, Chengliang Liu 0003, Wai Keung Wong, Jie Wen 0001, Xiaopeng Jin, Yong Xu 0001 |
AAAI | 1 |
| 2023 | Masked Two-channel Decoupling Framework for Incomplete Multi-view Weak Multi-label LearningabstractMulti-view learning has become a popular research topic in recent years, but research on the cross-application of classic multi-label classification and multi-view learning is still in its early stages. In this paper, we focus on the complex yet highly realistic task of incomplete multi-view weak multi-label learning and propose a masked two-channel decoupling framework based on deep neural networks to solve this problem. The core innovation of our method lies in decoupling the single-channel view-level representation, which is common in deep multi-view learning methods, into a shared representation and a view-proprietary representation. We also design a cross-channel contrastive loss to enhance the semantic property of the two channels. Additionally, we exploit supervised information to design a label-guided graph regularization loss, helping the extracted embedding features preserve the geometric structure among samples. Inspired by the success of masking mechanisms in image and text analysis, we develop a random fragment masking strategy for vector features to improve the learning ability of encoders. Finally, it is important to emphasize that our model is fully adaptable to arbitrary view and label absences while also performing well on the ideal full data. We have conducted sufficient and convincing experiments to confirm the effectiveness and advancement of our model. Chengliang Liu 0003, Jie Wen 0001, Chao Huang 0008, Zhihao Wu 0002, Xiaoling Luo 0001, Yong Xu 0001 |
NeurIPS | 6 |
| 2023 | Class-guided human motion prediction via multi-spatial-temporal supervision
Honghu Pan, Lian Wu, Chao Huang 0008, Xiaoling Luo 0001, Yong Xu 0001 |
Neural Comput. Appl. | 5 |
| 2023 | A biologically inspired auto-associative network with sparse temporal population coding
Xiaoling Luo 0001, Yi Chen 0034, Hong Qu 0002 |
Neural Networks | 3 |
| 2023 | Supervised Learning in Multilayer Spiking Neural Networks With Spike Temporal Error BackpropagationabstractThe brain-inspired spiking neural networks (SNNs) hold the advantages of lower power consumption and powerful computing capability. However, the lack of effective learning algorithms has obstructed the theoretical advance and applications of SNNs. The majority of the existing learning algorithms for SNNs are based on the synaptic weight adjustment. However, neuroscience findings confirm that synaptic delays can also be modulated to play an important role in the learning process. Here, we propose a gradient descent-based learning algorithm for synaptic delays to enhance the sequential learning performance of single spiking neuron. Moreover, we extend the proposed method to multilayer SNNs with spike temporal-based error backpropagation. In the proposed multilayer learning algorithm, information is encoded in the relative timing of individual neuronal spikes, and learning is performed based on the exact derivatives of the postsynaptic spike times with respect to presynaptic spike times. Experimental results on both synthetic and realistic datasets show significant improvements in learning efficiency and accuracy over the existing spike temporal-based learning algorithms. We also evaluate the proposed learning method in an SNN-based multimodal computational model for audiovisual pattern recognition, and it achieves better performance compared with its counterparts. Xiaoling Luo 0001, Hong Qu 0002, Zhang Yi 0001, Jilun Zhang, Malu Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Temporal-Sequential Learning with Columnar-Structured Spiking Neural Networks
Xiaoling Luo 0001, Yi Chen 0034, Malu Zhang, Hong Qu 0002 |
ICONIP (4) | 1 |
| 2022 | Correction to: PHR-search: a search framework for protein remote homology detection based on the predicted protein hierarchical relationships
Xiaopeng Jin, Xiaoling Luo 0001, Bin Liu 0014 |
Briefings Bioinform. | 2 |
| 2022 | PHR-search: a search framework for protein remote homology detection based on the predicted protein hierarchical relationshipsabstractProtein remote homology detection is one of the most fundamental research tool for protein structure and function prediction. Most search methods for protein remote homology detection are evaluated based on the Structural Classification of Proteins-extended (SCOPe) benchmark, but the diverse hierarchical structure relationships between the query protein and candidate proteins are ignored by these methods. In order to further improve the predictive performance for protein remote homology detection, a search framework based on the predicted protein hierarchical relationships (PHR-search) is proposed. In the PHR-search framework, the superfamily level prediction information is obtained by extracting the local and global features of the Hidden Markov Model (HMM) profile through a convolution neural network and it is converted to the fold level and class level prediction information according to the hierarchical relationships of SCOPe. Based on these predicted protein hierarchical relationships, filtering strategy and re-ranking strategy are used to construct the two-level search of PHR-search. Experimental results show that the PHR-search framework achieves the state-of-the-art performance by employing five basic search methods, including HHblits, JackHMMER, PSI-BLAST, DELTA-BLAST and PSI-BLASTexB. Furthermore, the web server of PHR-search is established, which can be accessed at http://bliulab.net/PHR-search. Xiaopeng Jin, Xiaoling Luo 0001, Bin Liu 0014 |
Briefings Bioinform. | 2 |
| 2021 | Bio-inspired Model Based on Global-Local Hybrid Learning in Spiking Neural NetworkabstractBringing machines up to human-level visual processing capabilities is an attractive research topic for decades. Deep neural networks (DNNs), inspired by the hierarchical structure of the human primary visual cortex at a macroscopic level, have achieved state-of-the-art performance in many applications. However, their practical applications remain limited due to the requisition of massive computing resources. Spiking neural networks (SNNs) simulate the spike-based information process of the biological neural system from the microscopic view and hold greater potential to ultra-low-power computations. In this paper, we imitate the human visual system from both the micro and macro scales and make the following contributions: (1) Inspired by the lateral effect between real neurons, we propose a Global-Local Hybrid Spike-Timing-Dependent Plasticity (GLHSTDP) algorithm that combines STDP with lateral synaptic learning mechanism, to train the spiking neural network. (2) We construct a deep spiking neural network (DSNN) to mimic the visual information processing mechanism in the human brain. Experimental results demonstrate that the proposed DSNN model equipped with the proposed learning algorithm works in a totally spike-based manner and achieve competitive accuracies on both the Caltech 101 and the MNIST datasets. Xiaobin Wang, Hong Qu 0002, Yi Chen 0034, Xiaoling Luo 0001 |
IJCNN | 6 |
| 2021 | A new recursive least squares-based learning algorithm for spiking neurons
Hong Qu 0002, Xiaoling Luo 0001, Yi Chen 0034, Malu Zhang, Zefang Li |
Neural Networks | 3 |
| 2021 | MVDRNet: Multi-view diabetic retinopathy detection by combining DCNNs and attention mechanisms
Xiaoling Luo 0001, Zuhui Pu, Yong Xu 0001, Wai Keung Wong, Jingyong Su, Xiaoyan Dou, Baikang Ye, Jiying Hu, Lisha Mou |
Pattern Recognit. | 1 |
| 2019 | MPD-AL: An Efficient Membrane Potential Driven Aggregate-Label Learning Algorithm for Spiking NeuronsabstractOne of the long-standing questions in biology and machine learning is how neural networks may learn important features from the input activities with a delayed feedback, commonly known as the temporal credit-assignment problem. The aggregate-label learning is proposed to resolve this problem by matching the spike count of a neuron with the magnitude of a feedback signal. However, the existing threshold-driven aggregate-label learning algorithms are computationally intensive, resulting in relatively low learning efficiency hence limiting their usability in practical applications. In order to address these limitations, we propose a novel membrane-potential driven aggregate-label learning algorithm, namely MPD-AL. With this algorithm, the easiest modifiable time instant is identified from membrane potential traces of the neuron, and guild the synaptic adaptation based on the presynaptic neurons’ contribution at this time instant. The experimental results demonstrate that the proposed algorithm enables the neurons to generate the desired number of spikes, and to detect useful clues embedded within unrelated spiking activities and background noise with a better learning efficiency over the state-of-the-art TDP1 and Multi-Spike Tempotron algorithms. Furthermore, we propose a data-driven dynamic decoding scheme for practical classification tasks, of which the aggregate labels are hard to define. This scheme effectively improves the classification accuracy of the aggregate-label learning algorithms as demonstrated on a speech recognition task. Malu Zhang, Jibin Wu, Yansong Chua, Xiaoling Luo 0001, Zihan Pan, Haizhou Li 0001 |
AAAI | 4 |
| 2019 | Multi-resolution dictionary learning for face recognition
Xiaoling Luo 0001, Yong Xu 0001, Jian Yang 0003 |
Pattern Recognit. | 1 |