Ao Li 0002

dblp:54/2788-2 · DBLP profile ↗
← Back
25ranked-venue papers
11as first author
24since 2021 · last 2026
0000-0003-0735-2917ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 14 since 2021Computer networks · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Sample-specific Modality Diagnosis and Cross-modal Enhancement for Incomplete Multimodal Representations
abstract
In multimodal sentiment analysis, modality missingness and quality degradation are common. Existing methods often rely on batch-level modality generation, generation but neglect sample-level missingness, hence their flexibility is limited severely in real-world scenarios. To address this, Sample-specific Modality Diagnosis and Cross-modal Enhancement for Incomplete Multimodal Representations (SMCIR) is proposed. Specifically, The Dynamic Multi-feature Fusion Detector (DMFD) is presented, which detects missingness and severity at the sample-level using indicators such as information entropy, modality similarity, and mutual information. Unlike batch-based methods, the DMFD provides fine-grained detection and adaptive responses, improving sensitivity to modality disturbances. Meanwhile, the Context-aware Modality Completion Generator (CMCG) is developed to restore missing modalities through context-guided reconstruction using multiscale feature fusion and cross-modal attention. In this way, the proposed CMCG method can avoid redundancy and inconsistency, enhancing the consistency and discriminativity of the fused representation. In CMCG, the text modality serves as a stable guide to improve context consistency. Experiments on the CMU-MOSI and CMU-MOSEI datasets show that SMCIR outperforms existing full-modal and non-recovery-based methods, well validating its efficacy and superiority in multimodal learning.
Junsong Chen, Jiyuan Liu 0003, Suyuan Liu, Wei Zhang 0049, Ao Li 0002, En Zhu, Xinwang Liu 0002
AAAI5
2026 Simplicity meets power: robust traffic flow prediction with ST-ConvLSTMNet model
Yuan Cheng 0004, Ao Li 0002
Appl. Intell.3
2026 Diffusion-imputation one-step incomplete multi-view clustering with anchor graph regularization
Ao Li 0002, Sanlin Mei, Dehua Miao, Fengwei Gu
Neurocomputing1
2026 Fine-tuned Whisper-based semantic-temporal aggregation networks for sound event classification
Chen Chen 0086, Ao Li 0002, Fengwei Gu, Liang Xi
Pattern Recognit.4
2026 Multi-way dual-aligned based deep incomplete multiview learning with application to single-cell clustering
Sanlin Mei, Ao Li 0002, Fengwei Gu
Pattern Recognit.2
2026 Auto-weighted multi-dimensional feature fusion for incomplete multi-view clustering
Ao Li 0002, Xinya Xu
Signal Process.1
2026 HALT: Hierarchical Attention Learning for Visual Tracking
abstract
Siamese tracking algorithms have gained widespread recognition due to their exceptional efficiency and scalability. However, they exhibit suboptimal performance when localizing arbitrary targets under various disturbances, particularly in complex environments involving challenges such as illumination variation, deformation, and background clutter. Therefore, this paper proposes a Hierarchical Attention Learning network (HAL) to enhance tracking performance. Drawing inspiration from the hybrid attention mechanism, HAL designs a Feature Mining Attention module (FMA), a Global Feature Attention module (GFA), and a hierarchical network structure. Concretely, FMA employs parallel branches to fully extract channel and spatial features, enabling preliminary enhancement of target features and establishing global associations. Since lower-level and higher-level features emphasize positional and semantic information, respectively, GFA comprehensively integrates multi-level features to obtain more accurate tracking predictions. In particular, to improve the model's representation capacity, a hierarchical network structure is developed to deepen the HAL network and strengthen the feature dependencies between the template and the search region. Finally, based on HAL, we propose a Hierarchical Attention Learning Tracker (HALT) for visual tracking in complex environments, which is capable of learning rich hierarchical features. Extensive experiments demonstrate that, compared to state-of-the-art trackers, our HALT achieves outstanding tracking performance across multiple benchmarks while maintaining a real-time speed of 48.5 fps.
Fengwei Gu, Ao Li 0002, Chen Chen 0086, Chengtao Cai, Renjie Qiao, Guangyao Zhai, Zhaojie Ju
IEEE Trans Autom. Sci. Eng.2
2025 Anchor-based incomplete multi-view clustering with graph convolution network
Ao Li 0002
Appl. Intell.1
2025 Deep spectral clustering network for incomplete multi-view clustering
Ao Li 0002, Sanlin Mei
Eng. Appl. Artif. Intell.1
2025 KBGCN: Enhance K-th bias GCN via K-1 hop neighborhood distribution
Hailu Yang 0001, Zhixin Lin, Ao Li 0002, Chen Chen 0086, Lili Wang 0008
Expert Syst. Appl.3
2025 Dual-consistency graph-spectral embedding joint learning for multi-view single-cell clustering
Ao Li 0002, Tongtong Ji, Chunrui Wang, Fengwei Gu
Knowl. Based Syst.1
2025 Trajectory Optimization for Tooth Preparation Robot Based on P-MRSD Algorithm
abstract
Oral diseases, including dental caries and cracked teeth, are leading causes of tooth loss and can pose significant health risks if left untreated. Manual tooth preparation by dentists is often prone to visual bias and positioning errors. To address these issues, tooth preparation robots, driven by automation and intelligence, have been proposed as replacements for repetitive tasks. However, existing tooth preparation robots with serial systems suffer from low preparation accuracy and poor safety when preparing real teeth of hard and brittle materials. This includes the problems of low accuracy of theoretical preparation trajectory planning and deformation of the end of the serial system due to force. In this study, we propose a preparation trajectory optimization strategy that includes methods for surface morphology optimization and end deformation optimization to enhance preparation accuracy and safety. Our approach utilizes the proposed prediction of material residue and stiffness deformation (P-MRSD) method to optimize tooth morphology based on five key parameters, such as the bur pose and preparation trajectory. Additionally, the extrusion force during preparation is optimized by considering factors such as the material removal rate, the Euclidean distance field of tool contacts (TC), and the system stiffness of the tooth preparation robot. The accuracy and safety of the execution trajectory are ensured by minimizing stiffness deformation. Finally, a tooth preparation robot hardware system is developed to verify the correlation between predicted and observed tooth morphology, with end deformation optimization based on the optimized trajectory. The feasibility of the robot for preparing hard and brittle teeth is demonstrated, and the proposed trajectory optimization method improves both accuracy and safety in preparation process. This provides a theoretical foundation and technical support for advancing automated robotic technology, particularly in the development of more accurate and safer serial robotic arm systems. Note to Practitioners—The motivation of this paper is to address the challenge of trajectory optimization for grinding hard and brittle materials using a serial robotic system, with potential applications in machining. The interaction between tool pose, trajectory, and material (specifically a cracked hard and brittle tooth) is analyzed to investigate its impact on preparation. Two optimization parameters are proposed to refine the theoretical preparation trajectory based on the surface morphology indexes. Real preparation experiments are then conducted using the proposed extrusion force optimization model, which enhances the motion accuracy of the theoretical trajectory. This improvement boosts the performance of the serial robotic arm system when preparing hard and brittle materials. The findings also suggest that computer-aided design and manufacturing systems could autonomously generate preparation trajectory optimization plans that consider both surface morphology and safety, offering insights and support for machining in other fields. Preliminary physical experiments demonstrate the feasibility of the proposed trajectory planning and kinematic parameter optimization methods, improving both preparation accuracy and safety. However, the research is still in the laboratory phase and has not yet been applied clinically. Future studies will focus on balancing safety and accuracy when the serial system operates intraorally, advancing the potential clinical application of tooth preparation robots.
Jianpeng Sun, Jingang Jiang 0001, Chunrui Wang, Zhonghao Xue, Ao Li 0002
IEEE Trans Autom. Sci. Eng.5
2025 Learning to Discriminate While Contrasting: Combating False Negative Pairs With Coupled Contrastive Learning for Incomplete Multi-View Clustering
abstract
The task of incomplete multi-view clustering (IMvC) aims to partition multi-view data with a lack of completeness into different clusters. The incompleteness can be typically categorized into the case of instance-missing and view-unaligned MvC. However, prior methods either consider each of them or struggle to pursue consistent latent representations among views. In this paper, we propose two forms of contrastive learning paradigms to jointly handle both cases for IMvC. Specifically, we design an instance-oriented contrastive (IOC) learning strategy to achieve intra-class consistency. As negative samples within different datasets can exhibit diverse distributions, we formulate a parameterized boundary for IOC learning to flexibly deal with such differing data modes. To preserve inter-view consistency, we further devise category-oriented contrastive (COC) learning such that data from different views can be seamlessly integrated into a combined semantic space. We also recover the missing instances with the learned latent representations in a reconstructing manner for realigning the incomplete multi-view data to facilitate clustering. Our approach unifies the solution to both incomplete cases into one formulation. To demonstrate the effectiveness of our model, we conduct four types of MvC tasks on six benchmark multi-view datasets and compare our method against state-of the-art IMvC methods. Extensive experiments show that our method achieves state-of-the-art performance, quantitatively and qualitatively.
Katsuya Hotta, Chunzhi Gu, Ao Li 0002, Jun Yu 0012, Chao Zhang 0030
IEEE Trans. Knowl. Data Eng.4
2025 Deep Incomplete Multiview Clustering via Local and Global Pseudo-Label Propagation
abstract
Since the rapid progress in multimedia and sensor technologies, multiview clustering (MVC) has become a prominent research area within machine learning and data mining, experiencing significant advancements over recent decades. MVC is distinguished from single-view clustering by its ability to integrate complementary information from multiple distinct data perspectives and enhance clustering performance. However, the efficacy of MVC methods is predicated on the availability of complete views for all samples-an assumption that frequently fails in practical scenarios where data views are often incomplete. To surmount this challenge, various approaches to incomplete MVC (IMVC) have been proposed, with deep neural networks emerging as a favored technique for their representation learning ability. Despite their promise, previous methods commonly adopt sample-level (e.g., features) or affinity-level (e.g., graphs) guidance, neglecting the discriminative label-level guidance (i.e., pseudo-labels). In this work, we propose a novel deep IMVC method termed pseudo-label propagation for deep IMVC (PLP-IMVC), which integrates high-quality pseudo-labels from the complete subset of incomplete data with deep label propagation networks to obtain improved clustering results. In particular, we first design a local model (PLP-L) that leverages pseudo-labels to their fullest extent. Then, we propose a global model (PLP-G) that exploits manifold regularization to mitigate the label noises, promote view-level information fusion, and learn discriminative unified representations. Experimental results across eight public benchmark datasets and three evaluation metrics prove our method's efficacy, demonstrating superior performance compared to 18 advanced baseline methods.
Ao Li 0002, Haoyue Xu, Hailu Yang 0001, Xinwang Liu 0002
IEEE Trans. Neural Networks Learn. Syst.2
2025 Contrastive Continual Multiview Clustering With Filtered Structural Fusion
abstract
Multiview clustering thrives in applications where views are collected in advance by extracting consistent and complementary information among views. However, it overlooks scenarios where data views are collected sequentially, i.e., real-time data. Due to privacy issues or memory burden, previous views are not available with time in these situations. Some methods are proposed to handle it but are trapped in a stability-plasticity dilemma. In specific, these methods undergo a catastrophic forgetting of prior knowledge when a new view is attained. Such a catastrophic forgetting problem (CFP) would cause the consistent and complementary information hard to get and affect the clustering performance. To tackle this, we propose a novel method termed contrastive continual multiview clustering with filtered structural fusion (CCMVC-FSF). Precisely, considering that data correlations play a vital role in clustering and prior knowledge ought to guide the clustering process of a new view, we develop a data buffer to store filtered structural information and utilize it to guide the generation of a robust partition matrix via contrastive learning. Additionally, to address the high complexity involved in acquiring and storing structural information, we propose a sampling strategy called clustering then sample. Furthermore, we theoretically connect CCMVC-FSF with semisupervised learning and knowledge distillation. Extensive experiments exhibit the excellence of the proposed method. Our code is publicly available at https://github.com/wanxinhang/CCMVC-FSF/.
Xinhang Wan, Jiyuan Liu 0003, Hao Yu 0017, Qian Qu, Ao Li 0002, Xinwang Liu 0002, Ke Liang 0006, Zhibin Dong, En Zhu
IEEE Trans. Neural Networks Learn. Syst.5
2024 HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection
abstract
Infrared small object detection is an important computer vision task involving the recognition and localization of tiny objects in infrared images, which usually contain only a few pixels. However, it encounters difficulties due to the diminutive size of the objects and the generally complex backgrounds in infrared images. In this paper, we propose a deep learning method, HCF-Net, that significantly improves infrared small object detection performance through multiple practical modules. Specifically, it includes the parallelized patch-aware attention (PPA) module, dimension-aware selective integration (DASI) module, and multi-dilated channel refiner (MDCR) module. The PPA module uses a multi-branch feature extraction strategy to capture feature information at different scales and levels. The DASI module enables adaptive channel selection and fusion. The MDCR module captures spatial features of different receptive field ranges through multiple depth-separable convolutional layers. Extensive experimental results on the SIRST infrared single-frame image dataset show that the proposed HCF-Net performs well, surpassing other traditional and deep learning models. Code is available at https://github.com/zhengshuchen/HCFNet.
Shibiao Xu, ShuChen Zheng, Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Xiaoqiang Teng, Ao Li 0002, Li Guo 0004
ICME8
2024 Incomplete multi-view clustering via local and global bagging of anchor graphs
Ao Li 0002, Haoyue Xu, Hailu Yang 0001, Shibiao Xu
Expert Syst. Appl.1
2024 Graph t-SNE multi-view autoencoder for joint clustering and completion of incomplete multi-view data
Ao Li 0002, Shibiao Xu, Yuan Cheng 0004
Knowl. Based Syst.1
2024 Anchor-based sparse subspace incomplete multi-view clustering
Ao Li 0002, Yuegong Sun
Wirel. Networks1
2024 Inception-Det: large aspect ratio rotating object detector for remote sensing images
Ao Li 0002, Yutong Niu, Zening Wang, Hailu Yang 0001
Wirel. Networks1
2023 Unsupervised multimodal domain adversarial network for time series classification
Liang Xi, Yujia Liang, Xunhua Huang, Ao Li 0002
Inf. Sci.5
2023 Unsupervised dimension-contribution-aware embeddings transformation for anomaly detection
Liang Xi, Chenchen Liang, Ao Li 0002
Knowl. Based Syst.4
2021 ST-CFSFDP algorithm based on Euclidean distance constraint
abstract
The spatial-temporal clustering by fast search and find of density peaks (ST-CFSFDP) has a better clustering effect on the spatiotemporal data set in a small space. However, there are some deficiencies in the spatiotemporal dataset with large data volume and far interval between sample points, the clustering results showed great differences, too many interference points during visualization. Given the above deficiencies, this paper proposes a spatial-temporal clustering by fast search and find of density peak algorithm based on Euclidean distance constraint, by increasing the partition constraint of some sample points, the problems existing in the spatiotemporal clustering algorithm of ST-CFSFDP are improved. Experimental results show that the improved algorithm has a better clustering effect than the original algorithm.
JunQiao Jiang, Yuan Cheng 0004, Ao Li 0002
MASS3
2021 Tensor-Based Reliable Multiview Similarity Learning for Robust Spectral Clustering on Uncertain Data
abstract
Similarity graph learning is the most key technique for multiview spectral clustering. However, existing methods fail when applied to uncertain data contaminated with various types of noise in an open environment. Due to the damaged structure by noise, unreliable similar relationships are learned, which extends similarity inconsistency among views. Moreover, the high-order correlation hidden in graphs are ignored generally. To address these problems, we propose a reliable similarity learning scheme for multiview clustering on uncertain data. This method can significantly improve spectral clustering performance in a noisy environment, and the contributions of our scheme include the following three aspects: 1) Uncertain data subspace reconstruction and adaptive graph learning are combined to construct a view-specific graph from high-quality recovered data, thus improving robustness. 2) A low-rank tensor constraint is utilized to facilitate multiview fusion, where the latent high-order correlation among view graphs will be fully explored when learning the consensus graph structure. 3) Data recovery, view-specific graphs, and latent consensus tensor structure are assembled into a unified framework, to be optimized jointly for mutual benefit. Our study also develops an efficient algorithm for obtaining overall solutions. The experimental results on several datasets demonstrate that our proposed approach shows significant improvements in robustness and evaluation metrics over the comparison methods.
Ao Li 0002, Jiajia Chen 0004, Mengke Yuan, Shibiao Xu, Guanglu Sun
IEEE Trans. Reliab.1
2020 Semi-Supervised Subspace Learning for Pattern Classification via Robust Low Rank Constraint
Ao Li 0002, Ruoqi An, Guanglu Sun, Xin Liu 0085, Qidi Wu, Hailong Jiang
Mob. Networks Appl.1