Xiaocui Li 0001

dblp:185/9845-1 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0002-5971-2331ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 2 since 2021
YearPublicationVenuePosition
2026 End-to-End Knowledge Distillation for Unsupervised Domain Adaptation with Large Vision-language Models
abstract
Knowledge distillation based on large vision-language models (VLMs) has recently emerged as a significant solution to transfer knowledge from the source domain to the target domain in unsupervised domain adaptation (UDA) tasks. However, existing methods employ a two-stage training pipeline, which not only complicates the training procedure but also lacks interactions between the source and target domains, severely hindering real-time cross-domain knowledge transfer. To address these challenges, we propose End-to-End Knowledge Distillation for UDA with large VLMs (termed as EKDA). (1) EKDA employs a lightweight prompt learning mechanism to first embed the knowledge from the source domain into VLMs, and then simultaneously utilize the image encoder and text encoder of VLMs to perform knowledge distillation on the target domain, significantly reducing the domain gap. (2) EKDA designs a teacher-student alternating training strategy to implement real-time collaborative interactions across domains, enabling an end-to-end paradigm to provide accurate source domain-aware supervision for the target domain. We conduct extensive experiments on 4 widely recognized benchmark datasets including Office-31, Office-Home, VisDA-2017, and Mini-DomainNet. Experimental results demonstrate that EKDA achieves significant performance improvement over the state-of-the-art UDA approaches, while maintaining a much lower model complexity. Take Office-Home for example, EKDA has gained at least 2.7% performance improvement while reducing the learnable parameters by over 80% compared with the state-of-the-art UDA baselines.
Yangtao Wang, Xingwei Deng, Yanzhao Xie, Weilong Peng, Siyuan Chen 0005, Xiaocui Li 0001, Maobin Tang, Meie Fang
AAAI6
2026 F3-SD: Focal feature fusion with self-distillation on large vision-language models for cross-modal retrieval
Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Xiaocui Li 0001, Maobin Tang, Meie Fang, Wensheng Zhang 0002
Pattern Recognit.5
2026 Prompt-affinity multi-modal class centroids for unsupervised domain adaption
abstract
In recent years, the advancements in large vision-language models (VLMs) like CLIP have sparked a renewed interest in leveraging the prompt learning mechanism to preserve semantic consistency between source and target domains in unsupervised domain adaption (UDA). While these approaches show promising results, they encounter fundamental limitations when quantifying the similarity between source and target domain data , primarily stemming from the redundant and modality-missing class centroids . To address these limitations, we propose P rompt-affinity M ulti-modal C lass C entroids for UDA (termed as PMCC). Firstly, we fuse the text class centroids (directly generated from the text encoder of CLIP with manual prompts for each class) and image class centroids (generated from the image encoder of CLIP for each class based on source domain images) to yield the multi-modal class centroids. Secondly, we conduct the cross-attention operation between each source or target domain image and these multi-modal class centroids. In this way, these class centroids that contain rich semantic information of each class will serve as a bridge to effectively measure the semantic similarity between different domains. Finally, we design a logit bias head and employ a multi-modal prompt learning mechanism to accurately predict the true class of each image for both source and target domains. We conduct extensive experiments on 4 popular UDA datasets including Office-31, Office-Home, VisDA-2017, and DomainNet. The experimental results validate our PMCC achieves higher performance with lower model complexity than the state-of-the-art (SOTA) UDA methods. The code of this project is available at GitHub: https://github.com/246dxw/PMCC .
Xingwei Deng, Yangtao Wang, Yanzhao Xie, Xiaocui Li 0001, Maobin Tang, Meie Fang, Wensheng Zhang 0002
Pattern Recognit.4
2026 RDS-Net: Recursive structure refinement network with Dual-awareness Shape transfer for point cloud completion
Xiaocui Li 0001, Weili Chen, Xinyu Zhang 0012, Yangtao Wang, Wei Liang 0005, Keqin Li 0001
Pattern Recognit.1
2026 Adaptive message passing mechanism for graph neural networks
Yangtao Wang, Linruo Liu, Yanzhao Xie, Maobin Tang, Xiaocui Li 0001
Pattern Recognit.7
2026 FedCD: Contrastive-distillation regularization for heterogeneous data in federated learning
Guodong Yi, Jianxu Zhang, Xinyu Zhang 0012, Wei Liang 0006, Xiaocui Li 0001
Pattern Recognit.6
2026 In-Network Load Balancing With Fast Congestion Flow Detection for Lossless Data Center Networks
abstract
To meet the high performance requirement of real time and critical applications in industrial Internet of Things, modern lossless ethernet data center networks (DCNs) deployed with remote direct memory access and priority-based flow control (PFC) are dedicated to delivering low latency and high bandwidth. However, existing load balancing schemes either lack sub-round-trip-time congestion sensing or fail to accurately detect and reroute flows that cause congestion in PFC-enabled lossless DCNs. Therefore, we propose LBoDSN, an in-network load balancing for lossless DCNs using direct switch notification (DSN) for fast congestion flow detection, to address above challenges. LBoDSN tracks the evolution of ingress queue lengths at destination switches to anticipate the initiation of PFC pause, precisely identifies congested flows before PFC pause, and then sends DSNs to source switches to perform rerouting. After rerouting, the congestion notification packet associated with the previous path is selectively discarded to improve transmission performance. Experiments under realistic workloads reveal that, LBoDSN outperforms CONGA by 13%–65%, and 25%–80% in average and tail Flow Completion Times (FCTs), respectively. Compared to ConWeave, LBoDSN achieves approximately 9% improvement in both average and tail FCTs, while reducing switch queue consumption for reordering.
Qingyu Shi 0001, Fangxue Jiang, Chuang Li 0004, Xiaocui Li 0001, Wenzhi Cao, Limei Liu
IEEE Trans. Ind. Informatics5
2026 Progressive Hybrid Pseudo-Labeling for Unsupervised Domain Adaptation With Ascending Low-Rank Adaptation
abstract
Unsupervised domain adaptation (UDA) based on large vision-language models (VLMs) has recently demonstrated strong generalization ability, yet it remains fundamentally challenged by noisy pseudo-labels and inefficient adaptation under large domain shifts. In this paper, we propose Progressive Hybrid Pseudo-Labeling for UDA with Ascending Low-Rank Adaptation (termed as PHPL), a parameter-efficient paradigm that addresses these challenges from two complementary perspectives. 1) We introduce a progressive hybrid pseudo-labeling strategy that constructs target-domain supervision by fusing predictions from a frozen teacher model and an adaptive student model with a progressive weighting scheme. By gradually transferring predictive responsibility from the teacher to the student during training, PHPL effectively mitigates early-stage pseudo-label noise and stabilizes self-training under large domain shifts. 2) To enable efficient and stable adaptation of large VLMs, we propose an ascending low-rank adaptation strategy that allocates LoRA capacity in a depth-aware manner. Specifically, larger low-rank updates are assigned to deeper, semantically richer layers, while shallow layers remain lightly parameterized, striking a favorable balance between parameter efficiency and representational expressiveness. We conduct extensive experiments on five widely-used UDA benchmarks, including Office-Home, Office-31, VisDA-2017, Mini-DomainNet, and DomainNet. Experimental results verify that PHPL consistently achieves higher performance across various cross-domain scenarios compared with existing CNN, Transformer, and VLMs-based solutions. Notably, PHPL demonstrates strong robustness on highly challenging large-scale conditions while requiring significantly less computational overhead, validating the effectiveness and scalability of the proposed lightweight adaptation paradigm. The code is available at https://github.com/el2k/PHPL.
Yangtao Wang, Mingxin Huang, Xingwei Deng, Yanzhao Xie, Xiaocui Li 0001
IEEE Trans. Image Process.5
2025 Enhancing Cross-modal Semantic Consistency via Key Token Alignment for Image-text Retrieval
abstract
Image-text retrieval (ITR) plays a pivotal role in advancing intelligent transportation systems, facilitating efficient retrieval and utilization of multimedia data to enhance traffic management and safety significantly. However, existing ITR solutions have not effectively addressed the issues of image patch redundancy and text word redundancy, leading to erroneous image-text matching. In this paper, we propose SCTA that enhances cross-modal semantic consistency via key token alignment for ITR. Firstly, SCTA evaluates the importance of each image patch by calculating the self-attention scores within patches and cross-attention scores between patches and words. Secondly, SCTA implements aggregation operations on image and text separately, aiming to generate information-rich key image patch embeddings and text word token embeddings. Finally, SCTA completes fine-grained alignment by maximizing the similarity between patch-to-word and word-to-patch. Therefore, SCTA simultaneously addresses image patch redundancy and text word redundancy issues, enhancing semantic consistency by aligning the core semantic information between image-text pairs. Extensive experiments on multiple datasets including Flickr30K and MS-COCO verify the superior performance of SCTA compared with the SOTA fine-grained ITR methods. The code of this paper is released at GitHub: https://github.com/ICME2025ITR/SCTA.
Huilong Lin, Yangtao Wang, Meie Fang, Yanzhao Xie, Xiaocui Li 0001, Weilong Peng, Siyuan Chen 0005, Maobin Tang, Ping Li 0016
ICME6
2025 Incomplete Multi-view Clustering via Local Reasoning and Correlation Analysis
abstract
In recent years, incomplete multi-view clustering (IMVC) has attracted considerable attention for its ability to acheieve effective clustering results through the integration of key information amidst missing view. However, the existing IMVC methods are still faced with 3 limitations: (1) They exhibit deficiencies in considering the weight distribution within views, (2) they ignore the varying contributions of different views to the common consistent representation, and (3) they struggle to sufficiently extract and recover the vital information within incomplete views. To address these limitations, we incorporates local reasoning and correlation analysis to design an incomplete multi-view clustering method(IMVCLRCA), which introduces a new strategy of feature learning and missing view recovery, fully exploiting local similarity and structural continuity within views and performing precise local reasoning recovery on missing data. By maximizing mutual information between views through contrastive learning, we achieve the consistent representation learning of multiple views. Furthermore, based on semantic consistency, we comprehensively consider the correlation between views, utilized a weight matrix to fuse cross-view data, and constructed a view with a correlation structure, ultimately obtaining a common consistent representation. We conduct extensive experiments on 4 public datasets including Caltech101-20, BBCSport, Scene-15, and LandUse-21. Experimental results demonstrate that IMVCLRCA has higher accuracy and robustness compared to the state-of-the-art IMVC methods. The anonymous code of this project is available on GitHub at https://github.com/ggg2111/2025WSDM-IMVCLRCA.
Xiaocui Li 0001, Xinyu Zhang 0012, Yangtao Wang, Qingyu Shi 0001, Wei Liang 0006
WSDM1
2025 CrossNet-VGA: Variational Collaboration and Graph Attention Fusion for Incomplete Multi-View Clustering
abstract
In recent years, multi-view data often suffer from incompleteness owing to environmental factors, equipment failures. Thus Incomplete Multi-View Clustering (IMVC) has become an important research focus, which aims to alleviate the adverse impacts of missing views and leverage inter-view complementary information to enhance clustering performance. However, existing IMVC methodologies suffer from three critical limitations: 1) Inadequate integration of cross-view learning and cross-instance learning; 2) Lack of explicit modeling for dynamic interactions between view-specific information and cross-view shared semantics; 3) Inability to dynamically capture high-order topological correlations under view-missing conditions, leading to semantic misalignment among samples. To address these challenges, we propose an IMVC framework CrossNet-VGA based on variational collaboration and graph attention fusion. Specifically, We formulate a novel multi-view evidence lower bound to explicitly separate view-specific latent variables and cross-view shared latent variables, and achieve inter-view semantic fusion by integrating variational distributions shared across views. Contrastive learning is employed to maximize mutual information and promote feature distribution uniformity, thereby achieving consistent representation learning. We employ dynamic $k$ -nearest neighbor graph construction and multi-head graph attention mechanisms to capture the inter-sample deep topological correlations, achieving robust structural alignment. Comprehensive experiments conducted on 6 public datasets demonstrate that CrossNet-VGA significantly outperforms the competing methods both on accuracy and robustness. The anonymous code of this work is available on GitHub at https://github.com/ggg2111/2025-TIP-CrossNet-VGA.
Xiaocui Li 0001, Xinyu Zhang 0012, Jie Wen 0001, Lian Wu
IEEE Trans. Image Process.1
2025 Efficient Algorithms for Approximate k-Radius Coverage Query on Large-Scale Road Networks
abstract
The challenge of optimally placing facilities to maximize coverage within road networks is a critical problem with significant implications for urban planning, emergency response, and the development of sustainable infrastructure. For instance, strategically locating fire stations or electric vehicle (EV) charging stations along a road network can greatly enhance public safety and support the adoption of clean transportation technologies. However, determining these optimal placements is computationally challenging, particularly when accounting for factors like road network distances and coverage radius. Traditional methods, such as greedy algorithms, offer a reasonable approximation but are limited by high computational complexity, making them less suitable for large-scale transportation networks. In response, our research introduces two novel algorithms designed to improve both the efficiency and scalability of the k-radius coverage problem. The first algorithm achieves a strong approximation with significantly reduced time complexity, while the second employs a sketch-based approach, offering a nearly linear time complexity relative to the number of edges. Although the second algorithm sacrifices some approximation accuracy, it offers substantial gains in computational speed, making it particularly valuable for large-scale transportation networks. Extensive experiments on large-scale real-world road networks demonstrate the superior performance of our proposed methods compared to existing solutions.
Xiaocui Li 0001, Dan He 0009, Xinyu Zhang 0012
IEEE Trans. Intell. Transp. Syst.1
2024 Domain Alignment with Large Vision-language Models for Cross-domain Remote Sensing Image Retrieval
abstract
Cross-domain remote sensing image retrieval has been a hotspot in the past few years. Most of the existing methods focus on combining semantic learning with domain adaptation on well-labeled source domain and unlabeled target domain. However, they face two serious challenges. (1) They cannot deal with practical scenarios where the source domain lacks sufficient label supervision. (2) They suffer from severe performance degradation when the data distribution between the source domain and target domain becomes highly inconsistent. To address these challenges, we propose D omain A lignment with L arge V ision-language models for cross-domain remote sensing image retrieval (termed as DALV). First, we design a dual-modality prototype guided pseudo-labeling mechanism, which leverages the pre-trained large vision-language model (i.e., CLIP) to assign pseudo-labels for all unlabeled source domain images and target domain images. Second, we compute the confidence scores for these pseudo-labels to distinguish their reliability. Next, we devise a loss reweighting strategy, which incorporates the confidence scores as weight values into the contrastive loss to mitigate the impact of noisy pseudo-labels. Finally, the low-rank adaptation fine-tuning means is adapted to update our model and achieve domain alignment to obtain class discriminative features. Extensive experiments on 12 cross-domain remote sensing image retrieval tasks show that our proposed DALV outperforms the state-of-the-art approaches. The source code is available at https://github.com/ptyy01/DALV.
Guocan Cai, Fufang Li, Yangtao Wang, Xin Tan 0002, Xiaocui Li 0001
CIKM6
2024 Image-text Retrieval with Main Semantics Consistency
abstract
Image-text retrieval (ITR) has been one of the primary tasks in cross-modal retrieval, serving as a crucial bridge between computer vision and natural language processing. Significant progress has been made to achieve global alignment and local alignment between images and texts by mapping images and texts into a common space to establish correspondences between these two modalities. However, the rich semantic content contained in each image may bring false matches, resulting in the matched text ignoring the main semantics but focusing on the secondary or other semantics of this image. To address this issue, this paper proposes a semantically optimized approach with a novel Main Semantics Consistency (MSC) loss function, which aims to rank the semantically most similar images (or texts) corresponding to the given query at the top position during the retrieval process. First, in each batch of image-text pairs, we separately compute (i) the image-image similarity, i.e., the similarity between every two images, (ii) the text-text similarity, i.e., the similarity between a group of texts (that belong to a certain image) and another group of texts (that belong to another image), and (iii) the image-text similarity, i.e., the similarity between each image and each text. Afterward, our proposed MSC effectively aligns the above image-image, image-text, and text-text similarity, since the main semantics of every two images will be highly close if their text descriptions remain highly semantically consistent. By this means, we can capture the main semantics of each image to be matched with its corresponding texts, prioritizing the semantically most related retrieval results. Extensive experiments on MSCOCO and FLICKR30K verify the superior performance of MSC compared with the SOTA image-text retrieval methods. The source code of this project is released at GitHub: https://github.com/xyi007/MSC.
Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Jingjing Li 0001, Xiaocui Li 0001, Weilong Peng, Maobin Tang, Meie Fang
CIKM6
2024 Multi-class Imbalanced Data Classification by Deep Multi-set Discriminant Metric Learning with Optimal Balance Sampling
Xinyu Zhang 0012, Xiaoyuan Jing, Xiaocui Li 0001, Jiagang Liu
DASFAA (2)3
2024 Adaptive Network Load Balancing at the End Host for Traffic Bursts in Data Centers
abstract
The network load balancing mechanism plays a pivotal role in enhancing transmission performance in modern cloud data centers. Conventional flowlet-based approaches at host side offer a balance between performance and deployment simplicity. However, their passive load balancing strategy restricts rerouting opportunities, and lacks precision in congestion detection as it necessitates at least one round-trip time (RTT) to acquire end-to-end congestion feedback. To overcome the performance loss caused by the above limitations, we propose BurstLoader, an enhanced flowlet-based mechanism that adapts to varying traffic burst intensities and improves congestion detection accuracy. BurstLoader proactively reroutes congested flows when no new flowlets are detected, while simultaneously avoiding the rerouting of flowlets that are in good transmission states. Furthermore, BurstLoader incorporates delay and its gradient for a more nuanced and precise congestion detection. The extensive experiments demonstrate that BurstLoader achieves a significant reduction in flow completion time (FCT) by up to 48% compared to other flowlet-based solutions deployed at the end host, while maintaining competitive performance even against schemes that require custom switches under realistic workloads.
Qingyu Shi 0001, Xiaocui Li 0001, Chuang Li 0004, Wenzhi Cao, Limei Liu
HPCC3
2024 LBoDSN: An In-Network Load Balancing Mechanism for Lossless Data Center Networks Based on Direct Switch Notification
Qingyu Shi 0001, Fangxue Jiang, Xiaocui Li 0001, Chuang Li 0004, Wenzhi Cao, Limei Liu
NPC (1)4
2024 ImMC-CSFL: Imbalanced Multi-view Clustering Algorithm Based on Common-Specific Feature Learning
Xiaocui Li 0001, Xinyu Zhang 0012, Qingyu Shi 0001, Xiance Tang
PAKDD (1)1
2024 Task graph offloading via deep reinforcement learning in mobile edge computing
Jiagang Liu, Yun Mi, Xinyu Zhang 0012, Xiaocui Li 0001
Future Gener. Comput. Syst.4
2021 Influence maximization in social graphs based on community structure and node coverage gain
Jingke Xi, Xiaocui Li 0001
Future Gener. Comput. Syst.4
2021 Multi-view clustering via neighbor domain correlation learning
Xiaocui Li 0001, Ke Zhou 0001, Chunhua Li 0002, Xinyu Zhang 0012, Yu Liu 0040, Yangtao Wang
Neural Comput. Appl.1
2020 Fast Graph Convolution Network Based Multi-label Image Recognition via Cross-modal Fusion
abstract
In multi-label image recognition, it has become a popular method to predict those labels that co-occur in an image via modeling the label dependencies. Previous works focus on capturing the correlation between labels, but neglect to effectively fuse the image features and label embeddings, which severely affects the convergence efficiency of the model and inhibits the further precision improvement of multi-label image recognition. To overcome this shortcoming, in this paper, we introduce Multi-modal Factorized Bilinear pooling (MFB) which works as an efficient component to fuse cross-modal embeddings and propose F-GCN, a fast graph convolution network (GCN) based multi-label image recognition model. F-GCN consists of three key modules: (1) an image representation learning module which adopts a convolution neural network (CNN) to learn and generate image representations, (2) a label co-occurrence embedding module which first obtains the label vectors via the word embeddings technique and then adopts GCN to capture label co-occurrence embeddings and (3) an MFB fusion module which efficiently fuses these cross-modal vectors to enable an end-to-end model with a multi-label loss function. We conduct extensive experiments on two multi-label datasets including MS-COCO and VOC2007. Experimental results demonstrate the MFB component efficiently fuses image representations and label co-occurrence embeddings and thus greatly improves the convergence efficiency of the model. In addition, the performance of image recognition has also been promoted compared with the state-of-the-art methods.
Yangtao Wang, Yanzhao Xie, Yu Liu 0040, Ke Zhou 0001, Xiaocui Li 0001
CIKM5
2020 A low cost and un-cancelled laplace noise based differential privacy algorithm for spatial decompositions
Xiaocui Li 0001, Yangtao Wang, Jingkuan Song, Yu Liu 0040, Xinyu Zhang 0012, Ke Zhou 0001, Chunhua Li 0002
World Wide Web1
2020 Semi-supervised clustering with deep metric learning and graph embedding
Xiaocui Li 0001, Hongzhi Yin, Ke Zhou 0001, Xiaofang Zhou 0001
World Wide Web1
2018 A More Secure Spatial Decompositions Algorithm via Indefeasible Laplace Noise in Differential Privacy
Xiaocui Li 0001, Yangtao Wang, Xinyu Zhang 0012, Ke Zhou 0001, Chunhua Li 0002
ADMA1