Xuhong Li 0002

dblp:76/5330-2 · DBLP profile ↗
← Back
25ranked-venue papers
13as first author
20since 2021 · last 2026
0000-0002-2582-8256ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 12 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SageCopilot: An LLM-Empowered Autonomous Agent for Data Science as a Service
abstract
While the field of natural language to SQL(NL2SQL) has made significant advancements in translating natural language instructions into executable SQL scripts for data querying and processing, achieving full automation within the broader data science pipeline–encompassing data querying, analysis, visualization, and reporting–remains a complex challenge. This study introduces SageCopilot, an advanced, industry-grade system that automates the data science pipeline by integrating Large Language Models (LLMs), Autonomous Agents (AutoAgents), and Language User Interfaces (LUIs). Designed with a two-phase architecture, SageCopilot uses an offline phase to generate high-quality demonstrations supporting In-Context Learning (ICL), which powers the online phase to transform user inputs into executable scripts for database queries, analysis, and visualization tasks. Leveraging specialized components such as NL2SQL, Text2Analyze, and Text2Viz, as well as chain-of-thought prompting for multi-turn interactions, SageCopilot achieves superior end-to-end automation. Rigorous experimentation with real-world datasets demonstrates the system's ability to minimize human intervention while ensuring correctness and user-friendly operation.
Yuan Liao 0003, Jiang Bian 0003, Yuhui Yun, Jiaming Chu, Yuchen Li 0006, Xuhong Li 0002, Shilei Ji, Haoyi Xiong
IEEE Trans. Serv. Comput.9
2025 SOLA-GCL: Subgraph-Oriented Learnable Augmentation Method for Graph Contrastive Learning
abstract
Graph contrastive learning has emerged as a powerful technique for learning graph representations that are robust and discriminative. However, traditional approaches often neglect the critical role of subgraph structures, particularly the intra-subgraph characteristics and inter-subgraph relationships, which are crucial for generating informative and diverse contrastive pairs. These subgraph features are crucial as they vary significantly across different graph types, such as social networks where they represent communities, and biochemical networks where they symbolize molecular interactions. To address this issue, our work proposes a novel subgraph-oriented learnable augmentation method for graph contrastive learning, termed SOLA-GCL, that centers around subgraphs, taking full advantage of the subgraph information for data augmentation. Specifically, SOLA-GCL initially partitions a graph into multiple densely connected subgraphs based on their intrinsic properties. To preserve and enhance the unique characteristics inherent to subgraphs, a graph view generator optimizes augmentation strategies for each subgraph, thereby generating tailored views for graph contrastive learning. This generator uses a combination of intra-subgraph and inter-subgraph augmentation strategies, including node dropping, feature masking, intra-edge perturbation, inter-edge perturbation, and subgraph swapping. Extensive experiments have been conducted on various graph learning applications, ranging from social networks to molecules, under semi-supervised learning, unsupervised learning, and transfer learning settings to demonstrate the superiority of our proposed approach.
Tianhao Peng 0002, Xuhong Li 0002, Haitao Yuan 0002, Yuchen Li 0006, Haoyi Xiong
AAAI2
2025 RankElectra: Semi-supervised Pre-training of Learning-to-Rank Electra for Web-scale Search
abstract
While representation learning has been used to boost the performance of Learning-to-Rank (LTR) models through distilling key features for webpage ranking, the weak supervision signals extracted from users' sparse click-through data lead to inadequate representation of query-webpage pairs for ranking score prediction. Recent studies in generative LTR pre-training demonstrate the feasibility of incorporating reconstruction loss for enhanced ranking score prediction. However, LTR is afterall a regression task and it might be reasonable to find an alternate route that pre-trains LTR models with discriminative losses. Following the success of Electra in representation learning for natural language processing (NLP), this work proposes RankElectra that pre-trains the LTR model as a discriminator module inside a generative learning framework. Specifically, RankElectra first structures sparsely-annotated query-webpage pairs into a bipartite graph, with query and webpage feature vectors as node types and ranking scores as the connecting edges, and then leverages positive and negative extension strategies to densify the graph by link predictions. Later, this work proposes a novel Electra module that pre-trains the LTR model as a discriminator module for node reconstruction tasks, where node features of selected edges would be randomly masked and reconstructed by a generator, and the discriminator learns to classify whether the reconstructed features are the original or replaced as well as perform correct ranking. Finally, the pre-trained discriminator module, rather than the generator, would be fine-tuned on the labeled graph. We carried out extensive offline and online evaluations using the real-world web traffic of Baidu search engine. The results show that RankElectra could significantly boost the ranking performance of Baidu Search compared with numbers of competitor systems.
Yuchen Li 0006, Haoyi Xiong, Jiang Bian 0003, Tianhao Peng 0002, Xuhong Li 0002, Shuaiqiang Wang, Linghe Kong, Dawei Yin 0001
KDD (1)6
2024 G-LIME: Statistical Learning for Local Interpretations of Deep Neural Networks Using Global Priors (Abstract Reprint)
abstract
To explain the prediction result of a Deep Neural Network (DNN) model based on a given sample, LIME [1] and its derivatives have been proposed to approximate the local behavior of the DNN model around the data point via linear surrogates. Though these algorithms interpret the DNN by finding the key features used for classification, the random interpolations used by LIME would perturb the explanation result and cause the instability and inconsistency between repetitions of LIME computations. To tackle this issue, we propose G-LIME that extends the vanilla LIME through high-dimensional Bayesian linear regression using the sparsity and informative global priors. Specifically, with a dataset representing the population of samples (e.g., the training set), G-LIME first pursues the global explanation of the DNN model using the whole dataset. Then, with a new data point, -LIME incorporates an modified estimator of ElasticNet-alike to refine the local explanation result through balancing the distance to the global explanation and the sparsity/feature selection in the explanation. Finally, G-LIME uses Least Angle Regression (LARS) and retrieves the solution path of a modified ElasticNet under varying -regularization, to screen and rank the importance of features [2] as the explanation result. Through extensive experiments on real world tasks, we show that the proposed method yields more stable, consistent, and accurate results compared to LIME.
Xuhong Li 0002, Haoyi Xiong, Xingjian Li 0002, Xiao Zhang 0001, Ji Liu 0003, Dejing Dou
AAAI1
2024 HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
abstract
Large language models (LLMs) have made significant progress in generating codes from textual prompts. However, existing benchmarks have mainly concentrated on translating English prompts to multilingual codes or have been constrained to very limited natural languages (NLs). These benchmarks have overlooked the vast landscape of massively multilingual NL to multilingual code, leaving a critical gap in the evaluation of multilingual LLMs. In response, we introduce HumanEval-XL, a massively multilingual code generation benchmark specifically crafted to address this deficiency. HumanEval-XL establishes connections between 23 NLs and 12 programming languages (PLs), and comprises of a collection of 22,080 prompts with an average of 8.33 test cases. By ensuring parallel data across multiple NLs and PLs, HumanEval-XL offers a comprehensive evaluation platform for multilingual LLMs, allowing the assessment of the understanding of different NLs. Our work serves as a pioneering step towards filling the void in evaluating NL generalization in the area of multilingual code generation. We make our evaluation code and data publicly available at https://github.com/FloatAI/HumanEval-XL.
Qiwei Peng 0002, Yekun Chai, Xuhong Li 0002
LREC/COLING3
2024 GiLOT: Interpreting Generative Language Models via Optimal Transport
abstract
While large language models (LLMs) surge with the rise of generative AI, algorithms to explain LLMs highly desire. Existing feature attribution methods adequate for discriminative language models like BERT often fail to deliver faithful explanations for LLMs, primarily due to two issues: (1) For every specific prediction, the LLM outputs a probability distribution over the vocabulary–a large number of tokens with unequal semantic distance; (2) As an autoregressive language model, the LLM handles input tokens while generating a sequence of probability distributions of various tokens. To address above two challenges, this work proposes GiLOT that leverages Optimal Transport to measure the distributional change of all possible generated sequences upon the absence of every input token, while taking into account the tokens’ similarity, so as to faithfully estimate feature attribution for LLMs. We have carried out extensive experiments on top of Llama families and their fine-tuned derivatives across various scales to validate the effectiveness of GiLOT for estimating the input attributions. The results show that GiLOT outperforms existing solutions on a number of faithfulness metrics under fair comparison settings. Source code is publicly available at https://github.com/holyseven/GiLOT.
Xuhong Li 0002, Yekun Chai, Haoyi Xiong
ICML1
2024 Towards accurate knowledge transfer via target-awareness representation disentanglement
abstract
Abstract Fine-tuning deep neural networks pre-trained on large scale datasets is one of the most practical transfer learning paradigm given limited quantity of training samples. To obtain better generalization, using the starting point as the reference (SPAR), either through weights or features, has been successfully applied to transfer learning as a regularizer. However, due to the domain discrepancy between the source and target task, there exists obvious risk of negative transfer in a straightforward manner of knowledge preserving. In this paper, we propose a novel transfer learning algorithm, introducing the idea of Target-awareness REpresentation Disentanglement ( $$\textrm{TRED}$$ TRED ), where the relevant knowledge with respect to the target task is disentangled from the original source model and used as a regularizer during fine-tuning the target model. Two alternative approaches, maximizing Maximum Mean Discrepancy (Max-MMD) and minimizing mutual information (Min-MI) are introduced to achieve the desired disentanglement. Experiments on various real world datasets show that our method stably improves the standard fine-tuning by more than 2% in average. $$\textrm{TRED}$$ TRED also outperforms related state-of-the-art transfer learning regularizers such as $$\mathrm {L^2\text {-}SP}$$ L 2 - SP , $$\textrm{AT}$$ AT , $$\textrm{DELTA}$$ DELTA , and $$\textrm{BSS}$$ BSS . Moreover, our solution is compatible with different choices of disentangling strategies. While the combination of Max-MMD and Min-MI typically achieves higher accuracy, only using Max-MMD can be a preferred choice in applications with low resource budgets.
Xingjian Li 0002, Di Hu 0001, Xuhong Li 0002, Haoyi Xiong, Cheng-Zhong Xu 0001, Dejing Dou
Mach. Learn.3
2024 P2ANet: A Large-Scale Benchmark for Dense Action Detection from Table Tennis Match Broadcasting Videos
abstract
While deep learning has been widely used for video analytics, such as video classification and action detection, dense action detection with fast-moving subjects from sports videos is still challenging. In this work, we release yet another sports video benchmark P 2 ANet for P ing P ong- A ction detection, which consists of 2,721 video clips collected from the broadcasting videos of professional table tennis matches in World Table Tennis Championships and Olympiads. We work with a crew of table tennis professionals and referees on a specially designed annotation toolbox to obtain fine-grained action labels (in 14 classes) for every ping-pong action that appeared in the dataset, and formulate two sets of action detection problems— action localization and action recognition . We evaluate a number of commonly seen action recognition (e.g., TSM, TSN, Video SwinTransformer, and Slowfast) and action localization models (e.g., BSN, BSN++, BMN, TCANet), using P 2 ANet for both problems, under various settings. These models can only achieve 48% area under the AR-AN curve for localization and 82% top-one accuracy for recognition since the ping-pong actions are dense with fast-moving subjects but broadcasting videos are with only 25 FPS. The results confirm that P 2 ANet is still a challenging task and can be used as a special benchmark for dense action detection from videos. We invite readers to examine our dataset by visiting the following link: https://github.com/Fred1991/P2ANET .
Jiang Bian 0003, Xuhong Li 0002, Tao Wang 0011, Qingzhong Wang, Feixiang Lu, Dejing Dou, Haoyi Xiong
ACM Trans. Multim. Comput. Commun. Appl.2
2024 When Search Engine Services Meet Large Language Models: Visions and Challenges
abstract
Combining Large Language Models (LLMs) with search engine services marks a significant shift in the field of services computing, opening up new possibilities to enhance how we search for and retrieve information, understand content, and interact with internet services. This paper conducts an in-depth examination of how integrating LLMs with search engines can mutually benefit both technologies. We focus on two main areas: using search engines to improve LLMs (Search4LLM) and enhancing search engine functions using LLMs (LLM4Search). For Search4LLM, we investigate how search engines can provide diverse high-quality datasets for pre-training of LLMs, how they can use the most relevant documents to help LLMs learn to answer queries more accurately, how training LLMs with Learning-To-Rank (LTR) tasks can enhance their ability to respond with greater precision, and how incorporating recent search results can make LLM-generated content more accurate and current. In terms of LLM4Search, we examine how LLMs can be used to summarize content for better indexing by search engines, improve query outcomes through optimization, enhance the ranking of search results by analyzing document relevance, and help in annotating data for learning-to-rank tasks in various learning contexts. However, this promising integration comes with its challenges, which include addressing potential biases and ethical issues in training models, managing the computational and other costs of incorporating LLMs into search services, and continuously updating LLM training with the ever-changing web content. We discuss these challenges and chart out required research directions to address them. We also discuss broader implications for service computing, such as scalability, privacy concerns, and the need to adapt search engine architectures for these advanced models.
Haoyi Xiong, Jiang Bian 0003, Yuchen Li 0006, Xuhong Li 0002, Mengnan Du, Shuaiqiang Wang, Dawei Yin 0001, Abdelsalam Helal
IEEE Trans. Serv. Comput.4
2023 Learning from Training Dynamics: Identifying Mislabeled Data beyond Manually Designed Features
abstract
While mislabeled or ambiguously-labeled samples in the training set could negatively affect the performance of deep models, diagnosing the dataset and identifying mislabeled samples helps to improve the generalization power. Training dynamics, i.e., the traces left by iterations of optimization algorithms, have recently been proved to be effective to localize mislabeled samples with hand-crafted features. In this paper, beyond manually designed features, we introduce a novel learning-based solution, leveraging a noise detector, instanced by an LSTM network, which learns to predict whether a sample was mislabeled using the raw training dynamics as input. Specifically, the proposed method trains the noise detector in a supervised manner using the dataset with synthesized label noises and can adapt to various datasets (either naturally or synthesized label-noised) without retraining. We conduct extensive experiments to evaluate the proposed method. We train the noise detector based on the synthesized label-noised CIFAR dataset and test such noise detector on Tiny ImageNet, CUB-200, Caltech-256, WebVision and Clothing1M. Results show that the proposed method precisely detects mislabeled samples on various datasets without further adaptation, and outperforms state-of-the-art methods. Besides, more experiments demonstrate that the mislabel identification can guide a label correction, namely data debugging, providing orthogonal improvements of algorithm-centric state-of-the-art techniques from the data aspect.
Qingrui Jia, Xuhong Li 0002, Jiang Bian 0003, Penghao Zhao, Haoyi Xiong, Dejing Dou
AAAI2
2023 ContRE: A Complementary Measure for Robustness Evaluation of Deep Networks via Contrastive Examples
abstract
Training images with data transformations, e.g., crops, shifts, rotations and color distortions, have been suggested as contrastive examples to evaluate the robustness of deep neural networks against data noises [1]. In this work, we propose a practical framework ContRE (which is the meaning of “against” in French) that uses Contrastive examples for DNN Robustness Estimation. Specifically, ContRE follows the assumption in [2], [3] that robust DNN models with good generalization performance are capable of extracting a consistent set of features and making consistent predictions from the same image under varying data transformations. Incorporating with a set of randomized strategies for well-designed data transformations over the training set, ContREadopts classification errors and Fisher ratios on the generated contrastive examples to assess and analyze the robustness of DNN models, which correlates to the models’ generalization performance. To show the effectiveness and efficiency of ContRE, extensive experiments have been done using various DNN models, e.g., ResNet, VGGNet, DenseNet, EfficientNet, etc., on three open source benchmark datasets, i.e., CIFAR-10, CIFAR-100, and ImageNet, with thorough ablation studies and applicability analyses. Our experiment results confirm that ❨1❩ behaviors of deep models on contrastive examples are strongly correlated to what on the testing set, and ❨2❩ the robustness that ContRE calculates is a robust measure of generalization performance complementing to the testing set in various settings. Codes is to be publicly available.
Xuhong Li 0002, Xuanyu Wu, Linghe Kong, Xiao Zhang 0001, Siyu Huang, Dejing Dou, Haoyi Xiong
ICDM1
2023 M4: A Unified XAI Benchmark for Faithfulness Evaluation of Feature Attribution Methods across Metrics, Modalities and Models
Xuhong Li 0002, Mengnan Du, Yekun Chai, Himabindu Lakkaraju, Haoyi Xiong
NeurIPS1
2023 G-LIME: Statistical learning for local interpretations of deep neural networks using global priors
Xuhong Li 0002, Haoyi Xiong, Xingjian Li 0002, Xiao Zhang 0001, Ji Liu 0003, Dejing Dou
Artif. Intell.1
2023 Cross-model consensus of explanations and beyond for image classification models: an empirical study
Xuhong Li 0002, Haoyi Xiong, Siyu Huang, Shilei Ji, Dejing Dou
Mach. Learn.1
2023 Distilling ensemble of explanations for weakly-supervised pre-training of image segmentation models
Xuhong Li 0002, Haoyi Xiong, Yi Liu 0040, Dingfu Zhou, Yaqing Wang 0002, Dejing Dou
Mach. Learn.1
2023 Feynman: Federated Learning-Based Advertising for Ecosystems-Oriented Mobile Apps Recommendation
abstract
While recommender systems have been ubiquitously used in digital marketing and online business development, the conversions of online advertising for mobile apps installation and activation sometimes are far from satisfactory, due to the lack of feedback from App-related activities, leading to a poor record of Return on Investment (RoI). Though the advertisers, e.g., App operators and App Store, are granted to log users’ app-related activities such as installation, activation, usages, and preferences per the agreement, they usually limit the access to such data from advertisement publishers, due to the privacy concerns. To improve conversions of online advertising under privacy controls, we proposeFeynman—afederated learning-based advertising platform for ecosystems-orientedmobileapps recommendation.Feynmanaims at improving the RoI of mobile app recommendation from an ecosystem's perspective, i.e., per investment in advertising an app (Goal. 1) increasing the number of new installs/users of the app, and then (Goal. 2) increasing the number of new active users (preferably with frequent in-app purchase activities). Incorporating with a federated computing platform,Feynmanleverages users’ records stored in advertisers to refine the pool of targeting users for ads distribution, and jointly builds the predictive models for users’ purchase activities forecasting using features from the Ads publisher and the advertiser. With refined target pools and more accurate models,Feynmanhas successfully helped several mobile apps in China by attracting more than 100 million users to further enlarge their user populations and revenues from in-app purchases. Note that rather than proposing new techniques for federated learning, the design ofFeynmandedicates to show its promising performance in the industrial practices of advertising using federated computing and privacy protected strategies. In three cases that we report in this paper,Feynmanoutperforms the state-of-the-art plans in terms of several key measurements, including Click-Through Rates (CTR), Conversion Rate (CVR), Cost per Action (CPA), and Non-targeting User Hit-Rates (NTHR).
Jiang Bian 0003, Jizhou Huang, Shilei Ji, Yuan Liao 0003, Xuhong Li 0002, Qingzhong Wang, Jingbo Zhou 0003, Dejing Dou, Yaqing Wang 0002, Haoyi Xiong
IEEE Trans. Serv. Comput.5
2022 MUSCLE: Multi-task Self-supervised Continual Learning to Pre-train Deep Models for X-Ray Images of Multiple Body Parts
Weibin Liao, Haoyi Xiong, Qingzhong Wang, Yan Mo, Xuhong Li 0002, Yi Liu 0040, Siyu Huang, Dejing Dou
MICCAI (8)5
2022 InterpretDL: Explaining Deep Models in PaddlePaddle
abstract
Techniques to explain the predictions of deep neural networks (DNNs) have been largely required for gaining insights into the black boxes. We introduce InterpretDL, a toolkit of explanation algorithms based on PaddlePaddle, with uniformed programming interfaces and "plug-and-play" designs. A few lines of codes are needed to obtain the explanation results without modifying the structure of the model. InterpretDL currently contains 16 algorithms, explaining training phases, datasets, global and local behaviors of post-trained deep models. InterpretDL also provides a number of tutorial examples and showcases to demonstrate the capability of InterpretDL working on a wide range of deep learning models, e.g., Convolutional Neural Networks (CNNs), Multi-Layer Preceptors (MLPs), Transformers, etc., for various tasks in both Computer Vision (CV) and Natural Language Processing (NLP). Furthermore, InterpretDL modularizes the implementations, making efforts to support the compatibility across frameworks. The project is available at https://github.com/PaddlePaddle/InterpretDL.
Xuhong Li 0002, Haoyi Xiong, Xingjian Li 0002, Xuanyu Wu, Dejing Dou
J. Mach. Learn. Res.1
2022 Interpretable deep learning: interpretation, interpretability, trustworthiness, and beyond
Xuhong Li 0002, Haoyi Xiong, Xingjian Li 0002, Xuanyu Wu, Xiao Zhang 0001, Ji Liu 0003, Jiang Bian 0003, Dejing Dou
Knowl. Inf. Syst.1
2022 From distributed machine learning to federated learning: a survey
Ji Liu 0003, Jizhou Huang, Yang Zhou 0001, Xuhong Li 0002, Shilei Ji, Haoyi Xiong, Dejing Dou
Knowl. Inf. Syst.4
2020 Cross-Task Transfer for Geotagged Audiovisual Aerial Scene Recognition
Di Hu 0001, Xuhong Li 0002, Lichao Mou, Pu Jin, Liping Jing, Xiao Xiang Zhu 0001, Dejing Dou
ECCV (24)2
2020 Transfer learning in computer vision tasks: Remember where you come from
Xuhong Li 0002, Yves Grandvalet, Franck Davoine, Jingchun Cheng, Yin Cui, Han Zhang 0005, Serge J. Belongie, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001
Image Vis. Comput.1
2020 A baseline regularization scheme for transfer learning with convolutional neural networks
Xuhong Li 0002, Yves Grandvalet, Franck Davoine
Pattern Recognit.1
2018 Explicit Inductive Bias for Transfer Learning with Convolutional Networks
abstract
In inductive transfer learning, fine-tuning pre-trained convolutional networks substantially outperforms training from scratch. When using fine-tuning, the underlying assumption is that the pre-trained model extracts generic features, which are at least partially relevant for solving the target task, but would be difficult to extract from the limited amount of data available on the target task. However, besides the initialization with the pre-trained model and the early stopping, there is no mechanism in fine-tuning for retaining the features learned on the source task. In this paper, we investigate several regularization schemes that explicitly promote the similarity of the final solution with the initial model. We show the benefit of having an explicit inductive bias towards the initial model, and we eventually recommend a simple $L^2$ penalty with the pre-trained model being a reference as the baseline of penalty for transfer learning tasks.
Xuhong Li 0002, Yves Grandvalet, Franck Davoine
ICML1
2018 A Simple Weight Recall for Semantic Segmentation: Application to Urban Scenes
abstract
In many learning tasks, including semantic image segmentation, performance can be effectively improved through the fine-tuning of a pre-trained convolutional network, instead of training from scratch. With fine-tuning, the underlying assumption is that the pre-trained model extracts generic features, which are at least partially relevant for solving a segmentation task, but that would be difficult to extract from the smaller amount of data that is available for training in urban driving scenes segmentation. However, besides the initialization with the pre-trained model and the early stopping, there is no mechanism in classical fine-tuning approaches for keeping the generic features. Even worse, the standard weight decay drives the parameters towards the origin and affects the learned features. In this paper, we show that a simple regularization that uses the pre-trained model as a reference consistently improves the performance when applied to semantic urban driving scene segmentation. Experiments are done on the Cityscapes dataset, with four different architectures of convolutional networks.
Xuhong Li 0002, Franck Davoine, Yves Grandvalet
Intelligent Vehicles Symposium1