Xubin Li

dblp:228/7857 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
20since 2021 · last 2026
0000-0002-1547-4052ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 15 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment
abstract
Normalizing Flows (NFs) are a class of generative models distinguished by a mathematically invertible architecture, where the forward pass transforms data into a latent space for density estimation, and the reverse pass generates new samples from this space. This characteristic creates an intrinsic synergy between representation learning and data generation. However, the generative quality of standard NFs is limited by poor semantic representations from log-likelihood optimization. To remedy this, we propose a novel alignment strategy that creatively leverages the invertibility of NFs: instead of regularizing the forward pass, we align the intermediate features of the generative (reverse) pass with representations from a powerful vision foundation model, demonstrating superior effectiveness over naive alignment. We also introduce a novel training-free, test-time optimization algorithm for classification, which provides a more intrinsic evaluation of the NF's embedded semantic knowledge. Comprehensive experiments demonstrate that our approach accelerates the training of NFs by over 3.3x, while simultaneously delivering significant improvements in both generative quality and classification accuracy. New state-of-the-art results for NFs are established on ImageNet 64 x 64 and 256 x 256.
Chenhui Zhu, Ruxue Wen, Xubin Li, Tiezheng Ge, Limin Wang 0002
AAAI6
2026 Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction
abstract
With the rise of smart personal devices, service-oriented human-agent interactions have become increasingly prevalent. This trend highlights the need for personalized dialogue assistants that can understand user-specific traits to accurately interpret requirements and tailor responses to individual preferences. However, existing approaches often overlook the complexities of long-term interactions and fail to capture users’ subjective characteristics. To address these gaps, we present PAL-Bench, a new benchmark designed to evaluate the personalization capabilities of service-oriented assistants in long-term user-agent interactions. In the absence of available real-world data, we develop a multi-step LLM-based synthesis pipeline, which is further verified and refined by human annotators. This process yields PAL-Set, the first Chinese dataset comprising multi-session user logs and dialogue histories, which serves as the foundation for PAL-Bench. Furthermore, to improve personalized service-oriented interactions, we propose H2Memory, a hierarchical and heterogeneous memory framework that incorporates retrieval-augmented generation to improve personalized response generation. Comprehensive experiments on both our PAL-Bench and an external dataset demonstrate the effectiveness of the proposed memory framework.
Zhaopei Huang, Qifeng Dai, Guozheng Wu, Xubin Li, Tiezheng Ge, Wenxuan Wang 0001, Qin Jin
AAAI5
2026 Multi-modal large language model-based image captioning algorithm in information and communication technology: Bridging the gap between general and industry domain
Lianying Chao, Xubin Li, Linfeng Yin, Haoran Cai, Sijie Wu, Dingcheng Shan
Eng. Appl. Artif. Intell.3
2026 An attack detection mechanism in smart contracts based on deep learning and feature fusion
abstract
The rapid growth of Ethereum has spurred widespread adoption of smart contracts, enabling substantial financial transactions. Once deployed on the blockchain, smart contracts are immutable, rendering them unmodifiable even if vulnerabilities are present. In recent years, numerous attacks exploiting these vulnerabilities have caused significant financial losses. Although prior research has improved vulnerability detection in source code or bytecode before deployment, identifying attacks that exploit vulnerabilities during the execution phase after deployment remains a significant challenge. These challenges arise from the limited adaptability of predefined detection rules and an overreliance on opcode sequence names, which often neglects a comprehensive analysis of opcode sequence properties. In this study, we propose an advanced multidimensional feature fusion technique designed to detect attacks during the execution phase of smart contracts. By leveraging deep learning, our approach enhances detection accuracy through a comprehensive analysis of attack behaviors across four dimensions: operation objects, action behaviors, functional categories, and gas consumption. Extensive experiments demonstrate that our method achieves a detection accuracy of 97.21% and a weighted F1-score of 97.21%, confirming its effectiveness in identifying attacks.
Peiqiang Li, Guojun Wang 0001, Wanyi Gu, Xubin Li, Yuheng Zhang 0001
Inf. Sci.4
2025 RHanDS: Refining Malformed Hands for Generated Images with Decoupled Structure and Style Guidance
abstract
Although diffusion models can generate high-quality human images, their applications are limited by the instability in generating hands with correct structures. In this paper, we introduce RHanDS, a conditional diffusion-based framework designed to refine malformed hands by utilizing decoupled structure and style guidance. The hand mesh reconstructed from the malformed hand offers structure guidance for correcting the structure of the hand, while the malformed hand itself provides style guidance for preserving the style of the hand. To alleviate the mutual interference between style and structure guidance, we introduce a two-stage training strategy and build a series of multi-style hand datasets. In the first stage, we use paired hand images for training to ensure stylistic consistency in hand refining. In the second stage, various hand images generated based on human meshes are used for training, enabling the model to gain control over the hand structure. Experimental results demonstrate that RHanDS can effectively refine hand structure while preserving consistency in hand style.
Ming Zeng 0008, Xubin Li, Tiezheng Ge, Bo Zheng 0007
AAAI5
2025 Do not Abstain! Identify and Solve the Uncertainty
abstract
Despite the widespread application of Large Language Models (LLMs) across various domains, they frequently exhibit overconfidence when encountering uncertain scenarios, yet existing solutions primarily rely on evasive responses (e.g., “I don’t know”) overlooks the opportunity of identifying and addressing the uncertainty to generate more satisfactory responses. To systematically investigate and improve LLMs’ ability of recognizing and addressing the source of uncertainty, we introduce ConfuseBench, a benchmark mainly focus on three types of uncertainty: document scarcity, limited capability, and query ambiguity. Experiments with ConfuseBench reveal that current LLMs struggle to accurately identify the root cause of uncertainty and solve it. They prefer to attribute uncertainty to query ambiguity while overlooking capability limitations, especially for those weaker models. To tackle this challenge, we first generate context-aware inquiries that highlight the confusing aspect of the original query. Then we judge the source of uncertainty based on the uniqueness of the inquiry’s answer. Further we use an on-policy training method, InteractDPO to generate better inquiries. Experimental results demonstrate the efficacy of our approach.
JingquanPeng JingquanPeng, Xubin Li, Tiezheng Ge, Bo Zheng 0007, Yong Liu 0018
ACL (1)4
2025 Minimal Impact ControlNet: Advancing Multi-ControlNet Integration
abstract
With the advancement of diffusion models, there is a growing demand for high-quality, controllable image generation, particularly through methods that utilize one or multiple control signals based on ControlNet. However, in current ControlNet training, each control is designed to influence all areas of an image, which can lead to conflicts when different control signals are expected to manage different parts of the image in practical applications. This issue is especially pronounced with edge-type control conditions, where regions lacking boundary information often represent low-frequency signals, referred to as silent control signals. When combining multiple ControlNets, these silent control signals can suppress the generation of textures in related areas, resulting in suboptimal outcomes. To address this problem, we propose Minimal Impact ControlNet. Our approach mitigates conflicts through three key strategies: constructing a balanced dataset, combining and injecting feature signals in a balanced manner, and addressing the asymmetry in the score function’s Jacobian matrix induced by ControlNet. These improvements enhance the compatibility of control signals, allowing for freer and more harmonious generation in areas with silent control signals.
Shikun Sun, Zixuan Wang 0026, Xubin Li, Tiezheng Ge, Zijie Ye, Xiaoyu Qin 0001, Junliang Xing, Bo Zheng 0007, Jia Jia 0001
ICLR4
2025 Differentiable Solver Search for Fast Diffusion Sampling
abstract
Diffusion models have demonstrated remarkable generation quality but at the cost of numerous function evaluations. Recently, advanced ODE-based solvers have been developed to mitigate the substantial computational demands of reverse-diffusion solving under limited sampling steps. However, these solvers, heavily inspired by Adams-like multistep methods, rely solely on t-related Lagrange interpolation. We show that t-related Lagrange interpolation is suboptimal for diffusion model and reveal a compact search space comprised of time steps and solver coefficients. Building on our analysis, we propose a novel differentiable solver search algorithm to identify more optimal solver. Equipped with the searched solver, rectified-flow models, e.g., SiT-XL/2 and FlowDCN-XL/2, achieve FID scores of 2.40 and 2.35, respectively, on ImageNet-$256\times256$ with only 10 steps. Meanwhile, DDPM model, DiT-XL/2, reaches a FID score of 2.33 with only 10 steps. Notably, our searched solver outperforms traditional solvers by a significant margin. Moreover, our searched solver demonstrates generality across various model architectures, resolutions, and model sizes.
Zexian Li, Qipeng Zhang, Tianhui Song, Xubin Li, Tiezheng Ge, Bo Zheng 0007, Limin Wang 0002
ICML5
2025 Mc-music algorithm for fast DOA estimation
Fulai Liu, Xubin Li, Zhibo Su, Yubiao Liu, Aiyi Zhang
Wirel. Networks2
2024 Accelerating Image Generation with Sub-path Linear Approximation Model
Tianhui Song, Weixin Feng, Xubin Li, Tiezheng Ge, Bo Zheng 0007, Limin Wang 0002
ECCV (53)4
2024 Exploring DCN-like architecture for fast image generation with arbitrary resolution
abstract
Arbitrary-resolution image generation still remains a challenging task in AIGC, as it requires handling varying resolutions and aspect ratios while maintaining high visual quality. Existing transformer-based diffusion methods suffer from quadratic computation cost and limited resolution extrapolation capabilities, making them less effective for this task. In this paper, we propose FlowDCN, a purely convolution-based generative model with linear time and memory complexity, that can efficiently generate high-quality images at arbitrary resolutions. Equipped with a new design of learnable group-wise deformable convolution block, our FlowDCN yields higher flexibility and capability to handle different resolutions with a single model. FlowDCN achieves the state-of-the-art 4.30 sFID on $256\times256$ ImageNet Benchmark and comparable resolution extrapolation results, surpassing transformer-based counterparts in terms of convergence speed (only $\frac{1}{5}$ images), visual quality, parameters ($8\%$ reduction) and FLOPs ($20\%$ reduction). We believe FlowDCN offers a promising solution to scalable and flexible image synthesis.
Zexian Li, Tianhui Song, Xubin Li, Tiezheng Ge, Bo Zheng 0007, Limin Wang 0002
NeurIPS4
2024 TransFront: Bi-path Feature Fusion for Detecting Front-running Attack in Decentralized Finance
Yuheng Zhang 0001, Guojun Wang 0001, Peiqiang Li, Xubin Li, Wanyi Gu, Mingfei Chen, Houji Chen
TrustCom4
2023 BOMGraph: Boosting Multi-scenario E-commerce Search with a Unified Graph Neural Network
abstract
Mobile Taobao Application delivers search services on multiple scenarios that take textual, visual, or product queries. This paper aims to propose a unified graph neural network for these search scenarios to leverage data from multiple scenarios and jointly optimize search performances with less training and maintenance costs. Towards this end, this paper proposes BOMGraph, BOosting Multi-scenario E-commerce Search with a unified Graph neural network. BOMGraph is embodied with several components to address challenges in multi-scenario search. It captures heterogeneous information flow across scenarios by inter-scenario and intra-scenario metapaths. It learns robust item representations by disentangling specific characteristics for different scenarios and encoding common knowledge across scenarios. It alleviates label scarcity and long-tail problems in scenarios with low traffic by contrastive learning with cross-scenario augmentation. BOMGraph has been deployed in production by Alibaba's E-commerce search advertising platform. Both offline evaluations and online A/B tests demonstrate the effectiveness of BOMGraph.
Shuai Fan 0007, Jinping Gou, Yang Li 0213, Jiaxing Bai, Chen Lin 0001, Wanxian Guan, Xubin Li, Hongbo Deng, Jian Xu 0015, Bo Zheng 0007
CIKM7
2023 Long Short-Term Planning for Conversational Recommendation Systems
Hongguang Shi, Yeqin Zhang, Xubin Li, Cam-Tu Nguyen
ICONIP (6)5
2023 E-commerce Search via Content Collaborative Graph Neural Network
abstract
Recently, many E-commerce search models are based on Graph Neural Networks (GNNs). Despite their promising performances, they are (1) lacking proper semantic representation of product contents; (2) less efficient for industry-scale graphs; and (3) less accurate on long-tail queries and cold-start products. To address these problems simultaneously, this paper proposes CC-GNN, a novel Content Collaborative Graph Neural Network. Firstly, CC-GNN enables content phrases to participate explicitly in graph propagation to capture the proper meaning of phrases and semantic drifts. Secondly, CC-GNN presents several efforts towards a more scalable graph learning framework, including efficient graph construction, MetaPath-guided Message Passing, and Difficulty-aware Representation Perturbation for graph contrastive learning. Furthermore, CC-GNN adopts Counterfactual Data Supplement at both supervised and contrastive learning to resolve the long-tail/cold-start problems. Extensive experiments on a real E-commerce dataset of 100-million-scale nodes show that CC-GNN produces significant improvements over existing methods (i.e., more than 10% improvements in terms of several key evaluation metrics for overall, long-tail queries and cold-start products) while reducing computational complexity. The proposed components of CC-GNN can be applied to other models for search and recommendation tasks. Experiments on a public dataset show that applying the proposed components can improve the performance of different recommendation models.
Guipeng Xv, Chen Lin 0001, Wanxian Guan, Jinping Gou, Xubin Li, Hongbo Deng, Jian Xu 0015, Bo Zheng 0007
KDD5
2023 Multi-Scenario Ranking with Adaptive Feature Learning
abstract
Recently, Multi-Scenario Learning (MSL) is widely used in recommendation and retrieval systems in the industry because it facilitates transfer learning from different scenarios, mitigating data sparsity and reducing maintenance cost. These efforts produce different MSL paradigms by searching more optimal network structure, such as Auxiliary Network, Expert Network, and Multi-Tower Network. It is intuitive that different scenarios could hold their specific characteristics, activating the user's intents quite differently. In other words, different kinds of auxiliary features would bear varying importance under different scenarios. With more discriminative feature representations refined in a scenario-aware manner, better ranking performance could be easily obtained without expensive search for the optimal network structure. Unfortunately, this simple idea is mainly overlooked but much desired in real-world systems.
Yu Tian 0008, Bofang Li, Si Chen 0010, Xubin Li, Hongbo Deng, Jian Xu 0015, Bo Zheng 0007, Qian Wang 0002, Chenliang Li 0005
SIGIR4
2023 Differential privacy protection method for trip-oriented shared data
abstract
Summary While location information sharing technology provides convenience for unmanned driving and journey navigation, user journey information sharing has also become a disaster for privacy information leakage. The traditional differential privacy method can only perturb the data entirely and cannot consider the design of data availability. In this paper, the difference privacy algorithm is improved by combining it with the Apriori algorithm, and the relevant perturbation is carried out after mining the associated data of the user's trip. In the face of possible data attacks, the privacy protection of the sensitive information of the user's actual data is ensured while the availability of the data is ensured. By testing 3000 trip data generated by experimental simulation, the results show that the correlation information between the original datasets is destroyed. However good availability is maintained after the Laplace data perturbation of the proposed algorithm for both simultaneous and multi‐person trips.
Danlei Du, Tao Peng 0011, Xubin Li, Shaobo Zhang 0001, Tian Wang 0001
Concurr. Comput. Pract. Exp.5
2022 Visual Encoding and Debiasing for CTR Prediction
abstract
Extracting expressive visual features is crucial for accurate Click-Through-Rate (CTR) prediction in visual search advertising systems. Current commercial systems use off-the-shelf visual encoders to facilitate fast online service. However, the extracted visual features are coarse-grained and/or biased. In this paper, we present a visual encoding framework for CTR prediction to overcome these problems. The framework is based on contrastive learning which pulls positive pairs closer and pushes negative pairs apart in the visual feature space. To obtain fine-grained visual features, we present contrastive learning supervised by click-through data to fine-tune the visual encoder. To reduce sample selection bias, firstly we train the visual encoder offline by leveraging both unbiased self-supervision and click supervision signals. Secondly, we incorporate a debiasing network in the online CTR predictor to adjust the visual features by contrasting high impression items with selected, low impression items. We deploy the framework in a mobile E-commerce app. Offline experiments on billion-scale datasets and online experiments demonstrate that the proposed framework can make accurate and unbiased predictions.
Guipeng Xv, Si Chen 0010, Chen Lin 0001, Wanxian Guan, Xingyuan Bu, Xubin Li, Hongbo Deng, Jian Xu 0015, Bo Zheng 0007
CIKM6
2022 Pretraining Representations of Multi-modal Multi-query E-commerce Search
abstract
The importance of modeling contextual information within a search session has been widely acknowledged. However, learning representations of multi-query multi-modal (MM) search, in which Mobile Taobao users repeatedly submit textual and visual queries, remains unexplored in literature. Previous work which learns task-specific representations of textual query sessions fails to capture diverse query types and correlations in MM search sessions. This paper presents to represent MM search sessions by heterogeneous graph neural network (HGN). A multi-view contrastive learning framework is proposed to pretrain the HGN, with two views to model different intra-query, inter-query, and inter-modality information diffusion in MM search. Extensive experiments demonstrate that, the pretrained session representation can benefit state-of-the-art baselines on various downstream tasks, such as personalized click prediction, query suggestion, and intent classification.
Wanxian Guan, Lianyun Li, Hui Li 0057, Chen Lin 0001, Xubin Li, Si Chen 0010, Jian Xu 0015, Hongbo Deng, Bo Zheng 0007
KDD6
2021 Training Generative Adversarial Networks in One Stage
abstract
Generative Adversarial Networks (GANs) have demonstrated unprecedented success in various image generation tasks. The encouraging results, however, come at the price of a cumbersome training process, during which the generator and discriminator are alternately updated in two stages. In this paper, we investigate a general training scheme that enables training GANs efficiently in only one stage. Based on the adversarial losses of the generator and discriminator, we categorize GANs into two classes, Symmetric GANs and Asymmetric GANs, and introduce a novel gradient decomposition method to unify the two, allowing us to train both classes in one stage and hence alleviate the training effort. We also computationally analyze the efficiency of the proposed method, and empirically demonstrate that, the proposed method yields a solid 1.5× acceleration across various datasets and network architectures. Furthermore, we show that the proposed method is readily applicable to other adversarial-training scenarios, such as data-free knowledge distillation. The code is available at https://github.com/zju-vipa/OSGAN.
Chengchao Shen, Youtan Yin, Xinchao Wang, Xubin Li, Jie Song 0011, Mingli Song
CVPR4
2020 Tumor Classification Based on Approximate Symmetry Using Dual-Branch Complementary Fusion Network
abstract
MRI technology is usually used to distinguish the grade of the tumor in the patient. Due to technical limitations, the classification of tumors (high-grade gliomas and metastases) on MRI images has become a problem for doctors. At present, the widely used neural network is gradually applied to the tumor classification of MRI images, which not only reduces the burden of human resources, but also shows good classification accuracy. Although different good experimental data are obtained under various neural networks, there is still a problem in using these neural networks for tumor classification: the semantic information expressed on the image by the deep features of the neural network is too scattered, and it is difficult to concentrate the lesion area. In this article, we propose a new strategy that combines the approximate symmetry properties of the MRI image with neural network, then uses a dual-branch network instead of the basic network for feature extraction, and adds complementary learning to the network, different features fusion and attention mechanism to enrich detailed information. Our method performs multiple comparison and ablation experiments on the dataset of glioma and metastasis, which proves that the proposed method is effective for tumor classification assisted by MRI.
Mei Yu 0004, Minyutong Cheng, Xubin Li, Zhiqiang Liu 0002, Jie Gao 0008, Xuzhou Fu, Xuewei Li 0001
BIBM3