Yuliang Yan

dblp:236/1967 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0003-1456-6224ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
abstract
Cost-aware routing dynamically dispatches user queries to models of varying capability to balance performance and inference cost.However, the routing strategy introduces a new security concern that adversaries may manipulate the router to consistently select expensive highcapability models.Existing routing attacks depend on either white-box access or heuristic prompts, rendering them ineffective in realworld black-box scenarios.In this work, we propose R 2 A, which aims to mislead black-box LLM routers to expensive models via adversarial suffix optimization.Specifically, R 2 A deploys a hybrid ensemble surrogate router to mimic the black-box router.A suffix optimization algorithm is further adapted for the ensemble-based surrogate.Extensive experiments on multiple open-source and commercial routing systems demonstrate that R 2 A significantly increases the routing rate to expensive models on queries of different distributions.Code and examples: https://github.com/ thcxiker/R2A-Attack.
Haochun Tang, Yuliang Yan, Jiahua Lu, Huaxiao Liu, Enyan Dai
ACL (1)2
2026 Beyond Existing Retrievals: Cross-Scenario Incremental Sample Learning Framework
Xun Luo, Jinlong Guo, Yuliang Yan
WSDM4
2026 PI2I: A Personalized Item-Based Collaborative Filtering Retrieval Framework
abstract
Efficiently selecting relevant content from vast candidate pools is a critical challenge in modern recommender systems. Traditional methods, such as item-to-item collaborative filtering (CF) and two-tower models, often fall short in capturing the complex user-item interactions due to uniform truncation strategies and overdue user-item crossing. To address these limitations, we propose Personalized Item-to-Item (PI2I), a novel two-stage retrieval framework that enhances the personalization capabilities of CF. In the first Indexer Building Stage (IBS), we optimize the retrieval pool by relaxing truncation thresholds to maximize Hit Rate, thereby temporarily retaining more items users might be interested in. In the second Personalized Retrieval Stage (PRS), we introduce an interactive scoring model to overcome the limitations of inner product calculations, allowing for richer modeling of intricate user-item interactions. Additionally, we construct negative samples based on the trigger-target (item-to-item) relationship, ensuring consistency between offline training and online inference. Offline experiments on large-scale real-world datasets demonstrate that PI2I outperforms traditional CF methods and rivals Two-Tower models. Deployed in the ''Guess You Like'' section on Taobao, PI2I achieved a 1.05% increase in online transaction rates. In addition, we have released a large-scale recommendation dataset collected from Taobao, containing 130 million real-world user interactions used in the experiments of this paper. The dataset is publicly available at https://huggingface.co/datasets/PI2I/PI2I, which could serve as a valuable benchmark for the research community.
Yingcai Ma, Kairui Fu, Dunxian Huang, Yuliang Yan, Jian Wu 0032
WWW6
2025 TBGRecall: A Generative Retrieval Model for E-commerce Recommendation Scenarios
abstract
Recommendation systems are essential tools in modern e-commerce, facilitating personalized user experiences by suggesting relevant products. Recent advancements in generative models have demonstrated potential in enhancing recommendation systems; however, these models often exhibit limitations in optimizing retrieval tasks, primarily due to their reliance on autoregressive generation mechanisms. Conventional approaches introduce sequential dependencies that impede efficient retrieval, as they are inherently unsuitable for generating multiple items without positional constraints within a single request session. To address these limitations, we propose TBGRecall, a framework integrating Next Session Prediction (NSP), designed to enhance generative retrieval models for e-commerce applications. Our framework reformulation involves partitioning input samples into multi-session sequences, where each sequence comprises a session token followed by a set of item tokens, and then further incorporate multiple optimizations tailored to the generative task in retrieval scenarios. In terms of training methodology, our pipeline integrates limited historical data pre-training with stochastic partial incremental training, significantly improving training efficiency and emphasizing the superiority of data recency over sheer data volume. Our extensive experiments, conducted on public benchmarks alongside a large-scale industrial dataset from TaoBao, show TBGRecall outperforms the state-of-the-art recommendation methods, and exhibits a clear scaling law trend. Ultimately, NSP represents a significant advancement in the effectiveness of generative recommendation systems for e-commerce applications.
Zida Liang, Changfa Wu, Dunxian Huang, Weiqiang Sun, Yuliang Yan, Jian Wu 0032, Yuning Jiang 0001, Bo Zheng 0007, Silu Zhou, Yu Zhang 0176
CIKM6
2024 SlimGPT: Layer-wise Structured Pruning for Large Language Models
abstract
Large language models (LLMs) have garnered significant attention for their remarkable capabilities across various domains, whose vast parameter scales present challenges for practical deployment. Structured pruning is an effective method to balance model performance with efficiency, but performance restoration under computational resource constraints is a principal challenge in pruning LLMs. Therefore, we present a low-cost and fast structured pruning method for LLMs named SlimGPT based on the Optimal Brain Surgeon framework. We propose Batched Greedy Pruning for rapid and near-optimal pruning, which enhances the accuracy of head-wise pruning error estimation through grouped Cholesky decomposition and improves the pruning efficiency of FFN via Dynamic Group Size, thereby achieving approximate local optimal pruning results within one hour. Besides, we explore the limitations of layer-wise pruning from the perspective of error accumulation and propose Incremental Pruning Ratio, a non-uniform pruning strategy to reduce performance degradation. Experimental results on the LLaMA benchmark show that SlimGPT outperforms other methods and achieves state-of-the-art results.
Gui Ling, Yuliang Yan, Qingwen Liu 0002
NeurIPS3
2023 Hallucination Detection for Generative Large Language Models by Bayesian Sequential Estimation
abstract
Large Language Models (LLMs) have made remarkable advancements in the field of natural language generation.However, the propensity of LLMs to generate inaccurate or non-factual content, termed "hallucinations", remains a significant challenge.Current hallucination detection methods often necessitate the retrieval of great numbers of relevant evidence, thereby increasing response times.We introduce a unique framework that leverages statistical decision theory and Bayesian sequential analysis to optimize the trade-off between costs and benefits during the hallucination detection process.This approach does not require a predetermined number of observations.Instead, the analysis proceeds in a sequential manner, enabling an expeditious decision towards "belief" or "disbelief" through a stop-or-continue strategy.Extensive experiments reveal that this novel framework surpasses existing methods in both efficiency and precision of hallucination detection.Furthermore, it requires fewer retrieval steps on average, thus decreasing response times 1 .
Yuliang Yan, Longtao Huang, Xiaoqing Zheng, Xuanjing Huang 0001
EMNLP2
2019 Domain-aware Neural Model for Sequence Labeling using Joint Learning
abstract
Recently, scholars have demonstrated empirical successes of deep learning in sequence labeling, and most of the prior works focused on the word representation inside the target sentence. Unfortunately, the global information, e.g., domain information of the target document, were ignored in the previous studies. In this paper, we propose an innovative joint learning neural network which can encapsulate the global domain knowledge and the local sentence/token information to enhance the sequence labeling model. Unlike existing studies, the proposed method employs domain labeling output as a latent evidence to facilitate tagging model and such joint embedding information is generated by an enhanced highway network. Meanwhile, a redesigned CRF layer is deployed to bridge the 'local output labels' and 'global domain information'. Various kinds of information can iteratively contribute to each other, and moreover, domain knowledge can be learnt in either supervised or unsupervised environment via the new model. Experiment with multiple data sets shows that the proposed algorithm outperforms classical and most recent state-of-the-art labeling methods.
Yuliang Yan
WWW3
2019 Quality-Sensitive Training! Social Advertisement Generation by Leveraging User Click Behavior
abstract
Social advertisement has emerged as a viable means to improve purchase sharing in the context of e-commerce. However, humanly generating lots of advertising scripts can be prohibitive to both e-platforms and online sellers, and moreover, developing the desired auto-generator will need substantial gold-standard training samples. In this paper, we put forward a novel seq2seq model to generate social advertisements automatically, in which a quality-sensitive loss function is proposed based on user click behavior to differentiate training samples of varied qualities. Our motivation is to leverage the clickthrough data as a kind of quality indicator to measure the textual fitness of each training sample quantitatively, and only those ground truths that satisfy social media users will be considered the eligible and able to optimize the social advertisement generation. Specifically, under the qualified case, the ground truth should be utilized to supervise the whole training phase as much as possible, whereas in the opposite situation, the generated result ought to preserve the semantics of original input to the greatest extent. Simulation experiments on a large-scale dataset demonstrate that our approach achieves a significant superiority over two existing methods of distant supervision and three state-of-the-art NLG solutions.
Yongzhen Wang 0002, Yuliang Yan, Xiaozhong Liu 0001
WWW3