Yueyang Su

dblp:260/8182 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-6154-8059ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Distilling Large Embeddings via Hyperspherical Householder Quantization
abstract
Yihang Wang, Bin Wu, Yueyang Su, Tianfu Zhang, Yiqi Du, Lei Yu, Jiafeng Guo, Xueqi Cheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yueyang Su, Tianfu Zhang, Yiqi Du, Lei Yu 0012, Jiafeng Guo, Xueqi Cheng 0001
ACL (1)3
2025 MDPO: Customized Direct Preference Optimization with a Metric-based Sampler for Question and Answer Generation
abstract
With the extensive use of large language models, automatically generating QA datasets for domain-specific fine-tuning has become crucial. However, considering the multifaceted demands for readability, diversity, and comprehensiveness of QA data, current methodologies fall short in producing high-quality QA datasets. Moreover, the dependence of existing evaluation metrics on ground truth labels further exacerbates the challenges associated with the selection of QA data. In this paper, we introduce a novel method for QA data generation, denoted as MDPO. We proposes a set of unsupervised evaluation metrics for QA data, enabling multidimensional assessment based on the relationships among context,question and answer. Furthermore, leveraging these metrics, we implement a customized direct preference optimization process that guides large language models to produce high-quality and domain-specific QA pairs. Empirical results on public datasets indicate that MDPO’s performance substantially surpasses that of state-of-the-art methods.
Yueyang Su, Yixing Fan, Jiafeng Guo
COLING3
2025 Improving multimodal named entity recognition via text-image relevance prediction with large language models
Qingyang Zeng, Minghui Yuan, Yueyang Su, Qianzi Che
Neurocomputing3
2024 CausalMMM: Learning Causal Structure for Marketing Mix Modeling
abstract
In online advertising, marketing mix modeling (MMM) is employed to predict the gross merchandise volume (GMV) of brand shops and help decision-makers to adjust the budget allocation of various advertising channels. Traditional MMM methods leveraging regression techniques can fail in handling the complexity of marketing. Although some efforts try to encode the causal structures for better prediction, they have the strict restriction that causal structures are prior-known and unchangeable. In this paper, we define a new causal MMM problem that automatically discovers the interpretable causal structures from data and yields better GMV predictions. To achieve causal MMM, two essential challenges should be addressed: (1) Causal Heterogeneity. The causal structures of different kinds of shops vary a lot. (2) Marketing Response Patterns. Various marketing response patterns i.e., carryover effect and shape effect, have been validated in practice. We argue that causal MMM needs dynamically discover specific causal structures for different shops and the predictions should comply with the prior known marketing response patterns. Thus, we propose CausalMMM that integrates Granger causality in a variational inference framework to measure the causal relationships between different channels and predict the GMV with the regularization of both temporal and saturation marketing response patterns. Extensive experiments show that CausalMMM can not only achieve superior performance of causal structure learning on synthetic datasets with improvements of 5.7%\sim 7.1%, but also enhance the GMV prediction results on a representative E-commerce platform.
Chang Gong 0001, Di Yao 0001, Lei Zhang 0206, Wenbin Li 0012, Yueyang Su, Jingping Bi
WSDM6
2023 Few-shot learning for trajectory outlier detection with only normal trajectories
abstract
Trajectory outlier detection has a profound impact on many real-world applications. Most existing methods, whether supervised or unsupervised, require adequate historical data as a prerequisite. However, due to the uneven distribution of the trajectories in space, the detection in some regions(e.g. rural areas) may suffer from data scarcity. Furthermore, because the occurrence of outliers is a small probability event, there are frequently no outlier trajectories in such limited data which makes it even worse. To handle such issues, we in this paper study a new problem, i.e. few-shot trajectory outlier detection with only normal trajectories. And we propose a novel model, named MetaTAD, which consists of a Multi-scale Encoder and an Ab-detection Meta Learner. The Multi-scale Encoder aims to learn the diverse features of trajectories for outlier detection, while the Ab-detection Meta Learner guides the model to detect with few labeled normal trajectories. Extensive experiments on two real taxi trajectory datasets show that MetaTAD achieves state-of-the-art performance compared with the baselines.
Yueyang Su, Di Yao 0001, Jingping Bi
IJCNN1
2023 TripSafe: Retrieving Safety-related Abnormal Trips in Real-time with Trajectory Data
abstract
Nowadays safety has become one of the most critical factors for ride-hailing service. Ride-hailing platforms have conducted meticulous background checks for drivers to minimize the risk of abnormal trips, e.g. violence and sexual assault. However, current methods are labor-consuming and highly rely on the personal information of drivers, which may harm the fairness of the order dispatching system. In this paper, we utilize the trip trajectories as inputs and propose a dual variational auto-encoder(VAE) framework, namely TripSafe, to estimate the probability of abnormal safety incidents. Specifically, TripSafe models the moving behavior and route information, as two independent components and employs VAEs to pre-train generative models for normal trips. Then, a fusion network is adopted to fine-tune the whole model with a few labeled samples. In practice, TripSafe monitors the data update and calculate the anomaly score of partial-observed trips in real-time. Experiments on real ridehailing data show that TripSafe is superior to the state-of-the-art baselines with about 14.2%~28.9% improvements on F1 score.
Yueyang Su, Di Yao 0001, Yunxia Fan, Jingping Bi
SIGIR1
2022 Trajectory-Based Mobile Game Bots Detection with Gaussian Mixture Model
Yueyang Su, Di Yao 0001, Jingping Bi, Runze Wu 0001, Jianrong Tao
ICANN (3)1
2022 GTAT: Adversarial Training with Generated Triplets
abstract
To circumvent the grave problem of present adversarial training methods, i.e. distortion of classification surface, we in this paper propose a generated Triplet-based adversarial training method-GTAT, in which a Generator generates a semi-hard Triplet by design, rather than directly invoking the existing clean examples and adversarial examples. Through this kind of generated semi-hard Triplet constraint, GTAT can reshape the classification boundaries appropriately across various classes, arising from two-facet synergies: i) pull the intra-class examples together with tight distances; and ii) push away the inter-class examples with broad distances. This synergy will simplify and broaden the classification surfaces across different classes. Extensive experiments on the popular MNIST and CIFAR-10 datasets show that our proposed GTAT significantly outperforms other state-of-the-art adversarial training methods. We believe GTAT opens a door for the adversarial training from a new horizon of rationally generating semi-hard Triplet-satisfied adversarial training (retraining) examples, instead of straightly performing retraining on the generated adversarial examples and existing clean examples, or on the generated adversarial examples only.
Xinxin Fan, Quanliang Jing, Yueyang Su, Jingping Bi
IJCNN4
2022 Few-shot Learning for Trajectory-based Mobile Game Cheating Detection
abstract
With the emerging of smartphones, mobile games have attracted billions of players and occupied most of the share for game companies. On the other hand, mobile game cheating, aiming to gain improper advantages by using programs that simulate the players' inputs, severely damages the game's fairness and harms the user experience. Therefore, detecting mobile game cheating is of great importance for mobile game companies. Many PC game-oriented cheating detection methods have been proposed in the past decades, however, they can not be directly adopted in mobile games due to the concern of privacy, power, and memory limitations of mobile devices. Even worse, in practice, the cheating programs are quickly updated, leading to the label scarcity for novel cheating patterns. To handle such issues, we in this paper introduce a mobile game cheating detection framework, namely FCDGame, to detect the cheats under the few-shot learning framework. FCDGame only consumes the screen sensor data, recording users' touch trajectories, which is less sensitive and more general for almost all mobile games. Moreover, a Hierarchical Trajectory Encoder and a Cross-pattern Meta Learner are designed in FCDGame to capture the intrinsic characters of mobile games and solve the label scarcity problem, respectively. Extensive experiments on two real online games show that FCDGame achieves almost 10% improvements in detection accuracy with only few fine-tuned samples.
Yueyang Su, Di Yao 0001, Xiaokai Chu, Wenbin Li 0012, Jingping Bi, Runze Wu 0001, Shize Zhang, Jianrong Tao
KDD1
2022 FingFormer: Contrastive Graph-based Finger Operation Transformer for Unsupervised Mobile Game Bot Detection
abstract
This paper studies the task of detecting bots for online mobile games. Considering the fact of lacking labeled cheating samples and restricted available data in the real detection systems, we aim to study the finger operations captured by screen sensors to infer the potential bots in an unsupervised way. In detail, we introduce a Transformer-style detection model, namely FingFormer. It studies the finger operations in the format of graph structure in order to capture the spatial and temporal relatedness between the two hands’ operations. To optimize the model in an unsupervised way, we introduce two contrastive learning strategies to refine both finger moving patterns and players’ operation habits. We conduct extensive experiments under different experimental environments, including the synthetic dataset, the offline dataset, as well as the large-scale online data flow from three mobile games. The multi-facet experiments illustrate the proposed model is both effective and general to detect the bots for different mobile games.
Wenbin Li 0012, Xiaokai Chu, Yueyang Su, Di Yao 0001, Runze Wu 0001, Shize Zhang, Jianrong Tao, Jingping Bi
WWW3