Yuanjun Liu 0001

dblp:115/9425-1 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0001-6983-8088ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Forgetting by Pruning: Data Deletion in Join Cardinality Estimation
abstract
Machine unlearning in learned cardinality estimation (CE) systems presents unique challenges due to the complex distributional dependencies in multi-table relational data. Specifically, data deletion, a core component of machine unlearning, faces three critical challenges in learned CE models: attribute-level sensitivity, inter-table propagation and domain disappearance leading to severe overestimation in multi-way joins. We propose Cardinality Estimation Pruning (CEP), the first unlearning framework specifically designed for multi-table learned CE systems. CEP introduces Distribution Sensitivity Pruning, which constructs semi-join deletion results and computes sensitivity scores to guide parameter pruning, and Domain Pruning, which removes support for value domains entirely eliminated by deletion. We evaluate CEP on state-of-the-art architectures NeuroCard and FACE across IMDb and TPC-H datasets. Results demonstrate CEP consistently achieves the lowest Q-error in multi-table scenarios, particularly under high deletion ratios, often outperforming full retraining. Furthermore, CEP significantly reduces convergence iterations, incurring negligible computational overhead of 0.3%-2.5% of fine-tuning time.
Chaowei He, Yuanjun Liu 0001, Qingzhi Ma, Shenyuan Ren, Xizhao Luo, Lei Zhao 0001, An Liu 0002
AAAI2
2026 ReTRE: Benchmarking LLM Transfer Robustness with Structure-Preserving Variants
abstract
ZhongDong Li, Weijie Shi, Yue Cui, Haolun MA, Yuanjun Liu, Jiawei Li, An Liu, Jia Zhu, Jiajie Xu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
ZhongDong Li, Yue Cui 0001, Haolun Ma, Yuanjun Liu 0001, An Liu 0002, Jia Zhu 0003, Jiajie Xu 0001
ACL (1)5
2025 GPE: Global Position Embedding for Trajectory Similarity Computation
abstract
Trajectory similarity computation is a fundamental functionality in trajectory data mining, with wide-ranging applications in location-based services. Position embedding, which transforms GPS points into embedding vectors, plays a critical role in learning-based trajectory similarity models. The quality of these embeddings significantly impacts the performance of the models on downstream tasks. Existing methods fail to satisfy all good properties, i.e., global, continuous, unique, and dynamic, thereby limiting the development of trajectory similarity computation in both local and global scenarios. Inspired by linear counting systems, such as the decimal system, we first propose the łambda-base circular system to embed positions on the circle, then introduce the multi-base global embedding method GPE to encode global positions into vectors. Experiments conducted on five real-world datasets with nine baseline methods demonstrate that the GPE achieves state-of-the-art performance across four key evaluations in downstream tasks.
Yuanjun Liu 0001, Guanfeng Liu 0001, Qingzhi Ma, Zhixu Li, Lei Zhao 0001, An Liu 0002
KDD (2)1
2024 KMCT: k-Means Clustering of Trajectories Efficiently in Location-Based Services
abstract
With the widespread use of GPS devices and the advancement of location-based services, a vast amount of trajectory data has been collected and mined for various applications. Trajectory clustering, which categorizes trajectories into distinct groups, is the fundamental functionality of trajectory data mining. The challenge is how to cluster on a mass of trajectory data efficiently and universally with satisfying results. The raw trajectory clustering algorithms are universal, but trapped in the dilemma between efficiency and desirable results. Other approaches, such as density-based, road network-based, and deep learning-based algorithms, encounter issues like high time complexity, loss of trajectory integrity, reliance on road networks, and data quality during training. To tackle these challenges, we first propose the efficient KMCT (k-Means Clustering of Trajectories) algorithm based on a semantic interpolation transformation to cluster raw trajectories and achieve satisfying results. Additionally, we introduce the DA-KMCT (Density Accelerated k-Means Clustering of Trajectories) algorithm to further boost the clustering process based on trajectory densities and an optimized centroid selecting strategy. Moreover, we present a novel clustering evaluation method called IOD, which efficiently estimates clustering results on large-scale datasets with linear time complexity. Experimental results on real-world datasets demonstrate that KMCT and DA-KMCT outperform five related methods in terms of clustering quality and time efficiency, and the proposed IOD evaluation shows a strong correlation with the Silhouette Coefficient, offering a reliable and efficient alternative for evaluating clustering results.
Yuanjun Liu 0001, Guanfeng Liu 0001, Qingzhi Ma, Zhixu Li, Shiting Wen, Lei Zhao 0001, An Liu 0002
CIKM1
2024 CLR2G: Cross modal Contrastive Learning on Radiology Report Generation
abstract
The automatic generation of radiological imaging reports aims to produce accurate and coherent clinical descriptions based on X-ray images. This facilitates clinicians in completing the arduous task of report writing and advances clinical automation. The primary challenge in radiological imaging report generation lies in accurately capturing and describing abnormal regions in the images under data bias conditions, resulting in the generation of lengthy texts containing image details. Existing methods mostly rely on prior knowledge such as medical knowledge graphs, corpora, and image databases to assist models in generating more precise textual descriptions. However, these methods still struggle to identify rare anomalies in the images. To address this issue, we propose a two-stage training model, named CLR2G, based on cross-modal contrastive learning. This model delegates the task of capturing anomalies, particularly those challenging for the generative model trained with cross-entropy loss under data bias conditions, to a specialized abnormality capture component. Specifically, we employ a semantic matching loss function to train additional abnormal image and text encoders through cross-modal contrastive learning, facilitating the capture of 13 common anomalies. We utilize the anomalous image features, text features and their confidence probabilities as a posteriori knowledge to help the model generate accurate image reports. Experimental results demonstrate the state-of-the-art performance of our method on two widely used public datasets, IU-Xray and MIMIC-CXR.
Hongchen Xue, Qingzhi Ma, Guanfeng Liu 0001, Jianfeng Qu, Yuanjun Liu 0001, An Liu 0002
CIKM5
2024 Beyond SweepLine: Efficient MaxRS Queries over Inaccurate Location Data
Yuanjun Liu 0001, Zhengcao Zhang, Jianfeng Qu, Guanfeng Liu 0001, An Liu 0002
DASFAA (1)1
2024 Improving Aspect-Based Sentiment Analysis via Tuple-Order Learning
abstract
In the field of Natural Language Processing (NLP), Aspect-Based Sentiment Analysis (ABSA) has gained significant attention in recent years due to its ability to perform fine-grained sentiment analysis. Generative methods tackle various ABSA tasks by autoregressively generating the target sequence of sentiment tuples in a specified format. However, the sentiment tuple is intrinsically an unordered set, and the method introduces an order bias between the generated sequence and the original target. Therefore, to investigate the impact of sentiment tuples order on model performance, we conduct a pilot experiment, unveiling that the order of tuples significantly influences the learning outcomes of the Seq2Seq model. Thus, we propose a novel tuple-order learning method that prioritizes tuples from simple to complex, facilitated by a discrete evaluation method that assesses the difficulty of each individual tuple. Specifically, we incorporate positional information on tuples and employ an effective strategy to expedite the assessment of individual tuples. The method optimizes the learning process while maintaining the structural integrity of existing generative models. Extensive experiments show that our approach significantly advances the performance on 14 datasets of 5 benchmark tasks. We will release our code at https://github.com/gongzhenhu/TOL.
Gongzhen Hu, Yuanjun Liu 0001, Xiabing Zhou, Min Zhang 0005
ECAI2
2023 Towards Effective Trajectory Similarity Measure in Linear Time
Yuanjun Liu 0001, An Liu 0002, Guanfeng Liu 0001, Zhixu Li, Lei Zhao 0001
DASFAA (1)1
2021 A novel deep recommend model based on rating matrix and item attributes
Yuanjun Liu 0001, Tao Wang 0084, Liangmin Guo, Xiaoyao Zheng, Yonglong Luo
J. Intell. Inf. Syst.3