VLDB 2026 Research / reviewers in the wild / expert
Tianlong Xu
dblp:337/6008
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Security and privacy · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math
Dingjie Song, Tianlong Xu, Yifan Zhang 0004, Hang Li 0007, Zhiling Yan, Haoyang Li 0018, Lichao Sun 0001, Qingsong Wen |
AIED (1) | 2 |
| 2026 | Breaking the Boundary Barrier: Robust Model Fingerprinting via Unlearnable Examples in Model-Parameter SpaceabstractDeep learning models represent valuable intellectual property due to their high development costs. To protect model ownership, existing fingerprinting techniques have been proposed to use adversarial examples to fingerprint a model's decision boundaries. However, these fingerprints are inherently fragile, as model decision boundaries are highly sensitive to common model modifications such as fine-tuning, pruning, and adversarial training. In this paper, we propose MFUE (Model Fingerprinting via Unlearnable Examples), a novel fingerprinting methodology that leverages the stable unlearnability of unlearnable examples to fingerprint arbitrary modified models in parameter space, fundamentally circumventing the inherent vulnerability of decision boundaries. To achieve robust model fingerprinting in parameter space, we are the first to identify that unlearnable examples, owing to their persistent training resistance, can serve as stable fingerprints beyond the model's decision boundaries. To endow unlearnable examples with robustness against arbitrary model modifications, we introduce adversarial training that simulates the randomness of model modifications by jointly optimizing the unlearnable examples over models at different training stages. We evaluate the performance of MFUE against six different attack types, including both model and input tampering. Through extensive experiments, we demonstrate that MFUE outperforms four existing methods in terms of robustness and uniqueness. Tianlong Xu, Zixiong Wang, Gaoyang Liu, Jian Chen 0046, Ahmed M. Abdelmoniem, Chen Wang 0011 |
KDD (1) | 1 |
| 2025 | Knowledge Tagging with Large Language Model Based Multi-Agent SystemabstractKnowledge tagging for questions is vital in modern intelligent educational applications, including learning progress diagnosis, practice question recommendations, and course content organization. Traditionally, these annotations have been performed by pedagogical experts, as the task demands not only a deep semantic understanding of question stems and knowledge definitions but also a strong ability to link problem-solving logic with relevant knowledge concepts. With the advent of advanced natural language processing (NLP) algorithms, such as pre-trained language models and large language models (LLMs), pioneering studies have explored automating the knowledge tagging process using various machine learning models. In this paper, we investigate the use of a multi-agent system to address the limitations of previous algorithms, particularly in handling complex cases involving intricate knowledge definitions and strict numerical constraints. By demonstrating its superior performance on the publicly available math question knowledge tagging dataset, MathKnowCT, we highlight the significant potential of an LLM-based multi-agent system in overcoming the challenges that previous methods have encountered. Finally, through an in-depth discussion of the implications of automating knowledge tagging, we underscore the promising future of deploying LLM-based algorithms in educational contexts. Hang Li 0007, Tianlong Xu, Ethan Chang, Qingsong Wen |
AAAI | 2 |
| 2025 | AI-Driven Virtual Teacher for Enhanced Educational Efficiency: Leveraging Large Pretrain Models for Autonomous Error Analysis and CorrectionabstractStudents frequently make mistakes while solving mathematical problems, and traditional error correction methods are both time-consuming and labor-intensive. This paper introduces an innovative Virtual AI Teacher system designed to autonomously analyze and correct student Errors (VATE). Leveraging advanced large language models (LLMs) like GPT-4, the system uses student drafts as a primary source for error analysis, which enhances understanding of the student's learning process. It incorporates sophisticated prompt engineering and maintains an error pool to reduce computational overhead. The AI-driven system also features a real-time dialogue component for efficient student interaction. Our approach demonstrates significant advantages over traditional and machine learning-based error correction methods, including reduced educational costs, high scalability, and superior generalizability. The system has been deployed in Squirrel AI's learning platform for elementary mathematics education, where it achieves 78.3% accuracy in error analysis and shows a marked improvement in student learning efficiency. Satisfaction surveys indicate a strong positive reception, highlighting the system's potential to transform educational practices. Tianlong Xu, Yifan Zhang 0004, Zhendong Chu, Shen Wang 0005, Qingsong Wen |
AAAI | 1 |
| 2025 | Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem SolutionsabstractHang Li, Tianlong Xu, Kaiqi Yang, Yucheng Chu, Yanling Chen, Yichi Song, Qingsong Wen, Hui Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Hang Li 0007, Tianlong Xu, Kaiqi Yang 0001, Yucheng Chu, Yichi Song, Qingsong Wen, Hui Liu 0031 |
ACL (1) | 2 |
| 2025 | Poisoning as a Post-Protection: Mitigating Membership Privacy Leakage From Gradient and Prediction of Federated ModelsabstractFederated learning (FL) is a distributed learning paradigm that enables multiple clients to train a unified model without sharing their private data. However, recent works demonstrate that FL models are vulnerable to membership inference attacks (MIAs), which can infer whether a data sample was used to train a given FL model. Existing countermeasures either require far-reaching modifications of FL training process or enforce extra processing in prediction phase, yielding them unlikely to be applied well in practice. In this paper, we design a post-protection mechanism, dubbedP$^{2}$-Protection, which degrades the inference performance of MIAs by simultaneously poisoning the prediction and gradient of the target FL model to reduce the privacy leakage of training data while keeping the model prediction accuracy.P$^{2}$-Protectiononly involves one additional training round to embed the poisoned prediction and gradient into the target FL model, without requiring model retraining or training process modification. We evaluateP$^{2}$-Protectionand compare it with two state-of-the-art defenses against three MIAs on five realistic datasets. Experimental results show thatP$^{2}$-Protectionoutperforms the existing defenses by offering limited implement overhead and improved utility-privacy trade-off. Gaoyang Liu, Tianlong Xu, Yang Yang 0060, Ahmed M. Abdelmoniem, Chen Wang 0011, Jiangchuan Liu |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | United We Stand, Divided We Fall: Fingerprinting Deep Neural Networks via Adversarial TrajectoriesabstractIn recent years, deep neural networks (DNNs) have witnessed extensive applications, and protecting their intellectual property (IP) is thus crucial. As a non-invasive way for model IP protection, model fingerprinting has become popular. However, existing single-point based fingerprinting methods are highly sensitive to the changes in the decision boundary, and may suffer from the misjudgment of the resemblance of sparse fingerprinting, yielding high false positives of innocent models. In this paper, we propose ADV-TRA, a more robust fingerprinting scheme that utilizes adversarial trajectories to verify the ownership of DNN models. Benefited from the intrinsic progressively adversarial level, the trajectory is capable of tolerating greater degree of alteration in decision boundaries. We further design novel schemes to generate a surface trajectory that involves a series of fixed-length trajectories with dynamically adjusted step sizes. Such a design enables a more unique and reliable fingerprinting with relatively low querying costs. Experiments on three datasets against four types of removal attacks show that ADV-TRA exhibits superior performance in distinguishing between infringing and innocent models, outperforming the state-of-the-art comparisons. Tianlong Xu, Chen Wang 0011, Gaoyang Liu, Yang Yang 0060, Kai Peng 0001, Wei Liu 0004 |
NeurIPS | 1 |
| 2024 | Gradient-Leaks: Enabling Black-Box Membership Inference Attacks Against Machine Learning ModelsabstractMachine Learning (ML) techniques have been applied to many real-world applications to perform a wide range of tasks. In practice, ML models are typically deployed as the black-box APIs to protect the model owner’s benefits and/or defend against various privacy attacks. In this paper, we present Gradient-Leaks as the first evidence showcasing the possibility of performing membership inference attacks (MIAs), with mere black-box access, which aim to determine whether a data record was utilized to train a given target ML model or not. The key idea of Gradient-Leaks is to construct a local ML model around the given record which locally approximates the target model’s prediction behavior. By extracting the membership information of the given record from the gradient of the substituted local model using an intentionally modified autoencoder, Gradient-Leaks can thus breach the membership privacy of the target model’s training data in an unsupervised manner, without any priori knowledge about the target model’s internals or its training data. Extensive experiments on different types of ML models with real-world datasets have shown that Gradient-Leaks can achieve a better performance compared with state-of-the-art attacks. Gaoyang Liu, Tianlong Xu, Rui Zhang 0066, Zixiong Wang, Chen Wang 0011, Ling Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Your Model Trains on My Data? Protecting Intellectual Property of Training Data via Membership Fingerprint AuthenticationabstractIn recent years, data has become the new oil that fuels various machine learning (ML) applications. Just as the oil refining, providing data to an ML model is a product of massive costs and expertise efforts. However, how to protect the intellectual property (IP) of the training data in ML remains largely open. In this paper, we present MeFA, a novel framework for detecting training data IP embezzlement via Membership Fingerprint Authentication, which is able to determine whether a suspect ML model is trained on the to be protected target data or not. The key observation is that a part of data has a similar influence on the prediction behavior of different ML models. On this basis, MeFA leverages membership inference techniques to extract these data as the fingerprints of the target data and constructs an authentication model to verify the data’s ownership by identifying the obtained membership fingerprints. MeFA has several salient features. It does not assume any knowledge of the suspect model except for its black-box prediction API, through which we can merely get the prediction output of a given input, and also does not require any modification to the dataset or the training process, since it takes advantage of the inherent membership property of the data. As a by-product, MeFA can also serve as a post-protection to verify the ownership of ML models, without modifying the training process of the model. Extensive experiments on three realistic datasets and seven types of ML models validate the effectiveness of MeFA, and demonstrate that it is also robust to scenarios when the training data is partially used or preprocessed with representative membership inference defenses. Gaoyang Liu, Tianlong Xu, Xiaoqiang Ma, Chen Wang 0011 |
IEEE Trans. Inf. Forensics Secur. | 2 |