VLDB 2026 Research / reviewers in the wild / expert
Tian Lu 0002
dblp:33/8478-2
· DBLP profile ↗
7ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0003-3730-1897ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Graph Data Collaboration in a Data-Sharing-Free Manner: A Novel Privacy-Preserving Graph Pretraining ModelabstractGraph data, prevalent in various domains such as telecommunication, supply chain, and social networks, holds significant potential for business, operations, and social administration. Collaborating on graph data across institutions or users can further unleash its value, making it a highly sought-after practice. However, such collaboration poses risks to information privacy and commercial confidentiality. In response, we introduce an innovative new model-sharing strategy for graph data collaboration. Here, a data owner pretrains a graph neural network (GNN) model on their private graph data and then provides model users with query access to this model. The pretrained GNN acts as an intermediary, encapsulating knowledge from the private data without exposing it directly. Two fundamental principles are essential for such a pretrained GNN model: model generalizability and privacy preservation. However, current efforts often fail to achieve both concurrently. To tackle this challenge and promote an open yet secure graph data collaboration framework, we propose a novel privacy-preserving operator. This operator integrates smoothly with graph data augmentation and graph contrastive learning, allowing the pretraining of a GNN that effectively eliminates private links at high risk of exposure while maintaining generalizability. Additionally, to improve model generalizability, we introduce a new method called generalizability learning to enhance the model’s adaptability when deployed on unseen data of model user. This approach is designed to simulate diverse environments and develop representations that remain invariant across these varied environments. Extensive experiments suggest that our model surpasses existing state-of-the-art approaches in striking an effective balance between privacy preservation and generalizability. History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning. Funding: This work was supported (to J. Xu) by the National Natural Science Foundation of China [Grants 62206056, 72271059, and 72442011] and the CIPSC-SMP-Zhipu Large Model Cross-Disciplinary Fund. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0115 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0115 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Jiarong Xu, Jiaan Wang, Zenan Zhou, Tian Lu 0002 |
INFORMS J. Comput. | 4 |
| 2025 | Social media meets FinTech platforms: How do online emotions support credit risk decision-making?
Zenan Zhou, Zhichen Chen, Yingjie Zhang 0003, Tian Lu 0002, Xianghua Lu |
Decis. Support Syst. | 4 |
| 2025 | CauseRuDi: Explaining Behavior Sequence Models by Causal Statistics Generation and Rule DistillationabstractRisk scoring systems have been widely deployed in many applications, which assign risk scores to users according to their behavior sequences. Though many deep learning methods with sophisticated designs have achieved promising results, the black-box nature hinders their applications due to fairness, explainability, and compliance consideration. Rule-based systems are considered reliable in these sensitive scenarios. However, building a rule system is labor-intensive. Experts need to find informative statistics from user behavior sequences, design rules based on statistics and assign weights to each rule. In this paper, we bridge the gap between effective but black-box models and transparent rule models. We propose a two-stage framework, CauseRuDi, that distills the knowledge of black-box teacher models into rule-based student models. We design a Monte Carlo tree search-based statistics generation method that maximizes the correlation or dependence between the generated statistics and the teacher model's outputs. We formulate a sequential move game and a simultaneous move coalitional game to generate multiple statistics. Then statistics are composed into logical rules with our proposed neural logical networks by mimicking the outputs of teacher models. We evaluate CauseRuDi on three real-world public datasets and an industrial dataset to demonstrate its effectiveness. Yao Zhang 0009, Yun Xiong, Yiheng Sun, Tian Lu 0002, Shengli Sun |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | TransBoost: A Boosting-Tree Kernel Transfer Learning Algorithm for Improving Financial InclusionabstractThe prosperity of mobile and financial technologies has bred and expanded various kinds of financial products to a broader scope of people, which contributes to financial inclusion. It brings non-trivial social benefits of diminishing financial inequality. However, the technical challenges in individual financial risk evaluation exacerbated by the unforeseen user characteristic distribution and limited credit history of new users, as well as the inexperience of newly-entered companies in handling complex data and obtaining accurate labels, impede further promotion of financial inclusion. To tackle these challenges, this paper develops a novel transfer learning algorithm (i.e., TransBoost) that combines the merits of tree-based models and kernel methods. The TransBoost is designed with a parallel tree structure and efficient weights updating mechanism with theoretical guarantee, which enables it to excel in tackling real-world data with high dimensional features and sparsity in O(n) time complexity. We conduct extensive experiments on two public datasets and a unique largescale dataset from Tencent Mobile Payment. The results show that the TransBoost outperforms other state-of-the- art benchmark transfer learning algorithms in terms of prediction accuracy with superior efficiency, demonstrate stronger robustness to data sparsity, and provide meaningful model interpretation. Besides, given a financial risk level, the TransBoost enables financial service providers to serve the largest number of users including those who would otherwise be excluded by other algorithms. That is, the TransBoost improves financial inclusion. Yiheng Sun, Tian Lu 0002, Cong Wang 0043, Huaiyu Fu, Jingran Dong, Yunjie Calvin Xu |
AAAI | 2 |
| 2022 | RuDi: Explaining Behavior Sequence Models by Automatic Statistics Generation and Rule DistillationabstractRisk scoring systems have been widely deployed in many applications, which assign risk scores to users according to their behavior sequences. Though many deep learning methods with sophisticated designs have achieved promising results, the black-box nature hinders their applications due to fairness, explainability, and compliance consideration. Rule-based systems are considered reliable in these sensitive scenarios. However, building a rule system is labor-intensive. Experts need to find informative statistics from user behavior sequences, design rules based on statistics and assign weights to each rule. In this paper, we bridge the gap between effective but black-box models and transparent rule models. We propose a two-stage method, RuDi, that distills the knowledge of black-box teacher models into rule-based student models. We design a Monte Carlo tree search-based statistics generation method that can provide a set of informative statistics in the first stage. Then statistics are composed into logical rules with our proposed neural logical networks by mimicking the outputs of teacher models. We evaluate RuDi on three real-world public datasets and an industrial dataset to demonstrate its effectiveness. Yao Zhang 0009, Yun Xiong, Yiheng Sun, Tian Lu 0002, Yangyong Zhu |
CIKM | 5 |
| 2022 | Examining the spillover effect of sustainable consumption on microloan repayment: A big data-based research
Yuanqiang Ye, Xianghua Lu, Tian Lu 0002 |
Inf. Manag. | 3 |
| 2018 | Internet usage and patient's trust in physician during diagnoses: A knowledge power perspectiveabstractDoes patients’ Internet search of disease information affect their trust in physicians during diagnosis? This study proposes a research model from a knowledge power perspective, that is, Internet search affects patients’ perception of their knowledge level. Our empirical study of more than 400 subjects suggests that for patients who searched online for disease information, the inconsistency between their self‐diagnosis expectations and their physician's diagnosis reduces their trust in their physician. The effect is stronger for those who spent more time on Internet search. Patients with chronic conditions are less affected by the inconsistency, as are patients of physicians with a higher professional status. This study also found that physicians’ interaction quality in the diagnosis process—how well they communicate with their patient—still plays a dominant role in gaining patient's trust. This finding suggests that even in the high‐tech age, high‐touch remains an important factor to physician‐patient trust. Tian Lu 0002, Yunjie Calvin Xu, Scott Wallace 0004 |
J. Assoc. Inf. Sci. Technol. | 1 |