EDBT 2026 Demo / reviewers in the wild / expert
Jianbin Lin
dblp:90/10769
· DBLP profile ↗
11ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0004-1269-6620ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the Gap: An End-to-End Framework for Decoupled Alignment in Dynamic Semantic ID GenerationabstractSemantic IDs derived from Multi-modal Large Language Models (MLLMs) integrate rich content semantics into recommendation systems but often lack the collaborative signals crucial for capturing user behavior. Additionally, static generation fails to adapt to evolving data distributions, causing codebook drift. To address these limitations, we propose a dynamic End-to-End Semantic ID Generation Framework based on Decoupled Representation Alignment. Our method aligns shared and private components from both MLLM content and Collaborative Filtering (CF) embeddings, integrating them via a hierarchical adaptive fusion module into a Residual Quantized Variational Autoencoder (RQ-VAE). This joint optimization promotes both semantic granularity and collaborative awareness while supporting dynamic codebook updates. Extensive offline experiments and online A/B tests on Alipay's Tab3 short-video scenario demonstrate the superiority of our method, which is now deployed to serve all users. Yu Cheng 0031, Jianbin Lin, Can Ye |
SIGIR | 4 |
| 2026 | Generative Enhanced Modeling: A Collaborative Framework for Enhancing User Representations via Semantic IDabstractUser interest modeling is foundational to recommender systems. However, sparse and noisy behaviors make traditional item-level sequence models brittle, especially for new and low-activity users. Furthermore, relying solely on a user's own history limits exploration and reinforces the ''filter bubbles''. To address this, we propose GEM (Generative Enhanced Modeling). GEM shifts the paradigm from self-behavior induction to collective experience migration. Specifically, it constructs LLM-based semantic IDs and embeddings. Grounded in information theory, GEM performs multi-stage denoising at both the user and item levels. This design effectively suppresses reward-driven noise while preserving target-aware signals. We deployed GEM on the Alipay Tab3 video feed. Offline evaluations show significant GAUC gains. Online A/B tests demonstrate a 0.9% lift in watch time alongside stable video views and improved exposure diversity. These results confirm that GEM enhances recommendation quality and successfully broadens user interests. Yu Cheng 0031, Jianbin Lin, Can Ye |
SIGIR | 4 |
| 2026 | SCOPE: Scalable Cross-Task Orthogonal Progressive Experts for Multi-Task Learning in Recommendations
Zixian Yang, Zhaokai Huang, Jianbin Lin, Leon Wenliang Zhong, Can Ye |
SIGIR | 6 |
| 2025 | Towards Principled Learning for Re-ranking in Recommender SystemsabstractAs the final stage of recommender systems, re-ranking presents ordered item lists to users that best match their interests. It plays such a critical role and has become a trending research topic with much attention from both academia and industry. Recent advances of re-ranking are focused on attentive listwise modeling of interactions and mutual influences among items to be re-ranked. However, principles to guide the learning process of a re-ranker, and to measure the quality of the output of the re-ranker, have been always missing. In this paper, we study such principles to learn a good re-ranker. Two principles are proposed, including convergence consistency and adversarial consistency. These two principles can be applied in the learning of a generic re-ranker and improve its performance. We validate such a finding by various baseline methods over different datasets. Qunwei Li, Jianbin Lin, Leon Wenliang Zhong |
SIGIR | 3 |
| 2025 | FedBiF: Communication-Efficient Federated Learning via Bits FreezingabstractFederated learning (FL) is an emerging distributed machine learning paradigm that enables collaborative model training without sharing local data. Despite its advantages, FL suffers from substantial communication overhead, which can affect training efficiency. Recent efforts have mitigated this issue by quantizing model updates to reduce communication costs. However, most existing methods apply quantization only after local training, introducing quantization errors into the trained parameters and potentially degrading model accuracy. In this paper, we propose Federated Bit Freezing (FedBiF), a novel FL framework that directly learns quantized model parameters during local training. In each communication round, the server first quantizes the model parameters and transmits them to the clients. FedBiF then allows each client to update only a single bit of the multi-bit parameter representation, freezing the remaining bits. This bit-by-bit update strategy reduces each parameter update to one bit while maintaining high precision in parameter representation. Extensive experiments are conducted on five widely used datasets under both IID and Non-IID settings. The results demonstrate that FedBiF not only achieves superior communication compression but also promotes sparsity in the resulting models. Notably, FedBiF attains accuracy comparable to FedAvg, even when using only 1 bit-per-parameter (bpp) for uplink and 3 bpp for downlink communication. The code is available athttps://github.com/Leopold1423/fedbif-tpds25. Shiwei Li 0002, Qunwei Li, Haozhao Wang, Ruixuan Li 0001, Jianbin Lin, Leon Wenliang Zhong |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | Generating Natural Language Adversarial Examples on a Large Scale with Generative ModelsabstractToday text classification models have been widely used. However, these classifiers are found to be easily fooled by adversarial examples. Fortunately, standard attacking methods generate adversarial texts in a pair-wise way, that is, an adversarial text can only be created from a real-world text by replacing a few words. In many applications, these texts are limited in numbers, therefore their corresponding adversarial examples are often not diverse enough and sometimes hard to read, thus can be easily detected by humans and cannot create chaos at a large scale. In this paper, we propose an end to end solution to efficiently generate adversarial texts from scratch using generative models, which are not restricted to perturbing the given texts. We call it unrestricted adversarial text generation. Specifically, we train a conditional variational autoencoder (VAE) with an additional adversarial loss to guide the generation of adversarial examples. Moreover, to improve the validity of adversarial texts, we utilize discrimators and the training framework of generative adversarial networks (GANs) to make adversarial texts consistent with real data. Experimental results on sentiment analysis demonstrate the scalability and efficiency of our method. It can attack text classification models with a higher success rate than existing methods, and provide acceptable quality for humans in the meantime. Yankun Ren, Jianbin Lin, Siliang Tang, Jun Zhou 0011, Yuan Qi 0001, Xiang Ren 0001 |
ECAI | 2 |
| 2019 | InfDetect: a Large Scale Graph-based Fraud Detection System for E-Commerce InsuranceabstractThe insurance industry has been creating innovative products around the emerging online shopping activities. Such ecommerce insurance is designed to protect buyers from potential risks such as impulse purchases and counterfeits. Fraudulent claims towards online insurance typically involve multiple parties such as buyers, sellers, and express companies, and they could lead to heavy financial losses. In order to uncover the relations behind organized fraudsters and detect fraudulent claims, we developed a large-scale insurance fraud detection system, i.e., InfDetect, which provides interfaces for commonly used graphs, standard data processing procedures, and a uniform graph learning platform. InfDetect is able to process big graphs containing up to 100 millions of nodes and billions of edges.In this paper, we investigate different graphs to facilitate fraudster mining, such as a device-sharing graph, a transaction graph, a friendship graph, and a buyer-seller graph. These graphs are fed to a uniform graph learning platform containing supervised and unsupervised graph learning algorithms. Cases on widely applied e-commerce insurance are described to demonstrate the usage and capability of our system. InfDetect has successfully detected thousands of fraudulent claims and saved over tens of thousands of dollars daily. Cen Chen 0001, Jianbin Lin, Li Wang 0056, Xinxing Yang, Jun Zhou 0011, Yuan Qi 0001 |
IEEE BigData | 3 |
| 2019 | A Semi-Supervised Graph Attentive Network for Financial Fraud DetectionabstractWith the rapid growth of financial services, fraud detection has been a very important problem to guarantee a healthy environment for both users and providers. Conventional solutions for fraud detection mainly use some rule-based methods or distract some features manually to perform prediction. However, in financial services, users have rich interactions and they themselves always show multifaceted information. These data form a large multiview network, which is not fully exploited by conventional methods. Additionally, among the network, only very few of the users are labelled, which also poses a great challenge for only utilizing labeled data to achieve a satisfied performance on fraud detection. To address the problem, we expand the labeled data through their social relations to get the unlabeled data and propose a semi-supervised attentive graph neural network, named SemiGNN to utilize the multi-view labeled and unlabeled data for fraud detection. Moreover, we propose a hierarchical attention mechanism to better correlate different neighbors and different views. Simultaneously, the attention mechanism can make the model interpretable and tell what are the important factors for the fraud and why the users are predicted as fraud. Experimentally, we conduct the prediction task on the users of Alipay, one of the largest third-party online and offline cashless payment platform serving more than 4 hundreds of million users in China. By utilizing the social relations and the user attributes, our method can achieve a better accuracy compared with the state-of-the-art methods on two tasks. Moreover, the interpretable results also give interesting intuitions regarding the tasks. Daixin Wang, Yuan Qi 0001, Jianbin Lin, Peng Cui 0001, Quanhui Jia, Yanming Fang, Jun Zhou 0011 |
ICDM | 3 |
| 2019 | RNE: A Scalable Network Embedding for Billion-Scale Recommendation
Jianbin Lin, Daixin Wang, Lu Guan, Yin Zhao, Binqiang Zhao, Jun Zhou 0011, Xiaolong Li 0005, Yuan Qi 0001 |
PAKDD (2) | 1 |
| 2018 | NetDP: An Industrial-Scale Distributed Network Representation Framework for Default Prediction in Ant Credit PayabstractAnt Credit Pay is a consumer credit service in Ant Financial Service Group. Similar to credit card, loan default is one of the major risks of this credit product. Hence, effective algorithm for default prediction is the key to losses reduction and profits increment for the company. However, the challenges facing in our scenario are different from those in conventional credit card service. The first one is scalability. The huge volume of users and their behaviors in Ant Financial requires the ability to process industrial-scale data and perform model training efficiently. The second challenges is the cold-start problem. Different from the manual review for credit card application in conventional banks, the credit limit of Ant Credit Pay is automatically offered to users based on the knowledge learned from big data. However, default prediction for new users is suffered from lack of enough credit behaviors. It requires that the proposal should leverage other new data source to alleviate the cold-start problem. Considering the above challenges and the special scenario in Ant Financial, we try to incorporate default prediction with network information to alleviate the cold-start problem. In this paper, we propose an industrial-scale distributed network representation framework, termed NetDP, for default prediction in Ant Credit Pay. The proposal explores network information generated by various interaction between users, and blends unsupervised and supervised network representation in a unified framework for default prediction problem. Moreover, we present a parameter-server-based distributed implement of our proposal to handle the scalability challenge. Experimental results demonstrate the effectiveness of our proposal, especially in cold-start problem, as well as the efficiency for industrial-scale dataset. Jianbin Lin, Zhiqiang Zhang 0012, Jun Zhou 0011, Xiaolong Li 0005, Jingli Fang, Yanming Fang, Yuan Qi 0001 |
IEEE BigData | 1 |
| 2009 | An Anti-attack Watermarking Based on Synonym Substitution for Chinese TextabstractCurrent research of natural language steganographic algorithms based on synonymy substitution mostly focused on invisibility, but ignored robustness. However, automatic disambiguation of Chinese word senses (WSD) achieves high accuracy, the adversary could destroyed the watermark easily if he disambiguated the stego-text and did the synonym substitution again. In this paper, against the high accuracy of WSD algorithm, two indicators are proposed, which are Lexical similarity and sense similarity. For reducing the accuracy rate of WSD, we argued that it should choose the word, which is low Lexical similarity and high senses similarity, in the synonym substitution. Therefore, the automatically attack will failure. The experiment showed that the algorithm reduced the accuracy of WSD from 90.4% to 74.5%. The robustness of watermarking has been improved. Jianbin Lin, Tianzhi Li, Dingyi Fang |
IAS | 2 |