EDBT 2026 Demo / reviewers in the wild / expert
Shouling Ji
dblp:07/8388
· DBLP profile ↗
26ranked-venue papers in the field
1as first author
16since 2021 · last 2026
0000-0003-4268-372XORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 8 (1 first)Data Mining & Knowledge Discovery · 7Information Retrieval & Web Search · 6Other / Interdisciplinary · 3Big Data, Cloud & Distributed Data Systems · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Eminence in Shadow: Exploiting Feature Boundary Ambiguity for Robust Backdoor AttacksabstractDeep neural networks (DNNs) underpin critical applications yet remain vulnerable to backdoor attacks, typically reliant on heuristic brute-force methods. Despite significant empirical advancements in backdoor research, the lack of rigorous theoretical analysis limits understanding of underlying mechanisms, constraining attack predictability and adaptability. Therefore, we provide a theoretical analysis targeting backdoor attacks, focusing on how sparse decision boundaries enable disproportionate model manipulation. Based on this finding, we derive a closed-form ''ambiguous boundary region'' wherein negligible relabeled samples induce substantial misclassification. Influence function analysis further quantifies significant parameter shifts caused by these margin samples, with minimal impact on clean accuracy, formally grounding why such low poison rates suffice for efficacious attacks. Leveraging these insights, we propose Eminence, an explainable and robust black-box backdoor framework with provable theoretical guarantees and inherent stealth properties. Eminence optimizes a universal, visually subtle trigger that strategically exploits vulnerable decision boundaries and effectively achieves robust misclassification with exceptionally low poison rates (≤ 0.01%, compared to SOTA methods typically requiring ≥ 1 %). Comprehensive experiments validate our theoretical discussions and demonstrate the effectiveness of Eminence, confirming an exponential relationship between margin poisoning and adversarial boundary manipulation. Eminence maintains ≥ 90% attack success rate, exhibits negligible clean-accuracy loss, and demonstrates high transferability across diverse models, datasets and scenarios. Our code is available at https://github.com/NESA-Lab/Eminence Zhou Feng, Chunyi Zhou 0001, Yuwen Pu, Tianyu Du, Jianhai Chen, Shouling Ji |
KDD (1) | 8 |
| 2026 | APT-CGLP: Advanced Persistent Threat Hunting via Contrastive Graph-Language Pre-TrainingabstractProvenance-based threat hunting identifies Advanced Persistent Threats (APTs) on endpoints by correlating attack patterns described in Cyber Threat Intelligence (CTI) with provenance graphs derived from system audit logs. A fundamental challenge in this paradigm lies in the modality gap —the structural and semantic disconnect between provenance graphs and CTI reports. Prior work addresses this by framing threat hunting as a graph matching task: 1) extracting attack graphs from CTI reports, and 2) aligning them with provenance graphs. However, this pipeline incurs severe information loss during graph extraction and demands intensive manual curation, undermining scalability and effectiveness. Xuebo Qiu, Mingqi Lv, Yimei Zhang 0003, Tieming Chen, Tiantian Zhu 0001, Qijie Song, Shouling Ji |
KDD (1) | 7 |
| 2026 | FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks
Naen Xu, Jinghuai Zhang, Chunyi Zhou 0001, Jun Wang 0020, Zhihui Fu, Tianyu Du, Zhaoxiang Wang, Shouling Ji |
WWW | 9 |
| 2025 | Privacy-Preserving Triangle Counting in Directed GraphsabstractIn directed graphs, the relationship between users is asymmetric, resulting in two types of triangles: cycle triangles and flow triangles. This paper studies the problem of privacy-preserving triangle counting in directed graphs. Based on different applications, we consider two scenarios, i.e., trusted and untrusted servers. In the literature, privacy-preserving triangle counting in undirected graphs has been widely studied. However, directly applying these algorithms to address our problem suffers from many issues. Concretely, for the trusted server scenario, the differentially private triangle counting algorithms, designed for undirected graphs, exhibit suboptimal performance when applied to directed graphs. Hence, we propose a new centralized differentially private algorithm that adds Laplacian noise to the exact numbers by analyzing global sensitivity. Furthermore, for the untrusted server scenario, the existing techniques cannot be used to count cycle and flow triangles with differential privacy because the local view of each user in directed graphs is limited to out-neighbors rather than all neighbors. Therefore, we design a novel locally differentially private algorithm to provide local unbiased estimation, which implies that after aggregating all the local estimations on the central server side, an unbiased estimation for the numbers of cycle and flow triangles is deduced. Empirical experiments on six real-world graph datasets demonstrate that our proposed algorithms achieve high efficiency and utility. Ziyao Wei, Qing Liu 0008, Zhikun Zhang 0001, Shouling Ji, Yunjun Gao |
ICDE | 4 |
| 2025 | FLMarket: Enabling Privacy-preserved Pre-training Data Pricing for Federated Learning
Zhenyu Wen, Wanglei Feng, Di Wu 0065, Haozhen Hu, Chang Xu 0031, Bin Qian 0002, Zhen Hong, Cong Wang 0006, Shouling Ji |
KDD (1) | 9 |
| 2025 | Enhancing Adversarial Transferability via Self-Ensemble Feature AlignmentabstractDeep neural networks (DNNs) have demonstrated remarkable success in tasks such as image classification and object detection, but remain vulnerable to adversarial attacks. To enhance the adversarial transferability across different architectures (e.g., from CNNs to ViTs), existing attacks leverage various strategies such as input transformations, gradient rectification, custom optimization objectives, and model ensembles, but struggle with limited effectiveness under minimal knowledge (e.g., the number of surrogate models). In this work, we propose a novel self-ensemble feature alignment(SEFA) strategy that significantly boosts adversarial transferability with minimal resource overhead. Motivated by the observation that adversarial transferability correlates with feature similarity across models, we leverage Centered Kernel Alignment (CKA) to measure and investigate intermediate features in both inter- and intra-model representation spaces. By splitting a single model into multiple sub-networks and aligning their feature spaces, our method effectively enhances adversarial transferability without relying on additional surrogate models. Experiments on the ImageNet dataset demonstrate that our approach achieves an average ASR of 80.0% (ResNet-50 surrogate) and 75.9% (Inc-v3 surrogate) on ImageNet, which outperform the second-best method (i.e., BSR) by +8.3% and +4.7% respectively. Further, it can seamlessly integrate with existing attacks to further increase cross architecture transferability. Zhiming Zhao, Qingming Li, Chunyi Zhou 0001, Shouling Ji |
ICMR | 5 |
| 2025 | ArtistAuditor: Auditing Artist Style Pirate in Text-to-Image Generation ModelsabstractText-to-image models based on diffusion processes, such as DALL-E, Stable Diffusion, and Midjourney, are capable of transforming texts into detailed images and have widespread applications in art and design. As such, amateur users can easily imitate professional-level paintings by collecting an artist's work and fine-tuning the model, leading to concerns about artworks' copyright infringement. To tackle these issues, previous studies either add visually imperceptible perturbation to the artwork to change its underlying styles (perturbation-based methods) or embed post-training detectable watermarks in the artwork (watermark-based methods). However, when the artwork or the model has been published online, i.e., modification to the original artwork or model retraining is not feasible, these strategies might not be viable. Linkang Du, Min Chen 0032, Zhou Su 0001, Shouling Ji, Peng Cheng 0001, Jiming Chen 0001, Zhikun Zhang 0001 |
WWW | 5 |
| 2025 | F$^{2}$2AT: Feature-Focusing Adversarial Training via Disentanglement of Natural and Perturbed PatternsabstractDeep neural networks (DNNs) are vulnerable to adversarial examples crafted by well-designed perturbations. This could lead to disastrous results on critical applications such as self-driving cars, surveillance security, and medical diagnosis. At present, adversarial training is one of the most effective defenses against adversarial examples. However, in traditional adversarial training, it is still difficult to achieve a good trade-off between clean accuracy and robustness since DNNs still learn spurious features. The intrinsic reason is that traditional adversarial training makes it difficult to fully learn core features from adversarial examples when noise and examples cannot be disentangled. In this paper, we disentangle the adversarial examples into natural and perturbed patterns by bit-plane slicing. We assume the higher bit-planes represent natural patterns and the lower bit-planes represent perturbed patterns, respectively. We propose Feature-Focusing Adversarial Training (F$^{2}$AT), which differs from previous work in that it enforces the model to focus on the core features from natural patterns and reduce the impact of spurious features from perturbed patterns. The experimental results demonstrated that the clean accuracy and adversarial robustness with our F$^{2}$AT can be significantly improved. Yaguan Qian, Zhaoquan Gu, Bin Wang 0062, Shouling Ji, Wei Wang 0012, Yanchun Zhang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | GIFT: Graph-guIded Feature Transfer for Cold-Start Video Click-Through Rate PredictionabstractShort video has witnessed rapid growth in the past few years in e-commerce platforms like Taobao. To ensure the freshness of the content, platforms need to release a large number of new videos every day, making conventional click-through rate (CTR) prediction methods suffer from the item cold-start problem. In this paper, we propose GIFT, an efficient Graph-guIded Feature Transfer system, to fully take advantages of the rich information of warmed-up videos to compensate for the cold-start ones. Specifically, we establish a heterogeneous graph that contains physical and semantic linkages to guide the feature transfer process from warmed-up video to cold-start videos.Specifically, we establish a heterogeneous graph that contains physical and semantic linkages to guide the feature transfer process. The physical linkages consist of the explicit relationships (e.g., produced by the same author, or showcasing the same product etc.), and the semantic linkages measure the proximity of multi-modal representations of two videos. We elaborately design the feature transfer function to make aware of different parts of transferred features (e.g., id representations and historical statistics) from different types of nodes and edges along the metapath on the graph. We conduct extensive experiments on a large real-world dataset, and the results show that our GIFT system outperforms SOTA methods significantly and brings a 6.82% lift on CTR in the homepage of Taobao App. Sihao Hu, Zhao Li 0007, Yazheng Yang, Qingwen Liu 0002, Shouling Ji |
CIKM | 7 |
| 2022 | DetectS ec: Evaluating the robustness of object detection models to adversarial attacksabstractDespite their tremendous success in various machine learning tasks, deep neural networks (DNNs) are inherently vulnerable to adversarial examples, which are maliciously crafted inputs to cause DNNs to misbehave. Intensive research has been conducted on this phenomenon in simple tasks (e.g., image classification). However, little is known about this adversarial vulnerability for object detection, a much more complicated task, which often requires specialized DNNs and multiple additional components. In this paper, we present DetectSec, a uniform platform for robustness analysis of object detection models. Currently, DetectSec implements 13 representative adversarial attacks with 7 utility metrics and 13 defenses on 18 standard object detection models. Leveraging DetectSec, we conduct the first rigorous evaluation of adversarial attacks on the state-of-the-art object detection models. We analyze the impact of the factors including DNN architecture and capacity on the model robustness. We show that many conclusions about adversarial attacks and defenses in image classification tasks do not transfer to object detection tasks, for example, the targeted attack is stronger than the untargeted attack for two-stage detectors. Our findings will aid future efforts in understanding and defending against adversarial attacks in complicated tasks. In addition, we compare the robustness of different detection models and discuss their relative strengths and weaknesses. The platform DetectSec will be open source as a unique facility for further research on adversarial attacks and defenses in object detection tasks. Tianyu Du, Shouling Ji, Bo Li 0026, Tao Wei 0002, Yunhan Jia, Raheem A. Beyah, Ting Wang 0006 |
Int. J. Intell. Syst. | 2 |
| 2022 | An interpretable outcome prediction model based on electronic health records and hierarchical attentionabstractOutcome prediction aims to predict the future health condition of patients from Electronic Health Record (EHR) data. Because of the sequential characteristic of EHR data, recurrent neural network (RNN)-based outcome prediction methods have achieved state-of-the-art results. However, the major drawback of RNN-based outcome prediction methods is lack of interpretability, which would lead to trust issues. Aiming at this problem, this paper proposes interpretable outcome prediction model with hierarchical attention (IoHAN), an interpretable outcome prediction model by leveraging attention mechanism. The main novelty of IoHAN is that it can pinpoint the fine-grained influence on the final prediction result of each medical component by decomposing the attention weights hierarchically into hospital visits, medical variables, and interactions between medical variables. We evaluated IoHAN on MIMIC-III, a large real-world EHR data set. The experiment results demonstrate that IoHAN can achieve higher prediction accuracy than state-of-the-art outcome prediction models. In addition, the hierarchical decomposed attention weights can interpret the prediction results in a more natural and understandable way. Dajian Zeng, Zhao Li 0007, Mingqi Lv, Ling Chen 0001, Shouling Ji |
Int. J. Intell. Syst. | 8 |
| 2022 | Focus : Function clone identification on cross-platformabstractAutomatic identification of function clones on cross-platform aims at determining whether two functions are identical or not without access to the source code, which is a fundamental challenge in vulnerability search, code plagiarism detection, and malware classification. With the rapid development of deep neural network in program analysis, the state-of-the-art neural network-based function clone identification methods propose to represent functions as embeddings by graph neural network (GNN). However, such a novel representation of functions brings in two challenges. (1) The feature engineering that accurately maps the raw data of binary code to machine learning features is complicated. (2) A highly accurate embedding of functions requires a customized GNN to focus on the most critical features to identify binary code. To the best of our knowledge, currently, a comprehensive work that can overcome the above challenges is still missing. In this paper, we propose a novel prototype named as Focus, which is designed to accurately and efficiently identify similar functions. Specifically, inspired by natural language processing techniques which effectively learns text semantic across natural languages, Focus can learn representative semantic features of functions by a customized learning model. To address the second challenge, a multi-head attention mechanism can be employed to capture the critical features of a function. Through extensive experiments, we demonstrate that Focus achieves high accuracy of function clone identification on a broad range of eight architectures. In particular, the identification performance (AUC value) of Focus is 97% and 99% for cross-platform and single-platform, respectively. Furthermore, the evaluation in real world applications shows that our Focus identifies 24 vulnerable functions among the top-30 candidates, which is one time higher than the baseline approaches. Lirong Fu, Shouling Ji, Changchang Liu, Peiyu Liu 0003, Fuzheng Duan, Zonghui Wang, Whenzhi Chen, Ting Wang 0006 |
Int. J. Intell. Syst. | 2 |
| 2022 | Exploiting Heterogeneous Graph Neural Networks with Latent Worker/Task Correlation Information for Label Aggregation in CrowdsourcingabstractCrowdsourcing has attracted much attention for its convenience to collect labels from non-expert workers instead of experts. However, due to the high level of noise from the non-experts, a label aggregation model that infers the true label from noisy crowdsourced labels is required. In this article, we propose a novel framework based on graph neural networks for aggregating crowd labels. We construct a heterogeneous graph between workers and tasks and derive a new graph neural network to learn the representations of nodes and the true labels. Besides, we exploit the unknown latent interaction between the same type of nodes (workers or tasks) by adding a homogeneous attention layer in the graph neural networks. Experimental results on 13 real-world datasets show superior performance over state-of-the-art models. Hanlu Wu, Tengfei Ma 0001, Lingfei Wu 0001, Fangli Xu, Shouling Ji |
ACM Trans. Knowl. Discov. Data | 5 |
| 2021 | Turbo: Fraud Detection in Deposit-free Leasing Service via Real-Time Behavior Network MiningabstractOnline deposit-free leasing service has witnessed rapid growth in China and shows a promising market in the future. While eliminating the requirement of a deposit does attract more users to the service, it also lowers the cost for fraudsters. Since the emergence of this service is relatively new, there are few works in literature focusing on detecting fraud transactions in it. Existing efforts mainly fall into hard-coded solutions such as block-listing or scorecard methods, which can be impotent in the face of the diverse fraud tactics, e.g., identity theft, or even suffering concept drift problem as the tactics evolve. In this paper, we contribute Turbo, an efficient graph-based anti-fraud system, to fully exploit the abundant user behavior logs in a real-time manner. Turbo is able to additionally make use of the implicit user relationships beyond the user features in the logs. To capture the user relationships, we first propose a novel algorithm to construct a time-evolving user behavior network called BN. Empirical analysis demonstrates that fraudsters in BN exhibit unique temporal aggregation and homophilic patterns, which inspires us to develop a novel heterogeneous adaptive graph neural network algorithm called HAG. Specifically, in HAG two graph operators are presented to mitigate the over-smoothing problem and make better use of the heterogeneous behavior relations in BN. Extensive experiments on a real-world dataset show that our method outperforms state-of-the-art methods significantly and can give a response in seconds for each detection request. Sihao Hu, Xuhong Zhang 0002, Junfeng Zhou, Shouling Ji, Zhao Li 0007, Qinming He, Liming Fang 0001 |
ICDE | 4 |
| 2021 | ACT-Detector: Adaptive channel transformation-based light-weighted detector for adversarial attacks
Jinyin Chen, Haibin Zheng, Wenchang Shangguan, Liangying Liu, Shouling Ji |
Inf. Sci. | 5 |
| 2021 | Deep Graph Matching and Searching for Semantic Code RetrievalabstractCode retrieval is to find the code snippet from a large corpus of source code repositories that highly matches the query of natural language description. Recent work mainly uses natural language processing techniques to process both query texts (i.e., human natural language) and code snippets (i.e., machine programming language), however, neglecting the deep structured features of query texts and source codes, both of which contain rich semantic information. In this article, we propose an end-to-end deep graph matching and searching (DGMS) model based on graph neural networks for the task of semantic code retrieval. To this end, we first represent both natural language query texts and programming language code snippets with the unified graph-structured data, and then use the proposed graph matching and searching model to retrieve the best matching code snippet. In particular, DGMS not only captures more structural information for individual query texts or code snippets, but also learns the fine-grained similarity between them by cross-attention based semantic matching operations. We evaluate the proposed DGMS model on two public code retrieval datasets with two representative programming languages (i.e., Java and Python). Experiment results demonstrate that DGMS significantly outperforms state-of-the-art baseline models by a large margin on both datasets. Moreover, our extensive ablation studies systematically investigate and illustrate the impact of each part of DGMS. Xiang Ling 0001, Lingfei Wu 0001, Saizhuo Wang, Tengfei Ma 0001, Fangli Xu, Alex X. Liu, Chunming Wu 0001, Shouling Ji |
ACM Trans. Knowl. Discov. Data | 9 |
| 2020 | Towards Fighting Cybercrime: Malicious URL Attack Type Detection using Multiclass ClassificationabstractMalicious Uniform Resource Locators (URLs) re-main one of the most common threats to cybersecurity. They are commonly spread through phishing, malware and spam. One popular way to detect malicious URLs is through black-lists. Blacklists maintain records of previously known malicious URL reputations. These lists are however shortcoming when there is need to detect newly generated malicious URLs. For that reason, modern research has resorted to training machine learning algorithms to detect malicious URLs. In this paper, we contributed towards the detection of malicious URLs using URL based features in a multiclass classification setting. We focused on three popular URL attack types which are phishing, spam and malware. Our work can be used as a supplementary tool in new or existing anti-phishing, anti-spam and anti-malware detection platforms. We compared the performance of the following ensemble learners: Extreme Gradient Boosting (XGBoost), Adaptive Boosting (AdaBoost), Light Gradient Boosting (LightGBM) and Categorical Boosting (CatBoost). We evaluated the performance of some URL features that we referred to as our features. These included priority features like Kullback-Leibler Divergence (KL divergence), bag of words segmentation and other word-based features. Results showed that our features performed better when compared to experiments we conducted without our features. We trained these algorithms on 126 983 URLs from benchmark datasets and all four learners returned an overall accuracy above 0.95. Tariro Manyumwa, Phillip Francis Chapita, Hanlu Wu, Shouling Ji |
IEEE BigData | 4 |
| 2020 | Attention with Long-Term Interval-Based Gated Recurrent Units for Modeling Sequential User Behaviors
Zhao Li 0007, Chenyi Lei, Pengcheng Zou, Donghui Ding, Shichang Hu, Zehong Hu, Shouling Ji, Jianliang Gao |
DASFAA (1) | 7 |
| 2020 | De-Health: All Your Online Health Information Are Belong to UsabstractIn this paper, we study the privacy of online health data. We present a novel online health data De-Anonymization (DA) framework, named De-Health. Leveraging two real world online health datasets WebMD and HealthBoards, we validate the DA efficacy of De-Health. We also present a linkage attack framework which can link online health/medical information to real world people. Through a proof-of-concept attack, we link 347 out of 2805 WebMD users to real world people, and find the full names, medical/health information, birthdates, phone numbers, and other sensitive information for most of the re-identified users. This clearly illustrates the fragility of the privacy of those who use online health forums. Shouling Ji, Qinchen Gu, Haiqin Weng, Qianjun Liu, Pan Zhou 0001, Jing Chen 0003, Zhao Li 0007, Raheem A. Beyah, Ting Wang 0006 |
ICDE | 1 |
| 2020 | AdvMind: Inferring Adversary Intent of Black-Box AttacksabstractDeep neural networks (DNNs) are inherently susceptible to adversarial attacks even under black-box settings, in which the adversary only has query access to the target models. In practice, while it may be possible to effectively detect such attacks (e.g., observing massive similar but non-identical queries), it is often challenging to exactly infer the adversary intent (e.g., the target class of the adversarial example the adversary attempts to craft) especially during early stages of the attacks, which is crucial for performing effective deterrence and remediation of the threats in many scenarios. Ren Pang, Xinyang Zhang 0001, Shouling Ji, Xiapu Luo, Ting Wang 0006 |
KDD | 3 |
| 2020 | Fast and parameter-light rare behavior detection in maritime trajectoriesabstractRare behaviors indicate important events and situations in maritime surveillance applications. State-of-the-art methods provide many effective solutions to detect anomalous behaviors. Meanwhile, most solutions are parameter-laden and too costly to identify useful rare behaviors with human knowledge in a visual analytics manner. This paper is concerned with a scheme cross trajectories, vessel attributes and the movement context for detecting rare behaviors through preprocessing, kNN-based clustering, and verification. Although the scheme involves several parameters, we demonstrate that they are able to be tackled in thresholds. As a result, a rare behavior factor is the single parameter that affect the detecting results. The proposed scheme is evaluated via a simulated data set for performance and a real life AIS data for effectiveness. Results show that high accuracy to labelled anomalies and useful rare behaviors can be achieved. Yifan Lei, Zhenguang Liu, Xun Wang 0007, Shouling Ji, Anthony K. H. Tung |
Inf. Process. Manag. | 5 |
| 2019 | CATS: Cross-Platform E-Commerce Fraud DetectionabstractNowadays, the popularity of e-commerce has brought huge economic benefits to factories, third-party merchants, and e-commerce service providers. Driven by such huge economic benefits, malicious merchants attempt to promote items through inserting fraudulent purchases, fake review scores, and/or feedback, into them. Mitigating this threat is challenging due to the difficulty of obtaining internal e-commerce data, the variance of e-commerce services used by malicious merchants, and the reluctance of service providers in cooperation. In this paper, we present an efficient, platform-independent, and robust e-commerce fraud detection system, CATS, to detect frauds for different large-scale e-commerce platforms. We implement the design of CATS into a prototype system and evaluate this prototype on the world's popular e-commerce platform Taobao. The evaluation result on Taobao shows that CATS can achieve a high accuracy of 91% in detecting frauds. Based on this success, we then apply CATS on another large-scale e-commerce platforms, and again CATS achieves an accuracy of 96%, suggesting that CATS is very effective on real e-commerce platforms. Based on the cross-platform evaluation results, we conduct a comprehensive analysis on the reported frauds and reveal several abnormal yet interesting behaviors of those reported frauds. Our study in this paper is expected to shed light on defending against frauds for various e-commerce platforms. Haiqin Weng, Shouling Ji, Fuzheng Duan, Zhao Li 0007, Jianhai Chen, Qinming He, Ting Wang 0006 |
ICDE | 2 |
| 2019 | Efficient Global String Kernel with Random Features: Beyond Counting SubstructuresabstractAnalysis of large-scale sequential data has been one of the most crucial tasks in areas such as bioinformatics, text, and audio mining. Existing string kernels, however, either (i) rely on local features of short substructures in the string, which hardly capture long discriminative patterns, (ii) sum over too many substructures, such as all possible subsequences, which leads to diagonal dominance of the kernel matrix, or (iii) rely on non-positive-definite similarity measures derived from the edit distance. Furthermore, while there have been works addressing the computational challenge with respect to the length of string, most of them still experience quadratic complexity in terms of the number of training samples when used in a kernel-based classifier. In this paper, we present a new class of global string kernels that aims to (i) discover global properties hidden in the strings through global alignments, (ii) maintain positive-definiteness of the kernel, without introducing a diagonal dominant kernel matrix, and (iii) have a training cost linear with respect to not only the length of the string but also the number of training string samples. To this end, the proposed kernels are explicitly defined through a series of different random feature maps, each corresponding to a distribution of random strings. We show that kernels defined this way are always positive-definite, and exhibit computational benefits as they always produce Random String Embeddings (RSE) that can be directly used in any linear classification models. Our extensive experiments on nine benchmark datasets corroborate that RSE achieves better or comparable accuracy in comparison to state-of-the-art baselines, especially with the strings of longer lengths. In addition, we empirically show that RSE scales linearly with the increase of the number and the length of string. Lingfei Wu 0001, Ian En-Hsu Yen, Siyu Huo, Liang Zhao 0002, Kun Xu 0005, Liang Ma 0002, Shouling Ji, Charu C. Aggarwal |
KDD | 7 |
| 2019 | TiSSA: A Time Slice Self-Attention Approach for Modeling Sequential User BehaviorsabstractModeling user behaviors as sequences provides critical advantages in predicting future user actions, such as predicting the next product to purchase or the next song to listen to, for personalized search and recommendation. Recently, recurrent neural networks (RNNs) have been adopted to leverage their power in modeling sequences. However, most of the previous RNN-based work suffers from the complex dependency problem, which may lose the integrity of highly correlated behaviors and may introduce noises derived from unrelated behaviors. In this paper, we propose to integrate a novel Time Slice Self-Attention (TiSSA) mechanism into RNNs for better modeling sequential user behaviors, which utilizes the time-interval-based gated recurrent units to exploit the temporal dimension when encoding user actions, and has a specially designed time slice hierarchical self-attention function to characterize both local and global dependency of user actions, while the final context-aware user representations can be used for downstream applications. We have performed experiments on a huge dataset collected from one of the largest e-commerce platforms in the world. Experimental results show that the proposed TiSSA achieves significant improvement over the state-of-the-art. TiSSA is also adopted in this large e-commerce platform, and the results of online A/B test further indicate its practical value. Chenyi Lei, Shouling Ji, Zhao Li 0007 |
WWW | 2 |
| 2018 | Online E-Commerce Fraud: A Large-Scale Detection and AnalysisabstractNowadays, e-commerce has become prevalent world-wide. With the big success of e-commerce, many malicious promotion services also rise: with the goal of increasing sales, malicious merchants attempt to promote their target items by illegally optimizing the search results using fake visits, purchases, etc. In this paper, we study the fraud detection problem on large-scale e-commerce platforms. First, we develop an efficient and scalable AnTi-Fraud system (ATF) to detect e-commerce frauds for large-scale e-commerce platforms, and implement it in parallel on a large-scale computing platform, called Open Data Processing Service (ODPS). Then, we evaluate ATF using two real large-scale e-commerce datasets (with tens of millions users and items). The results demonstrate that both the precision and the recall of ATF can achieve 0.97+, which suggests that ATF is very effective. More importantly, we deploy ATF on the Taobao platform of Alibaba, which is one of the world's largest e-commerce platforms. The evaluation results show that ATF can also achieve an accuracy of 98.16% on Taobao, which again suggests that ATF is very effective and deployable in practice. Our study in this paper is expected to shed light on defending against online frauds for practical e-commerce platforms. Haiqin Weng, Zhao Li 0007, Shouling Ji, Chen Chu, Haifeng Lu, Tianyu Du, Qinming He |
ICDE | 3 |
| 2016 | Sapprox: Enabling Efficient and Accurate Approximations on Sub-datasets with Distribution-aware Online SamplingabstractIn this paper, we aim to enable both efficient and accurate approximations on arbitrary sub-datasets of a large dataset. Due to the prohibitive storage overhead of caching offline samples for each sub-dataset, existing offline sample based systems provide high accuracy results for only a limited number of sub-datasets, such as the popular ones. On the other hand, current online sample based approximation systems, which generate samples at runtime, do not take into account the uneven storage distribution of a sub-dataset. They work well for uniform distribution of a sub-dataset while suffer low sampling efficiency and poor estimation accuracy on unevenly distributed sub-datasets. To address the problem, we develop a distribution aware method calledSapprox. Our idea is to collect the occurrences of a sub-dataset at each logical partition of a dataset (storage distribution) in the distributed system, and make good use of such information to facilitate online sampling. There are three thrusts in Sapprox. First, we develop a probabilistic map to reduce the exponential number of recorded sub-datasets to a linear one. Second, we apply thecluster sampling with unequal probability theoryto implement a distribution-aware sampling method for efficient online sub-dataset sampling. Third, we quantitatively derive the optimal sampling unit size in a distributed file system by associating it with approximation costs and accuracy. We have implemented Sapprox into Hadoop ecosystem as an example system and open sourced it on GitHub. Our comprehensive experimental results show that Sapprox can achieve a speedup by up to 20× over the precise execution. Xuhong Zhang 0002, Jun Wang 0001, Jiangling Yin, Shouling Ji |
Proc. VLDB Endow. | 4 |