EDBT 2026 Demo / reviewers in the wild / expert
Wenzhuo Yang
dblp:18/11088
· DBLP profile ↗
22ranked-venue papers
10as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 first-author · 2 since 2021Computer networks · 6 · 1 first-author · 6 since 2021Security and privacy · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep Learning Approaches for Anti-Money Laundering on Mobile Transactions: Review, Framework, and DirectionsabstractMoney laundering is a financial crime that obscures the origin of illicit funds, necessitating the development and enforcement of anti-money laundering (AML) policies by governments and organizations. The proliferation of mobile payment platforms and smart IoT devices has significantly complicated AML investigations. As payment networks become more interconnected, there is an increasing need for efficient real-time detection to process large volumes of transaction data on heterogeneous payment systems by different operators such as digital currencies, cryptocurrencies and account-based payments. Most of these mobile payment networks are supported by connected devices, many of which are considered loT devices in the FinTech space that constantly generate data. Furthermore, the growing complexity and unpredictability of transaction patterns across these networks contribute to a higher incidence of false positives. While machine learning solutions have the potential to enhance detection efficiency, their application in AML faces unique challenges, such as addressing privacy concerns tied to sensitive financial data and managing the real-world constraint of limited data availability due to data regulations. Existing surveys in the AML literature broadly review machine learning approaches for money laundering detection, but they often lack an in-depth exploration of advanced deep learning techniques—an emerging field with significant potential. To address this gap, this paper conducts a comprehensive review of deep learning solutions and the challenges associated with their use in AML. Additionally, we propose a novel framework that applies the least-privilege principle by integrating machine learning techniques, codifying AML red flags, and employing account profiling to provide context for predictions and enable effective fraud detection under limited data availability. Specifically, our approach defines AML-relevant financial profile characteristics and risk indicators to contextualize transactions and assess their associated risks. The proposed context-risk-predict AML (CRP-AML) model demonstrates notable success, achieving an F1 score of 82.51% on the minority class and nearly doubling the performance of other pattern detection models when the proportion of money laundering records in the dataset drops as low as 0.0005. Jiani Fan, Lwin Khin Shar, Ruichen Zhang 0001, Ziyao Liu, Wenzhuo Yang, Dusit Niyato, Kwok-Yan Lam |
IEEE Internet Things J. | 5 |
| 2024 | Effective Intrusion Detection in Heterogeneous Internet-of-Things Networks via Ensemble Knowledge Distillation-Based Federated LearningabstractWith the rapid development of low-cost consumer electronics and cloud computing, Internet-of- Things (IoT) devices are widely adopted for supporting next-generation distributed systems such as smart cities and industrial control systems. IoT devices are often susceptible to cyber attacks due to their open deployment environment and limited computing capabilities for stringent security controls. Hence, Intrusion Detection Systems (IDS) have emerged as one of the effective ways of securing IoT networks by monitoring and detecting abnormal activities. However, existing IDS approaches rely on centralized servers to generate behaviour profiles and detect anomalies, causing high response time and large operational costs due to communication overhead. Besides, sharing of behaviour data in an open and distributed IoT network environment may violate on-device privacy requirements. Additionally, various IoT devices tend to capture heterogeneous data, which complicates the training of behaviour models. In this paper, we introduce Federated Learning (FL) to collaboratively train a decentralized shared model of IDS, without exposing training data to others. Furthermore, we propose an effective method called Federated Learning Ensemble Knowledge Distillation (FLEKD) to mitigate the heterogeneity problems across various clients. FLEKD enables a more flexible aggregation method than conventional model fusion techniques. Experiment results on the public dataset CICIDS2019 demonstrate that the proposed approach outperforms local training and traditional FL in terms of both speed and performance and significantly improves the system's ability to detect unknown attacks. Finally, we evaluate our proposed framework's performance in three potential real-world scenarios and show FLEKD has a clear advantage in experimental results. Jiyuan Shen, Wenzhuo Yang, Zhaowei Chu, Jiani Fan, Dusit Niyato, Kwok-Yan Lam |
ICC | 2 |
| 2024 | LEO Satellite-Enabled Networks: A Privacy-Preserving Framework for Spectrum Pricing and Power Control OptimizationabstractLow Earth orbit (LEO) satellite systems are receiving increasing attention as they provide extensive global coverage. Secure and efficient management of limited spectrum bands and power resources are crucial for controlling operational costs and ensuring reliable communication in LEO satellite systems. However, spectrum pricing and power control optimization are challenging tasks. First, dynamic pricing is needed for leasing idle satellite spectrum to terrestrial users, as it must consider user mobility and real-time demand changes. Additionally, there is a trust concern that when utilizing the leased spectrum, terrestrial users may maliciously exceed limited transmit power to improve the quality of service (QoS). Moreover, users' privacy should be protected because the data collected by satellites often contain sensitive information such as location, budget, and QoS needs. In this paper, we propose a hybrid spectrum pricing and power control framework for LEO satellite-enabled networks to mitigate the above concerns by combining blockchain technology and Federated Learning (FL). We first design a local deep reinforcement learning algorithm for LEO satellite systems to learn a revenue-maximizing pricing strategy and power control scheme. Subsequently, these individual agents collaborate to establish an FL system without sharing their sensitive raw data. We also propose a reputation-based blockchain used in the global model aggregation phase to further enhance the traceability of the network and guarantee the trust. We conduct simulation tests to evaluate the efficacy of the proposed scheme, and our results show its capability to efficiently find the maximum revenue scheme for LEO satellite systems while preserving the privacy of each participating agent in an auditable mode. Bowen Shen, Kwok-Yan Lam, Wenzhuo Yang, Ziyao Liu, Feng Li 0008 |
MSN | 3 |
| 2024 | AutoFL: A Bayesian Game Approach for Autonomous Client Participation in Federated Edge LearningabstractGiven that devices (i.e., clients) participating in federated edge learning (FEL) are autonomous and resource-constrained in nature, it is critical to design effective incentive mechanisms to encourage client participation so as to improve the performance of FEL. In this article, we aim to boost the FEL training efficiency by answering how much compute resource should clients autonomously contribute to maximize their utilities. To this end, we develop AutoFL, an autonomous client participation decision framework for federated learning at the network edge without assuming that each client possesses complete information. We first model the problem of autonomous client participation as a Bayesian game with incomplete information, where each player in the game is associated with a set of types according to network conditions. We optimize an individual client's decision based on the dynamics of the population estimated following the Bayes rule. We prove that AutoFL can converge to a unique Bayesian Nash equilibrium point. Empirical results on three real datasets show that AutoFL achieves a higher model accuracy with only 15.5-24.5% model aggregation time per global training round, and its energy cost saving on mobile devices is 82.2-86.8% compared to the state-of-the-art algorithms. Moreover, we can achieve a 2.75-3.2x long-term fairness compared to classical solutions. Miao Hu 0001, Wenzhuo Yang, Zhenxiao Luo, Xuezheng Liu, Yipeng Zhou, Xu Chen 0004, Di Wu 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Understanding Security in Smart City Domains From the ANT-Centric PerspectiveabstractA city is a large human settlement that serves the people who live there, and a smart city is a concept of how cities might better serve their residents through new forms of technology. In this article, we focus on four major smart city domains according to Maslow’s hierarchy of needs: smart utility, smart transportation, smart homes, and smart healthcare. Numerous Internet of Things (IoT) applications have been developed to achieve the intelligence that we desire in our smart domains, ranging from personal gadgets, such as health trackers and smart watches to large-scale industrial IoT systems, such as nuclear and energy management systems. However, many of the existing smart city IoT solutions can be made better by considering the suitability of their security strategies. Inappropriate system security designs generally occur in two scenarios: first, system designers recognize the importance of security but are unsure of where, when, or how to implement it and second, system designers try to fit traditional security designs to meet the smart city security context. Thus, the objective of this article is to provide application designers with the missing security link they may need in order to improve their security designs. By evaluating the specific context of each smart city domain and the context-specific security requirements, we aim to provide directions on when, where, and how they should implement security strategies and the possible security challenges they need to consider. In addition, we present a new perspective on security issues in smart cities from a data-centric viewpoint by referring to the reference architecture, the activity-network-things (ANTs)-centric architecture. This architecture is built upon the concept of “security in a zero-trust environment,” to achieve end-to-end data security. By doing so, we reduce the security risks posed by new system interactions or unanticipated user behaviors while avoiding the hassle of regularly upgrading security models. Jiani Fan, Wenzhuo Yang, Ziyao Liu, Jiawen Kang 0001, Dusit Niyato, Kwok-Yan Lam, Hongyang Du 0001 |
IEEE Internet Things J. | 2 |
| 2023 | Merlion: End-to-End Machine Learning for Time SeriesabstractWe introduce Merlion, an open-source machine learning library for time series. It features a unified interface for many commonly used models and datasets for forecasting and anomaly detection on both univariate and multivariate time series, along with standard pre/post-processing layers. It has several modules to improve ease-of-use, including a no-code visual dashboard, anomaly score calibration to improve interpetability, AutoML for hyperparameter tuning and model selection, and model ensembling. Merlion also provides an evaluation framework that simulates the live deployment of a model in production, and a distributed computing backend to run time series models at industrial scale. This library aims to provide engineers and researchers a one-stop solution to rapidly develop models for their specific time series needs and benchmark them across multiple datasets. Aadyot Bhatnagar, Paul Kassianik, Tian Lan 0006, Wenzhuo Yang, Rowan Cassius, Doyen Sahoo, Devansh Arpit, Sri Subramanian, Gerald Woo, Amrita Saha, Arun Kumar Jagota, Gokulakrishnan Gopalakrishnan, K. C. Krithika, Sukumar Maddineni, Dae-ki Cho, Bo Zong, Yingbo Zhou 0002, Caiming Xiong, Silvio Savarese, Steven C. H. Hoi, Huan Wang 0016 |
J. Mach. Learn. Res. | 5 |
| 2022 | Differentiated Security Architecture for Secure and Efficient Infotainment Data Communication in IoV Networks
Jiani Fan, Lwin Khin Shar, Wenzhuo Yang, Dusit Niyato, Kwok-Yan Lam |
NSS | 4 |
| 2022 | Gain Without Pain: Offsetting DP-Injected Noises Stealthily in Cross-Device Federated LearningabstractFederated learning (FL) is an emerging paradigm through which decentralized devices can collaboratively train a common model. However, a serious concern is the leakage of privacy from exchanged gradient information between clients and the parameter server (PS) in FL. To protect gradient information, clients can adopt differential privacy (DP) to add additional noises and distort original gradients before they are uploaded to the PS. Nevertheless, the model accuracy will be significantly impaired by DP noises, making DP impracticable in real systems. In this work, we propose a novel noise information secretly sharing (NISS) algorithm to alleviate the disturbance of DP noises by sharing negated noises among clients. We theoretically prove that: 1) if clients are trustworthy, DP noises can be perfectly offset on the PS and 2) clients can easily distort negated DP noises to protect themselves in case that other clients are not totally trustworthy, though the cost lowers model accuracy. NISS is particularly applicable for FL across multiple Internet of Things (IoT) systems, in which all IoT devices need to collaboratively train a model. To verify the effectiveness and the superiority of the NISS algorithm, we conduct experiments with the MNIST and CIFAR-10 data sets. The experimental results verify our analysis and demonstrate that NISS can improve model accuracy by 19% on average and obtain better privacy protection if clients are trustworthy. Wenzhuo Yang, Yipeng Zhou, Miao Hu 0001, Di Wu 0001, James Xi Zheng, Hui Wang 0011, Song Guo 0001, Chao Li 0067 |
IEEE Internet Things J. | 1 |
| 2021 | On the Diversity and Explainability of Recommender Systems: A Practical Framework for Enterprise App RecommendationabstractThis paper introduces an enterprise app recommendation problem with a new "to-business'' use case, which aims to assist a sales team acting as the bridge connecting the applications and developers with the customers who apply these apps to solve their business problems. Our recommender system is an assistant to the sales team, helping recommend relevant apps to the customers for their businesses and increasing the likelihood of improving sales revenue. Besides recommendation accuracy, recommendation diversity and explainability are even more crucial since they provide more exposure opportunities for app developers and improve the transparency and trustworthiness of the recommender system. To allow the sales team to explore unpopular but relevant apps and understand why such apps are recommended, we propose a novel framework for improving aggregate recommendation diversity and generating recommendation explanations, which supports a wide variety of models for improving recommendation accuracy. The model in our framework is simple yet effective, which can be trained in an end-to-end manner and deployed as a recommendation service easily. Furthermore, our framework can also apply to other generic recommender systems for improving diversity and generating explanations. Experiments on public and private datasets demonstrate the effectiveness of our framework and solution. Wenzhuo Yang, Jia Li 0015, Latrice Barnett, Markus Anderle, Simo Arajärvi, Harshavardhan Utharavalli, Caiming Xiong, Steven C. H. Hoi |
CIKM | 1 |
| 2021 | Effective Anomaly Detection Model Training with only Unlabeled Data by Weakly Supervised Learning Techniques
Wenzhuo Yang, Kwok-Yan Lam |
ICICS (1) | 1 |
| 2020 | Unsupervised Learning of Disentangled Location EmbeddingsabstractLearning semantically coherent location embeddings can benefit downstream applications such as human mobility prediction. However, the conflation of geographic and semantic attributes of a location can harm such coherence, especially when semantic labels are not provided for the learning. To resolve this problem, in this paper, we present a novel unsupervised method for learning location embeddings from human trajectories. Our method advances traditional transition-based techniques in two ways: 1) we alleviate the disturbance of geographic attributes on the semantics by disentangling the two spaces; and 2) we incorporate spatio-temporal attributes and regular visiting patterns of trajectories to capture the semantics more accurately. Moreover, we present the first quantitative evaluation on location embeddings by introducing an original query-based metric, and we apply the metric in experiments on two Foursquare datasets, which demonstrate the improvement our model achieves on semantic coherence. We further apply the learned embeddings to two downstream applications, namely next point-of-interest recommendation and trajectory verification. Empirical results demonstrate the advantages of the disentangled embeddings over four state-of-the-art unsupervised location embedding methods. Kun Ouyang, Yuxuan Liang 0002, Ye Liu 0002, David S. Rosenblum, Wenzhuo Yang |
IJCNN | 5 |
| 2019 | Automated Cyber Threat Intelligence Reports Classification for Early Warning of Cyber Attacks in Next Generation SOC
Wenzhuo Yang, Kwok-Yan Lam |
ICICS | 1 |
| 2018 | Using Blockchain to Control Access to Cloud Data
Wenzhuo Yang, Kwok-Yan Lam, Xun Yi |
Inscrypt | 2 |
| 2018 | A Non-Parametric Generative Model for Human TrajectoriesabstractModeling human mobility and synthesizing realistic trajectories play a fundamental role in urban planning and privacy-preserving location data analysis. Due to its high dimensionality and also the diversity of its applications, existing trajectory generative models do not preserve the geometric (and more importantly) semantic features of human mobility, especially for longer trajectories. In this paper, we propose and evaluate a novel non-parametric generative model for location trajectories that tries to capture the statistical features of human mobility {\em as a whole}. This is in contrast with existing models that generate trajectories in a sequential manner. We design a new representation of locations, and use generative adversarial networks to produce data points in that representation space which will be then transformed to a time-series location trajectory form. We evaluate our method on realistic location trajectories and compare our synthetic traces with multiple existing methods on how they preserve geographic and semantic features of real traces at both aggregated and individual levels. The empirical results prove the capability of our model in preserving the utility of real data. Kun Ouyang, Reza Shokri, David S. Rosenblum, Wenzhuo Yang |
IJCAI | 4 |
| 2016 | Online Collaborative Learning for Open-Vocabulary Visual ClassifiersabstractWe focus on learning open-vocabulary visual classifiers, which scale up to a large portion of natural language vocabulary (e.g., over tens of thousands of classes). In particular, the training data are large-scale weakly labeled Web images since it is difficult to acquire sufficient well-labeled data at this category scale. In this paper, we propose a novel online learning paradigm towards this challenging task. Different from traditional N-way independent classifiers that generally fail to handle the extremely sparse and inter-related labels, our classifiers learn from continuous label embeddings discovered by collaboratively decomposing the sparse image-label matrix. Leveraging on the structure of the proposed collaborative learning formulation, we develop an efficient online algorithm that can jointly learn the label embeddings and visual classifiers. The algorithm can learn over 30,000 classes of 1,000 training images within 1 second on a standard GPU. Extensively experimental results on four benchmarks demonstrate the effectiveness of our method. Hanwang Zhang, Xindi Shang, Wenzhuo Yang, Huan Xu 0001, Huan-Bo Luan, Tat-Seng Chua |
CVPR | 3 |
| 2016 | A Transparent Learning Approach for Attack Prediction Based on User Behavior Analysis
Peizhi Shao, Jiuming Lu, Raymond K. Wong 0001, Wenzhuo Yang |
ICICS | 4 |
| 2015 | A Unified Framework for Outlier-Robust PCA-like AlgorithmsabstractWe propose a unified framework for making a wide range of PCA-like algorithms – including the standard PCA, sparse PCA and non-negative sparse PCA, etc. – robust when facing a constant fraction of arbitrarily corrupted outliers. Our theoretic analysis establishes solid performance guarantees of the proposed framework: its estimation error is upper bounded by a term depending on the intrinsic parameters of the data model, the selected PCA-like algorithm and the fraction of outliers. Comprehensive experiments on synthetic and real-world datasets demonstrate that the outlier-robust PCA-like algorithms derived from our framework have outstanding performance. Wenzhuo Yang, Huan Xu 0001 |
ICML | 1 |
| 2015 | Streaming Sparse Principal Component AnalysisabstractThis paper considers estimating the leading k principal components with at most s non-zero attributes from p-dimensional samples collected sequentially in memory limited environments. We develop and analyze two memory and computational efficient algorithms called streaming sparse PCA and streaming sparse ECA for analyzing data generated according to the spike model and the elliptical model respectively. In particular, the proposed algorithms have memory complexity O(pk), computational complexity O(pk mink,slogp) and sample complexity Θ(s \log p). We provide their finite sample performance guarantees, which implies statistical consistency in the high dimensional regime. Numerical experiments on synthetic and real-world datasets demonstrate good empirical performance of the proposed algorithms. Wenzhuo Yang, Huan Xu 0001 |
ICML | 1 |
| 2015 | A Divide and Conquer Framework for Distributed Graph ClusteringabstractGraph clustering is about identifying clusters of closely connected nodes, and is a fundamental technique of data analysis with many applications including community detection, VLSI network partitioning, collaborative filtering, and many others. In order to improve the scalability of existing graph clustering algorithms, we propose a novel divide and conquer framework for graph clustering, and establish theoretical guarantees of exact recovery of the clusters. One additional advantage of the proposed framework is that it can identify small clusters – the size of the smallest cluster can be of size o(\sqrtn), in contrast to Ω(\sqrtn) required by standard methods. Extensive experiments on synthetic and real-world datasets demonstrate the efficiency and effectiveness of our framework. Wenzhuo Yang, Huan Xu 0001 |
ICML | 1 |
| 2014 | The Coherent Loss Function for ClassificationabstractA prediction rule in binary classification that aims to achieve the lowest probability of misclassification involves minimizing over a non-convex, 0-1 loss function, which is typically a computationally intractable optimization problem. To address the intractability, previous methods consider minimizing the cumulative loss – the sum of convex surrogates of the 0-1 loss of each sample. In this paper, we revisit this paradigm and develop instead an axiomatic framework by proposing a set of salient properties on functions for binary classification and then propose the coherent loss approach, which is a tractable upper-bound of the empirical classification error over the entire sample set. We show that the proposed approach yields a strictly tighter approximation to the empirical classification error than any convex cumulative loss approach while preserving the convexity of the underlying optimization problem, and this approach for binary classification also has a robustness interpretation which builds a connection to robust SVMs. The experimental results show that our approach outperforms the standard SVM when additional constraints are imposed. Wenzhuo Yang, Melvyn Sim, Huan Xu 0001 |
ICML | 1 |
| 2013 | A Unified Robust Regression Model for Lasso-like AlgorithmsabstractWe develop a unified robust linear regression model and show that it is equivalent to a general regularization framework to encourage sparse-like structure that contains group Lasso and fused Lasso as specific examples. This provides a robustness interpretation of these widely applied Lasso-like algorithms, and allows us to construct novel generalizations of Lasso-like algorithms by considering different uncertainty sets. Using this robustness interpretation, we present new sparsity results, and establish the statistical consistency of the proposed regularized linear regression. This work extends a classical result from Xu et al. (2010) that relates standard Lasso with robust linear regression to learning problems with more general sparse-like structures, and provides new robustness-based tools to to understand learning problems with sparse-like structures. Wenzhuo Yang, Huan Xu 0001 |
ICML (3) | 1 |
| 2012 | Consistent depth maps recovery from a trinocular video sequenceabstractIn this paper, we propose a novel dense depth recovery method for a trinocular video sequence. Specifically, we contribute a novel trinocular stereo matching model, which can effectively utilize the advantages of trinocular stereo images, and incorporate the visibility term with segmentation prior for robust depth estimate. In order to make the recovered depth maps more accurate and temporally consistent, we propose to first classify the pixels to static and dynamic ones, and then perform spatio-temporal depth optimization for them in different ways. Especially, we propose two motion models for handling dynamic pixels. The traditional bundle optimization model and our spatio-temporal optimization model are softly combined in a probabilistic way, so that the depths of both static and dynamic pixels can be effectively refined. Our automatic depth recovery approach is evaluated using a variety of challenging trinocular video sequences. Wenzhuo Yang, Guofeng Zhang 0001, Hujun Bao |
CVPR | 1 |