VLDB 2026 Research / reviewers in the wild / expert
Sin G. Teo
dblp:71/11239 · also Sin Gee Teo
· DBLP profile ↗
23ranked-venue papers
3as first author
13since 2021 · last 2027
0000-0003-1090-505XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Security and privacy · 5 · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Shadowed set-based three-way clustering with bilevel aggregation for personalized federated learning
Shubao Zhao, Sin G. Teo, Zengxiang Li |
Inf. Sci. | 2 |
| 2026 | Assessing Privacy Disclosure Compliance of Android Third-Party SDKs
Mark Huasong Meng, Chuan Yan, Zhang Qing cnwatcher, Kailong Wang 0001, Sin G. Teo, Guangdong Bai, Jin Song Dong 0001 |
IEEE Trans. Software Eng. | 6 |
| 2025 | Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code RepositoriesabstractAlperen Yildiz, Sin G Teo, Yiling Lou, Yebo Feng, Chong Wang, Dinil Mon Divakaran. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Alperen Yildiz, Sin G. Teo, Yiling Lou, Yebo Feng, Chong Wang 0013, Dinil Mon Divakaran |
ACL (1) | 2 |
| 2025 | High-accuracy AoA-based Localization using Hierarchical ML Classifiers in Outdoor EnvironmentsabstractAccurate and reliable localization is a key requirement for 6G network operations, but it can be particularly challenging in outdoor environments. In this paper, we propose a machine learning (ML)–based localization framework that leverages angle of arrival (AoA) as a feature extracted from channel state information (CSI). The proposed approach employs high-resolution AoA estimation algorithms, including multiple signal classification (MUSIC) and estimation of signal parameters via rotational invariance techniques (ESPRIT), which feed a hierarchical, two-stage classifier to identify specific trajectories (hereafter referred to as tracks) in a given outdoor environment. The first stage of the classifier is a binary line-of-sight (LoS) / non-line-of-sight (NLoS) classifier, followed by region-specific multi-class classifiers for fine-grained identification of the specific LoS or NLoS tracks. We evaluate our approach using a real-world massive multiple-input multiple-output (mMIMO) orthogonal frequency division multiplexing (OFDM) outdoor CSI dataset collected at the Nokia campus in Stuttgart, Germany. Experimental results show that i) the LoS / NLoS identification accuracy can reach 100%, and ii) the proposed two-stage approach significantly outperforms a single-stage multi-class baseline, achieving accuracy over 98% in LoS regions and 95% in NLoS regions. These findings demonstrate the potential of combining AoA with ML for robust localization in outdoor mMIMO propagation environments. Bac Trinh-Nguyen, Sara Berri, Sin G. Teo, Tram Truong Huu, Arsenia Chorti |
GLOBECOM | 3 |
| 2025 | FETA: A systematic and efficient approach for feature engineering on anti-static and anti-dynamic malware analysisabstractMalware detection is a critical but very challenging task in cybersecurity. The eternal competition between malware authors (cyber attackers) and security analysts (detectors) is a never-ending game in which malware evolves rapidly and becomes more sophisticated as cyber attackers constantly evolve their tactics to evade detection. Such competition raises the demand for new automated malware detection techniques to keep pace with malware evolution and address sophisticated malware. This paper presents an empirical study that analyzes the effectiveness of static and dynamic features using machine learning algorithms. We propose FETA, a systematic approach for F eature E ngineering on anti-s T atic and anti-dyn A mic malware analysis. FETA combines static and dynamic features through feature aggregation and model integration techniques to improve detection accuracy and robustness. Extensive experiments on a real-world dataset show that the aggregation of static and dynamic features outperforms individual feature sets, achieving a detection rate of 98.06%. Additionally, we provide insights into feature selection and conduct a deep analysis of misclassified samples. This research contributes to the development of more effective and efficient malware detection techniques for enhanced cybersecurity. Dima Rabadi, Jia Yi Loo, Amudha Narayanan, Sin G. Teo, Tram Truong Huu |
J. Inf. Secur. Appl. | 5 |
| 2025 | Subkv: Quantizing Long Context KV Cache for Sub-Billion Parameter Language Models on Edge DevicesabstractABSTRACT Background Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their substantial computational and memory requirements present significant challenges for widespread deployment on edge devices. Motivation In long‐context scenarios, even sub‐billion parameter LLMs face unavoidable memory and performance bottlenecks due to inefficient KV Cache utilization. Existing quantization methods fail to address these challenges effectively. Method This paper addresses these challenges by introducing advanced quantization techniques tailored for sub‐billion parameter LLMs. It specifically targets reducing memory consumption through the conversion of the model's KV Cache to lower‐bit integers. We present SubKV, a quantization method specifically designed to optimize the KV Cache in sub‐billion parameter LLMs. Our analysis reveals distinct distributional differences in the magnitude of key and value caches. Leveraging this insight, we apply Per‐Channel Quantization to the key cache and Per‐Token Quantization to the value cache. Furthermore, we introduce the Dynamic Window Quantization method to enhance attention computations. To mitigate the extreme sensitivity of the first token, we also introduce Attention Sink‐Aware Quantization. Results Experimental results demonstrate that SubKV significantly reduces the KV Cache size during long context inference while maintaining model performance, offering superior results to existing KV Cache quantization methods. Ziqian Zeng, Tao Zhang 0019, Zhengdong Lu, Huiping Zhuang, Hongen Shao, Sin G. Teo, Xiaofeng Zou |
Softw. Pract. Exp. | 7 |
| 2025 | Guest Editorial: Special Issue on Trustworthy Federated Learning
Qiang Yang 0001, Han Yu 0001, Sin G. Teo, Bo Li 0001, Guodong Long, Lixin Fan, Yang Liu 0165, Le Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Fast and Efficient Malware Detection with Joint Static and Dynamic Features Through Transfer Learning
Mao V. Ngo, Tram Truong Huu, Dima Rabadi, Jia Yi Loo, Sin G. Teo |
ACNS (1) | 5 |
| 2023 | Supervised Robustness-preserving Data-free Neural Network PruningabstractWhen deploying pre-trained neural network models in real-world applications, model consumers often encounter resource-constraint platforms such as mobile and smart devices. They typically use the pruning technique to reduce the size and complexity of the model, generating a lighter one with less resource consumption. Nonetheless, most existing pruning methods are proposed with a premise that the model after being pruned has a chance to be fine-tuned or even retrained based on the original training data. This may be unrealistic in practice, as the data controllers are often reluctant to provide their model consumers with the original data. In this work, we study the neural network pruning in the data-free context, aiming to yield lightweight models that are not only accurate in prediction but also robust against undesired inputs in open-world deployments. Considering the absence of fine-tuning and retraining that can fix the mis-pruned units, we replace the traditional aggressive one-shot strategy with a conservative one that treats model pruning as a progressive process. We propose a pruning method based on stochastic optimization that uses robustness-related metrics to guide the pruning process. Our method is evaluated with a series of experiments on diverse neural network models. The experimental results show that it significantly outperforms existing one-shot data-free pruning approaches in terms of robustness preservation and accuracy. Mark Huasong Meng, Guangdong Bai, Sin G. Teo, Jin Song Dong 0001 |
ICECCS | 3 |
| 2023 | Post-GDPR Threat Hunting on Android Phones: Dissecting OS-level Safeguards of User-unresettable Identifiers
Mark Huasong Meng, Zhang Qing cnwatcher, Guangshuai Xia, Yuwei Zheng, Yanjun Zhang 0002, Guangdong Bai, Sin G. Teo, Jin Song Dong 0001 |
NDSS | 8 |
| 2023 | Enhancing Federated Learning Robustness Using Data-Agnostic Model Pruning
Mark Huasong Meng, Sin G. Teo, Guangdong Bai, Kailong Wang 0001, Jin Song Dong 0001 |
PAKDD (2) | 2 |
| 2023 | Privacy-Preserving Cross-Environment Human Activity RecognitionabstractRecent studies have demonstrated the success of using the channel state information (CSI) from the WiFi signal to analyze human activities in a fixed and well-controlled environment. Those systems usually degrade when being deployed in new environments. A straightforward solution to solve this limitation is to collect and annotate data samples from different environments with advanced learning strategies. Although workable as reported, those methods are often privacy sensitive because the training algorithms need to access the data from different environments, which may be owned by different organizations. We present a practical method for the WiFi-based privacy-preserving cross-environment human activity recognition (HAR). It collects and shares information from different environments, while maintaining the privacy of individual person being involved. At the core of our approach is the utilization of the Johnson-Lindenstrauss transform, which is theoretically shown to be differentially private. Based on that, we further design an adversarial learning strategy to generate environment-invariant representations for HAR. We demonstrate the effectiveness of the proposed method with different data modalities from two real-life environments. More specifically, on the raw CSI dataset, it shows 2.18% and 1.24% improvements over challenging baselines for two environments, respectively. Moreover, with the discrete wavelet transform features, it further yields 5.71% and 1.55% improvements, respectively. Le Zhang 0001, Wei Cui 0002, Bing Li 0002, Zhenghua Chen, Min Wu 0008, Sin G. Teo |
IEEE Trans. Cybern. | 6 |
| 2022 | Detecting Contradictions from CoAP RFC Based on Knowledge Graph
Xinguo Feng, Yanjun Zhang 0002, Mark Huasong Meng, Sin G. Teo |
NSS | 4 |
| 2020 | Advanced Windows Methods on Malware Detection and ClassificationabstractApplication Programming Interfaces (APIs) are still considered the standard accessible data source and core wok of the most widely adopted malware detection and classification techniques. API-based malware detectors highly rely on measuring API’s statistical features, such as calculating the frequency counter of calling specific API calls or finding their malicious sequence pattern (i.e., signature-based detectors). Using simple hooking tools, malware authors would help in failing such detectors by interrupting the sequence and shuffling the API calls or deleting/inserting irrelevant calls (i.e., changing the frequency counter). Moreover, relying on API calls (e.g., function names) alone without taking into account their function parameters is insufficient to understand the purpose of the program. For example, the same API call (e.g., writing on a file) would act in two ways if two different arguments are passed (e.g., writing on a system versus user file). However, because of the heterogeneous nature of API arguments, most of the available API-based malicious behavior detectors would consider only the API calls without taking into account their argument information (e.g., function parameters). Alternatively, other detectors try considering the API arguments in their techniques, but they acquire having proficient knowledge about the API arguments or powerful processors to extract them. Such requirements demand a prohibitive cost and complex operations to deal with the arguments. To overcome the above limitations, with the help of machine learning and without any expert knowledge of the arguments, we propose a light-weight API-based dynamic feature extraction technique, and we use it to implement a malware detection and type classification approach. To evaluate our approach, we use reasonable datasets of 7774 benign and 7105 malicious samples belonging to ten distinct malware types. Experimental results show that our type classification module could achieve an accuracy of , where our malware detection module could reach an accuracy of over , and outperforms many state-of-the-art API-based malware detectors. Dima Rabadi, Sin G. Teo |
ACSAC | 2 |
| 2020 | DAG: A General Model for Privacy-Preserving Data Mining : (Extended Abstract)abstractSecure multi-party computation (SMC) allows parties to jointly compute a function over their inputs, while keeping every input confidential. SMC has been extensively applied in tasks with privacy requirements, such as privacy-preserving data mining (PPDM), to learn task output and at the same time protect input data privacy. However, existing SMC-based solutions are ad-hoc - they are proposed for specific applications, and thus cannot be applied to other applications directly. To address this issue, we propose a privacy model DAG (Directed Acyclic Graph) that consists of a set of fundamental secure operators (e.g., +, -, ×, /, and power). Our model is general - its operators, if pipelined together, can implement various functions, even complicated ones. The experimental results also show that our DAG model can run in acceptable time. Sin G. Teo, Jianneng Cao, Vincent Cheng-Siong Lee |
ICDE | 1 |
| 2020 | Citywide Traffic Flow Prediction Based on Multiple Gated Spatio-temporal Convolutional Neural NetworksabstractTraffic flow prediction is crucial for public safety and traffic management, and remains a big challenge because of many complicated factors, e.g., multiple spatio-temporal dependencies, holidays, and weather. Some work leveraged 2D convolutional neural networks (CNNs) and long short-term memory networks (LSTMs) to explore spatial relations and temporal relations, respectively, which outperformed the classical approaches. However, it is hard for these work to model spatio-temporal relations jointly. To tackle this, some studies utilized LSTMs to connect high-level layers of CNNs, but left the spatio-temporal correlations not fully exploited in low-level layers. In this work, we propose novel spatio-temporal CNNs to extract spatio-temporal features simultaneously from low-level to high-level layers, and propose a novel gated scheme to control the spatio-temporal features that should be propagated through the hierarchy of layers. Based on these, we propose an end-to-end framework, multiple gated spatio-temporal CNNs (MGSTC), for citywide traffic flow prediction. MGSTC can explore multiple spatio-temporal dependencies through multiple gated spatio-temporal CNN branches, and combine the spatio-temporal features with external factors dynamically. Extensive experiments on two real traffic datasets demonstrates that MGSTC outperforms other state-of-the-art baselines. Cen Chen 0002, Kenli Li 0001, Sin G. Teo, Xiaofeng Zou, Keqin Li 0001, Zeng Zeng |
ACM Trans. Knowl. Discov. Data | 3 |
| 2020 | DAG: A General Model for Privacy-Preserving Data MiningabstractSecure multi-party computation (SMC) allows parties to jointly compute a function over their inputs, while keeping every input confidential. It has been extensively applied in tasks with privacy requirements, such as privacy-preserving data mining (PPDM), to learn task output and at the same time protect input data privacy. However, existing SMC-based solutions are ad-hoc - they are proposed for specific applications, and thus cannot be applied to other applications directly. To address this issue, we propose a privacy model DAG (Directed Acyclic Graph) that consists of a set of fundamental secure operators (e.g., +, -, ×, /, and power). Our model is general - its operators, if pipelined together, can implement various functions, even complicated ones like Naı̈ve Bayes classifier. It is also extendable - new secure operators can be defined to expand the functions that the model supports. For case study, we have applied our DAG model to two data mining tasks: kernel regression and Naı̈ve Bayes. Experimental results show that DAG generates outputs that are almost the same as those by non-private setting, where multiple parties simply disclose their data. The experimental results also show that our DAG model runs in acceptable time, e.g., in kernel regression, when training data size is 683,093, one prediction in non-private setting takes 5.93 sec, and that by our DAG model takes 12.38 sec. Sin G. Teo, Jianneng Cao, Vincent Cheng-Siong Lee |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | Gated Residual Recurrent Graph Neural Networks for Traffic PredictionabstractTraffic prediction is of great importance to traffic management and public safety, and very challenging as it is affected by many complex factors, such as spatial dependency of complicated road networks and temporal dynamics, and many more. The factors make traffic prediction a challenging task due to the uncertainty and complexity of traffic states. In the literature, many research works have applied deep learning methods on traffic prediction problems combining convolutional neural networks (CNNs) with recurrent neural networks (RNNs), which CNNs are utilized for spatial dependency and RNNs for temporal dynamics. However, such combinations cannot capture the connectivity and globality of traffic networks. In this paper, we first propose to adopt residual recurrent graph neural networks (Res-RGNN) that can capture graph-based spatial dependencies and temporal dynamics jointly. Due to gradient vanishing, RNNs are hard to capture periodic temporal correlations. Hence, we further propose a novel hop scheme into Res-RGNN to utilize the periodic temporal dependencies. Based on Res-RGNN and hop Res-RGNN, we finally propose a novel end-to-end multiple Res-RGNNs framework, referred to as “MRes-RGNN”, for traffic prediction. Experimental results on two traffic datasets have demonstrated that the proposed MRes-RGNN outperforms state-of-the-art methods significantly. Cen Chen 0002, Kenli Li 0001, Sin G. Teo, Xiaofeng Zou, Jie Wang 0042, Zeng Zeng |
AAAI | 3 |
| 2018 | Exploiting Spatio-Temporal Correlations with Multiple 3D Convolutional Neural Networks for Citywide Vehicle Flow PredictionabstractPredicting vehicle flows is of great importance to traffic management and public safety in smart cities, and very challenging as it is affected by many complex factors, such as spatio-temporal dependencies with external factors (e.g., holidays, events and weather). Recently, deep learning has shown remarkable performance on traditional challenging tasks, such as image classification, due to its powerful feature learning capabilities. Some works have utilized LSTMs to connect the high-level layers of 2D convolutional neural networks (CNNs) to learn the spatio-temporal features, and have shown better performance as compared to many classical methods in traffic prediction. However, these works only build temporal connections on the high-level features at the top layer while leaving the spatio-temporal correlations in the low-level layers not fully exploited. In this paper, we propose to apply 3D CNNs to learn the spatio-temporal correlation features jointly from low-level to high-level layers for traffic data. We also design an end-to-end structure, named as MST3D, especially for vehicle flow prediction. MST3D can learn spatial and multiple temporal dependencies jointly by multiple 3D CNNs, combine the learned features with external factors and assign different weights to different branches dynamically. To the best of our knowledge, it is the first framework that utilizes 3D CNNs for traffic prediction. Experiments on two vehicle flow datasets Beijing and New York City have demonstrated that the proposed framework, MST3D, outperforms the state-of-the-art methods. Cen Chen 0002, Kenli Li 0001, Sin G. Teo, Guizi Chen, Xiaofeng Zou, Xulei Yang, Ramaseshan C. Vijay, Jiashi Feng, Zeng Zeng |
ICDM | 3 |
| 2018 | Deep Learning for Practical Image Recognition: Case Study on Kaggle CompetitionsabstractIn past years, deep convolutional neural networks (DCNN) have achieved big successes in image classification and object detection, as demonstrated on ImageNet in academic field. However, There are some unique practical challenges remain for real-world image recognition applications, e.g., small size of the objects, imbalanced data distributions, limited labeled data samples, etc. In this work, we are making efforts to deal with these challenges through a computational framework by incorporating latest developments in deep learning. In terms of two-stage detection scheme, pseudo labeling, data augmentation, cross-validation and ensemble learning, the proposed framework aims to achieve better performances for practical image recognition applications as compared to using standard deep learning methods. The proposed framework has recently been deployed as the key kernel for several image recognition competitions organized by Kaggle. The performance is promising as our final private scores were ranked 4 out of 2293 teams for fish recognition on the challenge "The Nature Conservancy Fisheries Monitoring" and 3 out of 834 teams for cervix recognition on the challenge "Intel &MobileODT Cervical Cancer Screening", and several others. We believe that by sharing the solutions, we can further promote the applications of deep learning techniques. Xulei Yang, Zeng Zeng, Sin G. Teo, Li Wang 0057, Vijay Chandrasekhar 0001, Steven C. H. Hoi |
KDD | 3 |
| 2017 | Privacy and Utility Preservation for Location Data Using Stay Region Analysis
Manoranjan Dash, Sin G. Teo |
ADMA | 2 |
| 2017 | Cloud-of-clouds based resource provisioning strategy for continuous write applicationsabstractNowadays, more and more online services based on cloud computing have taken the places of some traditional applications (e.g., Health Care) that continuously generate large volume of data and require data storage and analysis in time. Such applications can be categorized as “Continuous Writing Applications” (CWA) that have particular requirements on bandwidth, storage, computation, and service reliability. In the meanwhile, they are very sensitive to the cost. In this paper, we present an architecture of multiple cloud service providers (CSPs) or “Cloud-of-Clouds” to provide services to the CWA and propose a novel resource scheduling algorithm to minimize the cost of entire systems. Difference from many research efforts that focus on a single resource, we take many factors into considerations that include user's requirements of bandwidth, storage and computation, the resources of CSPs that can provide, CSPs for data backup, the configurations of Cloud-of-Clouds, system models of CSPs, and many more. We first present the system models of classic CWA applications to capture the resource requirements of users on Cloud-of-Clouds. We then present the problem formulation and our optimal strategy of user scheduling based on Minimum First Derivative Length (MFDL) of load paths among the systems. Through theoretical analysis, we prove that our proposed algorithm Optimal user Scheduling for Cloud-of-Clouds (OSCC) can achieve the optimal solution. Zeng Zeng, Bharadwaj Veeravalli, Samee Ullah Khan, Sin G. Teo |
APCC | 4 |
| 2015 | DAG: A Model for Privacy Preserving ComputationabstractSecure multi-party computation (SMC) allows parties to jointly compute a function over their inputs, while keeping every input confidential. It has been extensively applied in privacy-preserving computation, such as privacy-preserving data mining (PPDM), to protect data privacy. However, most SMC-based solutions are ad-hoc. They are proposed for specific applications, and thus cannot be applied to other applications directly. To address this issue, we propose a privacy model DAG (Directed A cyclic Graph) that consists of a set of secure operators (e.g., Multiplication and division). Our DAG model is general -- its operators, if pipelined together, can implement various functions. It is also extendable -- secure operators can be defined to add new features to the model. As an application study, we have applied our DAG to kernel regression. Experiments on datasets of more than 680,000 tuples show that our DAG model is effective and its running time is nearly thrice that of non-privacy setting, where parties directly disclose data. Sin G. Teo, Jianneng Cao, Vincent Cheng-Siong Lee |
ICWS | 1 |