Vincent Cheng-Siong Lee

dblp:50/10222 · also Vincent C. S. Lee · DBLP profile ↗
← Back
44ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0001-5976-4601ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 17 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 A dynamic balanced training regime for mitigating class imbalance in thyroid nodule classification
abstract
Handling class imbalance remains a major challenge in developing intelligent diagnostic systems for ultrasound-based thyroid nodule classification. Synthetic sample generation often distorts data distributions and weakens texture fidelity, a critical concern in ultrasound, where fine anatomical details carry diagnostic value. Loss reweighting may destabilize optimization under severe imbalance, while recent approaches favor architectural complexity over training-level refinement, limiting clinical deployment. These issues hinder accurate minority-class detection in screening, where precise classification is essential to avoid missed diagnoses or unnecessary interventions. In this paper, a lightweight adaptive framework extending the Dynamic Balanced Training Regime (DBTR) is proposed for class-imbalanced thyroid ultrasound classification. First, balanced subsets are constructed directly from original ultrasound images, preserving diagnostic texture without synthetic augmentation. Then, a recall-guided learning-rate adjustment is introduced to modulate the classifier head with a fixed backbone, strengthening minority-class sensitivity under speckle noise. Experiments were conducted on the imbalanced DDTI benchmark using 10-fold cross-validation and further tested on two recent independent external cohorts, TN5000 and ThyUS2Path, acquired under different clinical settings. Internal evaluation demonstrates that the proposed framework achieves the highest AUC (0.714) among the evaluated CNN architectures, while maintaining high sensitivity (0.984) and improving specificity from 0.233 to 0.250 over the Xception baseline, reaching balanced accuracy (0.617). Under external evaluation, a 5.5% relative increase in specificity was obtained on ThyUS2Path despite domain shift, while sensitivity remained preserved across both cohorts. Computational benchmarking confirmed efficient deployment, supporting the integration of intelligent diagnostic support systems into point-of-care ultrasound screening workflows.
Fatimah Ali S. Aljalis, Vincent Cheng-Siong Lee, Xinyu Zhang 0017
Expert Syst. Appl.2
2025 The Distributed Intelligent Collaboration to AAV-Assisted VEC: Joint Position Optimization and Task Scheduling
abstract
Deploying autonomous aerial vehicles (AAVs) as aerial base stations enhances the coverage and performance of communication networks in vehicular edge computing scenarios. However, due to the limited communication range and energy capacity of AAVs, they cannot continuously cover entire areas or sustain long flights. Therefore, achieving full communication coverage of a target area with a minimal number of AAVs and efficient task offloading remains a significant challenge. To address this problem, the AAV-assisted two-stage intelligent collaboration (UTIC) method is proposed in this article to tackle the joint position optimization and task scheduling issue. First, a AAV-assisted two-stage task scheduling system model is designed to optimize the allocation process. Second, an Enhanced Particle Swarm Optimization algorithm is designed to determine the optimal positions of AAVs, ensuring complete coverage of all mobile vehicles (MVs) with the minimum number of AAVs. Third, deep deterministic policy gradient method is employed to find the optimal scheduling decisions for MVs, considering energy consumption, delay, and task priorities. Simulation results demonstrate that the proposed UTIC method can achieve nearly 20% reduction in AAV deployment and outperform three other classical reinforcement learning algorithms in terms of reducing system cost.
Meng Yi, Vincent Cheng-Siong Lee, Peng Yang 0014, Peisong Li, Yifan Zhang 0039, Wei Wei 0006, Honghao Gao
IEEE Internet Things J.2
2024 Computational Intelligence for Optimizing UAV Positioning and Task Scheduling in UAV-Assisted MEC Systems
Meng Yi, Vincent Cheng-Siong Lee, Yifan Zhang 0039, Peisong Li, Peng Yang 0014
ICONIP (4)2
2023 Beyond Smoothing: Unsupervised Graph Representation Learning with Edge Heterophily Discriminating
abstract
Unsupervised graph representation learning (UGRL) has drawn increasing research attention and achieved promising results in several graph analytic tasks. Relying on the homophily assumption, existing UGRL methods tend to smooth the learned node representations along all edges, ignoring the existence of heterophilic edges that connect nodes with distinct attributes. As a result, current methods are hard to generalize to heterophilic graphs where dissimilar nodes are widely connected, and also vulnerable to adversarial attacks. To address this issue, we propose a novel unsupervised Graph Representation learning method with Edge hEterophily discriminaTing (GREET) which learns representations by discriminating and leveraging homophilic edges and heterophilic edges. To distinguish two types of edges, we build an edge discriminator that infers edge homophily/heterophily from feature and structure information. We train the edge discriminator in an unsupervised way through minimizing the crafted pivot-anchored ranking loss, with randomly sampled node pairs acting as pivots. Node representations are learned through contrasting the dual-channel encodings obtained from the discriminated homophilic and heterophilic edges. With an effective interplaying scheme, edge discriminating and representation learning can mutually boost each other during the training phase. We conducted extensive experiments on 14 benchmark datasets and multiple learning scenarios to demonstrate the superiority of GREET.
Yixin Liu 0001, Yizhen Zheng, Daokun Zhang, Vincent Cheng-Siong Lee, Shirui Pan
AAAI4
2023 Finding the Missing-half: Graph Complementary Learning for Homophily-prone and Heterophily-prone Graphs
abstract
Real-world graphs generally have only one kind of tendency in their connections. These connections are either homophilic-prone or heterophily-prone. While graphs with homophily-prone edges tend to connect nodes with the same class (i.e., intra-class nodes), heterophily-prone edges tend to build relationships between nodes with different classes (i.e., inter-class nodes). Existing GNNs only take the original graph as input during training. The problem with this approach is that it forgets to take into consideration the ''missing-half'' structural information, that is, heterophily-prone topology for homophily-prone graphs and homophily-prone topology for heterophily-prone graphs. In our paper, we introduce Graph cOmplementAry Learning, namely GOAL, which consists of two components: graph complementation and complemented graph convolution. The first component finds the missing-half structural information for a given graph to complement it. The complemented graph has two sets of graphs including both homophily- and heterophily-prone topology. In the latter component, to handle complemented graphs, we design a new graph convolution from the perspective of optimisation. The experiment results show that GOAL consistently outperforms all baselines in eight real-world datasets.
Yizhen Zheng, He Zhang 0012, Vincent Cheng-Siong Lee, Yu Zheng 0013, Xiao Wang 0017, Shirui Pan
ICML3
2023 Learning Strong Graph Neural Networks with Weak Information
abstract
Graph Neural Networks (GNNs) have exhibited impressive performance in many graph learning tasks. Nevertheless, the performance of GNNs can deteriorate when the input graph data suffer from weak information, i.e., incomplete structure, incomplete features, and insufficient labels. Most prior studies, which attempt to learn from the graph data with a specific type of weak information, are far from effective in dealing with the scenario where diverse data deficiencies exist and mutually affect each other. To fill the gap, in this paper, we aim to develop an effective and principled approach to the problem of graph learning with weak information (GLWI). Based on the findings from our empirical analysis, we derive two design focal points for solving the problem of GLWI, i.e., enabling long-range propagation in GNNs and allowing information propagation to those stray nodes isolated from the largest connected component. Accordingly, we propose D2PT, a dual-channel GNN framework that performs long-range information propagation not only on the input graph with incomplete structure, but also on a global graph that encodes global semantic similarities. We further develop a prototype contrastive alignment algorithm that aligns the class-level prototypes learned from two channels, such that the two different information propagation processes can mutually benefit from each other and the finally learned model can well handle the GLWI problem. Extensive experiments on eight real-world benchmark datasets demonstrate the effectiveness and efficiency of our proposed methods in various GLWI scenarios.
Yixin Liu 0001, Kaize Ding, Jianling Wang, Vincent Cheng-Siong Lee, Huan Liu 0001, Shirui Pan
KDD4
2023 Compliance Analyses of Australia's Online Household Appliances
abstract
Commercially sold electrical or gas products must comply with the safety standards imposed within a country and get registered and certified by a regulated body. However, with the increasing transition of businesses to e-commerce platforms, it becomes challenging to govern the compliance status of online products. This can increase the risk of purchasing non-compliant products which may be unsafe to use. Additionally, examining the compliance status before purchasing can be strenuous because the relevant compliance information can be ambiguous and not always directly available. Therefore, we collaborated with a regulated body from Australia, Energy Safe Victoria, and conducted compliance analyses for household appliances sold on multiple online platforms. A fully autonomous method shown in this public repository is also introduced to check the compliance status of any online product. In this talk, we discuss the compliance check process, which incorporates fuzzy logic for textual matching and a Convolutional Neural Network (CNN) model to classify the product listing based on the images listed. Subsequently, we studied the results with the business users and found that many online listings are non-compliant, signifying that online-shopping consumers are highly susceptible to buying unsafe products. We hope this talk can inspire more follow-up works that collaborate with regulated bodies to introduce a user-friendly compliance check platform that assists in educating consumers to purchase compliant products.
Chang How Tan, Vincent Cheng-Siong Lee, Jessie Nghiem, Priya Laxman
WSDM2
2023 Anomaly Detection in Dynamic Graphs via Transformer
abstract
Detecting anomalies for dynamic graphs has drawn increasing attention due to their wide applications in social networks, e-commerce, and cybersecurity. Recent deep learning-based approaches have shown promising results over shallow methods. However, they fail to address two core challenges of anomaly detection in dynamic graphs: the lack of informative encoding for unattributed nodes and the difficulty of learning discriminate knowledge from coupled spatial-temporal dynamic graphs. To overcome these challenges, in this paper, we present a novelTransformer-basedAnomalyDetection framework forDYnamic graphs (TADDY). Our framework constructs a comprehensive node encoding strategy to better represent each node’s structural and temporal roles in an evolving graphs stream. Meanwhile, TADDY captures informative representation from dynamic graphs with coupled spatial-temporal patterns via a dynamic graph transformer model. The extensive experimental results demonstrate that our proposed TADDY framework outperforms the state-of-the-art methods by a large margin on six real-world datasets.
Yixin Liu 0001, Shirui Pan, Yu Guang Wang 0001, Liang Wang 0017, Qingfeng Chen, Vincent Cheng-Siong Lee
IEEE Trans. Knowl. Data Eng.7
2022 SFPAD: Sentiment-driven Firm Pricing Anomaly Detection Algorithm and Application
abstract
In January of 2021, the stock price of GameStop Corp ascended 20 times under the influence of social media in a mere two weeks’ time, becoming the first-ever case of sentiment-driven short squeeze. This caused certain stock prices to deviate significantly from their fundamental value, and hence, had an impact on the market efficiency. To help finance policymakers, regulators, and investment analysts to prepare for cases alike in the future, our study investigates a Machine Learning method of detecting sentiment-driven firm pricing anomaly using data collected from the GameStop frenzy. The proposed method combines anomaly pattern detection with sentiment analysis, which the latter is used to identify conditions that are considered crucial to the presence of sentiment-driven pricing. In our experiments, the adapted Machine Learning (SFPAD) method was able to outperform traditional detection technique when applied to historical time series.
Senyuan Zheng, Vincent Cheng-Siong Lee, Zhaolin Guan
IEEE Big Data2
2022 Unifying Graph Contrastive Learning with Flexible Contextual Scopes
abstract
Graph contrastive learning (GCL) has recently emerged as an effective learning paradigm to alleviate the reliance on labelling information for graph representation learning. The core of GCL is to maximise the mutual information between the representation of a node and its contextual representation (i.e., the corresponding instance with similar semantic information) summarised from the contextual scope (e.g., the whole graph or 1-hop neighbourhood). This scheme distils valuable self-supervision signals for GCL training. However, existing GCL methods still suffer from limitations, such as the incapacity or inconvenience in choosing a suitable contextual scope for different datasets and building biased contrastiveness. To address aforementioned problems, we present a simple self-supervised learning method termed Unifying Graph Contrastive Learning with Flexible Contextual Scopes (UGCL for short). Our algorithm builds flexible contextual representations with tunable contextual scopes by controlling the power of an adjacency matrix. Additionally, our method ensures contrastiveness is built within connected components to reduce the bias of contextual representations. Based on representations from both local and contextual scopes, UGCL optimises a very simple contrastive loss function for graph representation learning. Essentially, the architecture of UGCL can be considered as a general framework to unify existing GCL methods. We have conducted intensive experiments and achieved new state-of-the-art performance in six out of eight benchmark datasets compared with self-supervised graph representation learning baselines. Our code has been open sourced1.1https://github.com/zyzisastudyreallyhardguy/UGCL
Yizhen Zheng, Yu Zheng 0013, Xiaofei Zhou 0002, Chen Gong 0002, Vincent Cheng-Siong Lee, Shirui Pan
ICDM5
2022 Rethinking and Scaling Up Graph Contrastive Learning: An Extremely Efficient Approach with Group Discrimination
abstract
Graph contrastive learning (GCL) alleviates the heavy reliance on label information for graph representation learning (GRL) via self-supervised learning schemes. The core idea is to learn by maximising mutual information for similar instances, which requires similarity computation between two node instances. However, GCL is inefficient in both time and memory consumption. In addition, GCL normally requires a large number of training epochs to be well-trained on large-scale datasets. Inspired by an observation of a technical defect (i.e., inappropriate usage of Sigmoid function) commonly used in two representative GCL works, DGI and MVGRL, we revisit GCL and introduce a new learning paradigm for self-supervised graph representation learning, namely, Group Discrimination (GD), and propose a novel GD-based method called Graph Group Discrimination (GGD). Instead of similarity computation, GGD directly discriminates two groups of node samples with a very simple binary cross-entropy loss. In addition, GGD requires much fewer training epochs to obtain competitive performance compared with GCL methods on large-scale datasets. These two advantages endow GGD with very efficient property. Extensive experiments show that GGD outperforms state-of-the-art self-supervised methods on eight datasets. In particular, GGD can be trained in 0.18 seconds (6.44 seconds including data preprocessing) on ogbn-arxiv, which is orders of magnitude (10,000+) faster than GCL baselines while consuming much less memory. Trained with 9 hours on ogbn-papers100M with billion edges, GGD outperforms its GCL counterparts in both accuracy and efficiency.
Yizhen Zheng, Shirui Pan, Vincent Cheng-Siong Lee, Yu Zheng 0013, Philip S. Yu
NeurIPS3
2022 Information resources estimation for accurate distribution-based concept drift detection
Chang How Tan, Vincent Cheng-Siong Lee, Mahsa Salehi
Inf. Process. Manag.2
2021 Heterogeneous Graph Attention Network for Small and Medium-Sized Enterprises Bankruptcy Prediction
Yizhen Zheng, Vincent Cheng-Siong Lee, Zonghan Wu, Shirui Pan
PAKDD (1)2
2021 Robotic Hierarchical Graph Neurons. A novel implementation of HGN for swarm robotic behaviour control
Phillip Smith, Aldeida Aleti, Vincent Cheng-Siong Lee, Robert A. Hunjet
Expert Syst. Appl.3
2020 DAG: A General Model for Privacy-Preserving Data Mining : (Extended Abstract)
abstract
Secure multi-party computation (SMC) allows parties to jointly compute a function over their inputs, while keeping every input confidential. SMC has been extensively applied in tasks with privacy requirements, such as privacy-preserving data mining (PPDM), to learn task output and at the same time protect input data privacy. However, existing SMC-based solutions are ad-hoc - they are proposed for specific applications, and thus cannot be applied to other applications directly. To address this issue, we propose a privacy model DAG (Directed Acyclic Graph) that consists of a set of fundamental secure operators (e.g., +, -, ×, /, and power). Our model is general - its operators, if pipelined together, can implement various functions, even complicated ones. The experimental results also show that our DAG model can run in acceptable time.
Sin G. Teo, Jianneng Cao, Vincent Cheng-Siong Lee
ICDE3
2020 DAG: A General Model for Privacy-Preserving Data Mining
abstract
Secure multi-party computation (SMC) allows parties to jointly compute a function over their inputs, while keeping every input confidential. It has been extensively applied in tasks with privacy requirements, such as privacy-preserving data mining (PPDM), to learn task output and at the same time protect input data privacy. However, existing SMC-based solutions are ad-hoc - they are proposed for specific applications, and thus cannot be applied to other applications directly. To address this issue, we propose a privacy model DAG (Directed Acyclic Graph) that consists of a set of fundamental secure operators (e.g., +, -, ×, /, and power). Our model is general - its operators, if pipelined together, can implement various functions, even complicated ones like Naı̈ve Bayes classifier. It is also extendable - new secure operators can be defined to expand the functions that the model supports. For case study, we have applied our DAG model to two data mining tasks: kernel regression and Naı̈ve Bayes. Experimental results show that DAG generates outputs that are almost the same as those by non-private setting, where multiple parties simply disclose their data. The experimental results also show that our DAG model runs in acceptable time, e.g., in kernel regression, when training data size is 683,093, one prediction in non-private setting takes 5.93 sec, and that by our DAG model takes 12.38 sec.
Sin G. Teo, Jianneng Cao, Vincent Cheng-Siong Lee
IEEE Trans. Knowl. Data Eng.3
2019 Nearest-Neighbour-Induced Isolation Similarity and Its Impact on Density-Based Clustering
abstract
A recent proposal of data dependent similarity called Isolation Kernel/Similarity has enabled SVM to produce better classification accuracy. We identify shortcomings of using a tree method to implement Isolation Similarity; and propose a nearest neighbour method instead. We formally prove the characteristic of Isolation Similarity with the use of the proposed method. The impact of Isolation Similarity on densitybased clustering is studied here. We show for the first time that the clustering performance of the classic density-based clustering algorithm DBSCAN can be significantly uplifted to surpass that of the recent density-peak clustering algorithm DP. This is achieved by simply replacing the distance measure with the proposed nearest-neighbour-induced Isolation Similarity in DBSCAN, leaving the rest of the procedure unchanged. A new type of clusters called mass-connected clusters is formally defined. We show that DBSCAN, which detects density-connected clusters, becomes one which detects mass-connected clusters, when the distance measure is replaced with the proposed similarity. We also provide the condition under which mass-connected clusters can be detected, while density-connected clusters cannot.
Kai Ming Ting, Ye Zhu 0002, Vincent Cheng-Siong Lee
AAAI4
2019 A Study of Educational Data Mining: Evidence from a Thai University
abstract
Educational data mining provides a way to predict student academic performance. A psychometric factor like time management is one of the major issues affecting Thai students’ academic performance. Current data sources used to predict students’ performance are limited to the manual collection of data or data from a single unit of study which cannot be generalised to indicate overall academic performance. This study uses an additional data source from a university log file to predict academic performance. It investigates the browsing categories and the Internet access activities of students with respect to their time management during their studies. A single source of data is insufficient to identify those students who are at-risk of failing in their academic studies. Furthermore, there is a paucity of recent empirical studies in this area to provide insights into the relationship between students’ academic performance and their Internet access activities. To contribute to this area of research, we employed two datasets such as web-browsing categories and Internet access activity types to select the best outcomes, and compared different weights in the time and frequency domains. We found that the random forest technique provides the best outcome in these datasets to identify those students who are at-risk of failure. We also found that data from their Internet access activities reveals more accurate outcomes than data from browsing categories alone. The combination of two datasets reveals a better picture of students’ Internet usage and thus identifies students who are academically at-risk of failure. Further work involves collecting more Internet access log file data, analysing it over a longer period and relating the period of data collection with events during the academic year.
Ruangsak Trakunphutthirak, Yen Cheung, Vincent Cheng-Siong Lee
AAAI3
2019 An Application of Convolutional Neural Networks for the Early Detection of Late-onset Neonatal Sepsis
abstract
Preterm newborns are vulnerable and easily infected due to the immature immune system. Late-onset neonatal sepsis occurring 48 hours after birth is a widespread disease among preterm newborns leading to high mortality and morbidity rates. The diagnosis is primarily based on biochemistry test, and the prescribed treatment is to use antibiotics. Risk averse clinicians, often applied overdose to reduce the mortality. A non-invasive method on monitoring vital sign signals deterioration to predict late-onset neonatal sepsis is proposed in this paper. First, we set up collectors within the local networks in Neonatal Intensive Care Units (NICUs) where bedside monitoring machines locate to capture the necessary data. Then they were transformed to images according to specific rules. Finally, a convolutional neural network was built to predict the onset of sepsis. Pilot experiments conducted on data we have collected demonstrated the feasibility of this deep learning model. This method could be incorporated into the current clinical workflow as a decision support system and provide useful information for clinicians.
Yifei Hu, Vincent Cheng-Siong Lee, Kenneth Tan
IJCNN2
2015 DAG: A Model for Privacy Preserving Computation
abstract
Secure multi-party computation (SMC) allows parties to jointly compute a function over their inputs, while keeping every input confidential. It has been extensively applied in privacy-preserving computation, such as privacy-preserving data mining (PPDM), to protect data privacy. However, most SMC-based solutions are ad-hoc. They are proposed for specific applications, and thus cannot be applied to other applications directly. To address this issue, we propose a privacy model DAG (Directed A cyclic Graph) that consists of a set of secure operators (e.g., Multiplication and division). Our DAG model is general -- its operators, if pipelined together, can implement various functions. It is also extendable -- secure operators can be defined to add new features to the model. As an application study, we have applied our DAG to kernel regression. Experiments on datasets of more than 680,000 tuples show that our DAG model is effective and its running time is nearly thrice that of non-privacy setting, where parties directly disclose data.
Sin G. Teo, Jianneng Cao, Vincent Cheng-Siong Lee
ICWS3
2012 Towards objective data selection in bankruptcy prediction
abstract
This paper proposes and tests a methodology for selecting features and test cases with the goal of improving medium term bankruptcy prediction accuracy in large uncontrolled datasets of financial records. We propose a Genetic Programming and Neural Network based objective feature selection methodology to identify key inputs, and then use those inputs to combine multi-level Self-Organising Maps with Spectral Clustering to build clusters. Performing objective feature selection within each of those clusters, this research was able to increase out-of-sample classification accuracy from 71.3% and 69.8% on the Genetic Programming and Neural Network models respectively to 80.0% and 77.3%.
Sverre Gunnersen, Kate Smith-Miles, Vincent Cheng-Siong Lee
IEEE Congress on Evolutionary Computation3
2012 Resilient Identity Crime Detection
abstract
Identity crime is well known, prevalent, and costly; and credit application fraud is a specific case of identity crime. The existing nondata mining detection system of business rules and scorecards, and known fraud matching have limitations. To address these limitations and combat identity crime in real time, this paper proposes a new multilayered detection system complemented with two additional layers: communal detection (CD) and spike detection (SD). CD finds real social relationships to reduce the suspicion score, and is tamper resistant to synthetic social relationships. It is the whitelist-oriented approach on a fixed set of attributes. SD finds spikes in duplicates to increase the suspicion score, and is probe-resistant for attributes. It is the attribute-oriented approach on a variable-size set of attributes. Together, CD and SD can detect more types of attacks, better account for changing legal behavior, and remove the redundant attributes. Experiments were carried out on CD and SD with several million real credit applications. Results on the data support the hypothesis that successful credit application fraud patterns are sudden and exhibit sharp spikes in duplicates. Although this research is specific to credit application fraud detection, the concept of resilience, together with adaptivity and quality data discussed in the paper, are general to the design, implementation, and evaluation of all detection systems.
Clifton Phua, Kate Smith-Miles, Vincent Cheng-Siong Lee, Ross W. Gayler
IEEE Trans. Knowl. Data Eng.3
2011 An Integrated BPM-SOA Framework for Agile Enterprises
Vincent Cheng-Siong Lee
ACIIDS (1)2
2010 Twittering for Earth: A Study on the Impact of Microblogging Activism on Earth Hour 2009 in Australia
Marc Cheong, Vincent Cheng-Siong Lee
ACIIDS (2)2
2010 A Study on Detecting Patterns in Twitter Intra-topic User and Message Clustering
abstract
Timely detection of hidden patterns is the key for the analysis and estimating of driving determinants for mission critical decision making. This study applies Cheong and Lee's “context-aware” content analysis framework to extract latent properties from Twitter messages (tweets). In addition, we incorporate an unsupervised Self-organizing Feature Map (SOM) as a machine learning-based clustering tool that has not been investigated in the context of opinion mining and sentimental analysis using microblogging. Our experimental results reveal the detection of interesting patterns for topics of interest which are latent and cannot be easily detected from the observed tweets without the aid of machine learning tools.
Marc Cheong, Vincent Cheng-Siong Lee
ICPR2
2010 A Competence-Based Collaborative Network: The West Midlands Collaborative Commerce Marketplace
Yen Cheung, Helana Scheepers, Mark Swift, Vincent Cheng-Siong Lee, Jay Bal
PRO-VE4
2010 Privacy-Preserving Gradient-Descent Methods
abstract
Gradient descent is a widely used paradigm for solving many optimization problems. Gradient descent aims to minimize a target function in order to reach a local minimum. In machine learning or data mining, this function corresponds to a decision model that is to be discovered. In this paper, we propose a preliminary formulation of gradient descent with data privacy preservation. We present two approaches-stochastic approach and least square approach-under different assumptions. Four protocols are proposed for the two approaches incorporating various secure building blocks for both horizontally and vertically partitioned data. We conduct experiments to evaluate the scalability of the proposed secure building blocks and the accuracy and efficiency of the protocols for four different scenarios. The excremental results show that the proposed secure building blocks are reasonably scalable and the proposed protocols allow us to determine a better secure protocol for the applications for each scenario.
Shuguo Han, Wee Keong Ng, Vincent Cheng-Siong Lee
IEEE Trans. Knowl. Data Eng.4
2009 An EM-Based Algorithm for Clustering Data Streams in Sliding Windows
Xuan-Hong Dang, Vincent Cheng-Siong Lee, Wee Keong Ng, Arridhana Ciptadi, Kok-Leong Ong
DASFAA2
2009 Incremental and Adaptive Clustering Stream Data over Sliding Window
Xuan-Hong Dang, Vincent Cheng-Siong Lee, Wee Keong Ng, Kok-Leong Ong
DEXA2
2007 Feature Selection Techniques, Company Wealth Assessment and Intra-sectoral Firm Behaviours
Mark B. Barnes, Vincent Cheng-Siong Lee
ICIC (1)2
2007 Privacy-preservation for gradient descent methods
abstract
Gradient descent is a widely used paradigm for solving many optimization problems. Stochastic gradient descent performs a series of iterations to minimize a target function in order to reach a local minimum. In machine learning or data mining, this function corresponds to a decision model that is to be discovered. The gradient descent paradigm underlies many commonly used techniques in data mining and machine learning, such as neural networks, Bayesian networks, genetic algorithms, and simulated annealing. To the best of our knowledge, there has not been any work that extends the notion of privacy preservation or secure multi-party computation to gradient-descent-based techniques. In this paper, we propose a preliminary approach to enable privacy preservation in gradient descent methods in general and demonstrate its feasibility in specific gradient descent methods.
Wee Keong Ng, Shuguo Han, Vincent Cheng-Siong Lee
KDD4
2007 A multivariate neuro-fuzzy system for foreign currency risk management decision making
Vincent Cheng-Siong Lee, Hsiao Tshung Wong
Neurocomputing1
2007 An empirical study of AI-based multiple domain models for selecting attributes that drive company wealth
Mark B. Barnes, Vincent Cheng-Siong Lee
Soft Comput.2
2006 Modeling Supply Chain Complexity Using a Distributed Multi-objective Genetic Algorithm
Khalid Al-Mutawah, Vincent Cheng-Siong Lee, Yen Cheung
ICCSA (1)2
2006 An Intelligent Algorithm for Modeling and Optimizing Dynamic Supply Chains Complexity
Khalid Al-Mutawah, Vincent Cheng-Siong Lee, Yen Cheung
ICIC (1)2
2005 Development and Test of an Artificial-Immune- Abnormal-Trading-Detection System for Financial Markets
Vincent Cheng-Siong Lee, Xingjian Yang
ICIC (1)1
2005 Volatility Transmission Between Stock and Bond Markets: Evidence from US and Australia
Victor Fang, Vincent Cheng-Siong Lee, Yee Choon Lim
IDEAL2
2004 Credit Risks of Interest Rate Swaps: A Comparative Study of CIR and Monte Carlo Simulation Approach
Victor Fang, Vincent Cheng-Siong Lee
IDEAL2
2004 An Algorithm for Artificial Intelligence-Based Model Adaptation to Dynamic Data Distribution
Vincent Cheng-Siong Lee, Alex Tze Hiang Sim
IDEAL1
2004 A Hybrid Fuzzy-Neuro Model for Preference-Based Decision Analysis
Vincent Cheng-Siong Lee, Alex Tze Hiang Sim
IDEAL1
2004 An Alternative Methodology for Mining Seasonal Pattern Using Self-Organizing Map
Denny, Vincent Cheng-Siong Lee
PAKDD2
2003 A Simplified Neuro-Fuzzy Inference System (S-NFIS) Tool for Criteria Matching
Alex Tze Hiang Sim, Vincent Cheng-Siong Lee
HIS2
2003 Testing the Expectations Hypothesis for Interest Rate Term Structure: Some Australian Evidence
Victor Fang, Vincent Cheng-Siong Lee
ICCSA (3)2
2003 A Fuzzy Multi-criteria Decision Model for Information System Security Investment
Vincent Cheng-Siong Lee
IDEAL1