VLDB 2026 Research / reviewers in the wild / expert
Hsing-Kuo Kenneth Pao
dblp:p/HsingKuoKennethPao · also Hsing-Kuo Pao
· DBLP profile ↗
15ranked-venue papers in the field
1as first author
7since 2021 · last 2024
0000-0002-5518-9475ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 13Data Mining & Knowledge Discovery · 1Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Reconsider Time Series Analysis for Insider Threat DetectionabstractInsider threat detection (ITD) presents a significant challenge in cybersecurity, particularly within large and complex organizations. Traditionally, ITD has been overshadowed by the focus of external threats, resulting in less attention and development in this critical area. Conventional ITD approaches often rely heavily on event-driven approaches. On top of that, researchers developed various rule-based methods to conquer the tasks. Based on that, we often ignore the intrinsic temporal relationships that are naturally built in between events that occur in different moments. For instance, we may easily understand events with causality such as one anomalous event followed by another specific event to complete a malicious action; however, may not be aware of events that occur around 9 am every morning during working hours. In our opinion, we attempt to re-consider the temporal behavior to extract the information hidden in cyberspace activities. Specifically, some effective sentence embeddings can assist us in providing informative internal representations to summarize temporal behaviors in the temporal activity sequences to make the right judgment on insider threat detection. In this paper, we propose a novel methodology for insider threat detection that emphasizes temporal relationship modeling on top of already-matured event sequence analysis to effectively catch insider threats. The proposed approach leverages contrastive sentence embeddings to learn users’ intentions in sequences, followed by the deployment of a user-level and event-level Contrastive Learning (euCL) model to incorporate temporal behaviors with user behavior embeddings. To validate the proposed methodology, we conduct extensive analyses and experiments using the publicly available CERT dataset. The results demonstrate the effectiveness and robustness of the proposed method in detecting insider threats and identifying malicious scenarios, highlighting its potential for enhancing cybersecurity measures in complex organizational environments. Chia-Cheng Chen, Hsing-Kuo Kenneth Pao |
IEEE Big Data | 2 |
| 2023 | Density-Based Prototypical Contrastive Learning on Visual RepresentationsabstractSelf-supervised learning is closing the gap between the fully-labeled supervised learning tasks and the scarce-labeled ones rapidly. The effort goes to various recent achievements of contrastive learning methods. However, researchers learn the weakness of contrastive learning. First, contrastive learning is typically an instance discrimination task in practice, which could be targeted on exploiting low-level image differences regardless the understanding of the semantic information. Second, contrastive learning does not encode the dissimilarity between the augmented views from the same image, which may result in potential under-clustering outcomes. Furthermore, contrastive learning treats two representation samples as a negative pair as long as they are from different image instances, which may lead to an over-clustering result. In this paper, we introduce Density-Based Prototypical Contrastive Learning (DBPCL), a self-supervised visual representation learning method that combines the instance discrimination task with density-based clustering to encode the semantic structures of data. By the proposed ProtoXent loss, the proposed method encourages samples to be closer to their assigned prototypes and push their negative prototypes apart at the same time. With the computed prototypes, DBPCL is able to discover both false-positive and false-negative pairs, and then address the issues of under-clustering and over-clustering. With the aforementioned reasons, the proposed DBPCL is superior to the state-of-the-art methods by improving the quality of learned representations. As an example, the proposed method has its prediction accuracy superior to the best group of methods of the same kind on CIFAR-100. The code of the proposed DBPCL is available at https://github.com/YuehLinChung/DBPCL for reference. Yueh-Lin Chung, Hsing-Kuo Kenneth Pao |
IEEE Big Data | 2 |
| 2023 | Self-supervised Federated Learning for Anomaly DetectionabstractFederated learning-based strategy can confirm data privacy, as well as other related ones like split learning and differential privacy. When enjoying secured data sharing, we also worry that the data privacy preserving may harm the model’s overall performance. Given the most recent progress in machine learning and AI methodology, we incorporate self-supervised learning to plug in the federated learning framework and the integrated system can guarantee the model performance and data privacy preservation simultaneously. In the integrated framework, we have different clients to keep their own data, and the data are well separated into the attribute half and the label half, for enhanced privacy, not to mention the additional privacy-preserving skill like differential privacy. Given all the aforementioned components, we can still have the privacy-preserving components equipped model performed superior to or is compatible with the performance without the privacy-preserving components. That is, one client can perform better when adding more information from other clients, without the data sharing in between, and no clients own both the attributes and the corresponding labels. The focused topic is anomaly detection and we pay attention to the imbalanced nature of the data which shows additional challenges to the problem. Given the problem, self-supervised learning is especially useful when obtaining the label information is considered non-trivial. After all, we demonstrate the overall model effectiveness when compared to methods without any federated learning components. Tze-Qian Eng, Hsing-Kuo Kenneth Pao, Chi-Chen Liao |
IEEE Big Data | 2 |
| 2023 | Fake News Detection via Sentiment NeutralizationabstractEverybody knows how fake news or disinformation can have big impact on deciding our daily lives. As people rely on social networks rather than reliable news channels to receive information, we no longer have enough protection mechanisms to help us judge between genuine and fake news. In this work, we propose a sentiment-based analysis for fake news detection. The approach leverages a few of the most recent acclaimed techniques, namely, Large Language Models and Artificial Intelligent Generated Content for sentiment analysis and the subsequent fake news detection. We have ChatGPT to help us rewrite a text for its neutralized version. By comparing the two, before and after the neutralization, we figure out their difference, such as different frequencies of sentiment word usage which can help us for fake news detection. Also, we observe an imbalanced improvement between the fake news and genuine ones. Fake news rather than genuine news in general contains more emotional words, or negative sentiment words to be specific. Taking advantage of such observation can indeed help to further improve the detection result. After all, the evaluation confirms the aforementioned statements and shows the superiority of the proposed method to other state-of-the-art methods, in datasets that own different languages. Hanjuan Huang, Hsuan-Ting Peng, Hsing-Kuo Kenneth Pao |
IEEE Big Data | 3 |
| 2022 | Pseudo slicer on three dimensional brain tumor segmentationabstractA constructive way to assist surgeons before performing brain tumor surgery is by visualizing a three-dimensional (3D) MRI image to determine the brain tumor volume when a pre-operative examination. However, the availability of 3D MRI brain tumors is more limited than 2D, especially for meningioma type. Therefore, we attempt to propose a model to construct a pseudo-3D volume of a meningioma tumor from a few 2D MRI slices. In the proposed model, we first generate pseudo slices for each given 2D MRI slice until satisfying 3D MRI construction from interpolation techniques, called pseudoslicer. Next, we segment the pseudo-3D MRIs to capture the 3D meningioma tumor using 3D volume-to-volume generative adversarial networks (3D-V2V-GAN). That is, we call the proposed model 3D-slicer-V2V-GAN or 3DS-V2V-GAN. In this work, the proposed model was evaluated from 27 meningioma patients at a private university hospital in Indonesia. The proposed model outperformed the SOTA model in terms of the Jaccard Index and dice coefficient. Additionally, we have radiologists to affirm the segmentation results in good quality. Mohammad Iqbal, Fitria Urbach, Hsing-Kuo Kenneth Pao, Anggraini Dwi Sensusiati, Nurul Hidayat, Imam Mukhlash |
IEEE Big Data | 3 |
| 2021 | Active Learning with Numerical Feature AnnotationabstractActive learning is an algorithm that aims to minimize the labeling effort in model training by wisely choosing the key data for labeling in an iterative and alternately querying labeled data and training procedure. While most of studies focus on the query strategy such as selecting the informative instances for labeling in each run, some others consider the query on features, or the query on both of the features and instances simultaneously to build effective model more efficiently than before. A typical strategy of querying features can take care only the categorical features. In this work, we extend the possibility of querying numerical features as well. We consider a Vector Quantization (VQ) technique to transform the numerical features to categorical ones and then all the features, as well as instances can all be evaluated using a unified criterion in the feature/instance selection step. Moreover, we study multivariate VQ as well as univariate VQ to find the best approach for the qunatization before sending the numerical features to categorical ones. After all, we can have active learning that deals with the selection on informative instances and features at the same time. We need to point out that not all numerical features could be appropriate to become categorical ones in the aforementioned transformation. As in our evaluation, some data from medical analysis and others are perfect to be one of such active learning applications. Along this line, we study the difference between applying the proposed feature/instance combined active learning to various data sets. A set of different vector quantization methods shall also be examined to find the best setting for effective active learning experience. Zhiye Fu, Hsing-Kuo Kenneth Pao, Jiabin He |
IEEE BigData | 2 |
| 2021 | A Dropout Style Model Augmentation for Cross Domain Few-Shot LearningabstractWe focus on few-shot learning with cross-domain shift. While most existing few-shot learning assumed similar distributions between the source and target data, we aim to deal with the scenarios when we need to experience domain shift, the scenarios that fit to many real-world applications such as customization from lab-made toys to industrial products. More specifically, we consider the domain shift when we do not have much overlap between the source and the target domains, such as they do not have many overlapped classes or they do not have many common features. In fact, we experience significant improvement from existing work especially when we do not have much overlap between the source and the target domains. We propose a simple yet effective dropout-style method given a model that has been trained based on low-complexity concept from the source domain. The main idea is to sample several sub-networks by dropping neurons (or feature maps) to construct a bunch of models with diverse features for the target domain. Afterward, we choose the most suitable sub-networks to construct the ensemble for the target domain learning. The proposed method requires almost no external storage other than the original model for the source domain learning. As another advantage over other learning methods is the computation in the method is highly parallelizable for the purpose of efficiency. To evaluate the proposed method, we conduct experiments and show that the proposed method can be in cooperated with many metric-based few-shot learning base models, with consistent improvement under various settings. Pei-Cheng Tu, Hsing-Kuo Kenneth Pao |
IEEE BigData | 2 |
| 2020 | Virtual Adversarial Active LearningabstractIn traditional active learning, one of the most well-known strategies is to select the most uncertain data for annotation. By doing that, we acquire as most as we can obtain from the labeling oracle so that the training in the next run can be much more effective than the one from this run once the informative labeled data are added to training. The strategy however, may not be suitable when deep learning become one of the dominant modeling techniques. Deep learning is notorious for its failure to achieve a certain degree of effectiveness under the adversarial environment. Often we see the sparsity in deep learning training space which gives us a result with low confidence. Moreover, to have some adversarial inputs to fool the deep learners, we should have an active learning strategy that can deal with the aforementioned difficulties. We propose a novel Active Learning strategy based on Virtual Adversarial Training (VAT) and the computation of local distributional roughness (LDR). Instead of selecting the data that are closest to the decision boundaries, we select the data that are located in a place with rough enough surface if measured by the posterior probability. The proposed strategy called Virtual Adversarial Active Learning (VAAL) can help us to find the data with rough surface, reshape the model with smooth posterior distribution output thanks to the active learning framework. Moreover, we shall prefer the labeling data that own enough confidence once they are annotated from an oracle. In VAAL, we have the VAT that can not only be used as a regularization term but also helps us effectively and actively choose the valuable samples for active learning labeling. Experiment results show that the proposed VAAL strategy can guide the convolutional networks model converging efficiently on several well-known datasets. Chin-Feng Yu, Hsing-Kuo Kenneth Pao |
IEEE BigData | 2 |
| 2019 | Extracting Explainable Deep Representation for Machine TutoringabstractIn this work, we propose a machine tutoring framework through a modified Adversarial AutoEncoder(AAE) for Interactive Machine Learning(IML). In IML, masters can teach machines for valuable knowledge and skills and machines can transfer such knowledge and skills to novices. Generally in IML, we can discuss transferring knowledge, something to be known and skills, something needs to know and to practice to learn it well from masters to machines and from machines to novices. In this work, we focus on one part of IML techniques, transferring knowledge and skills from machines to novices, called machine tutoring and realize it, in a Taiko-drum playing game platform. In Taiko-drum playing, skill learning is considered more important than knowledge learning. In particular, we adopt Internal Measurement Unit (IMU), attached on novice's forearms to collect the necessary data to understand novice's motions. We assume that the IMU data can help us understand how well novices learn on some particular motions from machines and from masters eventually. On the side, we capture game scores in each novice's play for the evaluation purpose. Four songs with different levels from easy to hard are used in the experiments. When a novice plays the game, we extract the explainable values from latent features in a deep learner for the representation of the novice's play. The machine tutoring is operated as we show the visualization that is built given the latent representation to the novice to suggest on how we may perform a task better and better as we play more. We demonstrate the effectiveness of the proposed machine tutoring framework by using three different deep neural network structures stacked on the Autoencoder included in the framework including MLP, CNN, and LSTM. The results show that AAE-regression CNN (AAER-CNN) performs the best on latent information representation when the song level is from easy to normal, and AAER-LSTM also performs the best when the song level is difficult on two tasks namely score prediction subject classification. Afterwards, we illustrate how the representation can help on machine tutoring. Ming-Chen Wang, Vahid Golderzahi, Hsing-Kuo Kenneth Pao |
IEEE BigData | 3 |
| 2017 | Big active learningabstractActive learning is a common strategy to deal with large-scale data with limited labeling effort. In each iteration of active learning, a query is ready for oracle to answer such as what the label is for a given unlabeled data. Given the method, we can request the labels only for those data that are essential and save the labeling effort from oracle. We focus on pool-based active learning where a set of unlabeled data is selected for querying in each run of active learning. To apply pool-based active learning to massive high-dimensional data, especially when the unlabeled data set is much larger than the labeled set, we propose the APRAL and MLP strategies so that the computation for active learning can be dramatically reduced while keeping the model power more or less the same. In APRAL, we avoid unnecessary data re-ranking given an unlabeled data selection criteria. To further improve the efficiency, with MLP, we organize the unlabeled data in a multi-layer pool based on a dimensionality reduction technique and the most valuable data to know their label information are more likely to store in the top layers. Given the APRAL and MLP strategies, the active learning computation time is reduced by about 83% if compared to the traditional active learning ones; at the same time, the model effectiveness remains. Er-Chen Huang, Hsing-Kuo Kenneth Pao, Yuh-Jye Lee |
IEEE BigData | 2 |
| 2016 | Compressed learning for time series classificationabstractThe time series classification has been studied for various applications in the last decades. In the time series classification problem, we decide the class information based on a small piece of the time series inputs. In general, the approaches to time series classification can be categorized into three types, distance-based, model-based, and feature-based approaches. In this research, we focus on the feature-based methods, which represent time series as a set of characterized values. It is quite often the case that features generated by existing representation techniques are not transparent to domain experts and the feature that are selected for classification are not completely interpretable. We aim to propose a novel time series representation, called Envelope to solve the problem. The proposed supervised feature extraction method transforms time series into simple 1/0/-1 values. A heuristic is introduced to determine the most appropriate representation which includes the features that are the best to discriminate data of different labels. Moreover, this new representation enjoys the characteristic of sparsity which is an essential property when we need to apply compressed sensing techniques. With this advantage, we can benefit from high transmission efficiency, the reduction of required storage and model complexity. We conduct a series of tests on various benchmark time series data to show the effectiveness of the proposed method. Other than the classification effectiveness, we demonstrate how to visualize the similarity between time series of the same and different kinds from the proposed Envelope method. Yuh-Jye Lee, Hsing-Kuo Kenneth Pao, Shueh-Han Shih, Jing-Yao Lin, Xin-Rong Chen |
IEEE BigData | 2 |
| 2014 | Efficient traffic speed forecasting based on massive heterogenous historical dataabstractDrivers dream of foreseeing traffic condition to enjoy efficient driving experience at all times. Given the historical patterns for different locations and different time, people should be able to guess the possible traffic speed in a near future moment. What is difficult and interesting for this task is that we need to filter the useful data that could help us for the next moment traffic speed prediction from a massive amount of historical data. On the other hand, the traffic condition could be highly dynamic and we can only give a reliable traffic prediction by using the most updated model for prediction. This implies that frequent retraining is necessary. To conquer the task, we propose a lazy learning approach for traffic speed prediction given massive historical data. The approach integrates the kNN and Gaussian process regression for efficient and robust traffic speed prediction. kNN can help us to select the most informative data for Gaussian process Regression using a big data framework. Thanks for the most recent progress of big data research, the processing of massive data for prediction in close to real time has become possible now compared to any time in the past. We aim at using a Hadoop framework for the prediction given heterogeneous data including traffic data such as speed, flow, occupancy, and weather data. Xing-Yu Chen, Hsing-Kuo Kenneth Pao, Yuh-Jye Lee |
IEEE BigData | 2 |
| 2013 | Malicious URL filtering - A big data applicationabstractMalicious URLs have become a channel for Internet criminal activities such as drive-by-download, spamming and phishing. Applications for the detection of malicious URLs are accurate but slow (because they need to download the content or query some Internet host information). In this paper we present a novel lightweight filter based only on the URL string itself to use before existing processing methods. We run experiments on a large dataset and demonstrate a 75% reduction in workload size while retaining at least 90% of malicious URLs. Existing methods do not scale well with the hundreds of millions of URLs encountered every day as the problem is a heavily-imbalanced, large-scale binary classification problem. Our proposed method is able to handle nearly two million URLs in less than five minutes. We generate two filtering models by using lexical features and descriptive features, and then combine the filtering results. The on-line learning algorithms are applied here not only for dealing with large-scale data sets but also for fitting the very short lifetime characteristics of malicious URLs. Our filter can significantly reduce the volume of URL queries on which further analysis needs to be performed, saving both computing time and bandwidth used for content retrieval. Min-Sheng Lin, Chien-Yi Chiu, Yuh-Jye Lee, Hsing-Kuo Kenneth Pao |
IEEE BigData | 4 |
| 2012 | Malicious URL Detection Based on Kolmogorov Complexity EstimationabstractMalicious URL detection has drawn a significant research attention in recent years. It is helpful if we can simply use the URL string to make precursory judgment about how dangerous a Web site is. By doing that, we can save efforts on the Web site content analysis and bandwidth for content retrieval. We propose a detection method that is based on an estimation of the conditional Kolmogorov complexity of URL strings. To overcome the incomputability of Kolmogorov complexity, we adopt a compression method for its approximation, called conditional Kolmogorov measure. As a single significant feature for detection, we can achieve a decent performance that can not be achieved by any other single feature that we know. Moreover, the proposed Kolmogorov measure can work together with other features for a successful detection. The experiment has been conducted using a private dataset from a commercial company which can collect more than one million unclassified URLs in a typical hour. On average, the proposed measure can process such hourly data in less than a few minutes. Hsing-Kuo Kenneth Pao, Yan-Lin Chou, Yuh-Jye Lee |
Web Intelligence | 1 |
| 2011 | Unifying Guilt-by-Association Approaches: Theorems and Fast Algorithms
Danai Koutra, Tai-You Ke, U Kang, Polo Chau, Hsing-Kuo Kenneth Pao, Christos Faloutsos |
ECML/PKDD (2) | 5 |