EDBT 2026 Demo / reviewers in the wild / expert
Jia-Chun Lin
dblp:98/139
· DBLP profile ↗
35ranked-venue papers
8as first author
15since 2021 · last 2025
0000-0003-3374-8536ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 1 first-author · 4 since 2021Security and privacy · 8 · 7 since 2021Systems, architecture and hardware · 7 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RePAD3: Advanced Lightweight Adaptive Anomaly Detection for Univariate Time Series of Any Pattern
Ming-Chang Lee, Jia-Chun Lin, Sokratis K. Katsikas |
ICAART (2) | 2 |
| 2025 | PacketZapper: A Scalable and Automated Platform for IoT Traffic Collection and Analysis
Mathias Fredrik Hedberg, Jia-Chun Lin, Ming-Chang Lee |
IoTBDS | 2 |
| 2024 | Impact of Recurrent Neural Networks and Deep Learning Frameworks on Real-Time Lightweight Time Series Anomaly Detection
Ming-Chang Lee, Jia-Chun Lin, Sokratis K. Katsikas |
ICICS (1) | 2 |
| 2024 | Investigating the Privacy Risk of Using Robot Vacuum Cleaners in Smart Environments
Benjamin Ulsmåg, Jia-Chun Lin, Ming-Chang Lee |
ICICS (1) | 2 |
| 2024 | Evaluation of K-Means Time Series Clustering Based on Z-Normalization and NP-Free
Ming-Chang Lee, Jia-Chun Lin, Volker Stolz |
ICPRAM | 2 |
| 2024 | IoTective: Automated Penetration Testing for Smart Home EnvironmentsabstractDen raske utbredelsen av Internet of Things (IoT) har reist betydelige bekymringer angående sikkerheten og personvernet til sammenkoblede smarte miljøer. Denne masteroppgaven presenterer IoTective, et automatisert verktøy for penetrasjonstesting designet for å vurdere sikkerhetstilstanden til IoT-enheter og -systemer. IoTective benytter seg av ulike skanningsteknikker, inkludert Wi-Fi, Bluetooth og ZigBee, for å identifisere sårbarheter, oppdage enheter og samle verdifull informasjon for analyse. Verktøyets brukervennlige grensesnitt og automatiseringsfunksjoner minimerer behovet for manuell innblanding, noe som gjør det tilgjengelig både for nybegynnere og erfarne sikkerhetsanalytikere. Gjennom vår Proof-of-Concept (PoC) demonstrerer IoTective sin effektivitet i å identifisere sårbarheter i et bredt spekter av smarte hjemmeenheter og -systemer. Verktøyets regelmessige oppdateringer, komplementaritet med manuell testing og funksjoner for prioritering bidrar til suksessen som et omfattende verktøy for vurdering av IoT-sikkerhet. Fremtidig arbeid inkluderer utvidelse av testing for spesifikke produsenter for å inkorporere støtte for flere leverandørers APIs, slik at det blir mulig med grundigere analyser og målrettet sårbarhetsdeteksjon. Med den kontinuerlige utviklingen av IoT-teknologier, bidrar IoTectives bidrag til feltet for vurdering av IoT-sikkerhet\nog dets potensial for videre forbedring til å gjøre det til en verdifull ressurs for sikring av IoT-miljøer. Kevin Nordnes, Jia-Chun Lin, Ming-Chang Lee, Victor Chang 0001 |
IoTBDS | 2 |
| 2024 | UoCAD: An Unsupervised Online Contextual Anomaly Detection Approach for Multivariate Time Series from Smart HomesabstractIn the context of time series data, a contextual anomaly is considered an event or action that causes a deviation in the data values from the norm. This deviation may appear normal if we do not consider the timestamp associated with it. Detecting contextual anomalies in real-world time series data poses a challenge because it often requires domain knowledge and an understanding of the surrounding context. In this paper, we propose UoCAD, an online contextual anomaly detection approach for multivariate time series data. UoCAD employs a sliding window method to (re)train a Bi-LSTM model in an online manner. UoCAD uses the model to predict the upcoming value for each variable/feature and calculates the model's prediction error value for each feature. To adapt to minor pattern changes, UoCAD employs a double-check approach without immediately triggering an anomaly notification. Two criteria, individual and majority, are explored for anomaly detection. The individual criterion identifies an anomaly if any feature is detected as anomalous, while the majority criterion triggers an anomaly when more than half of the features are identified as anomalous. We evaluate UoCAD using an air quality dataset containing a contextual anomaly. The results show UoCAD's effectiveness in detecting the contextual anomaly across different sliding window sizes but with varying false positives and detection time consumption. Aafan Ahmad Toor, Jia-Chun Lin, Ming-Chang Lee, Ernst Gunnar Gran |
IoTBDS | 2 |
| 2024 | GAD: A Real-Time Gait Anomaly Detection System with Online Adaptive Learning
Ming-Chang Lee, Jia-Chun Lin, Sokratis K. Katsikas |
SEC | 2 |
| 2024 | Exploring the effects of RNNs and deep learning frameworks on real-time, lightweight, adaptive time series anomaly detectionabstractSummary Real‐time, lightweight, adaptive time series anomaly detection is increasingly critical in cybersecurity, industrial control, finance, healthcare, and many other domains due to its capability to promptly process time series and detect anomalies without requiring extensive computation resources. While numerous anomaly detection approaches have emerged recently, they generally employ a single type of recurrent neural network (RNN) and are implemented using a single type of deep learning framework. The impacts of using various RNN types across different deep learning frameworks on the performance of these approaches remain unclear due to a lack of comprehensive evaluations. In this article, we aim to investigate the impact of different RNN variants and deep learning frameworks on real‐time, lightweight, and adaptive time series anomaly detection. We reviewed several state‐of‐the‐art anomaly detection approaches and implemented a representative approach using several RNN variants supported by three popular deep learning frameworks. A thorough evaluation was conducted to analyze the detection accuracy, time efficiency, and resource consumption of each implementation using four real‐world, open‐source time series datasets. The results show that RNN variants and deep learning frameworks have a significant impact. Therefore, it is crucial to carefully select appropriate RNN variants and deep learning frameworks for the implementation. Ming-Chang Lee, Jia-Chun Lin, Sokratis K. Katsikas |
Concurr. Comput. Pract. Exp. | 2 |
| 2023 | NP-Free: A Real-Time Normalization-free and Parameter-tuning-free Representation Approach for Open-ended Time SeriesabstractTo help analyze time series in data mining applications, many time series representation approaches have been proposed to convert a raw time series into another series for representing the original time series. However, existing approaches are not designed for open-ended time series (which is a sequence of data points being continuously collected at a fixed interval without any length limit) because these approaches need to know the total length of the target time series in advance and preprocess the entire time series using a normalization method. Furthermore, many representation approaches require users to configure and tune some parameters beforehand in order to achieve satisfactory representation results. In this paper, we propose NP-Free, a real-time Normalization-free and Parameter-tuning-free representation approach for open-ended time series. Without needing to use any normalization method or tune any parameter, NP-Free can generate a representation for a raw time series on the fly by converting each data point of the time series into a root-mean-square error (RMSE) value based on Long Short-Term Memory (LSTM) and a Look-Back and Predict-Forward strategy. To demonstrate the capability of NP-Free in representing time series, we conducted several experiments based on real-world open-source time series datasets. We also evaluated the time consumption of NP-Free in generating representations. Ming-Chang Lee, Jia-Chun Lin, Volker Stolz |
COMPSAC | 2 |
| 2023 | Impact of Deep Learning Libraries on Online Adaptive Lightweight Time Series Anomaly Detection
Ming-Chang Lee, Jia-Chun Lin |
ICSOFT | 2 |
| 2023 | RoLA: A Real-Time Online Lightweight Anomaly Detection System for Multivariate Time Series
Ming-Chang Lee, Jia-Chun Lin |
ICSOFT | 2 |
| 2023 | RePAD2: Real-Time Lightweight Adaptive Anomaly Detection for Open-Ended Time Series
Ming-Chang Lee, Jia-Chun Lin |
IoTBDS | 2 |
| 2021 | How Far Should We Look Back to Achieve Effective Real-Time Time-Series Anomaly Detection?
Ming-Chang Lee, Jia-Chun Lin, Ernst Gunnar Gran |
AINA (1) | 2 |
| 2021 | SALAD: Self-Adaptive Lightweight Anomaly Detection for Real-time Recurrent Time SeriesabstractProviding a lightweight self-adaptive approach that does not need offline training in advance and meanwhile is able to detect anomalies in real time could be highly beneficial. Such an approach could be immediately applied and deployed on any commodity machine to provide timely anomaly alerts. To facilitate such an approach, this paper introduces SALAD, which is a Self-Adaptive Lightweight Anomaly Detection approach based on a special type of recurrent neural networks called Long Short-Term Memory (LSTM). Instead of using offline training, SALAD converts a target time series into a series of average absolute relative error (AARE) values on the fly and predicts an AARE value for every upcoming data point based on short-term historical AARE values. If the difference between a calculated AARE value and its corresponding forecast AARE value is higher than a self-adaptive detection threshold, the corresponding data point is considered anomalous. Otherwise, the data point is considered normal. Experiments based on a real-world time series dataset demonstrates that SALAD outperforms five other state-of-the-art anomaly detection approaches in terms of detection accuracy. In addition, the results also show that SALAD is lightweight and can be deployed on a commodity machine. Ming-Chang Lee, Jia-Chun Lin, Ernst Gunnar Gran |
COMPSAC | 2 |
| 2020 | DALC: Distributed Automatic LSTM Customization for Fine-Grained Traffic Speed Prediction
Ming-Chang Lee, Jia-Chun Lin |
AINA | 2 |
| 2020 | RePAD: Real-Time Proactive Anomaly Detection for Time Series
Ming-Chang Lee, Jia-Chun Lin, Ernst Gunnar Gran |
AINA | 2 |
| 2020 | ReRe: A Lightweight Real-Time Ready-to-Go Anomaly Detection Approach for Time SeriesabstractAnomaly detection is an active research topic in many different fields such as intrusion detection, network monitoring, system health monitoring, IoT healthcare, etc. However, many existing anomaly detection approaches require either human intervention or domain knowledge, and may suffer from high computation complexity, consequently hindering their applicability in real-world scenarios. Therefore, a lightweight and ready-to-go approach that is able to detect anomalies in real-time is highly sought-after. Such an approach could be easily and immediately applied to perform time series anomaly detection on any commodity machine. The approach could provide timely anomaly alerts and by that enable appropriate countermeasures to be undertaken as early as possible. With these goals in mind, this paper introduces ReRe, which is a Real-time Ready-to-go proactive Anomaly Detection algorithm for streaming time series. ReRe employs two lightweight Long Short-Term Memory (LSTM) models to predict and jointly determine whether or not an upcoming data point is anomalous based on short-term historical data points and two long-term self-adaptive thresholds. Our experiment based on real-world time-series datasets demonstrates the good performance of ReRe in real-time anomaly detection without requiring human intervention or domain knowledge. Ming-Chang Lee, Jia-Chun Lin, Ernst Gunner Gan |
COMPSAC | 2 |
| 2020 | Distributed Fine-Grained Traffic Speed Prediction for Large-Scale Transportation Networks Based on Automatic LSTM Customization and Sharing
Ming-Chang Lee, Jia-Chun Lin, Ernst Gunnar Gran |
Euro-Par | 2 |
| 2019 | A Flexible Framework for Program Evolution and VerificationabstractWe propose a flexible framework for modeling of distributed systems, supporting evolution by means of unrestricted modifications in such systems, and with support of verification and re-verification. We focus on the setting of concurrent and object-oriented programs, and consider a core high-level modeling language supporting active, concurrent objects. We show that our framework can deal with verification of software changes that are not possible to verify in comparable frameworks. We demonstrate the approach by variations over a simple example. Olaf Owe, Jia-Chun Lin, Elahe Fazeldehkordi |
MODELSWARD | 2 |
| 2018 | EasyChoose: A Continuous Feature Extraction and Review Highlighting Scheme on Hadoop YARNabstractToday the Internet offers a massive amount of reviews and user experiences about a variety of products from different manufacturers, ranging from smartphones, automobiles, and home appliances to Internet services such as hotel booking and airplane booking. For a careful customer it is time-consuming to make good purchasing decisions due to a variety of similar products, lots of reviews for each product, and distributed reviews on the Internet. To alleviate this situation, this paper proposes EasyChoose, which is a distributed scheme based on Hadoop YARN to continuously collect product reviews from the Internet, extract representative product features based on previous customers' reviews, and highlight the main point of the reviews. In this paper, we use online hotel booking as an example to demonstrate the effectiveness of EasyChoose. The results show that EasyChoose is able to automatically extract representative product features and highlight reviews without losing the original meanings. Furthermore, EasyChoose is able to continuously provide such service to keep up with changes in recent customers' reviews. Ming-Chang Lee, Jia-Chun Lin, Olaf Owe |
AINA | 2 |
| 2018 | Modeling and Simulation of Spark StreamingabstractAs more and more devices connect to Internet of Things, unbounded streams of data will be generated, which have to be processed "on the fly" in order to trigger automated actions and deliver real-time services. Spark Streaming is a popular realtime stream processing framework. To make efficient use of Spark Streaming and achieve stable stream processing, it requires a careful interplay between different parameter configurations. Mistakes may lead to significant resource overprovisioning and bad performance. To alleviate such issues, this paper develops an executable and configurable model named SSP (stands for Spark Streaming Processing) to model and simulate Spark Streaming. SSP is written in ABS, which is a formal, executable, and object-oriented language for modeling distributed systems by means of concurrent object groups. SSP allows users to rapidly evaluate and compare different parameter configurations without deploying their applications on a cluster/cloud. The simulation results show that SSP is able to mimic Spark Streaming in different scenarios. Jia-Chun Lin, Ming-Chang Lee, Ingrid Chieh Yu, Einar Broch Johnsen |
AINA | 1 |
| 2016 | ABS-YARN: A Formal Framework for Modeling Hadoop YARN Clusters
Jia-Chun Lin, Ingrid Chieh Yu, Einar Broch Johnsen, Ming-Chang Lee |
FASE | 1 |
| 2016 | Comparing AWS Deployments Using Model-Based Predictions
Einar Broch Johnsen, Jia-Chun Lin, Ingrid Chieh Yu |
ISoLA (2) | 2 |
| 2016 | Impacts of Task Re-Execution Policy on MapReduce JobsabstractMapReduce is a popular distributed programming framework for large-scale data processing. To prevent MapReduce jobs from being interrupted by node failures that occur frequently in a MapReduce cluster consisting of a set of commodity machines/nodes, the most well-known MapReduce implementation, i.e. Hadoop, adopts a task re-execution policy (TR policy). When a map/reduce task of a job crashes, the TR policy assigns another node to reperform the task. However, the impact of the TR policy on MapReduce jobs in terms of reliability, job turnaround time (JTT) and energy consumption are not clear, particularly when jobs have different features, e.g. different filtering percentages, different input-data sizes, and different numbers of reduce tasks. In this paper, we formally analyze the job completion reliability (JCR) of a job based on Poisson distributions, and then derive the expected JTT and job energy consumption (JEC) based on the universal generation function. Extensive analyses are further conducted to explore the impact of the TR policy on JCR, JTT and JEC of jobs with different features. The results show that employing the TR policy can dramatically improve JCR for a large MapReduce job. Moreover, if the JCR of a job is highly improved by the TR policy, the expected JTT and JEC will not be significantly prolonged and increased, respectively. Jia-Chun Lin, Fang-Yie Leu, Ying-Ping Chen |
Comput. J. | 1 |
| 2016 | Performance evaluation of job schedulers on Hadoop YARNabstractSummary To solve the limitation of Hadoop on scalability, resource sharing, and application support, the open‐source community proposes the next generation of Hadoop's compute platform called Yet Another Resource Negotiator (YARN) by separating resource management functions from the programming model. This separation enables various application types to run on YARN in parallel. To achieve fair resource sharing and high resource utilization, YARN provides the capacity scheduler and the fair scheduler. However, the performance impacts of the two schedulers are not clear when mixed applications run on a YARN cluster. Therefore, in this paper, we study four scheduling‐policy combinations (SPCs for short) derived from the two schedulers and then evaluate the four SPCs in extensive scenarios, which consider not only four application types, but also three different queue structures for organizing applications. The experimental results enable YARN managers to comprehend the influences of different SPCs and different queue structures on mixed applications. The results also help them to select a proper SPC and an appropriate queue structure to achieve better application execution performance. Copyright © 2016 John Wiley & Sons, Ltd. Jia-Chun Lin, Ming-Chang Lee |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Hybrid Job-Driven Scheduling for Virtual MapReduce ClustersabstractIt is cost-efficient for a tenant with a limited budget to establish a virtual MapReduce cluster by renting multiple virtual private servers (VPSs) from a VPS provider. To provide an appropriate scheduling scheme for this type of computing environment, we propose in this paper a hybrid job-driven scheduling scheme (JoSS for short) from a tenant's perspective. JoSS provides not only job-level scheduling, but also map-task level scheduling and reduce-task level scheduling. JoSS classifies MapReduce jobs based on job scale and job type and designs an appropriate scheduling policy to schedule each class of jobs. The goal is to improve data locality for both map tasks and reduce tasks, avoid job starvation, and improve job execution performance. Two variations of JoSS are further introduced to separately achieve a better map-data locality and a faster task assignment. We conduct extensive experiments to evaluate and compare the two variations with current scheduling algorithms supported by Hadoop. The results show that the two variations outperform the other tested algorithms in terms of map-data locality, reduce-data locality, and network overhead without incurring significant overhead. In addition, the two variations are separately suitable for different MapReduce-workload scenarios and provide the best job performance among all tested algorithms. Ming-Chang Lee, Jia-Chun Lin, Ramin Yahyapour |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Analyzing job completion reliability and job energy consumption for a heterogeneous MapReduce cluster under different intermediate-data replication policies
Jia-Chun Lin, Fang-Yie Leu, Ying-Ping Chen |
J. Supercomput. | 1 |
| 2015 | Impact of MapReduce Policies on Job Completion Reliability and Job Energy ConsumptionabstractRecently, MapReduce has been widely employed by many companies/organizations to tackle data-intensive problems over a large-scale MapReduce cluster. To solve machine/node failure which is inevitable in a MapReduce cluster, MapReduce employs several policies, such as input-data replication and intermediate-data replication policies. To speed up job execution, MapReduce allows reduce tasks to early fetch their required intermediate data. However, the impact of these policy combinations on the job completion reliability (JCR for short) and job energy consumption (JEC for short) of a MapReduce cluster was not clear, where JCR is the reliability with which a MapReduce job can be completed by the cluster, whereas JEC is the energy consumed by the cluster to complete the job. Therefore, in this study, we analyze the JCR and JEC of a MapReduce cluster on four policy combinations (POCs for short) derived from two typical intermediate-data replication policies and two typical reduce-task assignment policies. The four POCs are further compared in extensive scenarios, which not only consider jobs at different scales with various parameters, but also give a MapReduce cluster two extreme parallel execution capabilities and diverse bandwidths. The analytical results enable MapReduce managers to comprehend how these POCs impact the JCR and JEC of a cluster and then select an appropriate POC based on the characteristics of their own MapReduce jobs and clusters. Jia-Chun Lin, Fang-Yie Leu, Ying-Ping Chen |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Impact of MapReduce Task Re-execution Policy on Job Completion Reliability and Job Completion TimeabstractMapReduce has been a worldwide accepted framework for solving data-intensive applications. To prevent MapReduce jobs from being interrupted by node failures which occur frequently in a large-scale MapReduce cluster, current MapReduce implementations, e.g., Hadoop, employ a task re-execution policy (TR policy for short) for MapReduce jobs, i.e., when a map/reduce task of a job fails due to node failure, this policy reperforms the task on another node. However, the impact of the TR policy on job completion reliability and job completion time have not been studied from a theoretical viewpoint, especially when the job is given different characteristics, e.g., different input data sizes, different numbers of reduce tasks, and different intermediate data sizes. In this study, we derive the job completion reliability (JCR for short) of a MapReduce job based on Poisson distributions and analyze the expected job completion time (JCT for short) based on the universal generation function. We use nine settings of task re-execution factor (TR factor for short) to explore the impact of the TR policy on the JCR and JCT of jobs. The results show that the TR policy can effectively improve JCR without significantly prolonging JCT. But there is no single TR factor with which all jobs can achieve a high JCR. Jia-Chun Lin, Fang-Yie Leu, Ying-Ping Chen, Waqaas Munawar |
AINA | 1 |
| 2014 | Scheduling MapReduce tasks on virtual MapReduce clusters from a tenant's perspectiveabstractRenting a set of virtual private servers (VPSs for short) from a VPS provider to establish a virtual MapReduce cluster is cost-efficient for a company/organization. To shorten job turnaround time and keep data locality as high as possible in this type of environment, this paper proposes a Best-Fit Task Scheduling scheme (BFTS for short) from a tenant's perspective. BFTS schedules each map task to a VPS that can finish the task earlier than the other VPSs by predicting and comparing the time required by every VPS to retrieve the map-input data, execute the map task, and become idle in an online manner. Furthermore, BFTS schedules each reduce task to a VPS that is close to most VPSs that execute the related map tasks. We conduct extensive experiments to compare BFTS with several scheduling algorithms employed by Hadoop. The experimental results show that BFTS is better than the other tested algorithms in terms of map-data locality, reduce-data locality, and job turnaround time. The overhead incurred by BFTS is also evaluated, which is inevitable but acceptable compared with the other algorithms. Jia-Chun Lin, Ming-Chang Lee, Ramin Yahyapour |
IEEE BigData | 1 |
| 2008 | Detection workload in a dynamic grid-based intrusion detection environment
Fang-Yie Leu, Ming-Chang Li, Jia-Chun Lin, Chao-Tung Yang |
J. Parallel Distributed Comput. | 3 |
| 2007 | An Enhanced DGIDE Platform for Intrusion Detection
Fang-Yie Leu, Fuu-Cheng Jiang, Ming-Chang Li, Jia-Chun Lin |
ATC | 4 |
| 2005 | Integrating Grid with Intrusion DetectionabstractIn recent years, distributed denial-of-service (DDoS) and denial-of-service (DoS) are the most dreadful network threats. Single-node IDS often suffers from losing its detection effectiveness and capability when processing enormous network traffic. To solve the drawbacks, we propose grid-based IDS, called grid intrusion detection system (GIDS), which uses grid computing resources to detect intrusion packets. For balancing detection load, score subtraction approach (SSA) and score addition approach (SAA) are deployed. Furthermore, to effectively detect intrusions, a two-phase packet detection process is proposed. The first phase detects logical and momentary attacks. Chronic attacks are detected in the second phase. Experiments are also performed and the results show that GIDS is truly an outstanding system in detecting attacks. Fang-Yie Leu, Jia-Chun Lin, Ming-Chang Li, Chao-Tung Yang, Po-Chi Shih |
AINA | 2 |
| 2005 | A Performance-Based Grid Intrusion Detection SystemabstractDistributed denial-of-service (DDoS) and denial-of-service (DoS) are the most dreadful network threats in recent years. In this paper, we propose a grid-based IDS, called performance-based grid intrusion detection system (PGIDS), which exploits grid's abundant computing resources to detect enormous intrusion packets and improve the drawbacks of traditional IDSs which suffer from losing their detection effectiveness and capability when processing massive network traffic. For balancing detection load and accelerating the performance of allocating detection node (DN), we use exponential average to predict network traffic and then assign the collected actual traffic to the most suitable DN. In addition, score subtraction algorithm (SSA) and score addition algorithm (SAA) are deployed to update and reflect the current performance of a DN. PGIDS detects not only DoS/DDoS attacks but also logical attacks. Experimental results show that PGIDS is truly an outstanding system in detecting attacks. Fang-Yie Leu, Jia-Chun Lin, Ming-Chang Li, Chao-Tung Yang |
COMPSAC (1) | 2 |