Ming-Chang Lee

dblp:18/3398 · DBLP profile ↗
← Back
34ranked-venue papers
25as first author
15since 2021 · last 2025
0000-0003-2484-4366ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 6 first-author · 4 since 2021Systems, architecture and hardware · 7 · 6 first-author · 1 since 2021Security and privacy · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 RePAD3: Advanced Lightweight Adaptive Anomaly Detection for Univariate Time Series of Any Pattern
Ming-Chang Lee, Jia-Chun Lin, Sokratis K. Katsikas
ICAART (2)1
2025 PacketZapper: A Scalable and Automated Platform for IoT Traffic Collection and Analysis
Mathias Fredrik Hedberg, Jia-Chun Lin, Ming-Chang Lee
IoTBDS3
2024 Impact of Recurrent Neural Networks and Deep Learning Frameworks on Real-Time Lightweight Time Series Anomaly Detection
Ming-Chang Lee, Jia-Chun Lin, Sokratis K. Katsikas
ICICS (1)1
2024 Investigating the Privacy Risk of Using Robot Vacuum Cleaners in Smart Environments
Benjamin Ulsmåg, Jia-Chun Lin, Ming-Chang Lee
ICICS (1)3
2024 Evaluation of K-Means Time Series Clustering Based on Z-Normalization and NP-Free
Ming-Chang Lee, Jia-Chun Lin, Volker Stolz
ICPRAM1
2024 IoTective: Automated Penetration Testing for Smart Home Environments
abstract
Den raske utbredelsen av Internet of Things (IoT) har reist betydelige bekymringer angående sikkerheten og personvernet til sammenkoblede smarte miljøer. Denne masteroppgaven presenterer IoTective, et automatisert verktøy for penetrasjonstesting designet for å vurdere sikkerhetstilstanden til IoT-enheter og -systemer. IoTective benytter seg av ulike skanningsteknikker, inkludert Wi-Fi, Bluetooth og ZigBee, for å identifisere sårbarheter, oppdage enheter og samle verdifull informasjon for analyse. Verktøyets brukervennlige grensesnitt og automatiseringsfunksjoner minimerer behovet for manuell innblanding, noe som gjør det tilgjengelig både for nybegynnere og erfarne sikkerhetsanalytikere. Gjennom vår Proof-of-Concept (PoC) demonstrerer IoTective sin effektivitet i å identifisere sårbarheter i et bredt spekter av smarte hjemmeenheter og -systemer. Verktøyets regelmessige oppdateringer, komplementaritet med manuell testing og funksjoner for prioritering bidrar til suksessen som et omfattende verktøy for vurdering av IoT-sikkerhet. Fremtidig arbeid inkluderer utvidelse av testing for spesifikke produsenter for å inkorporere støtte for flere leverandørers APIs, slik at det blir mulig med grundigere analyser og målrettet sårbarhetsdeteksjon. Med den kontinuerlige utviklingen av IoT-teknologier, bidrar IoTectives bidrag til feltet for vurdering av IoT-sikkerhet\nog dets potensial for videre forbedring til å gjøre det til en verdifull ressurs for sikring av IoT-miljøer.
Kevin Nordnes, Jia-Chun Lin, Ming-Chang Lee, Victor Chang 0001
IoTBDS3
2024 UoCAD: An Unsupervised Online Contextual Anomaly Detection Approach for Multivariate Time Series from Smart Homes
abstract
In the context of time series data, a contextual anomaly is considered an event or action that causes a deviation in the data values from the norm. This deviation may appear normal if we do not consider the timestamp associated with it. Detecting contextual anomalies in real-world time series data poses a challenge because it often requires domain knowledge and an understanding of the surrounding context. In this paper, we propose UoCAD, an online contextual anomaly detection approach for multivariate time series data. UoCAD employs a sliding window method to (re)train a Bi-LSTM model in an online manner. UoCAD uses the model to predict the upcoming value for each variable/feature and calculates the model's prediction error value for each feature. To adapt to minor pattern changes, UoCAD employs a double-check approach without immediately triggering an anomaly notification. Two criteria, individual and majority, are explored for anomaly detection. The individual criterion identifies an anomaly if any feature is detected as anomalous, while the majority criterion triggers an anomaly when more than half of the features are identified as anomalous. We evaluate UoCAD using an air quality dataset containing a contextual anomaly. The results show UoCAD's effectiveness in detecting the contextual anomaly across different sliding window sizes but with varying false positives and detection time consumption.
Aafan Ahmad Toor, Jia-Chun Lin, Ming-Chang Lee, Ernst Gunnar Gran
IoTBDS3
2024 GAD: A Real-Time Gait Anomaly Detection System with Online Adaptive Learning
Ming-Chang Lee, Jia-Chun Lin, Sokratis K. Katsikas
SEC1
2024 Exploring the effects of RNNs and deep learning frameworks on real-time, lightweight, adaptive time series anomaly detection
abstract
Summary Real‐time, lightweight, adaptive time series anomaly detection is increasingly critical in cybersecurity, industrial control, finance, healthcare, and many other domains due to its capability to promptly process time series and detect anomalies without requiring extensive computation resources. While numerous anomaly detection approaches have emerged recently, they generally employ a single type of recurrent neural network (RNN) and are implemented using a single type of deep learning framework. The impacts of using various RNN types across different deep learning frameworks on the performance of these approaches remain unclear due to a lack of comprehensive evaluations. In this article, we aim to investigate the impact of different RNN variants and deep learning frameworks on real‐time, lightweight, and adaptive time series anomaly detection. We reviewed several state‐of‐the‐art anomaly detection approaches and implemented a representative approach using several RNN variants supported by three popular deep learning frameworks. A thorough evaluation was conducted to analyze the detection accuracy, time efficiency, and resource consumption of each implementation using four real‐world, open‐source time series datasets. The results show that RNN variants and deep learning frameworks have a significant impact. Therefore, it is crucial to carefully select appropriate RNN variants and deep learning frameworks for the implementation.
Ming-Chang Lee, Jia-Chun Lin, Sokratis K. Katsikas
Concurr. Comput. Pract. Exp.1
2023 NP-Free: A Real-Time Normalization-free and Parameter-tuning-free Representation Approach for Open-ended Time Series
abstract
To help analyze time series in data mining applications, many time series representation approaches have been proposed to convert a raw time series into another series for representing the original time series. However, existing approaches are not designed for open-ended time series (which is a sequence of data points being continuously collected at a fixed interval without any length limit) because these approaches need to know the total length of the target time series in advance and preprocess the entire time series using a normalization method. Furthermore, many representation approaches require users to configure and tune some parameters beforehand in order to achieve satisfactory representation results. In this paper, we propose NP-Free, a real-time Normalization-free and Parameter-tuning-free representation approach for open-ended time series. Without needing to use any normalization method or tune any parameter, NP-Free can generate a representation for a raw time series on the fly by converting each data point of the time series into a root-mean-square error (RMSE) value based on Long Short-Term Memory (LSTM) and a Look-Back and Predict-Forward strategy. To demonstrate the capability of NP-Free in representing time series, we conducted several experiments based on real-world open-source time series datasets. We also evaluated the time consumption of NP-Free in generating representations.
Ming-Chang Lee, Jia-Chun Lin, Volker Stolz
COMPSAC1
2023 Impact of Deep Learning Libraries on Online Adaptive Lightweight Time Series Anomaly Detection
Ming-Chang Lee, Jia-Chun Lin
ICSOFT1
2023 RoLA: A Real-Time Online Lightweight Anomaly Detection System for Multivariate Time Series
Ming-Chang Lee, Jia-Chun Lin
ICSOFT1
2023 RePAD2: Real-Time Lightweight Adaptive Anomaly Detection for Open-Ended Time Series
Ming-Chang Lee, Jia-Chun Lin
IoTBDS1
2021 How Far Should We Look Back to Achieve Effective Real-Time Time-Series Anomaly Detection?
Ming-Chang Lee, Jia-Chun Lin, Ernst Gunnar Gran
AINA (1)1
2021 SALAD: Self-Adaptive Lightweight Anomaly Detection for Real-time Recurrent Time Series
abstract
Providing a lightweight self-adaptive approach that does not need offline training in advance and meanwhile is able to detect anomalies in real time could be highly beneficial. Such an approach could be immediately applied and deployed on any commodity machine to provide timely anomaly alerts. To facilitate such an approach, this paper introduces SALAD, which is a Self-Adaptive Lightweight Anomaly Detection approach based on a special type of recurrent neural networks called Long Short-Term Memory (LSTM). Instead of using offline training, SALAD converts a target time series into a series of average absolute relative error (AARE) values on the fly and predicts an AARE value for every upcoming data point based on short-term historical AARE values. If the difference between a calculated AARE value and its corresponding forecast AARE value is higher than a self-adaptive detection threshold, the corresponding data point is considered anomalous. Otherwise, the data point is considered normal. Experiments based on a real-world time series dataset demonstrates that SALAD outperforms five other state-of-the-art anomaly detection approaches in terms of detection accuracy. In addition, the results also show that SALAD is lightweight and can be deployed on a commodity machine.
Ming-Chang Lee, Jia-Chun Lin, Ernst Gunnar Gran
COMPSAC1
2020 DALC: Distributed Automatic LSTM Customization for Fine-Grained Traffic Speed Prediction
Ming-Chang Lee, Jia-Chun Lin
AINA1
2020 RePAD: Real-Time Proactive Anomaly Detection for Time Series
Ming-Chang Lee, Jia-Chun Lin, Ernst Gunnar Gran
AINA1
2020 ReRe: A Lightweight Real-Time Ready-to-Go Anomaly Detection Approach for Time Series
abstract
Anomaly detection is an active research topic in many different fields such as intrusion detection, network monitoring, system health monitoring, IoT healthcare, etc. However, many existing anomaly detection approaches require either human intervention or domain knowledge, and may suffer from high computation complexity, consequently hindering their applicability in real-world scenarios. Therefore, a lightweight and ready-to-go approach that is able to detect anomalies in real-time is highly sought-after. Such an approach could be easily and immediately applied to perform time series anomaly detection on any commodity machine. The approach could provide timely anomaly alerts and by that enable appropriate countermeasures to be undertaken as early as possible. With these goals in mind, this paper introduces ReRe, which is a Real-time Ready-to-go proactive Anomaly Detection algorithm for streaming time series. ReRe employs two lightweight Long Short-Term Memory (LSTM) models to predict and jointly determine whether or not an upcoming data point is anomalous based on short-term historical data points and two long-term self-adaptive thresholds. Our experiment based on real-world time-series datasets demonstrates the good performance of ReRe in real-time anomaly detection without requiring human intervention or domain knowledge.
Ming-Chang Lee, Jia-Chun Lin, Ernst Gunner Gan
COMPSAC1
2020 Distributed Fine-Grained Traffic Speed Prediction for Large-Scale Transportation Networks Based on Automatic LSTM Customization and Sharing
Ming-Chang Lee, Jia-Chun Lin, Ernst Gunnar Gran
Euro-Par1
2019 Adaptive Write Interference Management with Efficient Mapping for Shingled Recording Disks
Ming-Chang Lee, Li-Pin Chang, Sung-Ming Wu, Wei-Shang Yui
ICCD1
2018 EasyChoose: A Continuous Feature Extraction and Review Highlighting Scheme on Hadoop YARN
abstract
Today the Internet offers a massive amount of reviews and user experiences about a variety of products from different manufacturers, ranging from smartphones, automobiles, and home appliances to Internet services such as hotel booking and airplane booking. For a careful customer it is time-consuming to make good purchasing decisions due to a variety of similar products, lots of reviews for each product, and distributed reviews on the Internet. To alleviate this situation, this paper proposes EasyChoose, which is a distributed scheme based on Hadoop YARN to continuously collect product reviews from the Internet, extract representative product features based on previous customers' reviews, and highlight the main point of the reviews. In this paper, we use online hotel booking as an example to demonstrate the effectiveness of EasyChoose. The results show that EasyChoose is able to automatically extract representative product features and highlight reviews without losing the original meanings. Furthermore, EasyChoose is able to continuously provide such service to keep up with changes in recent customers' reviews.
Ming-Chang Lee, Jia-Chun Lin, Olaf Owe
AINA1
2018 Modeling and Simulation of Spark Streaming
abstract
As more and more devices connect to Internet of Things, unbounded streams of data will be generated, which have to be processed "on the fly" in order to trigger automated actions and deliver real-time services. Spark Streaming is a popular realtime stream processing framework. To make efficient use of Spark Streaming and achieve stable stream processing, it requires a careful interplay between different parameter configurations. Mistakes may lead to significant resource overprovisioning and bad performance. To alleviate such issues, this paper develops an executable and configurable model named SSP (stands for Spark Streaming Processing) to model and simulate Spark Streaming. SSP is written in ABS, which is a formal, executable, and object-oriented language for modeling distributed systems by means of concurrent object groups. SSP allows users to rapidly evaluate and compare different parameter configurations without deploying their applications on a cluster/cloud. The simulation results show that SSP is able to mimic Spark Streaming in different scenarios.
Jia-Chun Lin, Ming-Chang Lee, Ingrid Chieh Yu, Einar Broch Johnsen
AINA2
2016 ABS-YARN: A Formal Framework for Modeling Hadoop YARN Clusters
Jia-Chun Lin, Ingrid Chieh Yu, Einar Broch Johnsen, Ming-Chang Lee
FASE4
2016 Performance evaluation of job schedulers on Hadoop YARN
abstract
Summary To solve the limitation of Hadoop on scalability, resource sharing, and application support, the open‐source community proposes the next generation of Hadoop's compute platform called Yet Another Resource Negotiator (YARN) by separating resource management functions from the programming model. This separation enables various application types to run on YARN in parallel. To achieve fair resource sharing and high resource utilization, YARN provides the capacity scheduler and the fair scheduler. However, the performance impacts of the two schedulers are not clear when mixed applications run on a YARN cluster. Therefore, in this paper, we study four scheduling‐policy combinations (SPCs for short) derived from the two schedulers and then evaluate the four SPCs in extensive scenarios, which consider not only four application types, but also three different queue structures for organizing applications. The experimental results enable YARN managers to comprehend the influences of different SPCs and different queue structures on mixed applications. The results also help them to select a proper SPC and an appropriate queue structure to achieve better application execution performance. Copyright © 2016 John Wiley & Sons, Ltd.
Jia-Chun Lin, Ming-Chang Lee
Concurr. Comput. Pract. Exp.2
2016 Hybrid Job-Driven Scheduling for Virtual MapReduce Clusters
abstract
It is cost-efficient for a tenant with a limited budget to establish a virtual MapReduce cluster by renting multiple virtual private servers (VPSs) from a VPS provider. To provide an appropriate scheduling scheme for this type of computing environment, we propose in this paper a hybrid job-driven scheduling scheme (JoSS for short) from a tenant's perspective. JoSS provides not only job-level scheduling, but also map-task level scheduling and reduce-task level scheduling. JoSS classifies MapReduce jobs based on job scale and job type and designs an appropriate scheduling policy to schedule each class of jobs. The goal is to improve data locality for both map tasks and reduce tasks, avoid job starvation, and improve job execution performance. Two variations of JoSS are further introduced to separately achieve a better map-data locality and a faster task assignment. We conduct extensive experiments to evaluate and compare the two variations with current scheduling algorithms supported by Hadoop. The results show that the two variations outperform the other tested algorithms in terms of map-data locality, reduce-data locality, and network overhead without incurring significant overhead. In addition, the two variations are separately suitable for different MapReduce-workload scenarios and provide the best job performance among all tested algorithms.
Ming-Chang Lee, Jia-Chun Lin, Ramin Yahyapour
IEEE Trans. Parallel Distributed Syst.1
2015 ReMBF: A Reliable Multicast Brute-Force Co-allocation Scheme for Multi-user Data Grids
abstract
In this paper we propose a novel co-allocation scheme, called a Reliable Multicast Brute-Force co-allocation scheme (ReMBF for short), which employs a reliable multicast (RM for short) technique with the Brute-Force (BF for short) scheme to accelerate data retrieval and delivery, and reliably transmit data to its users for data grids. Several types of data access patterns, including Zipf-like, geometric, and uniform distributions, are utilized to model user access behaviors and evaluate the performance of ReMBF. The simulation results demonstrate that ReMBF can efficiently deliver a bulk of data in a shorter time period compared with two state-of-the-art schemes.
Ming-Chang Lee, Fang-Yie Leu, Ying-Ping Chen
COMPSAC1
2015 Pareto-based cache replacement for YouTube
Ming-Chang Lee, Fang-Yie Leu, Ying-Ping Chen
World Wide Web1
2014 Cache Replacement Algorithms for YouTube
abstract
In recent years, many social network systems like, YouTube, Facebook, Twitter, etc. have been a part of our everyday life. Among these systems, YouTube which plays video programs of different interesting themes for users has been one of the most attractive ones. Basically, when the space of Memcached in YouTube is full, the Least Recently Used algorithm (LRU for short) is employed to evict a least recently watched video. However, the LRU, due to its own property, may cause more miss counts of Memcached for YouTube, consequently increasing the network bandwidth and energy consumptions. To solve these problems, in this paper we proposed two cache replacement algorithms, the Pareto Least Recently Used algorithm (PLRU for short) and Pareto Least Frequently Used algorithm (PLFU for short), in which videos are classified into different popularity categories, and those videos in the top 10% and 20% of the top two popular categories of YouTube, based on Pareto principle, are then chosen to serve users' requests without removing them from Memcached. The simulation results show that the PLFU algorithm can significantly reduce miss counts of Memcached compared with those when the LRU and Least Frequently Used algorithm (LFU for short) are employed, thus achieving great performance improvements for the top two popular categories of videos. Also, when a large space of Memcached is available, the miss counts of the PLRU are higher than those of the LRU. But when comparing them in a small available Memcached space, the miss counts of the LRU on the contrary is higher than those of the PLRU.
Ming-Chang Lee, Fang-Yie Leu, Ying-Ping Chen
AINA1
2014 Scheduling MapReduce tasks on virtual MapReduce clusters from a tenant's perspective
abstract
Renting a set of virtual private servers (VPSs for short) from a VPS provider to establish a virtual MapReduce cluster is cost-efficient for a company/organization. To shorten job turnaround time and keep data locality as high as possible in this type of environment, this paper proposes a Best-Fit Task Scheduling scheme (BFTS for short) from a tenant's perspective. BFTS schedules each map task to a VPS that can finish the task earlier than the other VPSs by predicting and comparing the time required by every VPS to retrieve the map-input data, execute the map task, and become idle in an online manner. Furthermore, BFTS schedules each reduce task to a VPS that is close to most VPSs that execute the related map tasks. We conduct extensive experiments to compare BFTS with several scheduling algorithms employed by Hadoop. The experimental results show that BFTS is better than the other tested algorithms in terms of map-data locality, reduce-data locality, and job turnaround time. The overhead incurred by BFTS is also evaluated, which is inevitable but acceptable compared with the other algorithms.
Jia-Chun Lin, Ming-Chang Lee, Ramin Yahyapour
IEEE BigData2
2012 PFRF: An adaptive data replication algorithm based on star-topology data grids
Ming-Chang Lee, Fang-Yie Leu, Ying-Ping Chen
Future Gener. Comput. Syst.1
2009 A Divide-and-Conquer Strategy and PVM Computation Environment for the Matrix Multiplication
Ming-Chang Lee
ICA3PP1
2009 Enterprise Financial Status Synthetic Evaluation Based on Fuzzy Rough Set Theory
Ming-Chang Lee, Jui-Fang Chang
KES-AMSTA1
2005 Statistical Data Analysis for Software Metrics Validation
Ming-Chang Lee
KES (4)1
2004 An object-oriented analysis method for customer relationship management information systems
Jyhjong Lin, Ming-Chang Lee
Inf. Softw. Technol.2