VLDB 2026 Research / reviewers in the wild / expert
Ying-Feng Hsu
dblp:30/10267
· DBLP profile ↗
15ranked-venue papers
9as first author
7since 2021 · last 2025
0000-0002-5335-4510ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scalability and Versatility of Energy-Aware Workload Allocation Optimizer (WAO) on KubernetesabstractThe Workload Allocation Optimizer (WAO), an energy-aware method for workload scheduling and load balancing, along with its implementation on Kubernetes, achieves substantial energy savings in data center operations without requiring hardware modifications or infrastructure changes. This paper evaluates the scalability and versatility of WAO through experimental analysis conducted in an actual data center environment. To facilitate energy-aware workload placement, power consumption models were developed for each CPU frequency governor, enabling precise estimation of allocation impacts. Furthermore, by integrating caching mechanisms, the WAO Scheduler maintains scheduling performance comparable to the default Kubernetes scheduler; combined with data center experiments and regression-based projections, this confirms its practicality even at the Kubernetes scalability limit of 5,000 Nodes. Experimental results further demonstrate WAO's effectiveness across diverse environments, ranging from small server rooms to large-scale data centers with heterogeneous hardware. In addition, by incorporating both power consumption and processing time into the scheduling criteria, WAO consistently delivers substantial energy savings without compromising computational performance. These findings establish WAO as a practical and effective energy-saving solution for Kubernetes-based data center environments, and suggest its effectiveness for compute-intensive tasks, such as AI model training and inference. Shunsuke Ise, Chizuko Mizumoto, Ying-Feng Hsu, Kazuhiro Matsuda, Morito Matsuoka |
JCC | 3 |
| 2023 | Historical Redundant Process Data Recovery based on Genetic AlgorithmabstractIn many other domains, the data may have such issues as conflicts, duplicates, and missing values, and it must be cleaned before utilization. For instance, historical data reports on numerous events for overlapping time intervals may have data conflicts caused by database redundancy. These conflicts can prevent researchers from obtaining the correct answers from data. In this paper, we investigated redundant process data recovery (RPDR) approaches to recover the individual values of time intervals from redundant process data. There are three major contributions to this study. First, we explore RPDR approaches from the areas of statistical analysis, evolutionary computation (Genetic Algorithm), and probabilistic value estimations (Bayesian method). Second, we explore the applicability of the proposed RPDR algorithms to the case of having additional information from redundant data. Third, we utilize the concept of optimal CD (Conflict Degree) further to reduce data aggregation error in the integrated historical database. In general, it is challenging to estimate an accurate individual value within a given time interval. With the help of optimal CD, our experimental results demonstrate the high efficiency of the proposed approach by using the genetic algorithm to minimize the misestimation of those sequence values in those individual time spans. Ying-Feng Hsu |
COMPSAC | 1 |
| 2023 | Comprehensive Analysis of Dieting Apps: Effectiveness, Design, and Frequency UsageabstractDieting mobile health (mHealth) apps are becoming more common in response to the prevalence of obesity. Obesity and overweight have become more widespread in recent years in correspondence to unbalanced diets, lack of exercise, and unhealthy eating habits, all of which can contribute to an increased chance of developing chronic diseases. However, despite the rising numbers of obesity and overweight cases, diet consultations costs have not decreased; instead, they have increased. Therefore, due to the affordability and accessibility issues of healthcare services, dieting apps have gained popularity as an alternative option. Dieting apps cost less than traditional healthcare, more convenient, easily understandable, and generally provides similar feedback to that of a nutritionist, making them more accessible to many individuals. This systematic review evaluates dieting apps from three aspects: effectiveness, design, and frequency of usage. This paper analyzes 11 studies, revealing that dieting apps are successful in certain aspects but lacked quality in other cases, suggesting that there is room for improvement. The effectiveness of dieting apps was found to be mediocre, with some studies indicating neutral or partial success in weight control or reduction. Users generally liked the design aesthetics, but felt that feedback could be more personalized, while the frequency of usage was mainly determined by the app's popularity. Allison Hsu, Ying-Feng Hsu |
COMPSAC | 2 |
| 2022 | Comparison of Teenagers' Writing Using Word Clouds and Analyzing EnginesabstractWorldwide, there are writing competitions for teenagers. However, usually, it is only the winning articles that are presented and remarked off. In order to understand the gap between winning and unselected articles, we compare the selected-winning articles and unselected articles to find the writing quality difference. We do this by using word clouds and various analyzing engines. On Write the World (WtW), while selected-winning pieces are published on WtW Reviews, the users (teenagers between the age of 13-18) can also self-share their articles with the public on the WtW platform. Our finding shows that to get unbiased results, the analyzing engines require the raw text to be inputted, rather than using only the 50 most frequently appeared words from the word cloud. Additionally, there is a notable difference in writing quality between selected-winning articles and self-shared articles. Allison Hsu, Ying-Feng Hsu |
IEEE Big Data | 2 |
| 2022 | Computer Education of the Primary Years Programme Exhibition at International Baccalaureate Schools During the COVID-19 PandemicabstractThis paper investigates computer education and learning of the Primary Years Programme Exhibition (PYPX) project held at the end of the transitional grade of elementary to middle school at International Baccalaureate (IB) Schools during the COVID-19 pandemic. During the COVID-19 pandemic, most of the schools' study and work activities were moved online, which brought significant challenges to elementary school students who were new to computers. On the other hand, it was also a great chance for students to learn more computer skills by digitally completing their PYPX project. We researched 33 11-year-old students who completed the PYPX projects for the pandemic years 2020 and 2021. Allison Hsu, Ying-Feng Hsu |
COMPSAC | 2 |
| 2022 | Real Network DDoS Pattern Analysis and DetectionabstractThe exponential growth of computer networks and network applications has also increased the incidence of cyberattacks. Network intrusion detection systems (NIDS) are essential for organizations to ensure the safety and security of their communications and information. Many machine learning approaches have been proposed for DoS/DDoS detection and mitigation. However, most of these are based on using synthetic benchmark datasets, which may not thoroughly reflect real network DDoS attack patterns and fail to consider their DDoS detection applicability in real network scenarios. In this paper, we explore machine learning-based DDoS detection in realtime campus network environments. This study includes three major contributions. First, we show the differences in network traffic and DDoS attack patterns between real networks and synthetic benchmark datasets from various aspects. Second, we propose a system architecture for DDoS detection using machine learning models and ensure that they are applicable to large-scale networks on a realtime basis. Third, to demonstrate the feasibility of our approach, we explore seven machine learning models and evaluate their performance from the aspects of DDoS detection capability, system processing time, and early DDoS detection. Our evaluation was conducted based on our real campus traffic log data, which contains about 400 million session logs per day. Ying-Feng Hsu, Araki Ryusei, Morito Matsuoka |
COMPSAC | 1 |
| 2022 | Early Detection of Campus Network DDoS Attacks using Predictive ModelsabstractDDoS attacks are one of the most threatening types of cyberattacks in the growing number of Internet-based services. In late 2016, a DDoS attack by IoT botnets of up to 1.5 Tbps caused many U.S. websites, including Twitter and Facebook, to become inaccessible. In addition, DDoS attacks are increasing every year, and the volume of attacks is expected to double in 2023, as compared to 2018. To protect services from DDoS attacks, much research has been done on IDS and has discussed methods with higher and more accurate detection. However, many studies use public benchmark datasets rather than real network traffic data, and as a result, their practicality is unknown. Threshold detection is already in place on our campus firewalls, but threshold detection cannot detect attacks until they actually come. In order to detect attacks before they actually come, we propose a system that uses machine learning to detect signs of attacks. In this study, we examined machine learning models for early detection of DDoS attacks using actual logs generated by servers at our campus, which contains about 400 million daily session logs. To ensure the feasibility and applicability of our proposed approach, we tested seven different machine learning methods, including GBDT, which has received much attention recently. A sliding window was also used for feature creation to improve the accuracy of predictive detection. Araki Ryusei, Ying-Feng Hsu, Morito Matsuoka |
GLOBECOM | 2 |
| 2020 | High-Performance Virus Detection System by using Deep LearningabstractMetagenomic shotgun sequencing enables us to explore diverse DNA sequences from viruses, bacteria, and eukaryotic microbes in complex samples. As the continuous advancement of sequencing technology generates a massive amount of sequencing data, its overall computational complexity has become a major challenge for traditional database sequence comparison methods. Studies have shown that deep learningoriented methods have been widely adopted to solve many classification problems, including those in the bioinformatics field, and have demonstrated this method's accuracy and efficiency for analyzing large-scale datasets. The aim of this study attempts to investigate how deep learning (LSTM model) can be used to learn sequential genome patterns through virus detection from metagenomic data. This study provides three major contributions. First, we provide the background and steps for the task of DNA sequencing classification from data collection, preprocessing, and normalization. Second, we analyze the effect of sequence length on LSTM classification accuracy and split the raw sequencing data to proper subsequences to improve the outcome of virus detection. Third, to enhance both the classification accuracy and processing speed, we introduce the concept of discrimination function that enables prediction results for multiple subsequences results and accelerated these processes through GPU parallel computing. Two case studies of HCV and influenza detection were conducted to elaborate upon the accuracy and computational efficiency of our proposed approach. Our test result showed that the proposed LSTM model obtained similar pathogen detection accuracy to the conventional BLAST method with a speed that was about 36 times faster. Ying-Feng Hsu, Makiko Ito, Takumi Maruyama, Morito Matsuoka, Nicolas Jung, Yuki Matsumoto, Daisuke Motooka, Shota Nakamura |
CEC | 1 |
| 2019 | Toward an Online Network Intrusion Detection System Based on Ensemble LearningabstractWith information technology growing and rapidly increasing, ubiquitous networking technology generates a massive amount of data and is integrated into our daily life. Network intrusion detection systems (NIDS) are essential for organizations to ensure the safety and security of their communication and information. In general, there are two types of NIDS: signature-based (SNIDS) and anomaly-based (ANDIS). Most modern NIDS solutions are signature-based techniques, which require a routine signature update and cannot detect unknown types of attacks. However, ANDIS has been extensively studied and is considered a better alternative to NIDS. In this paper, we present a stacked ensemble learning based ANIDS that consists of autoencoder (AE), support vector machine (SVM), and random forest (RF) models. To show the overall applicability of our approach, we demonstrate our work through two well-known NIDS benchmark datasets: NSL-KDD and UNSW-NB15 and a real campus network log, which includes about 300 million daily records. We compare our method to three different machine learning classical models and two other reported study results. Our test result implies that our proposed method can also limit both false positive and false negative predictions. Ying-Feng Hsu, Zhenyu He 0002, Yuya Tarutani, Morito Matsuoka |
CLOUD | 1 |
| 2019 | Real-Time Workload Allocation Optimizer for Computing Systems by Using Deep LearningabstractWith the increase of the Internet of Things (IoT) business, the number of edge computing systems are rapidly increasing. Reducing the power consumption of these computing systems has become a social issue. For that purpose, we proposed and demonstrated the power consumption reduction method by the optimal task assignment technology. Specifically, for sequential real-time jobs, we proposed a workload allocation optimizer (WAO) to minimize the power consumption of computing systems. This assignment algorithm achieved a 13% power reduction compared to the worst job assignment. Hayato Kuwahara, Ying-Feng Hsu, Kazuhiro Matsuda, Morito Matsuoka |
CLOUD | 2 |
| 2019 | Deep Learning Approach for Pathogen Detection Through Shotgun Metagenomics Sequence Classification
Ying-Feng Hsu, Makiko Ito, Takumi Maruyama, Morito Matsuoka, Nicolas Jung, Yuki Matsumoto, Daisuke Motooka, Shota Nakamura |
AIME | 1 |
| 2019 | Toward a Workload Allocation Optimizer for Power Saving in Data CentersabstractThe number and scale of data centers are both rapidly increasing due to a continuously growing demand for cloud computing services from many areas. Cloud computing infrastructure relies on a massive amount of HPC servers to process millions of tasks and consumes an enormous amount of power. The implementation of advanced task allocation technology provides a solution for energy efficiency and has therefore become an essential goal for data centers. In this paper, we propose a novel CPU-intensive workload allocation optimizer (WAO) for the task of power saving within data centers. There are three major contributions to this research. First, a data center monitoring module, which continually reports the latest status of the data center and stores operational data. Second, we propose an accurate and efficient server power prediction model for all servers in the HPC clusters. Third, we provide an optimal task assignment engine that evaluates and assigns tasks to the most appropriate server to facilitate minimal power consumption. Our experimental results show that our proposed WAO can obtain about 29.6% power savings and 26% more productivity in a real data center. Ying-Feng Hsu, Hayato Kuwahara, Kazuhiro Matsuda, Morito Matsuoka |
IC2E | 1 |
| 2018 | A Novel Automated Cloud Storage Tiering System through Hot-Cold Data ClassificationabstractWith information technology growing and rapidly increasing ICT equipment, a massive amount of data have been generated and stored in the cloud. However, the majority of them are infrequently accessed data. Data temperature describes the frequency of data access. Hot storage is dedicated to storing frequently accessed data while cold storage is designed for infrequently accessed data. To cope with the issue of exponential data growth in cloud, it is essential to allocate different categories of data to proper storage media. In this research, we propose an automated cloud storage tiering system for the task of data temperature prediction through hot-cold data classification and data migration, which allocates the predicted categorized data to the corresponding storage media. There are three major contributions in this paper. Firstly, the feasibility: by successfully predicting the infrequent access data and moving them to the cold storage, we obtain significant cost savings. Secondly, the reliability: while having the benefit of storage-cost saving, our proposed system also ensures customers satisfaction by enhancing the ratio of data access through hot storage. Lastly, the flexibility: operational strategy varies from cloud storage service providers. Our system characterizes different scenarios and provides the customized optimal solution. Ying-Feng Hsu, Ryo Irie, Shuuichirou Murata, Morito Matsuoka |
IEEE CLOUD | 1 |
| 2018 | A High-Performance Sequence Analysis Engine for Shotgun Metagenomics through GPU AccelerationabstractWith the continual growth of low-cost and high-throughput DNA sequence technology, the scale and amount of next-generation sequencing (NGS) datasets are continually increasing in many genomics research areas. Shotgun metagenomics sequencing provides comprehensive information on microorganisms, based on complex samples of the ecosystem. Due to challenges of its scale and computational complexity, efficient sequence processing and analyzing tools are needed. In this paper, we propose a novel high-performance shotgun metagenomics sequence analysis engine for the task of sequence comparison. It includes two major components. First, a customized shifting database, which is optimized from any existing DNA sequence dataset. Second, a high-performance sequence computation algorithm that utilizes the customized shifting reference database and accelerates GPU parallel computing. We elaborated upon the efficiency and computational complexity of our proposed approach in an HPC server, which has eight Nvidia Tesla P100 GPUs. We also conducted a case study to detect viral sequences from patients' blood samples. Our experimental result shows that we obtain similar accuracy to the conventional BLAST method, but with a computational speed that is about twenty times faster. Ying-Feng Hsu, Morito Matsuoka, Nicolas Jung, Yuki Matsumoto, Daisuke Motooka, Shota Nakamura |
BIBE | 1 |
| 2018 | Self-Aware Workload Forecasting in Data Center Power PredictionabstractThe number and scale of data centers are rapidly increasing, due to the growing demand for cloud computing services. Cloud computing infrastructure relies on a massive amount of information and communication technology (ICT) equipment, which consume an enormous amount of power. Power saving and energy optimization have therefore become essential goals for data centers. An enhanced data center energy management system (DEMS) provides a solution for data center power consumption based on its coordinative control of ICT equipment. An efficient power prediction model is essential for such a DEMS because it facilitates the proactive control of ICT equipment and reduces the total power consumption. In this paper, we propose a novel self-aware workload forecasting (SAWF) framework for total power consumption prediction in data centers. It includes three major components. First, there is a feature selection module, which evaluates the importance of variables from all ICT equipment in a data center and dynamically selects the most relevant variables for data input. Second, we propose an accurate and efficient neural network model to forecast future total power consumption. Third, we provide an online error monitoring and model updating module that continuously monitors prediction errors and updates the model when necessary. Ying-Feng Hsu, Kazuhiro Matsuda, Morito Matsuoka |
CCGrid | 1 |