EDBT 2026 Demo / reviewers in the wild / expert
Edward Chuah
dblp:78/1824
· DBLP profile ↗
16ranked-venue papers
12as first author
7since 2021 · last 2025
0000-0002-8659-1803ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 6 first-author · 3 since 2021Security and privacy · 7 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deep learning-based prediction of major page faults in cluster systems
Edward Chuah, Arshad Jhumka, Sai Narasimhamurthy, Aladdin Ayesh |
CCF Trans. High Perform. Comput. | 1 |
| 2025 | Deep learning-based prediction of reflection attacks using NetFlow data
Edward Chuah, Arshad Jhumka, Aladdin Ayesh |
Comput. Secur. | 1 |
| 2025 | A systematic literature review of log-correlation tools for cyberattack detection and prediction in large networks
Edward Chuah, Harsha K. Kalutarage, Kasim Tasdemir, Atnafu Abrham, Carsten Maple |
J. Inf. Secur. Appl. | 1 |
| 2024 | An empirical study of reflection attacks using NetFlow dataabstractAbstract Reflection attacks are one of the most intimidating threats organizations face. A reflection attack is a special type of distributed denial-of-service attack that amplifies the amount of malicious traffic by using reflectors and hides the identity of the attacker. Reflection attacks are known to be one of the most common causes of service disruption in large networks. Large networks perform extensive logging of NetFlow data, and parsing this data is an advocated basis for identifying network attacks. We conduct a comprehensive analysis of NetFlow data containing 1.7 billion NetFlow records and identified reflection attacks on the network time protocol (NTP) and NetBIOS servers. We set up three regression models including the Ridge, Elastic Net and LASSO. To the best of our knowledge, there is no work that studied different regression models to understand patterns of reflection attacks in a large network. In this paper, we (a) propose an approach for identifying correlations of reflection attacks, and (b) evaluate the three regression models on real NetFlow data. Our results show that (a) reflection attacks on the NTP servers are not correlated, (b) reflection attacks on the NetBIOS servers are not correlated, (c) the traffic generated by those reflection attacks did not overwhelm the NTP and NetBIOS servers, and (d) the dwell times of reflection attacks on the NTP and NetBIOS servers are too small for predicting reflection attacks on these servers. Our work on reflection attacks identification highlights recommendations that could facilitate better handling of reflection attacks in large networks. Edward Chuah, Neeraj Suri |
Cybersecur. | 1 |
| 2023 | An empirical study of major page faults for failure diagnosis in cluster systems
Edward Chuah, Arshad Jhumka, Sai Narasimhamurthy |
J. Supercomput. | 1 |
| 2021 | Sentiment Analysis based Error Detection for Large-Scale SystemsabstractToday's large-scale systems such as High Performance Computing (HPC) Systems are designed/utilized towards exascale computing, inevitably decreasing its reliability due to the increasing design complexity. HPC systems conduct extensive logging of their execution behaviour. In this paper, we leverage the inherent meaning behind the log messages and propose a novel sentiment analysis-based approach for the error detection in large-scale systems, by automatically mining the sentiments in the log messages. Our contributions are four-fold. (1) We develop a machine learning (ML) based approach to automatically build a sentiment lexicon, based on the system log message templates. (2) Using the sentiment lexicon, we develop an algorithm to detect system errors. (3) We develop an algorithm to identify the nodes and components with erroneous behaviors, based on sentiment polarity scores. (4) We evaluate our solution vs. other state-of-the-art machine/deep learning algorithms based on three representative supercomputers' system logs. Experiments show that our error detection algorithm can identify error messages with an average MCC score and f-score of 91% and 96% respectively, while state of the art ML/deep learning model (LSTM) obtains only 67% and 84%. To the best of our knowledge, this is the first work leveraging the sentiments embedded in log entries of large-scale systems for system health analysis. Khalid Ayedh Alharthi, Arshad Jhumka, Sheng Di, Franck Cappello, Edward Chuah |
DSN | 5 |
| 2021 | Challenges in Identifying Network Attacks Using Netflow DataabstractLarge networks often encounter attacks that can affect the network availability. While multiple techniques exist to detect network attacks, a comprehensive understanding of how an attack occurs considering the various layers and components of the network software stack, can be an important element to help improve network security. By performing correlation analysis on contemporary unlabeled Netflow data, this paper conducts a comprehensive study of network flow events to identify communication patterns that may precede an attack, thereby providing potentially useful attack signatures to network administrators. Our work shows that, surprisingly, the Netflow data is not strongly correlated to network attacks. We observe that while spoof requests trigger reflection attacks, only a small percentage of the network packets are associated with the attack. Furthermore, lead time enhancements are feasible for reflection attacks that show long dwell times. Our study on network event correlations highlights empirical observations that could facilitate better attack handling in large networks. Edward Chuah, Neeraj Suri, Arshad Jhumka, Samantha Alt |
NCA | 1 |
| 2019 | Towards comprehensive dependability-driven resource use and message log-analysis for HPC systems diagnosis
Edward Chuah, Arshad Jhumka, Samantha Alt, Daniel Balouek-Thomert, James C. Browne, Manish Parashar |
J. Parallel Distributed Comput. | 1 |
| 2017 | Enabling Dependability-Driven Resource Use and Message Log-Analysis for Cluster System DiagnosisabstractRecent work have used both failure logs and resource use data separately (and together) to detect system failure-inducing errors and to diagnose system failures. System failure occurs as a result of error propagation and the (unsuccessful) execution of error recovery mechanisms. Knowledge of error propagation patterns and unsuccessful error recovery is important for more accurate and detailed failure diagnosis, and knowledge of recovery protocols deployment is important for improving system reliability. This paper presents the CORRMEXT framework which carries failure diagnosis another significant step forward by analyzing and reporting error propagation patterns and degrees of success and failure of error recovery protocols. CORRMEXT uses both error messages and resource use data in its analyses. Application of CORRMEXT to data from the Ranger supercomputer have produced new insights. CORRMEXT has: (i) identified correlations between resource use counters that capture recovery attempts after an error, (ii) identified correlations between error events to capture error propagation patterns within the system, (iii) identified error propagation and recovery paths during system execution to explain system behaviour, (iv) showed that the earliest times of change in system behaviour can only be identified by analyzing both the correlated resource use counters and correlated errors. CORRMEXT will be installed on the HPC clusters at the Texas Advanced Computing Center in Autumn 2017. Edward Chuah, Arshad Jhumka, Samantha Alt, Theodoros Damoulas, Nentawe Gurumdimma, Marie-Christine Sawley, William L. Barth, Tommy Minyard, James C. Browne |
HiPC | 1 |
| 2016 | Using Message Logs and Resource Use Data for Cluster Failure DiagnosisabstractFailure diagnosis for large compute clusters using only message logs is known to be incomplete. Recent availability of resource use data provides another potentially useful source of data for failure detection and diagnosis. Early work combining message logs and resource use data for failure diagnosis has shown promising results. This paper describes the CRUMEL framework which implements a new approach to combining rationalized message logs and resource use data for failure diagnosis. CRUMEL identifies patterns of errors and resource use and correlates these patterns by time with system failures. Application of CRUMEL to data from the Ranger supercomputer has yielded improved diagnoses over previous research. CRUMEL has: (i) showed that more events correlated with system failures can only be identified by applying different correlation algorithms, (ii) confirmed six groups of errors, (iii) identified Lustre I/O resource use counters which are correlated with occurrence of Lustre faults which are potential flags for online detection of failures, (iv) matched the dates of correlated error events and correlated resource use with the dates of compute node hang-ups and (v) identified two more error groups associated with compute node hang-ups. The pre-processed data will be put on the public domain in September, 2016. Edward Chuah, Arshad Jhumka, James C. Browne, Nentawe Gurumdimma, Sai Narasimhamurthy, William L. Barth |
HiPC | 1 |
| 2016 | CRUDE: Combining Resource Usage Data and Error Logs for Accurate Error Detection in Large-Scale Distributed SystemsabstractThe use of console logs for error detection in large scale distributed systems has proven to be useful to system administrators. However, such logs are typically redundant and incomplete, making accurate detection very difficult. In an attempt to increase this accuracy, we complement these incomplete console logs with resource usage data, which captures the resource utilisation of every job in the system. We then develop a novel error detection methodology, the CRUDE approach, that makes use of both the resource usage data and console logs. We thus make the following specific technical contributions: we develop (i) a clustering algorithm to group nodes with similar behaviour, (ii) an anomaly detection algorithm to identify jobs with anomalous resource usage, (iii) an algorithm that links jobs with anomalous resource usage with erroneous nodes. We then evaluate our approach using console logs and resource usage data from the Ranger Supercomputer. Our results are positive: (i) our approach detects errors with a true positive rate of about 80%, and (ii) when compared with the well-known Nodeinfo error detection algorithm, our algorithm provides an average improvement of around 85% over Nodeinfo, with a best-case improvement of 250%. Nentawe Gurumdimma, Arshad Jhumka, Maria Liakata, Edward Chuah, James C. Browne |
SRDS | 4 |
| 2014 | Online failure prediction for HPC resources using decentralized clusteringabstractEnsuring high reliability of large-scale clusters is becoming more critical as the size of these machines continues to grow, since this increases the complexity and amount of interactions between different nodes and thus results in a high failure frequency. For this reason, predicting node failures in order to prevent errors from happening in the first place has become extremely valuable. A common approach for failure prediction is to analyze traces of system events to find correlations between event types or anomalous event patterns and node failures, and to use the types or patterns identified as failure predictors at run-time. However, typical centralized solutions for failure prediction in this manner suffer from high transmission and processing overheads at very large scales. We present a solution to the problem of predicting compute node soft-lockups in large scale clusters by using a decentralized online clustering algorithm (DOC) to detect anomalies in resource usage logs, which have been shown to correlate to particular types of node failures in supercomputer clusters. We demonstrate the effectiveness of this system by using the monitoring logs from the Ranger supercomputer at Texas Advanced Computing Center. Experiments shows that this approach can achieve similar accuracy as other related approaches, while maintaining low RAM and bandwidth usage, with a runtime impact to current running applications of less than 2%. Alejandro Pelaez, Andres Quiroz, James C. Browne, Edward Chuah, Manish Parashar |
HiPC | 4 |
| 2013 | Linking Resource Usage Anomalies with System Failures from Cluster Log DataabstractBursts of abnormally high use of resources are thought to be an indirect cause of failures in large cluster systems, but little work has systematically investigated the role of high resource usage on system failures, largely due to the lack of a comprehensive resource monitoring tool which resolves resource use by job and node. The recently developed TACC_Stats resource use monitor provides the required resource use data. This paper presents the ANCOR diagnostics system that applies TACC_Stats data to identify resource use anomalies and applies log analysis to link resource use anomalies with system failures. Application of ANCOR to first identify multiple sources of resource anomalies on the Ranger supercomputer, then correlate them with failures recorded in the message logs and diagnosing the cause of the failures, has identified four new causes of compute node soft lockups. ANCOR can be adapted to any system that uses a resource use monitor which resolves resource use by job. Edward Chuah, Arshad Jhumka, Sai Narasimhamurthy, John L. Hammond, James C. Browne, William L. Barth |
SRDS | 1 |
| 2011 | Establishing Hypothesis for Recurrent System Failures from Cluster Log FilesabstractA goal for the analysis of supercomputer logs is to establish causal relationships among events which reflect significant state changes in the system. Establishing these relationships is at the heart of failure diagnosis. In principle, a log analysis tool could automate many of the manual steps systems administrators must currently use to diagnose system failures. However, supercomputer logs are unstructured, incomplete and contain considerable ambiguity so that direct discovery of causal relationships is difficult. This paper describes the second generation FDiag log-based failure diagnostics framework that provides automation of the manual failure diagnosis process and determines with high confidence, the likely cause of the failure, the components involved and the event sequences which contain the times of the causal and terminal events. FDiag extracts relevant events from the system logs, performs correlation analysis on these events and from these correlations determines the components involved and the event sequences. The diagnostics capabilities of FDiag are validated by comparing its assessments on known instances of recurrent failures on the Ranger supercomputer at the University of Texas at Austin. We believe FDiag is the first log analyzer to demonstrate this level of diagnostics capability from the system logs of an open source software stack incorporating Linux and the Lustre file system. FDiag will be put into production use for support of failure diagnosis on Ranger in September, 2011. Edward Chuah, Gary Kee Khoon Lee, William-Chandra Tjhi, Shyh-Hao Kuo, Terence Hung, John L. Hammond, Tommy Minyard, James C. Browne |
DASC | 1 |
| 2010 | Diagnosing the root-causes of failures from cluster log filesabstractSystem event logs are often the primary source of information for diagnosing (and predicting) the causes of failures for cluster systems. Due to interactions among the system hardware and software components, the system event logs for large cluster systems are comprised of streams of interleaved events, and only a small fraction of the events over a small time span are relevant to the diagnosis of a given failure. Furthermore, the process of troubleshooting the causes of failures is largely manual and ad-hoc. In this paper, we present a systematic methodology for reconstructing event order and establishing correlations among events which indicate the root-causes of a given failure from very large syslogs. We developed a diagnostics tool, FDiag, to extract the log entries as structured message templates and uses statistical correlation analysis to establish probable cause and effect relationships for the fault being analyzed. We applied FDiag to analyze failures due to breakdowns in interactions between the Lustre file system and its clients on the Ranger supercomputer at the Texas Advanced Computing Center (TACC). The results are positive. FDiag is able to identify the dates and the time periods that contain the significant events which eventually led to the occurrence of compute node soft lockups. Edward Chuah, Shyh-Hao Kuo, Paul Hiew, William-Chandra Tjhi, Gary Kee Khoon Lee, John L. Hammond, Marek T. Michalewicz, Terence Hung, James C. Browne |
HiPC | 1 |
| 2008 | An optimal smooth QoS adaptation strategy for QoS differentiated scalable media streamingabstractDue to the advance of technologies in multimedia compression and network communications, scalable media streaming services have been availed to provide QoS differentiated services for heterogeneous users. However, it is still a big challenge to support consistent end-to-end Quality of Services (QoS) for the users due to the dynamic feature of the Internet, and abrupt variability of the network resources may severally affect the client perceived QoS. In this paper, we address the issues of smooth QoS adaptation for scalable streaming services. We propose an Optimal Smooth QoS Adaptation (OS-QA) strategy which allocates the server resource adaptively to cope with the variability of network bandwidth and protects the service quality of different quality classes under dynamic resource constraints. We analyze the quality variation caused by resource fluctuation and proposed OS-QA to minimize the average QoS variance under the resource constraints. Simulations are conducted to compare our proposed method with other QoS adaptation methods, and performance is analyzed in terms of QoS variance and PSNR. Results show that our proposed method is able to gracefully adapt the QoS and protect the client perceived QoS by minimizing the QoS variance under dynamic network resource constrains. Xiaorong Li, Edward Chuah, Jo Yew Tham, Kwong Huang Goh |
ICME | 2 |