EDBT 2026 Demo / reviewers in the wild / expert
Linh Ngo 0001
dblp:74/1895 · also Linh B. Ngo, Linh Bao Ngo
· DBLP profile ↗
5ranked-venue papers in the field
0as first author
2since 2021 · last 2024
0000-0002-9889-2742ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Predicting ChatGPT's Ability to Solve Complex Programming ChallengesabstractThe recent emergence of Large Language Model (LLM)-based tools such as OpenAI’s ChatGPT and Google’s Gemini has sparked excitement across the software development industry, and offered promises to transform the software development process. Despite the enthusiasm, it remains uncertain whether these tools are already good enough at coding to replace the role of software developers. Currently, no studies have provided insights into the performance of LLMs, such as understanding which characteristics of a programming task might affect an LLM's performance, or predicting how an LLM will handle new programming challenges. In this work, we address these challenges by first creating a data collection framework to gather 3,323 programming tasks from Kattis, a widely-used programming challenge platform. We then use OpenAI's ChatGPT to solve these programming tasks. The solutions obtained from ChatGPT are submitted back to Kattis to evaluate their correctness and effectiveness. Next, we use the collected data, including both problem and solution information, to analyze the task characteristics that significantly influence ChatGPT's performance. Building on the analysis, we develop predictive models that can forecast the efficacy of ChatGPT on new programming problems. Our analysis indicates that factors such as the difficulty level of a programming challenge, or the readability complexity of a problem description can significantly affect the efficacy of ChatGPT. Finally, the experimental results show that our predictive model can correctly predict ChatGPT performance with an accuracy of up to 90% for easy problems, and up to 79% for difficult problems. Nguyen Ho, James J. May, Bao Ngo, Jack Formato, Linh Ngo 0001, Van Long Ho, Hoang Bui |
IEEE Big Data | 5 |
| 2023 | Causal Associations between Temporal EventsabstractCausal inference from observational data has been widely studied to infer causal relations between causes and effects. Due to the popularity of event-based data, causal inference from event datasets has attracted increasing interest. However, inferring causalities from observational event sequences is challenging because of the heterogeneous and irregular nature of event-based data. Existing work on causal inference for temporal events disregards the event durations, and is thus unable to capture their impact on the causal relations. In the present paper, we overcome this limitation by proposing a new modeling approach for temporal events that captures and utilizes event durations. Based on this new temporal model, we propose a set of novel Duration-based Event Causality (DEC) scores, including the Duration-based Necessity and Sufficiency Trade-off score, and the Duration-based Conditional Intensity Rates scores that take into consideration event durations when inferring causal associations between temporal events. We conduct an extensive experimental evaluation using both synthetic datasets and real-world event datasets in the environmental domains to evaluate our proposed scores, and compare them against the closest baseline. The experimental results show that our proposed scores outperform the baseline with a large margin using the popular evaluation metric Hits@K. Nguyen Ho, Trinh Cong Le, Van Long Ho, Nguyen Tuong Huynh, Linh Ngo 0001 |
IEEE Big Data | 5 |
| 2014 | Synthetic data generation for the internet of thingsabstractThe concept of Internet of Things (IoT) is rapidly moving from a vision to being pervasive in our everyday lives. This can be observed in the integration of connected sensors from a multitude of devices such as mobile phones, healthcare equipment, and vehicles. There is a need for the development of infrastructure support and analytical tools to handle IoT data, which are naturally big and complex. But, research on IoT data can be constrained by concerns about the release of privately owned data. In this paper, we present the design and implementation results of a synthetic IoT data generation framework. The framework enables research on synthetic data that exhibit the complex characteristics of original data without compromising proprietary information and personal privacy. Jason W. Anderson, Ken E. Kennedy, Linh Ngo 0001, André Luckow, Amy W. Apon |
IEEE BigData | 3 |
| 2014 | Managing the academic data lifecycle: A case study of HPCCabstractAcademic data can be classified into multiple categories and come from a large number of sources. Many research areas require combining data from different sources into a unified set on which analytical techniques can be applied. In this research paper the authors introduce the High Performance Computing Cluster (HPCC) as a platform to streamline the process of ingesting, curating, integrating and transforming scholarly data from multiple sources and in varying formats, particularly when several of these datasets lack common attributes to support the integration process. Michael E. Payne, Linh Ngo 0001, Flavio Villanustre, Amy W. Apon |
IEEE BigData | 2 |
| 2013 | Academic publishing as a social media paradigmabstractThis work seeks to bridge areas of academic institutional research, social network analysis, and content analysis through the application of a social media paradigm to the academic research publishing environment. The concept is built upon the analysis of similarities and differences that exist in the structural and functional building blocks of academic publishing and social media. The potential impact in the form of new research directions that arise from this application are presented. Michael E. Payne, Linh Ngo 0001, Amy W. Apon |
IEEE BigData | 2 |