Minh Ngoc Dinh

dblp:19/8524 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
3since 2021 · last 2026
0000-0003-4026-6068ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-authorArtificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Information extraction and text analysis · 40% Language models and text generation · 40% Question answering and dialogue systems · 20%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
GPUs and heterogeneous computing · 35% Energy-efficient computing · 26% High-performance computing · 22%
Software engineering, system software, and programming languages
4 papers
Debugging and program repair · 94% Program analysis · 6%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
argument mining
1.012026
ARQUSUMM: Argument-aware Quantitative Summarization of Online Conversations · AAAI 2026
Information retrieval › text summarization
conversation summarization
1.012026
ARQUSUMM: Argument-aware Quantitative Summarization of Online Conversations · AAAI 2026
Information retrieval
text summarization
1.012026
ARQUSUMM: Argument-aware Quantitative Summarization of Online Conversations · AAAI 2026
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
product question answering
0.912025
QQSUM: A Novel Task and Model of Quantitative Query-Focused Summarization for Review-based Product Question Answering · ACL (1) 2025
Natural language and speech › Language models and text generation › text summarization › controllable summarization
query-focused summarization
0.912025
QQSUM: A Novel Task and Model of Quantitative Query-Focused Summarization for Review-based Product Question Answering · ACL (1) 2025
Natural language and speech › Language models and text generation
text summarization
0.912025
QQSUM: A Novel Task and Model of Quantitative Query-Focused Summarization for Review-based Product Question Answering · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › argument mining
key point analysis
0.812024
Prompted Aspect Key Point Analysis for Quantitative Review Summarization · ACL (1) 2024
Debugging and program repair › software debugging
relative debugging
0.532015
Relative debugging for a highly parallel hybrid computer system · SC 2015
Scalable Relative Debugging · IEEE Trans. Parallel Distributed Syst. 2014
Data centric highly parallel debugging · HPDC 2010
GPUs and heterogeneous computing
GPU computing
0.212015
Relative debugging for a highly parallel hybrid computer system · SC 2015
GPUs and heterogeneous computing › heterogeneous computing systems
hybrid computer
0.212015
Relative debugging for a highly parallel hybrid computer system · SC 2015
High-performance computing
large-scale parallel debugging
0.212014
Scalable Relative Debugging · IEEE Trans. Parallel Distributed Syst. 2014
Electronic design automation › hardware verification and test
debugging
0.112012
Scalable parallel debugging with statistical assertions · PPoPP 2012
Debugging and program repair › automated debugging
assertion-based debugging
0.112010
Data centric highly parallel debugging · HPDC 2010
Debugging and program repair › concurrent program debugging
parallel program debugging
0.112010
Data centric highly parallel debugging · HPDC 2010
Parallel and multicore computing
parallel programming models
0.112015
Relative debugging for a highly parallel hybrid computer system · SC 2015
High-performance computing › scientific computing systems
molecular dynamics simulation
0.012012
Scalable parallel debugging with statistical assertions · PPoPP 2012
High-performance computing
scientific computing
0.012012
Scalable parallel debugging with statistical assertions · PPoPP 2012

Methods — techniques the papers use, named apart from their topics

clustering · 2.0argumentation theory · 2.0LLM few-shot learning · 2.0retrieval-augmented generation · 0.9few-shot learning · 0.9in-context learning · 0.8aspect sentiment analysis · 0.8data model · 0.4case study · 0.4point-to-point comparison · 0.4hash-based comparison · 0.4statistical assertion · 0.3parallelization · 0.3hashing · 0.1assertions · 0.1
YearPublicationVenuePosition
2026 ARQUSUMM: Argument-aware Quantitative Summarization of Online Conversations
abstract
Online conversations have become more prevalent on public discussion platforms (e.g. Reddit). With growing controversial topics, it is desirable to summarize not only diverse arguments, but also their rationale and justification. Early studies on text summarization focus on capturing general salient information in source documents, overlooking the argumentative nature of online conversations. Recent research on conversation summarization although considers the argumentative relationship among sentences, fail to explicate deeper argument structure within sentences for summarization. In this paper, we propose a novel task of argument-aware quantitative summarization to reveal the claim-reason structure of arguments in conversations, with quantities measuring argument strength. We further propose ARQUSUMM, a novel framework to address the task. To reveal the underlying argument structure within sentences, ARQUSUMM leverages LLM few-shot learning grounded in the argumentation theory to identify propositions within sentences and their claim-reason relationships. For quantitative summarization, ARQUSUMM employs argument structure-aware clustering algorithms to aggregate arguments and quantify their support. Experiments show that ARQUSUMM outperforms existing conversation and quantitative summarization models and generate summaries representing argument structures that are more helpful to users, of high textual quality and quantification accuracy.
An Quang Tang, Xiuzhen Zhang 0001, Minh Ngoc Dinh, Zhuang Li 0001
AAAI3
2025 QQSUM: A Novel Task and Model of Quantitative Query-Focused Summarization for Review-based Product Question Answering
abstract
Review-based Product Question Answering (PQA) allows e-commerce platforms to automatically address customer queries by leveraging insights from user reviews.However, existing PQA systems generate answers with only a single perspective, failing to capture the diversity of customer opinions.In this paper we introduce a novel task Quantitative Query-Focused Summarization (QQSUM), which aims to summarize diverse customer opinions into representative Key Points (KPs) and quantify their prevalence to effectively answer user queries.While Retrieval-Augmented Generation (RAG) shows promise for PQA, its generated answers still fall short of capturing the full diversity of viewpoints.To tackle this challenge, our model QQSUM-RAG, which extends RAG, employs few-shot learning to jointly train a KP-oriented retriever and a KP summary generator, enabling KP-based summaries that capture diverse and representative opinions.Experimental results demonstrate that QQSUM-RAG achieves superior performance compared to state-of-the-art RAG baselines in both textual quality and quantification accuracy of opinions.Our source code is available at: https://github.com/ antangrocket1312/
An Quang Tang, Xiuzhen Zhang 0001, Minh Ngoc Dinh, Zhuang Li 0001
ACL (1)3
2024 Prompted Aspect Key Point Analysis for Quantitative Review Summarization
abstract
Key Point Analysis (KPA) aims for quantitative summarization that provides key points (KPs) as succinct textual summaries and quantities measuring their prevalence.KPA studies for arguments and reviews have been reported in the literature.A majority of KPA studies for reviews adopt supervised learning to extract short sentences as KPs before matching KPs to review comments for quantification of KP prevalence.Recent abstractive approaches still generate KPs based on sentences, often leading to KPs with overlapping and hallucinated opinions, and inaccurate quantification.In this paper, we propose Prompted Aspect Key Point Analysis (PAKPA) for quantitative review summarization.PAKPA employs aspect sentiment analysis and prompted in-context learning with Large Language Models (LLMs) to generate and quantify KPs grounded in aspects for business entities, which achieves faithful KPs with accurate quantification, and removes the need for large amounts of annotated data for supervised training.Experiments on the popular review dataset Yelp and the aspect-oriented review summarization dataset SPACE show that our framework achieves state-of-the-art performance.Source code and data are available at: https://github.com/ antangrocket1312/PAKPA Key Point Prevalence Matching Comments Friendly and helpful staff.46 From the minute I walked in the door the staff has treated me with such kindness and respect, there are not enough words I can say of how much gratitude I have for the staff here.It's a fine hotel, pretty basic, and all of the staff members we encountered were quite friendly.Excellent location for business travelers.37 Its a great location for business travelers since I can stay here and always be going the opposite way of traffic.*This hotel has an IDEAL location, the shuttle was perfect, and we got a great deal on priceline.Clean and comfortable rooms.32 *You will enjoy the breathtaking views from your spacious, clean and comfortable room.The rooms were nice and clean and the bed was comfortable.Convenient and helpful shuttle service.12 *This hotel has an IDEAL location, the shuttle was perfect, and we got a great deal on priceline.I would also like to give a shout out to the terrific shuttle drivers Jeff and Rod who were so great to my family group.Amazing view of Vanderbilt football stadium or The Parthenon.10 *You will enjoy the breathtaking views from your spacious, clean and comfortable room.Then the surprise came when we opened the curtains to see a full on view of the Vanderbilt football stadium.
An Quang Tang, Xiuzhen Zhang 0001, Minh Ngoc Dinh, Erik Cambria
ACL (1)3
2020 Tracking scientific simulation using online time-series modelling
abstract
The increase in compute power and complexity of supercomputing systems requires the decrease in the feature size and the supply voltage of internal components. Such development makes unintended errors such as soft errors, potentially caused by random bit flips, inevitable because of the huge size of the resources (such as CPU cores and memory). In this paper, we discuss a non-parametric statistical modelling technique to implement a soft error detector. By exploring temporal autocorrelation within key variables of a running scientific simulation, we introduce an automatic anomaly detection technique in which runtime data from a time-step based simulation can be converted into a time series, and a time series modelling technique can be used to identify soft errors at runtime. Experiments with LAMMPS, a high-performance molecular dynamics simulator, and with PLUTO, an open-source astrophysical code, reveal that the time-series based detector is subjected to less than 3% of both false-positive rate and false-negative rate while incurring only 6% performance overheads.
Minh Ngoc Dinh, Chien Trung Vo, David Abramson 0001
CCGRID1
2018 Energy efficiency modeling of parallel applications
Mark Endrei, Chao Jin 0001, Minh Ngoc Dinh, David Abramson 0001, Heidi Poxon, Luiz DeRose, Bronis R. de Supinski
SC3
2017 A Computational Pipeline for the IUCN Risk Assessment for Meso-American Reef Ecosystem
abstract
Coral reefs are of global economic and biological significance but are subject to increasing threats. As a result, it is essential to understand the risk of coral reef ecosystem collapse and to develop assessment process for those ecosystems. The International Union for Conservation of Nature (IUCN) Red List of Ecosystem (RLE) is a framework to assess the vulnerability of an ecosystem. Importantly, the assessment processes need to be repeatable as new monitoring data arises. The repeatability will also enhance transparency. In this paper, we discuss the evolution of a computational pipeline for risk assessment of the Meso-American reef ecosystem, a diverse reef ecosystem located in the Caribbean, with the focus on improving the execution time starting from sequential and parallel implementation and finally using Apache Spark. The final form of the pipeline is a scientific workflow to improve its repeatability and reproducibility.
Hoang Anh Nguyen, Lucie Bland, Tristan Roberts, Siddeswara Guru, Minh Ngoc Dinh, David Abramson 0001
eScience5
2015 Relative debugging for a highly parallel hybrid computer system
abstract
Relative debugging traces software errors by comparing two executions of a program concurrently - one code being a reference version and the other faulty. Relative debugging is particularly effective when code is migrated from one platform to another, and this is of significant interest for hybrid computer architectures containing CPUs accelerators or coprocessors. In this paper we extend relative debugging to support porting stencil computation on a hybrid computer. We describe a generic data model that allows programmers to examine the global state across different types of applications, including MPI/OpenMP, MPI/OpenACC, and UPC programs. We present case studies using a hybrid version of the `stellarator' particle simulation DELTA5D, on Titan at ORNL, and the UPC version of Shallow Water Equations on Crystal, an internal supercomputer of Cray. These case studies used up to 5,120 GPUs and 32,768 CPU cores to illustrate that the debugger is effective and practical.
Luiz De Rose, Andrew Gontarek, Aaron Vose, Bob Moench, David Abramson 0001, Minh Ngoc Dinh, Chao Jin 0001
SC6
2015 A data-centric framework for debugging highly parallel applications
abstract
Summary Contemporary parallel debuggers allow users to control more than one processing thread while supporting the same examination and visualisation operations of that of sequential debuggers. This approach restricts the use of parallel debuggers when it comes to large scale scientific applications run across hundreds of thousands compute cores. First, manually observing the runtime data to detect error becomes impractical because the data is too big. Second, performing expensive but useful debugging operations becomes infeasible as the computational codes become more complex, involving larger data structures, and as the machines become larger. This study explores the idea of a data‐centric debugging approach, which could be used to make parallel debuggers more powerful. It discusses the use ofad hocdebug‐time assertions that allow a user to reason about the state of a parallel computation. These assertions support the verification and validation of program state at runtime as a whole rather than focusing on that of only a single process state. Furthermore, the debugger's performance can be improved by exploiting the underlying parallel platform because the available compute cores can execute parallel debugging functions, while a program is idling at a breakpoint. We demonstrate the system with several case studies and evaluate the performance of the tool on a 20 000 cores Cray XE6. Copyright © 2013 John Wiley & Sons, Ltd.
Minh Ngoc Dinh, David Abramson 0001, Chao Jin 0001, Andrew Gontarek, Bob Moench, Luiz De Rose
Softw. Pract. Exp.1
2014 Scalable Relative Debugging
abstract
Detecting and isolating bugs that arise only at high processor counts is a challenging task. Over a number of years, we have implemented a special debugging method, called "relative debugging," that supports debugging applications as they evolve or are ported to larger machines. It allows a user to compare the state of a suspect program against another reference version even as the number of processors is increased. The innovative idea is the comparison of runtime data to reason about the state of the suspect program. While powerful, a naïve implementation of the comparison phase does not scale to large problems running on large machines. In this paper, we propose two different solutions including a hash-based scheme and a direct point-to-point scheme. We demonstrate the implementation, a case study, as well as the performance, of our techniques on 20K cores of a Cray XE6 system.
Minh Ngoc Dinh, David Abramson 0001, Chao Jin 0001
IEEE Trans. Parallel Distributed Syst.1
2012 A Scalable Parallel Debugging Library with Pluggable Communication Protocols
abstract
Parallel debugging faces challenges in both scalability and efficiency. A number of advanced methods have been invented to improve the efficiency of parallel debugging. As the scale of system increases, these methods highly rely on a scalable communication protocol in order to be utilized in large-scale distributed environments. This paper describes a debugging middleware that provides fundamental debugging functions supporting multiple communication protocols. Its pluggable architecture allows users to select proper communication protocols as plug-ins for debugging on different platforms. It aims to be utilized by various advanced debugging technologies across different computing platforms. The performance of this debugging middleware is examined on a Cray XE Supercomputer with 21,760 CPU cores.
Chao Jin 0001, David Abramson 0001, Minh Ngoc Dinh, Andrew Gontarek, Bob Moench, Luiz De Rose
CCGRID3
2012 Scalable parallel debugging with statistical assertions
abstract
Traditional debuggers are of limited value for modern scientific codes that manipulate large complex data structures. This paper discusses a novel debug-time assertion, called a "Statistical Assertion", that allows a user to reason about large data structures, and the primitives are parallelised to provide an efficient solution. We present the design and implementation of statistical assertions, and illustrate the debugging technique with a molecular dynamics simulation. We evaluate the performance of the tool on a 12,000 cores Cray XE6.
Minh Ngoc Dinh, David Abramson 0001, Chao Jin 0001, Andrew Gontarek, Bob Moench, Luiz De Rose
PPoPP1
2011 Assertion Based Parallel Debugging
abstract
Programming languages have advanced tremendously over the years, but program debuggers have hardly changed. Sequential debuggers do little more than allow a user to control the flow of a program and examine its state. Parallel ones support the same operations on multiple processes, which are adequate with a small number of processors, but become unwieldy and ineffective on very large machines. Typical scientific codes have enormous multi-dimensional data structures and it is impractical to expect a user to view the data using traditional display techniques. In this paper we discuss the use of debug-time assertions, and show that these can be used to debug parallel programs. The techniques reduce the debugging complexity because they reason about the state of large arrays without requiring the user to know the expected value of every element. Assertions can be expensive to evaluate, but their performance can be improved by running them in parallel. We demonstrate the system with a case study finding errors in a parallel version of the Shallow Water Equations, and evaluate the performance of the tool on a 4,096 cores Cray XE6.
Minh Ngoc Dinh, David Abramson 0001, Donny Kurniawan, Chao Jin 0001, Bob Moench, Luiz De Rose
CCGRID1
2010 Data centric highly parallel debugging
abstract
Debugging parallel programs is an order of magnitude more complex than sequential ones, and yet, most parallel debuggers provide little extra functionality than their sequential counterparts. This problem becomes more serious as computational codes become more complex, involving larger data structures, and as the machines become larger. Peta-scale machines consisting of millions of cores pose a significant challenge for existing techniques. We argue that debugging must become more data-centric, and believe that "assertions" provide a useful model. Assertions allow a user to declare their expectations about the program state as a whole rather than focusing on that of only a single process state. Previously, we have implemented a special type of assertion that supports debugging applications as they evolve or are ported to different platforms. They allow a user to compare the state of one program against another reference version. These 'relative debugging' assertions, whilst powerful, pose significant implementation challenges for large peta-scale machines. In this paper we discuss a hashing technique that provides a scalable solution for very large problems on very large machines. We illustrate the scheme on 65k cores of Kraken, a Cray XT5 at the University of Tennessee.
David Abramson 0001, Minh Ngoc Dinh, Donny Kurniawan, Bob Moench, Luiz De Rose
HPDC2
2009 Virtual Microscopy and Analysis Using Scientific Workflows
abstract
Most commercial microscopes are stand-alone instruments, controlled by dedicated computer systems. These provide limited storage and processing capabilities. Virtual microscopes, on the other hand, link the image capturing hardware and data analysis software into a wide area network of high performance computers, large storage devices and software systems. In this paper we discuss extensions to Grid workflow engines that allow them to execute scientific experiments on virtual microscopes. We demonstrate the utility of such a system in a biomedical case study concerning the imaging of cancer and antibody based therapeutics.
David Abramson 0001, Blair Bethwaite, Minh Ngoc Dinh, Colin Enticott, Stephen Firth, Slavisa Garic, Ian Harper, Martin Lackmann, Tirath Ramdas, A. B. M. Russel, Stefan Schek, Mary Vail
eScience3