Tung Dao

dblp:200/9083 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0006-6015-2358ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Pareto-Grid-Guided Large Language Models for Fast and High-Quality Heuristics Design in Multi-Objective Combinatorial Optimization
abstract
Multi-objective combinatorial optimization problems (MOCOP) frequently arise in practical applications that require the simultaneous optimization of conflicting objectives. Although traditional evolutionary algorithms can be effective, they typically depend on domain knowledge and repeated parameter tuning, limiting flexibility when applied to unseen MOCOP instances. Recently, integration of Large Language Models (LLMs) into evolutionary computation has opened new avenues for automatic heuristic generation, using their advanced language understanding and code synthesis capabilities. Nevertheless, most existing approaches predominantly focus on single-objective tasks, often neglecting key considerations such as runtime efficiency and heuristic diversity in multi-objective settings. To bridge this gap, we introduce Multi-heuristics for MOCOP via Pareto-Grid-guided Evolution of LLMs (MPaGE), a novel enhancement of the Simple Evolutionary Multiobjective Optimization (SEMO) framework that leverages LLMs and Pareto Front Grid (PFG) technique. By partitioning the objective space into grids and retaining top-performing candidates to guide heuristic generation, MPaGE utilizes LLMs to prioritize heuristics with semantically distinct logical structures during variation, thus promoting diversity and mitigating redundancy within the population. Through extensive evaluations, MPaGE demonstrates superior performance over existing LLM-based frameworks, and achieves competitive results to traditional Multi-objective evolutionary algorithms (MOEAs), with significantly faster runtime.
Ha Minh Hieu, Hung Phan, Tung Duy Doan, Tung Dao, Huynh Thi Thanh Binh
AAAI4
2026 MOTIF: Multi-strategy Optimization via Turn-based Interactive Framework
abstract
Designing effective algorithmic components remains a fundamental obstacle in tackling NP-hard combinatorial optimization problems (COPs), where solvers often rely on carefully hand-crafted strategies. Despite recent advances in using large language models (LLMs) to synthesize high-quality components, most approaches restrict the search to a single element—commonly a heuristic scoring function—thus missing broader opportunities for innovation. We introduce a broader formulation of solver design as a multi-strategy optimization problem, which seeks to jointly improve a set of interdependent components under a unified objective. To address this, we propose MOTIF—Multi-strategy Optimization via Turn-based Interactive Framework—a novel framework based on Monte Carlo Tree Search that facilitates turn-based optimization between two LLM agents. At each turn, an agent improves one component by leveraging the history of both its own and its opponent’s prior updates, promoting both competitive pressure and emergent cooperation. This structured interaction broadens the search landscape and encourages the discovery of diverse, high-performing solutions. Experiments across multiple COP domains show that MOTIF consistently outperforms state-of-the-art methods, highlighting the promise of turn-based, multi-agent prompting for fully automated solver design.
Nguyen Viet Tuan Kiet, Tung Dao, Huynh Thi Thanh Binh
AAAI2
2025 IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants
abstract
We introduce IndEgo, a multimodal egocentric and exocentric dataset addressing common industrial tasks, including assembly/disassembly, logistics and organisation, inspection and repair, woodworking, and others. The dataset contains 3,460 egocentric recordings (approximately 197 hours), along with 1,092 exocentric recordings (approximately 97 hours). A key focus of the dataset is collaborative work, where two workers jointly perform cognitively and physically intensive tasks. The egocentric recordings include rich multimodal data and added context via eye gaze, narration, sound, motion, and others. We provide detailed annotations (actions, summaries, mistake annotations, narrations), metadata, processed outputs (eye gaze, hand pose, semi-dense point cloud), and benchmarks on procedural and non-procedural task understanding, Mistake Detection, and reasoning-based Question Answering. Baseline evaluations for Mistake Detection, Question Answering and collaborative task understanding show that the dataset presents a challenge for the state-of-the-art multimodal models. Our dataset is available at: https://huggingface.co/datasets/FraunhoferIPK/IndEgo
Vivek Chavan, Yasmina Imgrund, Tung Dao, Sanwantri Bai, Bosong Wang, Ze Lu, Oliver Heimann, Jörg Krüger
NeurIPS3
2023 Triggering Modes in Spectrum-Based Multi-location Fault Localization
abstract
Spectrum-based fault localization (SBFL) techniques can aid in debugging, but their practicality in industrial settings has been limited due to the large number of tests needed to execute before applying SBFL. Previous research has explored different trigger modes for SBFL and found that applying it immediately after the first test failure is also effective. However, this study only considered single-location bugs, while multi-location bugs are prevalent in real-world scenarios and especially at our company Cvent, which is interested in integrating SBFL to its CI/CD workflow.
Tung Dao, Na Meng 0001, ThanhVu Nguyen
ESEC/SIGSOFT FSE1
2021 Exploring the Triggering Modes of Spectrum-Based Fault Localization: An Industrial Case
abstract
Fault localization is important for software development and maintenance. Among existing techniques, spectrum-based fault localization (SBFL) is effective to locate bugs based on the execution coverage of passed and failed tests. However, current SBFL tools require the execution of all tests before suggesting any ranked list of suspicious locations. In reality, such all-test execution can be very time-consuming; SBFL's outputs based on the whole-suite execution can significantly delay developers' debugging activities and jeopardize their productivity. For this paper, we were curious whether we can apply SBFL immediately after seeing one or several test failures, instead of waiting for all tests to finish their run. Specifically, with 28 injected bugs and 13 real bugs in a close-sourced software product, we collected the statement-level coverage for each test case, and investigated the usage of 25 alternative SBFL formulas. We triggered SBFL in five modes: (i) after the first test failure, (ii) after the first failure and some extra passed tests, (iii) after every test failure, (iv) at a specified time interval (e.g., every 2 minutes), or (v) after the complete execution of all tests. Our study shows interesting results. Compared with whole-suite execution, triggering SBFL formulas earlier based on partial execution helps locate bugs more effectively. Among the five modes, the first-failure-driven mode works best. Additionally, we conducted similar experiments on 57 real bugs from the Defects4J dataset and observed similar phenomena. Our observations imply that instead of waiting for the completion of all test runs, it is quite promising to apply SBFL formulas immediately after the initial test failure. In this way, developers are likely to get better suggestions within a shorter period of time. Our research will help developers better adopt SBFL in practice.
Tung Dao, Max Wang, Na Meng 0001
ICST1
2017 How does execution information help with information-retrieval based bug localization?
abstract
Bug localization is challenging and time-consuming. Given a bug report, a developer may spend tremendous time comprehending the bug description together with code in order to locate bugs. To facilitate bug report comprehension, information retrieval (IR)-based bug localization techniques have been proposed to automatically search for and rank potential buggy code elements (i.e., classes or methods). However, these techniques do not leverage any dynamic execution information of buggy programs. In this paper, we perform the first systematic study on how dynamic execution information can help with static IR-based bug localization. More specifically, with the fixing patches and bug reports of 157 real bugs, we investigated the impact of various execution information (i.e. coverage, slicing, and spectrum) on three IR-based techniques: the baseline technique, BugLocator, and BLUiR. Our experiments demonstrate that both the coverage and slicing information of failed tests can effectively reduce the search space and improve IR-based techniques at both class and method levels. Using additional spectrum information can further improve bug localization at the method but not the class level. Some of our investigated ways of augmenting IR-based bug localization with execution information even outperform a state-of-the-art technique, which merges spectrum with an IR-based technique in a complicated way. Different from prior work, by investigating various easy-to-understand ways to combine execution information with IR-based techniques, this study shows for the first time that execution information can generally bring considerable improvement to IR-based bug localization.
Tung Dao, Lingming Zhang 0001, Na Meng 0001
ICPC1