VLDB 2026 Research / reviewers in the wild / expert
Hung Viet Pham
dblp:163/1941
· DBLP profile ↗
20ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Consistency Meets Verification: Enhancing Test Generation Quality in Large Language Models Without Ground-Truth Solutions
Hamed Taherkhani, Alireza DaghighFarsoodeh, Mohammad Chowdhury, Hung Viet Pham, Hadi Hemmati |
ICST | 4 |
| 2026 | Evaluating API-Level Deep Learning Fuzzers: A Comprehensive Benchmarking StudyabstractIn recent years, the practice of fuzzing Deep Learning (DL) APIs has received significant attention in the software engineering community. Many API-level DL fuzzers have been proposed to test individual DL APIs by generating malformed input. Although these fuzzers have been effective in detecting bugs and outperforming prior work, there remains a gap in benchmarking them against ground-truth, real-world bugs in DL libraries. Existing comparisons among these API-level DL fuzzers primarily focus on the bugs detected but do not offer a comprehensive, in-depth evaluation of the fuzzers’ effectiveness. In this work, we perform the first in-depth evaluation of state-of-the-art API-level DL fuzzers that generate tests for single DL APIs, focusing on their effectiveness against real-world bugs. We manually created an extensive benchmark dataset, including 517 real-world DL bugs collected from PyTorch and TensorFlow libraries that can be triggered by malformed inputs. We then apply seven state-of-the-art DL fuzzers— FreeFuzz , DeepRel , NablaFuzz , DocTer , ACETest , TitanFuzz , and FuzzGPT —to our benchmark dataset, following their respective instructions. Our results show that these fuzzers detect only 6.5% (34 out of 517) of the unique real-world bugs in the dataset. Our analysis identifies two dominant factors that impact the effectiveness of these fuzzers in detecting real-world bugs. These findings suggest opportunities for improving the performance of fuzzers in future work. Overall, this study extends previous work on DL fuzzers by providing an extensive evaluation and benchmarking platform for fuzzing DL libraries. Nima Shiri Harzevili, Moshi Wei, Mohammad Mahdi Mohajer, Hung Viet Pham, Song Wang 0009 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2025 | History-Driven Fuzzing for Deep Learning LibrariesabstractRecently, many Deep Learning (DL) fuzzers have been proposed for API-level testing of DL libraries. However, they either perform unguided input generation (e.g., not considering the relationship between API arguments when generating inputs) or only support a limited set of corner-case test inputs. Furthermore, many developer APIs crucial for library development remain untested, as they are typically not well documented and lack clear usage guidelines, unlike end-user APIs. This makes them a more challenging target for automated testing. To fill this gap, we propose a novel fuzzer named Orion, which combines guided test input generation and corner-case test input generation based on a set of fuzzing heuristic rules constructed from historical data known to trigger critical issues in the underlying implementation of DL APIs. To extract the fuzzing heuristic rules, we first conduct an empirical study on the root cause analysis of 376 vulnerabilities in two of the most popular DL libraries, PyTorch and TensorFlow. We then construct the fuzzing heuristic rules based on the root causes of the extracted historical vulnerabilities. Using these fuzzing heuristic rules, Orion generates corner-case test inputs for API-level fuzzing. In addition, we extend the seed collection of existing studies to include test inputs for developer APIs. Our evaluation shows that Orion reports 135 vulnerabilities in the latest releases of TensorFlow and PyTorch, 76 of which were confirmed by the library developers. Among the 76 confirmed vulnerabilities, 69 were previously unknown, and 7 have already been fixed. The rest are awaiting further confirmation. For end-user APIs, Orion detected 45.58% and 90% more vulnerabilities in TensorFlow and PyTorch, respectively, compared to the state-of-the-art conventional fuzzer, DeepRel. When compared to the state-of-the-art LLM-based DL fuzzer, AtlasFuz, and Orion detected 13.63% more vulnerabilities in TensorFlow and 18.42% more vulnerabilities in PyTorch. Regarding developer APIs, Orion stands out by detecting 117% more vulnerabilities in TensorFlow and 100% more vulnerabilities in PyTorch compared to the most relevant fuzzer designed for developer APIs, such as FreeFuzz. Nima Shiri Harzevili, Mohammad Mahdi Mohajer, Moshi Wei, Hung Viet Pham, Song Wang 0009 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2025 | D3: Differential Testing of Distributed Deep Learning With Model GenerationabstractDeep Learning (DL) techniques have been widely deployed in many application domains. The growth of DL models’ size and complexity demands distributed training of DL models. Since DL training is complex, software implementing distributed DL training is error-prone. Thus, it is crucial to test distributed deep learning software to improve its reliability and quality. To address this issue, we propose adifferentialtesting technique—D3, which leverages adistributedequivalence rule that we create to test distributeddeeplearning software. The rationale is that the same model trained with the same model input under different distributed settings should produce equivalent prediction output within certain thresholds. The different output indicates potential bugs in the distributed deep learning software. D3automatically generates a diverse set of distributed settings, DL models, and model input to test distributed deep learning software. Our evaluation on two of the most popular DL libraries, i.e., PyTorch and TensorFlow, shows that D3detects 21 bugs, including 12 previously unknown bugs. Jiannan Wang 0002, Hung Viet Pham, Lin Tan 0001, Adnan Aziz, Erik Meijer 0001 |
IEEE Trans. Software Eng. | 2 |
| 2024 | Can ChatGPT Support Developers? An Empirical Evaluation of Large Language Models for Code GenerationabstractLarge language models (LLMs) have demonstrated notable proficiency in code generation, with numerous prior studies showing their promising capabilities in various development scenarios. However, these studies mainly provide evaluations in research settings, which leaves a significant gap in understanding how effectively LLMs can support developers in real-world. To address this, we conducted an empirical analysis of conversations in DevGPT, a dataset collected from developers' conversations with ChatGPT (captured with the Share Link feature on platforms such as GitHub). Our empirical findings indicate that the current practice of using LLM-generated code is typically limited to either demonstrating high-level concepts or providing examples in documentation, rather than to be used as production-ready code. These findings indicate that there is much future work needed to improve LLMs in code generation before they can be integral parts of modern software development. Kailun Jin, Chung-Yu Wang, Hung Viet Pham, Hadi Hemmati |
MSR | 3 |
| 2024 | CEDAR: Continuous Testing of Deep Learning LibrariesabstractSince Deep Learning (DL) libraries undergo rapid development with thousands of lines of code changes daily, they require continuous testing to detect software bugs and ensure code quality. In this paper, we explore DL testing approaches in a continuous testing setting. To make it feasible, we present the first continuous testing framework for DL libraries-CEDAR-that integrates two state-of-the-art DL testing approaches (DocTer and EAGLE) efficiently to test two popular DL libraries, PyTorch and TensorFlow. Through the application of CEDAR to 20 versions of PyTorch and TensorFlow, CEDAR detects 83 bugs in 140 APIs. Out of the 83 bugs, 23 are previously unknown bugs with 21 confirmed or fixed by the developers. The results also show CEDAR has effectively shortened the bug detection latency by almost a year (338.6 days) on average. In addition, CEDAR demonstrates its effectiveness in detecting new regression bugs and masked bugs. With three optimization strategies, CEDAR reduces the time and space overhead by a factor of 15.4 and 9.7. We share insights and lessons learned from our research, aiming to advance the development of more effective and efficient continuous testing for DL libraries, benefiting both developers and researchers. Danning Xie, Jiannan Wang 0002, Hung Viet Pham, Lin Tan 0001, Adnan Aziz, Erik Meijer 0001 |
SANER | 3 |
| 2023 | DisGUIDE: Disagreement-Guided Data-Free Model ExtractionabstractRecent model-extraction attacks on Machine Learning as a Service (MLaaS) systems have moved towards data-free approaches, showing the feasibility of stealing models trained with difficult-to-access data. However, these attacks are ineffective or limited due to the low accuracy of extracted models and the high number of queries to the models under attack. The high query cost makes such techniques infeasible for online MLaaS systems that charge per query. We create a novel approach to get higher accuracy and query efficiency than prior data-free model extraction techniques. Specifically, we introduce a novel generator training scheme that maximizes the disagreement loss between two clone models that attempt to copy the model under attack. This loss, combined with diversity loss and experience replay, enables the generator to produce better instances to train the clone models. Our evaluation on popular datasets CIFAR-10 and CIFAR-100 shows that our approach improves the final model accuracy by up to 3.42% and 18.48% respectively. The average number of queries required to achieve the accuracy of the prior state of the art is reduced by up to 64.95%. We hope this will promote future work on feasible data-free model extraction and defenses against such attacks. Jonathan Rosenthal, Eric Enouen, Hung Viet Pham, Lin Tan 0001 |
AAAI | 3 |
| 2023 | How Effective Are Neural Networks for Fixing Security VulnerabilitiesabstractSecurity vulnerability repair is a difficult task that is in dire need of automation. Two groups of techniques have shown promise: (1) large code language models (LLMs) that have been pre-trained on source code for tasks such as code completion, and (2) automated program repair (APR) techniques that use deep learning (DL) models to automatically fix software bugs. Yi Wu 0023, Nan Jiang 0012, Hung Viet Pham, Thibaud Lutellier, Jordan Davis, Lin Tan 0001, Petr Babkin, Sameena Shah |
ISSTA | 3 |
| 2022 | EAGLE: Creating Equivalent Graphs to Test Deep Learning LibrariesabstractTesting deep learning (DL) software is crucial and challenging. Recent approaches use differential testing to cross-check pairs of implementations of the same functionality across different libraries. Such approaches require two DL libraries implementing the same functionality, which is often unavailable. In addition, they rely on a high-level library, Keras, that implements missing functionality in all supported DL libraries, which is prohibitively expensive and thus no longer maintained. Jiannan Wang 0002, Thibaud Lutellier, Shangshu Qian, Hung Viet Pham, Lin Tan 0001 |
ICSE | 4 |
| 2022 | DocTer: documentation-guided fuzzing for testing deep learning API functionsabstractInput constraints are useful for many software development tasks. For example, input constraints of a function enable the generation of valid inputs, i.e., inputs that follow these constraints, to test the function deeper. API functions of deep learning (DL) libraries have DL-specific input constraints, which are described informally in the free-form API documentation. Existing constraint-extraction techniques are ineffective for extracting DL-specific input constraints. Danning Xie, Mijung Kim, Hung Viet Pham, Lin Tan 0001, Xiangyu Zhang 0001, Michael W. Godfrey |
ISSTA | 4 |
| 2021 | DEVIATE: A Deep Learning Variance Testing FrameworkabstractDeep learning (DL) training is nondeterministic and such nondeterminism was shown to cause significant variance of model accuracy (up to 10.8%). Such variance may affect the validity of the comparison of newly proposed DL techniques with baselines. To ensure such validity, DL researchers and practitioners must replicate their experiments multiple times with identical settings to quantify the variance of the proposed approaches and baselines. Replicating and measuring DL variances reliably and efficiently is challenging and understudied.We propose a ready-to-deploy framework DEVIATE that (1) measures DL training variance of a DL model with minimal manual efforts, and (2) provides statistical tests of both accuracy and variance. Specifically, DEVIATE automatically analyzes the DL training code and extracts monitored important metrics (such as accuracy and loss). In addition, DEVIATE performs popular statistical tests and provides users with a report of statistical p-values and effect sizes along with various confidence levels when comparing to selected baselines.We demonstrate the effectiveness of DEVIATE by performing case studies with adversarial training. Specifically, for an adversarial training process that uses the Fast Gradient Signed Method to generate adversarial examples as the training data, DEVIATE measures a max difference of accuracy among 8 identical training runs with fixed random seeds to be up to 5.1%.Tool and demo links: https://github.com/lin-tan/DEVIATE Hung Viet Pham, Mijung Kim, Lin Tan 0001, Yaoliang Yu, Nachiappan Nagappan |
ASE | 1 |
| 2021 | Are My Deep Learning Systems Fair? An Empirical Study of Fixed-Seed TrainingabstractDeep learning (DL) systems have been gaining popularity in critical tasks such as credit evaluation and crime prediction. Such systems demand fairness. Recent work shows that DL software implementations introduce variance: identical DL training runs (i.e., identical network, data, configuration, software, and hardware) with a fixed seed produce different models. Such variance could make DL models and networks violate fairness compliance laws, resulting in negative social impact. In this paper, we conduct the first empirical study to quantify the impact of software implementation on the fairness and its variance of DL systems. Our study of 22 mitigation techniques and five baselines reveals up to 12.6% fairness variance across identical training runs with identical seeds. In addition, most debiasing algorithms have a negative impact on the model such as reducing model accuracy, increasing fairness variance, or increasing accuracy variance. Our literature survey shows that while fairness is gaining popularity in artificial intelligence (AI) related conferences, only 34.4% of the papers use multiple identical training runs to evaluate their approach, raising concerns about their results’ validity. We call for better fairness evaluation and testing protocols to improve fairness and fairness variance of DL systems as well as DL research validity and reproducibility at large. Shangshu Qian, Hung Viet Pham, Thibaud Lutellier, Zeou Hu, Lin Tan 0001, Yaoliang Yu, Jiahao Chen 0001, Sameena Shah |
NeurIPS | 2 |
| 2020 | CoCoNuT: combining context-aware neural translation models using ensemble for program repairabstractAutomated generate-and-validate (GV) program repair techniques (APR) typically rely on hard-coded rules, thus only fixing bugs following specific fix patterns. These rules require a significant amount of manual effort to discover and it is hard to adapt these rules to different programming languages. Thibaud Lutellier, Hung Viet Pham, Lawrence Pang, Moshi Wei, Lin Tan 0001 |
ISSTA | 2 |
| 2020 | Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of VarianceabstractDeep learning (DL) training algorithms utilize nondeterminism to improve models' accuracy and training efficiency. Hence, multiple identical training runs (e.g., identical training data, algorithm, and network) produce different models with different accuracies and training times. In addition to these algorithmic factors, DL libraries (e.g., TensorFlow and cuDNN) introduce additional variance (referred to as implementation-level variance) due to parallelism, optimization, and floating-point computation. Hung Viet Pham, Shangshu Qian, Jiannan Wang 0002, Thibaud Lutellier, Jonathan Rosenthal, Lin Tan 0001, Yaoliang Yu, Nachiappan Nagappan |
ASE | 1 |
| 2019 | CRADLE: cross-backend validation to detect and localize bugs in deep learning librariesabstractDeep learning (DL) systems are widely used in domains including aircraft collision avoidance systems, Alzheimer's disease diagnosis, and autonomous driving cars. Despite the requirement for high reliability, DL systems are difficult to test. Existing DL testing work focuses on testing the DL models, not the implementations (e.g., DL software libraries) of the models. One key challenge of testing DL libraries is the difficulty of knowing the expected output of DL libraries given an input instance. Fortunately, there are multiple implementations of the same DL algorithms in different DL libraries. Thus, we propose CRADLE, a new approach that focuses on finding and localizing bugs in DL software libraries. CRADLE (1) performs cross-implementation inconsistency checking to detect bugs in DL libraries, and (2) leverages anomaly propagation tracking and analysis to localize faulty functions in DL libraries that cause the bugs. We evaluate CRADLE on three libraries (TensorFlow, CNTK, and Theano), 11 datasets (including ImageNet, MNIST, and KGS Go game), and 30 pre-trained models. CRADLE detects 12 bugs and 104 unique inconsistencies, and highlights functions relevant to the causes of inconsistencies for all 104 unique inconsistencies. Hung Viet Pham, Thibaud Lutellier, Weizhen Qi, Lin Tan 0001 |
ICSE | 1 |
| 2016 | Learning API usages from bytecode: a statistical approachabstractMobile app developers rely heavily on standard API frameworks and libraries. However, learning API usages is often challenging due to the fast-changing nature of API frameworks for mobile systems and the insufficiency of API documentation and source code examples. In this paper, we propose a novel approach to learn API usages from bytecode of Android mobile apps. Our core contributions include HAPI, a statistical model of API usages and three algorithms to extract method call sequences from apps' bytecode, to train HAPI based on those sequences, and to recommend method calls in code completion using the trained HAPIs. Our empirical evaluation shows that our prototype tool can effectively learn API usages from 200 thousand apps containing 350 million method sequences. It recommends next method calls with top-3 accuracy of 90% and outperforms baseline approaches on average 10--20%. Tam The Nguyen, Hung Viet Pham, Phong Minh Vu, Tung Thanh Nguyen |
ICSE | 2 |
| 2016 | Phrase-based extraction of user opinions in mobile app reviewsabstractMobile app reviews often contain useful user opinions like bug reports or suggestions. However, looking for those opinions manually in thousands of reviews is inefective and time- consuming. In this paper, we propose PUMA, an automated, phrase-based approach to extract user opinions in app reviews. Our approach includes a technique to extract phrases in reviews using part-of-speech (PoS) templates; a technique to cluster phrases having similar meanings (each cluster is considered as a major user opinion); and a technique to monitor phrase clusters with negative sentiments for their outbreaks over time. We used PUMA to study two popular apps and found that it can reveal severe problems of those apps reported in their user reviews. Phong Minh Vu, Hung Viet Pham, Tam The Nguyen, Tung Thanh Nguyen |
ASE | 2 |
| 2015 | Recommending API Usages for Mobile Apps with Hidden Markov ModelabstractMobile apps often rely heavily on standard API frameworks and libraries. However, learning to use those APIs is often challenging due to the fast-changing nature of API frameworks and the insufficiency of documentation and code examples. This paper introduces DroidAssist, a recommendation tool for API usages of Android mobile apps. The core of DroidAssist is HAPI, a statistical, generative model of API usages based on Hidden Markov Model. With HAPIs trained from existing mobile apps, DroidAssist could perform code completion for method calls. It can also check existing call sequences to detect and repair suspicious (i.e. unpopular) API usages. Tam The Nguyen, Hung Viet Pham, Phong Minh Vu, Tung Thanh Nguyen |
ASE | 2 |
| 2015 | Mining User Opinions in Mobile App Reviews: A Keyword-Based Approach (T)abstractUser reviews of mobile apps often contain complaints or suggestions which are valuable for app developers to improve user experience and satisfaction. However, due to the large volume and noisy-nature of those reviews, manually analyzing them for useful opinions is inherently challenging. To address this problem, we propose MARK, a keyword-based framework for semi-automated review analysis. MARK allows an analyst describing his interests in one or some mobile apps by a set of keywords. It then finds and lists the reviews most relevant to those keywords for further analysis. It can also draw the trends over time of those keywords and detect their sudden changes, which might indicate the occurrences of serious issues. To help analysts describe their interests more effectively, MARK can automatically extract keywords from raw reviews and rank them by their associations with negative reviews. In addition, based on a vector-based semantic representation of keywords, MARK can divide a large set of keywords into more cohesive subsets, or suggest keywords similar to the selected ones. Phong Minh Vu, Tam The Nguyen, Hung Viet Pham, Tung Thanh Nguyen |
ASE | 3 |
| 2015 | Tool Support for Analyzing Mobile App ReviewsabstractMobile app reviews often contain useful user opinions for app developers. However, manual analysis of those reviews is challenging due to their large volume and noisynature. This paper introduces MARK, a supporting tool for review analysis of mobile apps. With MARK, an analyst can describe her interests of one or more apps via a set of keywords. MARK then lists the reviews most relevant to those keywords for further analyses. It can also draw the trends over time of the selected keywords, which might help the analyst to detect sudden changes in the related user reviews. To help the analyst describe her interests more effectively, MARK can automatically extract and rank the keywords by their associations with negative reviews, divide a large set of keywords into more cohesive subgroups, or expand a small set into a broader one. Phong Minh Vu, Hung Viet Pham, Tam The Nguyen, Tung Thanh Nguyen |
ASE | 2 |