VLDB 2026 Research / reviewers in the wild / expert
Pak-Lok Poon
dblp:67/1531
· DBLP profile ↗
26ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-2840-2418ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 22 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MT-Boost: A metamorphic-testing based training method for enhancing the robustness of deep neural network classifiersabstractContext: In metamorphic testing (MT), a set of metamorphic relations (MRs) are identified to verify whether or not a trained deep neural network (DNN) can produce consistent performance when specific transformations are applied to its input. Most DNNs trained with existing methods often perform poorly with respect to MRs, thereby indicating that these DNNs are not robust. Objective: To improve DNN’s performance in the context of MT, a set of defined MRs is used to generate training inputs to retrain a DNN model. Our main objective is to develop a method to balance a DNN’s accuracy and robustness with less time consumption and having the capability to cater to multiple MRs. Methods: In this paper, we introduce our regularization-based method (known as MT-Boost), which uses reinforcement learning to search for the best way of using MRs to generate inputs and express them as loss function regularizers. When developing MT-Boost, we transform the robustness-improving problem into a reinforcement-learning agent’s training problem. Results: MT-Boost is evaluated on eight DNN models with four popular datasets. MT-Boost achieves the largest robustness improvement for each model and maintains relatively high accuracy performance when compared with seven other baseline methods. Our sensitivity analysis also shows the high stability performance of MT-Boost across four reinforcement-learning algorithms and other hyperparameters. Conclusion: Experimental results show that MT-Boost is effective and efficient for improving DNN’s robustness. Kun Qiu 0001, Yu Zhou 0067, Pak-Lok Poon, Tsong Yueh Chen |
Inf. Softw. Technol. | 3 |
| 2025 | Evaluating the effectiveness of neuron coverage metrics: a metamorphic-testing approach
Zenghui Zhou, Pak-Lok Poon, Tsong Yueh Chen, Kun Qiu 0001, Zheng Zheng 0001 |
Softw. Qual. J. | 2 |
| 2025 | Metamorphic Relation Generation: State of the Art and Research DirectionsabstractMetamorphic testing has become one mainstream technique to address the notorious oracle problem in software testing, thanks to its great successes in revealing real-life bugs in a wide variety of software systems. Metamorphic relations, the core component of metamorphic testing, have continuously attracted research interests from both academia and industry. In the last decade, a rapidly increasing number of studies have been conducted to systematically generate metamorphic relations from various sources and for different application domains. In this article, based on the systematic review on the state of the art for metamorphic relations’ generation, we summarize and highlight visions for further advancing the theory and techniques for identifying and constructing metamorphic relations and discuss promising research directions in related areas. Rui Li 0013, Huai Liu, Pak-Lok Poon, Dave Towey, Chang-Ai Sun, Zheng Zheng 0001, Zhiquan Zhou 0001, Tsong Yueh Chen |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2024 | Spreadsheet quality assurance: a literature reviewabstractAbstract Spreadsheets are very common for information processing to support decision making by both professional developers and non-technical end users. Moreover, business intelligence and artificial intelligence are increasingly popular in the industry nowadays, where spreadsheets have been used as, or integrated into, intelligent or expert systems in various application domains. However, it has been repeatedly reported that faults often exist in operational spreadsheets, which could severely compromise the quality of conclusions and decisions based on the spreadsheets. With a view to systematically examining this problem via survey of existing work, we have conducted a comprehensive literature review on the quality issues and related techniques of spreadsheets over a 35.5-year period (from January 1987 to June 2022) for target journals and a 10.5-year period (from January 2012 to June 2022) for target conferences. Among other findings, two major ones are: (a) Spreadsheet quality is best addressed throughout the whole spreadsheet life cycle, rather than just focusing on a few specific stages of the life cycle. (b) Relatively more studies focus on spreadsheet testing and debugging (related to fault detection and removal) when compared with spreadsheet specification, modeling, and design (related to development). As prevention is better than cure, more research should be performed on the early stages of the spreadsheet life cycle. Enlightened by our comprehensive review, we have identified the major research gaps as well as highlighted key research directions for future work in the area. Pak-Lok Poon, Man Fai Lau, Yuen-Tak Yu, Sau-Fun Tang |
Frontiers Comput. Sci. | 1 |
| 2024 | Improving the validation of multiple-object detection using a complex-network-community-based relevance metricabstractAlthough many of today’s object detectors (ODs) are fairly powerful and advanced, most of them still suffer from high detection failure rates. To address this issue, we have developed an innovative, multiple-object detection validation method using a complex-network-community-based relevance metric. This metric aims to measure the relevance of multiple objects in the same OD output, based on our observation that a faulty OD output generally includes objects that are irrelevant or unrelated to each other. To verify the effectiveness of our method, we formulated four research questions, and performed an experiment with statistical analyses to address these questions. Our experiment provides strong support that our method (particularly the relevance metric) is highly effective at helping human testers in identifying faulty OD outputs. Kun Qiu 0001, Pak-Lok Poon, Shijun Zhao, Dave Towey, Lanlin Yu |
Knowl. Based Syst. | 2 |
| 2022 | Theoretical and Empirical Analyses of the Effectiveness of Metamorphic Relation CompositionabstractMetamorphic Relations (MRs) play a key role in determining the fault detection capability of Metamorphic Testing (MT). As human judgement is required for MR identification, systematic MR generation has long been an important research area in MT. Additionally, due to the extra program executions required for follow-up test cases, some concerns have been raised about MT cost-effectiveness. Consequently, the reduction in testing costs associated with MT has become another important issue to be addressed. MR composition can address both of these problems. This technique can automatically generate new MRs by composing existing ones, thereby reducing the number of follow-up test cases. Despite this advantage, previous studies on MR composition have empirically shown that some composite MRs have lower fault detection capability than their corresponding component MRs. To investigate this issue, we performed theoretical and empirical analyses to identify what characteristics component MRs should possess so that their corresponding composite MR has at least the same fault detection capability as the component MRs do. We have also derived a convenient, but effective guideline so that the fault detection capability of MT will most likely not be reduced after composition. Kun Qiu 0001, Zheng Zheng 0001, Tsong Yueh Chen, Pak-Lok Poon |
IEEE Trans. Software Eng. | 4 |
| 2021 | METRIC$^{+}$+: A Metamorphic Relation Identification Technique Based on Input Plus Output DomainsabstractMetamorphic testing is well known for its ability to alleviate the oracle problem in software testing. The main idea ofmetamorphic testing is to test a software system by checking whether each identified metamorphic relation (MR) holds among severalexecutions. In this regard, identifying MRs is an essential task in metamorphic testing. In view of the importance of this identificationtask, METRIC (METamorphic Relation Identification based on Category-choice framework) was developed to help software testersidentify MRs from a given set of complete test frames. However, during MR identification, METRIC primarily focuses on the inputdomain without sufficient attention given to the output domain, thereby hindering the effectiveness of METRIC. Inspired by this problem,we have extended METRIC into METRIC+by incorporating the information derived from the output domain for MR identification. A toolimplementing METRIC+has also been developed. Two rounds of experiments, involving four real-life specifications, have beenconducted to evaluate the effectiveness and efficiency of METRIC+. The results have confirmed that METRIC+is highly effective andefficient in MR identification. Additional experiments have been performed to compare the fault detection capability of the MRsgenerated by METRIC+and those bymMT (another MR identification technique). The comparison results have confirmed that the MRsgenerated by METRIC+are highly effective in fault detection. Chang-Ai Sun, An Fu, Pak-Lok Poon, Xiaoyuan Xie, Huai Liu, Tsong Yueh Chen |
IEEE Trans. Software Eng. | 3 |
| 2020 | METTLE: A METamorphic Testing Approach to Assessing and Validating Unsupervised Machine Learning SystemsabstractUnsupervised machine learning is the training of an artificial intelligence system using information that is neither classified nor labeled, with a view to modeling the underlying structure or distribution in a dataset. Since unsupervised machine learning systems are widely used in many real-world applications, assessing the appropriateness of these systems and validating their implementations with respect to individual users' requirements and specific application scenarios/contexts are indisputably two important tasks. Such assessments and validation tasks, however, are fairly challenging due to the absence of a priori knowledge of the data. In view of this challenge, in this article, we develop a METamorphic Testing approach to assessing and validating unsupervised machine LEarning systems, abbreviated as mettle. Our approach provides a new way to unveil the (possibly latent) characteristics of various machine learning systems, by explicitly considering the specific expectations and requirements of these systems from individual users' perspectives. To support mettle, we have further formulated 11 generic metamorphic relations (MRs), covering users' generally expected characteristics that should be possessed by machine learning systems. We have performed an experiment and a user evaluation study to evaluate the viability and effectiveness of mettle. Our experiment and user evaluation study have shown that, guided by user-defined MR-based adequacy criteria, end users are able to assess, validate, and select appropriate clustering systems in accordance with their own specific needs. Our investigation has also yielded insightful understanding and interpretation of the behavior of the machine learning systems from an end-user software engineering's perspective, rather than a designer's or implementor's perspective, who normally adopts a theoretical approach. Xiaoyuan Xie, Zhiyi Zhang 0005, Tsong Yueh Chen, Yang Liu 0003, Pak-Lok Poon, Baowen Xu |
IEEE Trans. Reliab. | 5 |
| 2018 | Introduction to the special issue on test oracles
Zhiquan Zhou 0001, Dave Towey, Pak-Lok Poon, T. H. Tse |
J. Syst. Softw. | 3 |
| 2016 | METRIC: METamorphic Relation Identification based on the Category-choice framework
Tsong Yueh Chen, Pak-Lok Poon, Xiaoyuan Xie |
J. Syst. Softw. | 2 |
| 2015 | Poster: Enhancing Partition Testing through Output VariationabstractA major test case generation approach is to divide the input domain into disjoint partitions, from which test cases can be selected. However, we observe that in some traditional approaches to partition testing, the same partition may be associated with different output scenarios. Such an observation implies that the partitioning of the input domain may not be precise enough for effective software fault detection. To solve this problem, partition testing should be fine-tuned to additionally use the information of output scenarios in test case generation, such that these test cases are more fine-grained not only with respect to the input partitions but also from the perspective of output scenarios. Huai Liu, Pak-Lok Poon, Tsong Yueh Chen |
ICSE (2) | 2 |
| 2012 | An empirical evaluation of several test-a-few strategies for testing particular conditionsabstractSUMMARY Existing specification‐based testing techniques often generate comprehensive test suites to cover diverse combinations of test‐relevant aspects. Such a test suite can be prohibitively expensive to execute exhaustively because of its large size. A pragmatic strategy often adopted in practice, called test‐once strategy, is to identify certain particular conditions from the specification and to test each such condition once only. This strategy is implicitly based on the uniformity assumption that the implementation will process a particular condition uniformly, regardless of other parameters or inputs. As the decision of adopting the test‐once strategy is often based on the specification, whether the uniformity assumption actually holds in the implementation needs to be critically assessed, or else the risk of inadequate testing could be non‐negligible. As viable alternatives to reduce such a risk, a family of test‐a‐few strategies for the testing of particular conditions is proposed in this paper. Two rounds of experiments that evaluate the effectiveness of the test‐a‐few strategies as compared with the test‐once strategy are further reported. Our experiments do the following: (1) provide clear evidence that the uniformity assumption often, but not always, holds and that the assumption usually fails to hold when the implementation is faulty; (2) demonstrate that all our proposed test‐a‐few strategies are statistically more reliable than the test‐once strategy in revealing faulty programs; (3) show that random sampling is already substantially more effective than the test‐once strategy; and (4) indicate that, compared with other test‐a‐few strategies under study, choice coverage seems to achieve a better trade‐off between test effort and effectiveness. Copyright © 2011 John Wiley & Sons, Ltd. Eric Ying Kwong Chan, Wing Kwong Chan, Pak-Lok Poon, Yuen-Tak Yu |
Softw. Pract. Exp. | 3 |
| 2012 | DESSERT: a DividE-and-conquer methodology for identifying categorieS, choiceS, and choicE Relations for Test case generationabstractThis paper extends the choce relation framework, abbreviated as choc'late, which assists software testers in the application of category/choice methods to testing. choc'late assumes that the tester is able to construct a single choice relation table from the entire specification; this table then forms the basis for test case generation using the associated algorithms. This assumption, however, may not hold true when the specification is complex and contains many specification components. For such a specification, the tester may construct a preliminary choice relation table from each specification component, and then consolidate all the preliminary tables into a final table to be processed by choc'late for test case generation. However, it is often difficult to merge these preliminary tables because such merging may give rise to inconsistencies among choice relations or overlaps among choices. To alleviate this problem, we introduce a DividE-and-conquer methodology for identifying categorieS, choiceS, and choicE Relations for Test case generation, abbreviated as dessert. The theoretical framework and the associated algorithms are discussed. To demonstrate the viability and effectiveness of our methodology, we describe case studies using the specifications of three real-life commercial software systems. Tsong Yueh Chen, Pak-Lok Poon, Sau-Fun Tang, T. H. Tse |
IEEE Trans. Software Eng. | 2 |
| 2011 | Contributions of tester experience and a checklist guideline to the identification of categories and choices for software testingabstractAn early step for most black-box testing methods is to identify a set of categories and choices (or their equivalents) from the specification. The identification is often performed in an ad hoc manner, thus the quality of categories and choices is in doubt. Poorly identified categories and choices will affect the comprehensiveness of test cases. In this paper, we describe several comparative studies using three commercial specifications and discuss the major results. The objectives of our studies are (a) to investigate the differences in the types and amounts of mistakes made between inexperienced and experienced software testers in an ad hoc identification approach and (b) to determine the extent of mistake reduction after discussing the mistakes with the software testers and providing them with an identification checklist. Pak-Lok Poon, T. H. Tse, Sau-Fun Tang, Fei-Ching Kuo |
Softw. Qual. J. | 1 |
| 2010 | Investigating ERP systems procurement practice: Hong Kong and Australian experiences
Pak-Lok Poon, Yuen-Tak Yu |
Inf. Softw. Technol. | 1 |
| 2009 | On an Ant Colony-Based Approach for Business Fraud Detection
Ou Liu, Jian Ma 0008, Pak-Lok Poon, Jun Zhang 0003 |
ICIC (1) | 3 |
| 2006 | Procurement of enterprise resource planning systems: experiences with some Hong Kong companiesabstractMany cases of adoption of Enterprise Resource Planning (ERP) systems have been reported in the literature. Some of the adopted ERP systems fail to satisfy the customer's requirements, despite the high spending and substantial efforts that have been put into the adoption exercise. This is undoubtedly unsatisfactory. A way to avoid this problem is to adopt a well planned, managed, and controlled ERP procurement process. This paper describes our studies of three Chinese companies in Hong Kong which have adopted ERP systems. We report the experience of these companies, and discuss how the Chinese culture might have shaped the procurement practices in their ERP adoption exercises. Pak-Lok Poon, Yuen-Tak Yu |
ICSE | 1 |
| 2004 | On the Testing of Particular Input ConditionsabstractGenerating test cases from a specification can be done at an early stage. However, so many important aspects relevant to testing can be identified from the specification that exhaustively testing their combinations can be very costly. A common approach to reduce testing costs is to identify some particular input conditions and test each of them only once. We argue that such an approach should be used judiciously, or else inadequate tests may result. This paper explores several alternatives to assess the validity of the tester's hypothesis that a particular condition can be tested adequately with only one test case. These alternatives help to test the particular conditions more reliably and, hence, reduce the risk of not revealing the existence of faults. Eric Ying Kwong Chan, Pak-Lok Poon, Yuen-Tak Yu |
COMPSAC | 2 |
| 2004 | On the identification of categories and choices for specification-based test case generation
Tsong Yueh Chen, Pak-Lok Poon, Sau-Fun Tang, T. H. Tse |
Inf. Softw. Technol. | 2 |
| 2004 | On the testing methods used by beginning software testers
Yuen-Tak Yu, Sebastian Ng, Pak-Lok Poon, Tsong Yueh Chen |
Inf. Softw. Technol. | 3 |
| 2003 | A Choice Relation Framework for Supporting Category-Partition Test Case GenerationabstractWe describe in this paper a choice relation framework for supporting category-partition test case generation. We capture the constraints among various values (or ranges of values) of the parameters and environment conditions identified from the specification, known formally as choices. We express these constraints in terms of relations among choices and combinations of choices, known formally as test frames. We propose a theoretical backbone and techniques for consistency checks and automatic deductions of relations. Based on the theory, algorithms have been developed for generating test frames from the relations. These test frames can then be used as the basis for generating test cases. Our algorithms take into consideration the resource constraints specified by software testers, thus maintaining the effectiveness of the test frames (and hence test cases) generated. Tsong Yueh Chen, Pak-Lok Poon, T. H. Tse |
IEEE Trans. Software Eng. | 2 |
| 2001 | A Study on a Path-based Strategy for Selecting Black-box Generated Test CasesabstractVarious black-box methods for the generation of test cases have been proposed in the literature. Many of these methods, including the category-partition method and the classification-tree method, follow the approach of partition testing, in which the input domain is partitioned into subdomains according to important aspects of the specification, and test cases are then derived from the subdomains. Though comprehensive in terms of these important aspects, execution of all the test cases so generated may not be feasible under the constraint of tight testing resources. In such circumstances, there is a need to select a smaller subset of test cases from the original test suite for execution. In this paper, we propose the use of white-box information to guide the selection of test cases from the original test suite generated by a black-box testing method. Furthermore, we have developed some techniques and algorithms to facilitate the implementation of our approach, and demonstrated its viability and benefits by means of a case study. Yuen-Tak Yu, Sau-Fun Tang, Pak-Lok Poon, Tsong Yueh Chen |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2000 | An Integrated Classification-Tree Methodology for Test Case GenerationabstractThis paper describes an integrated methodology for the construction of test cases from functional specifications using the classification-tree method. It is an integration of our extensions to the classification-hierarchy table, the classification tree construction algorithm, and the classification tree restructuring technique. Based on the methodology, a prototype system ADDICT, which stands for AutomateD test Data generation system using the Integrated Classification-Tree method, has been built. Tsong Yueh Chen, Pak-Lok Poon, T. H. Tse |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 1998 | On the effectiveness of classification trees for test case construction
Tsong Yueh Chen, Pak-Lok Poon |
Inf. Softw. Technol. | 2 |
| 1997 | Construction of classification trees via the classification-hierarchy table
Tsong Yueh Chen, Pak-Lok Poon |
Inf. Softw. Technol. | 2 |
| 1996 | Improving the Quality of Classification Trees via RestructuringabstractThe classification-hierarchy table developed by Chen and Poon (1996) provides a systematic approach to construct classification trees from given sets of classifications and their associated classes. The paper enhances their study by defining a metric to measure the "quality" of a classification tree, and providing an algorithm to improve this quality. Tsong Yueh Chen, Pak-Lok Poon |
APSEC | 2 |