EDBT 2026 Demo / reviewers in the wild / expert
Parminder Flora
dblp:07/5582
· DBLP profile ↗
19ranked-venue papers
0as first author
1since 2021 · last 2023
0000-0003-0414-1145ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 1 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
4 papers |
Software maintenance and evolution · 48% Software testing · 18% Empirical software engineering · 18% | |
| Databases, data mining, and information retrieval
3 papers |
Database system architecture and tuning · 57% Data integration and cleaning · 22% Query processing and optimization · 22% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software maintenance and evolution › software performance engineering
performance antipattern detection |
0.4 | 2 | 2016 | Finding and Evaluating the Performance Impact of Redundant Data Access for Applications that are Developed Using Object-Relational Mapping Frameworks · IEEE Trans. Software Eng. 2016 Detecting performance anti-patterns for applications developed using object-relational mapping · ICSE 2014 |
Empirical software engineering › software analytics
performance regression detection |
0.2 | 1 | 2015 | An Industrial Case Study on the Automated Detection of Performance Regressions in Heterogeneous Environments · ICSE (2) 2015 |
Software testing
performance testing |
0.2 | 1 | 2015 | An Industrial Case Study on the Automated Detection of Performance Regressions in Heterogeneous Environments · ICSE (2) 2015 |
Query processing and optimization
database access optimization |
0.2 | 1 | 2014 | Detecting performance anti-patterns for applications developed using object-relational mapping · ICSE 2014 |
Data integration and cleaning › schema mapping
object-relational mapping |
0.2 | 1 | 2014 | Detecting performance anti-patterns for applications developed using object-relational mapping · ICSE 2014 |
Debugging and program repair › performance debugging
performance bug detection |
0.2 | 1 | 2014 | Detecting performance anti-patterns for applications developed using object-relational mapping · ICSE 2014 |
Software maintenance and evolution
software configuration |
0.1 | 1 | 2016 | CacheOptimizer: helping developers configure caching frameworks for hibernate-based database-centric web applications · SIGSOFT FSE 2016 |
Software maintenance and evolution
software evolution |
0.1 | 1 | 2015 | An Industrial Case Study on the Automated Detection of Performance Regressions in Heterogeneous Environments · ICSE (2) 2015 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2015 | An Industrial Case Study on the Automated Detection of Performance Regressions in Heterogeneous Environments · ICSE (2) 2015 |
Methods — techniques the papers use, named apart from their topics
static analysis · 0.9performance profiling · 0.5weighting · 0.4voting · 0.4ensemble modeling · 0.4performance assessment · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Towards reducing the time needed for load testingabstractAbstract The performance of large‐scale systems must be thoroughly tested under various levels of workload, as load‐related issues can have a disastrous impact on the system. However, load testing often requires a large amount of time, running from hours to even days. In our prior work, we reduced the execution time of a load test by detecting repetitiveness in individual performance metric values, such as CPU utilization, that are observed during the test. However, as we explain in this paper, disregarding combinations of performance metrics may miss important information about the load‐related behavior of a system. In this paper we revisit our prior approach, by proposing an approach that reduces the execution time of a load test by detecting whether a test no longer exercises new combinations of the observed performance metrics. We study three open source systems, in which we use our new and prior approaches to reduce the execution time of a 24‐hour load test. We show that our new approach is capable of reducing the execution time of the test to less than 8.5 hours, while preserving a coverage of at least 95% of the combinations that are observed between the performance metrics during the 24‐hour tests. Hammam M. AlGhamdi, Cor-Paul Bezemer, Weiyi Shang, Ahmed E. Hassan, Parminder Flora |
J. Softw. Evol. Process. | 5 |
| 2019 | A CTA-DCD Model to Determine Design Requirements for Technology to Support People with Mild Cognitive Impairment / Dementia at Work
Karan Shastri, Jennifer Boger, Parminder Flora, Arlene Astell, Ann-Charlotte Nedlund, Katja Karjalainen, Anna Mäki-Petäjä-Leinonen, Louise Nygård |
CogSci | 3 |
| 2016 | An empirical study on the practice of maintaining object-relational mapping code in Java systemsabstractDatabases have become one of the most important components in modern software systems. For example, web services, cloud computing systems, and online transaction processing systems all rely heavily on databases. To abstract the complexity of accessing a database, developers make use of Object-Relational Mapping (ORM) frameworks. ORM frameworks provide an abstraction layer between the application logic and the underlying database. Such abstraction layer automatically maps objects in Object-Oriented Languages to database records, which significantly reduces the amount of boilerplate code that needs to be written. Tse-Hsun (Peter) Chen, Weiyi Shang, Jinqiu Yang 0001, Ahmed E. Hassan, Michael W. Godfrey, Mohamed N. Nasser, Parminder Flora |
MSR | 7 |
| 2016 | CacheOptimizer: helping developers configure caching frameworks for hibernate-based database-centric web applicationsabstractTo help improve the performance of database-centric cloud-based web applications, developers usually use caching frameworks to speed up database accesses. Such caching frameworks require extensive knowledge of the application to operate effectively. However, all too often developers have limited knowledge about the intricate details of their own application. Hence, most developers find configuring caching frameworks a challenging and time-consuming task that requires extensive and scattered code changes. Furthermore, developers may also need to frequently change such configurations to accommodate the ever changing workload. Tse-Hsun (Peter) Chen, Weiyi Shang, Ahmed E. Hassan, Mohamed N. Nasser, Parminder Flora |
SIGSOFT FSE | 5 |
| 2016 | Finding and Evaluating the Performance Impact of Redundant Data Access for Applications that are Developed Using Object-Relational Mapping FrameworksabstractDevelopers usually leverage Object-Relational Mapping (ORM) to abstract complex database accesses for large-scale systems. However, since ORM frameworks operate at a lower-level (i.e., data access), ORM frameworks do not know how the data will be used when returned from database management systems (DBMSs). Therefore, ORM cannot provide an optimal data retrieval approach for all applications, which may result in accessing redundant data and significantly affect system performance. Although ORM frameworks provide ways to resolve redundant data problems, due to the complexity of modern systems, developers may not be able to locate such problems in the code; hence, may not proactively resolve the problems. In this paper, we propose an automated approach, which we implement as a Java framework, to locate redundant data problems. We apply our framework on one enterprise and two open source systems. We find that redundant data problems exist in 87 percent of the exercised transactions. Due to the large number of detected redundant data problems, we propose an automated approach to assess the impact and prioritize the resolution efforts. Our performance assessment result shows that by resolving the redundant data problems, the system response time for the studied systems can be improved by an average of 17 percent. Tse-Hsun (Peter) Chen, Weiyi Shang, Zhen Ming (Jack) Jiang, Ahmed E. Hassan, Mohamed N. Nasser, Parminder Flora |
IEEE Trans. Software Eng. | 6 |
| 2015 | An Industrial Case Study on the Automated Detection of Performance Regressions in Heterogeneous EnvironmentsabstractA key goal of performance testing is the detection of performance degradations (i.e., regressions) compared to previous releases. Prior research has proposed the automation of such analysis through the mining of historical performance data (e.g., CPU and memory usage) from prior test runs. Nevertheless, such research has had limited adoption in practice. Working with a large industrial performance testing lab, we noted that a major hurdle in the adoption of prior work (including our own work) is the incorrect assumption that prior tests are always executed in the same environment (i.e., labs). All too often, tests are performed in heterogenous environments with each test being run in a possibly different lab with different hardware and software configurations. To make automated performance regression analysis techniques work in industry, we propose to model the global expected behaviour of a system as an ensemble (combination) of individual models, one for each successful previous test run (and hence configuration). The ensemble of models of prior test runs are used to flag performance deviations (e.g., CPU counters showing higher usage) in new tests. The deviations are then aggregated using simple voting or more advanced weighting to determine whether the counters really deviate from the expected behaviour or whether it was simply due to an environment-specific variation. Case studies on two open-source systems and a very large scale industrial application show that our weighting approach outperforms a state-of-the-art environment-agnostic approach. Feedback from practitioners who used our approach over a 4 year period (across several major versions) has been very positive. King Chun Foo, Zhen Ming (Jack) Jiang, Bram Adams, Ahmed E. Hassan, Ying Zou 0001, Parminder Flora |
ICSE (2) | 6 |
| 2015 | Automated Detection of Performance Regressions Using Regression Models on Clustered Performance CountersabstractPerformance testing is conducted before deploying system updates in order to ensure that the performance of large software systems did not degrade (i.e., no performance regressions). During such testing, thousands of performance counters are collected. However, comparing thousands of performance counters across versions of a software system is very time consuming and error-prone. In an effort to automate such analysis, model-based performance regression detection approaches build a limited number (i.e., one or two) of models for a limited number of target performance counters (e.g., CPU or memory) and leverage the models to detect performance regressions. Such model-based approaches still have their limitations since selecting the target performance counters is often based on experience or gut feeling. In this paper, we propose an automated approach to detect performance regressions by analyzing all collected counters instead of focusing on a limited number of target counters. We first group performance counters into clusters to determine the number of performance counters needed to truly represent the performance of a system. We then perform statistical tests to select the target performance counters, for which we build regression models. We apply the regression models on new version of the system to detect performance regressions. Weiyi Shang, Ahmed E. Hassan, Mohamed N. Nasser, Parminder Flora |
ICPE | 4 |
| 2014 | Detecting performance anti-patterns for applications developed using object-relational mappingabstractObject-Relational Mapping (ORM) provides developers a conceptual abstraction for mapping the application code to the underlying databases. ORM is widely used in industry due to its convenience; permitting developers to focus on developing the business logic without worrying too much about the database access details. However, developers often write ORM code without considering the impact of such code on database performance, leading to cause transactions with timeouts or hangs in large-scale systems. Unfortunately, there is little support to help developers automatically detect suboptimal database accesses. In this paper, we propose an automated framework to detect ORM performance anti-patterns. Our framework automatically flags performance anti-patterns in the source code. Furthermore, as there could be hundreds or even thousands of instances of anti-patterns, our framework provides sup- port to prioritize performance bug fixes based on a statistically rigorous performance assessment. We have successfully evaluated our framework on two open source and one large-scale industrial systems. Our case studies show that our framework can detect new and known real-world performance bugs and that fixing the detected performance anti- patterns can improve the system response time by up to 98%. Tse-Hsun (Peter) Chen, Weiyi Shang, Zhen Ming (Jack) Jiang, Ahmed E. Hassan, Mohamed N. Nasser, Parminder Flora |
ICSE | 6 |
| 2014 | An industrial case study of automatically identifying performance regression-causesabstractEven the addition of a single extra field or control statement in the source code of a large-scale software system can lead to performance regressions. Such regressions can considerably degrade the user experience. Working closely with the members of a performance engineering team, we observe that they face a major challenge in identifying the cause of a performance regression given the large number of performance counters (e.g., memory and CPU usage) that must be analyzed. We propose the mining of a regression-causes repository (where the results of performance tests and causes of past regressions are stored) to assist the performance team in identifying the regression-cause of a newly-identified regression. We evaluate our approach on an open-source system, and a commercial system for which the team is responsible. The results show that our approach can accurately (up to 80% accuracy) identify performance regression-causes using a reasonably small number of historical test runs (sometimes as few as four test runs per regression-cause). Thanh H. D. Nguyen, Meiyappan Nagappan, Ahmed E. Hassan, Mohamed N. Nasser, Parminder Flora |
MSR | 5 |
| 2014 | Continuous validation of load test suitesabstractUltra-Large-Scale (ULS) systems face continuously evolving field workloads in terms of activated/disabled feature sets, varying usage patterns and changing deployment configurations. These evolving workloads often have a large impact, on the performance of a ULS system. Hence, continuous load testing is critical to ensuring the error-free operation of such systems. A common challenge facing performance analysts is to validate if a load test closely resembles the current field workloads. Such validation may be performed by comparing execution logs from the load test and the field. However, the size and unstructured nature of execution logs makes such a comparison unfeasible without automated support. In this paper, we propose an automated approach to validate whether a load test resembles the field workload and, if not, determines how they differ by compare execution logs from a load test and the field. Performance analysts can then update their load test cases to eliminate such differences, hence creating more realistic load test cases. We perform three case studies on two large systems: one open-source system and one enterprise system. Our approach identifies differences between load tests and the field with a precision of >75% compared to only >16% for the state-of-the-practice. Mark D. Syer, Zhen Ming (Jack) Jiang, Meiyappan Nagappan, Ahmed E. Hassan, Mohamed N. Nasser, Parminder Flora |
ICPE | 6 |
| 2014 | An exploratory study of the evolution of communicated information about the execution of large software systemsabstractSUMMARY Substantial research in software engineering focuses on understanding the dynamic nature of software systems in order to improve software maintenance and program comprehension. This research typically makes use of automated instrumentation and profiling techniques after the fact, that is, without considering domain knowledge. In this paper, we examine another source of dynamic information that is generated from statements that have been inserted into the code base during development to draw the system administrators' attention to important run‐time phenomena. We call this source communicated information (CI). Examples of CI include execution logs and system events. The availability of CI has sparked the development of an ecosystem ofLog Processing Apps(LPAs) that surround the software system under analysis to monitor and document various run‐time constraints. The dependence of LPAs on the timeliness, accuracy and granularity of the CI means that it is important to understand the nature of CI and how it evolves over time, both qualitatively and quantitatively. Yet, to our knowledge, little empirical analysis has been performed on CI and its evolution. In a case study on two large open source and one industrial software systems, we explore the evolution of CI by mining the execution logs of these systems and the logging statements in the source code. Our study illustrates the need for better traceability between CI and the LPAs that analyze the CI. In particular, we find that the CI changes at a high rate across versions, which could lead to fragile LPAs. We found that up to 70% of these changes could have been avoided and the impact of 15% to 80% of the changes can be controlled through the use of robust analysis techniques by LPAs. We also found that LPAs that track implementation‐level CI (e.g. performance analysis) and the LPAs that monitor error messages (system health monitoring) are more fragile than LPAs that track domain‐level CI (e.g. workload modelling), because the latter CI tends to be long‐lived. Copyright © 2013 John Wiley & Sons, Ltd. Weiyi Shang, Zhen Ming (Jack) Jiang, Bram Adams, Ahmed E. Hassan, Michael W. Godfrey, Mohamed N. Nasser, Parminder Flora |
J. Softw. Evol. Process. | 7 |
| 2013 | Leveraging Performance Counters and Execution Logs to Diagnose Memory-Related Performance IssuesabstractLoad tests ensure that software systems are able to perform under the expected workloads. The current state of load test analysis requires significant manual review of performance counters and execution logs, and a high degree of system-specific expertise. In particular, memory-related issues (e.g., memory leaks or spikes), which may degrade performance and cause crashes, are difficult to diagnose. Performance analysts must correlate hundreds of megabytes or gigabytes of performance counters (to understand resource usage) with execution logs (to understand system behaviour). However, little work has been done to combine these two types of information to assist performance analysts in their diagnosis. We propose an automated approach that combines performance counters and execution logs to diagnose memory-related issues in load tests. We perform three case studies on two systems: one open-source system and one large-scale enterprise system. Our approach flags ≤ 0.1% of the execution logs with a precision ≥ 80%. Mark D. Syer, Zhen Ming (Jack) Jiang, Meiyappan Nagappan, Ahmed E. Hassan, Mohamed N. Nasser, Parminder Flora |
ICSM | 6 |
| 2012 | Automated detection of performance regressions using statistical process control techniquesabstractThe goal of performance regression testing is to check for performance regressions in a new version of a software system. Performance regression testing is an important phase in the software development process. Performance regression testing is very time consuming yet there is usually little time assigned for it. A typical test run would output thousands of performance counters. Testers usually have to manually inspect these counters to identify performance regressions. In this paper, we propose an approach to analyze performance counters across test runs using a statistical process control technique called control charts. We evaluate our approach using historical data of a large software team as well as an open-source software project. The results show that our approach can accurately identify performance regressions in both software systems. Feedback from practitioners is very promising due to the simplicity and ease of explanation of the results. Thanh H. D. Nguyen, Bram Adams, Zhen Ming (Jack) Jiang, Ahmed E. Hassan, Mohamed N. Nasser, Parminder Flora |
ICPE | 6 |
| 2011 | Automated Verification of Load Tests Using Control ChartsabstractLoad testing is an important phase in the software development process. It is very time consuming but there is usually little time for it. As a solution to the tight testing schedule, software companies automate their testing procedures. However, existing automation only reduces the time required to run load tests. The analysis of the test results is still performed manually. A typical load test outputs thousands of performance counters. Analyzing these counters manually requires time and tacit knowledge of the system-under-test from the performance engineers. The goal of this study is to derive an approach to automatically verify load tests' results. We propose an approach based on a statistical quality control technique called control charts. Our approach can a) automatically determine if a test run passes or fails and b) identify the subsystem where performance problem originated. We conduct two case studies on a large commercial telecommunication software and an open-source software system to evaluate our approach. Our results warrant further development of control chart based techniques in performance verification. Thanh H. D. Nguyen, Bram Adams, Zhen Ming (Jack) Jiang, Ahmed E. Hassan, Mohamed N. Nasser, Parminder Flora |
APSEC | 6 |
| 2010 | Using Load Tests to Automatically Compare the Subsystems of a Large Enterprise SystemabstractEnterprise systems are load tested for every added feature, software updates and periodic maintenance to ensure that the performance demands on system quality, availability and responsiveness are met. In current practice, performance analysts manually analyze load test data to identify the components that are responsible for performance deviations. This process is time consuming and error prone due to the large volume of performance counter data collected during monitoring, the limited operational knowledge of analyst about all the subsystem involved and their complex interactions and the unavailability of up-to-date documentation in the rapidly evolving enterprise. In this paper, we present an automated approach based on a robust statistical technique, Principal Component Analysis (PCA) to identify subsystems that show performance deviations in load tests. A case study on load test data of a large enterprise application shows that our approach do not require any instrumentation or domain knowledge to operate, scales well to large industrial system, generate few false positives (89% average precision) and detects performance deviations among subsystems in limited time. Haroon Malik, Bram Adams, Ahmed E. Hassan, Parminder Flora, Gilbert Hamann |
COMPSAC | 4 |
| 2009 | Automated performance analysis of load testsabstractThe goal of a load test is to uncover functional and performance problems of a system under load. Performance problems refer to the situations where a system suffers from unexpectedly high response time or low throughput. It is difficult to detect performance problems in a load test due to the absence of formally defined performance objectives and the large amount of data that must be examined. In this paper, we present an approach which automatically analyzes the execution logs of a load test for performance problems. We first derive the system's performance baseline from previous runs. Then we perform an in-depth performance comparison against the derived performance baseline. Case studies show that our approach produces few false alarms (with a precision of 77%) and scales well to large industrial systems. Zhen Ming (Jack) Jiang, Ahmed E. Hassan, Gilbert Hamann, Parminder Flora |
ICSM | 4 |
| 2008 | Automatic identification of load testing problemsabstractMany software applications must provide services to hundreds or thousands of users concurrently. These applications must be load tested to ensure that they can function correctly under high load. Problems in load testing are due to problems in the load environment, the load generators, and the application under test. It is important to identify and address these problems to ensure that load testing results are correct and these problems are resolved. It is difficult to detect problems in a load test due to the large amount of data which must be examined. Current industrial practice mainly involves time-consuming manual checks which, for example, grep the logs of the application for error messages. In this paper, we present an approach which mines the execution logs of an application to uncover the dominant behavior (i.e., execution sequences) for the application and flags anomalies (i.e., deviations) from the dominant behavior. Using a case study of two open source and two large enterprise software applications, we show that our approach can automatically identify problems in a load test. Our approach flags < 0.01% of the log lines for closer analysis by domain experts. The flagged lines indicate load testing problems with a relatively small number of false alarms. Our approach scales well for large applications and is currently used daily in practice. Zhen Ming (Jack) Jiang, Ahmed E. Hassan, Gilbert Hamann, Parminder Flora |
ICSM | 4 |
| 2008 | Retrieving relevant reports from a customer engagement repositoryabstractCustomers of modern enterprise applications commonly engage the vendor of the application for on-site troubleshooting and fine tuning of large deployments. The results of these engagements are documented in customer engagement reports. The reports contain valuable information about the observed symptoms, identified problems, attempted workarounds and the final solution. Such information is valuable in supporting analysts in future engagements. Engagement reports are stored in a customer engagement repository. Retrieving relevant reports from such a repository is usually ad-hoc and is based on using basic text search. We present a technique to retrieve relevant reports from an engagement repository. The technique takes as input an execution log for a particular deployment and retrieves relevant engagement reports. The technique identifies relevant reports by comparing execution logs attached to the report stored in the engagement repository. The technique returns two types of relevant reports: (1) reports for engagement with similar operational profiles to identify prior engagements with similar workloads and (2) reports for engagement with similar signature profiles to identify prior engagements with similar problems. Using our technique, support analysts can locate relevant engagement reports and use the knowledge in them to quickly resolve problems at hand. To demonstrate the feasibility of our technique, we present two case studies: one case study uses an industry standard open source application, the Dell DVD Store, while the other case study uses a large enterprise application. The results of our case study show that our technique performs well in identifying relevant reports with high precision and high recall. Dharmesh Thakkar, Zhen Ming (Jack) Jiang, Ahmed E. Hassan, Gilbert Hamann, Parminder Flora |
ICSM | 5 |
| 2008 | An automated approach for abstracting execution logs to execution eventsabstractAbstract Execution logs are generated by output statements that developers insert into the source code. Execution logs are widely available and are helpful in monitoring, remote issue resolution, and system understanding of complex enterprise applications. There are many proposals for standardized log formats such as the W3C and SNMP formats. However, most applications usead hocnon‐standardized logging formats. Automated analysis of such logs is complex due to the loosely defined structure and a large non‐fixed vocabulary of words. The large volume of logs, produced by enterprise applications, limits the usefulness of manual analysis techniques. Automated techniques are needed to uncover the structure of execution logs. Using the uncovered structure, sophisticated analysis of logs can be performed. In this paper, we propose a log abstraction technique that recognizes the internal structure of each log line. Using the recovered structure, log lines can be easily summarized and categorized to help comprehend and investigate the complex behavior of large software applications. Our proposed approach handles free‐form log lines with minimal requirements on the format of a log line. Through a case study using log files from four enterprise applications, we demonstrate that our approach abstracts log files of different complexities with high precision and recall. Copyright © 2008 John Wiley & Sons, Ltd. Zhen Ming (Jack) Jiang, Ahmed E. Hassan, Gilbert Hamann, Parminder Flora |
J. Softw. Maintenance Res. Pract. | 4 |