EDBT 2026 Demo / reviewers in the wild / expert
Khaled El Emam
dblp:25/575
· DBLP profile ↗
56ranked-venue papers
29as first author
2since 2021 · last 2025
0000-0003-3325-4149ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 41 · 20 first-authorApplied, interdisciplinary, general and emerging computing · 15 · 9 first-author · 2 since 2021Artificial intelligence and machine learning · 2Human-computer interaction and ubiquitous computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
15 papers |
Empirical software engineering · 43% Software maintenance and evolution · 36% Requirements engineering and software design · 16% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software maintenance and evolution
software inspection |
0.1 | 4 | 2001 | An Internally Replicated Quasi-Experimental Comparison of Checklist and Perspective-Based Reading of Code Documents · IEEE Trans. Software Eng. 2001 Evaluating Capture-Recapture Models with Two Inspectors · IEEE Trans. Software Eng. 2001 A Comprehensive Evaluation of Capture-Recapture Models for Estimating Software Defect Content · IEEE Trans. Software Eng. 2000 |
Empirical software engineering
software effort estimation |
0.1 | 4 | 2001 | Software Cost Estimation with Incomplete Data · IEEE Trans. Software Eng. 2001 Explaining the Cost of European Space and Military Projects · ICSE 1999 An Assessment and Comparison of Common Software Cost Estimation Modeling Techniques · ICSE 1999 |
Empirical software engineering › mining software repositories
defect prediction |
0.1 | 1 | 2009 | An Investigation into the Functional Form of the Size-Defect Relationship for Software Modules · IEEE Trans. Software Eng. 2009 |
Software maintenance and evolution › software quality assurance
quality assurance prioritization |
0.1 | 1 | 2009 | An Investigation into the Functional Form of the Size-Defect Relationship for Software Modules · IEEE Trans. Software Eng. 2009 |
Empirical software engineering
software metrics |
0.1 | 2 | 2002 | The Optimal Class Size for Object-Oriented Software · IEEE Trans. Software Eng. 2002 The Confounding Effect of Class Size on the Validity of Object-Oriented Metrics · IEEE Trans. Software Eng. 2001 |
Requirements engineering and software design › object-oriented analysis and design
object-oriented design |
0.0 | 1 | 2002 | The Optimal Class Size for Object-Oriented Software · IEEE Trans. Software Eng. 2002 |
Empirical software engineering › experimental methodology
class size confounding |
0.0 | 1 | 2001 | The Confounding Effect of Class Size on the Validity of Object-Oriented Metrics · IEEE Trans. Software Eng. 2001 |
Software maintenance and evolution › software quality assurance › defect analysis
defect estimation |
0.0 | 1 | 2001 | Evaluating Capture-Recapture Models with Two Inspectors · IEEE Trans. Software Eng. 2001 |
Software maintenance and evolution › software inspection
reading techniques |
0.0 | 1 | 2001 | An Internally Replicated Quasi-Experimental Comparison of Checklist and Perspective-Based Reading of Code Documents · IEEE Trans. Software Eng. 2001 |
Compilers and program optimization
dead code elimination |
0.0 | 1 | 2000 | A Comprehensive Evaluation of Capture-Recapture Models for Estimating Software Defect Content · IEEE Trans. Software Eng. 2000 |
Requirements engineering and software design › software process
software process assessment |
0.0 | 1 | 2000 | Validating the ISO/IEC 15504 Measure of Software Requirements Analysis Process Capability · IEEE Trans. Software Eng. 2000 |
Requirements engineering and software design › risk management
risk assessment |
0.0 | 1 | 1998 | COBRA: A Hybrid Method for Software Cost Estimation, Benchmarking, and Risk Assessment · ICSE 1998 |
Requirements engineering and software design › software process
software process simulation |
0.0 | 1 | 1998 | Using Simulation to Build Inspection Efficiency Benchmarks for Development Projects · ICSE 1998 |
Empirical software engineering
fault prediction |
0.0 | 2 | 2002 | The Optimal Class Size for Object-Oriented Software · IEEE Trans. Software Eng. 2002 The Confounding Effect of Class Size on the Validity of Object-Oriented Metrics · IEEE Trans. Software Eng. 2001 |
Software maintenance and evolution › software evolution
corrective maintenance |
0.0 | 1 | 1997 | Characterizing and Modeling the Cost of Rework in a Library of Reusable Software Components · ICSE 1997 |
Software maintenance and evolution › software reuse
reusable software components |
0.0 | 1 | 1997 | Characterizing and Modeling the Cost of Rework in a Library of Reusable Software Components · ICSE 1997 |
Software maintenance and evolution
software reuse |
0.0 | 1 | 1997 | Characterizing and Modeling the Cost of Rework in a Library of Reusable Software Components · ICSE 1997 |
Empirical software engineering
benchmarking |
0.0 | 2 | 1999 | An Assessment and Comparison of Common Software Cost Estimation Modeling Techniques · ICSE 1999 COBRA: A Hybrid Method for Software Cost Estimation, Benchmarking, and Risk Assessment · ICSE 1998 |
Empirical software engineering
evidence-based software engineering |
0.0 | 1 | 2002 | Preliminary Guidelines for Empirical Research in Software Engineering · IEEE Trans. Software Eng. 2002 |
Requirements engineering and software design
software process |
0.0 | 1 | 1993 | A Comprehensive Process Model for Studying Software Process Papers · ICSE 1993 |
Requirements engineering and software design › software process
software process modeling |
0.0 | 1 | 1993 | A Comprehensive Process Model for Studying Software Process Papers · ICSE 1993 |
Software maintenance and evolution › code review
code inspection |
0.0 | 1 | 2001 | An Internally Replicated Quasi-Experimental Comparison of Checklist and Perspective-Based Reading of Code Documents · IEEE Trans. Software Eng. 2001 |
Software testing
fault detection |
0.0 | 1 | 2001 | An Internally Replicated Quasi-Experimental Comparison of Checklist and Perspective-Based Reading of Code Documents · IEEE Trans. Software Eng. 2001 |
Programming languages and type systems
functional programming |
0.0 | 1 | 2001 | The Confounding Effect of Class Size on the Validity of Object-Oriented Metrics · IEEE Trans. Software Eng. 2001 |
Empirical software engineering › software evaluation
inspection effectiveness |
0.0 | 1 | 2001 | Evaluating Capture-Recapture Models with Two Inspectors · IEEE Trans. Software Eng. 2001 |
Software maintenance and evolution
software quality assurance |
0.0 | 1 | 1998 | Using Simulation to Build Inspection Efficiency Benchmarks for Development Projects · ICSE 1998 |
Empirical software engineering
systematic literature review |
0.0 | 1 | 1993 | A Comprehensive Process Model for Studying Software Process Papers · ICSE 1993 |
Methods — techniques the papers use, named apart from their topics
regression analysis · 0.1class-level defect data analysis · 0.1capture-recapture estimation · 0.1simulation · 0.1systematic guideline development · 0.0empirical study · 0.0defect analysis · 0.0statistical confounding analysis · 0.0monte carlo simulation · 0.0empirical methodology · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Should we synthesize more than we need: impact of synthetic data generation for high-dimensional cross-sectional medical dataabstractOBJECTIVE: In medical research and education, generative artificial intelligence/machine learning (AI/ML) models to synthesize artificial medical data can enable the sharing of high-quality data while preserving the privacy of patients. Given that such data is often high-dimensional, a relevant consideration is whether to synthesize the entire dataset when only a task-relevant subset is needed. This study evaluates how the number of variables in training impacts fidelity, utility, and privacy of the synthetic data (SD). MATERIAL AND METHODS: We used 12 cross-sectional medical datasets, defined a downstream task with corresponding core variables, and derived 6354 variants by adding adjunct variables to the core. SD was generated using 7 different generative models and evaluated for fidelity, downstream utility, and privacy. Mixed-effect models were used to assess the effect of adjunct variables on the respective evaluation metric, accounting for the medical dataset as a random component. RESULTS: Fidelity was unaffected by the number of adjunct variables in 5/7 SDG models. Similarly, downstream utility remained stable in 6/7 (predictive task) and 5/7 (inferential task) SDG models. Where significant effects were observed, they were minimal, resulting, for example, in a 0.05 decrease in Area under the Receiver Operating Characteristic curve (AUROC) when adding 120 variables. Privacy was not impacted by the number of adjunct variables. DISCUSSION: Our findings show that fidelity, utility, and privacy are preserved when generating a more comprehensive medical dataset than the task-relevant subset. CONCLUSION: Our findings support a cost-effective, utility, and privacy-preserving way of implementing SDG into medical research and education. Lisa Pilgram, Samer El Kababji, Khaled El Emam |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Optimizing the synthesis of clinical trial data using sequential treesabstractOBJECTIVE: With the growing demand for sharing clinical trial data, scalable methods to enable privacy protective access to high-utility data are needed. Data synthesis is one such method. Sequential trees are commonly used to synthesize health data. It is hypothesized that the utility of the generated data is dependent on the variable order. No assessments of the impact of variable order on synthesized clinical trial data have been performed thus far. Through simulation, we aim to evaluate the variability in the utility of synthetic clinical trial data as variable order is randomly shuffled and implement an optimization algorithm to find a good order if variability is too high. MATERIALS AND METHODS: Six oncology clinical trial datasets were evaluated in a simulation. Three utility metrics were computed comparing real and synthetic data: univariate similarity, similarity in multivariate prediction accuracy, and a distinguishability metric. Particle swarm was implemented to optimize variable order, and was compared with a curriculum learning approach to ordering variables. RESULTS: As the number of variables in a clinical trial dataset increases, there is a pattern of a marked increase in variability of data utility with order. Particle swarm with a distinguishability hinge loss ensured adequate utility across all 6 datasets. The hinge threshold was selected to avoid overfitting which can create a privacy problem. This was superior to curriculum learning in terms of utility. CONCLUSIONS: The optimization approach presented in this study gives a reliable way to synthesize high-utility clinical trial datasets. Khaled El Emam, Lucy Mosquera, Chaoyi Zheng |
J. Am. Medical Informatics Assoc. | 1 |
| 2016 | Medical Data Privacy Handbook, A. Gkoulalas-Divanis, G. Loukides. Springer International Publishing, Switzerland (2015) 832 pp., ISBN: 978-3-319-23633-9
Khaled El Emam |
J. Biomed. Informatics | 1 |
| 2016 | A unified framework for evaluating the risk of re-identification of text de-identification toolsabstractOBJECTIVES: It has become regular practice to de-identify unstructured medical text for use in research using automatic methods, the goal of which is to remove patient identifying information to minimize re-identification risk. The metrics commonly used to determine if these systems are performing well do not accurately reflect the risk of a patient being re-identified. We therefore developed a framework for measuring the risk of re-identification associated with textual data releases. METHODS: We apply the proposed evaluation framework to a data set from the University of Michigan Medical School. Our risk assessment results are then compared with those that would be obtained using a typical contemporary micro-average evaluation of recall in order to illustrate the difference between the proposed evaluation framework and the current baseline method. RESULTS: We demonstrate how this framework compares against common measures of the re-identification risk associated with an automated text de-identification process. For the probability of re-identification using our evaluation framework we obtained a mean value for direct identifiers of 0.0074 and a mean value for quasi-identifiers of 0.0022. The 95% confidence interval for these estimates were below the relevant thresholds. The threshold for direct identifier risk was based on previously used approaches in the literature. The threshold for quasi-identifiers was determined based on the context of the data release following commonly used de-identification criteria for structured data. DISCUSSION: Our framework attempts to correct for poorly distributed evaluation corpora, accounts for the data release context, and avoids the often optimistic assumptions that are made using the more traditional evaluation approach. It therefore provides a more realistic estimate of the true probability of re-identification. CONCLUSIONS: This framework should be used as a basis for computing re-identification risk in order to more realistically evaluate future text de-identification tools. Martin Scaiano, Grant Middleton, Luk Arbuckle, Varada Kolhatkar, Liam Peyton, Moira Dowling, Debbie S. Gipson, Khaled El Emam |
J. Biomed. Informatics | 8 |
| 2015 | A privacy preserving protocol for tracking participants in phase I clinical trialsabstractOBJECTIVE: Some phase 1 clinical trials offer strong financial incentives for healthy individuals to participate in their studies. There is evidence that some individuals enroll in multiple trials concurrently. This creates safety risks and introduces data quality problems into the trials. Our objective was to construct a privacy preserving protocol to track phase 1 participants to detect concurrent enrollment. DESIGN: A protocol using secure probabilistic querying against a database of trial participants that allows for screening during telephone interviews and on-site enrollment was developed. The match variables consisted of demographic information. MEASUREMENT: The accuracy (sensitivity, precision, and negative predictive value) of the matching and its computational performance in seconds were measured under simulated environments. Accuracy was also compared to non-secure matching methods. RESULTS: The protocol performance scales linearly with the database size. At the largest database size of 20,000 participants, a query takes under 20s on a 64 cores machine. Sensitivity, precision, and negative predictive value of the queries were consistently at or above 0.9, and were very similar to non-secure versions of the protocol. CONCLUSION: The protocol provides a reasonable solution to the concurrent enrollment problems in phase 1 clinical trials, and is able to ensure that personal information about participants is kept secure. Khaled El Emam, Hanna Farah, Saeed Samet, Aleksander Essex, Elizabeth Jonker, Murat Kantarcioglu, Craig Earle |
J. Biomed. Informatics | 1 |
| 2014 | Real-World Data Set Parameters and Synthesization for Matching Identity in Clinical ProtocolsabstractA main challenge for clinical protocol evaluations is the lack of public real-world data sets due to the private nature of patient information. We studied the case of phase 1 clinical trials where the identity of participants is key in determining their eligibility to participate in a trial. Our objective is to use the experience from our study to present a list of parameters to help generate data sets that closely match their real-world counterparts. We also examine existing tools and address their limitations with a tool of our own. Through the development of our clinical trial protocol, we discovered a field selection that proved to be efficient to detect a participant's identity, which may be used by other researchers in their protocols. Hanna Farah, Daniel Amyot, Khaled El Emam |
CBMS | 3 |
| 2014 | A Tool for Simple and Efficient Clinical Protocol EvaluationabstractResearchers need to evaluate newly thought clinical protocols for validity and performance. A main challenge for such evaluations is the lack of real-world data sets due to the private nature of patient information. Researchers resort to synthetic data in order to simulate the real-world counterpart. We present a tool for researchers to increase their confidence in the evaluations they perform by allowing them to run their evaluations on a large number of synthetic data sets efficiently. This tool is used for distributing the evaluation on multiple machines in parallel, hence reducing processing time by the number of machines used, and for monitoring and merging the results. It allows the comparison of multiple algorithms and includes filters based on name length and record field selection to evaluate the effectiveness of algorithms in that regard. It also addresses the firewall limitations of current laboratory infrastructures, especially when security is of key importance. Hanna Farah, Daniel Amyot, Khaled El Emam |
CBMS | 3 |
| 2013 | A secure distributed logistic regression protocol for the detection of rare adverse drug eventsabstractBACKGROUND: There is limited capacity to assess the comparative risks of medications after they enter the market. For rare adverse events, the pooling of data from multiple sources is necessary to have the power and sufficient population heterogeneity to detect differences in safety and effectiveness in genetic, ethnic and clinically defined subpopulations. However, combining datasets from different data custodians or jurisdictions to perform an analysis on the pooled data creates significant privacy concerns that would need to be addressed. Existing protocols for addressing these concerns can result in reduced analysis accuracy and can allow sensitive information to leak. OBJECTIVE: To develop a secure distributed multi-party computation protocol for logistic regression that provides strong privacy guarantees. METHODS: We developed a secure distributed logistic regression protocol using a single analysis center with multiple sites providing data. A theoretical security analysis demonstrates that the protocol is robust to plausible collusion attacks and does not allow the parties to gain new information from the data that are exchanged among them. The computational performance and accuracy of the protocol were evaluated on simulated datasets. RESULTS: The computational performance scales linearly as the dataset sizes increase. The addition of sites results in an exponential growth in computation time. However, for up to five sites, the time is still short and would not affect practical applications. The model parameters are the same as the results on pooled raw data analyzed in SAS, demonstrating high model accuracy. CONCLUSION: The proposed protocol and prototype system would allow the development of logistic regression models in a secure manner without requiring the sharing of personal health information. This can alleviate one of the key barriers to the establishment of large-scale post-marketing surveillance programs. We extended the secure protocol to account for correlations among patients within sites through generalized estimating equations, and to accommodate other link functions by extending it to generalized linear models. Khaled El Emam, Saeed Samet, Luk Arbuckle, Robyn Tamblyn, Craig Earle, Murat Kantarcioglu |
J. Am. Medical Informatics Assoc. | 1 |
| 2013 | Biomedical data privacy: problems, perspectives, and recent advancesabstractThe notion of privacy in the healthcare domain is at least as old as the ancient Greeks. Several decades ago, as electronic medical record (EMR) systems began to take hold, the necessity of patient privacy was recognized as a core principle, or even a right, that must be upheld.1,2 This belief was re-enforced as computers and EMRs became more common in clinical environments.3–5 However, the arrival of ultra-cheap data collection and processing technologies is fundamentally changing the face of healthcare. The traditional boundaries of primary and tertiary care environments are breaking down and health information is increasingly collected through mobile devices,6 in personal domains (eg, in one's home7), and from sensors attached on or in the human body (eg, body area networks8–10). At the same time, the detail and diversity of information collected in the context of healthcare and biomedical research is increasing at an unprecedented rate, with clinical and administrative health data being complemented with a range of *omics data, where genomics11 and proteomics12 are currently leading the charge, with other types of molecular data on the horizon.13 Healthcare organizations (HCOs) are adopting and adapting information technologies to support an expanding array of activities designed to derive value from these growing data archives, in terms of enhanced health outcomes.14 The ready availability of such large volumes of detailed data has also been accompanied by privacy invasions. Recent breach notification laws at the US federal and state levels have brought to the public's attention the scope and frequency of these invasions. For example, there are cases of healthcare provider snooping on the medical records of famous people, family, and friends, use of personal information for identity fraud, and millions of records disclosed through lost and stolen unencrypted mobile devices.15 The danger is that such publicized incidents will erode patient trust over time, and lead to privacy protective behaviors. For example, between 15% and 17% of US adults have changed their behavior to protect the privacy of their health information, doing things such as: going to another doctor, paying out-of-pocket when insured to avoid disclosure, not seeking care to avoid disclosure to an employer, giving inaccurate or incomplete information on medical history, self-treating or self-medicating rather than seeing a provider, or asking a doctor not to write down the health problem or record a less serious or embarrassing condition.16–18 A survey of service members who had been on active duty found that respondents were concerned that if they received treatment for their mental health problems, it would not be kept confidential and would have a negative impact on future job assignments and career advancement.19 Specific vulnerable populations have reported similar privacy protective behaviors, such as adolescents, people with HIV or at high risk for HIV, women undergoing genetic testing, mental health patients, and victims of domestic violence.20–26 A survey of Californian residents found that discussing depression with their primary care physician was a barrier to 15% of the respondents because of privacy concerns.27 On the other hand, some legal scholars are questioning the survival of conventional privacy expectations.28 Privacy has conventionally been defined as an individual's ability to control the disclosure of personal facts.29,30 However, privacy is also a multi-dimensional concept31,32 and any shifts in privacy expectations are not homogeneous in direction and intensity across all of these dimensions. Furthermore, advances in informatics that may be eroding individuals' control over their information are being countered by advances in privacy enhancing technologies, as well as regulatory and policy changes that give individuals back control over their information. This special issue was established to solicit current research in privacy as it is currently understood and is being redefined for emerging biomedical systems. The selected articles consider the different dimensions of privacy, and describe some novel privacy enhancing technologies and their applications, as well as the governance, regulatory, and policy mechanisms that are being used to manage privacy risks. Privacy is a major patient, provider, regulator, and legislator concern today. There is therefore a need to address these concerns in a practical way that can be deployed in the short term. Deployment must be preceded by a convincing evidence base demonstrating the rationale, costs, and benefits of an intervention. At the same time, new theoretical models and novel approaches that still need to be evaluated and tested in the field, are also necessary to ensure that the field keeps evolving. In putting together this special issue we attempted to balance these two perspectives, with articles presenting results of immediate relevance and applicability, and material covering theoretical work that remains to be proven in practical settings. There were 53 papers submitted for consideration in this issue, of which 13 were accepted for publication, for an acceptance rate of 25%. All papers were subject to a rigorous review by at least two referees and oversight by one of the guest editors. The review process for papers authored by guest editors, as well as the editor-in-chief, was handled by an unaffiliated associate editor of the journal. In addition to peer-reviewed manuscripts, two invited papers were solicited for the special issue to address the topics of privacy policy and technical data protection mechanisms. Privacy is an overloaded and complex term.33 The concept of privacy often subsumes various constructs, such as anonymity (ie, the ability to hide one's identity), confidentiality (ie, the ability to share information with a second party without the information being publicly revealed), and solitude (ie, the right to be left alone). Even when the particular construct is unambiguous, it remains difficult to have discussions around the topic of privacy because it is highly contextual, such that the expectations of privacy are often specialized to the situation.34 For instance, a patient's expectation of privacy changes when disclosing information to a care provider versus a random person on the street. The expectation is further modified by the perceived sensitivity of the health information in question. And, the extent to which health information (eg, a positive assertion of an HIV diagnosis) is deemed to be sensitive varies from patient to patient. This special issue is organized to trace the lifecycle of biomedical information, which we coarsely partition into three zones that follow the general data lifecycle as follows. The first zone corresponds to the point at which health information is collected from patients. The collection may occur while an individual is physically located at a healthcare provider or beyond (eg, such as through a website on the internet or an application running on a mobile device). In this zone, privacy tends to be concerned with who can collect health information, how much information should be collected, at what time, and for what purposes. The specific notions of privacy addressed in this zone tend to be associated with anonymity (eg, Is the recipient of the data permitted to know the identity of the information from which it is being collected?), limiting content (eg, What is the minimal amount of information that will satisfy the purpose?), and consent (eg, Did the patient agree to the terms of the data collection?). The second zone corresponds to the context in which the data have left the control of the patient and are housed in a system controlled, or accessed by, those who provide a primary service (eg, provision of care, study of biomedical data explicitly solicited for a specific research project). In this zone, privacy tends to be realized through confidentiality (eg, Who is permitted to access or use the data and for what purposes?) and security (eg, How can we ensure that the data are protected from misuse or abuse while at rest in a database or in transit between authorized entities?). The third zone corresponds to the scenario in which biomedical data are utilized for purposes which are different from their primary use. The data may be used by the organization that initially collected the data (eg, repurposing of clinical data for research) or disseminated to external entities (eg, publication of public use datasets) for the performance of certain tasks (eg, evaluation of health policies). In this zone, the privacy issues that tend to arise are anonymity and consent for individuals and groups (eg, Can data collected from a particular ethnic group be reused to study a specific phenotype?). Each of these zones can be partitioned into two interacting, although conceptually distinct, tracks. In the first track, biomedical privacy is defined and regulated via socio-legal mechanisms. This is the arena where the public, ethicists, and policy and law makers come together to define what privacy rights and responsibilities exist. In the second track, technical controls are specified and realized in working information technologies to maintain societal expectations of privacy or requirements specified in policy and law. It is critical to integrate these tracks to ensure that privacy expectations are appropriately represented in technical controls and that policies are designed to realistically account for state-of-the-art technical capabilities. It is further important that societal expectations of privacy and privacy enhancing technologies are current and cognizant of shifts in expectation of technical sophistication. The socio-legal track of this special issue commences by delving into the desires and expectations society harbors for privacy. As mentioned earlier, privacy is a societal phenomenon, such that the extent to which it is realized is dependent on how society chooses to codify the concept in policy and law. This process often begins with field studies and sessions that engage stakeholders in their preferences.35 In this vein, Caine and Hanania report on a study that asks patients who should control access to health information and what granularity of control is desirable.36 Often, the expectations of privacy are dependent on the domain in which information is collected. As the traditional boundaries of the healthcare domain expand, it is important to determine how individuals' perspectives on privacy relate to new technologies. To begin to address this issue, van der Velden and El Emam focus on the use of social media by teenage patients, and how they perceive their health information privacy when interacting online.37 Insights gained here should inform the more general health data context. The next set of papers in the socio-legal track move beyond primary uses for biomedical data and into secondary settings. In this environment, it is assumed that policies and laws have been codified. However, policy and law is dependent on the locale, such that it is critical to understand how it guides data management practices. The first paper in this group, by Pencarrick Hertzman, Meagher, and McGrail, presents a case study about how the ‘Privacy by Design’ framework was applied in British Columbia (BC), Canada, to facilitate access to health information in Population Data BC.38 To date, Population Data BC has facilitated over 350 research studies. This work is followed by the first invited paper by McGraw, which takes a look at the de-identification strategy of the US Health Insurance Portability and Accountability Act (HIPAA).39 This strategy enables HCOs to disclose information about patients in a manner that is no longer subject to oversight by the regulatory authorities because the risk that they would be individually identifiable is deemed very small. Clarifications to what de-identification means and how it can be achieved in accordance with HIPAA were recently published by the U.S. federal government.40 The paper by McGraw reports on a workshop on various stakeholders' support for the current HIPAA de-identification strategy held by the Center for Democracy and Technology and discusses policy proposals to address concerns and improve trust in the process. Peterson and colleagues then recount a recent case before the Supreme Court, Sorrell vs IMS Health, and illustrate the challenges associated with selling prescription records for various purposes, such as post-market effectiveness.41 They highlight some of the concerns associated with the dissemination of identifiable prescriber and de-identified patient information. While the previous papers focus on traditional health information, the last paper in this track, by Kosseim and colleagues, addresses privacy issues associated with the management of *omics data in particular, with a specific focus on genomics.42 This paper describes the legal and ethical principles and practices adopted in the Canadian provinces of Newfoundland and Labrador to enable research with genomic, phenomic, and genealogical data. While law and policy codify the rights and requirements for managing data privacy in the biomedical domain, information technology is necessary to uphold and ensure their realization in practice. In this regard, certain aspects of privacy can be achieved through information security. The Security Rule of HIPAA specifies various administrative, physical, and technical safeguards that covered entities must have in place (ie, required controls) or document why such protections are not prudent (ie, addressable controls). For instance, it is required that all covered entities ensure that appropriate authorization is provided before employees of an HCO access a patient's EMR. By contrast, the encryption of health information at rest within the HCO is addressable, but is not required. Along these lines, the technical track of this special issue begins with a paper by Kwon and Johnson that assesses the extent to which 250 HCOs in the USA have (or have not) adopted various security practices.43 Their analysis demonstrates patterns of leaders, followers, and laggers in their adoption, and provides recommendations for improving regulatory compliance. Although this work provides a high-level assessment of the adoption of security practices, it does not provide specific guidance on data management strategies. Thus, the next paper, by Fabbri and LeFevre, discusses a privacy threat encountered on a daily basis in primary care settings, specifically the insider threat.44 This threat is particularly important to study in the healthcare domain because traditional information security controls (eg, role-based access control) are difficult to realize in care settings due to the highly dynamic nature of healthcare teams. An increasing number of publications have proposed auditing strategies for EMRs,45,46 however, this line of work is unique in that it suggests EMR users can be ‘explained’ by the diagnoses that are assigned to the patient records they access. This work suggests that data-driven auditing strategies may help winnow the set of accesses to patient records to a manageable size for review by administrative officials (eg, privacy officers) of HCOs. The next set of papers in the technical section focus on various strategies that can be invoked to protect patients' privacy when data are shared for secondary use. However, before presenting specific protection methodologies, this section begins with an illustration of types of research studies that can be enabled through de-identified data. The paper by White and Horvitz integrates web search data from Bing and geocoded data from mobile devices to show how search for health information online correlates with an individual's physical presence at a healthcare providing facility.47 This research is performed on data that are stripped of user identifiers and location information prior to the analysis. However, there may be times when it is beneficial to link a patient's record across multiple healthcare institutions, or within a single institution. To support such efforts without revealing a patient's identity, there has been a flurry of research in private record linkage48–51 (or entity resolution). Such linkage is increasingly based on hashed versions of patient identifiers (eg, personal names) or quasi-identifiers (eg, demographics). The paper from Cassa, Miller, and Mandl suggests a protocol to derive a secure fingerprint from genomic data.52 They indicate how this approach may be applied to track a patient's record across the research enterprise in place of explicitly identifying information. The final set of papers focus on strategies for de-identifying various types of health information. Biomedical data can take a wide array of forms, ranging from free text (eg, natural language clinical notes) to structured information (eg, such as discharge databases) to high-dimensional information (eg, genome-wide scans of single nucleotide polymorphisms). The majority of health information is in free text form and so a significant amount of research53,54 over the past several years has investigated how to detect and redact a prespecified set of potential identifiers (such as the list of 18 features in the HIPAA Safe Harbor de-identification standard). The paper by Ferrandez and colleagues provides an example of how rule-based (eg, dictionaries, regular expressions, and rules) and machine (eg, random support and can be to construct a free text for over different types of Health clinical The paper by and colleagues how machine text de-identification has impact on clinical concept in the form of from over types from the Although the previous and in the illustrate how identifiers in clinical text can be information in the text may still or of the patient (or their As have to de-identification strategies that are more in their strategies are often applied to more data such as database The paper by and colleagues how a patient's of results can be unique and used as a to track a patient back to their To this they a which to the An analysis patients' records that such it highly that a patient's record be to a group of less than individuals while minimal on clinical Although some this of addition does not an In this regard, a significant number of publications have (ie, strategies be applied to ensure that record corresponds to at least patients (ie, the it has been that de-identification strategies based on the of a prespecified list of features or more strategies may not be an appropriate of privacy protection because they can about the patients from the data were While this general notion has been models have been In the second invited paper, and describe the notion of privacy from a theoretical In this of are permitted to of a which with a (eg, a of may be reported as This is such that it is that the determine a specific individual to the database within a certain and that the is within a certain of the This has a number of important but also a number of and practical to in healthcare The final paper of the special issue, by and colleagues, provides a more practical on how privacy be applied to of health They how this privacy protection be applied to a from but there are challenges to this approach to high-dimensional data. As this special issue the of data privacy in the biomedical domain is and It and technical boundaries and is specialized to the of data and process being it is not to review the field in this As we this to a we that topics (eg, access consent disclosure and policy to manage health information were not but are no less important than those reported on in this At the same time, we that new and technologies are new challenges to privacy that the biomedical will need to in the not technology that we to highlight is As and the amount of data by healthcare it is increasingly the case the health information is being in systems beyond the control and oversight of and in with different privacy laws and this issue demonstrates that appropriate protections can be defined for emerging systems and that there is a working on and are that research in this area will lead to new appropriate that balance privacy and data and system Bradley A. Malin, Khaled El Emam, Christine M. O'Keefe |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | "Not all my friends need to know": a qualitative study of teenage patients, privacy, and social mediaabstractBACKGROUND: The literature describes teenagers as active users of social media, who seem to care about privacy, but who also reveal a considerable amount of personal information. There have been no studies of how they manage personal health information on social media. OBJECTIVE: To understand how chronically ill teenage patients manage their privacy on social media sites. DESIGN: A qualitative study based on a content analysis of semistructured interviews with 20 hospital patients (12-18 years). RESULTS: Most teenage patients do not disclose their personal health information on social media, even though the study found a pervasive use of Facebook. Facebook is a place to be a "regular", rather than a sick teenager. It is a place where teenage patients stay up-to-date about their social life-it is not seen as a place to discuss their diagnosis and treatment. The majority of teenage patients don't use social media to come into contact with others with similar conditions and they don't use the internet to find health information about their diagnosis. CONCLUSIONS: Social media play an important role in the social life of teenage patients. They enable young patients to be "regular" teenagers. Teenage patients' online privacy behavior is an expression of their need for self-definition and self-protection. Maja van der Velden, Khaled El Emam |
J. Am. Medical Informatics Assoc. | 2 |
| 2011 | A secure protocol for protecting the identity of providers when disclosing data for disease surveillanceabstractBACKGROUND: Providers have been reluctant to disclose patient data for public-health purposes. Even if patient privacy is ensured, the desire to protect provider confidentiality has been an important driver of this reluctance. METHODS: Six requirements for a surveillance protocol were defined that satisfy the confidentiality needs of providers and ensure utility to public health. The authors developed a secure multi-party computation protocol using the Paillier cryptosystem to allow the disclosure of stratified case counts and denominators to meet these requirements. The authors evaluated the protocol in a simulated environment on its computation performance and ability to detect disease outbreak clusters. RESULTS: Theoretical and empirical assessments demonstrate that all requirements are met by the protocol. A system implementing the protocol scales linearly in terms of computation time as the number of providers is increased. The absolute time to perform the computations was 12.5 s for data from 3000 practices. This is acceptable performance, given that the reporting would normally be done at 24 h intervals. The accuracy of detection disease outbreak cluster was unchanged compared with a non-secure distributed surveillance protocol, with an F-score higher than 0.92 for outbreaks involving 500 or more cases. CONCLUSION: The protocol and associated software provide a practical method for providers to disclose patient data for sentinel, syndromic or other indicator-based surveillance while protecting patient privacy and the identity of individual providers. Khaled El Emam, Jay Mercer, Liam Peyton, Murat Kantarcioglu, Bradley A. Malin, David L. Buckeridge, Saeed Samet, Craig Earle |
J. Am. Medical Informatics Assoc. | 1 |
| 2010 | Testing the theory of relative defect proneness for closed-source software
Akif Günes Koru, Dongsong Zhang, Khaled El Emam |
Empir. Softw. Eng. | 4 |
| 2010 | The inadvertent disclosure of personal health information through peer-to-peer file sharing programsabstractOBJECTIVE: There has been a consistent concern about the inadvertent disclosure of personal information through peer-to-peer file sharing applications, such as Limewire and Morpheus. Examples of personal health and financial information being exposed have been published. We wanted to estimate the extent to which personal health information (PHI) is being disclosed in this way, and compare that to the extent of disclosure of personal financial information (PFI). DESIGN: After careful review and approval of our protocol by our institutional research ethics board, files were downloaded from peer-to-peer file sharing networks and manually analyzed for the presence of PHI and PFI. The geographic region of the IP addresses was determined, and classified as either USA or Canada. MEASUREMENT: We estimated the proportion of files that contain personal health and financial information for each region. We also estimated the proportion of search terms that return files with personal health and financial information. We ascertained and discuss the ethical issues related to this study. RESULTS: Approximately 0.4% of Canadian IP addresses had PHI, as did 0.5% of US IP addresses. There was more disclosure of financial information, at 1.7% of Canadian IP addresses and 4.7% of US IP addresses. An analysis of search terms used in these file sharing networks showed that a small percentage of the terms would return PHI and PFI files (ie, there are people successfully searching for PFI and PHI on the peer-to-peer file sharing networks). CONCLUSION: There is a real risk of inadvertent disclosure of PHI through peer-to-peer file sharing networks, although the risk is not as large as for PFI. Anyone keeping PHI on their computers should avoid installing file sharing applications on their computers, or if they have to use such tools, actively manage the risks of inadvertent disclosure of their, their family's, their clients', or patients' PHI. Khaled El Emam, Emilio Neri, Elizabeth Jonker, Marina Sokolova, Liam Peyton, Angelica Neisa, Teresa Scassa |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Model Formulation: Evaluating Predictors of Geographic Area Population Size Cut-offs to Manage Re-identification RiskabstractOBJECTIVE: In public health and health services research, the inclusion of geographic information in data sets is critical. Because of concerns over the re-identification of patients, data from small geographic areas are either suppressed or the geographic areas are aggregated into larger ones. Our objective is to estimate the population size cut-off at which a geographic area is sufficiently large so that no data suppression or further aggregation is necessary. DESIGN: The 2001 Canadian census data were used to conduct a simulation to model the relationship between geographic area population size and uniqueness for some common demographic variables. Cut-offs were computed for geographic area population size, and prediction models were developed to estimate the appropriate cut-offs. MEASUREMENTS: Re-identification risk was measured using uniqueness. Geographic area population size cut-offs were estimated using the maximum number of possible values in the data set and a traditional entropy measure. RESULTS: The model that predicted population cut-offs using the maximum number of possible values in the data set had R2 values around 0.9, and relative error of prediction less than 0.02 across all regions of Canada. The models were then applied to assess the appropriate geographic area size for the prescription records provided by retail and hospital pharmacies to commercial research and analysis firms. CONCLUSIONS: To manage re-identification risk, the prediction models can be used by public health professionals, health researchers, and research ethics boards to decide when the geographic area population size is sufficiently large. Khaled El Emam, Ann Brown, Philip AbdelMalik |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Research Paper: A Globally Optimal k-Anonymity Method for the De-Identification of Health DataabstractBACKGROUND: Explicit patient consent requirements in privacy laws can have a negative impact on health research, leading to selection bias and reduced recruitment. Often legislative requirements to obtain consent are waived if the information collected or disclosed is de-identified. OBJECTIVE: The authors developed and empirically evaluated a new globally optimal de-identification algorithm that satisfies the k-anonymity criterion and that is suitable for health datasets. DESIGN: Authors compared OLA (Optimal Lattice Anonymization) empirically to three existing k-anonymity algorithms, Datafly, Samarati, and Incognito, on six public, hospital, and registry datasets for different values of k and suppression limits. Measurement Three information loss metrics were used for the comparison: precision, discernability metric, and non-uniform entropy. Each algorithm's performance speed was also evaluated. RESULTS: The Datafly and Samarati algorithms had higher information loss than OLA and Incognito; OLA was consistently faster than Incognito in finding the globally optimal de-identification solution. CONCLUSIONS: For the de-identification of health datasets, OLA is an improvement on existing k-anonymity algorithms in terms of information loss and performance. Khaled El Emam, Fida Dankar, Romeo Issa, Elizabeth Jonker, Daniel Amyot, Elise Cogo, Jean-Pierre Corriveau, Mark Walker, Sadrul Chowdhury, Regis Vaillancourt, Tyson Roffey, Jim Bottomley |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | An Investigation into the Functional Form of the Size-Defect Relationship for Software ModulesabstractThe importance of the relationship between size and defect proneness of software modules is well recognized. Understanding the nature of that relationship can facilitate various development decisions related to prioritization of quality assurance activities. Overall, the previous research only drew a general conclusion that there was a monotonically increasing relationship between module size and defect proneness. In this study, we analyzed class-level size and defect data in order to increase our understanding of this crucial relationship. In order to obtain validated and more generalizable results, we studied four large-scale object-oriented products, Mozilla, Cn3d, JBoss, and Eclipse. Our results consistently revealed a significant effect of size on defect proneness; however, contrary to common intuition, the size-defect relationship took a logarithmic form, indicating that smaller classes were proportionally more problematic than larger classes. Therefore, practitioners should consider giving higher priority to smaller modules when planning focused quality assurance activities with limited resources. For example, in Mozilla and Eclipse, an inspection strategy investing 80% of available resources on 100-LOC classes and the rest on 1,000-LOC classes would be more than twice as cost effective as the opposite strategy. These results should be immediately useful to guide focused quality assurance activities in large-scale software projects. Akif Günes Koru, Dongsong Zhang, Khaled El Emam |
IEEE Trans. Software Eng. | 3 |
| 2008 | Theory of relative defect proneness
Akif Günes Koru, Khaled El Emam, Dongsong Zhang, Divya Mathew |
Empir. Softw. Eng. | 2 |
| 2008 | Research Paper: Protecting Privacy Using k-AnonymityabstractOBJECTIVE: There is increasing pressure to share health information and even make it publicly available. However, such disclosures of personal health information raise serious privacy concerns. To alleviate such concerns, it is possible to anonymize the data before disclosure. One popular anonymization approach is k-anonymity. There have been no evaluations of the actual re-identification probability of k-anonymized data sets. DESIGN: Through a simulation, we evaluated the re-identification risk of k-anonymization and three different improvements on three large data sets. MEASUREMENT: Re-identification probability is measured under two different re-identification scenarios. Information loss is measured by the commonly used discernability metric. RESULTS: For one of the re-identification scenarios, k-Anonymity consistently over-anonymizes data sets, with this over-anonymization being most pronounced with small sampling fractions. Over-anonymization results in excessive distortions to the data (i.e., high information loss), making the data less useful for subsequent analysis. We found that a hypothesis testing approach provided the best control over re-identification risk and reduces the extent of information loss compared to baseline k-anonymity. CONCLUSION: Guidelines are provided on when to use the hypothesis testing approach instead of baseline k-anonymity. Khaled El Emam, Fida Dankar |
J. Am. Medical Informatics Assoc. | 1 |
| 2007 | SPICE in retrospect: Developing a standard for process assessment
Terence P. Rout, Khaled El Emam, Mario Fusani, Dennis R. Goldenson, Ho-Won Jung |
J. Syst. Softw. | 2 |
| 2004 | Applications of statistics in software engineering
Khaled El Emam, Anita D. Carleton |
J. Syst. Softw. | 1 |
| 2002 | The Optimal Class Size for Object-Oriented SoftwareabstractA growing body of literature suggests that there is an optimal size for software components. This means that components that are too small or too big will have a higher defect content (i.e., there is a U-shaped curve relating defect content to size). The U-shaped curve has become known as the "Goldilocks Conjecture." Recently, a cognitive theory has been proposed to explain this phenomenon and it has been expanded to characterize object-oriented software. This conjecture has wide implications for software engineering practice. It suggests 1) that designers should deliberately strive to design classes that are of the optimal size, 2) that program decomposition is harmful, and 3) that there exists a maximum (threshold) class size that should not be exceeded to ensure fewer faults in the software. The purpose of the current paper is to evaluate this conjecture for object-oriented systems. We first demonstrate that the claims of an optimal component/class size (1) above) and of smaller components/classes having a greater defect content (2) above) are due to a mathematical artifact in the analyses performed previously. We then empirically test the threshold effect claims of this conjecture (3) above). To our knowledge, the empirical test of size threshold effects for object-oriented systems has not been performed thus far. We performed an initial study with an industrial C++ system and repeated it twice on another C++ system and on a commercial Java application. Our results provide unambiguous evidence that there is no threshold effect of class size. We obtained the same result for three systems using four different size measures. These findings suggest that there is a simple continuous relationship between class size and faults, and that, optimal class size, smaller classes are better and threshold effects conjectures have no sound theoretical nor empirical basis. Khaled El Emam, Saïda Benlarbi, Nishith Goel, Walcélio L. Melo, Hakim Lounis, Shesh N. Rai |
IEEE Trans. Software Eng. | 1 |
| 2002 | Preliminary Guidelines for Empirical Research in Software EngineeringabstractEmpirical software engineering research needs research guidelines to improve the research and reporting processes. We propose a preliminary set of research guidelines aimed at stimulating discussion among software researchers. They are based on a review of research guidelines developed for medical researchers and on our own experience in doing and reviewing software engineering research. The guidelines are intended to assist researchers, reviewers, and meta-analysts in designing, conducting, and evaluating empirical studies. Editorial boards of software engineering journals may wish to use our recommendations as a basis for developing guidelines for reviewers and for framing policies for dealing with the design, data collection, and analysis and reporting of empirical studies. Barbara A. Kitchenham, Shari Lawrence Pfleeger, Lesley Pickard, Peter Jones 0002, David C. Hoaglin, Khaled El Emam, Jarrett Rosenberg |
IEEE Trans. Software Eng. | 6 |
| 2001 | Ethics and Open Source
Khaled El Emam |
Empir. Softw. Eng. | 1 |
| 2001 | Modelling the Likelihood of Software Process Improvement: An Exploratory Study
Khaled El Emam, Dennis R. Goldenson, James McCurley, James D. Herbsleb |
Empir. Softw. Eng. | 1 |
| 2001 | Comparing case-based reasoning classifiers for predicting high risk software components
Khaled El Emam, Saïda Benlarbi, Nishith Goel, Shesh N. Rai |
J. Syst. Softw. | 1 |
| 2001 | An empirical evaluation of the ISO/IEC 15504 assessment model
Khaled El Emam, Ho-Won Jung |
J. Syst. Softw. | 1 |
| 2001 | The prediction of faulty classes using object-oriented design metrics
Khaled El Emam, Walcélio L. Melo, Javam C. Machado |
J. Syst. Softw. | 1 |
| 2001 | The Confounding Effect of Class Size on the Validity of Object-Oriented MetricsabstractMuch effort has been devoted to the development and empirical validation of object-oriented metrics. The empirical validations performed thus far would suggest that a core set of validated metrics is close to being identified. However, none of these studies allow for the potentially confounding effect of class size. We demonstrate a strong size confounding effect and question the results of previous object-oriented metrics validation studies. We first investigated whether there is a confounding effect of class size in validation studies of object-oriented metrics and show that, based on previous work, there is reason to believe that such an effect exists. We then describe a detailed empirical methodology for identifying those effects. Finally, we perform a study on a large C++ telecommunications framework to examine if size is really a confounder. This study considered the Chidamber and Kemerer metrics and a subset of the Lorenz and Kidd metrics. The dependent variable was the incidence of a fault attributable to a field failure (fault-proneness of a class). Our findings indicate that, before controlling for size, the results are very similar to previous studies. The metrics that are expected to be validated are indeed associated with fault-proneness. Khaled El Emam, Saïda Benlarbi, Nishith Goel, Shesh N. Rai |
IEEE Trans. Software Eng. | 1 |
| 2001 | Evaluating Capture-Recapture Models with Two InspectorsabstractCapture-recapture (CR) models have been proposed as an objective method for controlling software inspections. CR models were originally developed to estimate the size of animal populations. In software, they have been used to estimate the number of defects in an inspected artifact. This estimate can be another source of information for deciding whether the artifact requires a reinspection to ensure that a minimal inspection effectiveness level has been attained. Little evaluative research has been performed thus far on the utility of CR models for inspections with two inspectors. We report on an extensive Monte Carlo simulation that evaluated capture-recapture models suitable for two inspectors assuming a code inspections context. We evaluate the relative error of the CR estimates as well as the accuracy of the reinspection decision made using the CR model. Our results indicate that the most appropriate capture-recapture model for two inspectors is an estimator that allows for inspectors with different capabilities. This model always produces an estimate (i.e., does not fail), has a predictable behavior (i.e., works well when its assumptions are met), will have a relatively high decision accuracy, and will perform better than the default decision of no reinspections. Furthermore, we identify the conditions under which this estimator will perform best. Khaled El Emam, Oliver Laitenberger |
IEEE Trans. Software Eng. | 1 |
| 2001 | An Internally Replicated Quasi-Experimental Comparison of Checklist and Perspective-Based Reading of Code DocumentsabstractThe basic premise of software inspections is that they detect and remove defects before they propagate to subsequent development phases where their detection and correction cost escalates. To exploit their full potential, software inspections must call for a close and strict examination of the inspected artifact. For this, reading techniques for defect detection may be helpful since these techniques tell inspection participants what to look for and, more importantly, how to scrutinize a software artifact in a systematic manner. Recent research efforts investigated the benefits of scenario-based reading techniques. A major finding has been that these techniques help inspection teams find more defects than existing state-of-the-practice approaches, such as, ad-hoc or checklist-based reading (CBR). We experimentally compare one scenario-based reading technique, namely, perspective-based reading (PBR), for defect detection in code documents with the more traditional CBR approach. The comparison was performed in a series of three studies, as a quasi experiment and two internal replications, with a total of 60 professional software developers at Bosch Telecom GmbH. Meta-analytic techniques were applied to analyze the data. Oliver Laitenberger, Khaled El Emam, Thomas G. Harbich |
IEEE Trans. Software Eng. | 2 |
| 2001 | Software Cost Estimation with Incomplete DataabstractThe construction of software cost estimation models remains an active topic of research. The basic premise of cost modeling is that a historical database of software project cost data can be used to develop a quantitative model to predict the cost of future projects. One of the difficulties faced by workers in this area is that many of these historical databases contain substantial amounts of missing data. Thus far, the common practice has been to ignore observations with missing data. In principle, such a practice can lead to gross biases and may be detrimental to the accuracy of cost estimation models. We describe an extensive simulation where we evaluate different techniques for dealing with missing data in the context of software cost modeling. Three techniques are evaluated: listwise deletion, mean imputation, and eight different types of hot-deck imputation. Our results indicate that all the missing data techniques perform well with small biases and high precision. This suggests that the simplest technique, listwise deletion, is a reasonable choice. However, this will not necessarily provide the best performance. Consistent best performance (minimal bias and highest precision) can be obtained by using hot-deck imputation with Euclidean distance and a z-score standardization. Kevin Strike, Khaled El Emam, Nazim H. Madhavji |
IEEE Trans. Software Eng. | 2 |
| 2000 | Thresholds for Object-Oriented MeasuresabstractA practical application of object oriented measures is to predict which classes are likely to contain a fault. This is contended to be meaningful because object oriented measures are believed to be indicators of psychological complexity, and classes that are more complex are likely to be faulty. Recently, a cognitive theory was proposed suggesting that there are threshold effects for many object oriented measures. This means that object oriented classes are easy to understand as long as their complexity is below a threshold. Above that threshold their understandability decreases rapidly, leading to an increased probability of a fault. This occurs, according to the theory, due to an overflow of short-term human memory. If this theory is confirmed, then it would provide a mechanism that would explain the introduction of faults into object oriented systems, and would also provide some practical guidance on how to design object oriented programs. The authors empirically test this theory on two C++ telecommunications systems. They test for threshold effects in a subset of the Chidamber and Kemerer (CK) suite of measures (S. Chidamber and C. Kemerer, 1994). The dependent variable was the incidence of faults that lead to field failures. The results indicate that there are no threshold effects for any of the measures studied. This means that there is no value for the studied CK measures where the fault-proneness changes from being steady to rapidly increasing. The results are consistent across the two systems. Therefore, we can provide no support to the posited cognitive theory. Saïda Benlarbi, Khaled El Emam, Nishith Goel, Shesh N. Rai |
ISSRE | 2 |
| 2000 | Validating the ISO/IEC 15504 measures of software development process capability
Khaled El Emam, Andreas Birk 0001 |
J. Syst. Softw. | 1 |
| 2000 | Estimating the extent of standards use: the case of ISO/IEC 15504
Khaled El Emam, Iñigo Garro |
J. Syst. Softw. | 1 |
| 2000 | The application of subjective estimates of effectiveness to controlling software inspections
Khaled El Emam, Oliver Laitenberger, Thomas G. Harbich |
J. Syst. Softw. | 1 |
| 2000 | An experimental comparison of reading techniques for defect detection in UML design documents
Oliver Laitenberger, Colin Atkinson 0001, Maud Schlich, Khaled El Emam |
J. Syst. Softw. | 4 |
| 2000 | A Comprehensive Evaluation of Capture-Recapture Models for Estimating Software Defect ContentabstractAn important requirement to control the inspection of software artifacts is to be able to decide, based on more objective information, whether the inspection can stop or whether it should continue to achieve a suitable level of artifact quality. A prediction of the number of remaining defects in an inspected artifact can be used for decision making. Several studies in software engineering have considered capture-recapture models to make a prediction. However, few studies compare the actual number of remaining defects to the one predicted by a capture-recapture model on real software engineering artifacts. The authors focus on traditional inspections and estimate, based on actual inspections data, the degree of accuracy of relevant state-of-the-art capture-recapture models for which statistical estimators exist. In order to assess their robustness, we look at the impact of the number of inspectors and the number of actual defects on the estimators' accuracy based on actual inspection data. Our results show that models are strongly affected by the number of inspectors, and therefore one must consider this factor before using capture-recapture models. When the number of inspectors is too small, no model is sufficiently accurate and underestimation may be substantial. In addition, some models perform better than others in a large number of conditions and plausible reasons are discussed. Based on our analyses, we recommend using a model taking into account that defects have different probabilities of being detected and the corresponding Jackknife Estimator. Furthermore, we calibrate the prediction models based on their relative error, as previously computed on other inspections. We identified theoretical limitations to this approach which were then confirmed by the data. Lionel C. Briand, Khaled El Emam, Bernd G. Freimut, Oliver Laitenberger |
IEEE Trans. Software Eng. | 2 |
| 2000 | Validating the ISO/IEC 15504 Measure of Software Requirements Analysis Process CapabilityabstractISO/IEC 15504 is an emerging international standard on software process assessment. It defines a number of software engineering processes and a scale for measuring their capability. One of the defined processes is software requirements analysis (SRA). A basic premise of the measurement scale is that higher process capability is associated with better project performance (i.e., predictive validity). The paper describes an empirical study that evaluates the predictive validity of SRA process capability. Assessments using ISO/IEC 15504 were conducted on 56 projects world-wide over a period of two years. Performance measures on each project were also collected using questionnaires, such as the ability to meet budget commitments and staff productivity. The results provide strong evidence of predictive validity for the SRA process capability measure used in ISO/IEC 15504, but only for organizations with more than 50 IT staff. Specifically, a strong relationship was found between the implementation of requirements analysis practices as defined in ISO/IEC 15504 and the productivity of software projects. For smaller organizations, evidence of predictive validity was rather weak. This can be interpreted in a number of different ways: that the measure of capability is not suitable for small organizations or that the SRA process capability has less effect on project performance for small organizations. Khaled El Emam, Andreas Birk 0001 |
IEEE Trans. Software Eng. | 1 |
| 1999 | An Assessment and Comparison of Common Software Cost Estimation Modeling TechniquesabstractThis paper investigates two essential data-driven, software cost modeling: questions related to (1) What modeling _ _ techniques are likely to yield more accurate results when using typical software development cost data?and (2) What are the benefits and drawbacks of using organizationspecific data as compared to multi-organization databases?The former question is important in guiding software cost analysts in their choice of the right type of modeling technique, if at all possible.In order to address this issue, we assess and compare a selection of common cost modeling techniques fulfilling a number of important criteria using a large multi-organizational database in the business application domain.Namely, these are: ordinary least squares regression, stepwise ANOVA, CART, and analogy.The latter question is important in order to assess the feasibility of using multi-organization cost databases to build cost models and the benefits gained from local, company-specific data collection and modeling.As a large subset of the data in the multi-company database came from one organization, we were able to investigate this issue by comparing organization-specific models with models based on multi-organization data.Results show that the performances of the modeling techniques considered were not significantly different, with the exception of the analogy-based models which appear to be less accurate.Surprisingly, when using standard cost factors (e.g., COCOMO-like factors, Function Points), organization specific models did not yield better results than generic, multi-organization models.Permission to make digital or hard topics c~fall or part of this work fbl personal or classroom USC is granted without fee provided that topics are not made or distributed for profit or commercial advnntage and that copies hear this notice and the full citation cm the tkt page.To CojJy othcrwiso, to republish, to post on servers or to redistribute to lists.requires prior specific permission and/or a fee. Lionel C. Briand, Khaled El Emam, Dagmar Surmann, Isabella Wieczorek, Katrina Maxwell |
ICSE | 2 |
| 1999 | Explaining the Cost of European Space and Military ProjectsabstractThere has been much controversy in the literature on several issues underlying the construction of parametric software development cost models.For example, it has been argued whether (dis)economies of scale exist in software production, what functional form should be assumed between effort and product size, whether COCOMO factors were useful, and whether the COCOMO factors are independent.Answers to such questions should help software organizations define suitable data collection programs and well-specified cost models.The only way to address these issues and obtain a generalizable conclusion is to investigate them on a large number of consistent data sets.In this paper we use a data set collected by the European Space Agency to perform such an investigation.To ensure a certain degree of consistency in our data, we focus our analysis on a set of space and military projects that represent an important application domain and the largest subset in the database.These projects have been performed, however, by a variety of organizations.First, our results indicate that two functional forms are plausible between effort and product size: linear and log-linear.This also means that different project subpopulations are likely to follow different functional forms.Second, besides product size, the strongest factor influencing cost appears to be team size.Larger teams result in substantially lower productivity, which is interesting considering this attribute is rarely collected in software engineering cost data bases.Third, although some COCOMO factors appear to be useful and significant covariates, they play a minor role in explaining project effort.Overall, the most plausible model appears to be a log-linear model involving KLOC, team size, and a principal component influenced by three COCOMO factors: reliability requirements (RELY), Permission to make digital or hard copies of all or part of this work IhI pcrsond or classroom USC is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies hear this notice and the f'ull citation on the lirst page.TO COPY otherwise, to republish, to post on servers or to rcdistrihutc 10 lists.requires prior specific permission and/or a fee. Lionel C. Briand, Khaled El Emam, Isabella Wieczorek |
ICSE | 2 |
| 1999 | Benchmarking Kappa: Interrater Agreement in Software Process Assessments
Khaled El Emam |
Empir. Softw. Eng. | 1 |
| 1998 | COBRA: A Hybrid Method for Software Cost Estimation, Benchmarking, and Risk AssessmentabstractCurrent cost estimation techniques have a number of drawbacks. For example, developing algorithmic models requires extensive past project data. Also, off-the-shelf models have been found to be difficult to calibrate but inaccurate without calibration. Informal approaches based on experienced estimators depend on estimators' availability and are not easily repeatable, as well as not being much more accurate than algorithmic techniques. We present a method for cost estimation that combines aspects of algorithmic and experiential approaches (referred to as COBRA, COst estimation, Benchmarking, and Risk Assessment). We find through a case study that cost estimates using COBRA show an average ARE of 0.09. Although we do not have the room to describe the benchmarking and risk assessment parts, the reader will find detailed information in (Briand et al., 1997). Lionel C. Briand, Khaled El Emam, Frank Bomarius |
ICSE | 2 |
| 1998 | Using Simulation to Build Inspection Efficiency Benchmarks for Development ProjectsabstractIt is difficult for organizations introducing and using software inspections to evaluate how efficient they are. However, it is of practical importance to determine whether they have been efficiently implemented or whether further corrective actions are necessary to bring them up to standard. We present in this paper a procedure for building inspection efficiency benchmarks based on simulation and typical inspection data. Based on most of the data published in the literature, we build an industry-wide benchmark which intends to capture the current practice regarding inspection efficiency. This benchmark construction procedure can also be used to build enterprise specific benchmarks. Last, we assess how robust we can expect them to be by distorting their input distributions to reflect violations of the assumptions made. Lionel C. Briand, Khaled El Emam, Oliver Laitenberger, Thomas Fussbroich |
ICSE | 2 |
| 1998 | A comparison and integration of capture-recapture models and the detection profile methodabstractIn order to control inspections, the number of remaining defects in software artifacts after their inspection should be estimated. This would allow, for example, deciding whether a reinspection of supposedly faulty artifacts is necessary. Several studies in software engineering have considered capture-recapture models for performing such estimations. These models were initially developed for estimating animal abundance in wildlife research. In addition to these models, researchers in software engineering have recently proposed an alternative approach, namely the detection profile method (DPM), that makes less restrictive assumptions than some capture-recapture models and that show promise in terms of estimation accuracy. The authors investigate how to select between these two approaches for defect content estimation. As a result of this investigation they present a selection procedure taking into account the strength and weaknesses of the two methods. A weakness known for capture-recapture models is that they tend to provide extreme under/over estimation. The existence of such extreme outliers can discourage their use because their consequences in terms of wasted effort or defect slippage can be substantial, and therefore it is not clear,whether a particular estimate can be trusted. The evaluation of the selection procedure with actual inspection data indicates that this selection procedure provides the same accuracy as capture-recapture models alone and DPM alone, and most importantly does not exhibit extreme over/under estimation. Lionel C. Briand, Khaled El Emam, Bernd G. Freimut |
ISSRE | 2 |
| 1998 | The repeatability of code defect classificationsabstractCounts of defects found during the various defect defection activities in software projects and their classification provide a basis for product quality evaluation and process improvement. However, since defect classifications are subjective, it is necessary to ensure that they are repeatable (i.e., that the classification is not dependent on the individual). We evaluate a slight adaptation of a commonly used defect classification scheme that has been applied in IBM's Orthogonal Defect Classification work, and in the SEI's Personal Software Process. The evaluation utilizes the Kappa statistic. We use defect data from code inspections conducted during a development project. Our results indicate that the classification scheme is in general repeatable. We further evaluate classes of defects to find out if confusion between some categories is more common, and suggest a potential improvement to the scheme. Khaled El Emam, Isabella Wieczorek |
ISSRE | 1 |
| 1998 | The Internal Consistencies of the 1987 SEI Maturity Questionnaire and the SPICE Capability Dimension
Pierfrancesco Fusaro, Khaled El Emam, Bob Smith |
Empir. Softw. Eng. | 2 |
| 1997 | Characterizing and Modeling the Cost of Rework in a Library of Reusable Software ComponentsabstractIn this paper we characterize and model the cost of rework in a Component Factory (CF) organization.A CF is responsible for developing and packaging reusable software components.Data was collected on corrective maintenance activities for the Generalized Support Software reuse asset library located at the Flight Dynamics Division of NASA's GSFC.We then constructed a predictive model of the cost of rework using the C4.5 system for generating a logical classification model.The predictor variables for the model are measures of internal software product attributes.The model demonstrates good prediction accuracy, and can be used by managers to allocate resources for corrective maintenance activities.Furthermore, we used the model to generate proscriptive coding guidelines to improve programming practices so that the cost of rework can be reduced in the future.The general approach we have used is applicable to other environments. Victor R. Basili, Steven E. Condon, Khaled El Emam, Robert B. Hendrick, Walcélio L. Melo |
ICSE | 3 |
| 1997 | Causal Analysis of the Requirements Change Process for a Large SystemabstractImplementations of requirements change processes in large system projects face many difficulties. We present a method for analysing requirements change processes to identify implementation weaknesses and their causes. This method relies on prescriptive process models and tracing actual change proposals through the process model. We apply the method in the analysis of the requirements change process for a large real time system within a Canadian government agency. This allowed us to identify the process implementation problems, and the process, organisational, and people causes of these problems. Based on that experience, we draw general conclusions about the proposed method and its applicability Khaled El Emam, Dirk Höltje, Nazim H. Madhavji |
ICSM | 1 |
| 1997 | Quantitative evaluation of capture-recapture models to control software inspectionsabstractAn important requirement to control the inspection of software artifacts is to be able to decide, based on objective information, whether inspection can stop or whether it should continue to achieve a suitable level of artifact quality. Several studies in software engineering have considered the use of capture-recapture models to predict the number of remaining defects in an inspected document as a decision criterion about reinspection. However, no study on software engineering artifacts compares the actual number of remaining defects to the one predicted by a capture-recapture model. Simulations have been performed but no definite conclusions can be drawn regarding the degree of accuracy of such models under realistic inspection conditions, and the factors affecting this accuracy. Furthermore, none of these studies performed an exhaustive comparison of existing models. In this study, we focus on traditional inspections and estimate, based on actual inspection data, the degree of accuracy of all relevant, state-of-the-art, capture-recapture models for which statistical estimators exist. We compare the various models' accuracies and look at the impact of the number of inspectors on these accuracies. Results show that model accuracies are strongly affected by the number of inspectors and, therefore, one must consider this factor before using capture-recapture models. When the number of inspectors is below 4, no model is sufficiently accurate and underestimation may be substantial. In addition, some models perform better than others in a large number of conditions and plausible reasons are discussed. Based on our analyses, we recommend using a model taking into account different probabilities of detecting defects and a Jacknife estimator. Lionel C. Briand, Khaled El Emam, Bernd G. Freimut, Oliver Laitenberger |
ISSRE | 2 |
| 1997 | Reply to ''Comments to the Paper: Briand, El Emam, Morasca: On the Application of Measurement Theory in Software Engineering
Lionel C. Briand, Khaled El Emam, Sandro Morasca |
Empir. Softw. Eng. | 2 |
| 1996 | On the application of measurement theory in software engineering
Lionel C. Briand, Khaled El Emam, Sandro Morasca |
Empir. Softw. Eng. | 2 |
| 1996 | An instrument for measuring the success of the requirements engineering process in information systems development
Khaled El Emam, Nazim H. Madhavji |
Empir. Softw. Eng. | 1 |
| 1996 | User Participation in the Requirements Engineering Process: An Empirical Study
Khaled El Emam, Soizic Quintin, Nazim H. Madhavji |
Requir. Eng. | 1 |
| 1995 | A field study of requirements engineering practices in information systems developmentabstractTo make recommendations for improving requirements engineering processes, it is critical to understand the problems faced in contemporary practice. We describe a field study whose general objectives were to formulate recommendations to practitioners for improving requirements engineering processes, and to provide directions for future research on methods and tools. The results indicate that there are seven key issues of greatest concern in requirements engineering practice. These issues are discussed in terms of the problems they represent, how these problems are addressed successfully in practice, and impediments to the implementation of such good practices. Khaled El Emam, Nazim H. Madhavji |
RE | 1 |
| 1995 | Measuring the success of requirements engineering processesabstractCentral to understanding and improving requirements engineering processes is the ability to measure requirements engineering success. The paper describes a research study whose objective was to develop an instrument to measure the success of requirements engineering processes. The instrument developed consists of 32 indicators that cover the two most important dimensions of requirements engineering success. These two dimensions were identified during the study to be: quality of requirements engineering products and quality of requirements engineering service. Evidence is presented demonstrating that the instrument has desirable psychometric properties, such as high reliability and validity. Khaled El Emam, Nazim H. Madhavji |
RE | 1 |
| 1993 | A Comprehensive Process Model for Studying Software Process Papers
Rudolf K. Keller, Richard Lajoie, Nazim H. Madhavji, Tilmann F. W. Bruckhaus, Kamel Toubache, Won-Kook Hong, Khaled El Emam |
ICSE | 7 |