VLDB 2026 Research / reviewers in the wild / expert
Akif Günes Koru
dblp:41/5540 · also Günes Koru
· DBLP profile ↗
30ranked-venue papers
7as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 14 · 6 first-authorSecurity and privacy · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Software maintenance and evolution · 50% Empirical software engineering · 50% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Empirical software engineering › mining software repositories
defect prediction |
0.1 | 1 | 2009 | An Investigation into the Functional Form of the Size-Defect Relationship for Software Modules · IEEE Trans. Software Eng. 2009 |
Software maintenance and evolution › software quality assurance
quality assurance prioritization |
0.1 | 1 | 2009 | An Investigation into the Functional Form of the Size-Defect Relationship for Software Modules · IEEE Trans. Software Eng. 2009 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.1 | 1 | 2007 | AffyProbeMiner: a web resource for computing or retrieving accurately redefined Affymetrix probe sets · Bioinform. 2007 |
Empirical software engineering
mining software repositories |
0.1 | 1 | 2005 | Comparing High-Change Modules and Modules with the Highest Measurement Values in Two Large-Scale Open-Source Products · IEEE Trans. Software Eng. 2005 |
Bioinformatics and computational biology
transcriptomics |
0.0 | 1 | 2007 | AffyProbeMiner: a web resource for computing or retrieving accurately redefined Affymetrix probe sets · Bioinform. 2007 |
Methods — techniques the papers use, named apart from their topics
regression analysis · 0.1class-level defect data analysis · 0.1refseq · 0.1genbank · 0.1structural measures · 0.1hypothesis testing · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Challenges and Opportunities of Information Technology Adoption by the Caregivers of Home Care Patients
Onimi Jademi, Abir Rahman, Akif Günes Koru |
AMIA | 3 |
| 2020 | Development and Evaluation of the Data Quality Toolkit: Toward Collaborative Improvement of Data Quality in Healthcare Organizations
Abir Rahman, Forhan B. Emdad, Akif Günes Koru |
AMIA | 3 |
| 2020 | An Effective and Computationally Efficient Approach for Anonymizing Large-Scale Physical Activity Data: Multi-Level Clustering-Based AnonymizationabstractPublishing physical activity data can facilitate reproducible health-care research in several areas such as population health management, behavioral health research, and management of chronic health problems. However, publishing such data also brings high privacy risks related to re-identification which makes anonymization necessary. One of the challenges in anonymizing physical activity data collected periodically is its sequential nature. The existing anonymization techniques work sufficiently for cross-sectional data but have high computational costs when applied directly to sequential data. This article presents an effective anonymization approach, multi-level clustering-based anonymization to anonymize physical activity data. Compared with the conventional methods, the proposed approach improves time complexity by reducing the clustering time drastically. While doing so, it preserves the utility as much as the conventional approaches. Pooja Parameshwarappa, Zhiyuan Chen 0003, Akif Günes Koru |
Int. J. Inf. Secur. Priv. | 3 |
| 2020 | Understanding and detecting defects in healthcare administration data: Toward higher data quality to better support healthcare operations and decisionsabstractOBJECTIVE: Development of systematic approaches for understanding and assessing data quality is becoming increasingly important as the volume and utilization of health data steadily increases. In this study, a taxonomy of data defects was developed and utilized when automatically detecting defects to assess Medicaid data quality maintained by one of the states in the United States. MATERIALS AND METHODS: There were more than 2.23 million rows and 32 million cells in the Medicaid data examined. The taxonomy was developed through document review, descriptive data analysis, and literature review. A software program was created to automatically detect defects by using a set of constraints whose development was facilitated by the taxonomy. RESULTS: Five major categories and seventeen subcategories of defects were identified. The major categories are missingness, incorrectness, syntax violation, semantic violation, and duplicity. More than 3 million defects were detected indicating substantial problems with data quality. Defect density exceeded 10% in five tables. The majority of the data defects belonged to format mismatch, invalid code, dependency-contract violation, and implausible value types. Such contextual knowledge can support prioritized quality improvement initiatives for the Medicaid data studied. CONCLUSIONS: This research took the initial steps to understand the types of data defects and detect defects in large healthcare datasets. The results generally suggest that healthcare organizations can potentially benefit from focusing on data quality improvement. For those purposes, the taxonomy developed and the approach followed in this study can be adopted. Akif Günes Koru |
J. Am. Medical Informatics Assoc. | 2 |
| 2019 | Towards Optimizing Usability in e-Government Healthcare Administration Systems: Recommendations from a Comprehensive Evaluation Study
Forhan B. Emdad, Akif Günes Koru |
AMIA | 2 |
| 2019 | Addressing the Pressing Home Care Coordination Challenges and Opportunities: A Literature Review
Mohammad Ishtiaque Rahman, Akif Günes Koru |
AMIA | 2 |
| 2019 | Challenges and Opportunities of Data Quality Improvement in Healthcare Organizations
Urmita Banerjee, Akif Günes Koru |
AMIA | 3 |
| 2017 | Evaluation and Adoption of Security Standards to Improve Information Security in Healthcare Administration Systems
Akif Günes Koru |
AMIA | 1 |
| 2017 | Towards Improving the Quality of Medicaid Data: Identifying Data Defects in a Large Healthcare Administration System
Pratik Tamakuwala, Akif Günes Koru |
AMIA | 3 |
| 2016 | Investigating How Health IT Solutions Responded to Information Needs for Fall-Risk Management: A Case Study
Dari Alhuwail, Akif Günes Koru |
AMIA | 2 |
| 2016 | Informing the Requirements for IT-Based Fall Risk Management Solutions in Home Health Care
Onimi Jademi, Uchenna A. Uchidiuno, Dari Alhuwail, Akif Günes Koru |
AMIA | 4 |
| 2016 | Testing the theory of relative dependency from an evolutionary perspective: higher dependencies concentration in smaller modules over the lifetime of software productsabstractAbstract Recent studies conducted on the single releases of multiple software products showed that dependencies concentrate on smaller modules, that is, smaller modules have more dependencies per source line of code. This phenomenon, called the Theory of Relative Dependency, explains why some earlier studies reported that smaller modules were proportionally more defect prone. It is important to test the Theory of Relative Dependency from multiple perspectives so that it can be used as an explanatory argument when garnering organizational support to give a higher quality assurance (QA) priority to smaller modules. In this study, we test the validity of this theory from an evolutionary perspective by examining the consecutive releases of a number of software products. Dependencies do concentrate over smaller modules regardless of the product age. Furthermore, continuous refactoring efforts are associated with increasing concentration of dependencies on smaller modules over product lifetime. Based on the consistent results, software managers and developers should consider giving a higher QA priority to smaller modules over the lifetime of a software product. In the projects where refactoring is adopted continuously, the QA priority on smaller modules should be further increased as the software product ages. Copyright © 2016 John Wiley & Sons, Ltd. Yixin Bian, Mohammed Aziz Parande, Akif Günes Koru |
J. Softw. Evol. Process. | 3 |
| 2015 | Identifying Home Care Clinical Practices Most Associated with Hospital Readmissions and Non-Admitted ER Visit Rates: Secondary Data Analysis
Dari Alhuwail, Akif Günes Koru |
AMIA | 2 |
| 2014 | Health Information Technology Adoption in Home Health ... Research in Progress
Dari Alhuwail, Akif Günes Koru, Ahmad Alaiad 0001, Anthony F. Norcio, Maxim Topaz |
AMIA | 2 |
| 2014 | Corrigendum to: "SPAPE: A semantic-preserving amorphous procedure extraction method for near-miss clones": [J. Syst. Softw. 86 (2013) 2077-2093]
Yixin Bian, Akif Günes Koru, Xiaohong Su, Peijun Ma |
J. Syst. Softw. | 2 |
| 2013 | "We're all in our own little island": A Qualitative Exploration of Patient Information Exchange during Admission to Home Health Agency
Maxim Topaz, D. Molkina, Akif Günes Koru, Ruth M. Masterson Creber, O. Jarrin, Kavita Radhakrishnan, Melissa O'Connor, Kathryn H. Bowles |
AMIA | 3 |
| 2013 | Predictive models in software engineering
Tim Menzies, Akif Günes Koru |
Empir. Softw. Eng. | 2 |
| 2013 | SPAPE: A semantic-preserving amorphous procedure extraction method for near-miss clones
Yixin Bian, Akif Günes Koru, Xiaohong Su, Peijun Ma |
J. Syst. Softw. | 2 |
| 2012 | Studying volatility predictors in open source softwareabstractVolatile software modules, for the purposes of this work, are defined as those that are significantly more change-prone than other modules in the same system or subsystem. There is significant literature investigating models for predicting which modules in a system will become volatile, and/or are defect-prone. Much of this work focuses on using source code-related characteristics (e.g., complexity metrics) and simple change metrics (e.g., number of past changes) as inputs to the predictive models. Our work attempts to broaden the array of factors considered in such prediction approaches. To this end, we collected data directly from development personnel about the factors they rely on to foresee what parts of a system are going to become volatile. In this paper, we describe a focus group study conducted with the development team of a small but active open source project, in which we asked this very question. The results of the focus group indicate, among other things, that a period of volatility in a particular area of the system is often predicted by a pattern characterized by inactivity in a certain area (resulting in that area becoming less mature than others), increased communication between developers regarding opportunities for improvement in that area, and then the emergence of a champion who takes the initiative to start working on those improvements. The initial changes lead to more changes (both to extend the improvements already made and to fix problems introduced), thus leading to volatility. Brandt Braunschweig, Neha Dhage, Maria Jose Viera, Carolyn B. Seaman, Sreedevi Sampath, Akif Günes Koru |
ESEM | 6 |
| 2010 | A longitudinal analysis of the dependency concentration in smaller modules for open-source software productsabstractOur recent studies on single releases of multiple open-source software (OSS) products showed a higher concentration of dependencies in smaller modules. For one of the products, it was observed that an isolatable and observable refactoring initiative exacerbated this concentration inequality. In this paper, we empirically investigate the dependency concentration in smaller modules from a longitudinal perspective: (1) whether this concentration inequality always exists over product life time; (2) how it changes. We hypothesize that the concentration inequality should either remain at same levels or increase over time. This is because large-scale and long-lived software products usually go through some degree of continuous and intermittent refactoring. Our results show that dependencies concentrate in smaller classes in all releases, and this concentration inequality generally increases over successive releases. We suggest that software practitioners continuously pay a higher QA attention to smaller modules. We also recommend increasing such QA focus as a product matures and goes through refactoring activities. Mohammed Aziz Parande, Akif Günes Koru |
ICSM | 2 |
| 2010 | Testing the theory of relative defect proneness for closed-source software
Akif Günes Koru, Dongsong Zhang, Khaled El Emam |
Empir. Softw. Eng. | 1 |
| 2009 | Software Engineering Education for BioinformaticsabstractAs software engineering educators, it is important for us to realize the increasing domain-specificity of software, and incorporate these changes in our design of teaching material. Bioinformatics software is an example of immensely complex and critical scientific software and this domain provides an excellent illustration of the role of computing in the life sciences. To study bioinformatics from a software engineering standpoint, we conducted an exploratory survey of bioinformatics developers. The survey had a range of questions about people, processes and products. We learned that practices like extreme programming, requirements engineering and documentation. As software engineering educators, we realized that the survey results had important implications for the education of bioinformatics professionals. We also investigated the current status of software engineering education in bioinformatics, by examining the curricula of more than fifty bioinformatics programs and the contents of over fifteen textbooks. We observed that there was no mention of the role and importance of software engineering practices essential for creating dependable software systems. Based on our findings and existing literature we present a set of recommendations for improving software engineering education in bioinformatics. Medha Umarji, Carolyn B. Seaman, Akif Günes Koru |
CSEE&T | 3 |
| 2009 | An Investigation into the Functional Form of the Size-Defect Relationship for Software ModulesabstractThe importance of the relationship between size and defect proneness of software modules is well recognized. Understanding the nature of that relationship can facilitate various development decisions related to prioritization of quality assurance activities. Overall, the previous research only drew a general conclusion that there was a monotonically increasing relationship between module size and defect proneness. In this study, we analyzed class-level size and defect data in order to increase our understanding of this crucial relationship. In order to obtain validated and more generalizable results, we studied four large-scale object-oriented products, Mozilla, Cn3d, JBoss, and Eclipse. Our results consistently revealed a significant effect of size on defect proneness; however, contrary to common intuition, the size-defect relationship took a logarithmic form, indicating that smaller classes were proportionally more problematic than larger classes. Therefore, practitioners should consider giving higher priority to smaller modules when planning focused quality assurance activities with limited resources. For example, in Mozilla and Eclipse, an inspection strategy investing 80% of available resources on 100-LOC classes and the rest on 1,000-LOC classes would be more than twice as cost effective as the opposite strategy. These results should be immediately useful to guide focused quality assurance activities in large-scale software projects. Akif Günes Koru, Dongsong Zhang, Khaled El Emam |
IEEE Trans. Software Eng. | 1 |
| 2008 | Prioritizing User-Session-Based Test Cases for Web Applications TestingabstractWeb applications have rapidly become a critical part of business for many organizations. However, increased usage of Web applications has not been reciprocated with corresponding increases in reliability. Unique characteristics, such as quick turnaround time, coupled with growing popularity motivate the need for efficient and effective Web application testing strategies. In this paper, we propose several new test suite prioritization strategies for Web applications and examine whether these strategies can improve the rate of fault detection for three Web applications and their preexisting test suites. We prioritize test suites by test lengths, frequency of appearance of request sequences, and systematic coverage of parameter-values and their interactions. Experimental results show that the proposed prioritization criteria often improve the rate of fault detection of the test suites when compared to random ordering of test cases. In general, the best prioritization metrics either (1) consider frequency of appearance of sequences of requests or (2) systematically cover combinations of parameter-values as early as possible. Sreedevi Sampath, Renée C. Bryce, Gokulanand Viswanath, Vani Kandimalla, Akif Günes Koru |
ICST | 5 |
| 2008 | Theory of relative defect proneness
Akif Günes Koru, Khaled El Emam, Dongsong Zhang, Divya Mathew |
Empir. Softw. Eng. | 1 |
| 2007 | AffyProbeMiner: a web resource for computing or retrieving accurately redefined Affymetrix probe setsabstractMOTIVATION: Affymetrix microarrays are widely used to measure global expression of mRNA transcripts. That technology is based on the concept of a probe set. Individual probes within a probe set were originally designated by Affymetrix to hybridize with the same unique mRNA transcript. Because of increasing accuracy in knowledge of genomic sequences, however, a substantial number of the manufacturer's original probe groupings and mappings are now known to be inaccurate and must be corrected. Otherwise, analysis and interpretation of an Affymetrix microarray experiment will be in error. RESULTS: AffyProbeMiner is a computationally efficient platform-independent tool that uses all RefSeq mature RNA protein coding transcripts and validated complete coding sequences in GenBank to (1) regroup the individual probes into consistent probe sets and (2) remap the probe sets to the correct sets of mRNA transcripts. The individual probes are grouped into probe sets that are 'transcript-consistent' in that they hybridize to the same mRNA transcript (or transcripts) and, therefore, measure the same entity (or entities). About 65.6% of the probe sets on the HG-U133A chip were affected by the remapping. Pre-computed regrouped and remapped probe sets for many Affymetrix microarrays are made freely available at the AffyProbeMiner web site. Alternatively, we provide a web service that enables the user to perform the remapping for any type of short-oligo commercial or custom array that has an Affymetrix-format Chip Definition File (CDF). Important features that differentiate AffyProbeMiner from other approaches are flexibility in the handling of splice variants, computational efficiency, extensibility, customizability and user-friendliness of the interface. AVAILABILITY: The web interface and software (GPL open source license), are publicly-accessible at http://discover.nci.nih.gov/affyprobeminer. Barry Zeeberg, Gang Qu 0001, Akif Günes Koru, Alessandro Ferrucci, Ari B. Kahn, Michael C. Ryan, Antej Nuhanovic, Peter J. Munson, William C. Reinhold, David W. Kane, John N. Weinstein |
Bioinform. | 4 |
| 2007 | Identifying and characterizing change-prone classes in two large-scale open-source products
Akif Günes Koru |
J. Syst. Softw. | 1 |
| 2005 | Comparing High-Change Modules and Modules with the Highest Measurement Values in Two Large-Scale Open-Source ProductsabstractIdentifying change-prone modules can enable software developers to take focused preventive actions that can reduce maintenance costs and improve quality. Some researchers observed a correlation between change proneness and structural measures, such as size, coupling, cohesion, and inheritance measures. However, the modules with the highest measurement values were not found to be the most troublesome modules by some of our colleagues in industry, which was confirmed by our previous study of six large-scale industrial products. To obtain additional evidence, we identified and compared high-change modules and modules with the highest measurement values in two large-scale open-source products, Mozilla and OpenOffice, and we characterized the relationship between them. Contrary to common intuition, we found through formal hypothesis testing that the top modules in change-count rankings and the modules with the highest measurement values were different. In addition, we observed that high-change modules had fairly high places in measurement rankings, but not the highest places. The accumulated findings from these two open-source products, together with our previous similar findings for six closed-source products, should provide practitioners with additional guidance in identifying the change-prone modules. Akif Günes Koru, Jeff Tian |
IEEE Trans. Software Eng. | 1 |
| 2003 | A Hierarchical Strategy for Testing Web-Based Applications and Ensuring Their ReliabilityabstractAfter examining the specific problems of testing and quality assurance for Web-based applications, we propose a strategy by integrating existing testing techniques and reliability analyses in a hierarchical framework. This strategy combines various usage models for statistical testing to perform high level testing and to guide selective testing of critical and frequently used subparts or components using traditional coverage-based structural testing. Reliability analysis and risk identification form an integral part of this strategy to help assure and improve the overall reliability for Web-based applications. Some preliminary results are included to demonstrate the general viability and effectiveness of our approach. Jeff Tian, Akif Günes Koru |
COMPSAC | 4 |
| 2003 | An empirical comparison and characterization of high defect and high complexity modules
Akif Günes Koru, Jeff Tian |
J. Syst. Softw. | 1 |