EDBT 2026 Demo / reviewers in the wild / expert
Foyzur Rahman
dblp:44/8127
· DBLP profile ↗
15ranked-venue papers
8as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 13 · 8 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
9 papers |
Empirical software engineering · 78% Program analysis · 12% Software testing · 5% | |
| Databases, data mining, and information retrieval
2 papers |
Data integration and cleaning · 58% Query processing and optimization · 25% Distributed and cloud data management · 17% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% |
Topics — the 17 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Empirical software engineering › mining software repositories
defect prediction |
0.7 | 5 | 2014 | Sample size vs. bias in defect prediction · ESEC/SIGSOFT FSE 2013 How, and why, process metrics are better · ICSE 2013 Recalling the "imprecision" of cross-project defect prediction · SIGSOFT FSE 2012 |
Data integration and cleaning › data warehouse
cloud data warehouse |
0.6 | 1 | 2022 | Amazon Redshift Re-invented · SIGMOD Conference 2022 |
Empirical software engineering
mining software repositories |
0.4 | 4 | 2012 | Ownership, experience and defects: a fine-grained study of authorship · ICSE 2011 LINKSTER: enabling efficient manual inspection and annotation of mined data · SIGSOFT FSE 2010 The missing links: bugs and bug-fix commits · SIGSOFT FSE 2010 |
Empirical software engineering
software analytics |
0.3 | 2 | 2013 | Sample size vs. bias in defect prediction · ESEC/SIGSOFT FSE 2013 Recalling the "imprecision" of cross-project defect prediction · SIGSOFT FSE 2012 |
Query processing and optimization › query optimization
query optimizer architecture |
0.2 | 1 | 2014 | Orca: a modular query optimizer architecture for big data · SIGMOD Conference 2014 |
Program analysis › static analysis
bug detection |
0.2 | 1 | 2014 | Comparing static bug finders and statistical prediction · ICSE 2014 |
Empirical software engineering › software metrics
process metrics |
0.2 | 1 | 2013 | How, and why, process metrics are better · ICSE 2013 |
Software testing › fault detection
bug finding tools |
0.1 | 1 | 2012 | To what extent could we detect field defects? an empirical study of false negatives in static bug finding tools · ASE 2012 |
Empirical software engineering › mining software repositories › defect prediction
cross-project defect prediction |
0.1 | 1 | 2012 | Recalling the "imprecision" of cross-project defect prediction · SIGSOFT FSE 2012 |
Program analysis
static analysis |
0.1 | 1 | 2012 | To what extent could we detect field defects? an empirical study of false negatives in static bug finding tools · ASE 2012 |
Empirical software engineering › mining software repositories › developer activity analysis
code ownership |
0.1 | 1 | 2011 | Ownership, experience and defects: a fine-grained study of authorship · ICSE 2011 |
Empirical software engineering › developer studies
developer experience |
0.1 | 1 | 2011 | Ownership, experience and defects: a fine-grained study of authorship · ICSE 2011 |
Empirical software engineering › mining software repositories › commit analysis
bug-fixing commit identification |
0.1 | 1 | 2010 | The missing links: bugs and bug-fix commits · SIGSOFT FSE 2010 |
Query processing and optimization
analytical query processing |
0.1 | 1 | 2014 | Orca: a modular query optimizer architecture for big data · SIGMOD Conference 2014 |
Empirical software engineering › software metrics
code metrics |
0.0 | 1 | 2013 | How, and why, process metrics are better · ICSE 2013 |
Software maintenance and evolution › code review
code inspection |
0.0 | 1 | 2011 | BugCache for inspections: hit or miss? · SIGSOFT FSE 2011 |
Empirical software engineering
software repositories |
0.0 | 1 | 2010 | The missing links: bugs and bug-fix commits · SIGSOFT FSE 2010 |
Methods — techniques the papers use, named apart from their topics
statistical prediction · 0.2static analysis · 0.2statistical prediction modeling · 0.2simulation · 0.2meta-analysis · 0.2precision-recall analysis · 0.1logistic regression · 0.1empirical study · 0.1version history analysis · 0.1empirical evaluation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Amazon Redshift Re-inventedabstractIn 2013, AmazonWeb Services revolutionized the data warehousing industry by launching Amazon Redshift, the first fully-managed, petabyte-scale, enterprise-grade cloud data warehouse. Amazon Redshift made it simple and cost-effective to efficiently analyze large volumes of data using existing business intelligence tools. This cloud service was a significant leap from the traditional on-premise data warehousing solutions, which were expensive, not elastic, and required significant expertise to tune and operate. Customers embraced Amazon Redshift and it became the fastest growing service in AWS. Today, tens of thousands of customers use Redshift in AWS's global infrastructure to process exabytes of data daily. Nikos Armenatzoglou, Sanuj Basu, Naga Bhanoori, Mengchu Cai, Naresh Chainani, Kiran Chinta, Venkatraman Govindaraju, Todd J. Green, Monish Gupta, Sebastian Hillig, Eric Hotinger, Yan Leshinksy, Jintian Liang, Michael McCreedy, Fabian Nagel, Ippokratis Pandis, Panos Parchas, Rahul Pathak, Orestis Polychroniou, Foyzur Rahman, Gokul Soundararajan, Sriram Subramanian, Douglas B. Terry |
SIGMOD Conference | 20 |
| 2015 | To what extent could we detect field defects? An extended empirical study of false negatives in static bug-finding tools
Ferdian Thung, Lucia, David Lo 0001, Lingxiao Jiang, Foyzur Rahman, Premkumar T. Devanbu |
Autom. Softw. Eng. | 5 |
| 2014 | Comparing static bug finders and statistical predictionabstractThe all-important goal of delivering better software at lower cost has led to a vital, enduring quest for ways to find and remove defects efficiently and accurately. To this end, two parallel lines of research have emerged over the last years. Static analysis seeks to find defects using algorithms that process well-defined semantic abstractions of code. Statistical defect prediction uses historical data to estimate parameters of statistical formulae modeling the phenomena thought to govern defect occurrence and predict where defects are likely to occur. These two approaches have emerged from distinct intellectual traditions and have largely evolved independently, in “splendid isolation”. In this paper, we evaluate these two (largely) disparate approaches on a similar footing. We use historical defect data to apprise the two approaches, compare them, and seek synergies. We find that under some accounting principles, they provide comparable benefits; we also find that in some settings, the performance of certain static bug-finders can be enhanced using information provided by statistical defect prediction. Foyzur Rahman, Sameer Khatri, Earl T. Barr, Premkumar T. Devanbu |
ICSE | 1 |
| 2014 | Orca: a modular query optimizer architecture for big dataabstractThe performance of analytical query processing in data management systems depends primarily on the capabilities of the system's query optimizer. Increased data volumes and heightened interest in processing complex analytical queries have prompted Pivotal to build a new query optimizer. Mohamed A. Soliman, Lyublena Antova, Venkatesh Raghavan, Amr El-Helw, Zhongxian Gu, Entong Shen, George C. Caragea, Carlos Garcia-Alvarado, Foyzur Rahman, Michalis Petropoulos, F. Michael Waas, Sivaramakrishnan Narayanan, Konstantinos Krikellas, Rhonda Baldwin |
SIGMOD Conference | 9 |
| 2013 | How, and why, process metrics are betterabstractDefect prediction techniques could potentially help us to focus quality-assurance efforts on the most defect-prone files. Modern statistical tools make it very easy to quickly build and deploy prediction models. Software metrics are at the heart of prediction models; understanding how and especially why different types of metrics are effective is very important for successful model deployment. In this paper we analyze the applicability and efficacy of process and code metrics from several different perspectives. We build many prediction models across 85 releases of 12 large open source projects to address the performance, stability, portability and stasis of different sets of metrics. Our results suggest that code metrics, despite widespread use in the defect prediction literature, are generally less useful than process metrics for prediction. Second, we find that code metrics have high stasis; they don't change very much from release to release. This leads to stagnation in the prediction models, leading to the same files being repeatedly predicted as defective; unfortunately, these recurringly defective files turn out to be comparatively less defect-dense. Foyzur Rahman, Premkumar T. Devanbu |
ICSE | 1 |
| 2013 | Sample size vs. bias in defect predictionabstractMost empirical disciplines promote the reuse and sharing of datasets, as it leads to greater possibility of replication. While this is increasingly the case in Empirical Software Engineering, some of the most popular bug-fix datasets are now known to be biased. This raises two significant concerns: first, that sample bias may lead to underperforming prediction models, and second, that the external validity of the studies based on biased datasets may be suspect. This issue has raised considerable consternation in the ESE literature in recent years. However, there is a confounding factor of these datasets that has not been examined carefully: size. Biased datasets are sampling only some of the data that could be sampled, and doing so in a biased fashion; but biased samples could be smaller, or larger. Smaller data sets in general provide less reliable bases for estimating models, and thus could lead to inferior model performance. In this setting, we ask the question, what affects performance more, bias, or size? We conduct a detailed, large-scale meta-analysis, using simulated datasets sampled with bias from a high-quality dataset which is relatively free of bias. Our results suggest that size always matters just as much bias direction, and in fact much more than bias direction when considering information-retrieval measures such as AUCROC and F-score. This indicates that at least for prediction models, even when dealing with sampling bias, simply finding larger samples can sometimes be sufficient. Our analysis also exposes the complexity of the bias issue, and raises further issues to be explored in the future. Foyzur Rahman, Daryl Posnett, Israel Herraiz, Premkumar T. Devanbu |
ESEC/SIGSOFT FSE | 1 |
| 2012 | When would this bug get reported?abstractNot all bugs in software would be experienced and reported by end users right away: Some bugs manifest themselves quickly and may be reported by users a few days after they get into the code base; others manifest many months or even years later, and may only be experienced and reported by a small number of users. We refer to the period of time between the time when a bug is introduced into code and the time when it is reported by a user as bug reporting latency. Knowledge of bug reporting latencies has an implication on prioritization of bug fixing activities-bugs with low reporting latencies may be fixed earlier than those with high latencies to shift debugging resources towards bugs highly concerning users. To investigate bug reporting latencies, we analyze bugs from three Java software systems: AspectJ, Rhino, and Lucene. We extract bug reporting data from their version control repositories and bug tracking systems, identify bug locations based on bug fixes, and back-trace bug introducing time based on change histories of the buggy code. Also, we remove non-essential changes, and most importantly, recover root causes of bugs from their treatments/fixes. We then calculate the bug reporting latencies, and find that bugs have diverse reporting latencies. Based on the calculated reporting latencies and features we extract from bugs, we build classification models that can predict whether a bug would be reported early (within 30 days) or later, which may be helpful for prioritizing bug fixing activities. Our evaluation on the three software systems shows that our bug reporting latency prediction models could achieve an AUC (Area Under the Receiving Operating Characteristics Curve) of 70.869%. Ferdian Thung, David Lo 0001, Lingxiao Jiang, Lucia, Foyzur Rahman, Premkumar T. Devanbu |
ICSM | 5 |
| 2012 | To what extent could we detect field defects? an empirical study of false negatives in static bug finding toolsabstractSoftware defects can cause much loss. Static bug-finding tools are believed to help detect and remove defects. These tools are designed to find programming errors; but, do they in fact help prevent actual defects that occur in the field and reported by users? If these tools had been used, would they have detected these field defects, and generated warnings that would direct programmers to fix them? To answer these questions, we perform an empirical study that investigates the effectiveness of state-of-the-art static bug finding tools on hundreds of reported and fixed defects extracted from three open source programs: Lucene, Rhino, and AspectJ. Our study addresses the question: To what extent could field defects be found and detected by state-of-the-art static bug-finding tools? Different from past studies that are concerned with the numbers of false positives produced by such tools, we address an orthogonal issue on the numbers of false negatives. We find that although many field defects could be detected by static bug finding tools, a substantial proportion of defects could not be flagged. We also analyze the types of tool warnings that are more effective in finding field defects and characterize the types of missed defects. Ferdian Thung, Lucia, David Lo 0001, Lingxiao Jiang, Foyzur Rahman, Premkumar T. Devanbu |
ASE | 5 |
| 2012 | Recalling the "imprecision" of cross-project defect predictionabstractThere has been a great deal of interest in defect prediction: using prediction models trained on historical data to help focus quality-control resources in ongoing development. Since most new projects don't have historical data, there is interest in cross-project prediction: using data from one project to predict defects in another. Sadly, results in this area have largely been disheartening. Most experiments in cross-project defect prediction report poor performance, using the standard measures of precision, recall and F-score. We argue that these IR-based measures, while broadly applicable, are not as well suited for the quality-control settings in which defect prediction models are used. Specifically, these measures are taken at specific threshold settings (typically thresholds of the predicted probability of defectiveness returned by a logistic regression model). However, in practice, software quality control processes choose from a range of time-and-cost vs quality tradeoffs: how many files shall we test? how many shall we inspect? Thus, we argue that measures based on a variety of tradeoffs, viz., 5%, 10% or 20% of files tested/inspected would be more suitable. We study cross-project defect prediction from this perspective. We find that cross-project prediction performance is no worse than within-project performance, and substantially better than random prediction! Foyzur Rahman, Daryl Posnett, Premkumar T. Devanbu |
SIGSOFT FSE | 1 |
| 2012 | Clones: what is that smell?
Foyzur Rahman, Christian Bird, Premkumar T. Devanbu |
Empir. Softw. Eng. | 1 |
| 2011 | Ownership, experience and defects: a fine-grained study of authorshipabstractRecent research indicates that "people" factors such as ownership, experience, organizational structure, and geographic distribution have a big impact on software quality. Understanding these factors, and properly deploying people resources can help managers improve quality outcomes. This paper considers the impact of code ownership and developer experience on software quality. In a large project, a file might be entirely owned by a single developer, or worked on by many. Some previous research indicates that more developers working on a file might lead to more defects. Prior research considered this phenomenon at the level of modules or files, and thus does not tease apart and study the effect of contributions of different developers to each module or file. We exploit a modern version control system to examine this issue at a fine-grained level. Using version history, we examine contributions to code fragments that are actually repaired to fix bugs. Are these code fragments "implicated" in bugs the result of contributions from many? or from one? Does experience matter? What type of experience? We find that implicated code is more strongly associated with a single developer's contribution; our findings also indicate that an author's specialized experience in the target file is more important than general experience. Our findings suggest that quality control efforts could be profitably targeted at changes made by single developers with limited prior experience on that file. Foyzur Rahman, Premkumar T. Devanbu |
ICSE | 1 |
| 2011 | BugCache for inspections: hit or miss?abstractInspection is a highly effective but costly technique for quality control. Most companies do not have the resources to inspect all the code; thus accurate defect prediction can help focus available inspection resources. BugCache is a simple, elegant, award-winning prediction scheme that "caches" files that are likely to contain defects [12]. In this paper, we evaluate the utility of BugCache as a tool for focusing inspection, we examine the assumptions underlying BugCache with the aim of improving it, and finally we compare it with a simple, standard bug-prediction technique. We find that BugCache is, in fact, useful for focusing inspection effort; but surprisingly, we find that its performance, when used for inspections, is not much better than a naive prediction model -- viz., a model that orders files in the system by their count of closed bugs and chooses enough files to capture 20% of the lines in the system. Foyzur Rahman, Daryl Posnett, Abram Hindle, Earl T. Barr, Premkumar T. Devanbu |
SIGSOFT FSE | 1 |
| 2010 | Clones: What is that smell?abstractClones are generally considered bad programming practice in software engineering folklore. They are identified as a bad smell and a major contributor to project maintenance difficulties. Clones inherently cause code bloat, thus increasing project size and maintenance costs. In this work, we try to validate the conventional wisdom empirically to see whether cloning makes code more defect prone. This paper analyses relationship between cloning and defect proneness. We find that, first, the great majority of bugs are not significantly associated with clones. Second, we find that clones may be less defect prone than non-cloned code. Finally, we find little evidence that clones with more copies are actually more error prone. Our findings do not support the claim that clones are really a ¿bad smell¿. Perhaps we can clone, and breathe easy, at the same time. Foyzur Rahman, Christian Bird, Premkumar T. Devanbu |
MSR | 1 |
| 2010 | The missing links: bugs and bug-fix commitsabstractEmpirical studies of software defects rely on links between bug databases and program code repositories. This linkage is typically based on bug-fixes identified in developer-entered commit logs. Unfortunately, developers do not always report which commits perform bug-fixes. Prior work suggests that such links can be a biased sample of the entire population of fixed bugs. The validity of statistical hypotheses-testing based on linked data could well be affected by bias. Given the wide use of linked defect data, it is vital to gauge the nature and extent of the bias, and try to develop testable theories and models of the bias. To do this, we must establish ground truth: manually analyze a complete version history corpus, and nail down those commits that fix defects, and those that do not. This is a diffcult task, requiring an expert to compare versions, analyze changes, find related bugs in the bug database, reverse-engineer missing links, and finally record their work for use later. This effort must be repeated for hundreds of commits to obtain a useful sample of reported and unreported bug-fix commits. We make several contributions. First, we present Linkster, a tool to facilitate link reverse-engineering. Second, we evaluate this tool, engaging a core developer of the Apache HTTP web server project to exhaustively annotate 493 commits that occurred during a six week period. Finally, we analyze this comprehensive data set, showing that there are serious and consequential problems in the data. Adrian Bachmann, Christian Bird, Foyzur Rahman, Premkumar T. Devanbu, Abraham Bernstein |
SIGSOFT FSE | 3 |
| 2010 | LINKSTER: enabling efficient manual inspection and annotation of mined dataabstractWhile many uses of mined software engineering data are automatic in nature, some techniques and studies either require, or can be improved, by manual methods. Unfortunately, manually inspecting, analyzing, and annotating mined data can be difficult and tedious, especially when information from multiple sources must be integrated. Oddly, while there are numerous tools and frameworks for automatically mining and analyzing data, there is a dearth of tools which facilitate manual methods. To fill this void, we have developed LINKSTER, a tool which integrates data from bug databases, source code repositories, and mailing list archives to allow manual inspection and annotation. LINKSTER has already been used successfully by an OSS project lead to obtain data for one empirical study. Christian Bird, Adrian Bachmann, Foyzur Rahman, Abraham Bernstein |
SIGSOFT FSE | 3 |