VLDB 2026 Research / reviewers in the wild / expert
Ganesha Upadhyaya
dblp:156/9545 · also Ganesha B. Upadhyaya
· DBLP profile ↗
6ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0003-0764-2578ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
5 papers |
Program analysis · 46% Empirical software engineering · 42% Concurrent programming · 6% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Empirical software engineering
mining software repositories |
0.8 | 3 | 2020 | On Accelerating Source Code Analysis at Massive Scale · IEEE Trans. Software Eng. 2018 Are code examples on an online Q&A forum reliable?: a study of API misuse on stack overflow · ICSE 2018 BCFA: bespoke control flow analysis for CFA at scale · ICSE 2020 |
Program analysis
control flow analysis |
0.8 | 2 | 2020 | BCFA: bespoke control flow analysis for CFA at scale · ICSE 2020 On Accelerating Source Code Analysis at Massive Scale · IEEE Trans. Software Eng. 2018 |
Empirical software engineering › mining software repositories
large-scale code analysis |
0.5 | 2 | 2020 | On Accelerating Source Code Analysis at Massive Scale · IEEE Trans. Software Eng. 2018 BCFA: bespoke control flow analysis for CFA at scale · ICSE 2020 |
Program analysis
API misuse |
0.3 | 1 | 2018 | Are code examples on an online Q&A forum reliable?: a study of API misuse on stack overflow · ICSE 2018 |
Empirical software engineering › mining software repositories › source-code mining
API usage mining |
0.3 | 1 | 2018 | Are code examples on an online Q&A forum reliable?: a study of API misuse on stack overflow · ICSE 2018 |
Program analysis
data flow analysis |
0.3 | 1 | 2018 | Collective program analysis · ICSE 2018 |
Runtime systems and virtual machines › virtual machine implementation
java virtual machine |
0.2 | 1 | 2015 | Effectively mapping linguistic abstractions for message-passing concurrency to threads on the Java virtual machine · OOPSLA 2015 |
Concurrent programming
message passing |
0.2 | 1 | 2015 | Effectively mapping linguistic abstractions for message-passing concurrency to threads on the Java virtual machine · OOPSLA 2015 |
Methods — techniques the papers use, named apart from their topics
machine learning-based strategy prediction · 0.4sparse representation · 0.3distributed computing · 0.3data flow analysis · 0.3control flow analysis · 0.3clustering · 0.3canonical labeling · 0.3API usage pattern mining · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | BCFA: bespoke control flow analysis for CFA at scaleabstractMany data-driven software engineering tasks such as discovering programming patterns, mining API specifications, etc., perform source code analysis over control flow graphs (CFGs) at scale. Analyzing millions of CFGs can be expensive and performance of the analysis heavily depends on the underlying CFG traversal strategy. State-of-the-art analysis frameworks use a fixed traversal strategy. We argue that a single traversal strategy does not fit all kinds of analyses and CFGs and propose bespoke control flow analysis (BCFA). Given a control flow analysis (CFA) and a large number of CFGs, BCFA selects the most efficient traversal strategy for each CFG. BCFA extracts a set of properties of the CFA by analyzing the code of the CFA and combines it with properties of the CFG, such as branching factor and cyclicity, for selecting the optimal traversal strategy. We have implemented BCFA in Boa, and evaluated BCFA using a set of representative static analyses that mainly involve traversing CFGs and two large datasets containing 287 thousand and 162 million CFGs. Our results show that BCFA can speedup the large scale analyses by 1%-28%. Further, BCFA has low overheads; less than 0.2%, and low misprediction rate; less than 0.01%. Ramanathan Ramu, Ganesha Upadhyaya, Hoan Anh Nguyen, Hridesh Rajan |
ICSE | 2 |
| 2018 | Are code examples on an online Q&A forum reliable?: a study of API misuse on stack overflowabstractProgrammers often consult an online Q&A forum such as Stack Overflow to learn new APIs. This paper presents an empirical study on the prevalence and severity of API misuse on Stack Overflow. To reduce manual assessment effort, we design ExampleCheck, an API usage mining framework that extracts patterns from over 380K Java repositories on GitHub and subsequently reports potential API usage violations in Stack Overflow posts. We analyze 217,818 Stack Overflow posts using ExampleCheck and find that 31% may have potential API usage violations that could produce unexpected behavior such as program crashes and resource leaks. Such API misuse is caused by three main reasons---missing control constructs, missing or incorrect order of API calls, and incorrect guard conditions. Even the posts that are accepted as correct answers or upvoted by other programmers are not necessarily more reliable than other posts in terms of API misuse. This study result calls for a new approach to augment Stack Overflow with alternative API usage details that are not typically shown in curated examples. Tianyi Zhang 0001, Ganesha Upadhyaya, Anastasia Schaadhardt, Hridesh Rajan, Miryung Kim |
ICSE | 2 |
| 2018 | Collective program analysisabstractPopularity of data-driven software engineering has led to an increasing demand on the infrastructures to support efficient execution of tasks that require deeper source code analysis. While task optimization and parallelization are the adopted solutions, other research directions are less explored. We present collective program analysis (CPA), a technique for scaling large scale source code analyses, especially those that make use of control and data flow analysis, by leveraging analysis specific similarity. Analysis specific similarity is about, whether two or more programs can be considered similar for a given analysis. The key idea of collective program analysis is to cluster programs based on analysis specific similarity, such that running the analysis on one candidate in each cluster is sufficient to produce the result for others. For determining analysis specific similarity and clustering analysis-equivalent programs, we use a sparse representation and a canonical labeling scheme. Our evaluation shows that for a variety of source code analyses on a large dataset of programs, substantial reduction in the analysis time can be achieved; on average a 69% reduction when compared to a baseline and on average a 36% reduction when compared to a prior technique. We also found that a large amount of analysis-equivalent programs exists in large datasets. Ganesha Upadhyaya, Hridesh Rajan |
ICSE | 1 |
| 2018 | On Accelerating Source Code Analysis at Massive ScaleabstractEncouraged by the success of data-driven software engineering (SE) techniques that have found numerous applications e.g., in defect prediction, specification inference, the demand for mining and analyzing source code repositories at scale has significantly increased. However, analyzing source code at scale remains expensive to the extent that data-driven solutions to certain SE problems are beyond our reach today. Extant techniques have focused on leveraging distributed computing to solve this problem, but with a concomitant increase in computational resource needs. This work proposes a technique that reduces the amount of computation performed by the ultra-large-scale source code mining task, especially those that make use of control and data flow analyses. Our key idea is to analyze the mining task to identify and remove the irrelevant portions of the source code, prior to running the mining task. We show a realization of our insight for mining and analyzing massive collections of control flow graphs of source codes. Our evaluation using 16 classical control-/data-flow analyses that are typical components of mining tasks and 7 Million CFGs shows that our technique can achieve on average a 40 percent reduction in the task computation time. Our case studies demonstrates the applicability of our technique to massive scale source code mining tasks. Ganesha Upadhyaya, Hridesh Rajan |
IEEE Trans. Software Eng. | 1 |
| 2017 | Candoia: a platform for building and sharing mining software repositories tools as appsabstractWe propose Candoia, a novel platform and ecosystemfor building and sharing Mining Software Repositories(MSR) tools. Using Candoia, MSR tools are built as apps, and Candoia ecosystem, acting as an appstore, allows effective sharing. Candoia platform provides, data extraction tools for curating custom datasets for user projects, and data abstractions for enabling uniform access to MSR artifacts from disparate sources, which makes apps portable and adoptable across diverse software project settings of MSR researchers and practitioners. The structured design of a Candoia app and the languages selected for building various components of a Candoia app promotes easy customization. To evaluate Candoia we have built over two dozen MSR apps for analyzing bugs, software evolution, project management aspects, and source code and programming practices showing the applicability of the platform for buildinga variety of MSR apps. For testing portability of apps acrossdiverse project settings, we tested the apps using ten popularproject repositories, such as Apache Tomcat, JUnit, Node.js, etc, and found that apps required no changes to be portable. We performed a user study to test customizability and we found that five of eight Candoia users found it very easy to customize an existing app. Candoia is available for download. Nitin M. Tiwari, Ganesha Upadhyaya, Hoan Anh Nguyen, Hridesh Rajan |
MSR | 2 |
| 2015 | Effectively mapping linguistic abstractions for message-passing concurrency to threads on the Java virtual machineabstractEfficient mapping of message passing concurrency (MPC) abstractions to Java Virtual Machine (JVM) threads is critical for performance, scalability, and CPU utilization; but tedious and time consuming to perform manually. In general, this mapping cannot be found in polynomial time, but we show that by exploiting the local characteristics of MPC abstractions and their communication patterns this mapping can be determined effectively. We describe our MPC abstraction to thread mapping technique, its realization in two frame- works (Panini and Akka), and its rigorous evaluation using several benchmarks from representative MPC frameworks. We also compare our technique against four default mapping techniques: thread-all, round-robin-task-all, random-task-all and work-stealing. Our evaluation shows that our mapping technique can improve the performance by 30%-60% over default mapping techniques. These improvements are due to a number of challenges addressed by our technique namely: i) balancing the computations across JVM threads, ii) reducing the communication overheads, iii) utilizing information about cache locality, and iv) mapping MPC abstractions to threads in a way that reduces the contention between JVM threads. Ganesha Upadhyaya, Hridesh Rajan |
OOPSLA | 1 |