VLDB 2026 Research / reviewers in the wild / expert
Rahul Agarwal
dblp:08/3103
· DBLP profile ↗
12ranked-venue papers
9as first author
2since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 56% Vision and language · 44% | |
| Software engineering, system software, and programming languages
2 papers |
Program analysis · 49% Concurrent programming · 30% Programming languages and type systems · 18% | |
| Theoretical computer science
1 paper |
Information theory · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › video captioning
sports commentary generation |
0.8 | 1 | 2024 | Large Scale Generative AI Text Applied to Sports and Music · KDD 2024 |
Natural language and speech › Language models and text generation
text generation |
0.8 | 1 | 2024 | Large Scale Generative AI Text Applied to Sports and Music · KDD 2024 |
Information theory › estimation theory
density estimation |
0.3 | 1 | 2017 | A Novel Nonparametric Maximum Likelihood Estimator for Probability Density Functions · IEEE Trans. Pattern Anal. Mach. Intell. 2017 |
Natural language and speech › Language models and text generation › controllable text generation
personalized text generation |
0.2 | 1 | 2024 | Large Scale Generative AI Text Applied to Sports and Music · KDD 2024 |
Program analysis
type-based analysis |
0.1 | 2 | 2005 | Automated type-based analysis of data races and atomicity · PPoPP 2005 Optimized run-time race detection and atomicity checking using partial discovered types · ASE 2005 |
Bioinformatics and computational biology
neuroscience |
0.1 | 1 | 2017 | A Novel Nonparametric Maximum Likelihood Estimator for Probability Density Functions · IEEE Trans. Pattern Anal. Mach. Intell. 2017 |
Concurrent programming
concurrency bugs |
0.1 | 2 | 2005 | Optimized run-time race detection and atomicity checking using partial discovered types · ASE 2005 Automated type-based analysis of data races and atomicity · PPoPP 2005 |
Program analysis › concurrent program analysis
atomicity analysis |
0.1 | 1 | 2005 | Automated type-based analysis of data races and atomicity · PPoPP 2005 |
Programming languages and type systems › type systems › behavioral type systems
atomicity type system |
0.1 | 1 | 2005 | Automated type-based analysis of data races and atomicity · PPoPP 2005 |
Concurrent programming › concurrency bug detection
atomicity violation detection |
0.1 | 1 | 2005 | Optimized run-time race detection and atomicity checking using partial discovered types · ASE 2005 |
Program analysis › concurrent program analysis
data race analysis |
0.1 | 1 | 2005 | Automated type-based analysis of data races and atomicity · PPoPP 2005 |
Concurrent programming › concurrency bug detection
data race detection |
0.1 | 1 | 2005 | Optimized run-time race detection and atomicity checking using partial discovered types · ASE 2005 |
Program analysis
static analysis |
0.1 | 1 | 2005 | Optimized run-time race detection and atomicity checking using partial discovered types · ASE 2005 |
Programming languages and type systems
type systems |
0.1 | 1 | 2005 | Automated type-based analysis of data races and atomicity · PPoPP 2005 |
Program analysis
dynamic analysis |
0.0 | 1 | 2005 | Optimized run-time race detection and atomicity checking using partial discovered types · ASE 2005 |
Program verification › dynamic verification
runtime verification |
0.0 | 1 | 2005 | Optimized run-time race detection and atomicity checking using partial discovered types · ASE 2005 |
Methods — techniques the papers use, named apart from their topics
multimodal data fusion · 0.8generative AI · 0.8kernel density estimation · 0.6band-limited maximum likelihood · 0.6type inference · 0.1type discovery algorithm · 0.1type discovery · 0.1static interprocedural analysis · 0.1static analysis · 0.1runtime checking · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Finding Interest Needle in Popularity Haystack: Improving Retrieval by Modeling Item ExposureabstractRecommender systems operate in closed feedback loops, where user interactions reinforce popularity bias, leading to over-recommendation of already popular items while under-exposing niche or novel content. Existing bias mitigation methods, such as Inverse Propensity Scoring (IPS) and Off-Policy Correction (OPC), primarily operate at the ranking stage or during training, lacking explicit real-time control over exposure dynamics. In this work, we introduce an exposure-aware retrieval scoring approach, which explicitly models item exposure probability and adjusts retrieval-stage ranking at inference time. Unlike prior work, this method decouples exposure effects from engagement likelihood, enabling controlled trade-offs between fairness and engagement in large-scale recommendation platforms. We validate our approach through online A/B experiments in a real-world video recommendation system, demonstrating a 25% increase in uniquely retrieved items and a 40% reduction in the dominance of over-popular content, all while maintaining overall user engagement levels. Our results establish a scalable, deployable solution for mitigating popularity bias at the retrieval stage, offering a new paradigm for bias-aware personalization. Rahul Agarwal, Amit Jaspal, Omkar Vichare |
UMAP | 1 |
| 2024 | Large Scale Generative AI Text Applied to Sports and MusicabstractWe address the problem of scaling up the production of media content, including commentary and personalized news stories, for large-scale sports and music events worldwide. Our approach relies on generative AI models to transform a large volume of multimodal data (e.g., videos, articles, real-time scoring feeds, statistics, and fact sheets) into coherent and fluent text. Based on this approach, we introduce, for the first time, an AI commentary system, which was deployed to produce automated narrations for highlight packages at the 2023 US Open, Wimbledon, and Masters tournaments. In the same vein, our solution was extended to create personalized content for ESPN Fantasy Football and stories about music artists for the GRAMMY awards. These applications were built using a common software architecture achieved a 15x speed improvement with an average Rouge-L of 82.00 and perplexity of 6.6. Our work was successfully deployed at the aforementioned events, supporting 90 million fans around the world with 8 billion page views, continuously pushing the bounds on what is possible at the intersection of sports, entertainment, and AI. Aaron K. Baughman, Eduardo Morales, Rahul Agarwal, Gozde Akay, Rogério Feris, Tony Johnson, Stephen Hammer, Leonid Karlinsky |
KDD | 3 |
| 2017 | A Novel Nonparametric Maximum Likelihood Estimator for Probability Density FunctionsabstractParametric maximum likelihood (ML) estimators of probability density functions (pdfs) are widely used today because they are efficient to compute and have several nice properties such as consistency, fast convergence rates, and asymptotic normality. However, data is often complex making parametrization of the pdf difficult, and nonparametric estimation is required. Popular nonparametric methods, such as kernel density estimation (KDE), produce consistent estimators but are not ML and have slower convergence rates than parametric ML estimators. Further, these nonparametric methods do not share the other desirable properties of parametric ML estimators. This paper introduces a nonparametric ML estimator that assumes that the square-root of the underlying pdf is band-limited (BL) and hence "smooth". The BLML estimator is computed and shown to be consistent. Although convergence rates are not theoretically derived, the BLML estimator exhibits faster convergence rates than state-of-the-art nonparametric methods in simulations. Further, algorithms to compute the BLML estimator with lesser computational complexity than that of KDE methods are presented. The efficacy of the BLML estimator is shown by applying it to (i) density tail estimation and (ii) density estimation of complex neuronal receptive fields where it outperforms state-of-the-art methods used in neuroscience. Rahul Agarwal, Zhe Chen 0001, Sridevi V. Sarma |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | A Novel Nonparametric Approach for Neural Encoding and Decoding Models of Multimodal Receptive FieldsabstractPyramidal neurons recorded from the rat hippocampus and entorhinal cortex, such as place and grid cells, have diverse receptive fields, which are either unimodal or multimodal. Spiking activity from these cells encodes information about the spatial position of a freely foraging rat. At fine timescales, a neuron's spike activity also depends significantly on its own spike history. However, due to limitations of current parametric modeling approaches, it remains a challenge to estimate complex, multimodal neuronal receptive fields while incorporating spike history dependence. Furthermore, efforts to decode the rat's trajectory in one- or two-dimensional space from hippocampal ensemble spiking activity have mainly focused on spike history-independent neuronal encoding models. In this letter, we address these two important issues by extending a recently introduced nonparametric neural encoding framework that allows modeling both complex spatial receptive fields and spike history dependencies. Using this extended nonparametric approach, we develop novel algorithms for decoding a rat's trajectory based on recordings of hippocampal place cells and entorhinal grid cells. Results show that both encoding and decoding models derived from our new method performed significantly better than state-of-the-art encoding and decoding models on 6 minutes of test data. In addition, our model's performance remains invariant to the apparent modality of the neuron's receptive field. Rahul Agarwal, Zhe Chen 0001, Fabian Kloosterman, Matthew A. Wilson, Sridevi V. Sarma |
Neural Comput. | 1 |
| 2013 | An Automatic Approach to Treebank Error Detection Using a Dependency Parser
Bhasha Agrawal, Rahul Agarwal, Samar Husain, Dipti Misra Sharma |
CICLing (1) | 2 |
| 2012 | Using Linear Systems Theory to Study Nonlinear Dynamics of Relay Cells
Rahul Agarwal, Sridevi V. Sarma |
ICINCO (1) | 1 |
| 2012 | A GUI to Detect and Correct Errors in Hindi Dependency Treebank
Rahul Agarwal, Bharat Ram Ambati, Anil Kumar Singh 0001 |
LREC | 1 |
| 2012 | Performance Limitations of Relay NeuronsabstractRelay cells are prevalent throughout sensory systems and receive two types of inputs: driving and modulating. The driving input contains receptive field properties that must be transmitted while the modulating input alters the specifics of transmission. For example, the visual thalamus contains relay neurons that receive driving inputs from the retina that encode a visual image, and modulating inputs from reticular activating system and layer 6 of visual cortex that control what aspects of the image will be relayed back to visual cortex for perception. What gets relayed depends on several factors such as attentional demands and a subject's goals. In this paper, we analyze a biophysical based model of a relay cell and use systems theoretic tools to construct analytic bounds on how well the cell transmits a driving input as a function of the neuron's electrophysiological properties, the modulating input, and the driving signal parameters. We assume that the modulating input belongs to a class of sinusoidal signals and that the driving input is an irregular train of pulses with inter-pulse intervals obeying an exponential distribution. Our analysis applies to any [Formula: see text] order model as long as the neuron does not spike without a driving input pulse and exhibits a refractory period. Our bounds on relay reliability contain performance obtained through simulation of a second and third order model, and suggest, for instance, that if the frequency of the modulating input increases or the DC offset decreases, then relay increases. Our analysis also shows, for the first time, how the biophysical properties of the neuron (e.g. ion channel dynamics) define the oscillatory patterns needed in the modulating input for appropriately timed relay of sensory information. In our discussion, we describe how our bounds predict experimentally observed neural activity in the basal ganglia in (i) health, (ii) in Parkinson's disease (PD), and (iii) in PD during therapeutic deep brain stimulation. Our bounds also predict different rhythms that emerge in the lateral geniculate nucleus in the thalamus during different attentional states. Rahul Agarwal, Sridevi V. Sarma |
PLoS Comput. Biol. | 1 |
| 2006 | Designing an adaptive learning module to teach software testingabstractAdaptive learning systems aim to precisely tailor education and training to the individual needs of learners. Such systems use an internal model of a user's current knowledge to adjust the navigational affordances and presentation order of material. The user model is incrementally built and updated as the user demonstrates mastery by completing exercises and tests. Designing courses that are delivered adaptively involves addressing many complexities. This paper describes experiences designing the first adaptive module in a series intended to teach software testing skills. Experiences in using the first module and a preliminary evaluation of its effectiveness are presented. Rahul Agarwal, Stephen H. Edwards, Manuel A. Pérez-Quiñones |
SIGCSE | 1 |
| 2005 | Optimized run-time race detection and atomicity checking using partial discovered typesabstractConcurrent programs are notorious for containing errors that are difficult to reproduce and diagnose. Two common kinds of concurrency errors are data races and atomicity violations (informally, atomicity means that executing methods concurrently is equivalent to executing them serially). Several static and dynamic (run-time) analysis techniques exist to detect potential races and atomicity violations. Run-time checking may miss errors in unexecuted code and incurs significant run-time overhead. On the other hand, run-time checking generally produces fewer false alarms than static analysis; this is a significant practical advantage, since diagnosing all of the warnings from static analysis of large codebases may be prohibitively expensive. This paper explores the use of static analysis to significantly decrease the overhead of run-time checking. Our approach is based on a type system for analyzing data races and atomicity. A type discovery algorithm is used to obtain types for as much of the program as possible (complete type inference for this type system is NP-hard, and parts of the program might be untypable). Warnings from the typechecker are used to identify parts of the program from which run-time checking can safely be omitted. The approach is completely automatic, scalable to very large programs, and significantly reduces the overhead of run-time checking for data races and atomicity violations. 1. Rahul Agarwal, Amit Sasturkar, Scott D. Stoller |
ASE | 1 |
| 2005 | Automated type-based analysis of data races and atomicityabstractConcurrent programs are notorious for containing errors that are difficult to reproduce and diagnose at run-time. This motivated the development of type systems that statically ensure the absence of some common kinds of concurrent programming errors including data races and atomicity violations. A method is atomic if every execution of the concurrent program is equivalent to an execution in which the atomic method is executed without being interleaved with other concurrently executed methods. Atomicity is a common correctness requirement in concurrent programs; atomicity violations may indicate incorrect synchronization. This paper presents Extended Parameterized Atomic Java (EPAJ), a type system for specifying and verifying atomicity in Java programs. EPAJ combines Flanagan and Qadeer's atomicity types [11] with a new and significantly more expressive type system for analyzing data races, called Extended Parameterized Race-Free Java (EPRFJ), allowing a more accurate analysis of atomicity. The paper also presents a type discovery algorithm to automatically obtain EPRFJ types, and a static interprocedural type inference algorithm that, given EPRFJ types, infers atomicity types. These algorithms can be incorporated into testing and debugging tools, benefiting users who know nothing about type systems. We report our experience with a prototype implementation. Amit Sasturkar, Rahul Agarwal, Scott D. Stoller |
PPoPP | 2 |
| 2004 | Type Inference for Parameterized Race-Free Java
Rahul Agarwal, Scott D. Stoller |
VMCAI | 1 |