VLDB 2026 Research / reviewers in the wild / expert
Jaroslav M. Fowkes
dblp:132/4642
· DBLP profile ↗
7ranked-venue papers
6as first author
1since 2021 · last 2025
0000-0002-8048-4572ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Approximating Large-Scale Hessian Matrices Using Secant EquationsabstractLarge-scale optimization algorithms frequently require sparse Hessian matrices that are not readily available. Existing methods for approximating large sparse Hessian matrices either do not impose sparsity or are computationally prohibitive. To try and overcome these limitations, we propose a novel approach that seeks to satisfy as many componentwise secant equations as necessary to define each row of the Hessian matrix. A naive application of this approach is too expensive for Hessian matrices that have some relatively dense rows but, by carefully taking into account the symmetry and connectivity of the Hessian matrix, we are able devise an approximation algorithm that is fast and efficient with scope for parallelism. Example sparse Hessian matrices from the CUTEst test collection for optimization illustrate the effectiveness and robustness of our proposed method. Jaroslav M. Fowkes, Nicholas I. M. Gould, Jennifer A. Scott |
ACM Trans. Math. Softw. | 1 |
| 2017 | Autofolding for Source Code SummarizationabstractDevelopers spend much of their time reading and browsing source code, raising new opportunities for summarization methods. Indeed, modern code editors provide code folding, which allows one to selectively hide blocks of code. However this is impractical to use as folding decisions must be made manually or based on simple rules. We introduce the autofolding problem, which is to automatically create a code summary by folding less informative code regions. We present a novel solution by formulating the problem as a sequence of AST folding decisions, leveraging a scoped topic model for code tokens. On an annotated set of popular open source projects, we show that our summarizer outperforms simpler baselines, yielding a 28 percent error reduction. Furthermore, we find through a case study that our summarizer is strongly preferred by experienced developers. More broadly, we hope this work will aid program comprehension by turning code folding into a usable and valuable tool. Jaroslav M. Fowkes, Pankajan Chanthirasegaran, Razvan Ranca, Miltiadis Allamanis, Mirella Lapata, Charles Sutton |
IEEE Trans. Software Eng. | 1 |
| 2016 | A Subsequence Interleaving Model for Sequential Pattern MiningabstractRecent sequential pattern mining methods have used the minimum description length (MDL) principle to define an encoding scheme which describes an algorithm for mining the most compressing patterns in a database. We present a novel subsequence interleaving model based on a probabilistic model of the sequence database, which allows us to search for the most compressing set of patterns without designing a specific encoding scheme. Our proposed algorithm is able to efficiently mine the most relevant sequential patterns and rank them using an associated measure of interestingness. The efficient inference in our model is a direct result of our use of a structural expectation-maximization framework, in which the expectation-step takes the form of a submodular optimization problem subject to a coverage constraint. We show on both synthetic and real world datasets that our model mines a set of sequential patterns with low spuriousness and redundancy, high interpretability and usefulness in real-world applications. Furthermore, we demonstrate that the quality of the patterns from our approach is comparable to, if not better than, existing state of the art sequential pattern mining algorithms. Jaroslav M. Fowkes, Charles Sutton |
KDD | 1 |
| 2016 | A Bayesian Network Model for Interesting Itemsets
Jaroslav M. Fowkes, Charles Sutton |
ECML/PKDD (2) | 1 |
| 2016 | Parameter-free probabilistic API mining across GitHubabstractExisting API mining algorithms can be difficult to use as they require expensive parameter tuning and the returned set of API calls can be large, highly redundant and difficult to understand. To address this, we present PAM (Probabilistic API Miner), a near parameter-free probabilistic algorithm for mining the most interesting API call patterns. We show that PAM significantly outperforms both MAPO and UPMiner, achieving 69% test-set precision, at retrieving relevant API call sequences from GitHub. Moreover, we focus on libraries for which the developers have explicitly provided code examples, yielding over 300,000 LOC of hand-written API example code from the 967 client projects in the data set. This evaluation suggests that the hand-written examples actually have limited coverage of real API usages. Jaroslav M. Fowkes, Charles Sutton |
SIGSOFT FSE | 1 |
| 2015 | Branching and bounding improvements for global optimization algorithms with Lipschitz continuity properties
Coralia Cartis, Jaroslav M. Fowkes, Nicholas I. M. Gould |
J. Glob. Optim. | 2 |
| 2013 | A branch and bound algorithm for the global optimization of Hessian Lipschitz continuous functions
Jaroslav M. Fowkes, Nicholas I. M. Gould, Chris L. Farmer |
J. Glob. Optim. | 1 |