VLDB 2026 Research / reviewers in the wild / expert
Tom Jacobs
dblp:186/4261
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Efficient and distributed learning · 43% Learning theory · 23% Optimization for machine learning · 16% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
2.6 | 3 | 2025 | The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis · NeurIPS 2025 Sign-In to the Lottery: Reparameterizing Sparse Training · NeurIPS 2025 Mask in the Mirror: Implicit Sparsification · ICLR 2025 |
Machine learning › Optimization for machine learning
implicit regularization |
1.7 | 2 | 2025 | Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias? · ICML 2025 Mask in the Mirror: Implicit Sparsification · ICLR 2025 |
Machine learning › Learning theory › implicit bias
mirror flow |
1.7 | 2 | 2025 | Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias? · ICML 2025 Mask in the Mirror: Implicit Sparsification · ICLR 2025 |
Machine learning › Learning paradigms › continual learning › catastrophic forgetting
catastrophic forgetting mitigation |
0.9 | 1 | 2025 | Pay Attention to Small Weights · NeurIPS 2025 |
Machine learning › Optimization for machine learning › regularized optimization
explicit regularization |
0.9 | 1 | 2025 | Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias? · ICML 2025 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.9 | 1 | 2025 | Pay Attention to Small Weights · NeurIPS 2025 |
Machine learning › Learning theory
generalization |
0.9 | 1 | 2025 | Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias? · ICML 2025 |
Machine learning › Learning theory
implicit bias |
0.9 | 1 | 2025 | Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias? · ICML 2025 |
Machine learning › Efficient and distributed learning › model compression › sparsity
network sparsification |
0.9 | 1 | 2025 | Mask in the Mirror: Implicit Sparsification · ICLR 2025 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.9 | 1 | 2025 | Pay Attention to Small Weights · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.9 | 1 | 2025 | The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
selective parameter update |
0.9 | 1 | 2025 | Pay Attention to Small Weights · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression
sparse training |
0.9 | 1 | 2025 | Sign-In to the Lottery: Reparameterizing Sparse Training · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › regularization › norm-based regularization
weight decay |
0.9 | 1 | 2025 | Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias? · ICML 2025 |
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel |
0.3 | 1 | 2025 | The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
weight initialization |
0.3 | 1 | 2025 | Sign-In to the Lottery: Reparameterizing Sparse Training · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
sign flip · 0.9random matrix theory · 0.9mirror flow analysis · 0.9mirror flow · 0.9l2 regularization · 0.9l1 regularization · 0.9graphon analysis · 0.9graph limit theory · 0.9dynamic reparameterization · 0.9continuous sparsification · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mask in the Mirror: Implicit SparsificationabstractContinuous sparsification strategies are among the most effective methods for reducing the inference costs and memory demands of large-scale neural networks. A key factor in their success is the implicit $L_1$ regularization induced by jointly learning both mask and weight variables, which has been shown experimentally to outperform explicit $L_1$ regularization. We provide a theoretical explanation for this observation by analyzing the learning dynamics, revealing that early continuous sparsification is governed by an implicit $L_2$ regularization that gradually transitions to an $L_1$ penalty over time. Leveraging this insight, we propose a method to dynamically control the strength of this implicit bias. Through an extension of the mirror flow framework, we establish convergence and optimality guarantees in the context of underdetermined linear regression. Our theoretical findings may be of independent interest, as we demonstrate how to enter the rich regime and show that the implicit bias can be controlled via a time-dependent Bregman potential. To validate these insights, we introduce PILoT, a continuous sparsification approach with novel initialization and dynamic regularization, which consistently outperforms baselines in standard experiments. Tom Jacobs, Rebekka Burkholz |
ICLR | 1 |
| 2025 | Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?abstractImplicit bias plays an important role in explaining how overparameterized models generalize well. Explicit regularization like weight decay is often employed in addition to prevent overfitting. While both concepts have been studied separately, in practice, they often act in tandem. Understanding their interplay is key to controlling the shape and strength of implicit bias, as it can be modified by explicit regularization. To this end, we incorporate explicit regularization into the mirror flow framework and analyze its lasting effects on the geometry of the training dynamics, covering three distinct effects: positional bias, type of bias, and range shrinking. Our analytical approach encompasses a broad class of problems, including sparse coding, matrix sensing, single-layer attention, and LoRA, for which we demonstrate the utility of our insights. To exploit the lasting effect of regularization and highlight the potential benefit of dynamic weight decay schedules, we propose to switch off weight decay during training, which can improve generalization, as we demonstrate in experiments. Tom Jacobs, Rebekka Burkholz |
ICML | 1 |
| 2025 | Sign-In to the Lottery: Reparameterizing Sparse TrainingabstractThe performance gap between training sparse neural networks from scratch (PaI) and dense-to-sparse training presents a major roadblock for efficient deep learning. According to the Lottery Ticket Hypothesis, PaI hinges on finding a problem specific parameter initialization. As we show, to this end, determining correct parameter signs is sufficient. Yet, they remain elusive to PaI. To address this issue, we propose Sign-In, which employs a dynamic reparameterization that provably induces sign flips. Such sign flips are complementary to the ones that dense-to-sparse training can accomplish, rendering Sign-In as an orthogonal method. While our experiments and theory suggest performance improvements of PaI, they also carve out the main open challenge to close the gap between PaI and dense-to-sparse training. Advait Gadhikar, Tom Jacobs, Rebekka Burkholz |
NeurIPS | 2 |
| 2025 | The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width AnalysisabstractSparse neural networks promise efficiency, yet training them effectively remains a fundamental challenge. Despite advances in pruning methods that create sparse architectures, understanding why some sparse structures are better trainable than others with the same level of sparsity remains poorly understood. Aiming to develop a systematic approach to this fundamental problem, we propose a novel theoretical framework based on the theory of graph limits, particularly graphons, that characterizes sparse neural networks in the infinite-width regime. Our key insight is that connectivity patterns of sparse neural networks induced by pruning methods converge to specific graphons as networks' width tends to infinity, which encodes implicit structural biases of different pruning methods. We postulate the *Graphon Limit Hypothesis* and provide empirical evidence to support it. Leveraging this graphon representation, we derive a *Graphon Neural Tangent Kernel (Graphon NTK)* to study the training dynamics of sparse networks in the infinite width limit. Graphon NTK provides a general framework for the theoretical analysis of sparse networks. We empirically show that the spectral analysis of Graphon NTK correlates with observed training dynamics of sparse networks, explaining the varying convergence behaviours of different pruning methods. Our framework provides theoretical insights into the impact of connectivity patterns on the trainability of various sparse network architectures. The-Anh Ta, Tom Jacobs, Rebekka Burkholz, Long Tran-Thanh |
NeurIPS | 3 |
| 2025 | Pay Attention to Small WeightsabstractFinetuning large pretrained neural networks is known to be resource-intensive, both in terms of memory and computational cost. To mitigate this, a common approach is to restrict training to a subset of the model parameters. By analyzing the relationship between gradients and weights during finetuning, we observe a notable pattern: large gradients are often associated with small-magnitude weights. This correlation is more pronounced in fine-tuning settings than in training from scratch. Motivated by this observation, we propose \textsc{NanoAdam}, which dynamically updates only the small-magnitude weights during fine-tuning and offers several practical advantages: first, the criterion is \emph{gradient-free}—the parameter subset can be determined without gradient computation; second, it preserves large-magnitude weights, which are likely to encode critical features learned during pre-training, thereby reducing the risk of catastrophic forgetting; thirdly, it permits the use of larger learning rates and consistently leads to better generalization performance in experiments. We demonstrate this for both NLP and vision tasks. Tom Jacobs, Advait Gadhikar, Rebekka Burkholz |
NeurIPS | 2 |
| 2023 | MATWI: A Multimodal Automatic Tool Wear Inspection Dataset and Baseline Algorithms
Lars De Pauw, Tom Jacobs, Toon Goedemé |
ICVS | 2 |
| 2022 | ViSRE: A Unified Visual Analysis Dashboard for Proactive Cloud Outage ManagementabstractEfficient outage detection and remediation is crucial for effectively operating cloud computing systems. To remediate outages, system engineers must quickly identify the causal relationships between metrics and correlate events across multiple monitoring tools. In practice, this process largely remains reactive due to the complexity and general lack of interpretability within such monitoring environments. This work presents ViSRE: an integrated visual analytics system that integrates causal and predictive models with interactive visualizations to aid in proactive cloud outage management. We develop enhanced node representations for our causal graph representation to support system engineers in performing root cause analysis and reasoning about causality chains in multi-dimensional temporal data. We report the results of a quantitative assessment of the proposed predictive models, which show good performance guarantees. To evaluate and refine our system, we conduct a study with six cloud system engineers who verify that our proposed techniques can support proactive cloud maintenance by intuitively displaying temporal relationships between predicted and raw data. By correlating and presenting data from disparate sources, ViSRE also reduces context switching costs and reduces the time spent on manually correlating events during remediation of time-critical outages. Paula Kayongo, Jane Hoffswell, Shiv Kumar Saini, Shaddy Garg, Eunyee Koh, Tom Jacobs |
VISSOFT | 7 |
| 2014 | Development of iBsafe: A Collaborative, Theory-based Approach to Creating a Mobile Game Application for Child Safety
Cinnamon A. Dixon, Robert Ammerman, Judith W. Dexheimer, Heekyoung Jung, Boyd Johnson, Jennifer Elliott, Tom Jacobs, Wendy J. Pomerantz, E. Melinda Mahabee-Gittens |
AMIA | 8 |