EDBT 2026 Demo / reviewers in the wild / expert
Vaibhav Aggarwal
dblp:270/3883
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Representation and self-supervised learning · 32% Image recognition and object detection · 24% Efficient and distributed learning · 12% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 50% Debugging and program repair · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Reconfigurable computing and FPGAs · 77% Memory systems · 23% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder |
1.4 | 2 | 2024 | R-MAE: Regions Meet Masked Autoencoders · ICLR 2024 The effectiveness of MAE pre-pretraining for billion-scale pretraining · ICCV 2023 |
Debugging and program repair
automated program repair |
0.9 | 1 | 2025 | DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree Traversal · ACL (1) 2025 |
Program synthesis and code generation
code agent |
0.9 | 1 | 2025 | DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree Traversal · ACL (1) 2025 |
Computer vision › Segmentation and scene understanding
interactive segmentation |
0.8 | 1 | 2024 | R-MAE: Regions Meet Masked Autoencoders · ICLR 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | MobileNetV4: Universal Models for the Mobile Ecosystem · ECCV (40) 2024 |
Machine learning › Representation and self-supervised learning › pre-training
foundation model pretraining |
0.7 | 1 | 2023 | The effectiveness of MAE pre-pretraining for billion-scale pretraining · ICCV 2023 |
Machine learning › Deep learning architectures and training › transformer › vision transformer
hierarchical vision transformer |
0.7 | 1 | 2023 | Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles · ICML 2023 |
Computer vision › Image recognition and object detection
image classification |
0.7 | 1 | 2023 | Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles · ICML 2023 |
Computer vision › Video understanding and tracking
video classification |
0.7 | 1 | 2023 | Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles · ICML 2023 |
Computer vision › Image recognition and object detection
visual recognition |
0.7 | 1 | 2023 | The effectiveness of MAE pre-pretraining for billion-scale pretraining · ICCV 2023 |
Reconfigurable computing and FPGAs › FPGA accelerator
FPGA-based sorting |
0.4 | 1 | 2020 | Bonsai: High-Performance Adaptive Merge Tree Sorting · ISCA 2020 |
Computer vision › Image recognition and object detection › efficient visual recognition
mobile image recognition |
0.2 | 1 | 2024 | MobileNetV4: Universal Models for the Mobile Ecosystem · ECCV (40) 2024 |
Memory systems › memory hierarchy
memory hierarchy optimization |
0.1 | 1 | 2020 | Bonsai: High-Performance Adaptive Merge Tree Sorting · ISCA 2020 |
Methods — techniques the papers use, named apart from their topics
tree traversal · 0.9large language model · 0.9dynamic action re-sampling · 0.9neural architecture search · 0.8masked autoencoder · 0.8vision transformer · 0.7self-supervised pretraining · 0.7masked autoencoding · 0.7masked autoencoder pretraining · 0.7merge tree architecture · 0.4adaptive sorting · 0.4FPGA programmability · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree TraversalabstractLarge Language Models (LLMs) have revolutionized various domains, including natural language processing, data analysis, and software development, by enabling automation.In software engineering, LLM-powered coding agents have garnered significant attention due to their potential to automate complex development tasks, assist in debugging, and enhance productivity.However, existing approaches often struggle with sub-optimal decision-making, requiring either extensive manual intervention or inefficient compute scaling strategies.To improve coding agent performance, we present Dynamic Action Re-Sampling (DARS), a novel inference time compute scaling approach for coding agents, that is faster and more effective at recovering from sub-optimal decisions compared to baselines.While traditional agents either follow linear trajectories or rely on random sampling for scaling compute, our approach DARS works by branching out a trajectory at certain key decision points by taking an alternative action given the history of the trajectory and execution feedback of the previous attempt from that point.We evaluate our approach on SWE-Bench Lite benchmark, demonstrating that this scaling strategy achieves a pass@k score of 55% with Claude 3.5 Sonnet V2.Our framework achieves a pass@1 rate of 47%, outperforming state-of-the-art (SOTA) opensource frameworks. 1 Vaibhav Aggarwal, Ojasv Kamal, Abhinav Japesh, Zhijing Jin 0001, Bernhard Schölkopf |
ACL (1) | 1 |
| 2025 | Adopting Whisper for Confidence EstimationabstractRecent research on word-level confidence estimation for speech recognition systems has primarily focused on lightweight models known as Confidence Estimation Modules (CEMs), which rely on hand-engineered features derived from Automatic Speech Recognition (ASR) outputs. In contrast, we propose a novel end-to-end approach that leverages the ASR model itself (Whisper) to generate word-level confidence scores. Specifically, we introduce a method in which the Whisper model is fine-tuned to produce scalar confidence scores given an audio input and its corresponding hypothesis transcript. Our experiments demonstrate that the fine-tuned Whisper-tiny model, comparable in size to a strong CEM baseline, achieves similar performance on the in-domain dataset and surpasses the CEM baseline on eight out-of-domain datasets, whereas the fine-tuned Whisper-large model consistently outperforms the CEM baseline by a substantial margin across all the datasets. Vaibhav Aggarwal, Shabari S. Nair, Yash Verma, Yash Jogi |
ICASSP | 1 |
| 2024 | MobileNetV4: Universal Models for the Mobile Ecosystem
Danfeng Qin, Chas Leichner, Manolis Delakis, Marco Fornoni, Shixin Luo, Colby R. Banbury, Chengxi Ye, Berkin Akin, Vaibhav Aggarwal, Tenghui Zhu, Daniele Moro, Andrew G. Howard |
ECCV (40) | 11 |
| 2024 | R-MAE: Regions Meet Masked AutoencodersabstractIn this work, we explore regions as a potential visual analogue of words for self-supervised image representation learning. Inspired by Masked Autoencoding (MAE), a generative pre-training baseline, we propose masked region autoencoding to learn from groups of pixels or regions. Specifically, we design an architecture which efficiently addresses the one-to-many mapping between images and regions, while being highly effective especially with high-quality regions. When integrated with MAE, our approach (R-MAE) demonstrates consistent improvements across various pre-training datasets and downstream detection and segmentation benchmarks, with negligible computational overheads. Beyond the quantitative evaluation, our analysis indicates the models pre-trained with masked region autoencoding unlock the potential for interactive segmentation. The code is provided at https://github.com/facebookresearch/r-mae. Duy-Kien Nguyen, Yanghao Li, Vaibhav Aggarwal, Martin R. Oswald, Alexander Kirillov, Cees Snoek, Xinlei Chen |
ICLR | 3 |
| 2024 | Improving Rare-Word Recognition of Whisper in Zero-Shot SettingsabstractWhisper, despite being trained on 680 K hours of web-scaled audio data, faces difficulty in recognizing rare words like domain-specific terms, with a solution being contextual biasing through prompting. To improve upon this method, in this paper, we propose a supervised learning strategy to fine-tune Whisper for contextual biasing instruction. We demonstrate that by using only 670 hours of Common Voice English set for fine-tuning, our model generalizes to 11 diverse opensource English datasets, achieving a 45.6% improvement in recognition of rare words and 60.8% improvement in recognition of words unseen during fine-tuning over the baseline method. Surprisingly, our model’s contextual biasing ability generalizes even to languages unseen during fine-tuning. Yash Jogi, Vaibhav Aggarwal, Shabari S. Nair, Yash Verma, Aayush Kubba |
SLT | 2 |
| 2023 | The effectiveness of MAE pre-pretraining for billion-scale pretrainingabstractThis paper revisits the standard pretrain-then-finetune paradigm used in computer vision for visual recognition tasks. Typically, state-of-the-art foundation models are pretrained using large scale (weakly) supervised datasets with billions of images. We introduce an additional pre-pretraining stage that is simple and uses the self-supervised MAE technique to initialize the model. While MAE has only been shown to scale with the size of models, we find that it scales with the size of the training dataset as well. Thus, our MAE-based pre-pretraining scales with both model and data size making it applicable for training foundation models. Pre-pretraining consistently improves both the model convergence and the downstream transfer performance across a range of model scales (millions to billions of parameters), and dataset sizes (millions to billions of images). We measure the effectiveness of pre-pretraining on 10 different visual recognition tasks spanning image classification, video recognition, object detection, low-shot classification and zero-shot recognition. Our largest model achieves new state-of-the-art results on iNaturalist-18 (91.3%), 1-shot ImageNet-1k (62.1%), and zero-shot transfer on Food-101 (96.2%). Our study reveals that model initialization plays a significant role, even for web-scale pretraining with billions of images. Mannat Singh, Quentin Duval, Kalyan Vasudev Alwala, Haoqi Fan 0001, Vaibhav Aggarwal, Aaron Adcock, Armand Joulin, Piotr Dollár, Christoph Feichtenhofer, Ross B. Girshick, Rohit Girdhar, Ishan Misra |
ICCV | 5 |
| 2023 | Hiera: A Hierarchical Vision Transformer without the Bells-and-WhistlesabstractModern hierarchical vision transformers have added several vision-specific components in the pursuit of supervised classification performance. While these components lead to effective accuracies and attractive FLOP counts, the added complexity actually makes these transformers slower than their vanilla ViT counterparts. In this paper, we argue that this additional bulk is unnecessary. By pretraining with a strong visual pretext task (MAE), we can strip out all the bells-and-whistles from a state-of-the-art multi-stage vision transformer without losing accuracy. In the process, we create Hiera, an extremely simple hierarchical vision transformer that is more accurate than previous models while being significantly faster both at inference and during training. We evaluate Hiera on a variety of tasks for image and video recognition. Our code and models are available at https://github.com/facebookresearch/hiera. Chaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei 0005, Haoqi Fan 0001, Po-Yao Huang 0001, Vaibhav Aggarwal, Arkabandhu Chowdhury, Omid Poursaeed, Judy Hoffman, Jitendra Malik, Yanghao Li, Christoph Feichtenhofer |
ICML | 7 |
| 2020 | Bonsai: High-Performance Adaptive Merge Tree SortingabstractSorting is a key computational kernel in many big data applications. Most sorting implementations focus on a specific input size, record width, and hardware configuration. This has created a wide array of sorters that are optimized only to a narrow application domain.In this work we show that merge trees can be implemented on FPGAs to offer state-of-the-art performance over many problem sizes. We introduce a novel merge tree architecture and develop Bonsai, an adaptive sorting solution that takes into consideration the off-chip memory bandwidth and the amount of on-chip resources to optimize sorting time. FPGA programmability allows us to leverage Bonsai to quickly implement the optimal merge tree configuration for any problem size and memory hierarchy.Using Bonsai, we develop a state-of-the-art sorter which specifically targets DRAM-scale sorting on AWS EC2 F1 instances. For 4-32 GB array size, our implementation has a minimum of 2.3x, 1.3x, 1.2x and up to 2.5x, 3.7x, 1.3x speedup over the best designs on CPUs, FPGAs, and GPUs, respectively. Our design exhibits 3.3x better bandwidth-efficiency compared to the best previous sorting implementations. Finally, we demonstrate that Bonsai can tune our design over a wide range of problem sizes(megabyte to terabyte) and memory hierarchies including DDR DRAMs, high-bandwidth memories (HBMs) and solid-state disks (SSDs). Nikola Samardzic, Weikang Qiao, Vaibhav Aggarwal, Mau-Chung Frank Chang, Jason Cong |
ISCA | 3 |