VLDB 2026 Research / reviewers in the wild / expert
Shivam Aggarwal
dblp:155/4186
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-1748-9810ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Condensed Data Expansion Using Model Inversion for Knowledge DistillationabstractCondensed datasets offer a compact representation of larger datasets, but training models directly on them or using them to enhance model performance through knowledge distillation (KD) can result in suboptimal outcomes due to limited information. To address this, we propose a method that expands condensed datasets using model inversion, a technique for generating synthetic data based on the impressions of a pre-trained model on its training data. This approach is particularly well-suited for KD scenarios, as the teacher model is already pre-trained and retains knowledge of the original training data. By creating synthetic data that complements the condensed samples, we enrich the training set and better approximate the underlying data distribution, leading to improvements in student model accuracy during knowledge distillation. Our method demonstrates significant gains in KD accuracy compared to using condensed datasets alone and outperforms standard model inversion-based KD methods by up to 11.4% across various datasets and model architectures. Importantly, it remains effective even when using as few as one condensed sample per class, and can also enhance performance in few-shot scenarios where only limited real data samples are available. Kuluhan Binici, Shivam Aggarwal, Cihan Acar, Nam Trung Pham, Karianto Leman, Gim Hee Lee, Tulika Mitra |
AAAI | 2 |
| 2026 | HALO: Hardware-Aware Quantization with Low Critical-Path-Delay Weights for LLM AccelerationabstractQuantization is critical for efficiently deploying large language models (LLMs). Yet conventional methods remain hardware-agnostic, limited to bit-width constraints, and do not account for intrinsic circuit characteristics such as the timing behaviors and energy profiles of Multiply-Accumulate (MAC) units. This disconnect from circuit-level behavior limits the ability to exploit available timing margins and energy-saving opportunities, reducing the overall efficiency of deployment on modern accelerators. To address these limitations, we propose HALO, a versatile framework for Hardware-Aware Post-Training Quantization (PTQ). Unlike traditional methods, HALO explicitly incorporates detailed hardware characteristics, including critical-path timing and power consumption, into its quantization approach. HALO strategically selects weights with low critical-path-delays enabling higher operational frequencies and dynamic frequency scaling without disrupting the architecture's dataflow. Remarkably, HALO achieves these improvements with only a few dynamic voltage and frequency scaling (DVFS) adjustments, ensuring simplicity and practicality in deployment. Additionally, by reducing switching activity within the MAC units, HALO effectively lowers energy consumption. Evaluations on accelerators such as Tensor Processing Units (TPUs) and Graphics Processing Units (GPUs) demonstrate that HALO significantly enhances inference efficiency, achieving average performance improvements of 270% and energy savings of 51% over baseline quantization methods, all with minimal impact on accuracy. Rohan Juneja, Shivam Aggarwal, Safeen Huda, Tulika Mitra, Li-Shiuan Peh |
AAAI | 2 |
| 2025 | DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE InferenceabstractMixture-of-Experts (MoE) models, though highly effective for various machine learning tasks, face significant deployment challenges on memory-constrained devices. While GPUs offer fast inference, their limited memory compared to CPUs means not all experts can be stored on the GPU simultaneously, necessitating frequent, costly data transfers from CPU memory, often negating GPU speed advantages. To address this, we present DAOP, an on-device MoE inference engine to optimize parallel GPU-CPU execution. DAOP dynamically allocates experts between CPU and GPU based on per-sequence activation patterns, and selectively pre-calculates predicted experts on CPUs to minimize transfer latency. This approach enables efficient resource utilization across various expert cache ratios while maintaining model accuracy through a novel graceful degradation mechanism. Comprehensive evaluations across various datasets show that DAOP outperforms traditional expert caching and prefetching methods by up to 8.20x and offloading techniques by 1.35x while maintaining accuracy. Yujie Zhang 0007, Shivam Aggarwal, Tulika Mitra |
DATE | 2 |
| 2025 | Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion ModelsabstractText-to-image diffusion models often use a fixed number of denoising steps, balancing time costs and image quality. However, the optimal number of steps depends on the complexity of the input text prompt. We propose an adaptive diffusion controller that dynamically adjusts the number of steps to generate high-quality images efficiently, without additional model training. By leveraging a mixture of step schedules with varying step sizes and evaluating the error term discrepancy at each timestep, our method transitions between schedules to optimize performance. Experiments on COCO and DiffusionDB show that our approach reduces inference time while maintaining visual fidelity, offering a more efficient alternative for text-to-image diffusion models. Kuluhan Binici, Cihan Acar, Shivam Aggarwal, Tulika Mitra |
ICIP | 3 |
| 2024 | CRISP: Hybrid Structured Sparsity for Class-Aware Model PruningabstractMachine learning pipelines for classification tasks often train a universal model to achieve accuracy across a broad range of classes. However, a typical user encounters only a limited selection of classes regularly. This disparity provides an opportunity to enhance computational efficiency by tailoring models to focus on user-specific classes. Existing works rely on unstructured pruning, which introduces randomly distributed non-zero values in the model, making it unsuitable for hardware acceleration. Alternatively, some approaches employ structured pruning, such as channel pruning, but these tend to provide only minimal compression and may lead to reduced model accuracy. In this work, we propose CRISP, a novel pruning framework leveraging a hybrid structured sparsity pattern that combines both fine-grained N:m structured sparsity and coarse-grained block sparsity. Our pruning strategy is guided by a gradient-based class-aware saliency score, allowing us to retain weights crucial for user-specific classes. CRISP achieves high accuracy with minimal memory consumption for popular models like ResNet-50, VGG-16, and MobileNetV2 on ImageNet and CIFAR-100 datasets. Moreover, CRISP delivers up to 14x reduction in latency and energy consumption compared to existing pruning methods while maintaining comparable accuracy. Our code is available here. Shivam Aggarwal, Kuluhan Binici, Tulika Mitra |
DATE | 1 |
| 2024 | Evaluating Active-learning Based Performance Prediction of Parallel ApplicationsabstractMany scientific applications use Message Passing Interface for distributed-memory parallelism. These applications simulate complex physical phenomena and are routinely executed on supercomputers and high performance compute clusters. The applications may run for several hours or days and there may be long queue waiting times on the supercomputers. Often the application developers are unable to correctly estimate the runtime of their application for a new configuration and thus may overestimate the runtime requirements. This may increase the queue wait times along with sub-optimal job scheduling decisions. Thus accurate prediction of the performance of an application (performance modeling) is helpful. We have developed statistical models for performance prediction of MPI applications based on active learning algorithms and are able to achieve low Median Absolute Percentage Error (MdAPE) in many cases. We used five scientific applications and benchmarks – HPCG, LULESH, miniAMR, miniMD, and miniFE. We predict the execution times of these applications using various data sizes on up to 512 cores of the PARAM Sanganak supercomputer at IIT Kanpur. The statistical models gave MdAPE of 10.6, 3.5, 8.21, 12.62, and 32.25% for HPCG, LULESH, miniAMR, miniMD, and miniFE respectively. Shivam Aggarwal, Preeti Malakar |
e-Science | 1 |
| 2024 | Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAsabstractPost-training quantization (PTQ) is a powerful technique for model compression, reducing the numerical precision in neural networks without additional training overhead. Recent works have investigated adopting 8 -bit floating-point formats (FP8) in the context of PTQ for model inference. However, floating-point formats smaller than 8 bits and their relative comparison in terms of accuracy-hardware cost with integers remains unexplored on FPGAs. In this work, we present minifloats, which are reduced-precision floating-point formats capable of further reducing the memory footprint, latency, and energy cost of a model while approaching full-precision model accuracy. We implement a custom FPGA-based multiply-accumulate operator library and explore the vast design space, comparing minifloat and integer representations across 3 to 8 bits for both weights and activations. We also examine the applicability of various integer-based quantization techniques to minifloats. Our experiments show that minifloats offer a promising alternative for emerging workloads such as vision transformers. Shivam Aggarwal, Hans Jakob Damsgaard, Alessandro Pappalardo, Giuseppe Franco, Thomas B. Preußer, Michaela Blott, Tulika Mitra |
FPL | 1 |
| 2024 | Chameleon: Dual Memory Replay for Online Continual Learning on Edge DevicesabstractOnce deployed on edge devices, a deep neural network model should dynamically adapt to newly discovered environments and personalize its utility for each user. The system must be capable of continual learning, i.e., learning new information from a temporal stream of data in situ without forgetting previously acquired knowledge. However, creating a personalized continual learning framework poses significant challenges due to limited compute and storage resources on edge devices. Existing methods rely on large memory storage to preserve past data while learning from incoming streams, making them impractical for such devices. In this paper, we propose Chameleon as a hardware-friendly continual learning solution for user-centric continual learning with dual replay buffers. The strategy takes advantage of the hierarchical memory structure commonly found in edge devices, utilizing a short-term replay store in on-chip memory and a long-term replay store in off-chip memory. We also present an FPGA-based analytical model to estimate the compute and communication costs of the dual replay strategy on the hardware, making effective design choices considering various latent layer options. We conduct extensive experiments on four different models, demonstrating our method’s consistent performance across diverse model architectures. Our method achieves up to 7× speedup and improved energy efficiency on popular edge devices, including ZCU102 FPGA, NVIDIA Jetson Nano, and Google’s EdgeTPU. Our code is available at https://github.com/ecolab-nus/Chameleon. Shivam Aggarwal, Kuluhan Binici, Tulika Mitra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Chameleon: Dual Memory Replay for Online Continual Learning on Edge DevicesabstractOnce deployed on edge devices, a deep neural network model should dynamically adapt to newly discovered environments and personalize its utility for each user. The system must be capable of continual learning, i.e., learning new information from a temporal stream of data in situ without forgetting pre-viously acquired knowledge. However, the prohibitive intricacies of such a personalized continual learning framework stand at odds with limited compute and storage on edge devices. Existing continual learning methods rely on massive memory storage to preserve the past data while learning from the incoming data stream. We propose Chameleon, a hardware-friendly continual learning framework for user-centric training with dual replay buffers. The proposed strategy leverages the hierarchical memory structure available on most edge devices, introducing a short-term replay store in the on-chip memory and a long-term replay store in the off-chip memory to acquire new information while retaining past knowledge. Extensive experiments on two large-scale continual learning benchmarks demonstrate the efficacy of our proposed method, achieving better or comparable accuracy than existing state-of-the-art techniques while reducing the mem-ory footprint by roughly$16\times$. Our method achieves up to$7\times$speedup and energy efficiency on edge devices such as ZCU102 FPGA, NVIDIA Jetson Nano and Google's EdgeTPU. Our code is available at https://github.com/ecolab-nus/Chameleon. Shivam Aggarwal, Kuluhan Binici, Tulika Mitra |
DATE | 1 |
| 2022 | Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo ReplayabstractData-Free Knowledge Distillation (KD) allows knowledge transfer from a trained neural network (teacher) to a more compact one (student) in the absence of original training data. Existing works use a validation set to monitor the accuracy of the student over real data and report the highest performance throughout the entire process. However, validation data may not be available at distillation time either, making it infeasible to record the student snapshot that achieved the peak accuracy. Therefore, a practical data-free KD method should be robust and ideally provide monotonically increasing student accuracy during distillation. This is challenging because the student experiences knowledge degradation due to the distribution shift of the synthetic data. A straightforward approach to overcome this issue is to store and rehearse the generated samples periodically, which increases the memory footprint and creates privacy concerns. We propose to model the distribution of the previously observed synthetic samples with a generative network. In particular, we design a Variational Autoencoder (VAE) with a training objective that is customized to learn the synthetic data representations optimally. The student is rehearsed by the generative pseudo replay technique, with samples produced by the VAE. Hence knowledge degradation can be prevented without storing any samples. Experiments on image classification benchmarks show that our method optimizes the expected value of the distilled model accuracy while eliminating the large memory overhead incurred by the sample-storing methods. Kuluhan Binici, Shivam Aggarwal, Nam Trung Pham, Karianto Leman, Tulika Mitra |
AAAI | 2 |
| 2014 | Identification and Detection of Phishing Emails Using Natural Language Processing TechniquesabstractPhishing refers to fraudulent social engineering techniques used to elicit sensitive information from unsuspecting victims. In this paper, our scheme is aimed at detecting phishing mails which do not contain any links but bank on the victim's curiosity by luring them into replying with sensitive information. We exploit the common features among all such phishing emails such as non-mentioning of the victim's name in the email, a mention of monetary incentive and a sentence inducing the recipient to reply. This textual analysis can be further combined with header analysis of the email so that a final combined evaluation on the basis of both these scores can be done. We have shown that this method is far better than the existing Phishing Email Detection techniques as this covers emails without links while the pre-existing methods were based on the presumption of link(s). Shivam Aggarwal, S. D. Sudarsan |
SIN | 1 |