EDBT 2026 Demo / reviewers in the wild / expert
Anirudh Jain
dblp:235/7412
· DBLP profile ↗
8ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0004-8930-6512ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RASSM: Residue-based Acceleration of Single Sparse Matrix Computation via Adaptive Tiling
Anirudh Jain, Pulkit Gupta, Thomas M. Conte |
ASPLOS (1) | 1 |
| 2023 | Safety Hints for HTM Capacity Abort MitigationabstractHardware Transactional Memory (HTM) is a high-performance instantiation of the powerful programming abstraction of transactional memory, which simplifies the daunting— yet critically important—task of parallel programming. While many HTM implementations with variable complexity exist in the literature, commercially available HTMs impose rigid restrictions to transaction and system behavior, limiting their practical use. A key constraint is the limited size of supported transactions, implicitly capped by hardware buffering capacity. We identify the opportunity to expand the effective capacity of these limited hardware structures by being more selective in memory accesses that need to be tracked. We leverage compiler and virtual memory support to identify safe memory accesses, which can never cause a transaction abort, subsequently passed as safety hints to the underlying HTM. With minor extensions over a conventional HTM implementation, HinTM uses these hints to selectively allocate transactional state tracking resources to unsafe accesses only, thus expanding the HTM’s effective capacity, and conversely reducing capacity aborts. We demonstrate that HinTM effectively augments the performance of a range of baseline HTM configurations. When coupled with a POWER8 HTM implementation, HinTM eliminates 64% of transactional capacity aborts, achieving 1.4× average speedup, and up to 8.7×. Anirudh Jain, Divya Kiran Kadiyala, Alexandros Daglis |
HPCA | 1 |
| 2022 | Scalable Energy-Efficient Microarchitectures With Computational Error Tolerance Via Redundant Residue Number SystemsabstractDue to high leakage current and threshold voltage, Dennard scaling has reached its limit on conventional semiconductor technology. Energy reduction at the transistor level by simply lowering supply voltage has proven to be infeasible for these devices (e.g., MOSFETs). Some recently proposed millivolt switch techniques aim to mitigate these issues, by maintaining a high on/off ratio of drain currents with a much lower supply voltage. However,$V_{dd}$reduction is constrained by high intermittent error probabilities in millivolt switches. Energy-efficient microarchitectures that are computationally error-tolerant are therefore urgently needed. This article systematically leverages the error correction and checkpointing properties of Redundant Residue Number Systems (RRNS) by varying the number of non-redundant ($n$) and redundant ($r$) residues. The state-of-the-art of RRNS microarchitecture is confined to a fixed configuration point within such a($n$n,$r$r)-RRNSdesign plane, as it supports single error correction alone. Being able to efficiently handle resilience in this($n$n,$r$r)-RRNSplane significantly improves reliability, allowing further${V_{dd}}$reduction to save energy. To this end, first, we propose a scalable RRNS microarchitecture that simultaneously supports both, error-correction, as well as checkpointing with restart capabilities upon detecting uncorrectable errors. Second, we design a novel RRNS-based adaptive checkpointing&restart mechanisms that automatically guarantees reliability while minimizing the energy-delay product (EDP). To the best of our knowledge, these are the first set of checkpointing mechanisms targeting the RRNS infrastructure. Moreover, these mechanisms optimize the usage efficiency of memory capacity. Third, we systematically explore the RRNS design space to find the best ($n$,$r$) configuration point. For similar reliability when compared to a conventional binary core without computationally error-tolerant (runs at high$V_{dd}$), the proposed RRNS scalable microarchitecture reduces EDP by 53 percent on average for memory-intensive workloads and by 67 percent on average for non-memory-intensive workloads. Bobin Deng, Sriseshan Srikanth, Anirudh Jain, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook |
IEEE Trans. Computers | 3 |
| 2021 | SortCache: Intelligent Cache Management for Accelerating Sparse Data WorkloadsabstractSparse data applications have irregular access patterns that stymie modern memory architectures. Although hyper-sparse workloads have received considerable attention in the past, moderately-sparse workloads prevalent in machine learning applications, graph processing and HPC have not. Where the former can bypass the cache hierarchy, the latter fit in the cache. This article makes the observation that intelligent, near-processor cache management can improve bandwidth utilization for data-irregular accesses, thereby accelerating moderately-sparse workloads. We propose SortCache, a processor-centric approach to accelerating sparse workloads by introducing accelerators that leverage the on-chip cache subsystem, with minimal programmer intervention. Sriseshan Srikanth, Anirudh Jain, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook |
ACM Trans. Archit. Code Optim. | 2 |
| 2020 | Special Session: Exploring the Ultimate Limits of Adiabatic CircuitsabstractThe field of adiabatic circuits is rooted in electronics know-how stretching all the way back to the 1960s and has potential applications in vastly increasing the energy efficiency of far-future computing. But now, the field is experiencing an increased level of attention in part due to its potential to reduce the vulnerability of systems to side-channel attacks that exploit, e.g., unwanted EM emissions, power supply fluctuations, and so forth. In this context, one natural question is: Just how low can the energy dissipation from adiabatic circuits, and the associated extraneous signal emissions, be made to go? We argue that the ultimate limits of this approach lie much farther away than is commonly appreciated. Recent advances at Sandia National Laboratories in the design of fully static, fully adiabatic CMOS logic styles and high-quality energy-recovering resonant power-clock drivers offer the potential to reduce dynamic switching losses by multiple orders of magnitude, and, particularly for cryogenic applications, optimization of device structures can reduce the standby power consumption of inactive devices, and the ultimate dissipation limits of the adiabatic approach, by multiple orders of magnitude as well. In this paper, we review the above issues, and give a preliminary overview of our group's activities towards the demonstration of groundbreaking levels of energy efficiency for semiconductor-based logic, together with a broader exploration of the ultimate limits of physically realizable techniques for approaching the theoretical ideal of perfect thermodynamic reversibility in computing, and the study of the implications of this technology direction for practical computing architectures. Michael P. Frank, Robert W. Brocato, Thomas M. Conte, Alexander H. Hsia, Anirudh Jain, Nancy A. Missert, Karpur Shukla, Brian D. Tierney |
ICCD | 5 |
| 2020 | Cloud Removal in Satellite Images Using Spatiotemporal Generative NetworksabstractSatellite images hold great promise for continuous environmental monitoring and earth observation. Occlusions cast by clouds, however, can severely limit coverage, making ground information extraction more difficult. Existing pipelines typically perform cloud removal with simple temporal composites and hand-crafted filters. In contrast, we cast the problem of cloud removal as a conditional image synthesis challenge, and we propose a trainable spatiotemporal generator network (STGAN) to remove clouds. We train our model on a new large-scale spatiotemporal dataset that we construct, containing 97640 image pairs covering all continents. We demonstrate experimentally that the proposed STGAN model outperforms standard models and can generate realistic cloud-free images with high PSNR and SSIM values across a variety of atmospheric conditions, leading to improved performance in downstream tasks such as land cover classification. Vishnu Sarukkai, Anirudh Jain, Burak Uzkent, Stefano Ermon |
WACV | 2 |
| 2020 | MetaStrider: Architectures for Scalable Memory-centric Reduction of Sparse Data StreamsabstractReduction is an operation performed on the values of two or more key-value pairs that share the same key. Reduction of sparse data streams finds application in a wide variety of domains such as data and graph analytics, cybersecurity, machine learning, and HPC applications. However, these applications exhibit low locality of reference, rendering traditional architectures and data representations inefficient. This article presents MetaStrider, a significant algorithmic and architectural enhancement to the state-of-the-art, SuperStrider. Furthermore, these enhancements enable a variety of parallel, memory-centric architectures that we propose, resulting in demonstrated performance that scales near-linearly with available memory-level parallelism. Sriseshan Srikanth, Anirudh Jain, Joseph M. Lennon, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook |
ACM Trans. Archit. Code Optim. | 2 |
| 2019 | Practical Deep Learning with Bayesian PrinciplesabstractBayesian methods promise to fix many shortcomings of deep learning, but they are impractical and rarely match the performance of standard methods, let alone improve them. In this paper, we demonstrate practical training of deep networks with natural-gradient variational inference. By applying techniques such as batch normalisation, data augmentation, and distributed training, we achieve similar performance in about the same number of epochs as the Adam optimiser, even on large datasets such as ImageNet. Importantly, the benefits of Bayesian principles are preserved: predictive probabilities are well-calibrated, uncertainties on out-of-distribution data are improved, and continual-learning performance is boosted. This work enables practical deep learning while preserving benefits of Bayesian principles. A PyTorch implementation is available as a plug-and-play optimiser. Kazuki Osawa, Siddharth Swaroop, Mohammad Emtiyaz Khan, Anirudh Jain, Runa Eschenhagen, Richard E. Turner, Rio Yokota |
NeurIPS | 4 |