Ashitabh Misra

dblp:308/4372 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-2523-5856ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Computer networks
1 paper
Edge and fog computing · 100%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
0.912025
Latency-Constrained Input-Aware Quantization of Time Series Inference Workflows at the Edge · INFOCOM 2025
Edge and fog computing
edge inference
0.912025
Latency-Constrained Input-Aware Quantization of Time Series Inference Workflows at the Edge · INFOCOM 2025
Performance modeling and evaluation › cache performance modeling
cache miss equation
0.712023
BullsEye : Scalable and Accurate Approximation Framework for Cache Miss Calculation · ACM Trans. Archit. Code Optim. 2023
Performance modeling and evaluation
cache performance modeling
0.712023
BullsEye : Scalable and Accurate Approximation Framework for Cache Miss Calculation · ACM Trans. Archit. Code Optim. 2023
Performance modeling and evaluation › cache performance modeling
reuse distance analysis
0.712023
BullsEye : Scalable and Accurate Approximation Framework for Cache Miss Calculation · ACM Trans. Archit. Code Optim. 2023
Compilers and program optimization › loop transformation
polyhedral compilation
0.212023
BullsEye : Scalable and Accurate Approximation Framework for Cache Miss Calculation · ACM Trans. Archit. Code Optim. 2023

Methods — techniques the papers use, named apart from their topics

quantization · 2.6latency optimization · 2.6sub-polyhedral approximation · 1.3linearization · 1.3handelman's theorem · 1.3domain sampling · 1.3bernstein's representation · 1.3
YearPublicationVenuePosition
2025 Latency-Constrained Input-Aware Quantization of Time Series Inference Workflows at the Edge
Ashitabh Misra, Nurani Saoda, Tarek F. Abdelzaher
INFOCOM1
2025 The bottlenecks of AI: challenges for embedded and real-time research in a data-centric age
abstract
Abstract Recent advances in AI culminate a shift in science and engineering away from strong reliance on algorithmic and symbolic knowledge towards new data-driven approaches. How does the emerging intelligent data-centric world impact research on real-time and embedded computing? We argue for two effects: (1) new challenges in embedded system contexts, and (2) new opportunities for community expansion beyond the embedded domain. First, on the embedded system side, the shifting nature of computing towards data-centricity affects the types of bottlenecks that arise. At training time, the bottlenecks are generally data-related. Embedded computing relies on scarce sensor data modalities, unlike those commonly addressed in mainstream AI, necessitating solutions for efficient learning from scarce sensor data. At inference time, the bottlenecks are resource-related, calling for improved resource economy and novel scheduling policies. Further ahead, the convergence of AI around large language models (LLMs) introduces additional model-related challenges in embedded contexts. Second, on the domain expansion side, we argue that community expertise in handling resource bottlenecks is becoming increasingly relevant to a new domain: the cloud environment, driven by AI needs. The paper discusses the novel research directions that arise in the data-centric world of AI, covering data-, resource-, and model-related challenges in embedded systems as well as new opportunities in the cloud domain.
Tarek F. Abdelzaher, Yigong Hu, Denizhan Kara, Tomoyoshi Kimura, Ashitabh Misra, Vishakha Ramani, Olivier Tardieu, Tianshi Wang 0002, Maggie B. Wigness, Alaa Youssef
Real Time Syst.5
2023 ViX: Analysis-driven Compiler for Efficient Low-Precision Variational Inference
abstract
As large quantities of stochastic data are processed onboard tiny edge devices, these systems must constantly make decisions under uncertainty. This challenge necessitates principled embedded compiler support for time- and energy-efficient probabilistic inference. However, compiling probabilistic inference to run on the edge is significantly understudied, and the existing research is limited to computationally expensive MCMC algorithms. Hence, these works cannot leverage faster variational inference algorithms which can better scale to larger data sizes that are representative of realistic workloads in the edge setting. However, naively writing code for differentiable inference on resource-constrained edge devices is challenging due to the need for expensive floating point computations. Even when using reduced precision, a developer still faces the challenge of choosing the right quantization scheme, as gradients can be notoriously unstable in the face of low-precision. To address these challenges, we propose ViX which is the first compiler for low-precision probabilistic programming with variational inference. ViX generates optimized variational inference code in reduced precision by automatically exploiting Bayesian domain knowledge and analytical mathematical properties to ensure that low-precision gradients can still be effectively used. ViX can scale inference to much larger data-sets than previous compilers for resource-constrained probabilistic programming while attaining both high accuracy and significant speedup. Our evaluation of ViX across 7 benchmarks shows that ViX-generated code is up to 8.15× faster than performing the same variational inference in 32-bit floating point and also up to 22.67× faster than performing the variational inference in 64-bit double precision, all with minimal accuracy loss. Further, on a subset of our benchmarks, ViX can scale inference to data sizes between 16–80× larger than the existing state-of-the-art tool Statheros.
Ashitabh Misra, Jacob Laurel, Sasa Misailovic
DATE1
2023 BullsEye : Scalable and Accurate Approximation Framework for Cache Miss Calculation
abstract
For Affine Control Programs or Static Control Programs (SCoP), symbolic counting of reuse distances could induce polynomials for each reuse pair. These polynomials along with cache capacity constraints lead to non-affine (semi-algebraic) sets; and counting these sets is considered to be a hard problem. The state-of-the-art methods use various exact enumeration techniques relying on existing cardinality algorithms that can efficiently count affine sets. We propose BullsEye , a novel, scalable, accurate, and problem-size independent approximation framework. It is an analytical cache model for fully associative caches with LRU replacement policy focusing on sampling and linearization of non-affine stack distance polynomials. First, we propose a simple domain sampling method that can improve the scalability of exact enumeration. Second, we propose linearization techniques relying on Handelman’s theorem and Bernstein’s representation . To improve the scalability of the Handelman’s theorem linearization technique, we propose template (Interval or Octagon) sub-polyhedral approximations. Our methods obtain significant compile-time improvements with high-accuracy when compared to HayStack on important polyhedral compilation kernels such as nussinov , cholesky , and adi from PolyBench , and harris , gaussianblur from LLVM -TestSuite. Overall, on PolyBench kernels, our methods show up to 3.31× (geomean) speedup with errors below ≈ 0.08% (geomean) for the octagon sub-polyhedral approximation.
Nilesh Rajendra Shah, Ashitabh Misra, Antoine Miné, Rakesh Venkat, Ramakrishna Upadrasta
ACM Trans. Archit. Code Optim.2