Alish Kanani

dblp:272/2748 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0000-8585-9241ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 HeMu: Energy-Efficient DNN Inferencing via Heterogeneous Multichiplet Architectures
Alish Kanani, Janardhan Rao Doppa, Ümit Y. Ogras, Partha Pratim Pande
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 MFIT : Multi-FIdelity Thermal Modeling for 2.5D and 3D Multi-Chiplet Architectures
abstract
Rapidly evolving artificial intelligence and machine learning applications require ever-increasing computational capabilities, while monolithic 2D design technologies approach their limits. 2.5D/3D heterogeneous integration of smaller chiplets using advanced packaging has emerged as a promising paradigm for addressing this limit and meeting performance demands. These approaches offer a significant cost reduction and higher manufacturing yield than monolithic 2D integrated circuits. However, the compact arrangement and high compute density of these systems exacerbate thermal management challenges, potentially compromising performance. Addressing these thermal modeling challenges is critical, especially as system sizes grow and different design stages require varying levels of accuracy and speed. Since no single thermal modeling technique meets all these needs, this article introduces MFIT, a range of multi-fidelity thermal models that effectively balance accuracy and speed. These multi-fidelity models can enable efficient design space exploration and runtime thermal management. Our extensive testing on systems with 16, 36, and 64 2.5D integrated chiplets and 16×3 3D integrated chiplets demonstrates that these models can reduce execution times from days to mere seconds and milliseconds with negligible loss in accuracy.
Lukas Pfromm, Alish Kanani, Parth Solanki, Eric Tervo, Jaehyun Park 0005, Janardhan Rao Doppa, Partha Pratim Pande, Ümit Y. Ogras
ACM Trans. Design Autom. Electr. Syst.2
2025 THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
abstract
Chiplet-based integration enables large-scale systems that combine diverse technologies, enabling higher yield, lower costs, and scalability, making them well-suited to AI workloads. Processing-in-Memory (PIM) has emerged as a promising solution for AI inference, leveraging technologies such as ReRAM, SRAM, and FeFET, each offering unique advantages and tradeoffs. A heterogeneous chiplet-based PIM architecture can harness the complementary strengths of these technologies to enable higher performance and energy efficiency. However, scheduling AI workloads across such a heterogeneous system is challenging due to competing performance objectives, dynamic workload characteristics, and power and thermal constraints. To address this need, we propose THERMOS, a thermally-aware, multi-objective scheduling framework for AI workloads on heterogeneous multi-chiplet PIM architectures. THERMOS trains a single multi-objective reinforcement learning (MORL) policy that is capable of achieving Pareto-optimal execution time, energy, or a balanced objective at runtime, depending on the target preferences. Comprehensive evaluations show that THERMOS achieves up to 89% faster average execution time and 57% lower average energy consumption than baseline AI workload scheduling algorithms with only 0.14% runtime and 0.022% energy overhead.
Alish Kanani, Lukas Pfromm, Janardhan Rao Doppa, Partha Pratim Pande, Ümit Y. Ogras
ACM Trans. Embed. Comput. Syst.1
2025 eMamba: Efficient Acceleration Framework for Mamba Models in Edge Computing
abstract
State Space Model (SSM)-based machine learning architectures have recently gained significant attention for processing sequential data. Mamba, a recent sequence-to-sequence SSM, offers competitive accuracy with superior computational efficiency compared to state-of-the-art transformer models. While this advantage makes Mamba particularly promising for resource-constrained edge devices, no hardware acceleration frameworks are currently optimized for deploying it in such environments. This article presents eMamba, a comprehensive end-to-end hardware acceleration framework explicitly designed for deploying Mamba models on edge platforms. eMamba maximizes computational efficiency by replacing complex normalization layers with lightweight hardware-aware alternatives and approximating expensive operations, such as SiLU activation and exponentiation, considering the target applications. Then, it performs an approximation-aware neural architecture search (NAS) to tune the learnable parameters used during approximation. Evaluations with Fashion-MNIST, CIFAR-10, and MARS, an open-source human pose estimation dataset, show eMamba achieves comparable accuracy to state-of-the-art techniques using 1.63–19.9× fewer parameters. In addition, it generalizes well to large-scale natural language tasks, demonstrating stable perplexity across varying sequence lengths on the WikiText2 dataset. We also quantize and implement the entire eMamba pipeline on an AMD ZCU102 FPGA and ASIC using GlobalFoundries (GF) 22 nm technology. Experimental results show 4.95–5.62× lower latency and 2.22–9.95× higher throughput, with 4.77× smaller area, 9.84× lower power, and 48.6× lower energy consumption than baseline solutions while maintaining competitive accuracy.
Alish Kanani, Ümit Y. Ogras, Jaehyun Park 0005
ACM Trans. Embed. Comput. Syst.4
2024 Thermal Modeling and Management Challenges in Heterogenous Integration: 2.5D Chiplet Platforms and Beyond
abstract
Heterogeneous integration using 2.5D chiplet platforms provides a new avenue for compact scale-out implementations of emerging applications, such as deep learning (DL). Integrating multiple small chiplets using a Network-on-Interposer (NoI) offers not only significant cost reductions and higher manufacturing yield compared to 2D ICs but also better thermal efficiency than 3D ICs and easier heterogeneous integration. However, dense integration and substantial compute density exacerbate thermal design problems, threatening to undermine the potential performance and cost benefits. Due to the significant role of temperature in the operation and reliability of integrated systems, it is critical to understand the role of heat in this emerging design area. However, little work has considered the thermal consequences of closely packaging a large number of computational elements. This paper overviews the thermal modeling challenges for chiplet-based 2.5D platforms, overviews existing approaches, and discusses the opportunities enabled by fast and accurate thermal models.
Jaehyun Park 0005, Alish Kanani, Lukas Pfromm, Parth Solanki, Eric Tervo, Janardhan Rao Doppa, Partha Pratim Pande, Ümit Y. Ogras
VTS2
2024 Runtime Monitoring of ML-Based Scheduling Algorithms Toward Robust Domain-Specific SoCs
abstract
Machine learning (ML) algorithms are being rapidly adopted to perform dynamic resource management tasks in heterogeneous system on chips. For example, ML-based task schedulers can make quick, high-quality decisions at runtime. Like any ML model, these offline-trained policies depend critically on the representative power of the training data. Hence, their performance may diminish or even catastrophically fail under unknown workloads, especially new applications. This article proposes a novel framework to continuously monitor the system to detect unforeseen scenarios using a gradient-based generalization metric called coherence. The proposed framework accurately determines whether the current policy generalizes to new inputs. If not, it incrementally trains the ML scheduler to ensure the robustness of the task-scheduling decisions. The proposed framework is evaluated thoroughly with a domain-specific SoC and six real-world applications. It can detect whether the trained scheduler generalizes to the current workload with 88.75%–98.39% accuracy. Furthermore, it enables$1.1\times -14\times $faster execution time when the scheduler is incrementally trained. Finally, overhead analysis performed on an Nvidia Jetson Xavier NX board shows that the proposed framework can run as a real-time background task.
A. Alper Goksoy, Alish Kanani, Satrajit Chatterjee, Ümit Y. Ogras
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2