EDBT 2026 Demo / reviewers in the wild / expert
Kerem Akarvardar
dblp:22/153
· DBLP profile ↗
11ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0001-5957-826XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ultrafast Generative AI by Ultradense 3D Integration: A Case Study on LLM-based Edge InferenceabstractGenerative AI (GenAI) is one of the most critical applications today, continually challenging the limits of semiconductor technology. We introduce a very fine-grained 3D memory-on-logic architecture along with a novel data mapping strategy to support Large Language Model (LLM)-based GenAI, including both prefill and generation stages. Our conceptual analysis shows how ultradense 3D connectivity can enhance text generation speed and energy-efficiency well-beyond current limits. Preliminary findings from a basic analytical model indicate that the single batch autoregressive generation rate for Llama 3.2 1B could surpass 5K tokens/sec by maximizing weight locality and enhancing memory bandwidth through massively parallel 3D links between Multiply-Accumulate (MAC) units in the logic tier and their dedicated memory partitions in the 3D stack. We also explore the impact of advanced logic nodes and quantify their benefits in reducing prefill latency. Finally, we examine the challenges associated with memory access power and power density under extreme bandwidth conditions and present pipelined access strategies to address them. Kerem Akarvardar, Xiaoyu Sun 0001, Brian Crafton, Xiaochen Peng, Haruki Mori, Abhiroop Bhattacharjee, Hidehiro Fujiwara, H.-S. Philip Wong |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2025 | Finding the Pareto Frontier of Low-Precision Data Formats and MAC Architecture for LLM InferenceabstractTo accelerate AI applications, numerous data formats and physical implementations of matrix multiplication have been proposed, creating a complex design space. This paper studies the efficient MAC implementation of the integer, floating-point, posit, and logarithmic number system (LNS) data formats and Microscaling (MX) and VectorScaled Quantization (VSQ) block data formats. We evaluate the area, power, and numerical accuracy (evaluated as signal-to-quantization noise ratio) of $\mathbf{3 5, 0 0 0}$ MAC designs spanning each data format and several key design parameters such as the inner product size and accumulation width. We find that for the same numerical accuracy, pareto optimal MAC designs with emerging data formats (LNS16, MXINT8, VSQINT4) achieve $1.8 \times 2.2 \times$, and $1.9 \times$ TOPs/W improvement compared to FP16, FP8, and FP4 dot product implementations. Brian Crafton, Xiaochen Peng, Xiaoyu Sun 0001, Ashwin Sanjay Lele, Win-San Khwa, Kerem Akarvardar |
DAC | 7 |
| 2024 | Non-volatile Memory Technologies for Edge AI ApplicationsabstractEmbedded non-volatile memories (NVMs) hold significant promise for advancing edge AI applications by offering unique advantages in near/in-memory computing-based accelerators. This invited paper delivers an in-depth review of the use cases for NVMs, emphasizing their potential to enhance area- and energy-efficiency in edge devices. We begin by outlining the latest advancements in TSMC NVM technologies and examining several NVM-based accelerator test chips enabled by the TSMC University Shuttle Program. Additionally, we delve into the tradeoffs involved in optimizing NVM devices, explore their potential for approximate computing applications, and assess the impact of NVM non-idealities on inference accuracy. Xiaoyu Sun 0001, Win-San Khwa, Xiaochen Peng, Meng-Fan Chang, Kerem Akarvardar |
ICCAD | 5 |
| 2024 | Efficient Processing of MLPerf Mobile Workloads Using Digital Compute-In-Memory MacrosabstractCompute-in-memory (CIM) has recently emerged as a promising design paradigm to accelerate deep neural network (DNN) processing. Continuously better energy and area efficiency at the macrolevel had been reported through many testchips over the last few years. However, in those macro design-oriented studies, accelerator-level considerations, such as memory accesses and processing of entire DNN workloads have not been investigated in-depth. In this article, we aim to fill this gap starting with the characteristics of our latest CIM macro fabricated with cutting-edge FinFET CMOS technology at 4-nm node. We then study, through an accelerator simulator developed in-house, three key items that would determine the efficiency of our CIM macro in the accelerator context while running MLPerf Mobile suite: 1) dataflow optimization; 2) optimal selection of CIM macro dimensions to further improve macro utilization; and 3) optimal combination of multiple CIM macros. Although there is typically a stark contrast between macro-level peak and accelerator-level average throughput and energy efficiency, the aforementioned optimizations are shown to improve the macro utilization by$3.04\times $and reduce the energy-delay product (EDP) to$0.34\times $compared to the original macro on MLPerf Mobile inference workloads. While we exploit a digital CIM macro in this study, the findings and proposed methods remain valid for other types of CIM (such as analog CIM and analog–digital–hybrid CIM) as well. Xiaoyu Sun 0001, Weidong Cao 0001, Brian Crafton, Kerem Akarvardar, Haruki Mori, Hidehiro Fujiwara, Hiroki Noguchi, Yu-Der Chih, Meng-Fan Chang, Yih Wang, Tsung-Yung Jonathan Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | Estimating Power, Performance, and Area for On-Sensor Deployment of AR/VR Workloads Using an Analytical FrameworkabstractAugmented Reality and Virtual Reality have emerged as the next frontier of intelligent image sensors and computer systems. In these systems, 3D die stacking stands out as a compelling solution, enabling in situ processing capability of the sensory data for tasks such as image classification and object detection at low power, low latency, and a small form factor. These intelligent 3D CMOS Image Sensor (CIS) systems present a wide design space, encompassing multiple domains (e.g., computer vision algorithms, circuit design, system architecture, and semiconductor technology, including 3D stacking) that have not been explored in-depth so far. This article aims to fill this gap. We first present an analytical evaluation framework, STAR-3DSim, dedicated to rapid pre-RTL evaluation of 3D-CIS systems capturing the entire stack from the pixel layer to the on-sensor processor layer. With STAR-3DSim, we then propose several knobs for PPA (power, performance, area) improvement of the Deep Neural Network (DNN) accelerator that can provide up to 53%, 41%, and 63% reduction in energy, latency, and area, respectively, across a broad set of relevant AR/VR workloads. Last, we present full-system evaluation results by taking image sensing, cross-tier data transfer, and off-sensor communication into consideration. Xiaoyu Sun 0001, Xiaochen Peng, Sai Qian Zhang, Jorge Gomez 0002, Win-San Khwa, Syed Shakib Sarwar, Ziyun Li 0001, Weidong Cao 0001, Chiao Liu, Meng-Fan Chang, Barbara De Salvo, Kerem Akarvardar, H.-S. Philip Wong |
ACM Trans. Design Autom. Electr. Syst. | 13 |
| 2023 | Technology Prospects for Data-Intensive ComputingabstractFor many decades, progress in computing hardware has been closely associated with CMOS logic density, performance, and cost. As such, slowdown in 2-D scaling, frequency saturation in CPUs, and increased cost of design and chip fabrication for advanced technology nodes since the early 2000s have led to concerns about how semiconductor technology may evolve in the future. However, the last two decades have also witnessed a parallel development in the application landscape: the advent of big data and consequent rise of data-intensive computing, using techniques such as machine learning. In this article, we advance the idea that data-intensive computing would further cement semiconductor technology as a foundational technology with multidimensional pathways for growth. Continued progress of semiconductor technology in this new context would require the adoption of a system-centric perspective to holistically harness logic, memory, and packaging resources. After examining the performance metrics for data-intensive computing, we present the historical trends for general-purpose graphics processing unit (GPGPU) as a representative data-intensive computing hardware. Thereon, we estimate the values of the key data-intensive computing parameters for the next decade, and our projections may serve as a precursor for a dedicated technology roadmap. By analyzing the compiled data, we identify and discuss specific opportunities and challenges for data-intensive computing hardware technology. Kerem Akarvardar, H.-S. Philip Wong |
Proc. IEEE | 1 |
| 2020 | A Density Metric for Semiconductor Technology [Point of View]abstractSince its inception, the semiconductor industry has used a physical dimension (the minimum gate length of a transistor) as a means to gauge continuous technology advancement. This metric is all but obsolete today. As a replacement, we propose a density metric, which aims to capture how advances in semiconductor device technologies enable system-level benefits. The proposed metric can be used to gauge advances in future generations of semi-conductor technologies in a holistic way, by accounting for the progress in logic, memory, and packaging/integration technologies simultaneously. H.-S. Philip Wong, Kerem Akarvardar, Dimitri A. Antoniadis, Jeffrey Bokor, Chenming Hu, Tsu-Jae King Liu, Subhasish Mitra, James D. Plummer, Sayeef S. Salahuddin |
Proc. IEEE | 2 |
| 2020 | Scanning the IssueabstractThis month’s issue offers insight into efficient compression and execution of DNNs, the challenge of connecting rural areas, and the clique problem in wireless communication. which H.-S. Philip Wong, Kerem Akarvardar, Dimitri A. Antoniadis, Jeffrey Bokor, Chenming Hu, Tsu-Jae King Liu, Subhasish Mitra, James D. Plummer, Sayeef S. Salahuddin, Lei Deng 0003, Song Han 0003, Luping Shi, Yuan Xie 0001, Elias Yaacoub, Mohamed-Slim Alouini, Ahmed Douik, Hayssam Dahrouj, Tareq Y. Al-Naffouri |
Proc. IEEE | 2 |
| 2010 | Efficient FPGAs using nanoelectromechanical relaysabstractNanoelectromechanical (NEM) relays are promising candidates for programmable routing in Field-Programmable-Gate Arrays (FPGAs). This is due to their zero leakage and potentially low on-resistance. Moreover, NEM relays can be fabricated using a low-temperature process and, hence, may be monolithically integrated on top of CMOS circuits. Hysteresis characteristics of NEM relays can be utilized for designing programmable routing switches in FPGAs without requiring corresponding routing SRAM cells. Our simulation results demonstrate that the use of NEM relays for programmable routing in FPGAs can simultaneously provide 43.6% footprint area reduction, 37% leakage power reduction, and up to 28% critical path delay reduction compared to traditional SRAM-based CMOS FPGAs at the 22nm technology node. Chen Chen 0018, Roozbeh Parsa, Nishant Patil, Soogine Chong, Kerem Akarvardar, J. Provine, Jeff Watt, Roger T. Howe, H.-S. Philip Wong, Subhasish Mitra |
FPGA | 5 |
| 2009 | Nanoelectromechanical (NEM) relays integrated with CMOS SRAM for improved stability and low leakageabstractWe present a hybrid nanoelectromechanical (NEM)/CMOS static random access memory (SRAM) cell, in which the two pull-down transistors of a conventional CMOS six transistor (6T) SRAM cell are replaced with NEM relays. This SRAM cell utilizes the infinite subthreshold slope and hysteretic properties of NEM relays to dramatically increase the cell stability compared to the conventional CMOS 6T SRAM cells. It also utilizes the zero off-state leakage of NEM relays to significantly decrease static power dissipation. The structure is designed so that the relatively long mechanical delay of the NEM relays does not result in performance degradation. Circuit simulations are performed using a VerilogA model of a NEM relay. Compared to a 65nm CMOS 6T SRAM cell, when 10nm-gap NEM relays (pull-in voltage = 0.8V, pull-out voltage = 0.2V, on resistance = 1kΩ) are integrated, hold and read static noise margin (SNM) improve by ~110% and ~250%, respectively. In addition, static power dissipation decreases by ~85%. The write delay decreases by ~60%, while read delay decreases by ~10%. The advantages in SNM and static power dissipation are expected to increase with scaling. Soogine Chong, Kerem Akarvardar, Roozbeh Parsa, Jun-Bo Yoon, Roger T. Howe, Subhasish Mitra, H.-S. Philip Wong |
ICCAD | 2 |
| 2005 | The G4-FET: a universal and programmable logic gateabstractThe G4-FET, a four-gate transistor compatible with standard silicon-on-insulator (SOI) CMOS technology, provides unique opportunities as a logic device. Combining both JFET- and MOSFET-like actions within one transistor body, the G4-FET offers two side (lateral) junction-based gates, a top MOS gate, and a MOS back gate that is activated by SOI substrate biasing. The G4-FET's conduction characteristics are controlled by the combined interaction of these four gates. In this paper, the G4-FET is demonstrated as a logic device, resulting in a universal and programmable logic gate that can lead to the design of more efficient logic circuits. As an example, we present a new full adder design based on the G4-FET that is significantly more efficient than conventional designs. Amir Fijany, Farrokh Vatan, Mohammad M. Mojarradi, Nikzad Benny Toomarian, Benjamin J. Blalock, Kerem Akarvardar, Sorin Cristoloveanu, Pierre Gentil |
ACM Great Lakes Symposium on VLSI | 6 |