Jamin Seo

dblp:334/0541 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 AIRCHITECT v2: Learning the Hardware Accelerator Design Space Through Unified Representations
abstract
Design space exploration (DSE) plays a crucial role in enabling custom hardware architectures, particularly for emerging applications like AI, where optimized and specialized designs are essential. With the growing complexity of deep neural networks (DNNs) and the introduction of advanced large language models (LLMs), the design space for DNN accelerators is expanding at an exponential rate. Additionally, this space is highly non-uniform and non-convex, making it increasingly difficult to navigate and optimize. Traditional DSE techniques rely on search-based methods, which involve iterative sampling of the design space to find the optimal solution. However, this process is both time-consuming and often fails to converge to the global optima for such design spaces. Recently, AIRCHITECT vl, the first attempt to address the limitations of search-based techniques, transformed DSE into a constant-time classification problem using recommendation networks. However, AIRCHITECT v1 lacked generalizability and had poor performance in complex design spaces. In this work, we propose AIRCHITECT v2, a more accurate and generalizable learning-based DSE technique applicable to large-scale design spaces that overcomes the shortcomings of earlier approaches. Specifically, we devise an encoder-decoder transformer model that ($a$) encodes the complex design space into a uniform intermediate representation using contrastive learning and (b) leverages a novel unified representation blending the advantages of classification and regression to effectively explore the large DSE space without sacrificing accuracy. Experimental results evaluated on 105real DNN workloads demonstrate that, on average, AIRCHITECT v2 outperforms existing techniques by 15% in identifying optimal design points. Furthermore, to demonstrate the generalizability of our method, we evaluate performance on unseen model workloads and attain a 1.7 x improvement in inference latency on the identified hardware architecture. Code and dataset are available at: https://github.com/maestro-project/AIrchitect-v2.
Jamin Seo, Akshat Ramachandran, Yu-Chuan Chuang, Anirudh Itagi, Tushar Krishna
DATE1
2025 Exploring Constrained Dataflow Accelerators for Real-Time Multi-Task Multi-Model Ml Workloads
abstract
Emerging machine learning (ML) workloads, such as those in AR/VR applications or drones, exhibit real-time multitask multi-model (RT-MTMM) characteristics. These workloads demand efficient executions of diverse combinations of multiple models within strict real-time constraints and tight energy budgets. Consequently, optimizing hardware accelerator design and dataflow choices for such ML models becomes imperative. Flexible dataflows in hardware have been explored to enhance performance and energy efficiency for diverse models. However, this flexibility involves substantial hardware costs, potentially outweighing the performance benefits. While coarse-grained flexible accelerators, including heterogeneous dataflow accelerators [1], have been proposed to mitigate such costs, they may prove suboptimal depending on combinations of models, potentially leading to under-performance of the hardware. Hence, this paper investigates the desired balance between the benefits and costs of flexibility. We present a systematic and quantitative exploration of medium-grained flexible accelerators, demonstrating that constrained yet judicious domain-aware flexibility choices in tile sizes (T), loop order (O), parallel dimensions$(\mathrm{P})$, and array shape (S) can allow RT-MTMM accelerators to achieve comparable real-time performance ($\times 1.06 / \times 1.13$) and energy efficiency ($\times 1.09 / \times 1.07$) to fully flexible designs with significantly lower area overhead ($\times 0.67 \times 0.91$) on edge/mobilescale accelerators.
Jamin Seo, Jianming Tong, Tushar Krishna, Hyoukjun Kwon
ISPASS1