Chris Kjellqvist

dblp:260/6347 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-4792-1910ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 COCOSSim: A Cycle-Accurate Simulator for Heterogeneous Systolic Array Architectures
abstract
Performance modeling is an essential tool for enabling cost-effective and efficient exploration of architectural and microarchitectural design decisions for hardware development. However, existing simulators face limitations in architectural flexibility, memory hierarchy modeling, and support for modern accelerators that cater to the computational demands of contemporary machine learning models, where non-linear operations like softmax and other vectorized activations have become increasingly important. To address these gaps, we present COCOSSim, a cycleaccurate performance simulator designed for evaluating architectural and microarchitectural modifications in heterogeneous systolic array-based accelerators. COCOSSim supports a wide range of neural network inference workloads, integrating systolic arrays with vector units for non-linear operations and DRAMSim3 for memory modeling. By offering a PyTorch frontend for seamless model integration, flexibility in scheduling strategies, and the ability to make architectural and microarchitectural modifications, COCOSSim enables detailed performance analysis and design exploration. When validated against the Google TPU v3, COCOSSim achieves an average error rate of 13 %. Additionally, COCOSSim is highly scalable and offers simulation speeds an order of magnitude faster than prior work while providing significantly greater flexibility and enhanced modeling capabilities. We present two case studies demonstrating how COCOSSim can be used to evaluate architectural modifications to modern heterogeneous accelerators: model parallelism and operation fusion.
Mansi Choudhary, Chris Kjellqvist, Jiaao Ma, Lisa Wu Wills
ISPASS2
2025 Beethoven: A Heterogeneous Multi-Core Accelerator System Composer
abstract
Hardware Development is challenging in large part due to the complexity of incorporating realistic designs onto hardware devices (e.g., FPGAs, CGRAs, ASICs). This work proposes a multi-core, hardware-software accelerator design framework called Beethoven. Beethoven provides a flexible and reusable many-core accelerator System-On-Chip integration environment through programming abstractions for Register-Transfer Logic development, generation of software linkage between the host system and the accelerator, and provision of a host software runtime. We thoroughly evaluate Beethoven on microbenchmarks, MachSuite, and a compute-limited attention accelerator. We compare Beethoven generated accelerated system performance to High-Level Synthesis generated and hand-written RTL accelerators and show that Beethoven provides on-par or better-performing systems with marginal overheads. Beethoven is open-source and available at https://github.com/Composer-Team/Beethoven.
Chris Kjellqvist, Brendan Peercy, Alvin R. Lebeck, Lisa Wu Wills
ISPASS1
2025 BigLittleMCA: A Spatially-Optimal Tiled Hardware Accelerator for MCMC Image Processing
abstract
Markov-Chain Monte-Carlo (MCMC) algorithms offer a general framework for performing interpretable inference but have high overheads due to the computational complexity of the sampling process and the large number of samples required to produce an accurate result. Computer Vision is a common class of workloads that can be performed using MCMC methods. As computer vision workloads trend toward high-resolution real-time inference, it becomes challenging to perform inference in contexts such as edge computing, which operates under strict power and area budgets. Previous work explores hardware techniques for efficient sampling; however, MCMC algorithms still require many samples. We reduce the overheads of Gibbs Sampling, an MCMC algorithm, using an approach we call mixed-resolution sampling. This approach uses low-resolution inference to provide a starting point for full-resolution sampling. We evaluate this approach on three important computer vision tasks: stereo matching, optical flow, and blind source separation. Mixed-resolution sampling reduces root mean square error (RMSE) by an average of 19.6% for stereo-matching tasks, 13% for optical flow tasks, and 6.3% for blind source separation relative to traditional Gibbs Sampling. To enable real-time, explainable MCMC inference under edge power constraints, we exploit the structure of mixed-resolution sampling to architect and implement a hardware-software co-designed accelerator architecture, BigLittleMCA ( Big - Little MC MC A ccelerator). BigLittleMCA is a tiled MCMC accelerator architecture that uses a small sampler for low-resolution sampling and a large sampler for full-resolution sampling. Our results show that the architecture sustains real-time 720p inference at 30 FPS (frames per second) using 48.5% less power than prior work.
Chris Kjellqvist, Lisa Wu Wills, Alvin R. Lebeck
ACM Trans. Archit. Code Optim.1
2022 SNS's not a synthesizer: a deep-learning-based synthesis predictor
abstract
The number of transistors that can fit on one monolithic chip has reached billions to tens of billions in this decade thanks to Moore's Law. With the advancement of every technology generation, the transistor counts per chip grow at a pace that brings about exponential increase in design time, including the synthesis process used to perform design space explorations. Such a long delay in obtaining synthesis results hinders an efficient chip development process, significantly impacting time-to-market. In addition, these large-scale integrated circuits tend to have larger and higher-dimension design spaces to explore, making it prohibitively expensive to obtain physical characteristics of all possible designs using traditional synthesis tools.
Ceyu Xu, Chris Kjellqvist, Lisa Wu Wills
ISCA2
2020 Safe, Fast Sharing of memcached as a Protected Library
abstract
Memcached is a widely used key-value store. It is structured as a multithreaded user-level server, accessed over socket connections by a potentially distributed collection of clients. Because socket communication is so much more expensive than a single operation on a K-V store, much of the client library is devoted to batching of requests. Batching is not always feasible, however, and the cost of communication seems particularly unfortunate when—as is often the case—clients are co-located on a single machine with the server, and have access to the same physical memory.
Chris Kjellqvist, Mohammad Hedayati, Michael L. Scott
ICPP1
2020 Understanding and optimizing persistent memory allocation
abstract
The proliferation of fast, dense, byte-addressable nonvolatile memory suggests that data might be kept in pointer-rich "in-memory" format across program runs and even process and system crashes. For full generality, such data requires dynamic memory allocation, and while the allocator could in principle be "rolled into" each data structure, it is desirable to make it a separate abstraction.
Wentao Cai 0002, Haosen Wen, H. Alan Beadle, Chris Kjellqvist, Mohammad Hedayati, Michael L. Scott
ISMM4