Sam Likun Xi

dblp:135/8437 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
0since 2021 · last 2020
0009-0009-9759-1647ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-authorSoftware engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Hardware accelerators and domain-specific architectures · 28% Performance modeling and evaluation · 20% Electronic design automation · 19%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 10 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.412020
SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads · ACM Trans. Archit. Code Optim. 2020
Performance modeling and evaluation › simulation › simulation software
simulation infrastructure
0.412020
SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads · ACM Trans. Archit. Code Optim. 2020
Electronic design automation
high-level synthesis
0.312018
A modular digital VLSI flow for high-productivity SoC design · DAC 2018
Electronic design automation
physical design
0.312018
A modular digital VLSI flow for high-productivity SoC design · DAC 2018
Memory systems › memory management › memory allocation
dynamic memory allocation
0.312017
Mallacc: Accelerating Memory Allocation · ASPLOS 2017
Energy-efficient computing › power modeling
architectural-level power estimation
0.212015
Quantifying sources of error in McPAT and potential impacts on architectural studies · HPCA 2015
Energy-efficient computing
power modeling
0.212015
Quantifying sources of error in McPAT and potential impacts on architectural studies · HPCA 2015
Energy-efficient computing › power-performance modeling
power model validation
0.212015
Quantifying sources of error in McPAT and potential impacts on architectural studies · HPCA 2015
Integrated circuit design
system-on-chip
0.112018
A modular digital VLSI flow for high-productivity SoC design · DAC 2018
Cloud and datacenter computing
datacenter workloads
0.112017
Mallacc: Accelerating Memory Allocation · ASPLOS 2017

Methods — techniques the papers use, named apart from their topics

soc integration · 0.9end-to-end simulation · 0.9accelerator modeling · 0.9systemc · 0.3globally asynchronous locally synchronous clocking · 0.3in-core hardware accelerator · 0.3gem5-aladdin simulation · 0.2validation · 0.2power modeling toolchain · 0.2
YearPublicationVenuePosition
2020 SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads
abstract
In recent years, there has been tremendous advances in hardware acceleration of deep neural networks. However, most of the research has focused on optimizing accelerator microarchitecture for higher performance and energy efficiency on a per-layer basis. We find that for overall single-batch inference latency, the accelerator may only make up 25–40%, with the rest spent on data movement and in the deep learning software framework. Thus far, it has been very difficult to study end-to-end DNN performance during early stage design (before RTL is available), because there are no existing DNN frameworks that support end-to-end simulation with easy custom hardware accelerator integration. To address this gap in research infrastructure, we present SMAUG, the first DNN framework that is purpose-built for simulation of end-to-end deep learning applications. SMAUG offers researchers a wide range of capabilities for evaluating DNN workloads, from diverse network topologies to easy accelerator modeling and SoC integration. To demonstrate the power and value of SMAUG, we present case studies that show how we can optimize overall performance and energy efficiency for up to 1.8×–5× speedup over a baseline system, without changing any part of the accelerator microarchitecture, as well as show how SMAUG can tune an SoC for a camera-powered deep learning pipeline.
Sam Likun Xi, Yuan Yao 0006, Kshitij Bhardwaj, Paul N. Whatmough, Gu-Yeon Wei, David Brooks 0001
ACM Trans. Archit. Code Optim.1
2018 A modular digital VLSI flow for high-productivity SoC design
abstract
A high-productivity digital VLSI flow for designing complex SoCs is presented. The flow includes high-level synthesis tools, an object-oriented library of synthesizable SystemC and C++ components, and a modular VLSI physical design approach based on fine-grained globally asynchronous locally synchronous (GALS) clocking. The flow was demonstrated on a 16nm FinFET testchip targeting machine learning and computer vision.
Brucek Khailany, Evgeni Khmer, Rangharajan Venkatesan, Jason Clemons, Joel S. Emer, Matthew Fojtik, Alicia Klinefelter, Michael Pellauer, Nathaniel Ross Pinckney, Sophia Shao, Shreesha Srinath, Christopher Torng, Sam Likun Xi, Yanqing Zhang 0002, Brian Zimmer
DAC13
2017 Mallacc: Accelerating Memory Allocation
abstract
Recent work shows that dynamic memory allocation consumes nearly 7% of all cycles in Google datacenters. With the trend towards increased specialization of hardware, we propose Mallacc, an in-core hardware accelerator designed for broad use across a number of high-performance, modern memory allocators. The design of Mallacc is quite different from traditional throughput-oriented hardware accelerators. Because memory allocation requests tend to be very frequent, fast, and interspersed inside other application code, accelerators must be optimized for latency rather than throughput and area overheads must be kept to a bare minimum. Mallacc accelerates the three primary operations of a typical memory allocation request: size class computation, retrieval of a free memory block, and sampling of memory usage. Our results show that malloc latency can be reduced by up to 50% with a hardware cost of less than 1500 um2 of silicon area, less than 0.006% of a typical high-performance processor core.
Svilen Kanev, Sam Likun Xi, Gu-Yeon Wei, David Brooks 0001
ASPLOS2
2016 Co-designing accelerators and SoC interfaces using gem5-Aladdin
abstract
Increasing demand for power-efficient, high-performance computing has spurred a growing number and diversity of hardware accelerators in mobile and server Systems on Chip (SoCs). This paper makes the case that the co-design of the accelerator microarchitecture with the system in which it belongs is critical to balanced, efficient accelerator microarchitectures. We find that data movement and coherence management for accelerators are significant yet often unaccounted components of total accelerator runtime, resulting in misleading performance predictions and inefficient accelerator designs. To explore the design space of accelerator-system co-design, we develop gem5-Aladdin, an SoC simulator that captures dynamic interactions between accelerators and the SoC platform, and validate it to within 6% against real hardware. Our co-design studies show that the optimal energy-delay-product (EDP) of an accelerator microarchitecture can improve by up to 7.4× when system-level effects are considered compared to optimizing accelerators in isolation.
Sophia Shao, Sam Likun Xi, Vijayalakshmi Srinivasan, Gu-Yeon Wei, David Brooks 0001
MICRO2
2015 Beyond the Wall: Near-Data Processing for Databases
abstract
The continuous growth of main memory size allows modern data systems to process entire large scale datasets in memory. The increase in memory capacity, however, is not matched by proportional decrease in memory latency, causing a mismatch for in-memory processing. As a result, data movement through the memory hierarchy is now one of the main performance bottlenecks for main memory data systems. Database systems researchers have proposed several innovative solutions to minimize data movement and to make data access patterns hardware-aware. Nevertheless, all relevant rows and columns for a given query have to be moved through the memory hierarchy; hence, movement of large data sets is on the critical path.
Sam Likun Xi, Oreoluwa Babarinsa, Manos Athanassoulis, Stratos Idreos
DaMoN1
2015 Quantifying sources of error in McPAT and potential impacts on architectural studies
abstract
Architectural power modeling tools are widely used by the computer architecture community for rapid evaluations of high-level design choices and design space explorations. Currently, McPAT [31] is the de facto power model, but the literature does not yet contain a careful examination of its modeling accuracy. In addition, the issue of how greatly power modeling error can affect architectural-level studies has not been quantified before. In this work, we present the first rigorous assessment of McPAT's core power and area models with a detailed, validated power modeling toolchain used in current industrial practice. We find that McPAT's predictions can have significant error because some of the models are either incomplete, too high-level, or assume implementations of structures that differ from that of the core at hand. We demonstrate that large errors are possible when using McPAT's dynamic power estimates in the context of voltage noise and thermal hotspots, but for steady-state properties, accurately modeling leakage power is more important. Based on our analysis, we are able to provide guidelines for creating accurate McPAT models, even without access to detailed industrial power modeling tools. We conclude that in spite of its accuracy gaps, McPAT is still a very useful tool for many architectural studies, and its limitations can often be adequately addressed for a given research study of interest.
Sam Likun Xi, Hans M. Jacobson, Pradip Bose, Gu-Yeon Wei, David Brooks 0001
HPCA1
2013 Understanding the critical path in power state transition latencies
abstract
Increasing demands on datacenter computing prompts research in energy-efficient warehouse scale systems. In one approach, server activation policies invoke low-power sleep states but the power state transition latency must be small to produce effective energy savings. Chrome OS and Arch Linux require 50ms and 650ms, respectively, to enter sleep states. These states consume merely 4-6% of nominal power. By analyzing the critical path, we propose strategies for selecting hardware components and optimizing kernel resume sequences to make datacenter server activation viable. With fast transitions, server activation can provide better performance at lower energy than dynamic voltage and frequency scaling.
Sam Likun Xi, Marisabel Guevara, Jared Nelson, Patrick Pensabene, Benjamin C. Lee
ISLPED1