Brian C. Schwedock

dblp:198/2772 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0002-1754-3728ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 73% Processor architecture and microarchitecture · 13% Cloud and datacenter computing · 7%
Network and information security
1 paper
Hardware security and side channels · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
dataflow architecture
0.812024
The TYR Dataflow Architecture: Improving Locality by Taming Parallelism · MICRO 2024
Memory systems › processing-in-memory
near-cache computing
0.812024
Leviathan: A Unified System for General-Purpose Near-Data Computing · MICRO 2024
Memory systems › processing-in-memory
near-data processing
0.812024
Leviathan: A Unified System for General-Purpose Near-Data Computing · MICRO 2024
Memory systems
processing-in-memory
0.812024
Leviathan: A Unified System for General-Purpose Near-Data Computing · MICRO 2024
Memory systems › memory hierarchy
cache hierarchy
0.612022
täk¯: a polymorphic cache hierarchy for general-purpose optimization of data movement · ISCA 2022
Memory systems › memory access optimization
data movement reduction
0.612022
täk¯: a polymorphic cache hierarchy for general-purpose optimization of data movement · ISCA 2022
Memory systems
cache design
0.412020
Jumanji: The Case for Dynamic NUCA in the Datacenter · MICRO 2020
Memory systems › cache design
non-uniform cache architecture
0.412020
Jumanji: The Case for Dynamic NUCA in the Datacenter · MICRO 2020
Cloud and datacenter computing › quality of service
tail latency
0.412020
Jumanji: The Case for Dynamic NUCA in the Datacenter · MICRO 2020
Parallel and multicore computing
programming models
0.212024
Leviathan: A Unified System for General-Purpose Near-Data Computing · MICRO 2024
Hardware security and side channels › microarchitectural attacks
cache attacks
0.112020
Jumanji: The Case for Dynamic NUCA in the Datacenter · MICRO 2020

Methods — techniques the papers use, named apart from their topics

dynamic NUCA · 0.9cache partitioning · 0.9simulation · 0.8reactive programming interface · 0.8
YearPublicationVenuePosition
2024 The TYR Dataflow Architecture: Improving Locality by Taming Parallelism
abstract
Architectures should aim to maximize parallelism within a machine's finite memories, but prior designs tend to extremes, either maximizing parallelism or minimizing state. In particular, prior unordered dataflow architectures suffer from a parallelism explosion that creates unbounded state, requires prohibitively large associative memories, and risks deadlock. The few architectures that successfully navigate the parallelism-state tradeoff are limited to embarrassingly parallel programs. Tyr is a new, general-purpose unordered dataflow architecture that achieves high parallelism with bounded state. The key insight is that prior unordered dataflow architectures are overly conservative, unnecessarily allocating tags from a single, global tag space. Tyr exploits program structure to break up tags into local tag spaces that operate independently. Local tag spaces eliminate tag competition between co-dependent parts of the program, provably guaranteeing forward progress with only two tags per local tag space. Tyr thus opens the door to an efficient, scalable implementation of unordered dataflow. Simulation of parallel programs demonstrates that Tyr achieves parallelism nearly identical to a naïve unordered dataflow architecture with orders-of-magnitude less state.
Nikhil Agarwal, Mitchell Fream, Souradip Ghosh, Brian C. Schwedock, Nathan Beckmann
MICRO4
2024 Leviathan: A Unified System for General-Purpose Near-Data Computing
abstract
The rising cost of data movement poses a significant challenge to future computing systems. The call to arms for novel data-centric systems has spawned a wave of near-data computing (NDC) architectures that move compute closer to data. Despite large benefits promised by NDC, prior designs suffer from limited applicability and difficult programming. This paper identifies the commonalities and differences across NDC designs to develop Leviathan, a unified architecture and programming interface for near-cache NDC. We build a taxonomy of NDC and identify the key dimensions as what, where, and when to compute. Leviathan provides a simple reactive-programming interface and automatically executes actions near data at the right time and place. The ability to integrate multiple NDC paradigms makes Leviathan the only general-purpose system to support a variety of specialized NDC designs. Across a range of NDC-specialized applications, Leviathan improves performance by 1.5×-3.7× and reduces energy by 22%-77% vs. a baseline multicore, while adding only ≈6% area compared to the last-level cache.
Brian C. Schwedock, Nathan Beckmann
MICRO1
2022 täk¯: a polymorphic cache hierarchy for general-purpose optimization of data movement
abstract
Current systems hide data movement from software behind the load-store interface. Software's inability to observe and respond to data movement is the root cause of many inefficiencies, including the growing fraction of execution time and energy devoted to data movement itself. Recent specialized memory-hierarchy designs prove that large data-movement savings are possible. However, these designs require custom hardware, raising a large barrier to their practical adoption.
Brian C. Schwedock, Piratach Yoovidhya, Jennifer Seibert, Nathan Beckmann
ISCA1
2020 Jumanji: The Case for Dynamic NUCA in the Datacenter
abstract
The datacenter introduces new challenges for computer systems around tail latency and security. This paper argues that dynamic NUCA techniques are a better solution to these challenges than prior cache designs. We show that dynamic NUCA designs can meet tail-latency deadlines with much less cache space than prior work, and that they also provide a natural defense against cache attacks. Unfortunately, prior dynamic NUCAs have missed these opportunities because they focus exclusively on reducing data movement.We present Jumanji, a dynamic NUCA technique designed for tail latency and security. We show that prior last-level cache designs are vulnerable to new attacks and offer imperfect performance isolation. Jumanji solves these problems while significantly improving performance of co-running batch applications. Moreover, Jumanji only requires lightweight hardware and a few simple changes to system software, similar to prior D-NUCAs. At 20 cores, Jumanji improves batch weighted speedup by 14% on average, vs. just 2% for a non-NUCA design with weaker security, and is within 2% of an idealized design.
Brian C. Schwedock, Nathan Beckmann
MICRO1