Charlie Roth

dblp:207/1829 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 46% Hardware accelerators and domain-specific architectures · 23% Cloud and datacenter computing · 23%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
data movement
0.312017
A many-core architecture for in-memory data processing · MICRO 2017
Cloud and datacenter computing
data movement acceleration
0.312017
A many-core architecture for in-memory data processing · MICRO 2017
Hardware accelerators and domain-specific architectures › domain-specific accelerator
data processing accelerator
0.312017
A many-core architecture for in-memory data processing · MICRO 2017
Memory systems › in-memory computing
in-memory data processing
0.312017
A many-core architecture for in-memory data processing · MICRO 2017
Processor architecture and microarchitecture
many-core architecture
0.112017
A many-core architecture for in-memory data processing · MICRO 2017

Methods — techniques the papers use, named apart from their topics

hardware RPC · 0.3
YearPublicationVenuePosition
2017 A many-core architecture for in-memory data processing
abstract
For many years, the highest energy cost in processing has been data movement rather than computation, and energy is the limiting factor in processor design [21]. As the data needed for a single application grows to exabytes [56], there is clearly an opportunity to design a bandwidth-optimized architecture for big data computation by specializing hardware for data movement. We present the Data Processing Unit or DPU, a shared memory many-core that is specifically designed for high bandwidth analytics workloads. The DPU contains a unique Data Movement System (DMS), which provides hardware acceleration for data movement and partitioning operations at the memory controller that is sufficient to keep up with DDR bandwidth. The DPU also provides acceleration for core to core communication via a unique hardware RPC mechanism called the Atomic Transaction Engine. Comparison of a DPU chip fabricated in 40nm with a Xeon processor on a variety of data processing applications shows a 3× - 15× performance per watt advantage.
Sandeep R. Agrawal, Sam Idicula, Arun Raghavan, Evangelos Vlachos, Venkatraman Govindaraju, Venkatanathan Varadarajan, Cagri Balkesen, Georgios Giannikis, Charlie Roth, Nipun Agarwal, Eric Sedlar
MICRO9