Stephen Somogyi

dblp:86/4440 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 76% Processor architecture and microarchitecture · 12% Performance modeling and evaluation · 12%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
0.222009
Spatio-temporal memory streaming · ISCA 2009
Spatial Memory Streaming · ISCA 2006
Memory systems › cache
prefetching
0.222009
Spatio-temporal memory streaming · ISCA 2009
Spatial Memory Streaming · ISCA 2006
Memory systems › memory access optimization
memory streaming
0.112009
Spatio-temporal memory streaming · ISCA 2009
Processor architecture and microarchitecture
branch prediction
0.112008
Predictor virtualization · ASPLOS 2008
Memory systems
cache coherence
0.112005
Temporal Streaming of Shared Memory · ISCA 2005
Performance modeling and evaluation › simulation › architectural simulation
full-system simulation
0.022006
Spatial Memory Streaming · ISCA 2006
Temporal Streaming of Shared Memory · ISCA 2005
Performance modeling and evaluation
simulation
0.012009
Spatio-temporal memory streaming · ISCA 2009
Memory systems
lookup table
0.012008
Predictor virtualization · ASPLOS 2008
Memory systems
on-chip memory
0.012008
Predictor virtualization · ASPLOS 2008
Performance modeling and evaluation › simulation › discrete-event simulation
trace-driven simulation
0.012005
Temporal Streaming of Shared Memory · ISCA 2005

Methods — techniques the papers use, named apart from their topics

simulation · 0.1virtualization · 0.1cycle-accurate simulation · 0.1
YearPublicationVenuePosition
2009 Spatio-temporal memory streaming
abstract
Recent research advocates memory streaming techniques to alleviate the performance bottleneck caused by the high latencies of off-chip memory accesses. Temporal memory streaming replays previously observed miss sequences to eliminate long chains of dependent misses. Spatial memory streaming predicts repetitive data layout patterns within fixed-size memory regions. Because each technique targets a different subset of misses, their effectiveness varies across workloads and each leaves a significant fraction of misses unpredicted.
Stephen Somogyi, Thomas F. Wenisch, Anastasia Ailamaki, Babak Falsafi
ISCA1
2008 Predictor virtualization
abstract
Many hardware optimizations rely on collecting information about program behavior at runtime. This information is stored in lookup tables. To be accurate and effective, these optimizations usually require large dedicated on-chip tables. Although technology advances offer an increased amount of on-chip resources, these resources are allocated to increase the size of on-chip conventional cache hierarchies.
Ioana Burcea, Stephen Somogyi, Andreas Moshovos, Babak Falsafi
ASPLOS2
2006 Spatial Memory Streaming
abstract
Prior research indicates that there is much spatial variation in applications' memory access patterns. Modern memory systems, however, use small fixed-size cache blocks and as such cannot exploit the variation. Increasing the block size would not only prohibitively increase pin and interconnect bandwidth demands, but also increase the likelihood of false sharing in shared-memory multiprocessors. In this paper, we show that memory accesses in commercial workloads often exhibit repetitive layouts that span large memory regions (e.g., several kB), and these accesses recur in patterns that are predictable through code-based correlation. We propose spatial memory streaming, a practical on-chip hardware technique that identifies code-correlated spatial access patterns and streams predicted blocks to the primary cache ahead of demand misses. Using cycle-accurate full-system multiprocessor simulation of commercial and scientific applications, we demonstrate that spatial memory streaming can on average predict 58% of LI and 65% of off-chip misses, for a mean performance improvement of 37% and at best 307%
Stephen Somogyi, Thomas F. Wenisch, Anastasia Ailamaki, Babak Falsafi, Andreas Moshovos
ISCA1
2005 Temporal Streaming of Shared Memory
abstract
Coherent read misses in shared-memory multiprocessors account for a substantial fraction of execution time in many important scientific and commercial workloads. We propose temporal streaming, to eliminate coherent read misses by streaming data to a processor in advance of the corresponding memory accesses. Temporal streaming dynamically identifies address sequences to be streamed by exploiting two common phenomena in shared-memory access patterns: (1) temporal address correlation-groups of shared addresses tend to be accessed together and in the same order; and (2) temporal stream locality-recently-accessed address streams are likely to recur. We present a practical design for temporal streaming. We evaluate our design using a combination of trace-driven and cycle-accurate full-system simulation of a cache-coherent distributed shared-memory system. We show that temporal streaming can eliminate 98% of coherent read misses in scientific applications, and between 43% and 60% in database and Web server workloads. Our design yields speedups of 1.07 to 3.29 in scientific applications, and 1.06 to 1.21 in commercial workloads.
Thomas F. Wenisch, Stephen Somogyi, Nikos Hardavellas, Jangwoo Kim, Anastasia Ailamaki, Babak Falsafi
ISCA2