Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Gwendolyn Voskuilen

dblp:62/8277 · also Gwendolyn R. Voskuilen · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 3 first-author · 1 since 2021Systems, architecture and hardware · 4 · 3 first-authorComputer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Memory systems · 48% Processor architecture and microarchitecture · 39% Performance modeling and evaluation · 12%
Software engineering, system software, and programming languages
3 papers
Concurrent programming · 55% Requirements engineering and software design · 21% Debugging and program repair · 18%
Computer networks
1 paper
Internet architecture and protocols · 100%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
multicore design
0.322014
Fractal++: Closing the performance gap between fractal and conventional coherence · ISCA 2014
Timetraveler: exploiting acyclic races for optimizing memory race recording · ISCA 2010
Memory systems
cache coherence
0.212014
High-performance fractal coherence · ASPLOS 2014
Memory systems › cache coherence › cache coherence protocol
cache coherence protocol verification
0.212014
Fractal++: Closing the performance gap between fractal and conventional coherence · ISCA 2014
Memory systems › cache coherence
directory-based coherence
0.212014
High-performance fractal coherence · ASPLOS 2014
Requirements engineering and software design › inconsistency management
conflict resolution
0.212013
Wait-n-GoTM: improving HTM performance by serializing cyclic dependencies · ASPLOS 2013
Concurrent programming › transactional memory
hardware transactional memory
0.212013
Wait-n-GoTM: improving HTM performance by serializing cyclic dependencies · ASPLOS 2013
Concurrent programming
transactional memory
0.212013
Wait-n-GoTM: improving HTM performance by serializing cyclic dependencies · ASPLOS 2013
Internet architecture and protocols › packet processing › packet classification
decision-tree packet classification
0.112010
EffiCuts: optimizing packet classification for memory and throughput · SIGCOMM 2010
Internet architecture and protocols › packet processing
packet classification
0.112010
EffiCuts: optimizing packet classification for memory and throughput · SIGCOMM 2010
Debugging and program repair › record and replay
deterministic replay
0.112010
Timetraveler: exploiting acyclic races for optimizing memory race recording · ISCA 2010
Concurrent programming
memory models
0.112010
Timetraveler: exploiting acyclic races for optimizing memory race recording · ISCA 2010
Processor architecture and microarchitecture › multicore design
memory race recording
0.112010
Timetraveler: exploiting acyclic races for optimizing memory race recording · ISCA 2010
Program verification
model checking
0.112014
Fractal++: Closing the performance gap between fractal and conventional coherence · ISCA 2014
Performance modeling and evaluation › simulation › processor simulation
multicore simulation
0.112014
High-performance fractal coherence · ASPLOS 2014
Performance modeling and evaluation
simulation
0.112014
High-performance fractal coherence · ASPLOS 2014
Processor architecture and microarchitecture
chip multiprocessor
0.012013
Wait-n-GoTM: improving HTM performance by serializing cyclic dependencies · ASPLOS 2013
Debugging and program repair
fault localization
0.012010
Timetraveler: exploiting acyclic races for optimizing memory race recording · ISCA 2010

Methods — techniques the papers use, named apart from their topics

simulation · 0.6observational equivalence · 0.6fractal coherence · 0.4wait-n-go conflict resolution · 0.3hardware timestamps · 0.3time-delay buffer · 0.2post-dating · 0.2acyclic race exploitation · 0.2
YearPublicationVenuePosition
2024 SEFsim: A Statistically-Guided Fast DRAM Simulator
abstract
In academia and industry, computer architects rely heavily on performance models for design space exploration. However, performance models are now experiencing longer simulation times due to the increasing design complexity of modern computing systems. DDR memory, a critical component in a computing system, requires an accurate performance model to properly evaluate the instructions per cycle (IPC). However, a detailed DRAM simulator models each DDR event and, therefore, contributes a considerable simulation time. This paper proposes Satistically-guided Epoch-evolving Fixed-latency Simulator (SEFsim), an approximate and fast DRAM simulation model, to significantly improve the simulation speed. The key design principle of SEFsim is to statistically capture the performance model of DRAM using a large number of patterns, enabling the model to accurately predict the latency and behavior of new workloads. Based on our evaluation using a detailed memory model and 10 workloads, SEFsim captures the original model with 96.16% accuracy while speeding up the simulation by 10.3X and 8.25 % in the standalone and full system evaluations, respectively.
Debpratim Adak, Hyokeun Lee, Ben Feinberg, Gwendolyn Voskuilen, Clay Hughes, Huiyang Zhou, Amro Awad
ISPASS4
2014 High-performance fractal coherence
abstract
Bugs in cache coherence protocols can cause system failures. Despite many advances, verification runs into state explosion for even moderately-sized systems. As multicores' core counts increase, coherence verifiability continues to be a key problem. A recent proposal, called fractal coherence, avoids the state explosion problem by applying the idea of observational equivalence between a larger system and its smaller sub-systems. A fractal protocol for a larger system is verified by design if a minimal sub-system is verified completely. While fractal coherence is a significant step forward, there are two shortcomings: (1) Architectural limitation: To achieve fractal coherence's logical hierarchy, TreeFractal, the specific fractal protocol, employs a tree architecture where each miss traverses many levels up and down the tree and each level redundantly holds its sub-trees' coherence tags. (2) Protocol restrictions: TreeFractal imposes a restriction on responses to read requests that forces read requests to obtain clean blocks from the nearest sharer even if the shared L2 or L3 is faster. These limitations impose significant performance and coherence tag state overheads. In this paper, we propose architectural support for coherence protocols to achieve scalable performance and verifiability. To address the architectural limitation, we propose FlatFractal, a directory-based architecture which decouples fractal coherence's logical hierarchy from the architecture and eliminates redundant tag state. To address the protocol restriction, we propose a simple change to the protocol that, while preserving observational equivalence, allows read requests to obtain the blocks from the shared L2 or L3. Our simulations show that for 16 cores, FlatFractal performs, on average, 57% better than TreeFractal and within 3% of a conventional directory.
Gwendolyn Voskuilen, T. N. Vijaykumar
ASPLOS1
2014 Fractal++: Closing the performance gap between fractal and conventional coherence
abstract
Cache coherence protocol bugs can cause multicores to fail. Existing coherence verification approaches incur state explosion at small scales or require considerable human effort. As protocols' complexity and multicores' core counts increase, verification continues to be a challenge. Recently, researchers proposed fractal coherence which achieves scalable verification by enforcing observational equivalence between sub-systems in the coherence protocol. A larger sub-system is verified implicitly if a smaller sub-system has been verified. Unfortunately, fractal protocols suffer from two fundamental limitations: (1) indirect-communication: sub-systems cannot directly communicate and (2) partially-serial-invalidations: cores must be invalidated in a specific, serial order. These limitations disallow common performance optimizations used by conventional directory protocols: reply-forwarding where caches communicate directly and parallel invalidations. Therefore, fractal protocols lack performance scalability while directory protocols lack verification scalability. To enable both performance and verification scalability, we propose Fractal++ which employs a new class of protocol optimizations for verification-constrained architectures: decoupled-replies, contention-hints, and fully-parallel-fractal-invalidations. The first two optimizations allow reply-forwarding-like performance while the third optimization enables parallel invalidations in fractal protocols. Unlike conventional protocols, Fractal++ preserves observational equivalence and hence is scalably verifiable. In 32-core simulations of single- and four-socket systems, Fractal++ performs nearly as well as a directory protocol while providing scalable verifiability whereas the best-performing previous fractal protocol performs 8% on average and up to 26% worse with a single-socket and 12% on average and up to 34% worse with a longer-latency multi-socket system.
Gwendolyn Voskuilen, T. N. Vijaykumar
ISCA1
2013 Wait-n-GoTM: improving HTM performance by serializing cyclic dependencies
abstract
Transactional memory (TM) has been proposed to alleviate some key programmability problems in chip multiprocessors. Most TMs optimistically allow concurrent transactions, detecting read-write or write-write conflicts. Upon conflicts, existing hardware TMs (HTMs) use one of three conflict-resolution policies: (1) always-abort, (2) always-wait for some conflicting transactions to complete, or (3) always-go past conflicts and resolve acyclic conflicts at commit or abort upon cyclic dependencies. While each policy has advantages, the policies degrade performance under contention by limiting concurrency (always-abort, always-wait) or incurring late aborts due to cyclic dependencies (always-go). Thus, while always-go avoids acyclic aborts, no policy avoids cyclic aborts. We propose Wait-n-GoTM (WnGTM) to increase concurrency while avoiding cyclic aborts. We observe that most cyclic dependencies are caused by threads interleaving multiple accesses to a few heavily-read-write-shared delinquent data cache blocks. These accesses occur in code sections called cycle inducer sections (CISTs). Accordingly, we propose Wait-n-Go (WnG) conflict-resolution to avoid many cyclic aborts by predicting and serializing the CISTs. To support the WnG policy, we extend previous HTMs to (1) allow multiple readers and writers, (2) scalably identify dependencies, and (3) detect cyclic dependencies via new mechanisms, namely, conflict transactional state, order-capture, and hardware timestamps, respectively. In 16-core simulations of STAMP, WnGTM achieves average speedups of 46% for higher-contention benchmarks and 28% for all benchmarks over always-abort (TokenTM) with low-contention benchmarks remaining unchanged, compared to always-go (DATM) and always-wait (LogTM-SE), which perform worse than and 6% better than TokenTM, respectively.
Syed Ali Raza Jafri, Gwendolyn Voskuilen, T. N. Vijaykumar
ASPLOS2
2010 Timetraveler: exploiting acyclic races for optimizing memory race recording
abstract
As chip multiprocessors emerge as the prevalent microprocessor architecture, support for debugging shared-memory parallel programs becomes important. A key difficulty is the programs' nondeterministic semantics due to which replay runs of a buggy program may not reproduce the bug. The non-determinism stems from memory races where accesses from two threads, at least one of which is a write, go to the same memory location. Previous hardware schemes for memory race recording log the predecessor-successor thread ordering at memory races and enforce the same orderings in the replay run to achieve deterministic replay. To reduce the log size, the schemes exploit transitivity in the orderings to avoid recording redundant orderings. To reduce the log size further while requiring minimal hardware, we propose Timetraveler which for the first time exploits acyclicity of races based on the key observation that an acyclic race need not be recorded even if the race is not covered already by transitivity. Timetraveler employs a novel and elegant mechanism called post-dating which both ensures that acyclic races, including those through the L2, are eventually ordered correctly, and identifies cyclic races. To address false cycles through the L2, Timetraveler employs another novel mechanism called time-delay buffer which delays the advancement of the L2 banks' timestamps and thereby reduces the false cycles. Using simulations, we show that Timetraveler reduces the log size for commercial workloads by 88% over the best previous approach while using only a 696-byte time-delay buffer.
Gwendolyn Voskuilen, Faraz Ahmad, T. N. Vijaykumar
ISCA1
2010 EffiCuts: optimizing packet classification for memory and throughput
abstract
Packet Classification is a key functionality provided by modern routers. Previous decision-tree algorithms, HiCuts and HyperCuts, cut the multi-dimensional rule space to separate a classifier's rules. Despite their optimizations, the algorithms incur considerable memory overhead due to two issues: (1) Many rules in a classifier overlap and the overlapping rules vary vastly in size, causing the algorithms' fine cuts for separating the small rules to replicate the large rules. (2) Because a classifier's rule-space density varies significantly, the algorithms' equi-sized cuts for separating the dense parts needlessly partition the sparse parts, resulting in many ineffectual nodes that hold only a few rules. We propose EffiCuts which employs four novel ideas: (1) Separable trees: To eliminate overlap among small and large rules, we separate all small and large rules. We define a subset of rules to be separable if all the rules are either small or large in each dimension. We build a distinct tree for each such subset where each dimension can be cut coarsely to separate the large rules, or finely to separate the small rules without incurring replication. (2) Selective tree merging: To reduce the multiple trees' extra accesses which degrade throughput, we selectively merge separable trees mixing rules that may be small or large in at most one dimension. (3) Equi-dense cuts: We employ unequal cuts which distribute a node's rules evenly among the children, avoiding ineffectual nodes at the cost of a small processing overhead in the tree traversal. (4) Node Co-location: To achieve fewer accesses per node than HiCuts and HyperCuts, we co-locate parts of a node and its children. Using ClassBench, we show that for similar throughput EffiCuts needs factors of 57 less memory than HyperCuts and of 4-8 less power than TCAM.
Balajee Vamanan, Gwendolyn Voskuilen, T. N. Vijaykumar
SIGCOMM2