EDBT 2026 Demo / reviewers in the wild / expert
Gwendolyn Voskuilen
dblp:62/8277 · also Gwendolyn R. Voskuilen
· DBLP profile ↗
6ranked-venue papers
3as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 3 first-author · 1 since 2021Systems, architecture and hardware · 4 · 3 first-authorComputer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Memory systems · 48% Processor architecture and microarchitecture · 39% Performance modeling and evaluation · 12% | |
| Software engineering, system software, and programming languages
3 papers |
Concurrent programming · 55% Requirements engineering and software design · 21% Debugging and program repair · 18% | |
| Computer networks
1 paper |
Internet architecture and protocols · 100% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
multicore design |
0.3 | 2 | 2014 | Fractal++: Closing the performance gap between fractal and conventional coherence · ISCA 2014 Timetraveler: exploiting acyclic races for optimizing memory race recording · ISCA 2010 |
Memory systems
cache coherence |
0.2 | 1 | 2014 | High-performance fractal coherence · ASPLOS 2014 |
Memory systems › cache coherence › cache coherence protocol
cache coherence protocol verification |
0.2 | 1 | 2014 | Fractal++: Closing the performance gap between fractal and conventional coherence · ISCA 2014 |
Memory systems › cache coherence
directory-based coherence |
0.2 | 1 | 2014 | High-performance fractal coherence · ASPLOS 2014 |
Requirements engineering and software design › inconsistency management
conflict resolution |
0.2 | 1 | 2013 | Wait-n-GoTM: improving HTM performance by serializing cyclic dependencies · ASPLOS 2013 |
Concurrent programming › transactional memory
hardware transactional memory |
0.2 | 1 | 2013 | Wait-n-GoTM: improving HTM performance by serializing cyclic dependencies · ASPLOS 2013 |
Concurrent programming
transactional memory |
0.2 | 1 | 2013 | Wait-n-GoTM: improving HTM performance by serializing cyclic dependencies · ASPLOS 2013 |
Internet architecture and protocols › packet processing › packet classification
decision-tree packet classification |
0.1 | 1 | 2010 | EffiCuts: optimizing packet classification for memory and throughput · SIGCOMM 2010 |
Internet architecture and protocols › packet processing
packet classification |
0.1 | 1 | 2010 | EffiCuts: optimizing packet classification for memory and throughput · SIGCOMM 2010 |
Debugging and program repair › record and replay
deterministic replay |
0.1 | 1 | 2010 | Timetraveler: exploiting acyclic races for optimizing memory race recording · ISCA 2010 |
Concurrent programming
memory models |
0.1 | 1 | 2010 | Timetraveler: exploiting acyclic races for optimizing memory race recording · ISCA 2010 |
Processor architecture and microarchitecture › multicore design
memory race recording |
0.1 | 1 | 2010 | Timetraveler: exploiting acyclic races for optimizing memory race recording · ISCA 2010 |
Program verification
model checking |
0.1 | 1 | 2014 | Fractal++: Closing the performance gap between fractal and conventional coherence · ISCA 2014 |
Performance modeling and evaluation › simulation › processor simulation
multicore simulation |
0.1 | 1 | 2014 | High-performance fractal coherence · ASPLOS 2014 |
Performance modeling and evaluation
simulation |
0.1 | 1 | 2014 | High-performance fractal coherence · ASPLOS 2014 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2013 | Wait-n-GoTM: improving HTM performance by serializing cyclic dependencies · ASPLOS 2013 |
Debugging and program repair
fault localization |
0.0 | 1 | 2010 | Timetraveler: exploiting acyclic races for optimizing memory race recording · ISCA 2010 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.6observational equivalence · 0.6fractal coherence · 0.4wait-n-go conflict resolution · 0.3hardware timestamps · 0.3time-delay buffer · 0.2post-dating · 0.2acyclic race exploitation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SEFsim: A Statistically-Guided Fast DRAM SimulatorabstractIn academia and industry, computer architects rely heavily on performance models for design space exploration. However, performance models are now experiencing longer simulation times due to the increasing design complexity of modern computing systems. DDR memory, a critical component in a computing system, requires an accurate performance model to properly evaluate the instructions per cycle (IPC). However, a detailed DRAM simulator models each DDR event and, therefore, contributes a considerable simulation time. This paper proposes Satistically-guided Epoch-evolving Fixed-latency Simulator (SEFsim), an approximate and fast DRAM simulation model, to significantly improve the simulation speed. The key design principle of SEFsim is to statistically capture the performance model of DRAM using a large number of patterns, enabling the model to accurately predict the latency and behavior of new workloads. Based on our evaluation using a detailed memory model and 10 workloads, SEFsim captures the original model with 96.16% accuracy while speeding up the simulation by 10.3X and 8.25 % in the standalone and full system evaluations, respectively. Debpratim Adak, Hyokeun Lee, Ben Feinberg, Gwendolyn Voskuilen, Clay Hughes, Huiyang Zhou, Amro Awad |
ISPASS | 4 |
| 2014 | High-performance fractal coherenceabstractBugs in cache coherence protocols can cause system failures. Despite many advances, verification runs into state explosion for even moderately-sized systems. As multicores' core counts increase, coherence verifiability continues to be a key problem. A recent proposal, called fractal coherence, avoids the state explosion problem by applying the idea of observational equivalence between a larger system and its smaller sub-systems. A fractal protocol for a larger system is verified by design if a minimal sub-system is verified completely. While fractal coherence is a significant step forward, there are two shortcomings: (1) Architectural limitation: To achieve fractal coherence's logical hierarchy, TreeFractal, the specific fractal protocol, employs a tree architecture where each miss traverses many levels up and down the tree and each level redundantly holds its sub-trees' coherence tags. (2) Protocol restrictions: TreeFractal imposes a restriction on responses to read requests that forces read requests to obtain clean blocks from the nearest sharer even if the shared L2 or L3 is faster. These limitations impose significant performance and coherence tag state overheads. In this paper, we propose architectural support for coherence protocols to achieve scalable performance and verifiability. To address the architectural limitation, we propose FlatFractal, a directory-based architecture which decouples fractal coherence's logical hierarchy from the architecture and eliminates redundant tag state. To address the protocol restriction, we propose a simple change to the protocol that, while preserving observational equivalence, allows read requests to obtain the blocks from the shared L2 or L3. Our simulations show that for 16 cores, FlatFractal performs, on average, 57% better than TreeFractal and within 3% of a conventional directory. Gwendolyn Voskuilen, T. N. Vijaykumar |
ASPLOS | 1 |
| 2014 | Fractal++: Closing the performance gap between fractal and conventional coherenceabstractCache coherence protocol bugs can cause multicores to fail. Existing coherence verification approaches incur state explosion at small scales or require considerable human effort. As protocols' complexity and multicores' core counts increase, verification continues to be a challenge. Recently, researchers proposed fractal coherence which achieves scalable verification by enforcing observational equivalence between sub-systems in the coherence protocol. A larger sub-system is verified implicitly if a smaller sub-system has been verified. Unfortunately, fractal protocols suffer from two fundamental limitations: (1) indirect-communication: sub-systems cannot directly communicate and (2) partially-serial-invalidations: cores must be invalidated in a specific, serial order. These limitations disallow common performance optimizations used by conventional directory protocols: reply-forwarding where caches communicate directly and parallel invalidations. Therefore, fractal protocols lack performance scalability while directory protocols lack verification scalability. To enable both performance and verification scalability, we propose Fractal++ which employs a new class of protocol optimizations for verification-constrained architectures: decoupled-replies, contention-hints, and fully-parallel-fractal-invalidations. The first two optimizations allow reply-forwarding-like performance while the third optimization enables parallel invalidations in fractal protocols. Unlike conventional protocols, Fractal++ preserves observational equivalence and hence is scalably verifiable. In 32-core simulations of single- and four-socket systems, Fractal++ performs nearly as well as a directory protocol while providing scalable verifiability whereas the best-performing previous fractal protocol performs 8% on average and up to 26% worse with a single-socket and 12% on average and up to 34% worse with a longer-latency multi-socket system. Gwendolyn Voskuilen, T. N. Vijaykumar |
ISCA | 1 |
| 2013 | Wait-n-GoTM: improving HTM performance by serializing cyclic dependenciesabstractTransactional memory (TM) has been proposed to alleviate some key programmability problems in chip multiprocessors. Most TMs optimistically allow concurrent transactions, detecting read-write or write-write conflicts. Upon conflicts, existing hardware TMs (HTMs) use one of three conflict-resolution policies: (1) always-abort, (2) always-wait for some conflicting transactions to complete, or (3) always-go past conflicts and resolve acyclic conflicts at commit or abort upon cyclic dependencies. While each policy has advantages, the policies degrade performance under contention by limiting concurrency (always-abort, always-wait) or incurring late aborts due to cyclic dependencies (always-go). Thus, while always-go avoids acyclic aborts, no policy avoids cyclic aborts. We propose Wait-n-GoTM (WnGTM) to increase concurrency while avoiding cyclic aborts. We observe that most cyclic dependencies are caused by threads interleaving multiple accesses to a few heavily-read-write-shared delinquent data cache blocks. These accesses occur in code sections called cycle inducer sections (CISTs). Accordingly, we propose Wait-n-Go (WnG) conflict-resolution to avoid many cyclic aborts by predicting and serializing the CISTs. To support the WnG policy, we extend previous HTMs to (1) allow multiple readers and writers, (2) scalably identify dependencies, and (3) detect cyclic dependencies via new mechanisms, namely, conflict transactional state, order-capture, and hardware timestamps, respectively. In 16-core simulations of STAMP, WnGTM achieves average speedups of 46% for higher-contention benchmarks and 28% for all benchmarks over always-abort (TokenTM) with low-contention benchmarks remaining unchanged, compared to always-go (DATM) and always-wait (LogTM-SE), which perform worse than and 6% better than TokenTM, respectively. Syed Ali Raza Jafri, Gwendolyn Voskuilen, T. N. Vijaykumar |
ASPLOS | 2 |
| 2010 | Timetraveler: exploiting acyclic races for optimizing memory race recordingabstractAs chip multiprocessors emerge as the prevalent microprocessor architecture, support for debugging shared-memory parallel programs becomes important. A key difficulty is the programs' nondeterministic semantics due to which replay runs of a buggy program may not reproduce the bug. The non-determinism stems from memory races where accesses from two threads, at least one of which is a write, go to the same memory location. Previous hardware schemes for memory race recording log the predecessor-successor thread ordering at memory races and enforce the same orderings in the replay run to achieve deterministic replay. To reduce the log size, the schemes exploit transitivity in the orderings to avoid recording redundant orderings. To reduce the log size further while requiring minimal hardware, we propose Timetraveler which for the first time exploits acyclicity of races based on the key observation that an acyclic race need not be recorded even if the race is not covered already by transitivity. Timetraveler employs a novel and elegant mechanism called post-dating which both ensures that acyclic races, including those through the L2, are eventually ordered correctly, and identifies cyclic races. To address false cycles through the L2, Timetraveler employs another novel mechanism called time-delay buffer which delays the advancement of the L2 banks' timestamps and thereby reduces the false cycles. Using simulations, we show that Timetraveler reduces the log size for commercial workloads by 88% over the best previous approach while using only a 696-byte time-delay buffer. Gwendolyn Voskuilen, Faraz Ahmad, T. N. Vijaykumar |
ISCA | 1 |
| 2010 | EffiCuts: optimizing packet classification for memory and throughputabstractPacket Classification is a key functionality provided by modern routers. Previous decision-tree algorithms, HiCuts and HyperCuts, cut the multi-dimensional rule space to separate a classifier's rules. Despite their optimizations, the algorithms incur considerable memory overhead due to two issues: (1) Many rules in a classifier overlap and the overlapping rules vary vastly in size, causing the algorithms' fine cuts for separating the small rules to replicate the large rules. (2) Because a classifier's rule-space density varies significantly, the algorithms' equi-sized cuts for separating the dense parts needlessly partition the sparse parts, resulting in many ineffectual nodes that hold only a few rules. We propose EffiCuts which employs four novel ideas: (1) Separable trees: To eliminate overlap among small and large rules, we separate all small and large rules. We define a subset of rules to be separable if all the rules are either small or large in each dimension. We build a distinct tree for each such subset where each dimension can be cut coarsely to separate the large rules, or finely to separate the small rules without incurring replication. (2) Selective tree merging: To reduce the multiple trees' extra accesses which degrade throughput, we selectively merge separable trees mixing rules that may be small or large in at most one dimension. (3) Equi-dense cuts: We employ unequal cuts which distribute a node's rules evenly among the children, avoiding ineffectual nodes at the cost of a small processing overhead in the tree traversal. (4) Node Co-location: To achieve fewer accesses per node than HiCuts and HyperCuts, we co-locate parts of a node and its children. Using ClassBench, we show that for similar throughput EffiCuts needs factors of 57 less memory than HyperCuts and of 4-8 less power than TCAM. Balajee Vamanan, Gwendolyn Voskuilen, T. N. Vijaykumar |
SIGCOMM | 2 |