Wonsun Ahn

dblp:65/4407 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
7 papers
Concurrent programming · 45% Compilers and program optimization · 21% Program analysis · 17%
Computer architecture, parallel and distributed computing, and storage systems
7 papers
Processor architecture and microarchitecture · 44% Memory systems · 31% Parallel and multicore computing · 19%

Topics — the 24 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Concurrent programming
concurrency bugs
0.212014
Dynamically detecting and tolerating IF-Condition Data Races · HPCA 2014
Concurrent programming › concurrency bug detection
data race detection
0.212014
Dynamically detecting and tolerating IF-Condition Data Races · HPCA 2014
Concurrent programming › concurrency bugs
data race tolerance
0.212014
Dynamically detecting and tolerating IF-Condition Data Races · HPCA 2014
Compilers and program optimization › dynamic optimization
dynamic language optimization
0.212014
Improving JavaScript performance by deconstructing the type system · PLDI 2014
Runtime systems and virtual machines › virtual machine implementation
javascript engine
0.212014
Improving JavaScript performance by deconstructing the type system · PLDI 2014
Program analysis › static analysis
pointer analysis
0.212013
DeAliaser: alias speculation using atomic region support · ASPLOS 2013
Parallel and multicore computing
deterministic execution
0.112012
BulkCompactor: Optimized deterministic execution via Conflict-Aware commit of atomic blocks · HPCA 2012
Processor architecture and microarchitecture
multicore design
0.112012
BulkCompactor: Optimized deterministic execution via Conflict-Aware commit of atomic blocks · HPCA 2012
Processor architecture and microarchitecture
atomic block execution
0.112010
ScalableBulk: Scalable Cache Coherence for Atomic Blocks in a Lazy Environment · MICRO 2010
Memory systems
cache coherence
0.112010
ScalableBulk: Scalable Cache Coherence for Atomic Blocks in a Lazy Environment · MICRO 2010
Memory systems › cache coherence
directory-based coherence
0.112010
ScalableBulk: Scalable Cache Coherence for Atomic Blocks in a Lazy Environment · MICRO 2010
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.112010
ScalableBulk: Scalable Cache Coherence for Atomic Blocks in a Lazy Environment · MICRO 2010
Memory systems
memory consistency
0.112009
BulkCompiler: high-performance sequential consistency through cooperative compiler and hardware support · MICRO 2009
Memory systems › memory consistency › memory consistency model
sequential consistency
0.112009
BulkCompiler: high-performance sequential consistency through cooperative compiler and hardware support · MICRO 2009
Compilers and program optimization › parallelization
thread-level speculation
0.112006
POSH: a TLS compiler that exploits program structure · PPoPP 2006
Processor architecture and microarchitecture
chip multiprocessor
0.112006
POSH: a TLS compiler that exploits program structure · PPoPP 2006
Processor architecture and microarchitecture › multithreading
speculative multithreading
0.112006
POSH: a TLS compiler that exploits program structure · PPoPP 2006
Programming languages and type systems › type systems
dynamic typing
0.112014
Improving JavaScript performance by deconstructing the type system · PLDI 2014
Concurrent programming
memory models
0.012013
DeAliaser: alias speculation using atomic region support · ASPLOS 2013
Concurrent programming
synchronization
0.012012
BulkCompactor: Optimized deterministic execution via Conflict-Aware commit of atomic blocks · HPCA 2012
Memory systems › cache coherence
cache coherence protocol
0.012010
ScalableBulk: Scalable Cache Coherence for Atomic Blocks in a Lazy Environment · MICRO 2010
Compilers and program optimization
dynamic optimization
0.012008
SoftSig: software-exposed hardware signatures for code analysis and optimization · ASPLOS 2008
Compilers and program optimization › compiler analysis
pointer disambiguation
0.012008
SoftSig: software-exposed hardware signatures for code analysis and optimization · ASPLOS 2008
Parallel and multicore computing
task partitioning
0.012006
POSH: a TLS compiler that exploits program structure · PPoPP 2006

Methods — techniques the papers use, named apart from their topics

simulation · 0.4hardware support · 0.4code transformation · 0.4speculation · 0.3runtime checks · 0.3type system deconstruction · 0.2hardware signatures · 0.2profiling · 0.1lazy conflict detection · 0.1chunk commit · 0.1
YearPublicationVenuePosition
2022 The Effect of Animations Using Real-world Analogies on Diverse Computer Systems Students
abstract
It is a challenge to engage students when teaching them abstract and complex computer systems concepts, such as buffer overflow, memory management, concurrent execution, and process synchronization. Past research has shown that interactive animation and real-life analogies make STEM concepts more approachable and help students achieve better learning outcomes. Based on these findings, we introduce interactive analogies into learning the concept of buffer overflow. More specifically, we created a dry-cleaning shop animation tool (https://scratch.mit.edu/projects/571317697/) targeting K-12 and undergraduate students. To assess the effectiveness of our tool, we are in the process of conducting a user study, in which students use our animation tool to learn about buffer overflow and take pre- and post-assessment on the concept. Our goal is to make CS learning more accessible to diverse students, regardless of their background and age.
Rachel Puckett, Wonsun Ahn, Sherif M. Khattab, Luis Oliveira 0002, Vinicius Petrucci
SIGCSE (2)3
2022 Laundry Overflow: Engaging Diverse Students in CyberSecurity using Interactive Analogies
abstract
Past research has shown interactive animations, and those that use real-life analogies in particular, can play an important role in providing the intuition required to understand Computer Science concepts. Nonetheless, the use of analogies continues to be under-explored in CS education compared to other STEM fields. To break the impasse, we aim to create and evaluate a set of interactive animations based on analogies to understand their efficacy. As our first addition, we have created an animation explaining the Buffer Overflow computer systems concept. Explaining the concept abstractly has had a track record of ineffectiveness in our department, since the concept of computer memory as a series of contiguous storage locations is so foreign to students. Instead, the animation uses the analogy of a dry-cleaning shop with a series of hangers to provide a concrete mental picture of computer memory. Students explore various dry-cleaning scenarios, in which customers drop off and pick up their laundry, to understand at their own pace when buffer overflows cause harm and when they are silently ignored. This animation: https://scratch.mit.edu/projects/571317697/ (and others) will be provided as an open educational resource to instructors to encourage the use of interactive analogies in their teaching, and to undergraduate and K-12 students.
Rachel Puckett, Wonsun Ahn, Sherif M. Khattab, Luis Oliveira 0002, Vinicius Petrucci
SIGCSE (2)3
2014 Dynamically detecting and tolerating IF-Condition Data Races
abstract
An IF-Condition Invariance Violation (ICIV) occurs when, after a thread has computed the control expression of an IF statement and while it is executing the THEN or ELSE clauses, another thread updates variables in the IF's control expression. An ICIV can be easily detected, and is likely to be a sign of a concurrency bug in the code. Typically, the ICIV is caused by a data race, which we call IF-Condition Data Race (ICR). In this paper, we analyze the data races reported in the bug databases of popular software systems and show that ICRs occur relatively often. Then, we present two techniques to handle ICRs dynamically. They rely on simple code transformations and, in one case, additional hardware help. One of them (SW-IF) detects the races, while the other (HW-IF) detects and prevents them. We evaluate SW-IF and HW-IF using a variety of applica- tions. We show that these new techniques are effective at finding new data race bugs and run with low overhead. Specifically, HW-IF finds 5 new (unreported) race bugs and SW-IF finds 3 of them. In addition, 8-threaded executions of SPLASH-2 codes show that, on average, SW-IF adds 2% execution overhead, while HW-IF adds less than 1%.
Shanxiang Qi, Abdullah Muzahid, Wonsun Ahn, Josep Torrellas
HPCA3
2014 Improving JavaScript performance by deconstructing the type system
abstract
Increased focus on JavaScript performance has resulted in vast performance improvements for many benchmarks. However, for actual code used in websites, the attained improvements often lag far behind those for popular benchmarks.
Wonsun Ahn, Jiho Choi, Thomas Shull, María Jesús Garzarán, Josep Torrellas
PLDI1
2013 DeAliaser: alias speculation using atomic region support
abstract
Alias analysis is a critical component in many compiler optimizations. A promising approach to reduce the complexity of alias analysis is to use speculation. The approach consists of performing optimizations assuming the alias relationships that are true most of the time, and repairing the code when such relationships are found not to hold through runtime checks.
Wonsun Ahn, Yuelu Duan, Josep Torrellas
ASPLOS1
2012 BulkCompactor: Optimized deterministic execution via Conflict-Aware commit of atomic blocks
abstract
Recent proposals for determinism-enforcement architectures are able to honor the dependences between threads through a commit step that often becomes a performance bottleneck. As they commit code blocks (or chunks) in a round-robin order, if one chunk gets squashed due to a conflict, its successors also observe a stall. We call this effect transitive squash delay. This paper proposes a novel, high-performance approach to deterministic execution based on Conflict-Aware commit. Rather than committing chunks in strict round-robin order, the idea is to skip those chunks with conflicts and deterministically execute them slightly later. The scheme, called BulkCompactor, largely eliminates transitive squash delay, “compacts” the chunk commits, and substantially speeds-up execution. With BulkCompactor, the squash overhead is O(N) rather than O(N2) as in round-robin. We describe BulkCompactor designs for machines with centralized or distributed commit. Finally, a simulation-based evaluation shows that BulkCompactor delivers performance comparable to nondeter-ministic systems. For example, for 32 processors, BulkCompactor incurs an average execution overhead of 22% over a nondetermin-istic system. The round-robin scheme's average overhead is 133%.
Yuelu Duan, Wonsun Ahn, Josep Torrellas
HPCA3
2010 ScalableBulk: Scalable Cache Coherence for Atomic Blocks in a Lazy Environment
abstract
Recently-proposed architectures that continuously operate on atomic blocks of instructions (also called chunks) can boost the programmability and performance of shared-memory multiprocessing. However, they must support chunk operations very efficiently. In particular, in lazy conflict-detection environments, it is key that they provide scalable chunk commits. Unfortunately, current proposals typically fail to enable maximum overlap of conflict-free chunk commits. This paper presents a novel directory-based protocol that enables highly-overlapped, scalable chunk commits. The protocol, called Scalable Bulk, builds on the previously-proposed BulkSC protocol. It introduces three general hardware primitives for scalable commit: preventing access to a set of directory entries, grouping directory modules, and initiating the commit optimistically. Our results with SPLASH-2 and PARSEC codes with up to 64 processors show that Scalable Bulk enables highly-overlapped chunk commits and delivers scalable performance. Unlike previously proposed schemes, it removes practically all commit stalls.
Xuehai Qian, Wonsun Ahn, Josep Torrellas
MICRO2
2009 BulkCompiler: high-performance sequential consistency through cooperative compiler and hardware support
abstract
A platform that supported Sequential Consistency (SC) for all codes --- not only the well-synchronized ones --- would simplify the task of programmers. Recently, several hardware architectures that support high-performance SC by committing groups of instructions at a time have been proposed. However, for a platform to support SC, it is insufficient that the hardware does; the compiler has to support SC as well.
Wonsun Ahn, Shanxiang Qi, M. Nicolaides, Josep Torrellas, Jae-Woo Lee, Samuel P. Midkiff, David C. Wong 0001
MICRO1
2008 SoftSig: software-exposed hardware signatures for code analysis and optimization
abstract
Many code analysis techniques for optimization, debugging, or parallelization need to perform runtime disambiguation of sets of addresses. Such operations can be supported efficiently and with low complexity with hardware signatures.
James Tuck 0001, Wonsun Ahn, Luis Ceze, Josep Torrellas
ASPLOS2
2006 POSH: a TLS compiler that exploits program structure
abstract
As multi-core architectures with Thread-Level Speculation (TLS) are becoming better understood, it is important to focus on TLS compilation. TLS compilers are interesting in that, while they do not need to fully prove the independence of concurrent tasks, they make choices of where and when to generate speculative tasks that are crucial to overall TLS performance.This paper presents POSH, a new, fully automated TLS compiler built on top of gcc. POSH is based on two design decisions. First, to partition the code into tasks, it leverages the code structures created by the programmer, namely subroutines and loops. Second, it uses a simple profiling pass to discard ineffective tasks. With the code generated by POSH, a simulated TLS chip multiprocessor with 4 superscalar cores delivers an average speedup of 1.30 for the SPECint 2000 applications. Moreover, an estimated 26% of this speedup is a result of the implicit data prefetching provided by squashed tasks.
Wei Liu 0014, James Tuck 0001, Luis Ceze, Wonsun Ahn, Karin Strauss, Jose Renau, Josep Torrellas
PPoPP4