EDBT 2026 Demo / reviewers in the wild / expert
Shasha Wen
dblp:134/6090
· DBLP profile ↗
9ranked-venue papers
5as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Performance modeling and evaluation · 47% Memory systems · 30% Parallel and multicore computing · 18% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 54% Program analysis · 46% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program analysis
dynamic analysis |
0.6 | 2 | 2018 | Watching for Software Inefficiencies with Witch · ASPLOS 2018 REDSPY: Exploring Value Locality in Software · ASPLOS 2017 |
Compilers and program optimization › compiler optimization
redundancy elimination |
0.4 | 1 | 2019 | Redundant loads: a software inefficiency indicator · ICSE 2019 |
Compilers and program optimization › compiler optimization › redundancy elimination
redundant load elimination |
0.4 | 1 | 2019 | Redundant loads: a software inefficiency indicator · ICSE 2019 |
Memory systems › cache coherence
false sharing detection |
0.3 | 1 | 2018 | Featherlight on-the-fly false-sharing detection · PPoPP 2018 |
Performance modeling and evaluation › performance monitoring
hardware performance counters |
0.3 | 1 | 2018 | Featherlight on-the-fly false-sharing detection · PPoPP 2018 |
Program analysis › code quality analysis
redundancy detection |
0.3 | 1 | 2017 | REDSPY: Exploring Value Locality in Software · ASPLOS 2017 |
Memory systems
non-uniform memory access |
0.3 | 1 | 2017 | An Efficient Abortable-locking Protocol for Multi-level NUMA Systems · PPoPP 2017 |
Parallel and multicore computing
synchronization |
0.3 | 1 | 2017 | An Efficient Abortable-locking Protocol for Multi-level NUMA Systems · PPoPP 2017 |
Performance modeling and evaluation
profiling |
0.1 | 1 | 2019 | Redundant loads: a software inefficiency indicator · ICSE 2019 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2019 | Redundant loads: a software inefficiency indicator · ICSE 2019 |
Processor architecture and microarchitecture
performance monitoring unit |
0.1 | 1 | 2018 | Watching for Software Inefficiencies with Witch · ASPLOS 2018 |
Parallel and multicore computing › parallel programming models › message passing
MPI runtime |
0.1 | 1 | 2017 | An Efficient Abortable-locking Protocol for Multi-level NUMA Systems · PPoPP 2017 |
Methods — techniques the papers use, named apart from their topics
dynamic analysis · 1.3debug registers · 1.0whole-program profiling · 0.8sampling · 0.7hardware performance monitoring · 0.7performance monitoring unit · 0.3timeout · 0.3HMCS-T · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Redundant loads: a software inefficiency indicatorabstractModern software packages have become increasingly complex with millions of lines of code and references to many external libraries. Redundant operations are a common performance limiter in these code bases. Missed compiler optimization opportunities, inappropriate data structure and algorithm choices, and developers' inattention to performance are some common reasons for the existence of redundant operations. Developers mainly depend on compilers to eliminate redundant operations. However, compilers' static analysis often misses optimization opportunities due to ambiguities and limited analysis scope; automatic optimizations to algorithmic and data structural problems are out of scope. We develop LoadSpy, a whole-program profiler to pinpoint redundant memory load operations, which are often a symptom of many redundant operations. The strength of LoadSpy exists in identifying and quantifying redundant load operations in programs and associating the redundancies with program execution contexts and scopes to focus developers' attention on problematic code. LoadSpy works on fully optimized binaries, adopts various optimization techniques to reduce its overhead, and provides a rich graphic user interface, which make it a complete developer tool. Applying LoadSpy showed that a large fraction of redundant loads is common in modern software packages despite highest levels of automatic compiler optimizations. Guided by LoadSpy, we optimize several well-known benchmarks and real-world applications, yielding significant speedups. Pengfei Su 0001, Shasha Wen, Hailong Yang 0002, Milind Chabbi, Xu Liu 0001 |
ICSE | 2 |
| 2018 | Watching for Software Inefficiencies with WitchabstractInefficiencies abound in complex, layered software. A variety of inefficiencies show up as wasteful memory operations. Many existing tools instrument every load and store instruction to monitor memory, which significantly slows execution and consumes enormously extra memory. Our lightweight framework, Witch, samples consecutive accesses to the same memory location by exploiting two ubiquitous hardware features: the performance monitoring units (PMU) and debug registers. Witch performs no instrumentation. Hence, witchcraft---tools built atop Witch---can detect a variety of software inefficiencies while introducing negligible slowdown and insignificant memory consumption and yet maintaining accuracy comparable to exhaustive instrumentation tools. Witch allowed us to scale our analysis to a large number of code bases. Guided by witchcraft, we detected several performance problems in important code bases; eliminating these inefficiencies resulted in significant speedups. Shasha Wen, Xu Liu 0001, John Byrne, Milind Chabbi |
ASPLOS | 1 |
| 2018 | ProfDP: A Lightweight Profiler to Guide Data Placement in Heterogeneous Memory SystemsabstractNew memory technologies, such as non-volatile memory and stacked memory, have reformed the memory hierarchies in modern and emerging computer architectures. It becomes common to see memories of different types integrated into the same system, as known as heterogeneous memory. Typically, a heterogeneous memory system consists of a small fast component and a large slow component. This encourages new style of data processing and exposes developers with a new problem: given two memory types, how shall we redesign applications to benefit from this memory arrangement and decide on the efficient data placement? Existing methods perform detailed memory access pattern analysis to guide data placement. However, these methods are heavyweight and ignore the interactions between software and hardware. Shasha Wen, Ludmila Cherkasova, Felix Xiaozhu Lin, Xu Liu 0001 |
ICS | 1 |
| 2018 | Featherlight on-the-fly false-sharing detectionabstractShared-memory parallel programs routinely suffer from false sharing---a performance degradation caused by different threads accessing different variables that reside on the same CPU cacheline and at least one variable is modified. State-of-the-art tools detect false sharing via a heavyweight process of logging memory accesses and feeding the ensuing access traces to an offline cache simulator. We have developed Feather, a lightweight, on-the-fly false-sharing detection tool. Feather achieves low overhead by exploiting two hardware features ubiquitous in commodity CPUs: the performance monitoring units (PMU) and debug registers. Additionally, Feather is a first-of-its-kind tool to detect false sharing in multi-process applications that use shared memory. Feather allowed us to scale false-sharing detection to myriad codes. Feather detected several false-sharing cases in important multi-core and multi-process codes including previous PPoPP artifacts. Eliminating false sharing resulted in dramatic (up to 16x) speedups. Milind Chabbi, Shasha Wen, Xu Liu 0001 |
PPoPP | 2 |
| 2017 | REDSPY: Exploring Value Locality in SoftwareabstractComplex code bases with several layers of abstractions have abundant inefficiencies that affect the execution time. Value redundancy is a kind of inefficiency where the same values are repeatedly computed, stored, or retrieved over the course of execution. Not all redundancies can be easily detected or eliminated with compiler optimization passes due to the inherent limitations of the static analysis. Shasha Wen, Milind Chabbi, Xu Liu 0001 |
ASPLOS | 1 |
| 2017 | DR-BW: Identifying Bandwidth Contention in NUMA Architectures with Supervised LearningabstractNon-Uniform Memory Access (NUMA) architectures are widely used in mainstream multi-socket computer systems to scale memory bandwidth. Without a NUMA-aware design, programs can suffer from significant performance degradation due to inter-socket bandwidth contention. However, identifying bandwidth contention is challenging. Existing methods measure bandwidth consumption. However, consumption alone is insufficient to quantify bandwidth contention. Furthermore, existing methods diagnose bandwidth for the entire program execution, but lack the ability to associate bandwidth performance to the source code and data structures involved. To address these challenges, we propose DR-BW, a new tool based on machine learning to identify bandwidth contention in NUMA architectures and provide optimization guidance. DR-BW first trains a set of micro benchmarks and extracts useful features to identify bandwidth contention via a supervised machine learning model. Our experiments show that DR-BW achieves more than 96% accuracy. Second, DR-BW associates memory accesses that incur bandwidth contention with data objects, which provides intuitive guidance for optimization. Third, we apply DR-BW to a number of real benchmarks. Our optimization based on the insights obtained from DR-BW yields up to a 6.5× speedup in modern NUMA architectures. Hao Xu 0048, Shasha Wen, Alfredo Giménez, Todd Gamblin, Xu Liu 0001 |
IPDPS | 2 |
| 2017 | An Efficient Abortable-locking Protocol for Multi-level NUMA SystemsabstractThe popularity of Non-Uniform Memory Access (NUMA) architectures has led to numerous locality-preserving hierarchical lock designs, such as HCLH, HMCS, and cohort locks. Locality-preserving locks trade fairness for higher throughput. Hence, some instances of acquisitions can incur long latencies, which may be intolerable for certain applications. Few locks admit a waiting thread to abandon its protocol on a timeout. State-of-the-art abortable locks are not fully locality aware, introduce high overheads, and unsuitable for frequent aborts. Enhancing locality-aware locks with lightweight timeout capability is critical for their adoption. In this paper, we design and evaluate the HMCS-T lock, a Hierarchical MCS (HMCS) lock variant that admits a timeout. HMCS-T maintains the locality benefits of HMCS while ensuring aborts to be lightweight. HMCS-T offers the progress guarantee missing in most abortable queuing locks. Our evaluations show that HMCS-T offers the timeout feature at a moderate overhead over its HMCS analog. HMCS-T, used in an MPI runtime lock, mitigated the poor scalability of an MPI+OpenMP BFS code and resulted in 4.3x superior scaling. Milind Chabbi, Abdelhalim Amer, Shasha Wen, Xu Liu 0001 |
PPoPP | 3 |
| 2015 | Runtime Value Numbering: A Profiling Technique to Pinpoint Redundant ComputationsabstractRedundant computations can severely degrade performance in HPC applications. Redundant computations arise due to various causes such as developers' inattention to performance, inappropriate choice of algorithms, and inefficient code generation, among others. Aliasing, limited optimization scopes, and insensitivity to input and execution contexts act as severe deterrents to static program analysis. Furthermore, static analysis cannot quantify the benefit from redundancy elimination. Consequently, large optimization efforts may yield little or no benefit. To address these limitations, we develop a dynamic profiler to pinpoint and quantify redundant computations in an execution. Our methodology -- Runtime Value Numbering (RVN) -- is based on the classical value numbering technique but works at runtime instead of compile time. RVN works on unmodified, fully-optimized binaries. RVN provides insightful feedback about redundancies and helps developers tune their applications for high performance. Since RVN employs fine-grained instrumentation, it incurs high overhead. We apply several optimizations to reduce the profiling overhead. Guided by the feedback from RVN, we optimize four benchmarks from SPEC CPU2000/2006 suite, the Sweep3D, and NAS Multi Grid (MG). We speed up these programs up to 1.22X. RVN identifies computation redundancies that compilers failed to optimize even with profile guided optimization. Shasha Wen, Xu Liu 0001, Milind Chabbi |
PACT | 1 |
| 2012 | Measuring and Visualizing Thread Communications for Pthread ApplicationsabstractEntering the era of multi/many core processors, multithreading has been used by applications frequently to enhance performance. However, with the increasing of thread number, dynamic behaviors of thread executions become more complex as well as making performance tuning more difficult. In this paper, we present a way to analyze the performance with the communication graph which describes how threads in parallel programs communicate with each other. We obtain runtime information during the actual executions of real-world applications, generates thread interaction graph and provides multiple visualization methods to programmers as an assistance of performance-tuning. The graphs are useful for optimization of programs, optimization of scheduling and deterministic accessing analysis of shared data. Shasha Wen, Yi Liu 0013, Tao Liu 0033, Bo Li 0098, Depei Qian 0001 |
PDCAT | 1 |