Xi Yang 0021

dblp:13/1520-21 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
1since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 3 first-author · 1 since 2021Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Performance modeling and evaluation · 51% Processor architecture and microarchitecture · 11% Parallel and multicore computing · 8%
Software engineering, system software, and programming languages
3 papers
Runtime systems and virtual machines · 56% Empirical software engineering · 30% Operating systems · 14%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation › benchmarking
benchmarking methodology
0.912025
Rethinking Java Performance Analysis · ASPLOS (1) 2025
Parallel and multicore computing › task scheduling
fine-grain scheduling
0.212016
Elfen Scheduling: Fine-Grain Principled Borrowing from Latency-Critical Workloads Using Simultaneous Multithreading · USENIX ATC 2016
Cloud and datacenter computing › request scheduling
latency-sensitive scheduling
0.212016
Elfen Scheduling: Fine-Grain Principled Borrowing from Latency-Critical Workloads Using Simultaneous Multithreading · USENIX ATC 2016
Electronic design automation › high-level synthesis
scheduling
0.212016
Elfen Scheduling: Fine-Grain Principled Borrowing from Latency-Critical Workloads Using Simultaneous Multithreading · USENIX ATC 2016
Processor architecture and microarchitecture › multithreading
simultaneous multithreading
0.212016
Elfen Scheduling: Fine-Grain Principled Borrowing from Latency-Critical Workloads Using Simultaneous Multithreading · USENIX ATC 2016
Performance modeling and evaluation › profiling
continuous profiling
0.212015
Computer performance microscopy with Shim · ISCA 2015
Performance modeling and evaluation
profiling
0.212015
Computer performance microscopy with Shim · ISCA 2015
Runtime systems and virtual machines
garbage collection
0.212013
Taking off the gloves with reference counting Immix · OOPSLA 2013
Runtime systems and virtual machines › garbage collection
reference counting
0.212013
Taking off the gloves with reference counting Immix · OOPSLA 2013
Runtime systems and virtual machines › garbage collection
tracing garbage collection
0.212013
Taking off the gloves with reference counting Immix · OOPSLA 2013
Operating systems › resource management
memory management
0.112011
Why nothing matters: the impact of zeroing · OOPSLA 2011
Performance modeling and evaluation
benchmarking
0.112011
Looking back on the language and hardware revolutions: measured power, performance, and scaling · ASPLOS 2011
Memory systems
cache management
0.112011
Why nothing matters: the impact of zeroing · OOPSLA 2011
Energy-efficient computing
power measurement
0.112011
Looking back on the language and hardware revolutions: measured power, performance, and scaling · ASPLOS 2011
Performance modeling and evaluation › performance monitoring
hardware performance counters
0.112015
Computer performance microscopy with Shim · ISCA 2015
Memory systems
memory management
0.012013
Taking off the gloves with reference counting Immix · OOPSLA 2013
Embedded and real-time systems › real-time scheduling
resource reclaiming
0.012013
Taking off the gloves with reference counting Immix · OOPSLA 2013
Processor architecture and microarchitecture
chip multiprocessor
0.012011
Looking back on the language and hardware revolutions: measured power, performance, and scaling · ASPLOS 2011
Integrated circuit design
technology scaling
0.012011
Looking back on the language and hardware revolutions: measured power, performance, and scaling · ASPLOS 2011

Methods — techniques the papers use, named apart from their topics

workload analysis · 1.7immix · 0.3simultaneous multithreading · 0.2principled borrowing · 0.2performance evaluation · 0.2measurement methodology · 0.1
YearPublicationVenuePosition
2025 Rethinking Java Performance Analysis
abstract
Representative workloads and principled methodologies are the foundation of performance analysis, which in turn provides the empirical grounding for much of the innovation in systems research. However, benchmarks are hard to maintain, methodologies are hard to develop, and our field moves fast. The tension between our fast-moving fields and their need to maintain their methodological foundations is a serious challenge. This paper explores that challenge through the lens of Java performance analysis. Lessons we draw extend to other languages and other fields of computer science.
Steve Blackburn, Zixian Cai, Xi Yang 0021, John N. Zigman
ASPLOS (1)4
2016 Elfen Scheduling: Fine-Grain Principled Borrowing from Latency-Critical Workloads Using Simultaneous Multithreading
Xi Yang 0021, Steve Blackburn, Kathryn S. McKinley
USENIX ATC1
2015 Computer performance microscopy with Shim
abstract
Developers and architects spend a lot of time trying to understand and eliminate performance problems. Unfortunately, the root causes of many problems occur at a fine granularity that existing continuous profiling and direct measurement approaches cannot observe. This paper presents the design and implementation of Shim, a continuous profiler that samples at resolutions as fine as 15 cycles; three to five orders of magnitude finer than current continuous profilers. Shim's fine-grain measurements reveal new behaviors, such as variations in instructions per cycle (IPC) within the execution of a single function. A Shim observer thread executes and samples autonomously on unutilized hardware. To sample, it reads hardware performance counters and memory locations that store software state. Shim improves its accuracy by automatically detecting and discarding samples affected by measurement skew. We measure Shim's observer effects and show how to analyze them. When on a separate core, Shim can continuously observe one software signal with a 2% overhead at a ~1200 cycle resolution. At an overhead of 61%, Shim samples one software signal on the same core with SMT at a ~15 cycle resolution. Modest hardware changes could significantly reduce overheads and add greater analytical capability to Shim. We vary prefetching and DVFS policies in case studies that show the diagnostic power of fine-grain IPC and memory bandwidth results. By repurposing existing hardware, we deliver a practical tool for fine-grain performance microscopy for developers and architects.
Xi Yang 0021, Steve Blackburn, Kathryn S. McKinley
ISCA1
2013 Taking off the gloves with reference counting Immix
abstract
Despite some clear advantages and recent advances, reference counting remains a poor cousin to high-performance tracing garbage collectors. The advantages of reference counting include a) immediacy of reclamation, b) incrementality, and c) local scope of its operations. After decades of languishing with hopelessly bad performance, recent work narrowed the gap between reference counting and the fastest tracing collectors to within 10%. Though a major advance, this gap remains a substantial barrier to adoption in performance-conscious application domains.
Rifat Shahriyar, Steve Blackburn, Xi Yang 0021, Kathryn S. McKinley
OOPSLA3
2012 Barriers reconsidered, friendlier still!
abstract
Read and write barriers mediate access to the heap allowing the collector to control and monitor mutator actions. For this reason, barriers are a powerful tool in the design of any heap management algorithm, but the prevailing wisdom is that they impose significant costs. However, changes in hardware and workloads make these costs a moving target. Here, we measure the cost of a range of useful barriers on a range of modern hardware and workloads. We confirm some old results and overturn others. We evaluate the microarchitectural sensitivity of barrier performance and the differences among benchmark suites. We also consider barriers in context, focusing on their behavior when used in combination, and investigate a known pathology and evaluate solutions. Our results show that read and write barriers have average overheads as low as 5.4% and 0.9% respectively. We find that barrier overheads are more exposed on the workload provided by the modern DaCapo benchmarks than on old SPECjvm98 benchmarks. Moreover, there are differences in barrier behavior between in-order and out-of- order machines, and their respective memory subsystems, which indicate different barrier choices for different platforms. These changing costs mean that algorithm designers need to reconsider their design choices and the nature of their resulting algorithms in order to exploit the opportunities presented by modern hardware.
Xi Yang 0021, Steve Blackburn, Daniel Frampton, Antony L. Hosking
ISMM1
2011 Looking back on the language and hardware revolutions: measured power, performance, and scaling
abstract
This paper reports and analyzes measured chip power and performance on five process technology generations executing 61 diverse benchmarks with a rigorous methodology. We measure representative Intel IA32 processors with technologies ranging from 130nm to 32nm while they execute sequential and parallel benchmarks written in native and managed languages. During this period, hardware and software changed substantially: (1) hardware vendors delivered chip multiprocessors instead of uniprocessors, and independently (2) software developers increasingly chose managed languages instead of native languages. This quantitative data reveals the extent of some known and previously unobserved hardware and software trends. Two themes emerge.
Hadi Esmaeilzadeh, Xi Yang 0021, Steve Blackburn, Kathryn S. McKinley
ASPLOS3
2011 Why nothing matters: the impact of zeroing
abstract
Memory safety defends against inadvertent and malicious misuse of memory that may compromise program correctness and security. A critical element of memory safety is zero initialization. The direct cost of zero initialization is surprisingly high: up to 12.7%, with average costs ranging from 2.7 to 4.5% on a high performance virtual machine on IA32 architectures. Zero initialization also incurs indirect costs due to its memory bandwidth demands and cache displacement effects. Existing virtual machines either: a) minimize direct costs by zeroing in large blocks, or b) minimize indirect costs by zeroing in the allocation sequence, which reduces cache displacement and bandwidth. This paper evaluates the two widely used zero initialization designs, showing that they make different tradeoffs to achieve very similar performance. Our analysis inspires three better designs: (1) bulk zeroing with cache-bypassing (non-temporal) instructions to reduce the direct and indirect zeroing costs simultaneously, (2) concurrent non-temporal bulk zeroing that exploits parallel hardware to move work off the application's critical path, and (3) adaptive zeroing, which dynamically chooses between (1) and (2) based on available hardware parallelism. The new software strategies offer speedups sometimes greater than the direct overhead, improving total performance by 3% on average. Our findings invite additional optimizations and microarchitectural support.
Xi Yang 0021, Steve Blackburn, Daniel Frampton, Jennifer B. Sartor, Kathryn S. McKinley
OOPSLA1