Qingsen Wang

dblp:235/2613 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 56% Performance modeling and evaluation · 36% Parallel and multicore computing · 8%
Software engineering, system software, and programming languages
2 papers
Concurrent programming · 44% Compilers and program optimization · 44% Runtime systems and virtual machines · 13%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › dynamic optimization
profile-guided optimization
0.412019
Pinpointing performance inefficiencies in Java · ESEC/SIGSOFT FSE 2019
Concurrent programming
transactional memory
0.412019
Lightweight hardware transactional memory profiling · PPoPP 2019
Memory systems
data locality
0.412019
Featherlight Reuse-Distance Measurement · HPCA 2019
Performance modeling and evaluation
profiling
0.412019
Featherlight Reuse-Distance Measurement · HPCA 2019
Memory systems › memory referencing behavior
reuse distance
0.412019
Featherlight Reuse-Distance Measurement · HPCA 2019
Runtime systems and virtual machines › managed runtime
java runtime
0.112019
Pinpointing performance inefficiencies in Java · ESEC/SIGSOFT FSE 2019
Performance modeling and evaluation › performance monitoring
hardware performance counters
0.112019
Featherlight Reuse-Distance Measurement · HPCA 2019
Parallel and multicore computing › transactional memory
hardware transactional memory
0.112019
Lightweight hardware transactional memory profiling · PPoPP 2019

Methods — techniques the papers use, named apart from their topics

hardware debug registers · 1.1sampling-based profiling · 0.8hardware performance monitoring · 0.8hardware performance counter sampling · 0.4decision-tree model · 0.4decision tree model · 0.4
YearPublicationVenuePosition
2019 Featherlight Reuse-Distance Measurement
abstract
Data locality has a profound impact on program performance. Reuse distance-the number of distinct memory locations accessed between two consecutive accesses to the same location-is the de facto, machine-independent metric of data locality in a program. Reuse distance measurement, typically, requires exhaustive instrumentation (code or binary) to log every memory access, which results in orders of magnitude runtime slowdown and memory bloat. Such high overheads impede reuse distance tools from adoption in long-running, production applications despite their usefulness. We develop RDX, a lightweight profiling tool for characterizing reuse distance in an execution; RDX typically incurs negligible time (5%) and memory (7%) overheads. RDX performs no instrumentation whatsoever but uniquely combines hardware performance counter sampling with hardware debug registers, both available in commodity CPU processors, to produce reuse-distance histograms. RDX typically has more than 90% accuracy compared to the ground truth. With the help of RDX, we are the first to characterize memory performance of long-running SPEC CPU2017 benchmarks.
Qingsen Wang, Xu Liu 0001, Milind Chabbi
HPCA1
2019 Can we trust profiling results?: understanding and fixing the inaccuracy in modern profilers
abstract
Profilers are an indispensable component in modern software stack of data centers and supercomputers. Profilers collect detailed performance data during program execution and guide code optimization across the entire software stack. The accuracy of the profiling result proves to be vital for one to effectively gain performance insights. Unfortunately, inaccuracy may arise due to measurement techniques or hardware limits, which can waste optimization efforts.
Hao Xu 0048, Qingsen Wang, Shuang Song 0007, Lizy Kurian John, Xu Liu 0001
ICS2
2019 Lightweight hardware transactional memory profiling
abstract
Programs that use hardware transactional memory (HTM) demand sophisticated performance analysis tools when they suffer from performance losses. We have developed TxSampler---a lightweight profiler for programs that use HTM. TxSampler measures performance via sampling and provides a structured performance analysis to guide intuitive optimization with a novel decision-tree model. TxSampler computes metrics that drive the investigation process in a systematic way. It not only pinpoints hot transactions with time quantification of transactional and fallback paths, but also identifies causes of transaction aborts such as data contention, capacity overflow, false sharing, and problematic instructions. TxSampler associates metrics with full call paths that are even deeply embedded inside transactions and maps them to the program's source code. Our evaluation of more than 30 HTM benchmarks and applications shows that TxSampler incurs ~4% runtime overhead and negligible memory overhead for its insightful analyses. Guided by TxSampler, we are able to optimize several HTM programs and obtain nontrivial speedups.
Qingsen Wang, Pengfei Su 0001, Milind Chabbi, Xu Liu 0001
PPoPP1
2019 Pinpointing performance inefficiencies in Java
abstract
Many performance inefficiencies such as inappropriate choice of algorithms or data structures, developers' inattention to performance, and missed compiler optimizations show up as wasteful memory operations. Wasteful memory operations are those that produce/consume data to/from memory that may have been avoided. We present, JXPerf, a lightweight performance analysis tool for pinpointing wasteful memory operations in Java programs. Traditional byte code instrumentation for such analysis (1) introduces prohibitive overheads and (2) misses inefficiencies in machine code generation. JXPerf overcomes both of these problems. JXPerf uses hardware performance monitoring units to sample memory locations accessed by a program and uses hardware debug registers to monitor subsequent accesses to the same memory. The result is a lightweight measurement at the machine code level with attribution of inefficiencies to their provenance --- machine and source code within full calling contexts. JXPerf introduces only 7% runtime overhead and 7% memory overhead making it useful in production. Guided by JXPerf, we optimize several Java applications by improving code generation and choosing superior data structures and algorithms, which yield significant speedups.
Pengfei Su 0001, Qingsen Wang, Milind Chabbi, Xu Liu 0001
ESEC/SIGSOFT FSE2