Jimmy Su

dblp:25/3293 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
0since 2021 · last 2012
0000-0002-7088-0176ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 48% High-performance computing · 48% Performance modeling and evaluation · 5%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%
Software engineering, system software, and programming languages
1 paper
Concurrent programming · 67% Compilers and program optimization · 33%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel architecture
multicore optimization
0.112012
Optimization of Parallel Particle-to-Grid Interpolation on Leading Multicore Platforms · IEEE Trans. Parallel Distributed Syst. 2012
High-performance computing
scientific computing
0.112012
Optimization of Parallel Particle-to-Grid Interpolation on Leading Multicore Platforms · IEEE Trans. Parallel Distributed Syst. 2012
Computational science and engineering › numerical analysis
adaptive mesh refinement
0.112007
An adaptive mesh refinement benchmark for modern parallel programming languages · SC 2007
Computational science and engineering › numerical analysis
finite difference methods
0.112007
An adaptive mesh refinement benchmark for modern parallel programming languages · SC 2007
Computational science and engineering
numerical analysis
0.112007
An adaptive mesh refinement benchmark for modern parallel programming languages · SC 2007
High-performance computing › scientific computing systems
adaptive mesh refinement
0.112007
An adaptive mesh refinement benchmark for modern parallel programming languages · SC 2007
Parallel and multicore computing › parallel computing
parallel programming languages
0.112007
An adaptive mesh refinement benchmark for modern parallel programming languages · SC 2007
Compilers and program optimization
compiler analysis
0.112005
Making Sequential Consistency Practical in Titanium · SC 2005
Concurrent programming
memory models
0.112005
Making Sequential Consistency Practical in Titanium · SC 2005
Concurrent programming › memory models
sequential consistency
0.112005
Making Sequential Consistency Practical in Titanium · SC 2005
Performance modeling and evaluation
benchmarking
0.012007
An adaptive mesh refinement benchmark for modern parallel programming languages · SC 2007

Methods — techniques the papers use, named apart from their topics

low-level optimization · 0.1finite difference · 0.1elliptic partial differential equation solver · 0.1auto-tuning · 0.1adaptive mesh refinement · 0.1
YearPublicationVenuePosition
2012 Optimization of Parallel Particle-to-Grid Interpolation on Leading Multicore Platforms
abstract
We are now in the multicore revolution which is witnessing a rapid evolution of architectural designs due to power constraints and correspondingly limited microprocessor clock speeds. Understanding how to efficiently utilize these systems in the context of demanding numerical algorithms is an urgent challenge to meet the ever growing computational needs of high-end computing. In this work, we examine multicore parallel optimization of the particle-to-grid interpolation step in particle-mesh methods, an inherently complex optimization problem due to its low computation intensity, irregular data accesses, and potential fine-grained data hazards. Our evaluated kernels are derived from two important numerical computations: a biological simulation of the heart using the Immersed Boundary (IB) method, and a Gyrokinetic Particle-in-Cell (PIC)-based application for studying fusion plasma microturbulence. We develop several novel synchronization and grid decomposition schemes, as well as low-level optimization techniques to maximize performance on three modern multicore platforms: Intel's Xeon X5550 (Nehalem), AMD's Opteron 2356 (Barcelona), and Sun's UltraSparc T2+ (Niagara). Results show that our optimizations lead to significant performance improvements, achieving up to a 5.6× speedup compared to the reference parallel implementation. Our work also provides valuable insight into the design of future autotuning frameworks for particle-to-grid interpolation on next-generation systems.
Kamesh Madduri, Jimmy Su, Samuel Williams 0001, Leonid Oliker, Stéphane Ethier, Katherine A. Yelick
IEEE Trans. Parallel Distributed Syst.2
2007 An adaptive mesh refinement benchmark for modern parallel programming languages
abstract
We present an Adaptive Mesh Refinement benchmark for evaluating programmability and performance of modern parallel programming languages. Benchmarks employed today by language developing teams, originally designed for performance evaluation of computer architectures, do not fully capture the complexity of state-of-the-art computational software systems running on today’s parallel machines or to be run on the emerging ones from the multi-cores to the peta-scale High Productivity Computer Systems. This benchmark, extracted from a real application framework, presents challenges for a programming language in both expressiveness and performance. It consists of an infrastructure for finite difference calculations on block-structured adaptive meshes and a solver for elliptic Partial Differential Equations built on this infrastructure. Adaptive Mesh Refinement algorithms are challenging to implement due to the irregularity introduced by local mesh refinement. We describe those challenges posed by this benchmark through two reference implementations (C++/Fortran/MPI and Titanium) and in the context of three programming models. Categories and Subject Descriptors
Tong Wen, Jimmy Su, Phillip Colella, Katherine A. Yelick, Noel Keen
SC2
2005 Making Sequential Consistency Practical in Titanium
abstract
The memory consistency model in shared memory parallel programming controls the order in which memory operations performed by one thread may be observed by another. The most natural model for programmers is to have memory accesses appear to take effect in the order specified in the original program. Language designers have been reluctant to use this strong semantics, called sequential consistency, due to concerns over the performance of memory fence instructions and related mechanisms that guarantee order. In this paper, we provide evidence for the practicality of sequential consistency by showing that advanced compiler analysis techniques are sufficient to eliminate the need for most memory fences and enable high-level optimizations. Our analyses eliminated over 97% of the memory fences that were needed by a na¨ýve implementation, accounting for 87 to 100% of the dynamically encountered fences in all but one benchmark. The impact of the memory model and analysis on runtime performance depends on the quality of the optimizations: more aggressive optimizations are likely to be invalidated by a strong memory consistency semantics. We consider two specific optimizations pipelining of bulk memory copies and communication aggregation and scheduling for irregular accesses and show that our most aggressive analysis is able to obtain the same performance as the relaxed model when applied to two linear algebra kernels. While additional work on parallel optimizations and analyses is needed, we believe these results provide important evidence on the viability of using a simple memory consistency model without sacrificing performance.
Amir Kamil, Jimmy Su, Katherine A. Yelick
SC2
2004 Array Prefetching for Irregular Array Accesses in Titanium
abstract
Summary form only given. Compiling irregular applications, such as sparse matrix vector multiply and particle/mesh methods in a SPMD parallel language is a challenging problem. These applications contain irregular array accesses, for which the array access pattern is not known until runtime. Numerous research projects have approached this problem under the inspector executor paradigm in the last 15 years. The value added by the work described in this paper is in using performance modeling to choose the best data communication method in the inspector executor model. We explore our ideas in a compiler for Titanium, a dialect of Java designed for high performance computing. For a sparse matrix vector multiply benchmark, experimental results show that the optimized Titanium code has comparable performance to C code with MPI using the Aztec library.
Jimmy Su, Katherine A. Yelick
IPDPS1