VLDB 2026 Research / reviewers in the wild / expert
Matthew Larsen
dblp:163/0448
· DBLP profile ↗
9ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-8157-9621ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
High-performance computing · 84% Parallel and multicore computing · 14% Performance modeling and evaluation · 2% | |
| Computer graphics and multimedia
2 papers |
Rendering · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › scientific visualization
in situ visualization |
0.9 | 2 | 2023 | A Hybrid in Situ Approach for Cost Efficient Image Database Generation · IEEE Trans. Vis. Comput. Graph. 2023 Performance modeling of in situ rendering · SC 2016 |
High-performance computing › scientific visualization › parallel visualization
parallel particle tracing |
0.9 | 1 | 2025 | Parallelize Over Data Particle Advection: Participation, Ping Pong Particles, and Overhead · IEEE Trans. Vis. Comput. Graph. 2025 |
High-performance computing
scientific visualization |
0.9 | 1 | 2025 | Parallelize Over Data Particle Advection: Participation, Ping Pong Particles, and Overhead · IEEE Trans. Vis. Comput. Graph. 2025 |
Rendering
parallel rendering |
0.7 | 1 | 2023 | A Hybrid in Situ Approach for Cost Efficient Image Database Generation · IEEE Trans. Vis. Comput. Graph. 2023 |
Rendering › volume rendering
parallel volume rendering |
0.4 | 1 | 2019 | A Scalable Hybrid Scheme for Ray-Casting of Unstructured Volume Data · IEEE Trans. Vis. Comput. Graph. 2019 |
Rendering › volume rendering
ray casting |
0.4 | 1 | 2019 | A Scalable Hybrid Scheme for Ray-Casting of Unstructured Volume Data · IEEE Trans. Vis. Comput. Graph. 2019 |
Rendering › volume rendering
unstructured grid rendering |
0.4 | 1 | 2019 | A Scalable Hybrid Scheme for Ray-Casting of Unstructured Volume Data · IEEE Trans. Vis. Comput. Graph. 2019 |
Rendering
volume rendering |
0.4 | 1 | 2019 | A Scalable Hybrid Scheme for Ray-Casting of Unstructured Volume Data · IEEE Trans. Vis. Comput. Graph. 2019 |
Parallel and multicore computing
data parallelism |
0.3 | 1 | 2025 | Parallelize Over Data Particle Advection: Participation, Ping Pong Particles, and Overhead · IEEE Trans. Vis. Comput. Graph. 2025 |
High-performance computing
distributed memory systems |
0.3 | 1 | 2025 | Parallelize Over Data Particle Advection: Participation, Ping Pong Particles, and Overhead · IEEE Trans. Vis. Comput. Graph. 2025 |
Parallel and multicore computing
load balancing |
0.1 | 1 | 2019 | A Scalable Hybrid Scheme for Ray-Casting of Unstructured Volume Data · IEEE Trans. Vis. Comput. Graph. 2019 |
Parallel and multicore computing › parallel computing
parallel rendering |
0.1 | 1 | 2019 | A Scalable Hybrid Scheme for Ray-Casting of Unstructured Volume Data · IEEE Trans. Vis. Comput. Graph. 2019 |
Performance modeling and evaluation › statistical analysis › statistical performance analysis
statistical performance modeling |
0.1 | 1 | 2016 | Performance modeling of in situ rendering · SC 2016 |
Methods — techniques the papers use, named apart from their topics
probing · 1.3cost estimation · 1.3statistical analysis · 0.9performance profiling · 0.9hybrid object-order/image-order rendering · 0.8statistical modeling · 0.2algorithmic complexity analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | In Situ Workload Estimation for Block Assignment and Duplication in Parallelization-Over-Data Particle AdvectionabstractAbstract Particle advection is a foundational algorithm for analyzing a flow field. The commonly used Parallelization‐Over‐Data (POD) strategy for particle advection can become slow and inefficient when there are unbalanced workloads, which are particularly prevalent in in situ workflows. In this work, we present an in situ workflow containing workload estimation for block assignment and duplication in a parallelization‐over‐data algorithm. With tightly coupled workload estimation and load‐balanced block assignment strategy, our workflow offers a considerable improvement over the traditional round‐robin block assignment strategy. Our experiments demonstrate that particle advection is up to 3X faster and associated workflow saves approximately 30% of execution time after adopting strategies presented in this work. Zhe Wang 0059, Kenneth Moreland, Matthew Larsen, James Kress, Hank Childs, Guan Li 0002, Guihua Shan, David Pugmire |
Comput. Graph. Forum | 3 |
| 2025 | Parallelize Over Data Particle Advection: Participation, Ping Pong Particles, and OverheadabstractParticle advection is one of the foundational algorithms for visualization and analysis and is central to understanding vector fields common to scientific simulations. Achieving efficient performance with large data in a distributed memory setting is notoriously difficult. Because of its simplicity and minimized movement of large vector field data, the Parallelize over Data (POD) algorithm has become a de facto standard. Despite its simplicity and ubiquitous usage, the scaling issues with the POD algorithm are known and have been described throughout the literature. In this paper, we describe a set of in-depth analyses of the POD algorithm that shed new light on the underlying causes for the poor performance of this algorithm. We designed a series of representative workloads to study the performance of the POD algorithm and executed them on a supercomputer while collecting timing and statistical data for analysis. we then performed two different types of analysis. In the first analysis, we introduce two novel metrics for measuring algorithmic efficiency over the course of a workload run. The second analysis was from the perspective of the particles being advected. Using particle-centric analysis, we identify that the overheads associated with particle movement between processes (not the communication itself) have a dramatic impact on the overall execution time. These overheads become particularly costly when flow features span multiple blocks, resulting in repeated particle circulation (which we term "ping pong particles") between blocks. Our findings shed important light on the underlying causes of poor performance and offer directions for future research to address these limitations. Zhe Wang 0059, Kenneth Moreland, Matthew Larsen, James Kress, Hank Childs, David Pugmire |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | A Hybrid in Situ Approach for Cost Efficient Image Database GenerationabstractThe visualization of results while the simulation is running is increasingly common in extreme scale computing environments. We present a novel approach for in situ generation of image databases to achieve cost savings on supercomputers. Our approach, a hybrid between traditional inline and in transit techniques, dynamically distributes visualization tasks between simulation nodes and visualization nodes, using probing as a basis to estimate rendering cost. Our hybrid design differs from previous works in that it creates opportunities to minimize idle time from four fundamental types of inefficiency: variability, limited scalability, overhead, and rightsizing. We demonstrate our results by comparing our method against both inline and in transit methods for a variety of configurations, including two simulation codes and a scaling study that goes above 19 K cores. Our findings show that our approach is superior in many configurations. As in situ visualization becomes increasingly ubiquitous, we believe our technique could lead to significant amounts of reclaimed cycles on supercomputers. Valentin Bruder, Matthew Larsen, Thomas Ertl, Hank Childs, Steffen Frey |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Evaluating adaptive and predictive power management strategies for optimizing visualization performance on supercomputers
Stephanie Brink, Matthew Larsen, Hank Childs, Barry Rountree |
Parallel Comput. | 2 |
| 2021 | Minimizing development costs for efficient many-core visualization using MCD3
Kenneth Moreland, Robert Maynard, David Pugmire, Abhishek Yenpure, Allison Vacanti, Matthew Larsen, Hank Childs |
Parallel Comput. | 6 |
| 2019 | Power and Performance Tradeoffs for Visualization AlgorithmsabstractOne of the biggest challenges for leading-edge supercomputers is power usage. Looking forward, power is expected to become an increasingly limited resource, so it is critical to understand the runtime behaviors of applications in this constrained environment in order to use power wisely. Within this context, we explore the tradeoffs between power and performance specifically for visualization algorithms. With respect to execution behavior under a power limit, visualization algorithms differ from traditional HPC applications, like scientific simulations, because visualization is more data intensive. This data intensive characteristic lends itself to alternative strategies regarding power usage. In this study, we focus on a representative set of visualization algorithms, and explore their power and performance characteristics as a power bound is applied. The result is a study that identifies how future research efforts can exploit the execution characteristics of visualization applications in order to optimize performance under a power bound. Stephanie Labasan, Matthew Larsen, Hank Childs, Barry Rountree |
IPDPS | 2 |
| 2019 | A Scalable Hybrid Scheme for Ray-Casting of Unstructured Volume DataabstractWe present an algorithm for parallel volume rendering that is a hybrid between classical object order and image order techniques. The algorithm operates on unstructured grids (and structured ones), and thus can deal with block boundaries interleaving in complex ways. It also deals effectively with cases that are prone to load imbalance, i.e., cases where cell sizes differ dramatically, either because of the nature of the input data, or because of the effects of the camera transformation. The algorithm divides work over resources such that each phase of its processing is bounded in the amount of computation it can perform. We demonstrate its efficacy through a series of studies, varying over camera position, data set size, transfer function, image size, and processor count. At its biggest, our experiments scaled up to 8,192 processors and operated on data sets with more than one billion cells. In total, we find that our hybrid algorithm performs well in all cases. This is because our algorithm naturally adapts its computation based on workload, and can operate like either an object order technique or an image order technique in scenarios where those techniques are efficient. Roba Binyahib, Tom Peterka, Matthew Larsen, Kwan-Liu Ma, Hank Childs |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2016 | Performance modeling of in situ renderingabstractWith the push to exascale, in situ visualization and analysis will continue to play an important role in high performance computing. Tightly coupling in situ visualization with simulations constrains resources for both, and these constraints force a complex balance of trade-offs. A performance model that provides an a priori answer for the cost of using an in situ approach for a given task would assist in managing the trade-offs between simulation and visualization resources. In this work, we present new statistical performance models, based on algorithmic complexity, that accurately predict the run-time cost of a set of representative rendering algorithms, an essential in situ visualization task. To train and validate the models, we conduct a performance study of an MPI+X rendering infrastructure used in situ with three HPC simulation applications. We then explore feasibility issues using the model for selected in situ rendering questions. Matthew Larsen, Cyrus Harrison, James Kress, David Pugmire, Jeremy S. Meredith, Hank Childs |
SC | 1 |
| 2015 | Ray tracing within a data parallel frameworkabstractCurrent architectural trends on supercomputers have dramatic increases in the number of cores and available computational power per die, but this power is increasingly difficult for programmers to harness effectively. High-level language constructs can simplify programming many-core devices, but this ease comes with a potential loss of processing power, particularly for cross-platform constructs. Recently, scientific visualization packages have embraced language constructs centering around data parallelism, with familiar operators such as map, reduce, gather, and scatter. Complete adoption of data parallelism will require that central visualization algorithms be revisited, and expressed in this new paradigm while preserving both functionality and performance. This investment has a large potential payoff: portable performance in software bases that can span over the many architectures that scientific visualization applications run on. With this work, we present a method for ray tracing consisting of entirely of data parallel primitives. Given the extreme computational power on nodes now prevalent on supercomputers, we believe that ray tracing can supplant rasterization as the work-horse graphics solution for scientific visualization. Our ray tracing method is relatively efficient, and we describe its performance with a series of tests, and also compare to leading-edge ray tracers that are optimized for specific platforms. We find that our data parallel approach leads to results that are acceptable for many scientific visualization use cases, with the key benefit of providing a single code base that can run on many architectures. Matthew Larsen, Jeremy S. Meredith, Paul A. Navrátil, Hank Childs |
PacificVis | 1 |