VLDB 2026 Research / reviewers in the wild / expert
Dmitry Duplyakin
dblp:151/5510
· DBLP profile ↗
13ranked-venue papers
7as first author
3since 2021 · last 2026
0000-0001-5132-0168ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 6 first-author · 2 since 2021Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Performance modeling and evaluation · 51% Cloud and datacenter computing · 30% Parallel and multicore computing · 14% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% | |
| Computer networks
1 paper |
Network measurement and analytics · 100% |
Topics — the 11 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
benchmarking |
1.0 | 2 | 2023 | Avoiding the Ordering Trap in Systems Performance Measurement · USENIX ATC 2023 On Studying CPU Performance of CloudLab Hardware · ICNP 2019 |
Visualization and visual analytics
design study |
1.0 | 1 | 2026 | Same Data, Different Audiences: Using Personas to Scope a Supercomputing Job Queue Visualization · IEEE Trans. Vis. Comput. Graph. 2026 |
Performance modeling and evaluation
workload characterization |
0.4 | 2 | 2019 | Taming Performance Variability · OSDI 2018 On Studying CPU Performance of CloudLab Hardware · ICNP 2019 |
Cloud and datacenter computing › cloud infrastructure
cloud testbed |
0.4 | 1 | 2019 | On Studying CPU Performance of CloudLab Hardware · ICNP 2019 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.4 | 1 | 2019 | The Design and Operation of CloudLab · USENIX ATC 2019 |
Performance modeling and evaluation
performance variability |
0.3 | 1 | 2018 | Taming Performance Variability · OSDI 2018 |
Cloud and datacenter computing
job scheduling |
0.3 | 1 | 2026 | Same Data, Different Audiences: Using Personas to Scope a Supercomputing Job Queue Visualization · IEEE Trans. Vis. Comput. Graph. 2026 |
Parallel and multicore computing
graph partitioning |
0.3 | 1 | 2017 | Machine and Application Aware Partitioning for Adaptive Mesh Refinement Applications · HPDC 2017 |
Parallel and multicore computing
load balancing |
0.3 | 1 | 2017 | Machine and Application Aware Partitioning for Adaptive Mesh Refinement Applications · HPDC 2017 |
Distributed systems
experimental testbed |
0.1 | 1 | 2019 | The Design and Operation of CloudLab · USENIX ATC 2019 |
High-performance computing › scientific computing systems
adaptive mesh refinement |
0.1 | 1 | 2017 | Machine and Application Aware Partitioning for Adaptive Mesh Refinement Applications · HPDC 2017 |
Methods — techniques the papers use, named apart from their topics
task analysis · 2.0personas · 2.0case study · 2.0benchmarking · 0.9statistical analysis · 0.4empirical measurement · 0.4space-filling curves · 0.3architecture-aware partitioning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Same Data, Different Audiences: Using Personas to Scope a Supercomputing Job Queue VisualizationabstractDomain-specific visualizations sometimes focus on narrow, albeit important, tasks for one group of users. This focus limits the utility of a visualization to other groups working with the same data. While tasks elicited from other groups can present a design pitfall if not disambiguated, they also present a design opportunity-namely, the development of visualizations that support multiple groups. This development choice presents a trade-off of broadening the scope but limiting support for the more narrow tasks of any one group, which in some cases can enhance the overall utility of the visualization. We investigate this scenario through a design study where we develop Guidepost, a notebook-embedded visualization of data that helps scientists assess compute wait times, machine learning researchers understand prediction accuracy, and system maintainers analyze usage trends. We adapt the use of personas for visualization design from existing literature in the HCI and design domains, applying them to categorize tasks based on their uniqueness across stakeholder personas. Under this model, tasks shared between all groups should be supported by interactive visualizations and tasks unique to each group can be deferred to scripting with notebook-embedded visualization design. We evaluate our visualization through real-world case studies and a task-focused evaluation with nine participants. We observe that together, Guidepost's visual encodings, interactions, and export capabilities support the tasks of our differing personas. Connor Scully-Allison, Kevin Menear, Kristi Potter, Andrew M. McNutt, Katherine E. Isaacs, Dmitry Duplyakin |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job AnalysisabstractHPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc. Xiaoyu Chu, Daniel Hofstätter, Shashikant Ilager, Sacheendra Talluri, Duncan Kampert, Damian Podareanu, Dmitry Duplyakin, Ivona Brandic, Alexandru Iosup |
ICPADS | 7 |
| 2023 | Avoiding the Ordering Trap in Systems Performance Measurement
Dmitry Duplyakin, Nikhil Ramesh, Carina Imburgia, Hamza Fathallah Al Sheikh, Semil Jain, Prikshit Tekta, Aleksander Maricq, Gary Wong, Robert Ricci |
USENIX ATC | 1 |
| 2020 | In Datacenter Performance, The Only Constant Is ChangeabstractAll computing infrastructure suffers from performance variability, be it bare-metal or virtualized. This phenomenon originates from many sources: some transient, such as noisy neighbors, and others more permanent but sudden, such as changes or wear in hardware, changes in the underlying hypervisor stack, or even undocumented interactions between the policies of the computing resource provider and the active workloads. Thus, performance measurements obtained on clouds, HPC facilities, and, more generally, datacenter environments are almost guaranteed to exhibit performance regimes that evolve over time, which leads to undesirable nonstationarities in application performance. In this paper, we present our analysis of performance of the bare-metal hardware available on the CloudLab testbed where we focus on quantifying the evolving performance regimes using changepoint detection. We describe our findings, backed by a dataset with nearly 6.9M benchmark results collected from over 1600 machines over a period of 2 years and 9 months. These findings yield a comprehensive characterization of real-world performance variability patterns in one computing facility, a methodology for studying such patterns on other infrastructures, and contribute to a better understanding of performance variability in general. Dmitry Duplyakin, Alexandru Uta, Aleksander Maricq, Robert Ricci |
CCGRID | 1 |
| 2020 | Is Big Data Performance Reproducible in Modern Cloud Networks?
Alexandru Uta, Alexandru Custura, Dmitry Duplyakin, Ivo Jimenez, Jan S. Rellermeyer, Carlos Maltzahn, Robert Ricci, Alexandru Iosup |
NSDI | 3 |
| 2020 | 3rd Workshop on Hot Topics in Cloud Computing Performance (HotCloudPerf'20): Performance VariabilityabstractNo abstract available. Alexandru Uta, Dmitry Duplyakin, Cristina L. Abad, Nikolas Herbst, Alexandru Iosup |
ICPE | 2 |
| 2019 | On Studying CPU Performance of CloudLab HardwareabstractEmpirical performance measurements of computer systems almost always exhibit variability and anomalies. Run-to-run and server-to-server variations are common for CPU, memory, disk, and network performance characteristics. In our previous work, we focused on taming performance variability for memory, disk, and network [1] and established an interactive analysis service at: https://confirm.fyi/ to help users of the CloudLab testbed [2] better plan and conduct their experiments. In this paper, we describe our analysis of CPU variability based on over 1.3M performance measurements from nearly 1,800 servers and present our initial findings. Dmitry Duplyakin, Alexandru Uta, Aleksander Maricq, Robert Ricci |
ICNP | 1 |
| 2019 | The Design and Operation of CloudLab
Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson 0004, Kirk Webb, Aditya Akella, Kuang-Ching Wang, Glenn Ricart, Lawrence H. Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, Prabodh Mishra |
USENIX ATC | 1 |
| 2018 | Evaluating Active Learning with Cost and Memory AwarenessabstractActive Learning (AL) is a methodology from Machine Learning and Design of Experiments (DOE) in which the quantities of interest are measured sequentially and the corresponding surrogate models are constructed incrementally. AL provides compelling optimizations over static DOE in applications with engineering processes where the cost of individual experiments is significant. It also helps perform series of computer experiments in parameter sweeps and performance analysis studies. One of the non-trivial tasks in the design of AL systems is the selection of algorithms for cost-efficient exploration of the input spaces of interest: AL needs to balance ""exploitation"" of experiments with modest costs and careful ""exploration"" of expensive configurations. Finding this balance in an automatic and general manner is challenging yet desirable in practice. In this paper, we investigate the application of AL algorithms to Adaptive Mesh Refinement (AMR) performed on a supercomputer. We use AL in conjunction with Gaussian Process Regression for the incremental modeling of cost and memory usage of a series of AMR simulations of a shock-bubble interaction phenomenon. In the studied 5-dimensional input parameter space - with physical, numerical, and machine parameters - we allow AL to guide experimentation across hundreds of configurations. We develop and evaluate a novel multi-objective AL experiment selection algorithm which prioritizes cost-efficient exploration of available configurations and at the same time avoids simulations that violate memory constraints. Dmitry Duplyakin, Jed Brown, Donna Calhoun |
IPDPS | 1 |
| 2018 | Taming Performance Variability
Aleksander Maricq, Dmitry Duplyakin, Ivo Jimenez, Carlos Maltzahn, Ryan Stutsman, Robert Ricci, Ana Klimovic |
OSDI | 2 |
| 2017 | Machine and Application Aware Partitioning for Adaptive Mesh Refinement ApplicationsabstractLoad balancing and partitioning are critical when it comes to parallel computations. Popular partitioning strategies based on space filling curves focus on equally dividing work. The partitions produced are independent of the architecture or the application. Given the ever-increasing relative cost of data movement and increasing heterogeneity of our architectures, it is no longer sufficient to only consider an equal partitioning of work. Minimizing communication costs are equally if not more important. Our hypothesis is that an unequal partitioning that minimizes communication costs significantly can scale and perform better than conventional equal-work partitioning schemes. This tradeoff is dependent on the architecture as well as the application. We validate our hypothesis in the context of a finite-element computation utilizing adaptive mesh-refinement. Our central contribution is a new partitioning scheme that minimizes the overall runtime of subsequent computations by performing architecture and application-aware non-uniform work assignment in order to decrease time to solution, primarily by minimizing data-movement. We evaluate our algorithm by comparing it against standard space-filling curve based partitioning algorithms and observing time-to-solution as well as energy-to-solution for solving Finite Element computations on adaptively refined meshes. We demonstrate excellent scalability of our new partition algorithm up to $262,144$ cores on ORNL's Titan and demonstrate that the proposed partitioning scheme reduces overall energy as well as time-to-solution for application codes by up to 22.0% Milinda Fernando, Dmitry Duplyakin, Hari Sundar |
HPDC | 2 |
| 2016 | Active Learning in Performance AnalysisabstractActive Learning (AL) is a methodology from machine learning in which the learner interacts with the data source. In this paper, we investigate application of AL techniques to a new domain: regression problems in performance analysis. For computational systems with many factors, each of which can take on many levels, fixed experiment designs can require many experiments, and can explore the problem space inefficiently. We address these problems with a dynamic, adaptive experiment design, using AL in conjunction with Gaussian Process Regression (GPR). The performance analysis process is "seeded" with a small number of initial experiments, then GPR provides estimates of regression confidence across the full input space. AL is used to suggest follow-up experiments to run, in general, it will suggest experiments in areas where the GRP model indicates low confidence, and through repeated experiments, the process eventually achieves high confidence throughout the input space. We apply this approach to the problem of estimating performance and energy usage of HPGMG-FE, and create good-quality predictive models for the quantities of interest, with low error and reduced cost, using only a modest number of experiments. Our analysis shows that the error reduction achieved from replacing the basic AL algorithm with a cost-aware algorithm can be significant, reaching up to 38% for the same computational cost of experiments. Dmitry Duplyakin, Jed Brown, Robert Ricci |
CLUSTER | 1 |
| 2015 | Highly Available Cloud-Based Cluster ManagementabstractWe present an architecture that increases persistence and reliability of automated infrastructure management in the context of hybrid, cluster-cloud environments. We describe our highly available implementation that builds upon Chef configuration management system and infrastructure-as-a-service cloud resources from Amazon Web Services. We summarize our experience with managing a 20-node Linux cluster using this implementation. Our analysis of utilization and cost of necessary cloud resources indicates that the designed system is a low-cost alternative to acquiring additional physical hardware for hardening cluster management. Dmitry Duplyakin, Matthew Haney, Henry M. Tufo |
CCGRID | 1 |