Dmitry Duplyakin

dblp:151/5510 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
3since 2021 · last 2026
0000-0001-5132-0168ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 6 first-author · 2 since 2021Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Performance modeling and evaluation · 51% Cloud and datacenter computing · 30% Parallel and multicore computing · 14%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%
Computer networks
1 paper
Network measurement and analytics · 100%

Topics — the 11 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
1.022023
Avoiding the Ordering Trap in Systems Performance Measurement · USENIX ATC 2023
On Studying CPU Performance of CloudLab Hardware · ICNP 2019
Visualization and visual analytics
design study
1.012026
Same Data, Different Audiences: Using Personas to Scope a Supercomputing Job Queue Visualization · IEEE Trans. Vis. Comput. Graph. 2026
Performance modeling and evaluation
workload characterization
0.422019
Taming Performance Variability · OSDI 2018
On Studying CPU Performance of CloudLab Hardware · ICNP 2019
Cloud and datacenter computing › cloud infrastructure
cloud testbed
0.412019
On Studying CPU Performance of CloudLab Hardware · ICNP 2019
Cloud and datacenter computing
cluster resource management and scheduling
0.412019
The Design and Operation of CloudLab · USENIX ATC 2019
Performance modeling and evaluation
performance variability
0.312018
Taming Performance Variability · OSDI 2018
Cloud and datacenter computing
job scheduling
0.312026
Same Data, Different Audiences: Using Personas to Scope a Supercomputing Job Queue Visualization · IEEE Trans. Vis. Comput. Graph. 2026
Parallel and multicore computing
graph partitioning
0.312017
Machine and Application Aware Partitioning for Adaptive Mesh Refinement Applications · HPDC 2017
Parallel and multicore computing
load balancing
0.312017
Machine and Application Aware Partitioning for Adaptive Mesh Refinement Applications · HPDC 2017
Distributed systems
experimental testbed
0.112019
The Design and Operation of CloudLab · USENIX ATC 2019
High-performance computing › scientific computing systems
adaptive mesh refinement
0.112017
Machine and Application Aware Partitioning for Adaptive Mesh Refinement Applications · HPDC 2017

Methods — techniques the papers use, named apart from their topics

task analysis · 2.0personas · 2.0case study · 2.0benchmarking · 0.9statistical analysis · 0.4empirical measurement · 0.4space-filling curves · 0.3architecture-aware partitioning · 0.3
YearPublicationVenuePosition
2026 Same Data, Different Audiences: Using Personas to Scope a Supercomputing Job Queue Visualization
abstract
Domain-specific visualizations sometimes focus on narrow, albeit important, tasks for one group of users. This focus limits the utility of a visualization to other groups working with the same data. While tasks elicited from other groups can present a design pitfall if not disambiguated, they also present a design opportunity-namely, the development of visualizations that support multiple groups. This development choice presents a trade-off of broadening the scope but limiting support for the more narrow tasks of any one group, which in some cases can enhance the overall utility of the visualization. We investigate this scenario through a design study where we develop Guidepost, a notebook-embedded visualization of data that helps scientists assess compute wait times, machine learning researchers understand prediction accuracy, and system maintainers analyze usage trends. We adapt the use of personas for visualization design from existing literature in the HCI and design domains, applying them to categorize tasks based on their uniqueness across stakeholder personas. Under this model, tasks shared between all groups should be supported by interactive visualizations and tasks unique to each group can be deferred to scripting with notebook-embedded visualization design. We evaluate our visualization through real-world case studies and a task-focused evaluation with nine participants. We observe that together, Guidepost's visual encodings, interactions, and export capabilities support the tasks of our differing personas.
Connor Scully-Allison, Kevin Menear, Kristi Potter, Andrew M. McNutt, Katherine E. Isaacs, Dmitry Duplyakin
IEEE Trans. Vis. Comput. Graph.6
2024 Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
abstract
HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.
Xiaoyu Chu, Daniel Hofstätter, Shashikant Ilager, Sacheendra Talluri, Duncan Kampert, Damian Podareanu, Dmitry Duplyakin, Ivona Brandic, Alexandru Iosup
ICPADS7
2023 Avoiding the Ordering Trap in Systems Performance Measurement
Dmitry Duplyakin, Nikhil Ramesh, Carina Imburgia, Hamza Fathallah Al Sheikh, Semil Jain, Prikshit Tekta, Aleksander Maricq, Gary Wong, Robert Ricci
USENIX ATC1
2020 In Datacenter Performance, The Only Constant Is Change
abstract
All computing infrastructure suffers from performance variability, be it bare-metal or virtualized. This phenomenon originates from many sources: some transient, such as noisy neighbors, and others more permanent but sudden, such as changes or wear in hardware, changes in the underlying hypervisor stack, or even undocumented interactions between the policies of the computing resource provider and the active workloads. Thus, performance measurements obtained on clouds, HPC facilities, and, more generally, datacenter environments are almost guaranteed to exhibit performance regimes that evolve over time, which leads to undesirable nonstationarities in application performance. In this paper, we present our analysis of performance of the bare-metal hardware available on the CloudLab testbed where we focus on quantifying the evolving performance regimes using changepoint detection. We describe our findings, backed by a dataset with nearly 6.9M benchmark results collected from over 1600 machines over a period of 2 years and 9 months. These findings yield a comprehensive characterization of real-world performance variability patterns in one computing facility, a methodology for studying such patterns on other infrastructures, and contribute to a better understanding of performance variability in general.
Dmitry Duplyakin, Alexandru Uta, Aleksander Maricq, Robert Ricci
CCGRID1
2020 Is Big Data Performance Reproducible in Modern Cloud Networks?
Alexandru Uta, Alexandru Custura, Dmitry Duplyakin, Ivo Jimenez, Jan S. Rellermeyer, Carlos Maltzahn, Robert Ricci, Alexandru Iosup
NSDI3
2020 3rd Workshop on Hot Topics in Cloud Computing Performance (HotCloudPerf'20): Performance Variability
abstract
No abstract available.
Alexandru Uta, Dmitry Duplyakin, Cristina L. Abad, Nikolas Herbst, Alexandru Iosup
ICPE2
2019 On Studying CPU Performance of CloudLab Hardware
abstract
Empirical performance measurements of computer systems almost always exhibit variability and anomalies. Run-to-run and server-to-server variations are common for CPU, memory, disk, and network performance characteristics. In our previous work, we focused on taming performance variability for memory, disk, and network [1] and established an interactive analysis service at: https://confirm.fyi/ to help users of the CloudLab testbed [2] better plan and conduct their experiments. In this paper, we describe our analysis of CPU variability based on over 1.3M performance measurements from nearly 1,800 servers and present our initial findings.
Dmitry Duplyakin, Alexandru Uta, Aleksander Maricq, Robert Ricci
ICNP1
2019 The Design and Operation of CloudLab
Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson 0004, Kirk Webb, Aditya Akella, Kuang-Ching Wang, Glenn Ricart, Lawrence H. Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, Prabodh Mishra
USENIX ATC1
2018 Evaluating Active Learning with Cost and Memory Awareness
abstract
Active Learning (AL) is a methodology from Machine Learning and Design of Experiments (DOE) in which the quantities of interest are measured sequentially and the corresponding surrogate models are constructed incrementally. AL provides compelling optimizations over static DOE in applications with engineering processes where the cost of individual experiments is significant. It also helps perform series of computer experiments in parameter sweeps and performance analysis studies. One of the non-trivial tasks in the design of AL systems is the selection of algorithms for cost-efficient exploration of the input spaces of interest: AL needs to balance ""exploitation"" of experiments with modest costs and careful ""exploration"" of expensive configurations. Finding this balance in an automatic and general manner is challenging yet desirable in practice. In this paper, we investigate the application of AL algorithms to Adaptive Mesh Refinement (AMR) performed on a supercomputer. We use AL in conjunction with Gaussian Process Regression for the incremental modeling of cost and memory usage of a series of AMR simulations of a shock-bubble interaction phenomenon. In the studied 5-dimensional input parameter space - with physical, numerical, and machine parameters - we allow AL to guide experimentation across hundreds of configurations. We develop and evaluate a novel multi-objective AL experiment selection algorithm which prioritizes cost-efficient exploration of available configurations and at the same time avoids simulations that violate memory constraints.
Dmitry Duplyakin, Jed Brown, Donna Calhoun
IPDPS1
2018 Taming Performance Variability
Aleksander Maricq, Dmitry Duplyakin, Ivo Jimenez, Carlos Maltzahn, Ryan Stutsman, Robert Ricci, Ana Klimovic
OSDI2
2017 Machine and Application Aware Partitioning for Adaptive Mesh Refinement Applications
abstract
Load balancing and partitioning are critical when it comes to parallel computations. Popular partitioning strategies based on space filling curves focus on equally dividing work. The partitions produced are independent of the architecture or the application. Given the ever-increasing relative cost of data movement and increasing heterogeneity of our architectures, it is no longer sufficient to only consider an equal partitioning of work. Minimizing communication costs are equally if not more important. Our hypothesis is that an unequal partitioning that minimizes communication costs significantly can scale and perform better than conventional equal-work partitioning schemes. This tradeoff is dependent on the architecture as well as the application. We validate our hypothesis in the context of a finite-element computation utilizing adaptive mesh-refinement. Our central contribution is a new partitioning scheme that minimizes the overall runtime of subsequent computations by performing architecture and application-aware non-uniform work assignment in order to decrease time to solution, primarily by minimizing data-movement. We evaluate our algorithm by comparing it against standard space-filling curve based partitioning algorithms and observing time-to-solution as well as energy-to-solution for solving Finite Element computations on adaptively refined meshes. We demonstrate excellent scalability of our new partition algorithm up to $262,144$ cores on ORNL's Titan and demonstrate that the proposed partitioning scheme reduces overall energy as well as time-to-solution for application codes by up to 22.0%
Milinda Fernando, Dmitry Duplyakin, Hari Sundar
HPDC2
2016 Active Learning in Performance Analysis
abstract
Active Learning (AL) is a methodology from machine learning in which the learner interacts with the data source. In this paper, we investigate application of AL techniques to a new domain: regression problems in performance analysis. For computational systems with many factors, each of which can take on many levels, fixed experiment designs can require many experiments, and can explore the problem space inefficiently. We address these problems with a dynamic, adaptive experiment design, using AL in conjunction with Gaussian Process Regression (GPR). The performance analysis process is "seeded" with a small number of initial experiments, then GPR provides estimates of regression confidence across the full input space. AL is used to suggest follow-up experiments to run, in general, it will suggest experiments in areas where the GRP model indicates low confidence, and through repeated experiments, the process eventually achieves high confidence throughout the input space. We apply this approach to the problem of estimating performance and energy usage of HPGMG-FE, and create good-quality predictive models for the quantities of interest, with low error and reduced cost, using only a modest number of experiments. Our analysis shows that the error reduction achieved from replacing the basic AL algorithm with a cost-aware algorithm can be significant, reaching up to 38% for the same computational cost of experiments.
Dmitry Duplyakin, Jed Brown, Robert Ricci
CLUSTER1
2015 Highly Available Cloud-Based Cluster Management
abstract
We present an architecture that increases persistence and reliability of automated infrastructure management in the context of hybrid, cluster-cloud environments. We describe our highly available implementation that builds upon Chef configuration management system and infrastructure-as-a-service cloud resources from Amazon Web Services. We summarize our experience with managing a 20-node Linux cluster using this implementation. Our analysis of utilization and cost of necessary cloud resources indicates that the designed system is a low-cost alternative to acquiring additional physical hardware for hardening cluster management.
Dmitry Duplyakin, Matthew Haney, Henry M. Tufo
CCGRID1