VLDB 2026 Research / reviewers in the wild / expert
Julian M. Kunkel
dblp:61/1408 · also Julian Martin Kunkel
· DBLP profile ↗
30ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0002-6915-1179ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Architectures, Learning Loops, and Emergence in Agentic Services Computing
Sadaf Shafi, Michael Bidollahkhani, Julian M. Kunkel |
COMPSAC | 3 |
| 2026 | Ontological Knowledge Blocks: Executable Compliance and Profile-Based Validation for Trustworthy AI SystemsabstractAI-enabled services deployed in critical digital infrastructure are subject to governance obligations spanning transparency, accountability, fairness, and traceability. Compliance today remains documentation-centric: obligations are described in prose, audits rely on static checklists, and verification depends on manual review. Such approaches do not scale to automated AI systems. This paper introduces Ontological Knowledge Blocks (OKBs), a programmable governance infrastructure that compiles regulatory obligations into machine-checkable constraints over structured evidence graphs. We formalize an OKB as a 5-tuple that binds normative obligations to an RDF/OWL concept schema, executable SHACL validation rules, explicit evidence requirements, and PROV-O provenance links. A deterministic regulatory compiler translates structured Intermediate Representation (IR) records into composable KB modules, enabling profile-based governance reconfiguration without modifying service code. We implement two prototypes and evaluate them in an AI-assisted HPC resource allocation scenario across 24 validation runs and four governance profiles. Results demonstrate profile-sensitive validation, strictly additive violation accumulation, SHACL validation latency between 12.6 ms and 100.3 ms, and profile equivalence testing confirming Combined as the strictly most comprehensive profile. All artefacts are released as open source. Aasish Kumar Sharma, Julian M. Kunkel |
COMPSAC | 2 |
| 2026 | Global AI Regulations for FAIR and Ethics in High-Risk Use Cases: A Comparative ReviewabstractAI governance is shifting from voluntary ethics to enforceable, risk-based regulation, yet cross-jurisdictional divergence creates compliance uncertainty for operators of high-stakes AI. We present a comparative matrix for the EU, US, and China that maps (i) risk classification triggers, (ii) binding obligations, (iii) enforcement and accountability mechanisms, and (iv) the degree to which FAIR principles are operationalised in practice. We stress-test the matrix on three high-impact domains: Electroencephalography (EEG)-guided rehabilitation robotics, AI-enabled debt collection in prospective Central Bank Digital Currency (CBDC) ecosystems, and AI-driven allocation of scarce Graphics Processing Unit (GPU) resources in emerging AI Factory infrastructures. Using primary legal texts and implementation evidence, we identify three recurring gaps: weak interoperability mandates, difficult operationalisation of cross-regime obligations (AI + sector regulation + data protection), and under-specified governance for critical digital infrastructure use cases. To bridge the implementation gap, we outline Knowledge Blocks, a machine-checkable compliance artefact pattern based on Resource Description Framework/Web Ontology Language (RDF/OWL), Shapes Constraint Language (SHACL), and Provenance Ontology (PROV-O), enabling audit-ready compliance-by-design across multiple regimes. Aasish Kumar Sharma, Dimitar Koysev, Christopher Anich, Roshni Kumari Ojha, Julian M. Kunkel |
COMPSAC | 5 |
| 2026 | DECICE: AI-Driven Scheduling and Digital Twin Integration for the Cloud-HPC-Edge Compute ContinuumabstractThis paper presents the DECICE project (Device Edge Cloud Intelligent Collaboration framEwork), a Horizon Europe Research and Innovation Action (Grant No. 101092582, December 2022 to November 2025) that developed an open-source framework for intelligent workload scheduling across the cloud-HPC-edge compute continuum. A consortium of 12 partners across 6 European countries organized the work into six work packages covering AI-driven scheduling, digital twin infrastructure, system architecture and integration, monitoring, use case validation, and dissemination. The two core technical contributions are an Integrated AI Scheduler (IAIS) employing RNN-based prediction and formal workflow modeling for constraint-aware workload mapping, and a Digital Twin aggregating real-time metrics with carbon intensity and anomaly prediction for energy-aware scheduling. The framework operates within Kubernetes environments, supports unified workflow ingestion from multiple formats, and bridges cloud-native and HPC orchestration through a Slurm integration layer. We present the project vision, the overall architecture, contributions from each work package, quantitative evaluation results, and the open-source release. Aasish Kumar Sharma, Felix Stein, Mirac Aydin, Michael Bidollahkhani, Sachin P. Nanavati, Mohsen Seyedkazemi Ardebili, Giorgi Mamulashvili, Mojtaba Akbari, Jonathan Decker, Zoya Masih, Julian M. Kunkel |
COMPSAC | 11 |
| 2026 | Robust I/O Characterization of Machine Learning Workloads Across Performance Analysis ToolsabstractData-driven models can predict filesystem I/O time proportion, a key metric for guiding I/O tuning in HPC applications, from performance data. This work evaluates whether the XGBoost-based model we proposed in our prior work [13], remains effective when the underlying profiling tool changes. Using I/O metrics collected with Score-P and DFTracer, we analyze two video processing workloads and train predictive models. Results show that predictive accuracy remains stable across profiling tools while revealing workload-specific differences in I/O behavior, suggesting that the relationship between configuration parameters and the proportion of filesystem I/O time is preserved despite variations in measurement methodology. Furthermore, the findings indicate that configuration-driven models can approximate I/O time proportion without requiring runtime tracing for every execution, reducing dependence on instrumentation in future machine learning-driven performance analysis workflows. Zoya Masih, Radita Liem, Julian M. Kunkel |
HPDC | 3 |
| 2026 | Aligning Storage Benchmark Metrics with Application-Level PerformanceabstractThe current state of practice in HPC is that performance metrics reported by storage benchmarks are disconnected from those obtained through application-level performance analysis tools, making it difficult for users to determine whether performance tuning efforts are effective or whether observed performance indicates underutilization of the system. In this work, we propose an approach to align IO500 storage benchmark results with application-level performance by deconstructing benchmark components and recalculating their metrics. The results show that certain universal metrics, such as bandwidth, can be meaningfully aligned with application performance, enabling more consistent and interpretable evaluation. Our findings also identify metrics that remain missing or cannot be reconciled, highlighting the need for standardization to align metrics produced by benchmarks and performance analysis tools. Radita Liem, Julian M. Kunkel, Jay F. Lofstead, Sarah Neuwirth |
SSDBM | 3 |
| 2026 | SAIA: a seamless Slurm-native solution for HPC-based servicesabstractAbstract Recent developments indicate a shift toward web services that employ ever larger AI models, e.g., Large Language Models (LLMs), requiring powerful hardware for inference. High-Performance Computing (HPC) systems are commonly equipped with such hardware for the purpose of large scale computation tasks. However, HPC infrastructure is inherently unsuitable for hosting real-time web services due to network, security and scheduling constraints. While various efforts exist to integrate external scheduling solutions, these often require compromises in terms of security or usability for existing HPC users. In this paper, we present SAIA, a Slurm-native platform consisting of a scheduler and a proxy. The scheduler interacts with Slurm to ensure the availability and scalability of services, while the proxy provides external access, which is secured via confined SSH commands. We have demonstrated SAIA’s applicability by deploying a large-scale LLM web service that has served over 50,000 users. Ali Doosthosseini, Jonathan Decker, Hendrik Nolte, Julian M. Kunkel |
J. Supercomput. | 4 |
| 2025 | Workflow-Driven Modeling for the Compute Continuum: An Optimization Approach to Automated System and Workload SchedulingabstractThe convergence of IoT, edge, cloud, and HPC technologies creates a heterogeneous compute continuum requiring sophisticated workload management. Current tools like SLURM, Kubernetes, and Snakemake lack automated optimization for cross-platform resource allocation, forcing users to manually map workloads across diverse infrastructures. We present a comprehensive framework integrating heterogeneous system and workload modeling integration with Snakemake followed by different tools and techniques like Mixed Integer Linear Programming (MILP) for multi-objective optimization to automate task mapping and scheduling across the compute continuum. Our approach extends Snakemake scheduler with formal mathematical models that optimize resource utilization and makespan. Experimental evaluation demonstrates that MILP-based solution achieves optimal scheduling for small-scale workflows (5x5 tasks) in 0.02 seconds, while heuristic methods provide 99.9% faster solutions for large-scale scenarios (5000×5000 tasks) with only 5-10% deviation from optimal makespan. For parallel workflows, the optimization achieves up to 16.7% makespan reduction compared to sequential scheduling approaches. Aasish Kumar Sharma, Christian Boehme, Patrick Gelß, Ramin Yahyapour, Julian M. Kunkel |
COMPSAC | 5 |
| 2025 | Grapheon RL: A Graph Neural Network and Reinforcement Learning Framework for Constraint and Data-Aware Workflow Mapping and Scheduling in Heterogeneous HPC SystemsabstractEfficient workflow mapping and scheduling in heterogeneous HPC-Compute Continuum (HPC-CC) systems is critical for multi-objective optimization like optimizing resource utilization and minimizing makespan or energy efficiency. Existing approaches face fundamental trade-offs: Mixed-Integer Linear Programming (MILP) provides optimal solutions but becomes computationally intractable for large workflows exceeding (50x50) nodes by tasks, while heuristic methods sacrifice optimality for speed and struggle with complex constraint modeling. We present GrapheonRL, a novel Graph Neural Network (GNN) and Reinforcement Learning (RL)-based framework that can be embedded in Snakemake to model workflows as dependency-aware graphs, enabling RL agents to dynamically learn constraint-aware scheduling policies without mathematical reformulation. We evaluated GrapheonRL against MILP and heuristic baselines (HEFT, OLB) on Standard Task Graph Set workflows (11–90 tasks) and extended to synthetic workflows of up to (10,000x10,000) nodes by tasks, GrapheonRL matches MILP optimality while offering significantly improved scalability, achieving 76% faster inference with linear memory growth (0.87 MB per 1000 tasks). On complex workflows, GrapheonRL maintains optimal makespan (569) whereas heuristics degrade substantially (HEFT: 829, OLB: 1160), demonstrating that learning-based scheduling effectively bridges the optimality-scalability gap for dynamic HPC-CC environments. Aasish Kumar Sharma, Julian M. Kunkel |
COMPSAC | 2 |
| 2025 | Ethical AI: Towards Defining a Collective Evaluation FrameworkabstractArtificial Intelligence (AI) is transforming sectors such as healthcare, finance, and autonomous systems, offering powerful tools for innovation. Yet its rapid integration raises urgent ethical concerns related to data ownership, privacy, and systemic bias. Issues like opaque decision-making, misleading outputs, and unfair treatment in high-stakes domains underscore the need for transparent and accountable AI systems.This article addresses these challenges by proposing a modular ethical assessment framework built on ontological blocks of meaning—discrete, interpretable units that encode ethical principles such as fairness, accountability, and ownership. By integrating these blocks with FAIR (Findable, Accessible, Interoperable, Reusable) principles, the framework supports scalable, transparent, and legally aligned ethical evaluations, including compliance with the EU AI Act.Using a real-world use case in AI-powered investor profiling, the paper demonstrates how the framework enables dynamic, behavior-informed risk classification. The findings suggest that ontological blocks offer a promising path toward explainable and auditable AI ethics, though challenges remain in automation and probabilistic reasoning. Aasish Kumar Sharma, Dimitar Kyosev, Julian M. Kunkel |
COMPSAC | 3 |
| 2025 | Performance Analysis of Convolutional Neural Network By Applying Unconstrained Binary Quadratic ProgrammingabstractConvolutional Neural Networks (CNNs) are pivotal in computer vision and Big Data analytics but demand significant computational resources when trained on large-scale datasets. Conventional training via back-propagation (BP) with loss functions like Mean Squared Error or Cross-Entropy often requires extensive iterations and may converge sub-optimally. Quantum computing offers a promising alternative by leveraging superposition, tunneling, and entanglement to search complex optimization landscapes more efficiently.In this work, we propose a hybrid optimization method that combines an Unconstrained Binary Quadratic Programming (UBQP) formulation with Stochastic Gradient Descent (SGD) to accelerate CNN training. Evaluated on the MNIST dataset, our approach achieves improvement without compromising the accuracy while maintaining similar execution times. These results illustrate the potential of hybrid quantum-classical techniques in High-Performance Computing (HPC) environments for Big Data and Deep Learning. Fully realizing these benefits, however, requires a careful alignment of algorithmic structures with underlying quantum mechanisms. Aasish Kumar Sharma, Sanjeeb Prashad Pandey, Julian M. Kunkel |
COMPSAC | 3 |
| 2025 | Maximizing Insights, Minimizing Data: I/O Time Prediction Using Transfer Learning
Adrian Voß, Radita Liem, Julian M. Kunkel, Jay F. Lofstead, Philip H. Carns |
HiPC | 3 |
| 2025 | Factors Impacting I/O Time Proportion in AI WorkloadsabstractThe decision to optimize I/O in scientific applications often depends on the proportion of I/O time within an application's runtime. For AI workloads, such as machine learning (ML), we observe that different configurations with the same data size can lead to varying I/O impacts. In this work, we analyze common tunable parameters—batch size, number of samples, and number of files—typically adjusted when adapting ML models to new datasets or use cases. Using the XGBoost model, we predict the I/O percentage for ResNet50 and UNet3D workloads and identify feature combinations that influence I/O behavior. Our results provide practical guidance for optimizing I/O performance in AI workloads. Zoya Masih, Radita Liem, Julian M. Kunkel |
HPDC | 3 |
| 2025 | Ephemeral Kubernetes: dynamically deleting and recreating clusters using WarewulfabstractAbstract With the rise of LLMs, GPU acceleration has become essential for both training and serving AI models. This requires HPC systems to be highly flexible with assigning multi-GPU nodes while also maintaining high security standards. Existing approaches involve utilizing nodes with batch and service schedulers, e.g., Slurm and Kubernetes, by dynamically moving nodes between the schedulers either through negotiation between the systems or via an external system. However, such a multi-use approach also increases the attack surface as more scheduling components operate with root permission. Moreover, it becomes increasingly difficult to recover from a security incident as attackers might have infected parts of either scheduling system. In this work, we present Ephemeral Kubernetes as a way to dynamically deploy and remove Kubernetes clusters in Warewulf managed environments such that nodes can be booted to be either part of a Slurm or Kubernetes cluster while being wiped at shutdown. Jonathan Decker, Julian M. Kunkel |
J. Supercomput. | 2 |
| 2024 | HOSHMAND: Accelerated AI-Driven Scheduler Emulating Conventional Task Distribution Techniques for Cloud WorkloadsabstractCloud computing clusters, especially those handling cloud workloads, require efficient job scheduling to optimize resource utilization and minimize completion time. Traditional approaches often fall short in dynamic cloud environments. We propose “HOSHMAND” (High-performance Open sourced AI-based Scheduling Handler for MAnaging Node Distribution), an AI-driven framework using a custom-tailored Recurrent Neural Network (RNN) to rapidly predict the most suitable nodes for cloud workload execution. A key feature of HOSHMAND is its accelerated scheduling capability, which significantly reduces the time required for job allocation compared to traditional methods. This is particularly crucial for cloud environments with fluctuating workloads and diverse computational requirements. A distinct capability of HOSHMAND is its proficiency in managing heterogeneous resources, ensuring optimal allocation regardless of varying computational capabilities or resource types. This adaptability is crucial for contemporary cloud computing en-vironments, which often comprise a diverse array of hardware configurations, to maintain high efficiency and resource utilization. Moreover, HOSHMAND mitigates the overhead associated with repetitive scheduling computations in similar scenarios by leveraging its historical knowledge. Upon recognizing a con-figuration of jobs analogous to previously encountered situations, it promptly enacts the most effective scheduling strategy without redundant recalculations. This predictive capability not only conserves computational resources but also accelerates job execution. Our approach, tested on cloud-based datasets, demonstrates remarkable improvements in scheduling speed and efficiency, validated by reduced time-to-schedule and enhanced overall system throughput. Through its innovative handling of heterogeneous resources and intelligent avoidance of unnecessary scheduling computations, HOSHMAND sets a new benchmark for AI -driven job scheduling in cloud computing environments. Michael Bidollahkhani, Aasish Kumar Sharma, Julian M. Kunkel |
COMPSAC | 3 |
| 2024 | High-Quality I/O Bandwidth Prediction with Minimal Data via Transfer Learning WorkflowabstractProviding a high-quality performance prediction has the potential to enhance various aspects of a cluster, such as devising scheduling and provisioning policies, guiding procurement decisions, suggesting candidate applications for tuning, and identifying probable scaling and porting challenges. Creating such a prediction for the I/O metrics is still challenging, however, due to the intricate interplay of multiple cluster components, making this an ideal case for machine learning. Nevertheless, achieving the required accuracy level with machine learning calls for a substantial amount of high-quality data, which is often a difficult challenge for most HPC clusters. In this work we explore the use of transfer learning to predict the applications’ I/O bandwidth based on a public dataset. As a result, our experiment can provide an I/O bandwidth prediction for a different cluster comparable to the current state-of-the-art result while employing 100 times less data than needed to construct the base model. Furthermore, we evaluate potential future improvements of the proposed workflow. Dmytro Povaliaiev, Radita Liem, Julian M. Kunkel, Jay F. Lofstead, Philip H. Carns |
SBAC-PAD | 3 |
| 2023 | DECICE: Device-Edge-Cloud Intelligent Collaboration FrameworkabstractDECICE is a Horizon Europe project that is developing an AI-enabled open and portable management framework for automatic and adaptive optimization and deployment of applications in computing continuum encompassing from IoT sensors on the Edge to large-scale Cloud / HPC computing infrastructures. In this paper, we describe the DECICE framework and architecture. Furthermore, we highlight use-cases for framework evaluation: intelligent traffic intersection, magnetic resonance imaging, and emergency response. Julian M. Kunkel, Christian Boehme, Jonathan Decker, Fabrizio Magugliani, Dirk Pleiter, Bastian Koller, Karthee Sivalingam, Sabri Pllana, Alexander Nikolov, Müjdat Soytürk, Christian Racca, Andrea Bartolini, Adrian Tate, Berkay Yaman |
CF | 1 |
| 2023 | Secure HPC: A workflow providing a secure partition on an HPC system
Hendrik Nolte, Nicolai Spicher, Andrew Russel, Tim Ehlers, Sebastian Krey, Dagmar Krefting, Julian M. Kunkel |
Future Gener. Comput. Syst. | 7 |
| 2022 | A Secure Workflow for Shared HPC SystemsabstractDriven by the progress of data and compute-intensive methods in various scientific domains, there is an in-creasing demand from researchers working with highly sensitive data to have access to the necessary computational resources to be able to adapt those methods in their respective fields. To satisfy the computing needs of those researchers cost-effectively, it is an open quest to integrate reliable security measures on existing High Performance Computing (HPC) clusters. The fundamental problem with securely working with sensitive data is, that HPC systems are shared systems that are typically trimmed for the highest performance - not for high security. For instance, there are commonly no additional virtualization techniques employed, thus, users typically have access to the host operating system. Since new vulnerabilities are being continuously discovered, solely relying on the traditional Unix permissions is not secure enough. In this paper, we discuss a generic and secure workflow that can be implemented on typical HPC systems allowing users to transfer, store and analyze sensitive data. In our experiments, we see an advantage in the asynchronous execution of IO requests, while reaching 80 % of the ideal performance. Hendrik Nolte, Simon Hernan Sarmiento Sabater, Tim Ehlers, Julian M. Kunkel |
CCGRID | 4 |
| 2019 | A similarity study of I/O traces via string kernels
Raul Torres, Julian M. Kunkel, Manuel F. Dolz, Thomas Ludwig 0002 |
J. Supercomput. | 2 |
| 2018 | Towards Green Scientific Data Compression Through High-Level I/O InterfacesabstractEvery HPC system today has to cope with a deluge of data generated by scientific applications, simulations or large-scale experiments. The upscaling of supercomputer systems and infrastructures, generally results in a dramatic increase of their energy consumption. In this paper, we argue that techniques like data compression can lead to significant gains in terms of power efficiency by reducing both network and storage requirements. However, any data reduction is highly data specific and should comply with established requirements. Therefore, unsuitable or inappropriate compression strategy can utilize more resources and energy than necessary. To that end, we propose a novel methodology for achieving on-the-fly intelligent determination of energy efficient data reduction for a given data set by leveraging state-of-the-art compression algorithms and meta data at application-level I/O. We motivate our work by analyzing the energy and storage saving needs of data sets from real-life scientific HPC applications, and review the various lossless compression techniques that can be applied. We find that the resulting data reduction can decrease the data volume transferred and stored by as much as 80 % in some cases, consequently leading to significant savings in storage and networking costs. Yevhen Alforov, Thomas Ludwig 0002, Anastasiia Novikova, Michael Kuhn 0003, Julian M. Kunkel |
SBAC-PAD | 5 |
| 2012 | Simulation-Aided Performance Evaluation of Server-Side Input/Output OptimizationsabstractThe performance of parallel distributed file systems suffers from many clients executing a large number of operations in parallel, because the I/O subsystem can be easily overwhelmed by the sheer amount of incoming I/O operations. Many optimizations exist that try to alleviate this problem. Client-side optimizations perform preprocessing to minimize the amount of work the file servers have to do. Server-side optimizations use server-internal knowledge to improve performance. The HD Trace framework contains components to simulate, trace and visualize applications. It is used as a test bed to evaluate optimizations that could later be implemented in real-life projects. This paper compares existing client-side optimizations and newly implemented server-side optimizations and evaluates their usefulness for I/O patterns commonly found in HPC. Server-directed I/O chooses the order of non-contiguous I/O operations and tries to aggregate as many operations as possible to decrease the load on the I/O subsystem and improve overall performance. The results show that server-side optimizations beat client-side optimizations in terms of performance for many use cases. Integrating such optimizations into parallel distributed file systems could alleviate the need for sophisticated client-side optimizations. Due to their additional knowledge of internal workflows server-side optimizations may be better suited to provide high performance in general. Michael Kuhn 0003, Julian M. Kunkel, Thomas Ludwig 0002 |
PDP | 2 |
| 2012 | IOPm - Modeling the I/O Path with a Functional Representation of Parallel File System and Hardware ArchitectureabstractThe I/O path model (IOPm) is a graphical representation of the architecture of parallel file systems and the machine they are deployed on. With help of IOPm, file system and machine configurations can be quickly analyzed and distinguished from each other. Contrary to typical representations of the machine and file system architecture, the model visualizes the data or meta data path of client access. Abstract functionality of hardware components such as client and server nodes is covered as well as software aspects such as high-level I/O libraries, collective I/O and caches. Redundancy could be represented, too. Besides the advantage of a standardized representation for analysis IOPm assists to identify and communicate bottlenecks in the machine and file system configuration by highlighting performance relevant functionalities. By abstracting functionalities from the components they are hosted on, IOPm will enable to build interfaces to monitor file system activity. Julian M. Kunkel, Thomas Ludwig 0002 |
PDP | 1 |
| 2012 | A study on data deduplication in HPC storage systemsabstractDeduplication is a storage saving technique that is highly successful in enterprise backup environments. On a file system, a single data block might be stored multiple times across different files, for example, multiple versions of a file might exist that are mostly identical. With deduplication, this data replication is localized and redundancy is removed -- by storing data just once, all files that use identical regions refer to the same unique data. The most common approach splits file data into chunks and calculates a cryptographic fingerprint for each chunk. By checking if the fingerprint has already been stored, a chunk is classified as redundant or unique. Only unique chunks are stored. This paper presents the first study on the potential of data deduplication in HPC centers, which belong to the most demanding storage producers. We have quantitatively assessed this potential for capacity reduction for 4 data centers (BSC, DKRZ, RENCI, RWTH). In contrast to previous deduplication studies focusing mostly on backup data, we have analyzed over one PB (1212 TB) of online file system data. The evaluation shows that typically 20% to 30% of this online data can be removed by applying data deduplication techniques, peaking up to 70% for some data sets. This reduction can only be achieved by a subfile deduplication approach, while approaches based on whole-file comparisons only lead to small capacity savings. Dirk Meister, Jürgen Kaiser, André Brinkmann, Toni Cortes, Michael Kuhn 0003, Julian M. Kunkel |
SC | 6 |
| 2009 | Small-file access in parallel file systemsabstractToday's computational science demands have resulted in ever larger parallel computers, and storage systems have grown to match these demands. Parallel file systems used in this environment are increasingly specialized to extract the highest possible performance for large I/O operations, at the expense of other potential workloads. While some applications have adapted to I/O best practices and can obtain good performance on these systems, the natural I/O patterns of many applications result in generation of many small files. These applications are not well served by current parallel file systems at very large scale. This paper describes five techniques for optimizing small-file access in parallel file systems for very large scale systems. These five techniques are all implemented in a single parallel file system (PVFS) and then systematically assessed on two test platforms. A microbenchmark and the mdtest benchmark are used to evaluate the optimizations at an unprecedented scale. We observe as much as a 905% improvement in small-file create rates, 1,106% improvement in small-file stat rates, and 727% improvement in small-file removal rates, compared to a baseline PVFS configuration on a leadership computing platform using 16,384 cores. Philip H. Carns, Samuel Lang, Robert B. Ross, Murali Vilayannur, Julian M. Kunkel, Thomas Ludwig 0002 |
IPDPS | 5 |
| 2009 | Tracing Internal Communication in MPI and MPI-I/OabstractMPI implementations can realize MPI operations with any algorithm that fulfills the specified semantics. To provide optimal efficiency the MPI implementation might choose the algorithm dynamically, depending on the parameters given to the function call. However, this selection is not transparent to the user. While this abstraction is appropriate for common users, achieving best performance with fixed parameter sets requires knowledge of internal processing. Also, for developers of collective operations it might be useful to understand timing issues inside the communication or I/O call. In this paper we extended the PIOviz environment to trace MPI internal communication. Thus, this allows the user to see PVFS server behavior together with the behavior in the MPI application and inside MPI itself. We present some analysis results for these capabilities for MPICH2 on a Beowulf Cluster. Julian M. Kunkel, Yuichi Tsujita, Olga Mordvinova, Thomas Ludwig 0002 |
PDCAT | 1 |
| 2009 | Dynamic file system semantics to enable metadata optimizations in PVFSabstractAbstract Modern file systems maintain extensive metadata about stored files. While metadata typically is useful, there are situations when the additional overhead of such a design becomes a problem in terms of performance. This is especially true for parallel and cluster file systems, where every metadata operation is even more expensive due to their architecture. In this paper several changes made to the parallel cluster file system Parallel Virtual File System (PVFS) are presented. The changes target at the optimization of workloads with large numbers of small files. To improve the metadata performance, PVFS was modified such that unnecessary metadata is not managed anymore. Several tests with a large quantity of files were performed to measure the benefits of these changes. The tests have shown that common file system operations can be sped up by a factor of two even with relatively few changes. Copyright © 2009 John Wiley & Sons, Ltd. Michael Kuhn 0003, Julian M. Kunkel, Thomas Ludwig 0002 |
Concurr. Comput. Pract. Exp. | 2 |
| 2008 | Directory-Based Metadata Optimizations for Small Files in PVFS
Michael Kuhn 0003, Julian M. Kunkel, Thomas Ludwig 0002 |
Euro-Par | 2 |
| 2008 | Bottleneck Detection in Parallel File Systems with Trace-Based Performance Monitoring
Julian M. Kunkel, Thomas Ludwig 0002 |
Euro-Par | 1 |
| 2007 | Performance Evaluation of the PVFS2 ArchitectureabstractAs the complexity of parallel file systems' software stacks increases it gets harder to reveal the reasons for performance bottlenecks in these software layers. This paper introduces a method which eliminates the influence of the physical storage on performance analysis in order to find these bottlenecks. Also, the influence of the hardware components on the performance is modeled to estimate the maximum achievable performance of a parallel file system. The paper focuses on the parallel virtual file system 2 (PVFS2) and shows results for the functionality file creation, small contiguous I/O requests and large contiguous I/O requests Julian M. Kunkel, Thomas Ludwig 0002 |
PDP | 1 |