EDBT 2026 Demo / reviewers in the wild / expert
Sidharth Kumar
dblp:49/10367
· DBLP profile ↗
32ranked-venue papers
8as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 4 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Volume Encoding Gaussians: Transfer Function-Agnostic 3D Gaussians for Volume RenderingabstractVisualizing the large-scale datasets output by HPC resources presents a difficult challenge, as the memory and compute power required become prohibitively expensive for end user systems. Novel view synthesis techniques can address this by producing a small, interactive model of the data, requiring only a set of training images to learn from. While these models allow accessible visualization of large data and complex scenes, they do not provide the interactions needed for scientific volumes, as they do not support interactive selection of transfer functions and lighting parameters. To address this, we introduce Volume Encoding Gaussians (VEG), a 3D Gaussian-based representation for volume visualization that supports arbitrary color and opacity mappings. Unlike prior 3D Gaussian Splatting (3DGS) methods that store color and opacity for each Gaussian, VEG decouple the visual appearance from the data representation by encoding only scalar values, enabling transfer function-agnostic rendering of 3DGS models. To ensure complete scalar field coverage, we introduce an opacity-guided training strategy, using differentiable rendering with multiple transfer functions to optimize our data representation. This allows VEG to preserve fine features across a dataset's full scalar range while remaining independent of any specific transfer function. Across a diverse set of volume datasets, we demonstrate that our method outperforms the state-of-the-art on transfer functions unseen during training, while requiring a fraction of the memory and training time. Landon Dyken, Andres Sewell, Will Usher 0001, Nathan DeBardeleben, Steve Petruzza, Sidharth Kumar |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Column-Oriented Datalog on the GPUabstractDatalog is a logic programming language widely used in knowledge representation and reasoning (KRR), program analysis, and social media mining due to its expressiveness and high performance. Traditionally, Datalog engines use either row-oriented or column-oriented storage. Engines like VLog and Nemo favor column-oriented storage for efficiency on limited-resource machines, while row-oriented engines like Soufflé use advanced datastructures with locking to perform better on multi-core CPUs. The advent of modern datacenter GPUs, such as the NVIDIA H100 with its ability to run over 16k threads simultaneously and high memory bandwidth, has reopened the debate on which storage layout is more effective. This paper presents the first column-oriented Datalog engines tailored to the strengths of modern GPUs. We present VFLog, a CUDA-based Datalog runtime library with a column-oriented GPU datastructure that supports all necessary relational algebra operations. Our results demonstrate over 200x performance gains over SOTA CPU-based column-oriented Datalog engines and a 2.5x speedup over GPU Datalog engines in various workloads, including KRR. Yihao Sun 0003, Sidharth Kumar, Thomas Gilray, Kristopher K. Micinski |
AAAI | 2 |
| 2025 | Optimizing Datalog for the GPUabstractModern Datalog engines (e.g., LogicBlox, Soufflé, ddlog) enable their users to write declarative queries which compute recursive deductions over extensional facts, leaving high-performance operationalization (query planning, semi-naïve evaluation, and parallelization) to the engine. Such engines form the backbone of modern high-throughput applications in static analysis, network monitoring, and social-media mining. In this paper, we present a methodology for implementing a modern in-memory Datalog engine on data center GPUs, allowing us to achieve significant (up to 45×) gains compared to Soufflé (a modern CPU-based engine) on context-sensitive points-to analysis of httpd. We present GPUlog, a Datalog engine backend that implements iterated relational algebra kernels over a novel range-indexed data structure we call the hash-indexed sorted array (HISA). HISA combines the algorithmic benefits of incremental range-indexed relations with the raw computation throughput of operations over dense data structures. Our experiments show that GPUlog is significantly faster than CPU-based Datalog engines, while achieving favorable memory footprint compared to contemporary GPU-based joins. Yihao Sun 0003, Ahmedur Rahman Shovon, Thomas Gilray, Sidharth Kumar, Kristopher K. Micinski |
ASPLOS (1) | 4 |
| 2025 | Machine Learning Driven Auto-Tuning for Non-Uniform All-to-All CollectivesabstractNon-uniform all-to-all communication patterns present optimization challenges in parallel computing due to their irregular data distribution and dynamic behavior. While MPI_Alltoallv provides the standard interface for such exchanges, achieving optimal performance requires careful selection among multiple implementation variants and tuning of algorithm-specific parameters. This paper presents a data-driven autotuning framework that combines machine learning-based runtime prediction with a lookup-table mechanism for fast configuration selection. The ML model estimates the communication time of each algorithm configuration under a given system setup, allowing the framework to identify the optimal implementation and parameter set based on predicted performance. We validate our approach through comprehensive benchmarking of MPI_Alltoallv and two specialized algorithms across varying process counts, message sizes, and tunable parameters. Applied to a real MPI-based transitive closure application on the Fugaku supercomputer, our framework achieves up to$6.03 \times$reduction in communication time over the vendor implementation, providing a detailed understanding of non-uniform collective communication behavior and a practical framework for automatic performance optimization in HPC applications. Kunting Qi, Jens Domke, Seydou Ba, Venkatram Vishwanath, Michael E. Papka, Sidharth Kumar |
HiPC | 7 |
| 2025 | Parameterized Algorithms for Non-uniform All-to-allabstractMPI_Alltoallv generalizes the uniform all-to-all communication (MPI_Alltoall) by enabling the exchange of data-blocks of varied sizes among processes. This function plays a crucial role in facilitating many computational tasks, such as FFT calculations and graph mining operations. Popular MPI libraries, such as MPICH and OpenMPI, implement MPI_Alltoall using a combination of linear and logarithmic algorithms. However, MPI_Alltoallv typically relies only on variations of linear algorithms, missing the benefits of logarithmic approaches. Furthermore, current algorithms also overlook the intricacies of modern HPC system architectures, such as the significant performance gap between intra-node (local) and inter-node (global) communication. To address these problems, this paper presents two novel algorithms: Parameterized Logarithmic non-uniform All-to-all (ParLogNa) and Parameterized Linear nonuniform All-to-all (ParLinNa). ParLogNa is a tunable logarithmic time algorithm for non-uniform all-to-all, and ParLinNa is a hierarchical and tunable near-linear-time algorithm for non-uniform all-to-all. These algorithms efficiently address the trade-off between bandwidth maximization and latency minimization that existing implementations struggle to optimize. We show a performance improvement over the state-of-the-art implementations by factors of 42x and 138x on Polaris and Fugaku, respectively. Jens Domke, Seydou Ba, Sidharth Kumar |
HPDC | 4 |
| 2025 | Ambient Diffusion Posterior Sampling: Solving Inverse Problems with Diffusion Models Trained on Corrupted DataabstractWe provide a framework for solving inverse problems with diffusion models learned from linearly corrupted data. Firstly, we extend the Ambient Diffusion framework to enable training directly from measurements corrupted in the Fourier domain. Subsequently, we train diffusion models for MRI with access only to Fourier subsampled multi-coil measurements at acceleration factors R$=2, 4, 6, 8$. Secondly, we propose $\textit{Ambient Diffusion Posterior Sampling}$ (A-DPS), a reconstruction algorithm that leverages generative models pre-trained on one type of corruption (e.g. image inpainting) to perform posterior sampling on measurements from a different forward process (e.g. image blurring). For MRI reconstruction in high acceleration regimes, we observe that A-DPS models trained on subsampled data are better suited to solving inverse problems than models trained on fully sampled data. We also test the efficacy of A-DPS on natural image datasets (CelebA, FFHQ, and AFHQ) and show that A-DPS can sometimes outperform models trained on clean data for several image restoration tasks in both speed and performance. Asad Aali, Giannis Daras, Brett Levac, Sidharth Kumar, Alexandros G. Dimakis, Jonathan I. Tamir |
ICLR | 4 |
| 2025 | Multi-Node Multi-GPU DatalogabstractDatalog, a declarative logic programming language that operates bottom-up, has experienced increasing popularity due to its natural handling of recursive queries.Its applications span diverse fields, including graph mining, program analysis, deductive databases, and neuro-symbolic reasoning.While Datalog shares similarities with SQL in using relational algebra kernels, it uniquely employs iterative execution until reaching a fixed point to support recursion.Current Datalog engines like SLOG, LogicBlox, and Soufflé work well with multi-core and multi-threaded systems, but none have yet tackled multi-node, multi-GPU architectures.Our research addresses this gap by developing the first multi-GPU, multinode Datalog engine.This advancement is particularly for high-performance computing (HPC) systems, which typically feature multiple GPUs per node.Our implementation combines MPI for inter-node communication with CUDA for GPU parallelization, enabling the processing of massive datasets in real time.We have created novel data-parallel implementations of core relational algebra operations (join), while also optimizing deduplication and tuple materialization.To handle iterative execution, we have developed two novel GPU-accelerated methods for non-uniform all-to-all data exchange.Evaluating on Argonne National Lab's Polaris supercomputer demonstrated our engine's effectiveness, achieving performance improvements of up to 32× against state-of-the-art multi-node Datalog engine. Ahmedur Rahman Shovon, Yihao Sun 0003, Kristopher K. Micinski, Thomas Gilray, Sidharth Kumar |
ICS | 5 |
| 2025 | Faster Annotation for Elevation-Guided Flood Extent Mapping by Consistency-Enhanced Active LearningabstractFlood extent mapping is crucial for disaster response and damage assessment. While Earth imagery and terrain data (in the form of DEM) are now readily available, there are few flood annotation data for training machine learning models, which hinders the automated mapping of flooded areas. We propose ALFA, an interactive active-learning-based approach to minimize the annotators' efforts when preparing the ground-truth flood map in a satellite image. ALFA calibrates the prediction consistency of a segmentation model (1) across training cycles and (2) for various data augmentations. The two consistencies are integrated into the design of both the acquisition function and the loss function to enhance the robustness of active learning with limited annotation inputs. ALFA recommends those superpixels that the underlying model is most uncertain about, and users can annotate their pixels with minimal clicks with the help of elevation guidance. Extensive experiments on various regions hit by flooding show that we can improve the annotation time from hours to around 20 minutes. ALFA is open sourced at https://github.com/saugatadhikari/alfa. Saugat Adhikari, Da Yan 0001, Landon Dyken, Sidharth Kumar, Lyuheng Yuan, Akhlaque Ahmad, Yang Zhou 0001, Steve Petruzza |
IJCAI | 5 |
| 2025 | Accelerating Web-Based Graph Drawing with Bottom-Up GPU Quadtree ConstructionabstractGraph drawing, or graph layout creation, is a computationally difficult challenge in visualization that involves placing the vertices of a graph into a layout that provides insight into its structure. In order to visualize large-scale graphs, effective layouts are necessary for understanding. Previous work has shown the potential for graph drawing directly in the web browser by using WebGPU, a new API that brings the full capabilities of modern GPUs to the web. Compared to the existing state-of-the-art for web-based graph visualization, which rely on CPU-based graph drawing algorithms, WebGPU-accelerated work improves performance and scalability. However, we find that existing WebGPU solutions utilize suboptimal quadtree data structures for graph drawing. In this work, we implement a modified quadtree data structure that uses a Hilbert spatial ordering for a fully parallelizable bottom-up construction algorithm in WebGPU. We utilize this data structure, along with optimizations to the quadtree traversal, to propose a massively more performant graph drawing algorithm. We evaluate the performance of our work against the existing state-of-the-art and demonstrate up to 69.5 × speed-ups for layout creation of relevant graphs while enabling graph drawing for datasets of much larger size. Landon Dyken, Will Usher 0001, Steve Petruzza, Stavros Sintos, Sidharth Kumar |
PacificVis | 5 |
| 2025 | Enabling Fast and Accurate Crowdsourced Annotation for Elevation-Aware Flood Extent MappingabstractMapping the extent of flood events is a necessary and important aspect of disaster management. In recent years, deep learning methods have evolved as an effective tool to quickly label high-resolution imagery and provide necessary flood extent mappings. These methods, though, require large amounts of annotated training data to create models that are accurate and robust to new flooded imagery. In this work, we present FloodTrace, a web-based application that enables effective crowdsourcing of flooded region annotation for machine learning applications. To create this application, we conducted extensive interviews with domain experts to produce a set of formal requirements. Our work brings topological segmentation tools to the web and greatly improves annotation efficiency compared to the state-of-the-art. The user-friendliness of our solution allows researchers to outsource annotations to non-experts and utilize them to produce training data with equal quality to fully expert-labeled data. We conducted a user study to confirm our application’s effectiveness in which 266 graduate students annotated high-resolution aerial imagery from Hurricane Matthew in North Carolina. Experimental results show the efficiency benefits of our application for untrained users, with median annotation time less than half the state-of-the-art annotation method. In addition, using our application’s aggregation and correction framework, flood detection models trained on crowdsourced annotations were able to achieve performance equal to models trained on fully expert-labeled annotations, while requiring a fraction of the expert’s time. Landon Dyken, Saugat Adhikari, Pravin Poudel, Steve Petruzza, Da Yan 0001, Will Usher 0001, Sidharth Kumar |
PacificVis | 7 |
| 2025 | Interactive Isosurface Visualization in Memory Constrained Environments Using Deep Learning and Speculative RaycastingabstractNew web technologies have enabled the deployment of powerful GPU-based computational pipelines that run entirely in the web browser, opening a new frontier for accessible scientific visualization applications. However, these new capabilities do not address the memory constraints of lightweight end-user devices encountered when attempting to visualize the massive data sets produced by today's simulations and data acquisition systems. We propose a novel implicit isosurface rendering algorithm for interactive visualization of massive volumes within a small memory footprint. We achieve this by progressively traversing a wavefront of rays through the volume and decompressing blocks of the data on-demand to perform implicit ray-isosurface intersections, displaying intermediate results each pass. We improve the quality of these intermediate results using a pretrained deep neural network that reconstructs the output of early passes, allowing for interactivity with better approximates of the final image. To accelerate rendering and increase GPU utilization, we introduce speculative ray-block intersection into our algorithm, where additional blocks are traversed and intersected speculatively along rays to exploit additional parallelism in the workload. Our algorithm is able to trade-off image quality to greatly decrease rendering time for interactive rendering even on lightweight devices. Our entire pipeline is run in parallel on the GPU to leverage the parallel computing power that is available even on lightweight end-user devices. We compare our algorithm to the state of the art in low-overhead isosurface extraction and demonstrate that it achieves - reductions in memory overhead and up to reductions in data decompressed. Landon Dyken, Will Usher 0001, Sidharth Kumar |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Bruck Algorithm Performance Analysis for Multi-GPU All-to-All CommunicationabstractIn high-performance computing, collective communication is critical for facilitating comprehensive data exchange involving all processes within an MPI communicator. Due to their inherently global nature, many collective operations present scalability challenges, particularly the all-to-all data shuffle with its quadratic communication pattern. Using a logarithmic communication pattern, the Bruck algorithm was designed to provide communication efficiency for all-to-all data shuffles involving short-sized messages. The Bruck algorithm has been extensively used to facilitate global data shuffles in a multi-CPU environment and is also part of the MPICH and Open MPI implementations. This work presents the first investigation of using the Bruck algorithm for all-to-all communication in multi-GPU systems using the NVIDIA Collective Communications Library (NCCL). Our experimental study demonstrates that while the Bruck algorithm exhibits superior performance for small-sized messages in a multi-CPU environment, the same advantages are not evident for multi-GPU environments. Furthermore, we describe and compare an optimized Bruck algorithm implementation in NCCL and compare it to NCCL’s default all-to-all and MPI-based implementations. Finally, we discuss the challenges and opportunities of implementing new multi-GPU collectives using NCCL’s public-facing API. Andres Sewell, Ahmedur Rahman Shovon, Landon Dyken, Sidharth Kumar, Steve Petruzza |
HPC Asia | 5 |
| 2024 | Datalog with First-Class FactsabstractDatalog is a popular logic programming language for deductive reasoning tasks in a wide array of applications, including business analytics, program analysis, and ontological reasoning. However, Datalog's restriction to flat facts over atomic constants leads to challenges in working with tree-structured data, such as derivation trees or abstract syntax trees. To ameliorate Datalog's restrictions, popular extensions of Datalog support features such as existential quantification in rule heads (Datalog*, Datalog ∃ ) or algebraic data types (Soufflé). Unfortunately, these are imperfect solutions for reasoning over structured and recursive data types, with general existentials leading to complex implementations requiring unification, and ADTs unable to trigger rule evaluation and failing to support efficient indexing. We present D L ∃! , a Datalog with first-class facts, wherein every fact is identified with a Skolem term unique to the fact. We show that this restriction offers an attractive price point for Datalogbased reasoning over tree-shaped data, demonstrating its application to databases, artificial intelligence, and programming languages. We implemented D L ∃! as a system Slog, which leverages the uniqueness restriction of D L ∃! to enable a communication-avoiding, massively-parallel implementation built on MPI. We show that Slog outperforms leading systems (Nemo, Vlog, RDFox, and Soufflé) on a variety of benchmarks, with the potential to scale to thousands of threads. Thomas Gilray, Arash Sahebolamri, Yihao Sun 0003, Sowmith Kunapaneni, Sidharth Kumar, Kristopher K. Micinski |
Proc. VLDB Endow. | 5 |
| 2023 | Communication-Avoiding Recursive AggregationabstractRecursive aggregation has been of considerable interest due to its unifying a wide range of deductive-analytic workloads, including social-media mining and graph analytics. For example, Single-Source Shortest Paths (SSSP), Connected Components (CC), and PageRank may all be expressed via recursive aggregates. Implementing recursive aggregation has posed a serious algorithmic challenge, with state-of-the-art work identifying sufficient conditions (e.g., pre-mappability) under which implementations may push aggregation within recursion, avoiding the serious materialization overhead inherent to traditional reachability-based methods (e.g., Datalog).State-of-the-art implementations of engines supporting recursive aggregates focus on large unified machines, due to the challenges posed by mixing semi-naïve evaluation with distribution. In this work, we present an approach to implementing recursive aggregates on high-performance clusters which avoids the communication overhead inhibiting current-generation distributed systems to scale recursive aggregates to extremely high process counts. Our approach leverages the observation that aggregators form functional dependencies, allowing us to implement recursive aggregates via a high-parallel local aggregation to ensure maximal throughput. Additionally, we present a dynamic join planning mechanism, which customizes join order per-iteration based on dynamic relation sizes. We implemented our approach in PARALAGG, a library which allows the declarative implementation of queries which utilize recursive aggregates and executes them using our MPI-based runtime. We evaluate PARALAGG on a large unified node and leadership-class supercomputers, demonstrating scalability up to 16,384 processes. Yihao Sun 0003, Sidharth Kumar, Thomas Gilray, Kristopher K. Micinski |
CLUSTER | 2 |
| 2023 | Towards Iterative Relational Algebra on the GPU
Ahmedur Rahman Shovon, Thomas Gilray, Kristopher K. Micinski, Sidharth Kumar |
USENIX ATC | 4 |
| 2022 | Optimizing the Bruck Algorithm for Non-uniform All-to-all CommunicationabstractIn MPI, collective routines MPI_Alltoall and MPI_Alltoallv play an important role in facilitating all-to-all inter-process data exchange. MPI_Alltoallv is a generalization of MPI_Alltoall, supporting the exchange of non-uniform distributions of data. Popular implementations of MPI, such as MPICH and OpenMPI, implement MPI_Alltoall using a combination of techniques such as the Spread-out algorithm and the Bruck algorithm. Spread-out has a linear complexity in P, compared to Bruck's logarithmic complexity (P: process count); a selection between these two techniques is made at runtime based on the data block size. However, MPI_Alltoallv is typically implemented using only variants of the spread-out algorithm, and therefore misses out on the performance benefits that the log-time Bruck algorithm offers (especially for smaller data loads). Thomas Gilray, Valerio Pascucci, Xuan Huang 0007, Kristopher K. Micinski, Sidharth Kumar |
HPDC | 6 |
| 2022 | FSE Compensated Motion Correction for MRI Using Data Driven Methods
Brett Levac, Sidharth Kumar, Sofia Kardonik, Jonathan I. Tamir |
MICCAI (6) | 2 |
| 2021 | Compiling data-parallel DatalogabstractDatalog allows intuitive declarative specification of logical inference tasks while enjoying efficient implementation via state-of-the-art engines such as LogicBlox and Soufflé. These engines enable high-performance implementation of complex logical tasks including graph mining, program analysis, and business analytics. However, all efficient modern Datalog solvers make use of shared memory, and present inherent challenges scalability. In this paper, we leverage recent insights in parallel relational algebra and present a methodology for constructing data-parallel deductive databases. Our approach leverages recent developments in parallelizing relational algebra to create an efficient data-parallel semantics for Datalog. Based on our methodology, we have implemented the first MPI-based data-parallel Datalog solver. Our experiments demonstrate comparable performance and improved single-node scalability versus Soufflé, a state-of-art solver. Thomas Gilray, Sidharth Kumar, Kristopher K. Micinski |
CC | 2 |
| 2021 | Load-balancing Parallel I/O of Compressed Hierarchical LayoutsabstractScientific phenomena are being simulated at ever-increasing resolution and fidelity thanks to advances in modern supercomputers. These simulations produce a deluge of data, putting unprecedented demand on the end-to-end data-movement pipeline that consists of parallel writes for checkpoint and analysis dumps and parallel localized reads for exploratory analysis and visualization tasks. Parallel I/O libraries are often optimized for uniformly distributed large-sized accesses, whereas reads for analysis and visualization benefit from data layouts that enable random-access and multiresolution queries. While multiresolution layouts enable interactive exploration of massive datasets, efficiently writing such layouts in parallel is challenging, and straightforward methods for creating a multiresolution hierarchy can lead to inefficient memory and disk access. In this paper, we propose a compressed, hierarchical layout that facilitates efficient parallel writes, while being efficient at serving random access, multiresolution read queries for post-hoc analysis and visualization. To efficiently write data to such a layout in parallel is challenging due to potential load-balancing issues at both the data transformation and disk I/O steps. Data is often not readily distributed in a way that facilitates efficient transformations necessary for creating a multi resolution hierar-chy. Further, when compression or data reduction is applied, the compressed data chunks may end up with different sizes, confounding efficient parallel I/O. To overcome both these issues, we present a novel two-phase load-balancing strategy to optimize both memory and disk access patterns unique to writing non-uniform multiresolution data. We implement these strategies in a parallel I/O library and evaluate the efficacy of our approach by using real-world simulation data and a novel approach to micro benchmarking on the Theta Supercomputer of Argonne National Laboratory. Duong Hoang, Steve Petruzza, Thomas Gilray, Valerio Pascucci, Sidharth Kumar |
HiPC | 6 |
| 2021 | Adaptive Spatially Aware I/O for Multiresolution Particle Data LayoutsabstractLarge-scale simulations on nonuniform particle distributions that evolve over time are widely used in cosmology, molecular dynamics, and engineering. Such data are often saved in an unstructured format that neither preserves spatial locality nor provides metadata for accelerating spatial or attribute subset queries, leading to poor performance of visualization tasks. Furthermore, the parallel I/O strategy used typically writes a file per process or a single shared file, neither of which is portable or scalable across different HPC systems. We present a portable technique for scalable, spatially aware adaptive aggregation that preserves spatial locality in the output. We evaluate our approach on two supercomputers, Stampede2 and Summit, and demonstrate that it outperforms prior approaches at scale, achieving up to 2.5 x faster writes and reads for nonuniform distributions. Furthermore, the layout written by our method is directly suitable for visual analytics, supporting low-latency reads and attribute-based filtering with little overhead. Will Usher 0001, Xuan Huang 0007, Steve Petruzza, Sidharth Kumar, Stuart R. Slattery, Samuel Temple Reeve, Feng Wang 0013, Chris R. Johnson 0001, Valerio Pascucci |
IPDPS | 4 |
| 2021 | Dynamic Resource Allocation in UAV-Enabled mmWave Communication NetworksabstractUnmanned aerial vehicle (UAV)-enabled cellular architecture over the millimeter-wave (mmWave) frequency band is likely to be the best solution for on-demand high data rate service provisioning in the next generation communication networks. The beam formed by the mmWave antenna array is highly directional and requires multiple-beam scans to cover the entire area. This work presents a novel sectoring approach to ensure coverage of the whole area. The side lobe gain of the antenna array is taken into consideration, which generates substantial interference in other sectors. The expression for the probability distribution of the signal-to-interference-plus-noise ratio due to simultaneous transmissions in different sectors is derived in the downlink communication scenario. To limit interference in the concurrent transmission strategy, a threshold on power spillage from adjacent sectors is placed. For this topology, a resource allocation problem is formulated aiming to maximize the sum rate while ensuring a minimum rate guarantee to each user. It is observed that sum-rate variation with height is unimodal. The sum power and backhaul capacity constraints are accounted. This optimization problem is mixed-integer nonconvex programming. Hence, it is solved using the Lagrangian dual decomposition method, which provides an asymptotic global optimal solution. Since this method is computationally intensive, a suboptimal solution is proposed. Simulation results demonstrate convergence to an optimal solution, and it is observed that backhaul link capacity restricts the sum rate. Numerical results are presented for multiple representative field environments consisting of different types of built-up areas. It is observed that the transmitter antenna array sidelobe has a strong impact on the performance as compared to the ideal scenario without sidelobe, which overestimates the total sum rate by a factor of 3. Sidharth Kumar, Suraj Suman, Swades De |
IEEE Internet Things J. | 1 |
| 2019 | Distributed Relational Algebra at ScaleabstractRelational algebra forms a basis of primitive operations suitable for applications in graphs and networks, program analysis, deductive databases, and constraint logic programming. Despite its expressive power, relational algebra has not received the same attention in high-performance-computing research as more common primitives like stencil computations, floating-point operations, numerical integration, and sparse linear algebra. Furthermore, specific challenges in addressing representation and communication among distributed portions of a relation, especially for inherently imbalanced relations, have previously thwarted successful scaling of relational algebra applications to supercomputers. In this paper, we present a set of efficient algorithms to effectively parallelize and scale key relational algebra primitives. We introduce a hybrid hash-tree approach to representing distributed imbalanced relations and permitting efficient communication. Finally, we demonstrate the scalability of our implementation with a fixed-point algorithm computing the transitive closure of a large graph (generating over 276 billion edges) on 32,768 processes. Thomas Gilray, Sidharth Kumar |
HiPC | 2 |
| 2019 | Spatially-aware Parallel I/O for Particle DataabstractParticle data are used across a diverse set of large scale simulations, for example, in cosmology, molecular dynamics and combustion. At scale these applications generate tremendous amounts of data, which is often saved in an unstructured format that does not preserve spatial locality; resulting in poor read performance for post-processing analysis and visualization tasks, which typically make spatial queries. In this work, we explore some of the challenges of large scale particle data management, and introduce new techniques to perform scalable, spatially-aware write and read operations. We propose an adaptive aggregation technique to improve the performance of data aggregation, for both uniform and non-uniform particle distributions. Furthermore, we enable efficient read operations by employing a level of detail re-ordering and a multi-resolution layout. Finally, we demonstrate the scalability of our techniques with experiments on large scale simulation workloads up to 256K cores on two different leadership supercomputers, Mira and Theta. Sidharth Kumar, Steve Petruzza, Will Usher 0001, Valerio Pascucci |
ICPP | 1 |
| 2019 | Capttery: Scalable Battery-like Room-level Wireless PowerabstractInternet-of-things (IoT) devices are becoming widely adopted, but they increasingly suffer from limited power, as power cords cannot reach the billions and batteries do not last forever. Existing systems address the issue with ultra-low-power designs and energy scavenging, which inevitably limit functionality. To unlock the full potential of ubiquitous computing and connectivity, our solution uses capacitive power transfer (CPT) to provide battery-like wireless power delivery, henceforth referred to as "Capttery". Capttery presents the first room-level (~5 m) CPT system, which delivers continuous milliwatt-level wireless power to multiple IoT devices concurrently. Unlike conventional one-to-one CPT systems that target kilowatt power in a controlled and potentially hazardous setup, Capttery is designed to be human-safe and invariant in a practical and dynamic environment. Our evaluation shows that Capttery can power end-to-end IoT applications across a typical room, where new receivers can be easily added in a plug-and-play manner. Chi Zhang 0018, Sidharth Kumar, Dinesh Bharadia |
MobiSys | 2 |
| 2018 | UAV-Assisted RF Energy TransferabstractLimited battery capacity is one of the major hurdles towards perpetual operation of wireless sensor networks (WSNs). In this work, a framework for Unmanned Aerial Vehicle (UAV) based wireless charging of sensor nodes using radio frequency energy transfer (RFET) is presented. First, RFET zone is conceptualized where energy transfer is possible such that the received power is above a sensitivity level of power harvester. Next, two novel strategies Static Charging Time Allocation (SCTA) and Optimal Charging Time Allocation (OCTA) are proposed for charging the sensors. The UAV remains static in SCTA, whereas in OCTA it hovers above each sensor node and replenishes depleted energy of the sensors that are within its RFET zone. Further, a model for evaluating the energy consumption of UAV is presented in order to determine the number of rounds required by the UAV for charging. Different gas sensors are planted emulating a practical deployment scenario, and with this arrangement charging time with the proposed strategies are evaluated. Our numerical results justify applicability of the UAV-assisted RFET. Suraj Suman, Sidharth Kumar, Swades De |
ICC | 2 |
| 2018 | RF Energy Transfer Channel Models for Sustainable IoTabstractSelf-sustainability of wireless nodes in Internet-of-Things applications can be realized with the help of controlled radio frequency energy transfer (RF-ET). However, due to significant energy loss in wireless dissipation, there is a need for novel schemes to improve the end-to-end RF-ET efficiency. In this paper, first we propose a new channel model for accurately characterizing the harvested dc power at the receiver. This model incorporates the effects of nonline of sight (NLOS) component along with the other factors, such as radiation pattern of transmit and receive antennas, losses associated with different polarization of transmitting field, and efficiency of power harvester circuit. Accuracy of the model is verified via experimental studies in an anechoic chamber (a controlled environment). Supported by experiments in controlled environment, we also formulate an optimization problem by accounting for the effect of NLOS component to maximize the RF-ET efficiency, which cannot be captured by the Friis formula. To solve this nonconvex problem, we present a computationally efficient golden section-based iterative algorithm. Finally, through extensive RF-ET measurements in different practical field environments we obtain the statistical parameters for Rician fading as well as path loss factor associated with shadow fading model, which also asserts the fact that Rayleigh fading is not well suited for RF-ET due to presence of a strong line of sight component. Sidharth Kumar, Swades De, Deepak Mishra 0001 |
IEEE Internet Things J. | 1 |
| 2017 | Reducing Network Congestion and Synchronization Overhead During Aggregation of Hierarchical DataabstractHierarchical data representations have been shown to be effective tools for coping with large-scale scientific data. Writing hierarchical data on supercomputers, however, is challenging as it often involves all-to-one communication during aggregation of low-resolution data which tends to span the entire network domain, resulting in several bottlenecks. We introduce the concept of indexing templates, which succinctly describe data organization and can be used to alter movement of data in beneficial ways. We present two techniques, domain partitioning and localized aggregation, that leverage indexing templates to alleviate congestion and synchronization overheads during data aggregation. We report experimental results that show significant I/O speedup using our proposed schemes on two of today's fastest supercomputers, Mira and Shaheen II, using the Uintah and S3D simulation frameworks. Sidharth Kumar, Duong Hoang, Steve Petruzza, John Edwards 0002, Valerio Pascucci |
HiPC | 1 |
| 2016 | Evaluation of In-Situ Analysis Strategies at Scale for Power Efficiency and ScalabilityabstractThe increasing gap between available compute power and I/O capabilities is resulting in simulation pipelines running on leadership computing facilities being reformulated. In particular, in-situ processing is complementing conventional post-process analysis, however, it can be performed by using the same compute resources as the simulation or using secondary dedicated resources. In this paper, we focus on three different in-situ analysis strategies, which use the same compute resources as the ongoing simulation but different data movement strategies. We evaluate the costs incurred by these strategies in terms of run time, scalability and power/energy consumption. Furthermore, we extrapolate power behavior to peta-scale and investigate different design choices through projections. Experimental evaluation at full machine scale on Titan supports that using fewer cores per node for in-situ analysis is the optimum choice in terms of scalability. Hence, further research effort should be devoted towards developing in-situ analysis techniques following this strategy in future high-end systems. Ivan Rodero, Manish Parashar, Aaditya G. Landge, Sidharth Kumar, Valerio Pascucci, Peer-Timo Bremer |
CCGrid | 4 |
| 2014 | Efficient I/O and Storage of Adaptive-Resolution DataabstractWe present an efficient, flexible, adaptive-resolution I/O framework that is suitable for both uniform and Adaptive Mesh Refinement (AMR) simulations. In an AMR setting, current solutions typically represent each resolution level as an independent grid which often results in inefficient storage and performance. Our technique coalesces domain data into a unified, multiresolution representation with fast, spatially aggregated I/O. Furthermore, our framework easily extends to importance-driven storage of uniform grids, for example, by storing regions of interest at full resolution and nonessential regions at lower resolution for visualization or analysis. Our framework, which is an extension of the PIDX framework, achieves state of the art disk usage and I/O performance regardless of resolution of the data, regions of interest, and the number of processes that generated the data. We demonstrate the scalability and efficiency of our framework using the Uintah and S3D large-scale combustion codes on the Mira and Edison supercomputers. Sidharth Kumar, John Edwards 0002, Peer-Timo Bremer, Aaron Knoll, Cameron Christensen, Venkatram Vishwanath, Philip H. Carns, John A. Schmidt, Valerio Pascucci |
SC | 1 |
| 2013 | Characterization and modeling of PIDX parallel I/O for performance optimizationabstractParallel I/O library performance can vary greatly in response to user-tunable parameter values such as aggregator count, file count, and aggregation strategy. Unfortunately, manual selection of these values is time consuming and dependent on characteristics of the target machine, the underlying file system, and the dataset itself. Some characteristics, such as the amount of memory per core, can also impose hard constraints on the range of viable parameter values. In this work we address these problems by using machine learning techniques to model the performance of the PIDX parallel I/O library and select appropriate tunable parameter values. We characterize both the network and I/O phases of PIDX on a Cray XE6 as well as an IBM Blue Gene/P system. We use the results of this study to develop a machine learning model for parameter space exploration and performance prediction. Sidharth Kumar, Avishek Saha, Venkatram Vishwanath, Philip H. Carns, John A. Schmidt, Giorgio Scorzelli, Hemanth Kolla, Ray W. Grout, Robert Latham, Robert B. Ross, Michael E. Papka, Jacqueline Chen, Valerio Pascucci |
SC | 1 |
| 2012 | Efficient data restructuring and aggregation for I/O acceleration in PIDXabstractHierarchical, multiresolution data representations enable interactive analysis and visualization of large-scale simulations. One promising application of these techniques is to store high performance computing simulation output in a hierarchical Z (HZ) ordering that translates data from a Cartesian coordinate scheme to a one-dimensional array ordered by locality at different resolution levels. However, when the dimensions of the simulation data are not an even power of 2, parallel HZ ordering produces sparse memory and network access patterns that inhibit I/O performance. This work presents a new technique for parallel HZ ordering of simulation datasets that restructures simulation data into large (power of 2) blocks to facilitate efficient I/O aggregation. We perform both weak and strong scaling experiments using the S3D combustion application on both Cray-XE6 (65,536 cores) and IBM Blue Gene/P (131,072 cores) platforms. We demonstrate that data can be written in hierarchical, multiresolution format with performance competitive to that of native data-ordering methods. Sidharth Kumar, Venkatram Vishwanath, Philip H. Carns, Joshua A. Levine, Robert Latham, Giorgio Scorzelli, Hemanth Kolla, Ray W. Grout, Robert B. Ross, Michael E. Papka, Jacqueline Chen, Valerio Pascucci |
SC | 1 |
| 2011 | PIDX: Efficient Parallel I/O for Multi-resolution Multi-dimensional Scientific DatasetsabstractThe IDX data format provides efficient, cache oblivious, and progressive access to large-scale scientific datasets by storing the data in a hierarchical Z (HZ) order. Data stored in IDX format can be visualized in an interactive environment allowing for meaningful explorations with minimal resources. This technology enables real-time, interactive visualization and analysis of large datasets on a variety of systems ranging from desktops and laptop computers to portable devices such as iPhones/iPads and over the web. While the existing ViSUS API for writing IDX data is serial, there are obvious advantages of applying the IDX format to the output of large scale scientific simulations. We have therefore developed PIDX - a parallel API for writing data in an IDX format. With PIDX it is now possible to generate IDX datasets directly from large scale scientific simulations with the added advantage of real-time monitoring and visualization of the generated data. In this paper, we provide an overview of the IDX file format and how it is generated using PIDX. We then present a data model description and a novel aggregation strategy to enhance the scalability of the PIDX library. The S3D combustion application is used as an example to demonstrate the efficacy of PIDX for a real-world scientific simulation. S3D is used for fundamental studies of turbulent combustion requiring exceptionally high fidelity simulations. PIDX achieves up to 18 GiB/s I/O throughput at 8,192 processes for S3D to write data out in the IDX format. This allows for interactive analysis and visualization of S3D data, thus, enabling in situ analysis of S3D simulation. Sidharth Kumar, Venkatram Vishwanath, Philip H. Carns, Brian Summa, Giorgio Scorzelli, Valerio Pascucci, Robert B. Ross, Jacqueline Chen, Hemanth Kolla, Ray W. Grout |
CLUSTER | 1 |