EDBT 2026 Demo / reviewers in the wild / expert
Kalin Kanov
dblp:56/10106
· DBLP profile ↗
7ranked-venue papers
4as first author
0since 2021 · last 2018
0000-0002-6323-9081ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 89% Hardware accelerators and domain-specific architectures · 11% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Computational science and engineering · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
parallel i/o |
0.4 | 2 | 2015 | Particle tracking in open simulation laboratories · SC 2015 Data-intensive spatial filtering in large numerical simulation datasets · SC 2012 |
High-performance computing
scientific data management |
0.3 | 2 | 2012 | Data-intensive spatial filtering in large numerical simulation datasets · SC 2012 I/O streaming evaluation of batch queries for data-intensive computational turbulence · SC 2011 |
High-performance computing
streaming i/o |
0.3 | 2 | 2012 | Data-intensive spatial filtering in large numerical simulation datasets · SC 2012 I/O streaming evaluation of batch queries for data-intensive computational turbulence · SC 2011 |
High-performance computing
scientific computing |
0.2 | 1 | 2015 | Particle tracking in open simulation laboratories · SC 2015 |
Hardware accelerators and domain-specific architectures
query processing |
0.1 | 1 | 2012 | Data-intensive spatial filtering in large numerical simulation datasets · SC 2012 |
Computational science and engineering
computational fluid dynamics |
0.1 | 1 | 2015 | Particle tracking in open simulation laboratories · SC 2015 |
Computational science and engineering › computational fluid dynamics
turbulence simulation |
0.1 | 1 | 2015 | Particle tracking in open simulation laboratories · SC 2015 |
Methods — techniques the papers use, named apart from their topics
task-parallel advection · 0.4data-parallel advection · 0.4summed volumes · 0.3decomposable kernel evaluation · 0.3partial sums · 0.2distributed query evaluation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Remote visual analysis of large turbulence databases at multiple scales
Jesus Pulido, Daniel Livescu, Kalin Kanov, Randal C. Burns, Curtis Canada, James P. Ahrens, Bernd Hamann |
J. Parallel Distributed Comput. | 3 |
| 2015 | Efficient evaluation of threshold queries of derived fields in a numerical simulation databaseabstractIn this paper, we present a method for the ecient evaluation of threshold queries of derived fields for large numerical simulation datasets stored in a cluster of relational databases. The datasets produced by these simulations are in the TB and even PB ranges. Data-intensive computations that examine entire time-steps of the simulation data are impractical to perform locally by the user, taking days or months to iterate over the entire dataset. The integrated method for the evaluation of threshold queries that we have developed achieves scalability through data-parallel execution of the computations on the nodes of an analysis database cluster. We extend the scientific analysis environment with the introduction of an application-aware cache for query results, building on the concept of semantic caching. The cache has little overhead and improves query performance by over an order of magnitude for queries that hit the cache. Caching the results of threshold queries preserves both the I/O and computation e↵ort used to obtain them. In the case of computational turbulence, this allows scientists to quickly focus on the most intense events and interesting regions in any time-step or the dataset as a whole, which greatly speeds up the rate of scientific exploration and discovery. Kalin Kanov, Randal C. Burns, Cristian Constantin Lalescu |
EDBT | 1 |
| 2015 | Particle tracking in open simulation laboratoriesabstractParticle tracking along streamlines and pathlines is a common scientific analysis technique, which has demanding data, computation and communication requirements. It has been studied in the context of high-performance computing due to the difficulty in its efficient parallelization and its high demands on communication and computational load. In this paper, we study efficient evaluation methods for particle tracking in open simulation laboratories. Simulation laboratories have a fundamentally different architecture from today's supercomputers and provide publicly-available analysis functionality. We focus on the I/O demands of particle tracking for numerical simulation datasets 100s of TBs in size. We compare data-parallel and task-parallel approaches for the advection of particles and show scalability results on data-intensive workloads from a live production environment. We have developed particle tracking capabilities for the Johns Hopkins Turbulence Databases, which store computational fluid dynamics simulation data, including forced isotropic turbulence, magnetohydrodynamics, channel flow turbulence and homogeneous buoyancy-driven turbulence. Kalin Kanov, Randal C. Burns |
SC | 1 |
| 2013 | Run-time creation of the turbulent channel flow database by an HPC simulation using MPI-DBabstractWe demonstrate a method based on MPI client-server implementation for storing the output of computations directly into the database. Our method automates the previously used inefficient ingestion process which required development of special tools for each simulation. In large-scale channel flow simulation experiments we were able to ingest the output data sets in real time, without delaying the simulation process. This was accomplished by building a Fortran interface to the MPI-DB software library and using it within the simulation code. The channel flow simulation data set will be exposed for analysis to researchers using the JHU Public Turbulence Database [7]. Jason Graham, Edward Givelberg, Kalin Kanov |
EuroMPI | 3 |
| 2012 | Data-intensive spatial filtering in large numerical simulation datasetsabstractWe present a query processing framework for the efficient evaluation of spatial filters on large numerical simulation datasets stored in a data-intensive cluster. Previously, filtering of large numerical simulations stored in scientific databases has been impractical owing to the immense data requirements. Rather, filtering is done during simulation or by loading snapshots into the aggregate memory of an HPC cluster. Our system performs filtering within the database and supports large filter widths. We present two complementary methods of execution: I/O streaming computes a batch filter query in a single sequential pass using incremental evaluation of decomposable kernels, summed volumes generates an intermediate data set and evaluates each filtered value by accessing only eight points in this dataset. We dynamically choose between these methods depending upon workload characteristics. The system allows us to perform filters against large data sets with little overhead: query performance scales with the cluster's aggregate I/O throughput. Kalin Kanov, Randal C. Burns, Gregory L. Eyink, Charles Meneveau, Alex Szalay |
SC | 1 |
| 2011 | MPI-DB, A Parallel Database Services Software Library for Scientific Computing
Edward Givelberg, Alex Szalay, Kalin Kanov, Randal C. Burns |
EuroMPI | 3 |
| 2011 | I/O streaming evaluation of batch queries for data-intensive computational turbulenceabstractWe describe a method for evaluating computational turbulence queries, including Lagrange Polynomial interpolation, based on partial sums that allows the underlying data to be accessed in any order and in parts. We exploit these properties to stream data from disk in a single pass and concurrently evaluate batch queries. The combination of sequential I/O and data sharing improves performance by an order of magnitude when compared with direct evaluation of each query. The technique also supports distributed evaluation of queries in a database cluster, assembling the partial sums from each node at the query mediator. Interpolation is fundamental to computational turbulence, over 95% of queries use these routines, and the partial sums method allows the JHU Turbulence Database Cluster to realize scale and throughput for our scientists' data-intensive workloads. Kalin Kanov, Eric A. Perlman, Randal C. Burns, Yanif Ahmad, Alex Szalay |
SC | 1 |