Milena Ivanova

dblp:i/MilenaIvanova · also Milena Gateva Koparanova · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
0since 2021 · last 2015
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 10 first-authorSoftware engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Query processing and optimization · 25% Spatial and temporal data management · 20% Indexing and storage engines · 19%
Theoretical computer science
1 paper
Coding theory · 100%

Topics — the 13 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Indexing and storage engines
column store
0.322015
GIS Navigation Boosted by Column Stores · Proc. VLDB Endow. 2015
An architecture for recycling intermediates in a column-store · SIGMOD Conference 2009
Spatial and temporal data management
spatial query processing
0.212015
GIS Navigation Boosted by Column Stores · Proc. VLDB Endow. 2015
Coding theory › error-correcting codes › block codes › linear code › self-dual codes
binary self-dual codes
0.212015
Self-Dual Codes With an Automorphism of Order 11 · IEEE Trans. Inf. Theory 2015
Coding theory › error-correcting codes
code classification
0.212015
Self-Dual Codes With an Automorphism of Order 11 · IEEE Trans. Inf. Theory 2015
Coding theory › error-correcting codes › block codes › linear code
self-dual codes
0.212015
Self-Dual Codes With an Automorphism of Order 11 · IEEE Trans. Inf. Theory 2015
Coding theory › error-correcting codes › block codes › linear code › self-dual codes
self-dual code classification
0.212015
Self-Dual Codes With an Automorphism of Order 11 · IEEE Trans. Inf. Theory 2015
Query processing and optimization › result reuse
intermediate result reuse
0.222010
An architecture for recycling intermediates in a column-store · ACM Trans. Database Syst. 2010
An architecture for recycling intermediates in a column-store · SIGMOD Conference 2009
Data integration and cleaning
extract-transform-load
0.212013
Lazy ETL in Action: ETL Technology Dates Scientific Data · Proc. VLDB Endow. 2013
Data models and query languages › multidimensional database
array DBMS
0.112012
TELEIOS: A Database-Powered Virtual Earth Observatory · Proc. VLDB Endow. 2012
Knowledge graphs
semantic web
0.112012
TELEIOS: A Database-Powered Virtual Earth Observatory · Proc. VLDB Endow. 2012
Spatial and temporal data management
spatial databases
0.112015
GIS Navigation Boosted by Column Stores · Proc. VLDB Endow. 2015
Data stream processing
continuous query processing
0.112005
Customizable Parallel Execution of Scientific Stream Queries · VLDB 2005
High-performance computing
scientific computing systems
0.012005
Customizable Parallel Execution of Scientific Stream Queries · VLDB 2005

Methods — techniques the papers use, named apart from their topics

stSPARQL · 0.3stRDF · 0.3SciQL · 0.3imprints index · 0.2column storage · 0.2automorphism-based construction · 0.2lazy transformation · 0.2lazy loading · 0.2lazy extraction · 0.2column-store · 0.1column store · 0.1materialized views · 0.1
YearPublicationVenuePosition
2015 A Round Table for Multi-disciplinary Research on Geospatial and Climate Data
abstract
Earth observation sciences produce large sets of data which are inherently rich in spatial and geo-spatial information. Together with live data collected from monitoring systems and large collections of semantically rich objects they provide new opportunities for advanced eScience research on climatology, urban planing and smart cities. Such combination of heterogeneous data sets forms a new source of knowledge. Efficient knowledge extraction from them is an eScience challenge. It requires efficient bulk data injection from both static and streaming data sources, dynamic adaptation of the physical and logical schema, efficient methods to correlate spatial and temporal data, and flexibility to (re-)formulate the research question at any time. In this work, we present a data management layer over a column-oriented relational data management system that provides efficient analysis of spatiotemporal data. It provides fast data ingestion through different data loaders, tabular and array based storage, and a dynamic step-wise exploration.
Romulo Goncalves, Milena Ivanova, Foteini Alvanaki, Jason Maassen, Kostis Kyzirakos, Oscar Martinez-Rubi, Hannes Mühleisen
e-Science2
2015 Massive point cloud data management: Design, implementation and execution of a point cloud benchmark
Peter van Oosterom, Oscar Martinez-Rubi, Milena Ivanova, Mike Hörhammer, Daniel Geringer, Siva Ravada, Theo Tijssen, Martin Kodde, Romulo Goncalves
Comput. Graph.3
2015 GIS Navigation Boosted by Column Stores
abstract
Earth observation sciences, astronomy, and seismology have large data sets which have inherently rich spatial and geospatial information. In combination with large collections of semantically rich objects which have a large number of thematic properties, they form a new source of knowledge for urban planning, smart cities and natural resource management. Modeling and storing these properties indicating the relationships between them is best handled in a relational database. Furthermore, the scalability requirements posed by the latest 26-attribute light detection and ranging (LIDAR) data sets are a challenge for file-based solutions. In this demo we show how to query a 640 billion point data set using a column store enriched with GIS functionality. Through a lightweight and cache conscious secondary index called Imprints, spatial queries performance on a flat table storage is comparable to traditional file-based solutions. All the results are visualised in real time using QGIS.
Foteini Alvanaki, Romulo Goncalves, Milena Ivanova, Martin L. Kersten, Kostis Kyzirakos
Proc. VLDB Endow.3
2015 Self-Dual Codes With an Automorphism of Order 11
abstract
Using a method for constructing self-dual codes having an automorphism of odd prime order, we classify up to equivalence all binary self-dual codes with an automorphism of order 11 with 6 cycles and minimum distance 12. This classification gives new [72, 36, 12] codes with weight enumerator that was previously not obtained as well as many [66, 33, 12], [68, 34, 12], and [70, 35, 12] codes with new values of the parameters in their respective weight enumerators.
Nikolay I. Yankov, Moon Ho Lee, Müberra Gürel, Milena Ivanova
IEEE Trans. Inf. Theory4
2013 Data vaults: a database welcome to scientific file repositories
abstract
Efficient management and exploration of high-volume scientific file repositories have become pivotal for advancement in science. We propose to demonstrate the Data Vault, an extension of the database system architecture that transparently opens scientific file repositories for efficient in-database processing and exploration.
Milena Ivanova, Yagiz Kargin, Martin L. Kersten, Stefan Manegold, Ying Zhang 0027, Mihai Datcu, Daniela Espinoza-Molina
SSDBM1
2013 Lazy ETL in Action: ETL Technology Dates Scientific Data
abstract
Both scientific data and business data have analytical needs. Analysis takes place after a scientific data warehouse is eagerly filled with all data from external data sources (repositories). This is similar to the initial loading stage of Extract, Transform, and Load (ETL) processes that drive business intelligence. ETL can also help scientific data analysis. However, the initial loading is a time and resource consuming operation. It might not be entirely necessary, e.g. if the user is interested in only a subset of the data. We propose to demonstrate Lazy ETL, a technique to lower costs for initial loading. With it, ETL is integrated into the query processing of the scientific data warehouse. For a query, only the required data items are extracted, transformed, and loaded transparently on-the-fly. The demo is built around concrete implementations of Lazy ETL for seismic data analysis. The seismic data warehouse is ready for query processing, without waiting for long initial loading. The audience fires analytical queries to observe the internal mechanisms and modifications that realize each of the steps; lazy extraction, transformation, and loading.
Yagiz Karæz, Milena Ivanova, Ying Zhang 0027, Stefan Manegold, Martin L. Kersten
Proc. VLDB Endow.2
2012 Just-In-Time Data Distribution for Analytical Query Processing
Milena Ivanova, Martin L. Kersten, Fabian Groffen
ADBIS1
2012 Data Vaults: A Symbiosis between Database Technology and Scientific File Repositories
Milena Ivanova, Martin L. Kersten, Stefan Manegold
SSDBM1
2012 TELEIOS: A Database-Powered Virtual Earth Observatory
abstract
TELEIOS is a recent European project that addresses the need for scalable access to petabytes of Earth Observation data and the discovery and exploitation of knowledge that is hidden in them. TELEIOS builds on scientific database technologies (array databases, SciQL, data vaults) and Semantic Web technologies (stRDF and stSPARQL) implemented on top of a state of the art column store database system (MonetDB). We demonstrate a first prototype of the TELEIOS Virtual Earth Observatory (VEO) architecture, using a forest fire monitoring application as example.
Manolis Koubarakis, Kostis Kyzirakos, Manos Karpathiotakis, Charalampos Nikolaou, Stavros Vassos, George Garbis, Michael Sioutis, Konstantina Bereta, Dimitrios Michail 0001, Charalambos Kontoes, Ioannis Papoutsis, Themos Herekakis, Stefan Manegold, Martin L. Kersten, Milena Ivanova, Holger Pirk, Ying Zhang 0027, Mihai Datcu, Gottfried Schwarz, Corneliu Octavian Dumitru, Daniela Espinoza-Molina, Katrin Molch, Ugo Di Giammatteo, Manuela Sagona, Sergio Perelli, Thorsten Reitz, Eva Klien, Robert Gregor
Proc. VLDB Endow.15
2011 SciQL: bridging the gap between science and relational DBMS
abstract
Scientific discoveries increasingly rely on the ability to efficiently grind massive amounts of experimental data using database technologies. To bridge the gap between the needs of the Data-Intensive Research fields and the current DBMS technologies, we propose SciQL (pronounced as 'cycle'), the first SQL-based query language for scientific applications with both tables and arrays as first class citizens. It provides a seamless symbiosis of array-, set- and sequence-interpretations. A key innovation is the extension of value-based grouping of SQL:2003 with structural grouping, i.e., fixed-sized and unbounded groups based on explicit relationships between elements positions. This leads to a generalisation of window-based query processing with wide applicability in science domains. This paper describes the main language features of SciQL and illustrates it using time-series concepts.
Ying Zhang 0027, Martin L. Kersten, Milena Ivanova, Niels Nes
IDEAS3
2010 An architecture for recycling intermediates in a column-store
abstract
Automatic recycling of intermediate results to improve both query response time and throughput is a grand challenge for state-of-the-art databases. Tuples are loaded and streamed through a tuple-at-a-time processing pipeline, avoiding materialization of intermediates as much as possible. This limits the opportunities for reuse of overlapping computations to DBA-defined materialized views and function/result cache tuning. In contrast, the operator-at-a-time execution paradigm produces fully materialized results in each step of the query plan. To avoid resource contention, these intermediates are evicted as soon as possible. In this article we study an architecture that harvests the byproducts of the operator-at-a-time paradigm in a column-store system using a lightweight mechanism, the recycler. The key challenge then becomes the selection of the policies to admit intermediates to the resource pool, to determine their retention period, and devise the eviction strategy when facing resource limitations. The proposed recycling architecture has been implemented in an open-source system. An experimental analysis against the TPC-H ad-hoc decision support benchmark and a complex, real-world application (SkyServer) demonstrates its effectiveness in terms of self-organizing behavior and its significant performance gains. The results indicate the potentials of recycling intermediates and charts a route for further development of database kernels.
Milena Ivanova, Martin L. Kersten, Niels Nes, Romulo Goncalves
ACM Trans. Database Syst.1
2009 An architecture for recycling intermediates in a column-store
abstract
Automatically recycling (intermediate) results is a grand challenge for state-of-the-art databases to improve both query response time and throughput. Tuples are loaded and streamed through a tuple-at-a-time processing pipeline avoiding materialization of intermediates as much as possible. This limits the opportunities for reuse of overlapping computations to DBA-defined materialized views and function/result cache tuning.
Milena Ivanova, Martin L. Kersten, Niels Nes, Romulo Goncalves
SIGMOD Conference1
2008 Self-organizing strategies for a column-store database
abstract
Column-store database systems open new vistas for improved maintenance through self-organization. Individual columns are the focal point, which simplify balancing conflicting requirements. This work presents two workload-driven self-organizing techniques in a column-store, i.e. adaptive segmentation and adaptive replication. Adaptive segmentation splits a column into non-overlapping segments based on the actual query load. Likewise, adaptive replication creates segment replicas. The strategies can support different application requirements by trading off the reorganization overhead for storage cost. Both techniques can significantly improve system performance as demonstrated in an evaluation of different scenarios.
Milena Ivanova, Martin L. Kersten, Niels Nes
EDBT1
2008 Adaptive Segmentation for Scientific Databases
abstract
In this paper we explore database segmentation in the context of a column-store DBMS targeted at a scientific database. We present a novel hardware- and scheme-oblivious segmentation algorithm, which learns and adapts to the workload immediately. The approach taken is to capitalize on (intermediate) query results, such that future queries benefit from a more appropriate data layout. The algorithm is implemented as an extension of a complete DBMS and evaluated against a real-life workload. It demonstrates significant performance gains without DBA assistance.
Milena Ivanova, Martin L. Kersten, Niels Nes
ICDE1
2007 MonetDB/SQL Meets SkyServer: the Challenges of a Scientific Database
abstract
This paper presents our experiences in porting the Sloan Digital Sky Survey(SDSS)/ SkyServer to the state-of- the-art open source database system MonetDB/SQL. SDSS acts as a well-documented benchmark for scientific database management. We have achieved a fully functional prototype for the personal SkyServer, to be downloaded from our site. The lessons learned are 1) the column store approach of MonetDB demonstrates a great potential in the world of scientific databases. However, the application also challenged the functionality of our implementation and revealed that a fully operational SQL environment is needed, e.g. including persistent stored modules; 2) the initial performance is competitive to the reference platform, MS SQL Server 2005, and 3) the analysis of SDSS query traces hints at several techniques to boost performance by utilizing repetitive behavior and zoom-in/zoom-out access patterns, that are currently not captured by the system.
Milena Ivanova, Niels Nes, Romulo Goncalves, Martin L. Kersten
SSDBM1
2005 Customizable Parallel Execution of Scientific Stream Queries
Milena Ivanova, Tore Risch
VLDB1
2002 Completing CAD Data Queries for Visualization
abstract
A system has been developed permitting database queries over data extracted from a CAD system where the query result is returned back to the CAD for visualization and analysis. This has several challenges. First, CAD data representations use complex object-oriented schemas and the query language must be object-oriented too. Second, the query system resides outside the CAD system and must therefore use standardized data exchange formats for interoperability with the CAD. ISO STEP standard exchange formats are used for the exchange. Third, a CAD system cannot import an arbitrary object structure but places restrictions on the imported objects to be acceptable. Therefore, the query system must complement the query results in order to produce an acceptable CAD model, called the model completion of the query. These problems have been solved using an extensible object-relational query processor. The system also supports queries combining CAD data with data from other data sources.
Milena Ivanova, Tore Risch
IDEAS1