Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ioana Ileana

dblp:133/8432 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0003-0554-8748ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 53% Database system architecture and tuning · 17% Indexing and storage engines · 17%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
similarity search
1.122025
Evaluating and Generating Query Workloads for High Dimensional Vector Similarity Search · KDD (2) 2025
LeaFi: Data Series Indexes on Steroids with Learned Filters · Proc. ACM Manag. Data 2025
Indexing and storage engines › temporal indexing
data series indexing
0.912025
LeaFi: Data Series Indexes on Steroids with Learned Filters · Proc. ACM Manag. Data 2025
Information retrieval › similarity search
high-dimensional similarity search
0.912025
Evaluating and Generating Query Workloads for High Dimensional Vector Similarity Search · KDD (2) 2025
Database system architecture and tuning
query workload generation
0.912025
Evaluating and Generating Query Workloads for High Dimensional Vector Similarity Search · KDD (2) 2025
Information retrieval › evaluation
benchmark
0.312025
Evaluating and Generating Query Workloads for High Dimensional Vector Similarity Search · KDD (2) 2025
Information retrieval › similarity search › sequence similarity search
time series similarity search
0.312025
LeaFi: Data Series Indexes on Steroids with Learned Filters · Proc. ACM Manag. Data 2025
Query processing and optimization › query rewriting
query answering using views
0.212014
Complete yet practical search for minimal query reformulations under constraints · SIGMOD Conference 2014
Information retrieval
query reformulation
0.212014
Complete yet practical search for minimal query reformulations under constraints · SIGMOD Conference 2014
Query processing and optimization › semantic query processing
semantic query optimization
0.212014
Complete yet practical search for minimal query reformulations under constraints · SIGMOD Conference 2014

Methods — techniques the papers use, named apart from their topics

simulated annealing · 0.9machine learning models · 0.9gradient-based optimization · 0.9provenance · 0.2
YearPublicationVenuePosition
2025 Evaluating and Generating Query Workloads for High Dimensional Vector Similarity Search
abstract
Similarity search lies at the heart of many modern applications, ranging from databases to deep learning to data series analysis. As such, a vast effort has been invested in developing algorithms, data structures and implementations to speed up this crucial subroutine. To empirically validate these approaches, several benchmarking efforts have been initiated covering a wide array of datasets. In this paper, we observe that usually little control is exercised on the hardness of the workloads with which methods are tested and compared. To address this issue, we first evaluate several query hardness measures with respect to their ability to capture the empirical hardness of a query, i.e. the effort invested by an index data structure to provide an answer. Then, we propose two methods, deemed Hephaestus-Annealing and Hephaestus-Gradient, for synthesizing query workloads so that they meet a user-specified hardness target. Both methods allow to produce workloads with the desired hardness: we find that Hephaestus-Gradient is faster, while Hephaestus-Annealing makes fewer assumptions on the target hardness measure. The resulting workloads can be used to gain insights into the behavior of similarity search algorithms.
Matteo Ceccarello, Alexandra Levchenko, Ioana Ileana, Themis Palpanas
KDD (2)3
2025 LeaFi: Data Series Indexes on Steroids with Learned Filters
abstract
The ever-growing collections of data series create a pressing need for efficient similarity search, which serves as the backbone for various analytics pipelines. Recent studies have shown that tree-based series indexes excel in many scenarios. However, we observe a significant waste of effort during search, due to suboptimal pruning. To address this issue, we introduce LeaFi, a novel framework that uses machine learning models to boost pruning effectiveness of tree-based data series indexes. These models act as learned filters, which predict tight node-wise distance lower bounds that are used to make pruning decisions, thus, improving pruning effectiveness. We describe the LeaFi-enhanced index building algorithm, which selects leaf nodes and generates training data to insert and train machine learning models, as well as the LeaFi-enhanced search algorithm, which calibrates learned filters at query time to support the user-defined quality target of each query. Our experimental evaluation, using two different tree-based series indexes and five diverse datasets, demonstrates the advantages of the proposed approach. LeaFi-enhanced data-series indexes improve pruning ratio by up to 20x and search time by up to 32x, while maintaining a target recall of 99%.
Qitong Wang 0003, Ioana Ileana, Themis Palpanas
Proc. ACM Manag. Data2
2018 Generating data series query workloads
Kostas Zoumpatianos, Yin Lou, Ioana Ileana, Themis Palpanas, Johannes Gehrke
VLDB J.3
2017 ChaseFUN: a Data Exchange Engine for Functional Dependencies at Scale
abstract
International audience
Angela Bonifati, Ioana Ileana, Michele Linardi
EDBT2
2016 Functional Dependencies Unleashed for Scalable Data Exchange
abstract
We address the problem of efficiently evaluating target functional dependencies (fds) in the Data Exchange (DE) process. Target fds naturally occur in many DE scenarios, including the ones in Life Sciences in which multiple source relations need to be structured under a constrained target schema. However, despite their wide use, target fds' evaluation is still a bottleneck in the state-of-the-art DE engines. Systems relying on an all-SQL approach typically do not support target fds unless additional information is provided. Alternatively, DE engines that do include these dependencies typically pay the price of a significant drop in performance and scalability. In this paper, we present a novel chase-based algorithm that can efficiently handle arbitrary fds on the target. Our approach essentially relies on exploiting the interactions between source-to-target (s-t) tuple-generating dependencies (tgds) and target fds. This allows us to tame the size of the intermediate chase results, by playing on a careful ordering of chase steps interleaving fds and (chosen) tgds. As a direct consequence, we importantly diminish the fd application scope, often a central cause of the dramatic overhead induced by target fds. Moreover, reasoning on dependency interaction further leads us to interesting parallelization opportunities, yielding additional scalability gains. We provide a proof-of-concept implementation of our chase-based algorithm and an experimental study aimed at gauging its scalability and efficiency. Finally, we empirically compare with the latest DE engines, and show that our algorithm outperforms them.
Angela Bonifati, Ioana Ileana, Michele Linardi
SSDBM2
2015 Invisible Glue: Scalable Self-Tunning Multi-Stores
Francesca Bugiotti, Damian Bursztyn, Alin Deutsch, Ioana Ileana, Ioana Manolescu
CIDR4
2014 Complete yet practical search for minimal query reformulations under constraints
abstract
We revisit the Chase&Backchase (C&B) algorithm for query reformulation under constraints, which provides a uniform solution to such particular-case problems as view-based rewriting under constraints, semantic query optimization, and physical access path selection in query optimization. For an important class of queries and constraints, C&B has been shown to be complete, i.e. guaranteed to find all (join-)minimal reformulations under constraints. C&B is based on constructing a canonical rewriting candidate called a universal plan, then inspecting its exponentially many sub-queries in search for minimal reformulations, essentially removing redundant joins in all possible ways. This inspection involves chasing the subquery. Because of the resulting exponentially many chases, the conventional wisdom has held that completeness is a concept of mainly theoretical interest. We show that completeness can be preserved at practically relevant cost by introducing Prov-C&B, a novel reformulation algorithm that instruments the chase to maintain provenance information connecting the joins added during the chase to the universal plan subqueries responsible for adding these joins. This allows it to directly "read off" the minimal reformulations from the result of a single chase of the universal plan, saving exponentially many chases of its subqueries. We exhibit natural scenarios yielding speedups of over two orders of magnitude between the execution of the best view-based rewriting found by a commercial query optimizer and that of the best rewriting found by Prov-C&B (which the optimizer misses because of limited reasoning about constraints).
Ioana Ileana, Bogdan Cautis, Alin Deutsch, Yannis Katsis
SIGMOD Conference1