Marco Patella

dblp:p/MPatella · DBLP profile ↗
← Back
27ranked-venue papers
0as first author
1since 2021 · last 2026
0000-0003-2655-0759ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 19Artificial intelligence and machine learning · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Applied, interdisciplinary, general and emerging computing · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 67% Language models and text generation · 33%
Databases, data mining, and information retrieval
7 papers
Query processing and optimization · 72% Data models and query languages · 11% Indexing and storage engines · 11%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Smart cities and intelligent transportation · 34% Bioinformatics and computational biology · 34% Computational social science and digital humanities · 31%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 20 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › large language model trustworthiness
large language model robustness
1.012026
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? · ACL (1) 2026
Natural language and speech › Language models and text generation › prompting
prompt sensitivity
1.012026
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? · ACL (1) 2026
Machine learning › Trustworthy machine learning
robustness
1.012026
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? · ACL (1) 2026
Query processing and optimization › preference query
skyline query
0.432014
Domination in the Probabilistic World: Computing Skylines for Arbitrary Correlations and Ranking Semantics · ACM Trans. Database Syst. 2014
The Skyline of a Probabilistic Relation · IEEE Trans. Knowl. Data Eng. 2013
Efficient sort-based skyline evaluation · ACM Trans. Database Syst. 2008
Query processing and optimization › preference query › skyline query
probabilistic skyline
0.422014
Domination in the Probabilistic World: Computing Skylines for Arbitrary Correlations and Ranking Semantics · ACM Trans. Database Syst. 2014
The Skyline of a Probabilistic Relation · IEEE Trans. Knowl. Data Eng. 2013
Computational social science and digital humanities › legal informatics
legal text analysis
0.312026
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? · ACL (1) 2026
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis
0.212016
Pattern Similarity Search in Genomic Sequences · IEEE Trans. Knowl. Data Eng. 2016
Algorithms and data structures › sequence algorithms
string algorithms
0.212016
Pattern Similarity Search in Genomic Sequences · IEEE Trans. Knowl. Data Eng. 2016
Data models and query languages › uncertain data
probabilistic data
0.212013
The Skyline of a Probabilistic Relation · IEEE Trans. Knowl. Data Eng. 2013
Indexing and storage engines
metric space indexing
0.142002
Searching in metric spaces with user-defined and approximate distances · ACM Trans. Database Syst. 2002
PAC Nearest Neighbor Queries: Approximate and Controlled Search in High-Dimensional and Metric Spaces · ICDE 2000
A Cost Model for Similarity Queries in Metric Spaces · PODS 1998
Bioinformatics and computational biology › gene regulation › regulatory element discovery
enhancer prediction
0.112016
Pattern Similarity Search in Genomic Sequences · IEEE Trans. Knowl. Data Eng. 2016
User interface design and tools › multimedia interface design
image browsing interface
0.112007
PIBE: Manage Your Images the Way You Want! · ICDE 2007
Multimedia analysis and retrieval › multimedia retrieval › content-based retrieval
shape retrieval
0.112005
WARP: Accurate Retrieval of Shapes Using Phase of Fourier Descriptors and Time Warping Distance · IEEE Trans. Pattern Anal. Mach. Intell. 2005
Information retrieval
similarity search
0.022000
PAC Nearest Neighbor Queries: Approximate and Controlled Search in High-Dimensional and Metric Spaces · ICDE 2000
M-tree: An Efficient Access Method for Similarity Search in Metric Spaces · VLDB 1997
Query processing and optimization
similarity query processing
0.012002
Searching in metric spaces with user-defined and approximate distances · ACM Trans. Database Syst. 2002
Indexing and storage engines
vector index
0.012002
Searching in metric spaces with user-defined and approximate distances · ACM Trans. Database Syst. 2002
Query processing and optimization › similarity query processing
approximate nearest neighbor query
0.012000
PAC Nearest Neighbor Queries: Approximate and Controlled Search in High-Dimensional and Metric Spaces · ICDE 2000
Indexing and storage engines › multidimensional indexing
high-dimensional indexing
0.012000
PAC Nearest Neighbor Queries: Approximate and Controlled Search in High-Dimensional and Metric Spaces · ICDE 2000
Spatial and temporal data management › spatial query processing
nearest neighbor query
0.012000
PAC Nearest Neighbor Queries: Approximate and Controlled Search in High-Dimensional and Metric Spaces · ICDE 2000
Query processing and optimization
cost estimation
0.011998
A Cost Model for Similarity Queries in Metric Spaces · PODS 1998

Methods — techniques the papers use, named apart from their topics

large language model prompting · 2.0window-based technique · 0.5dynamic programming · 0.5multidisciplinary approach · 0.3optimization · 0.2dominance rules · 0.2order theory · 0.2expected score semantics · 0.2expected rank semantics · 0.2presorting · 0.1dominance testing · 0.1persistent storage · 0.1hierarchical organization · 0.1phase of fourier coefficients · 0.1fourier descriptors · 0.1dynamic time warping · 0.1metric space indexing · 0.0cost modeling · 0.0
YearPublicationVenuePosition
2026 Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards?
abstract
Lorenzo Molfetta, Alessio Cocchieri, Luca Ragazzi, Ilaria Bartolini, Marco Patella, Gianluca Moro. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lorenzo Molfetta, Alessio Cocchieri, Luca Ragazzi, Ilaria Bartolini, Marco Patella, Gianluca Moro
ACL (1)5
2020 Rammed, or What RAM3S Taught Us
abstract
RAM3S (Real-time Analysis of Massive MultiMedia Streams) is a framework that acts as a middleware software layer between multimedia stream analysis techniques and Big Data streaming platforms, so as to facilitate the implementation of the former on top of the latter. Indeed, the use of Big Data platforms can give way to the efficient management and analysis of large data amounts, but they require the user to concentrate on issues related to distributed computing, since their services are often too raw. The use of RAM3S greatly simplifies deploying non-parallel techniques to platforms like Apache Storm or Apache Flink, a fact that is demonstrated by the four different use cases we describe here. We detail the lessons we learned from exploiting RAM3S to implement the detailed use cases.
Ilaria Bartolini, Marco Patella
iiWAS2
2020 Search and comparison of (epi)genomic feature patterns in multiple genome browser tracks
abstract
BACKGROUND: Genome browsers are widely used for locating interesting genomic regions, but their interactive use is obviously limited to inspecting short genomic portions. An ideal interaction is to provide patterns of regions on the browser, and then extract other genomic regions over the whole genome where such patterns occur, ranked by similarity. RESULTS: We developed SimSearch, an optimized pattern-search method and an open source plugin for the Integrated Genome Browser (IGB), to find genomic region sets that are similar to a given region pattern. It provides efficient visual genome-wide analytics computation in large datasets; the plugin supports intuitive user interactions for selecting an interesting pattern on IGB tracks and visualizing the computed occurrences of similar patterns along the entire genome. SimSearch also includes functions for the annotation and enrichment of results, and is enhanced with a Quickload repository including numerous epigenomic feature datasets from ENCODE and Roadmap Epigenomics. The paper also includes some use cases to show multiple genome-wide analyses of biological interest, which can be easily performed by taking advantage of the presented approach. CONCLUSIONS: The novel SimSearch method provides innovative support for effective genome-wide pattern search and visualization; its relevance and practical usefulness is demonstrated through a number of significant use cases of biological interest. The SimSearch IGB plugin, documentation, and code are freely available at https://deib-geco.github.io/simsearch-app/ and https://github.com/DEIB-GECO/simsearch-app/ .
Arnaud Céol, Piero Montanari, Ilaria Bartolini, Stefano Ceri, Paolo Ciaccia, Marco Patella, Marco Masseroli
BMC Bioinform.6
2019 Multiple Instance Classification in the Image Domain
Ilaria Bartolini, Pietro Pascarella, Marco Patella
SISAP3
2018 A general framework for real-time analysis of massive multimedia streams
Ilaria Bartolini, Marco Patella
Multim. Syst.2
2018 Windsurf: the best way to SURF - (and SIFT/BRISK/ORB/FREAK, too)
Ilaria Bartolini, Marco Patella
Multim. Syst.2
2018 The Need of Multidisciplinary Approaches and Engineering Tools for the Development and Implementation of the Smart City Paradigm
abstract
This paper is motivated by the concept that the successful, effective, and sustainable implementation of the smart city paradigm requires a close cooperation among researchers with different, complementary interests and, in most cases, a multidisciplinary approach. It first briefly discusses how such a multidisciplinary methodology, transversal to various disciplines such as architecture, computer science, civil engineering, electrical, electronic and telecommunication engineering, social science and behavioral science, etc., can be successfully employed for the development of suitable modeling tools and real solutions of such sociotechnical systems. Then, the paper presents some pilot projects accomplished by the authors within the framework of some major European Union (EU) and national research programs, also involving the Bologna municipality and some of the key players of the smart city industry. Each project, characterized by different and complementary approaches/modeling tools, is illustrated along with the relevant contextualization and the advancements with respect to the state of the art.
Oreste Andrisano, Ilaria Bartolini, Paolo Bellavista, Andrea Boeri, Luciano Bononi, Alberto Borghetti, Armando Brath, Giovanni Emanuele Corazza, Antonio Corradi, Stefano de Miranda, Fabio Fava, Luca Foschini 0001, Giovanni Leoni 0002, Danila Longo, Michela Milano, Fabio Napolitano, Carlo Alberto Nucci, Gianni Pasolini, Marco Patella, Tullio Salmon Cinotti, Daniele Tarchi, Francesco Ubertini, Daniele Vigo
Proc. IEEE19
2017 The Power of Distance Distributions: Cost Models and Scheduling Policies for Quality-Controlled Similarity Queries
Paolo Ciaccia, Marco Patella
SISAP2
2016 Pattern Similarity Search in Genomic Sequences
abstract
Genomics, with the high amount of heterogeneous data that it is generating, is opening many interesting practical and theoretical computational problems; one of them is the search for a collection of genomic regions at given distances from each other, i.e., a pattern of genomic regions, along the whole genome. In this paper, we present an optimized pattern-search algorithm able to find efficiently, within a large set of genomic data, genomic region sequences which are similar to a given pattern. We start with a base version of the problem, which is solved using dynamic programming enhanced with an efficient window-based technique; then, we extend the algorithm to more complex scenarios with practical applications in revealing interesting and unknown regions of the genome, thus, making it an important ingredient in supporting biological research. We apply our algorithm to enhancer detection, a relevant biological problem, showing that the method is both efficient and accurate.
Piero Montanari, Ilaria Bartolini, Paolo Ciaccia, Marco Patella, Stefano Ceri, Marco Masseroli
IEEE Trans. Knowl. Data Eng.4
2014 Domination in the Probabilistic World: Computing Skylines for Arbitrary Correlations and Ranking Semantics
abstract
In a probabilistic database, deciding if a tuple u is better than another tuple v has not a univocal solution, rather it depends on the specific Probabilistic Ranking Semantics (PRS) one wants to adopt so as to combine together tuples' scores and probabilities. In deterministic databases it is known that skyline queries are a remarkable alternative to (top- k ) ranking queries, because they remove from the user the burden of specifying a scoring function that combines values of different attributes into a single score. The skyline of a deterministic relation R is the set of undominated tuples in R -- tuple u dominates tuple v iff on all the attributes of interest u is better than or equal to v and strictly better on at least one attribute. Domination is equivalent to having s ( u ) ≥ s ( v ) for all monotone scoring functions s (). The skyline of a probabilistic relation R p can be similarly defined as the set of P-undominated tuples in R p , where now u P-dominates v iff, whatever monotone scoring function one would use to combine the skyline attributes, u is reputed better than v by the PRS at hand. This definition, which is applicable to arbitrary ranking semantics and probabilistic correlation models, is parametric in the adopted PRS, thus it ensures that ranking and skyline queries will always return consistent results. In this article we provide an overall view of the problem of computing the skyline of a probabilistic relation. We show how, under mild conditions that indeed hold for all known PRSs, checking P-domination can be cast into an optimization problem, whose complexity we characterize for a variety of combinations of ranking semantics and correlation models. For each analyzed case we also provide specific P-domination rules , which are exploited by the algorithm we detail for the case where the probabilistic model is known to the query processor. We also consider the case in which the probability of tuple events can only be obtained through an oracle, and describe another skyline algorithm for this loosely integrated scenario. Our experimental evaluation of P-domination rules and skyline algorithms confirms the theoretical analysis.
Ilaria Bartolini, Paolo Ciaccia, Marco Patella
ACM Trans. Database Syst.3
2013 SHIATSU: tagging and retrieving videos without worries
Ilaria Bartolini, Marco Patella, Corrado Romani
Multim. Tools Appl.2
2013 The Skyline of a Probabilistic Relation
abstract
In a deterministic relation, tuple u dominates tuple v if u is no worse than v on all attributes, and better than v on at least one attribute. This concept is at the heart of skyline queries, that return the set of undominated tuples. In this paper we extend the notion of skyline to probabilistic relations by generalizing to this context the definition of tuple domination. Our approach is parametric in the semantics for ranking probabilistic tuples and, being it based on order-theoretic principles, preserves the three properties the skyline has in the deterministic case: it equals the union of all top1 results of monotone scoring functions, it requires no additional parameter, and it is insensitive to attribute scales. We then show how domination among probabilistic tuples can be efficiently checked by means of a set of rules. We detail rules for the cases in which tuples are ranked using either the "expected rank" or the "expected score" semantics, and explain how the approach can be applied to other semantics as well. Since computing the skyline of a probabilistic relation is a time-consuming task, we introduce algorithms for checking domination rules in an optimized way. Experiments show that these algorithms can reduce execution times with respect to a naive evaluation.
Ilaria Bartolini, Paolo Ciaccia, Marco Patella
IEEE Trans. Knowl. Data Eng.3
2011 Metric information filtering
Paolo Ciaccia, Marco Patella
Inf. Syst.2
2010 Query processing issues in region-based image databases
Ilaria Bartolini, Paolo Ciaccia, Marco Patella
Knowl. Inf. Syst.3
2009 The Panda framework for Comparing Patterns
Ilaria Bartolini, Paolo Ciaccia, Eirini Ntoutsi, Marco Patella, Yannis Theodoridis
Data Knowl. Eng.4
2008 Efficient sort-based skyline evaluation
abstract
Skyline queries compute the set of Pareto-optimal tuples in a relation, that is, those tuples that are not dominated by any other tuple in the same relation. Although several algorithms have been proposed for efficiently evaluating skyline queries, they either necessitate the relation to have been indexed or have to perform the dominance tests on all the tuples in order to determine the result. In this article we introduce salsa, a novel skyline algorithm that exploits the idea of presorting the input data so as to effectively limit the number of tuples to be read and compared. This makes salsa also attractive when skyline queries are executed on top of systems that do not understand skyline semantics, or when the skyline logic runs on clients with limited power and/or bandwidth. We prove that, if one considers symmetric sorting functions, the number of tuples to be read is minimized by sorting data according to a “minimum coordinate,” minC, criterion, and that performance can be further improved if data distribution is known and an asymmetric sorting function is used. Experimental results obtained on synthetic and real datasets show that salsa consistently outperforms state-of-the-art sequential skyline algorithms and that its performance can be accurately predicted.
Ilaria Bartolini, Paolo Ciaccia, Marco Patella
ACM Trans. Database Syst.3
2007 PIBE: Manage Your Images the Way You Want!
abstract
A customizable system for image browsing, named PIBE, is proposed. In details, PIBE provides the user with a set of browsing and personalization facilities that enable an effective and efficient exploration of the image collection. The approach is novel and appealing because: 1) the personalization actions over the hierarchical organization of images are local, 2) the storage of the browsing structure is persistent, and 3) the provided GUI makes browsing and personalization facilities extremely intuitive and "easy-to-use".
Ilaria Bartolini, Paolo Ciaccia, Marco Patella
ICDE3
2006 SaLSa: computing the skyline without scanning the whole sky
abstract
Skyline queries compute the set of Pareto-optimal tuples in a relation, ie those tuples that are not dominated by any other tuple in the same relation. Although several algorithms have been proposed for efficiently evaluating skyline queries, they either require to extend the relational server with specialized access methods (which is not always feasible) or have to perform the dominance tests on all the tuples in order to determine the result. In this paper we introduce SaLSa (Sort and Limit Skyline algorithm), which exploits the sorting machinery of a relational engine to order tuples so that only a subset of them needs to be examined for computing the skyline result. This makes SaLSa particularly attractive when skyline queries are executed on top of systems that do not understand skyline semantics or when the skyline logic runs on clients with limited power and/or bandwidth.
Ilaria Bartolini, Paolo Ciaccia, Marco Patella
CIKM3
2006 Adaptively browsing image databases with PIBE
Ilaria Bartolini, Paolo Ciaccia, Marco Patella
Multim. Tools Appl.3
2005 WARP: Accurate Retrieval of Shapes Using Phase of Fourier Descriptors and Time Warping Distance
abstract
Effective and efficient retrieval of similar shapes from large image databases is still a challenging problem in spite of the high relevance that shape information can have in describing image contents. In this paper, we propose a novel Fourier-based approach, called WARP, for matching and retrieving similar shapes. The unique characteristics of WARP are the exploitation of the phase of Fourier coefficients and the use of the Dynamic Time Warping (DTW) distance to compare shape descriptors. While phase information provides a more accurate description of object boundaries than using only the amplitude of Fourier coefficients, the DTW distance permits us to accurately match images even in the presence of (limited) phase shiftings. In terms of classical precision/recall measures, we experimentally demonstrate that WARP can gain, say, up to 35 percent in precision at a 20 percent recall level with respect to Fourier-based techniques that use neither phase nor DTW distance.
Ilaria Bartolini, Paolo Ciaccia, Marco Patella
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 A Unified and Flexible Framework for Comparing Simple and Complex Patterns
Ilaria Bartolini, Paolo Ciaccia, Eirini Ntoutsi, Marco Patella, Yannis Theodoridis
PKDD4
2002 String Matching with Metric Trees Using an Approximate Distance
Ilaria Bartolini, Paolo Ciaccia, Marco Patella
SPIRE3
2002 Searching in metric spaces with user-defined and approximate distances
abstract
Novel database applications, such as multimedia, data mining, e-commerce, and many others, make intensive use of similarity queries in order to retrieve the objects that better fit a user request. Since the effectiveness of such queries improves when the user is allowed to personalize the similarity criterion according to which database objects are evaluated and ranked, the development of access methods able to efficiently support user-defined similarity queries becomes a basic requirement. In this article we introduce the first index structure, called the QIC-M-tree, that can process user-defined queries in generic metric spaces, that is, where the only information about indexed objects is their relative distances. The QIC-M-tree is a metric access method that can deal with several distinct distances at a time: (1) a query (user-defined) distance , (2) an index distance (used to build the tree), and (3) a comparison (approximate) distance (used to quickly discard from the search uninteresting parts of the tree). We develop an analytical cost model that accurately characterizes the performance of the QIC-M-tree and validate such model through extensive experimentation on real metric data sets. In particular, our analysis is able to predict the best evaluation strategy (i.e., which distances to use) under a variety of configurations, by properly taking into account relevant factors such as the distribution of distances, the cost of computing distances, and the actual index structure. We also prove that the overall saving in CPU search costs when using an approximate distance can be estimated by using information on the data set only (thus such measure is independent of the underlying access method) and show that performance results are closely related to a novel "indexing" error measure.
Paolo Ciaccia, Marco Patella
ACM Trans. Database Syst.2
2000 PAC Nearest Neighbor Queries: Approximate and Controlled Search in High-Dimensional and Metric Spaces
abstract
In high-dimensional and complex metric spaces, determining the nearest neighbor (NN) of a query object q can be a very expensive task, because of the poor partitioning operated by index structures-the so-called "curse of dimensionality". This also affects approximately correct (AC) algorithms, which return as results a point whose distance from q is less than (1+/spl epsiv/) times the distance between q and its true NN. In this paper we introduce a new approach to approximate similarity search, called PAC-NN queries, where the error bound /spl epsiv/ can be exceeded with probability /spl delta/ and both /spl epsiv/ and /spl delta/ parameters can be tuned at query time to trade the quality of the result for the cost of the search. We describe sequential and index-based PAC-NN algorithms that exploit the distance distribution of the query object in order to determine a stopping condition that respects the error bound. Analysis and experimental evaluation of the sequential algorithm confirm that, for moderately large data sets and suitable /spl epsiv/ and /spl delta/ values, PAC-NN queries can be efficiently solved and the error controlled. Then, we provide experimental evidence that indexing can further speed-up the retrieval process by up to 1-2 orders of magnitude without giving up the accuracy of the result.
Paolo Ciaccia, Marco Patella
ICDE2
1998 Processing Complex Similarity Queries with Distance-Based Access Methods
Paolo Ciaccia, Marco Patella, Pavel Zezula
EDBT2
1998 A Cost Model for Similarity Queries in Metric Spaces
abstract
We consider the problem of estimating CPU (distance computations) and I/O costs for processing range and k-nearest neighbors queries over metric spaces. Unlike the specific case of vector spaces, where information on data distribution has been exploited to derive cost models for predicting the performance of multi-dimensional access methods, in a generic metric space there is no such a possibility, which makes the problem quite different and requires a novel approach. We insist that the distance distribution of objects can be profitably used to solve the problem, and consequently develop a concrete cost model for the M-tree access method [10]. Our results rely on the assumption that the indexed dataset comes from a metric space which is "homogeneous" enough (in a probabilistic sense) to allow reliable cost estimations even if the distance distribution with respect to a specific query object is unknown. We experimentally validate the model over both real and synthetic datasets, and sho...
Paolo Ciaccia, Marco Patella, Pavel Zezula
PODS2
1997 M-tree: An Efficient Access Method for Similarity Search in Metric Spaces
Paolo Ciaccia, Marco Patella, Pavel Zezula
VLDB2