Alexe Dumitru-Bogdan

dblp:05/6762 · also Bogdan Alexe · DBLP profile ↗
← Back
31ranked-venue papers
18as first author
4since 2021 · last 2026
0009-0004-9732-8770ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 17 · 13 first-authorArtificial intelligence and machine learning · 13 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
14 papers
Data integration and cleaning · 88% Database theory · 7% Data models and query languages · 2%
Artificial intelligence
10 papers
Image recognition and object detection · 67% Video understanding and tracking · 13% Information extraction and text analysis · 7%

Topics — the 21 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
schema mapping
1.1112012
MapMerge: correlating independent schema mappings · VLDB J. 2012
Characterizing schema mappings via data examples · ACM Trans. Database Syst. 2011
EIRENE: Interactive Design and Refinement of Schema Mappings via Data Examples · Proc. VLDB Endow. 2011
Computer vision › Image recognition and object detection › object localization
weakly supervised object localization
0.422016
How Hard Can It Be? Estimating the Difficulty of Visual Search in an Image · CVPR 2016
Weakly Supervised Localization and Learning with Generic Knowledge · Int. J. Comput. Vis. 2012
Computer vision › Image recognition and object detection
object detection
0.322012
Measuring the Objectness of Image Windows · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Searching for objects driven by context · NIPS 2012
Computer vision › Video understanding and tracking
video anomaly detection
0.312017
Unmasking the Abnormal Events in Video · ICCV 2017
Computer vision › Image recognition and object detection › object detection
objectness estimation
0.322012
Measuring the Objectness of Image Windows · IEEE Trans. Pattern Anal. Mach. Intell. 2012
What is an object? · CVPR 2010
Data integration and cleaning
temporal data integration
0.212014
Preference-aware Integration of Temporal Data · Proc. VLDB Endow. 2014
Database theory › dependency theory
tuple-generating dependencies
0.112011
Characterizing schema mappings via data examples · ACM Trans. Database Syst. 2011
Computer vision › Image recognition and object detection
object localization
0.112010
Localizing Objects While Learning Their Appearance · ECCV (4) 2010
Machine learning › Learning paradigms
unsupervised learning
0.112010
Localizing Objects While Learning Their Appearance · ECCV (4) 2010
Data integration and cleaning › schema mapping
mapping composition
0.112010
MapMerge: Correlating Independent Schema Mappings · Proc. VLDB Endow. 2010
Data integration and cleaning › schema mapping
mapping refinement
0.112008
Muse: Mapping Understanding and deSign by Example · ICDE 2008
Data integration and cleaning
data exchange
0.122010
Characterizing schema mappings via data examples · PODS 2010
Comparing and evaluating mapping systems with STBenchmark · Proc. VLDB Endow. 2008
Query processing and optimization › query rewriting › query transformation
query composition
0.112005
Sub-document queries over XML with XSQirrel · WWW 2005
Data models and query languages
XML query languages
0.112005
Sub-document queries over XML with XSQirrel · WWW 2005
Medical and health informatics
electronic health records
0.012004
An Electronic Patient Record "on Steroids": Distributed, Peer-to-Peer, Secure and Privacy-conscious · VLDB 2004
Computer vision › 3D vision
approximate nearest neighbor
0.012011
Exploiting spatial overlap to efficiently compute appearance distances between image windows · NIPS 2011
Computer vision › Segmentation and scene understanding
saliency detection
0.012010
What is an object? · CVPR 2010
Data integration and cleaning
schema matching
0.012010
MapMerge: Correlating Independent Schema Mappings · Proc. VLDB Endow. 2010
Image and video processing
image segmentation
0.012010
ClassCut for Unsupervised Class Segmentation · ECCV (5) 2010
Data integration and cleaning › schema mapping
schema mapping discovery
0.012008
STBenchmark: towards a benchmark for mapping systems · Proc. VLDB Endow. 2008
Debugging and program repair
fault localization
0.012006
SPIDER: a Schema mapPIng DEbuggeR · VLDB 2006

Methods — techniques the papers use, named apart from their topics

entity resolution · 0.3unmasking · 0.3binary classifier · 0.3regression · 0.2deep features · 0.2preference rules · 0.2polynomial-delay enumeration · 0.2sequential observation · 0.1saliency · 0.1information extraction · 0.1context modeling · 0.1bayesian framework · 0.1HOG detector · 0.1model theory · 0.1homomorphism duality · 0.1example-driven mapping refinement · 0.1complexity analysis · 0.1GLAV constraints · 0.1
YearPublicationVenuePosition
2026 Referenceless Evaluation of Machine Translation Models by Ranking Performance in Romanian to English Translate-train Settings
Mihail Feraru, Alexandra Diaconu, Alexe Dumitru-Bogdan
LREC3
2025 Map Generation From Overlapping Microscopy Images Using Stitching Methods
abstract
High-resolution microscopy images play an important role in histopathology and biomedical research. Their limited field of view necessitates stitching multiple overlapping image tiles to form larger, coherent image maps. In this work, we propose and evaluate a suite of image stitching methods, ranging from classical feature-based approaches such as SIFT and KAZE to modern deep learning-based techniques like SuperPoint and LightGlue, for generating high-quality composite microscopy images. We introduce optimizations for stitching order and patch matching to improve alignment robustness and computational efficiency. The proposed methods are benchmarked on two microscopy datasets, including one with ground-truth labels, using both traditional and learned similarity metrics. Experimental results demonstrate that deep learning-based pipelines, particularly those integrating LightGlue, consistently outperform classical alternatives in both visual quality and robustness, while maintaining practical execution times.
Gabriel-Sebastian Buta, Ciprian-Mihai Ceausescu, Alexe Dumitru-Bogdan
ICTAI3
2025 Multi-Dataset Cross-Domain Knowledge Distillation for Medical Image Segmentation
abstract
We propose a novel cross-domain transfer learning framework that leverages knowledge from multiple datasets to improve medical image segmentation on a target task. Our method employs a teacher-student learning paradigm, where a joint teacher model aggregates domain-invariant features from diverse datasets, and a dataset-specific student model is trained via knowledge distillation. We validate our approach on six medical imaging datasets —BrainMetShare, ISLES, BraTS (MRI-based) and Lung MSD, LiTS, KiTS (CT-based)— demonstrating its effectiveness in addressing distributional shifts and enhancing segmentation accuracy across heterogeneous tasks. The results show consistent improvements over baseline models. These findings underscore the potential of multi-dataset knowledge distillation for robust and generalizable segmentation in medical imaging applications.
Ciprian-Mihai Ceausescu, Alexe Dumitru-Bogdan
KES2
2025 Multi-Level Feature Distillation of Joint Teachers Trained on Distinct Image Datasets
abstract
We propose a novel teacher-student framework to distill knowledge from multiple teachers trained on distinct datasets. Each teacher is first trained from scratch on its own dataset. Then, the teachers are combined into a joint architecture, which fuses the features of all teachers at multiple representation levels. The joint teacher architecture is fine-tuned on samples from all datasets, thus gathering useful generic information from all data samples. Finally, we employ a multi-level feature distillation procedure to transfer the knowledge to a student model for each of the considered datasets. We conduct image classification experiments on seven benchmarks, and action recognition experiments on three benchmarks. To illustrate the power of our feature distillation procedure, the student architectures are chosen to be identical to those of the individual teachers. To demonstrate the flexibility of our approach, we combine teachers with distinct architectures. We show that our novel Multi-Level Feature Distillation (MLFD) can significantly surpass equivalent architectures that are either trained on individual datasets, or jointly trained on all datasets at once. Furthermore, we confirm that each step of the proposed training procedure is well motivated by a comprehensive ablation study. We publicly release our code at https://github.com/AdrianIordache/MLFD.
Adrian Iordache, Alexe Dumitru-Bogdan, Radu Tudor Ionescu
WACV2
2019 Detecting Abnormal Events in Video Using Narrowed Normality Clusters
abstract
We formulate the abnormal event detection problem as an outlier detection task and we propose a two-stage algorithm based on k-means clustering and one-class Support Vector Machines (SVM) to eliminate outliers. In the feature extraction stage, we propose to augment spatio-temporal cubes with deep appearance features extracted from the last convolutional layer of a pre-trained neural network. After extracting motion and appearance features from the training video containing only normal events, we apply k-means clustering to find clusters representing different types of normal motion and appearance features. In the first stage, we consider that clusters with fewer samples (with respect to a given threshold) contain mostly outliers, and we eliminate these clusters altogether. In the second stage, we shrink the borders of the remaining clusters by training a one-class SVM model on each cluster. To detected abnormal events in the test video, we analyze each test sample and consider its maximum normality score provided by the trained one-class SVM models, based on the intuition that a test sample can belong to only one cluster of normality. If the test sample does not fit well in any narrowed normality cluster, then it is labeled as abnormal. We compare our method with several state-of-the-art methods on three benchmark data sets. The empirical results indicate that our abnormal event detection framework can achieve better results in most cases, while processing the test video in real-time at 24 frames per second on a single CPU.
Radu Tudor Ionescu, Sorina Smeureanu, Marius Popescu, Alexe Dumitru-Bogdan
WACV4
2017 Unmasking the Abnormal Events in Video
abstract
We propose a novel framework for abnormal event detection in video that requires no training sequences. Our framework is based on unmasking, a technique previously used for authorship verification in text documents, which we adapt to our task. We iteratively train a binary classifier to distinguish between two consecutive video sequences while removing at each step the most discriminant features. Higher training accuracy rates of the intermediately obtained classifiers represent abnormal events. To the best of our knowledge, this is the first work to apply unmasking for a computer vision task. We compare our method with several state-of-the-art supervised and unsupervised methods on four benchmark data sets. The empirical results indicate that our abnormal event detection framework can achieve state-of-the-art results, while running in real-time at 20 frames per second.
Radu Tudor Ionescu, Sorina Smeureanu, Alexe Dumitru-Bogdan, Marius Popescu
ICCV3
2016 How Hard Can It Be? Estimating the Difficulty of Visual Search in an Image
abstract
We address the problem of estimating image difficulty defined as the human response time for solving a visual search task. We collect human annotations of image difficulty for the PASCAL VOC 2012 data set through a crowd-sourcing platform. We then analyze what human interpretable image properties can have an impact on visual search difficulty, and how accurate are those properties for predicting difficulty. Next, we build a regression model based on deep features learned with state of the art convolutional neural networks and show better results for predicting the ground-truth visual search difficulty scores produced by human annotators. Our model is able to correctly rank about 75% image pairs according to their difficulty score. We also show that our difficulty predictor generalizes well to new classes not seen during training. Finally, we demonstrate that our predicted difficulty scores are useful for weakly supervised object localization (8% improvement) and semi-supervised object classification (1% improvement).
Radu Tudor Ionescu, Alexe Dumitru-Bogdan, Marius Leordeanu, Marius Popescu, Dim P. Papadopoulos, Vittorio Ferrari
CVPR2
2014 Preference-aware Integration of Temporal Data
abstract
A complete description of an entity is rarely contained in a single data source, but rather, it is often distributed across different data sources. Applications based on personal electronic health records, sentiment analysis, and financial records all illustrate that significant value can be derived from integrated, consistent, and queryable profiles of entities from different sources. Even more so, such integrated profiles are considerably enhanced if temporal information from different sources is carefully accounted for. We develop a simple and yet versatile operator, called prawn, that is typically called as a final step of an entity integration workflow. Prawn is capable of consistently integrating and resolving temporal conflicts in data that may contain multiple dimensions of time based on a set of preference rules specified by a user (hence the name prawn for preference-aware union ). In the event that not all conflicts can be resolved through preferences, one can enumerate each possible consistent interpretation of the result returned by prawn at a given time point through a polynomial-delay algorithm. In addition to providing algorithms for implementing prawn, we study and establish several desirable properties of prawn. First, prawn produces the same temporally integrated outcome, modulo representation of time, regardless of the order in which data sources are integrated. Second, prawn can be customized to integrate temporal data for different applications by specifying application-specific preference rules. Third, we show experimentally that our implementation of prawn is feasible on both "small" and "big" data platforms in that it is efficient in both storage and execution time. Finally, we demonstrate a fundamental advantage of prawn: we illustrate that standard query languages can be immediately used to pose useful temporal queries over the integrated and resolved entity repository.
Alexe Dumitru-Bogdan, Mary Roth, Wang Chiew Tan
Proc. VLDB Endow.1
2013 Constructing consumer profiles from social media data
abstract
Social media is playing a growing role in providing consumer feedback to companies about their products and services. To maximize the benefit of this feedback, companies want to know how different consumer-segments they are interested in, such as parents, frequent travelers, and comic book fans react to their products and campaigns. In this paper, we describe how constructing consumer profiles is valuable to obtain such insights. We present the challenges in analyzing noisy social media data and the techniques we employ for building the profiles. We also present detailed experimental results from the analysis of over seven billion messages to construct profiles of over 100 million consumers. We demonstrate how consumer profiles can help in understanding consumer feedback by different key segments using a TV show analysis scenario.
Mauricio A. Hernández, Kirsten Hildrum, Prateek Jain 0001, Rohit Wagle, Alexe Dumitru-Bogdan, Rajasekar Krishnamurthy, Ioana Stanoi, Chitra Venkatramani
IEEE BigData5
2012 Searching for objects driven by context
abstract
The dominant visual search paradigm for object class detection is sliding windows. Although simple and effective, it is also wasteful, unnatural and rigidly hardwired. We propose strategies to search for objects which intelligently explore the space of windows by making sequential observations at locations decided based on previous observations. Our strategies adapt to the class being searched and to the content of a particular test image. Their driving force is exploiting context as the statistical relation between the appearance of a window and its location relative to the object, as observed in the training set. In addition to being more elegant than sliding windows, we demonstrate experimentally on the PASCAL VOC 2010 dataset that our strategies evaluate two orders of magnitude fewer windows while at the same time achieving higher detection accuracy.
Alexe Dumitru-Bogdan, Nicolas Heess, Yee Whye Teh, Vittorio Ferrari
NIPS1
2012 Surfacing time-critical insights from social media
abstract
We propose to demonstrate an end-to-end framework for leveraging time-sensitive and critical social media information for businesses. More specifically, we focus on identifying, structuring, integrating, and exposing timely insights that are essential to marketing services and monitoring reputation over social media. Our system includes components for information extraction from text, entity resolution and integration, analytics, and a user interface.
Alexe Dumitru-Bogdan, Mauricio A. Hernández, Kirsten Hildrum, Rajasekar Krishnamurthy, Georgia Koutrika, Meena Nagarajan, Haggai Roitman, Michal Shmueli-Scheuer, Ioana Stanoi, Chitra Venkatramani, Rohit Wagle
SIGMOD Conference1
2012 Weakly Supervised Localization and Learning with Generic Knowledge
Thomas Deselaers, Alexe Dumitru-Bogdan, Vittorio Ferrari
Int. J. Comput. Vis.2
2012 Measuring the Objectness of Image Windows
abstract
We present a generic objectness measure, quantifying how likely it is for an image window to contain an object of any class. We explicitly train it to distinguish objects with a well-defined boundary in space, such as cows and telephones, from amorphous background elements, such as grass and road. The measure combines in a Bayesian framework several image cues measuring characteristics of objects, such as appearing different from their surroundings and having a closed boundary. These include an innovative cue to measure the closed boundary characteristic. In experiments on the challenging PASCAL VOC 07 dataset, we show this new cue to outperform a state-of-the-art saliency measure, and the combined objectness measure to perform better than any cue alone. We also compare to interest point operators, a HOG detector, and three recent works aiming at automatic object segmentation. Finally, we present two applications of objectness. In the first, we sample a small numberof windows according to their objectness probability and give an algorithm to employ them as location priors for modern class-specific object detectors. As we show experimentally, this greatly reduces the number of windows evaluated by the expensive class-specific model. In the second application, we use objectness as a complementary score in addition to the class-specific model, which leads to fewer false positives. As shown in several recent papers, objectness can act as a valuable focus of attention mechanism in many other applications operating on image windows, including weakly supervised learning of object categories, unsupervised pixelwise segmentation, and object tracking in video. Computing objectness is very efficient and takes only about 4 sec. per image.
Alexe Dumitru-Bogdan, Thomas Deselaers, Vittorio Ferrari
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 MapMerge: correlating independent schema mappings
Alexe Dumitru-Bogdan, Mauricio A. Hernández, Lucian Popa 0001, Wang Chiew Tan
VLDB J.1
2011 Exploiting spatial overlap to efficiently compute appearance distances between image windows
abstract
We present a computationally efficient technique to compute the distance of high-dimensional appearance descriptor vectors between image windows. The method exploits the relation between appearance distance and spatial overlap. We derive an upper bound on appearance distance given the spatial overlap of two windows in an image, and use it to bound the distances of many pairs between two images. We propose algorithms that build on these basic operations to efficiently solve tasks relevant to many computer vision applications, such as finding all pairs of windows between two images with distance smaller than a threshold, or finding the single pair with the smallest distance. In experiments on the PASCAL VOC 07 dataset, our algorithms accurately solve these problems while greatly reducing the number of appearance distances computed, and achieve larger speedups than approximate nearest neighbour algorithms based on trees [18]and on hashing [21]. For example, our algorithm finds the most similar pair of windows between two images while computing only 1% of all distances on average.
Alexe Dumitru-Bogdan, Viviana Petrescu, Vittorio Ferrari
NIPS1
2011 Designing and refining schema mappings via data examples
abstract
A schema mapping is a specification of the relationship between a source schema and a target schema. Schema mappings are fundamental building blocks in data integration and data exchange and, as such, obtaining the right schema mapping constitutes a major step towards the integration or exchange of data. Up to now, schema mappings have typically been specified manually or have been derived using mapping-design systems that automatically generate a schema mapping from a visual specification of the relationship between two schemas. We present a novel paradigm and develop a system for the interactive design of schema mappings via data examples. Each data example represents a partial specification of the semantics of the desired schema mapping. At the core of our system lies a sound and complete algorithm that, given a finite set of data examples, decides whether or not there exists a GLAV schema mapping (i.e., a schema mapping specified by Global-and-Local-As-View constraints) that "fits" these data examples. If such a fitting GLAV schema mapping exists, then our system constructs the "most general" one. We give a rigorous computational complexity analysis of the underlying decision problem concerning the existence of a fitting GLAV schema mapping, given a set of data examples. Specifically, we prove that this problem is complete for the second level of the polynomial hierarchy, hence, in a precise sense, harder than NP-complete. This worst-case complexity analysis notwithstanding, we conduct an experimental evaluation of our prototype implementation that demonstrates the feasibility of interactively designing schema mappings using data examples. In particular, our experiments show that our system achieves very good performance in real-life scenarios.
Alexe Dumitru-Bogdan, Balder ten Cate, Phokion G. Kolaitis, Wang Chiew Tan
SIGMOD Conference1
2011 EIRENE: Interactive Design and Refinement of Schema Mappings via Data Examples
Alexe Dumitru-Bogdan, Balder ten Cate, Phokion G. Kolaitis, Wang Chiew Tan
Proc. VLDB Endow.1
2011 Characterizing schema mappings via data examples
abstract
Schema mappings are high-level specifications that describe the relationship between two database schemas; they are considered to be the essential building blocks in data exchange and data integration, and have been the object of extensive research investigations. Since in real-life applications schema mappings can be quite complex, it is important to develop methods and tools for understanding, explaining, and refining schema mappings. A promising approach to this effect is to use “good” data examples that illustrate the schema mapping at hand. We develop a foundation for the systematic investigation of data examples and obtain a number of results on both the capabilities and the limitations of data examples in explaining and understanding schema mappings. We focus on schema mappings specified by source-to-target tuple generating dependencies (s-t tgds) and investigate the following problem: which classes of s-t tgds can be “uniquely characterized” by a finite set of data examples? Our investigation begins by considering finite sets of positive and negative examples, which are arguably the most natural choice of data examples. However, we show that they are not powerful enough to yield interesting unique characterizations. We then consider finite sets of universal examples, where a universal example is a pair consisting of a source instance and a universal solution for that source instance. We first show that unique characterizations via universal examples is, in a precise sense, equivalent to the existence of Armstrong bases (a relaxation of the classical notion of Armstrong databases). After this, we show that every schema mapping specified by LAV s-t tgds is uniquely characterized by a finite set of universal examples with respect to the class of LAV s-t tgds. Moreover, this positive result extends to the much broader classes of n -modular schema mappings, n a positive integer. Finally, we study the unique characterizability of GAV schema mappings. It turns out that some GAV schema mappings are uniquely characterizable by a finite set of universal examples with respect to the class of GAV s-t tgds, while others are not. By unveiling a tight connection with homomorphism dualities, we establish an effective, sound, and complete criterion for determining whether or not a GAV schema mapping is uniquely characterizable by a finite set of universal examples with respect to the class of GAV s-t tgds.
Alexe Dumitru-Bogdan, Balder ten Cate, Phokion G. Kolaitis, Wang Chiew Tan
ACM Trans. Database Syst.1
2010 What is an object?
abstract
We present a generic objectness measure, quantifying how likely it is for an image window to contain an object of any class. We explicitly train it to distinguish objects with a well-defined boundary in space, such as cows and telephones, from amorphous background elements, such as grass and road. The measure combines in a Bayesian framework several image cues measuring characteristics of objects, such as appearing different from their surroundings and having a closed boundary. This includes an innovative cue measuring the closed boundary characteristic. In experiments on the challenging PASCAL VOC 07 dataset, we show this new cue to outperform a state-of-the-art saliency measure, and the combined measure to perform better than any cue alone. Finally, we show how to sample windows from an image according to their objectness distribution and give an algorithm to employ them as location priors for modern class-specific object detectors. In experiments on PASCAL VOC 07 we show this greatly reduces the number of windows evaluated by class-specific object detectors.
Alexe Dumitru-Bogdan, Thomas Deselaers, Vittorio Ferrari
CVPR1
2010 ClassCut for Unsupervised Class Segmentation
Alexe Dumitru-Bogdan, Thomas Deselaers, Vittorio Ferrari
ECCV (5)1
2010 Localizing Objects While Learning Their Appearance
Thomas Deselaers, Alexe Dumitru-Bogdan, Vittorio Ferrari
ECCV (4)2
2010 Characterizing schema mappings via data examples
abstract
Schema mappings are high-level specifications that describe the relationship between two database schemas; they are considered to be the essential building blocks in data exchange and data integration, and have been the object of extensive research investigations. Since in real-life applications schema mappings can be quite complex, it is important to develop methods and tools for understanding, explaining, and refining schema mappings. A promising approach to this effect is to use "good" data examples that illustrate the schema mapping at hand.
Alexe Dumitru-Bogdan, Phokion G. Kolaitis, Wang Chiew Tan
PODS1
2010 MapMerge: Correlating Independent Schema Mappings
abstract
One of the main steps towards integration or exchange of data is to design the mappings that describe the (often complex) relationships between the source schemas or formats and the desired target schema. In this paper, we introduce a new operator, called MapMerge, that can be used to correlate multiple, independently designed schema mappings of smaller scope into larger schema mappings. This allows a more modular construction of complex mappings from various types of smaller mappings such as schema correspondences produced by a schema matcher or pre-existing mappings that were designed by either a human user or via mapping tools. In particular, the new operator also enables a new "divide-and-merge" paradigm for mapping creation, where the design is divided (on purpose) into smaller components that are easier to create and understand, and where MapMerge is used to automatically generate a meaningful overall mapping. We describe our MapMerge algorithm and demonstrate the feasibility of our implementation on several real and synthetic mapping scenarios. In our experiments, we make use of a novel similarity measure between two database instances with different schemas that quantifies the preservation of data associations. We show experimentally that MapMerge improves the quality of the schema mappings, by significantly increasing the similarity between the input source instance and the generated target instance.
Alexe Dumitru-Bogdan, Mauricio A. Hernández, Lucian Popa 0001, Wang Chiew Tan
Proc. VLDB Endow.1
2008 Muse: Mapping Understanding and deSign by Example
abstract
A fundamental problem in information integration is that of designing the relationships, called schema mappings, between two schemas. The specification of a semantically correct schema mapping is typically a complex task. Automated tools can suggest potential mappings, but few tools are available for helping a designer understand mappings and design alternative mappings. We describe Muse, a mapping design wizard that uses data examples to assist designers in understanding and refining a schema mapping towards the desired specification. We present novel algorithms behind Muse and show how Muse systematically guides the designer on two important components of a mapping design: the specification of the desired grouping semantics for sets of data and the choice among alternative interpretations for semantically ambiguous mappings. In every component, Muse infers the desired semantics based on the designer's actions on a short sequence of small examples. Whenever possible, Muse draws examples from a familiar database, thus facilitating the design process even further. We report our experience with Muse on some publicly available schemas.
Alexe Dumitru-Bogdan, Laura Chiticariu, Renée J. Miller, Wang Chiew Tan
ICDE1
2008 Muse: a system for understanding and designing mappings
abstract
Schema mappings are logical assertions that specify the relationships between a source and a target schema in a declarative way. The specification of such mappings is a fundamental problem in information integration. Mappings can be generated by existing mapping systems (semi-)automatically from a visual specification between two schemas. In general, the well-known 80-20 rule applies for mapping generation tools. They can automate 80% of the work, covering common cases and creating a mapping that is close to correct. However, ensuring complete correctness can still require intricate manual work to perfect portions of the mapping.
Alexe Dumitru-Bogdan, Laura Chiticariu, Renée J. Miller, Daniel Pepper, Wang Chiew Tan
SIGMOD Conference1
2008 STBenchmark: towards a benchmark for mapping systems
abstract
A fundamental problem in information integration is to precisely specify the relationships, called mappings, between schemas. Designing mappings is a time-consuming process. To alleviate this problem, many mapping systems have been developed to assist the design of mappings. However, a benchmark for comparing and evaluating these systems has not yet been developed. We present STBenchmark, a solution towards a much needed benchmark for mapping systems. We first describe the challenges that are unique to the development of benchmarks for mapping systems. After this, we describe the three components of STBenchmark: (1) a basic suite of mapping scenarios that we believe represents a minimum set of transformations that should be readily supported by any mapping system, (2) a mapping scenario generator as well as an instance generator that can produce complex mapping scenarios and, respectively, instances of varying sizes of a given schema, (3) a simple usability model that can be used as a first-cut measure on the case of use of a mapping system. We use STBenchmark to evaluate four mapping systems and report our results, as well as describe some interesting observations.
Alexe Dumitru-Bogdan, Wang Chiew Tan, Yannis Velegrakis
Proc. VLDB Endow.1
2008 Comparing and evaluating mapping systems with STBenchmark
abstract
Schema mappings are fundamental building blocks in many information integration applications. Designing mappings is a time-consuming process and for that reason many mapping systems have been developed to assist in the task of designing mappings. However, to the best of our knowledge, a benchmark for comparing and evaluating these systems has not yet been developed. We demonstrate STBenchmark, a benchmark that we have developed for evaluating mapping systems. Our demonstration will showcase the different aspects of mapping systems that STBenchmark evaluates, highlight the results of our comparison and evaluation of four mapping systems, as well as make a case for the need for a standard specification input mechanism to mapping systems in order to make progress towards the development of a uniform testbed or repository for schema mappings and data exchange tasks.
Alexe Dumitru-Bogdan, Wang Chiew Tan, Yannis Velegrakis
Proc. VLDB Endow.1
2006 SPIDER: a Schema mapPIng DEbuggeR
Alexe Dumitru-Bogdan, Laura Chiticariu, Wang Chiew Tan
VLDB1
2005 User Profile Management in Converged Networks (Episode II): "Share your Data, Keep your Secrets"
Arnaud Sahuguet, Alexe Dumitru-Bogdan, Irini Fundulaki, Pierre-Yves Lalilgand, Abdullatif Shikfa, Antoine Arnail
CIDR2
2005 Sub-document queries over XML with XSQirrel
abstract
This paper describes XSQirrel, a new XML query language that transforms a document into a sub-document, i.e. a tree where the root-to-leaf paths are a subset of the root-to-leaf paths from the original document.We show that this type of queries is extremely useful for various applications (e.g. web services) and that the currently existing query languages are poorly equipped to express, reason and evaluate such queries. In particular, we emphasize the need to be able to compose such queries. We present the XSQirrel language with its syntax, semantics and two language specific operators, union and composition.For the evaluation of the language, we leverage well established query technologies by translating XSQirrel expressions into XPath programs, XQuery queries or XSLT stylesheets.We provide some experimental results that compare our various evaluation strategies. We also show the runtime benefits of query composition over sequential evaluation.
Arnaud Sahuguet, Alexe Dumitru-Bogdan
WWW2
2004 An Electronic Patient Record "on Steroids": Distributed, Peer-to-Peer, Secure and Privacy-conscious
Serge Abiteboul, Alexe Dumitru-Bogdan, Omar Benjelloun, Bogdan Cautis, Irini Fundulaki, Tova Milo, Arnaud Sahuguet
VLDB2