VLDB 2026 Research / reviewers in the wild / expert
Mihaela A. Bornea
dblp:45/6755
· DBLP profile ↗
14ranked-venue papers
8as first author
4since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 7 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Question answering and dialogue systems · 37% Information extraction and text analysis · 28% Deep learning architectures and training · 14% | |
| Databases, data mining, and information retrieval
5 papers |
Query processing and optimization · 29% Transaction processing and concurrency control · 19% Graph data management · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
extractive question answering |
0.7 | 1 | 2023 | GAAMA 2.0: An Integrated System That Answers Boolean and Extractive Questions · AAAI 2023 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.7 | 1 | 2023 | GAAMA 2.0: An Integrated System That Answers Boolean and Extractive Questions · AAAI 2023 |
Natural language and speech › Information extraction and text analysis › relation extraction
biomedical relation extraction |
0.5 | 1 | 2021 | Exploring the Efficacy of Generic Drugs in Treating Cancer · AAAI 2021 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.5 | 1 | 2021 | Multilingual Transfer Learning for QA using Translation as Data Augmentation · AAAI 2021 |
Machine learning › Deep learning architectures and training
data augmentation |
0.5 | 1 | 2021 | Multilingual Transfer Learning for QA using Translation as Data Augmentation · AAAI 2021 |
Natural language and speech › Question answering and dialogue systems
multilingual question answering |
0.5 | 1 | 2021 | Multilingual Transfer Learning for QA using Translation as Data Augmentation · AAAI 2021 |
Bioinformatics and computational biology › drug discovery
drug repurposing |
0.5 | 1 | 2021 | Exploring the Efficacy of Generic Drugs in Treating Cancer · AAAI 2021 |
Natural language and speech › Information extraction and text analysis › text classification › low-resource text classification
few-shot text classification |
0.4 | 1 | 2019 | Combining Unsupervised Pre-training and Annotator Rationales to Improve Low-shot Text Classification · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Information extraction and text analysis
text classification |
0.4 | 1 | 2019 | Combining Unsupervised Pre-training and Annotator Rationales to Improve Low-shot Text Classification · EMNLP/IJCNLP (1) 2019 |
Machine learning › Representation and self-supervised learning › pre-training
unsupervised pre-training |
0.4 | 1 | 2019 | Combining Unsupervised Pre-training and Annotator Rationales to Improve Low-shot Text Classification · EMNLP/IJCNLP (1) 2019 |
Query processing and optimization › join processing
join algorithms |
0.2 | 2 | 2011 | Semi-Streamed Index Join for near-real time execution of ETL transformations · ICDE 2011 Double Index NEsted-Loop Reactive Join for Result Rate Optimization · ICDE 2009 |
Query processing and optimization › join processing › join algorithms
adaptive join |
0.2 | 2 | 2010 | Adaptive Join Operators for Result Rate Optimization on Streaming Inputs · IEEE Trans. Knowl. Data Eng. 2010 Double Index NEsted-Loop Reactive Join for Result Rate Optimization · ICDE 2009 |
Machine learning › Deep learning architectures and training
transformer |
0.2 | 1 | 2023 | GAAMA 2.0: An Integrated System That Answers Boolean and Extractive Questions · AAAI 2023 |
Information retrieval › cross-language information retrieval
query translation |
0.2 | 1 | 2013 | Building an efficient RDF store over a relational database · SIGMOD Conference 2013 |
Graph data management › RDF data management
RDF query processing |
0.2 | 1 | 2013 | Building an efficient RDF store over a relational database · SIGMOD Conference 2013 |
Graph data management › RDF data management
relational storage of RDF |
0.2 | 1 | 2013 | Building an efficient RDF store over a relational database · SIGMOD Conference 2013 |
Distributed and cloud data management
data replication |
0.1 | 1 | 2011 | One-copy serializability with snapshot isolation under the hood · ICDE 2011 |
Transaction processing and concurrency control › serializability
one-copy serializability |
0.1 | 1 | 2011 | One-copy serializability with snapshot isolation under the hood · ICDE 2011 |
Transaction processing and concurrency control › isolation levels › snapshot isolation
serializable snapshot isolation |
0.1 | 1 | 2011 | One-copy serializability with snapshot isolation under the hood · ICDE 2011 |
Transaction processing and concurrency control › isolation levels
snapshot isolation |
0.1 | 1 | 2011 | One-copy serializability with snapshot isolation under the hood · ICDE 2011 |
Natural language and speech › Information extraction and text analysis › data annotation
annotator rationales |
0.1 | 1 | 2019 | Combining Unsupervised Pre-training and Annotator Rationales to Improve Low-shot Text Classification · EMNLP/IJCNLP (1) 2019 |
Machine learning › Trustworthy machine learning
rationale-based training |
0.1 | 1 | 2019 | Combining Unsupervised Pre-training and Annotator Rationales to Improve Low-shot Text Classification · EMNLP/IJCNLP (1) 2019 |
Query processing and optimization
join processing |
0.1 | 1 | 2010 | Adaptive Join Operators for Result Rate Optimization on Streaming Inputs · IEEE Trans. Knowl. Data Eng. 2010 |
Data stream processing
stream join |
0.1 | 1 | 2010 | Adaptive Join Operators for Result Rate Optimization on Streaming Inputs · IEEE Trans. Knowl. Data Eng. 2010 |
Distributed systems › replication
replication and fault tolerance |
0.0 | 1 | 2011 | One-copy serializability with snapshot isolation under the hood · ICDE 2011 |
Query processing and optimization › join processing
multi-way join |
0.0 | 1 | 2010 | Adaptive Join Operators for Result Rate Optimization on Streaming Inputs · IEEE Trans. Knowl. Data Eng. 2010 |
Data integration and cleaning
heterogeneous data source integration |
0.0 | 1 | 2009 | Double Index NEsted-Loop Reactive Join for Result Rate Optimization · ICDE 2009 |
Methods — techniques the papers use, named apart from their topics
natural language processing · 1.0machine learning · 1.0continuous integration · 1.0transformer models · 0.7adapter · 0.7machine translation · 0.5language model · 0.5adversarial training · 0.5unsupervised pre-training · 0.4annotator rationales · 0.4readset extraction · 0.2multi-version concurrency control · 0.2certification · 0.2shredding of RDF into relational · 0.2query translation · 0.2semi-streaming index join · 0.1cache replacement policy · 0.1reentrant join technique · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | GAAMA 2.0: An Integrated System That Answers Boolean and Extractive QuestionsabstractRecent machine reading comprehension datasets include extractive and boolean questions but current approaches do not offer integrated support for answering both question types. We present a front-end demo to a multilingual machine reading comprehension system that handles boolean and extractive questions. It provides a yes/no answer and highlights the supporting evidence for boolean questions. It provides an answer for extractive questions and highlights the answer in the passage. Our system, GAAMA 2.0, achieved first place on the TyDi QA leaderboard at the time of submission. We contrast two different implementations of our approach: including multiple transformer models for easy deployment, and a shared transformer model utilizing adapters to reduce GPU memory footprint for a resource-constrained environment. J. Scott McCarley, Mihaela A. Bornea, Sara Rosenthal, Anthony Ferritto, Md. Arafat Sultan, Avirup Sil, Radu Florian |
AAAI | 2 |
| 2021 | Exploring the Efficacy of Generic Drugs in Treating CancerabstractThousands of scientific publications discuss evidence on the efficacy of non-cancer generic drugs being tested for cancer. However, trying to manually identify and extract such evidence is intractable at scale. We introduce a natural language processing pipeline to automate the identification of relevant studies and facilitate the extraction of therapeutic associations between generic drugs and cancers from PubMed abstracts. We annotate datasets of drug-cancer evidence and use them to train models to identify and characterize such evidence at scale. To make this evidence readily consumable, we incorporate the results of the models in a web application that allows users to browse documents and their extracted evidence. Users can provide feedback on the quality of the evidence extracted by our models. This feedback is used to improve our datasets and the corresponding models in a continuous integration system. We describe the natural language processing pipeline in our application and the steps required to deploy services based on the machine learning models. Ioana Baldini, Mariana Bernagozzi, Sulbha Aggarwal, Mihaela A. Bornea, Saksham Chawla, Joppe Geluykens, Dmitriy Katz, Pratik Mukherjee, Smruthi Ramesh, Sara Rosenthal, Jagrati Sharma, Kush R. Varshney, Laura B. Kleiman, Pradeep Mangalath, Catherine Del Vecchio Fitz |
AAAI | 4 |
| 2021 | Multilingual Transfer Learning for QA using Translation as Data AugmentationabstractPrior work on multilingual question answering has mostly focused on using large multilingual pre-trained language models (LM) to perform zero-shot language-wise learning: train a QA model on English and test on other languages. In this work, we explore strategies that improve cross-lingual transfer by bringing the multilingual embeddings closer in the semantic space. Our first strategy augments the original English training data with machine translation-generated data. This results in a corpus of multilingual silver-labeled QA pairs that is 14 times larger than the original training set. In addition, we propose two novel strategies, language adversarial training and language arbitration framework, which significantly improve the (zero-resource) cross-lingual transfer performance and result in LM embeddings that are less language-variant. Empirically, we show that the proposed models outperform the previous zero-shot baseline on the recently introduced multilingual MLQA and TyDiQA datasets. Mihaela A. Bornea, Lin Pan 0003, Sara Rosenthal, Radu Florian, Avirup Sil |
AAAI | 1 |
| 2021 | Generative Relation Linking for Question Answering over Knowledge Bases
Gaetano Rossiello, Nandana Mihindukulasooriya, Ibrahim Abdelaziz, Mihaela A. Bornea, Alfio Massimiliano Gliozzo, Tahira Naseem, Pavan Kapanipathi |
ISWC | 4 |
| 2019 | Combining Unsupervised Pre-training and Annotator Rationales to Improve Low-shot Text ClassificationabstractOren Melamud, Mihaela Bornea, Ken Barker. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Oren Melamud, Mihaela A. Bornea, Ken Barker 0002 |
EMNLP/IJCNLP (1) | 2 |
| 2016 | An Executable Specification for SPARQL
Mihaela A. Bornea, Julian Dolby, Achille Fokoue, Anastasios Kementsietsidis, Kavitha Srinivas, Mandana Vaziri |
WISE (2) | 1 |
| 2015 | Relational Path Mining in Structured KnowledgeabstractLarge sources of structured knowledge are available in many domains, enabling the construction of applications requiring relational knowledge. But in spite of the apparent availability of relational content, the semantics and granularity of these sources don't always match the requirements of specific tasks. Yet even when the coverage of explicit relational knowledge in a source seems inadequate, there may be implicit knowledge in the complete space of relation instances. Mihaela A. Bornea, Ken Barker 0002 |
K-CAP | 1 |
| 2014 | Problem-oriented patient record summary: An early report on a Watson applicationabstractAs the use of Electronic Medical Records (EMRs) becomes widespread, the amount of data in an EMR becomes a challenge for its comprehension. We developed problem-oriented EMR summarization to address this issue, as a part of a larger effort of adapting IBM Watson to the medical domain. The problem-orientation refers to the central role of a patient's medical problems in the summary. The summarization uses a generated problem list, relates these generated medical problems to relevant clinical data, and organizes the clinical data in a medically meaningful manner. Watson analytics are used for creating the summarization. This is a step in building the next generation EMR, one that is based not on just keeping record but instead on a conceptual understanding of medicine, thereby crossing the threshold from record storage to an intelligent entity for clinical decision making. Murthy V. Devarakonda, Ching-Huei Tsou, Mihaela A. Bornea |
Healthcom | 4 |
| 2014 | An Offline Optimal SPARQL Query Planning Approach to Evaluate Online Heuristic Planners
Achille Fokoue, Mihaela A. Bornea, Julian Dolby, Anastasios Kementsietsidis, Kavitha Srinivas |
WISE (1) | 2 |
| 2013 | Building an efficient RDF store over a relational databaseabstractEfficient storage and querying of RDF data is of increasing importance, due to the increased popularity and widespread acceptance of RDF on the web and in the enterprise. In this paper, we describe a novel storage and query mechanism for RDF which works on top of existing relational representations. Reliance on relational representations of RDF means that one can take advantage of 35+ years of research on efficient storage and querying, industrial-strength transaction support, locking, security, etc. However, there are significant challenges in storing RDF in relational, which include data sparsity and schema variability. We describe novel mechanisms to shred RDF into relational, and novel query translation techniques to maximize the advantages of this shredded representation. We show that these mechanisms result in consistently good performance across multiple RDF benchmarks, even when compared with current state-of-the-art stores. This work provides the basis for RDF support in DB2 v.10.1. Mihaela A. Bornea, Julian Dolby, Anastasios Kementsietsidis, Kavitha Srinivas, Patrick Dantressangle, Octavian Udrea, Bishwaranjan Bhattacharjee |
SIGMOD Conference | 1 |
| 2011 | Semi-Streamed Index Join for near-real time execution of ETL transformationsabstractActive data warehouses have emerged as a new business intelligence paradigm where data in the integrated repository is refreshed in near real-time. This shift of practices achieves higher consistency between the stored information and the latest updates, which in turn influences crucially the output of decision making processes. In this paper we focus on the changes required in the implementation of Extract Transform Load (ETL) operations which now need to be executed in an online fashion. In particular, the ETL transformations frequently include the join between an incoming stream of updates and a disk-resident table of historical data or metadata. In this context we propose a novel Semi-Streaming Index Join (SSIJ) algorithm that maximizes the throughput of the join by buffering stream tuples and then judiciously selecting how to best amortize expensive disk seeks for blocks of the stored relation among a large number of stream tuples. The relation blocks required for joining with the stream are loaded from disk based on an optimal plan. In order to maximize the utilization of the available memory space for performing the join, our technique incorporates a simple but effective cache replacement policy for managing the retrieved blocks of the relation. Moreover, SSIJ is able to adapt to changing characteristics of the stream (i.e. arrival rate, data distribution) by dynamically adjusting the allocated memory between the cached relation blocks and the stream. Our experiments with a variety of synthetic and real data sets demonstrate that SSIJ consistently outperforms the state-of-the-art algorithm in terms of the maximum sustainable throughput of the join while being also able to accommodate deadlines on stream tuple processing. Mihaela A. Bornea, Antonios Deligiannakis, Yannis Kotidis, Vasilis Vassalos |
ICDE | 1 |
| 2011 | One-copy serializability with snapshot isolation under the hoodabstractThis paper presents a method that allows a replicated database system to provide a global isolation level stronger than the isolation level provided on each individual database replica. We propose a new multi-version concurrency control algorithm called, serializable generalized snapshot isolation (SGSI), that targets middleware replicated database systems. Each replica runs snapshot isolation locally and the replication middleware guarantees global one-copy serializability. We introduce novel techniques to provide a stronger global isolation level, namely readset extraction and enhanced certification that prevents read-write and write-write conflicts in a replicated setting. We prove the correctness of the proposed algorithm, and build a prototype replicated database system to evaluate SGSI performance experimentally. Extensive experiments with an 8 replica database system under the TPC-W workload mixes demonstrate the practicality and low overhead of the algorithm. Mihaela A. Bornea, Orion Hodson, Sameh Elnikety, Alan D. Fekete |
ICDE | 1 |
| 2010 | Adaptive Join Operators for Result Rate Optimization on Streaming InputsabstractAdaptive join algorithms have recently attracted a lot of attention in emerging applications where data are provided by autonomous data sources through heterogeneous network environments. Their main advantage over traditional join techniques is that they can start producing join results as soon as the first input tuples are available, thus, improving pipelining by smoothing join result production and by masking source or network delays. In this paper, we first propose Double Index NEsted-loops Reactive join (DINER), a new adaptive two-way join algorithm for result rate maximization. DINER combines two key elements: an intuitive flushing policy that aims to increase the productivity of in-memory tuples in producing results during the online phase of the join, and a novel reentrant join technique that allows the algorithm to rapidly switch between processing in-memory and disk-resident tuples, thus, better exploiting temporary delays when new data are not available. We then extend the applicability of the proposed technique for a more challenging setup: handling more than two inputs. Multiple Index NEsted-loop Reactive join (MINER) is a multiway join operator that inherits its principles from DINER. Our experiments using real and synthetic data sets demonstrate that DINER outperforms previous adaptive join algorithms in producing result tuples at a significantly higher rate, while making better use of the available memory. Our experiments also shows that in the presence of multiple inputs, MINER manages to produce a high percentage of early results, outperforming existing techniques for adaptive multiway join. Mihaela A. Bornea, Vasilis Vassalos, Yannis Kotidis, Antonios Deligiannakis |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2009 | Double Index NEsted-Loop Reactive Join for Result Rate OptimizationabstractAdaptive join algorithms have recently attracted a lot of attention in emerging applications where data is provided by autonomous data sources through heterogeneous network environments. Their main advantage over traditional join techniques is that they can start producing join results as soon as the first input tuples are available, thus improving pipelining by smoothing join result production and by masking source or network delays. In this paper we propose double index nested loops reactive join (DINER), a new adaptive join algorithm for result rate maximization. DINER combines two key elements: an intuitive flushing policy that aims to increase the productivity of in-memory tuples in producing results during the online phase of the join, and a novel re-entrant join technique that allows the algorithm to rapidly switch between processing in-memory and disk-resident tuples, thus better exploiting temporary delays when new data is not available. Our experiments using real and synthetic data sets demonstrate that DINER outperforms previous adaptive join algorithms in producing result tuples at a significantly higher rate, while making better use of the available memory. Mihaela A. Bornea, Vasilis Vassalos, Yannis Kotidis, Antonios Deligiannakis |
ICDE | 1 |