VLDB 2026 Research / reviewers in the wild / expert
Rares Vernica
dblp:88/5139
· DBLP profile ↗
16ranked-venue papers
3as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 14 · 3 first-authorArtificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
7 papers |
Query processing and optimization · 22% Database system architecture and tuning · 20% Data models and query languages · 20% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Cloud and datacenter computing · 82% Parallel and multicore computing · 18% | |
| Artificial intelligence
2 papers |
Face, body and person analysis · 64% Information extraction and text analysis · 36% |
Topics — the 20 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Indexing and storage engines
LSM-tree |
0.2 | 1 | 2014 | AsterixDB: A Scalable, Open Source BDMS · Proc. VLDB Endow. 2014 |
Data models and query languages › query language
semistructured query language |
0.2 | 1 | 2014 | AsterixDB: A Scalable, Open Source BDMS · Proc. VLDB Endow. 2014 |
Cloud and datacenter computing › cloud platform
cloud service platform |
0.2 | 1 | 2013 | Cloud based multimedia analytic platform · ACM Multimedia 2013 |
Data models and query languages › query language
declarative query language |
0.1 | 1 | 2012 | ASTERIX: An Open Source System for "Big Data" Management and Analysis · Proc. VLDB Endow. 2012 |
Database system architecture and tuning
parallel database system |
0.1 | 1 | 2012 | ASTERIX: An Open Source System for "Big Data" Management and Analysis · Proc. VLDB Endow. 2012 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.1 | 1 | 2012 | On the optimization of schedules for MapReduce workloads in the presence of shared scans · VLDB J. 2012 |
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
mapreduce scheduling |
0.1 | 1 | 2012 | On the optimization of schedules for MapReduce workloads in the presence of shared scans · VLDB J. 2012 |
Data integration and cleaning
entity resolution |
0.1 | 1 | 2010 | Efficient parallel set-similarity joins using MapReduce · SIGMOD Conference 2010 |
Query processing and optimization › similarity join
set similarity join |
0.1 | 1 | 2010 | Efficient parallel set-similarity joins using MapReduce · SIGMOD Conference 2010 |
Natural language and speech › Information extraction and text analysis
entity typing |
0.1 | 1 | 2008 | Entity categorization over large document collections · KDD 2008 |
Query processing and optimization
selectivity estimation |
0.1 | 1 | 2008 | SEPIA: estimating selectivities of approximate string predicates in large Databases · VLDB J. 2008 |
Query processing and optimization › query rewriting
query relaxation |
0.1 | 1 | 2006 | Relaxing Join and Selection Queries · VLDB 2006 |
Computer vision › Face, body and person analysis › facial attribute analysis
demographic estimation |
0.0 | 1 | 2013 | Cloud based multimedia analytic platform · ACM Multimedia 2013 |
Computer vision › Face, body and person analysis
face recognition |
0.0 | 1 | 2013 | Cloud based multimedia analytic platform · ACM Multimedia 2013 |
Computer vision › Face, body and person analysis › face recognition
face verification |
0.0 | 1 | 2013 | Cloud based multimedia analytic platform · ACM Multimedia 2013 |
Spatial and temporal data management › spatial query processing
geo-spatial query |
0.0 | 1 | 2012 | ASTERIX: An Open Source System for "Big Data" Management and Analysis · Proc. VLDB Endow. 2012 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.0 | 1 | 2011 | Hyracks: A flexible and extensible foundation for data-intensive computing · ICDE 2011 |
Distributed and cloud data management
parallel data processing |
0.0 | 1 | 2010 | Efficient parallel set-similarity joins using MapReduce · SIGMOD Conference 2010 |
Information retrieval › text analysis
document collection analysis |
0.0 | 1 | 2008 | Entity categorization over large document collections · KDD 2008 |
Information retrieval › similarity measure
string similarity |
0.0 | 1 | 2008 | SEPIA: estimating selectivities of approximate string predicates in large Databases · VLDB J. 2008 |
Methods — techniques the papers use, named apart from their topics
distributed computing · 0.3partitioned-parallel execution · 0.2DAG-based dataflow · 0.2partitioning · 0.2mapreduce · 0.2sampling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Similarity query support in big data management systems
Taewoo Kim 0001, Wenhai Li, Alexander Behm, Inci Cetindil, Rares Vernica, Vinayak R. Borkar, Michael J. Carey 0001, Chen Li 0001 |
Inf. Syst. | 5 |
| 2019 | Using text mining for personalization and recommendation for an enriched hybrid learning experienceabstractAbstract Students are nowadays given many options to consume educational content in digital formats as alternatives to printed material. Previous research suggests that while digital content has advantages, printed media still provides other benefits that cannot be matched by digital. Therefore, technology should leverage the benefits of both. In this paper, we present the Meaningful Education and Training Information System, a multifaceted hybrid textbook learning platform. The goal of the system is to provide an easy digital‐to‐print‐to‐digital content creation and reading service. The Meaningful Education and Training Information System incorporates technologies for layout, personalization, cocreation, and assessments. These facilitate common teacher/student tasks and help provide a richer, more effective learning experience. Our system has been demonstrated in multiple international education events, partner engagements, and pilots with local universities and high schools. Rares Vernica, Tamir Hassan, Niranjan Damera-Venkata |
Comput. Intell. | 2 |
| 2018 | Supporting Similarity Queries in Apache AsterixDB
Taewoo Kim 0001, Wenhai Li, Alexander Behm, Inci Cetindil, Rares Vernica, Vinayak R. Borkar, Michael J. Carey 0001, Chen Li 0001 |
EDBT | 5 |
| 2016 | METIS: A Multi-faceted Hybrid Book Learning PlatformabstractToday, students are offered a wide variety of alternatives to printed material for the consumption of educational content. Previous research suggests that, while digital content has its advantages, printed content still offers benefits that cannot be matched by digital media. This paper introduces the Meaningful Education and Training Information System (METIS), a multi-faceted hybrid book learning platform. The goal of the system is to provide an easy digital-to-print-to-digital content creation and reading service. METIS incorporates technology for layout, personalization, co-creation and assessment. These facilitate and, in many cases, significantly simplify common teacher/student tasks. Our system has been demonstrated at several international education events, partner engagements, and pilots with local universities and high schools. We present the system and discuss how it enables hybrid learning. Rares Vernica, Tamir Hassan, Niranjan Damera-Venkata, Jian Fan, Jerry Liu, Steven J. Simske, Shanchan Wu |
DocEng | 2 |
| 2015 | AERO: An Extensible Framework for Adaptive Web Layout SynthesisabstractWe present AERO, an extensible framework for adaptive web layout synthesis. The goal is to provide an underlying software architecture to allow general adaptive layout behaviors. The framework consists of a 1) a suite of templates specified in HTML/CSS, 2) A hierarchical, highly customizable scoring function specification and 3) An evaluation engine that leverages native browser rendering to rapidly render content and apply the scoring functions. Unlike current responsive layout frameworks for web (e.g., Twitter Bootstrap) that have pre-configured grid layouts that adapt in a manually pre-encoded content-independent manner, AERO allows layout to adapt automatically based on multiple content-dependent criteria like aesthetic quality, cropability of individual images, layout A/B testing results, Ad placement etc. Rares Vernica, Niranjan Damera-Venkata |
DocEng | 1 |
| 2014 | AsterixDB: A Scalable, Open Source BDMSabstractAsterixDB is a new, full-function BDMS (Big Data Management System) with a feature set that distinguishes it from other platforms in today's open source Big Data ecosystem. Its features make it well-suited to applications like web data warehousing, social data storage and analysis, and other use cases related to Big Data. AsterixDB has a flexible NoSQL style data model; a query language that supports a wide range of queries; a scalable runtime; partitioned, LSM-based data storage and indexing (including B + -tree, R-tree, and text indexes); support for external as well as natively stored data; a rich set of built-in types; support for fuzzy, spatial, and temporal types and queries; a built-in notion of data feeds for ingestion of data; and transaction support akin to that of a NoSQL store. Development of AsterixDB began in 2009 and led to a mid-2013 initial open source release. This paper is the first complete description of the resulting open source AsterixDB system. Covered herein are the system's data model, its query language, and its software architecture. Also included are a summary of the current status of the project and a first glimpse into how AsterixDB performs when compared to alternative technologies, including a parallel relational DBMS, a popular NoSQL store, and a popular Hadoop-based SQL data analytics platform, for things that both technologies can do. Also included is a brief description of some initial trials that the system has undergone and the lessons learned (and plans laid) based on those early "customer" engagements. Sattam Alsubaiee, Yasser Altowim, Hotham Altwaijry, Alexander Behm, Vinayak R. Borkar, Yingyi Bu, Michael J. Carey 0001, Inci Cetindil, Madhusudan Cheelangi, Khurram Faraaz, Eugenia Gabrielova, Raman Grover, Zachary Heilbron, Young-Seok Kim, Chen Li 0001, Guangqiang Li, Ji Mahn Ok, Nicola Onose, Pouria Pirzadeh, Vassilis J. Tsotras, Rares Vernica, Till Westmann |
Proc. VLDB Endow. | 21 |
| 2013 | Cloud based multimedia analytic platformabstractMultimedia Analytic Platform is a cloud based service to expose state-of-the-art multimedia technologies for mobile and web application development. As a product-quality service platform, it offers comprehensive API documentation, code example, service description and sandbox for trial for each multimedia technology. The utilization of the cloud storage and distributed computing framework allows the service platform to run with robustness and efficiency. The current technologies supported by the platform include face detection, face verification, face demographic estimation, feature extraction, image matching, and image collage. Since its initial public launch in October 2012, it has been adopted by universities and third party companies for course support and application development. Rares Vernica, Qian Lin 0001 |
ACM Multimedia | 2 |
| 2012 | Adaptive MapReduce using situation-aware mappersabstractWe propose new adaptive runtime techniques for MapReduce that improve performance and simplify job tuning. We implement these techniques by breaking a key assumption of MapReduce that mappers run in isolation. Instead, our mappers communicate through a distributed meta-data store and are aware of the global state of the job. However, we still preserve the fault-tolerance, scalability, and programming API of MapReduce. We utilize these "situation-aware mappers" to develop a set of techniques that make MapReduce more dynamic: (a) Adaptive Mappers dynamically take multiple data partitions (splits) to amortize mapper start-up costs; (b) Adaptive Combiners improve local aggregation by maintaining a cache of partial aggregates for the frequent keys; (c) Adaptive Sampling and Partitioning sample the mapper outputs and use the obtained statistics to produce balanced partitions for the reducers. Our experimental evaluation shows that adaptive techniques provide up to 3x performance improvement, in some cases, and dramatically improve performance stability across the board. Rares Vernica, Andrey Balmin, Kevin S. Beyer, Vuk Ercegovac |
EDBT | 1 |
| 2012 | ASTERIX: An Open Source System for "Big Data" Management and AnalysisabstractAt UC Irvine, we are building a next generation parallel database system, called ASTERIX, as our approach to addressing today's "Big Data" management challenges. ASTERIX aims to combine time-tested principles from parallel database systems with those of the Web-scale computing community, such as fault tolerance for long running jobs. In this demo, we present a whirlwind tour of ASTERIX, highlighting a few of its key features. We will demonstrate examples of our data definition language to model semi-structured data, and examples of interesting queries using our declarative query language. In particular, we will show the capabilities of ASTERIX for answering geo-spatial queries and fuzzy queries, as well as ASTERIX' data feed construct for continuously ingesting data. Sattam Alsubaiee, Yasser Altowim, Hotham Altwaijry, Alexander Behm, Vinayak R. Borkar, Yingyi Bu, Michael J. Carey 0001, Raman Grover, Zachary Heilbron, Young-Seok Kim, Chen Li 0001, Nicola Onose, Pouria Pirzadeh, Rares Vernica |
Proc. VLDB Endow. | 14 |
| 2012 | On the optimization of schedules for MapReduce workloads in the presence of shared scans
Joel L. Wolf, Andrey Balmin, Deepak Rajan, Kirsten Hildrum, Rohit Khandekar, Sujay S. Parekh, Kun-Lung Wu, Rares Vernica |
VLDB J. | 8 |
| 2011 | Hyracks: A flexible and extensible foundation for data-intensive computingabstractHyracks is a new partitioned-parallel software platform designed to run data-intensive computations on large shared-nothing clusters of computers. Hyracks allows users to express a computation as a DAG of data operators and connectors. Operators operate on partitions of input data and produce partitions of output data, while connectors repartition operators' outputs to make the newly produced partitions available at the consuming operators. We describe the Hyracks end user model, for authors of dataflow jobs, and the extension model for users who wish to augment Hyracks' built-in library with new operator and/or connector types. We also describe our initial Hyracks implementation. Since Hyracks is in roughly the same space as the open source Hadoop platform, we compare Hyracks with Hadoop experimentally for several different kinds of use cases. The initial results demonstrate that Hyracks has significant promise as a next-generation platform for data-intensive applications. Vinayak R. Borkar, Michael J. Carey 0001, Raman Grover, Nicola Onose, Rares Vernica |
ICDE | 5 |
| 2011 | ASTERIX: towards a scalable, semistructured data platform for evolving-world models
Alexander Behm, Vinayak R. Borkar, Michael J. Carey 0001, Raman Grover, Chen Li 0001, Nicola Onose, Rares Vernica, Alin Deutsch, Yannis Papakonstantinou, Vassilis J. Tsotras |
Distributed Parallel Databases | 7 |
| 2010 | Efficient parallel set-similarity joins using MapReduceabstractIn this paper we study how to efficiently perform set-similarity joins in parallel using the popular MapReduce framework. We propose a 3-stage approach for end-to-end set-similarity joins. We take as input a set of records and output a set of joined records based on a set-similarity condition. We efficiently partition the data across nodes in order to balance the workload and minimize the need for replication. We study both self-join and R-S join cases, and show how to carefully control the amount of data kept in main memory on each node. We also propose solutions for the case where, even if we use the most fine-grained partitioning, the data still does not fit in the main memory of a node. We report results from extensive experiments on real datasets, synthetically increased in size, to evaluate the speedup and scaleup properties of the proposed algorithms using Hadoop. Rares Vernica, Michael J. Carey 0001, Chen Li 0001 |
SIGMOD Conference | 1 |
| 2008 | Entity categorization over large document collectionsabstractExtracting entities (such as people, movies) from documents and identifying the categories (such as painter, writer) they belong to enable structured querying and data analysis over unstructured document collections. In this paper, we focus on the problem of categorizing extracted entities. Most prior approaches developed for this task only analyzed the local document context within which entities occur. In this paper, we significantly improve the accuracy of entity categorization by (i) considering an entity's context across multiple documents containing it, and (ii) exploiting existing large lists of related entities (e.g., lists of actors, directors, books). These approaches introduce computational challenges because (a) the context of entities has to be aggregated across several documents and (b) the lists of related entities may be very large. We develop techniques to address these challenges. We present a thorough experimental study on real data sets that demonstrates the increase in accuracy and the scalability of our approaches. Venkatesh Ganti, Arnd Christian König, Rares Vernica |
KDD | 3 |
| 2008 | SEPIA: estimating selectivities of approximate string predicates in large Databases
Chen Li 0001, Rares Vernica |
VLDB J. | 3 |
| 2006 | Relaxing Join and Selection Queries
Nick Koudas, Chen Li 0001, Anthony K. H. Tung, Rares Vernica |
VLDB | 4 |