Rares Vernica

dblp:88/5139 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 3 first-authorArtificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Query processing and optimization · 22% Database system architecture and tuning · 20% Data models and query languages · 20%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Cloud and datacenter computing · 82% Parallel and multicore computing · 18%
Artificial intelligence
2 papers
Face, body and person analysis · 64% Information extraction and text analysis · 36%

Topics — the 20 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Indexing and storage engines
LSM-tree
0.212014
AsterixDB: A Scalable, Open Source BDMS · Proc. VLDB Endow. 2014
Data models and query languages › query language
semistructured query language
0.212014
AsterixDB: A Scalable, Open Source BDMS · Proc. VLDB Endow. 2014
Cloud and datacenter computing › cloud platform
cloud service platform
0.212013
Cloud based multimedia analytic platform · ACM Multimedia 2013
Data models and query languages › query language
declarative query language
0.112012
ASTERIX: An Open Source System for "Big Data" Management and Analysis · Proc. VLDB Endow. 2012
Database system architecture and tuning
parallel database system
0.112012
ASTERIX: An Open Source System for "Big Data" Management and Analysis · Proc. VLDB Endow. 2012
Cloud and datacenter computing
cluster resource management and scheduling
0.112012
On the optimization of schedules for MapReduce workloads in the presence of shared scans · VLDB J. 2012
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
mapreduce scheduling
0.112012
On the optimization of schedules for MapReduce workloads in the presence of shared scans · VLDB J. 2012
Data integration and cleaning
entity resolution
0.112010
Efficient parallel set-similarity joins using MapReduce · SIGMOD Conference 2010
Query processing and optimization › similarity join
set similarity join
0.112010
Efficient parallel set-similarity joins using MapReduce · SIGMOD Conference 2010
Natural language and speech › Information extraction and text analysis
entity typing
0.112008
Entity categorization over large document collections · KDD 2008
Query processing and optimization
selectivity estimation
0.112008
SEPIA: estimating selectivities of approximate string predicates in large Databases · VLDB J. 2008
Query processing and optimization › query rewriting
query relaxation
0.112006
Relaxing Join and Selection Queries · VLDB 2006
Computer vision › Face, body and person analysis › facial attribute analysis
demographic estimation
0.012013
Cloud based multimedia analytic platform · ACM Multimedia 2013
Computer vision › Face, body and person analysis
face recognition
0.012013
Cloud based multimedia analytic platform · ACM Multimedia 2013
Computer vision › Face, body and person analysis › face recognition
face verification
0.012013
Cloud based multimedia analytic platform · ACM Multimedia 2013
Spatial and temporal data management › spatial query processing
geo-spatial query
0.012012
ASTERIX: An Open Source System for "Big Data" Management and Analysis · Proc. VLDB Endow. 2012
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.012011
Hyracks: A flexible and extensible foundation for data-intensive computing · ICDE 2011
Distributed and cloud data management
parallel data processing
0.012010
Efficient parallel set-similarity joins using MapReduce · SIGMOD Conference 2010
Information retrieval › text analysis
document collection analysis
0.012008
Entity categorization over large document collections · KDD 2008
Information retrieval › similarity measure
string similarity
0.012008
SEPIA: estimating selectivities of approximate string predicates in large Databases · VLDB J. 2008

Methods — techniques the papers use, named apart from their topics

distributed computing · 0.3partitioned-parallel execution · 0.2DAG-based dataflow · 0.2partitioning · 0.2mapreduce · 0.2sampling · 0.1
YearPublicationVenuePosition
2020 Similarity query support in big data management systems
Taewoo Kim 0001, Wenhai Li, Alexander Behm, Inci Cetindil, Rares Vernica, Vinayak R. Borkar, Michael J. Carey 0001, Chen Li 0001
Inf. Syst.5
2019 Using text mining for personalization and recommendation for an enriched hybrid learning experience
abstract
Abstract Students are nowadays given many options to consume educational content in digital formats as alternatives to printed material. Previous research suggests that while digital content has advantages, printed media still provides other benefits that cannot be matched by digital. Therefore, technology should leverage the benefits of both. In this paper, we present the Meaningful Education and Training Information System, a multifaceted hybrid textbook learning platform. The goal of the system is to provide an easy digital‐to‐print‐to‐digital content creation and reading service. The Meaningful Education and Training Information System incorporates technologies for layout, personalization, cocreation, and assessments. These facilitate common teacher/student tasks and help provide a richer, more effective learning experience. Our system has been demonstrated in multiple international education events, partner engagements, and pilots with local universities and high schools.
Rares Vernica, Tamir Hassan, Niranjan Damera-Venkata
Comput. Intell.2
2018 Supporting Similarity Queries in Apache AsterixDB
Taewoo Kim 0001, Wenhai Li, Alexander Behm, Inci Cetindil, Rares Vernica, Vinayak R. Borkar, Michael J. Carey 0001, Chen Li 0001
EDBT5
2016 METIS: A Multi-faceted Hybrid Book Learning Platform
abstract
Today, students are offered a wide variety of alternatives to printed material for the consumption of educational content. Previous research suggests that, while digital content has its advantages, printed content still offers benefits that cannot be matched by digital media. This paper introduces the Meaningful Education and Training Information System (METIS), a multi-faceted hybrid book learning platform. The goal of the system is to provide an easy digital-to-print-to-digital content creation and reading service. METIS incorporates technology for layout, personalization, co-creation and assessment. These facilitate and, in many cases, significantly simplify common teacher/student tasks. Our system has been demonstrated at several international education events, partner engagements, and pilots with local universities and high schools. We present the system and discuss how it enables hybrid learning.
Rares Vernica, Tamir Hassan, Niranjan Damera-Venkata, Jian Fan, Jerry Liu, Steven J. Simske, Shanchan Wu
DocEng2
2015 AERO: An Extensible Framework for Adaptive Web Layout Synthesis
abstract
We present AERO, an extensible framework for adaptive web layout synthesis. The goal is to provide an underlying software architecture to allow general adaptive layout behaviors. The framework consists of a 1) a suite of templates specified in HTML/CSS, 2) A hierarchical, highly customizable scoring function specification and 3) An evaluation engine that leverages native browser rendering to rapidly render content and apply the scoring functions. Unlike current responsive layout frameworks for web (e.g., Twitter Bootstrap) that have pre-configured grid layouts that adapt in a manually pre-encoded content-independent manner, AERO allows layout to adapt automatically based on multiple content-dependent criteria like aesthetic quality, cropability of individual images, layout A/B testing results, Ad placement etc.
Rares Vernica, Niranjan Damera-Venkata
DocEng1
2014 AsterixDB: A Scalable, Open Source BDMS
abstract
AsterixDB is a new, full-function BDMS (Big Data Management System) with a feature set that distinguishes it from other platforms in today's open source Big Data ecosystem. Its features make it well-suited to applications like web data warehousing, social data storage and analysis, and other use cases related to Big Data. AsterixDB has a flexible NoSQL style data model; a query language that supports a wide range of queries; a scalable runtime; partitioned, LSM-based data storage and indexing (including B + -tree, R-tree, and text indexes); support for external as well as natively stored data; a rich set of built-in types; support for fuzzy, spatial, and temporal types and queries; a built-in notion of data feeds for ingestion of data; and transaction support akin to that of a NoSQL store. Development of AsterixDB began in 2009 and led to a mid-2013 initial open source release. This paper is the first complete description of the resulting open source AsterixDB system. Covered herein are the system's data model, its query language, and its software architecture. Also included are a summary of the current status of the project and a first glimpse into how AsterixDB performs when compared to alternative technologies, including a parallel relational DBMS, a popular NoSQL store, and a popular Hadoop-based SQL data analytics platform, for things that both technologies can do. Also included is a brief description of some initial trials that the system has undergone and the lessons learned (and plans laid) based on those early "customer" engagements.
Sattam Alsubaiee, Yasser Altowim, Hotham Altwaijry, Alexander Behm, Vinayak R. Borkar, Yingyi Bu, Michael J. Carey 0001, Inci Cetindil, Madhusudan Cheelangi, Khurram Faraaz, Eugenia Gabrielova, Raman Grover, Zachary Heilbron, Young-Seok Kim, Chen Li 0001, Guangqiang Li, Ji Mahn Ok, Nicola Onose, Pouria Pirzadeh, Vassilis J. Tsotras, Rares Vernica, Till Westmann
Proc. VLDB Endow.21
2013 Cloud based multimedia analytic platform
abstract
Multimedia Analytic Platform is a cloud based service to expose state-of-the-art multimedia technologies for mobile and web application development. As a product-quality service platform, it offers comprehensive API documentation, code example, service description and sandbox for trial for each multimedia technology. The utilization of the cloud storage and distributed computing framework allows the service platform to run with robustness and efficiency. The current technologies supported by the platform include face detection, face verification, face demographic estimation, feature extraction, image matching, and image collage. Since its initial public launch in October 2012, it has been adopted by universities and third party companies for course support and application development.
Rares Vernica, Qian Lin 0001
ACM Multimedia2
2012 Adaptive MapReduce using situation-aware mappers
abstract
We propose new adaptive runtime techniques for MapReduce that improve performance and simplify job tuning. We implement these techniques by breaking a key assumption of MapReduce that mappers run in isolation. Instead, our mappers communicate through a distributed meta-data store and are aware of the global state of the job. However, we still preserve the fault-tolerance, scalability, and programming API of MapReduce. We utilize these "situation-aware mappers" to develop a set of techniques that make MapReduce more dynamic: (a) Adaptive Mappers dynamically take multiple data partitions (splits) to amortize mapper start-up costs; (b) Adaptive Combiners improve local aggregation by maintaining a cache of partial aggregates for the frequent keys; (c) Adaptive Sampling and Partitioning sample the mapper outputs and use the obtained statistics to produce balanced partitions for the reducers. Our experimental evaluation shows that adaptive techniques provide up to 3x performance improvement, in some cases, and dramatically improve performance stability across the board.
Rares Vernica, Andrey Balmin, Kevin S. Beyer, Vuk Ercegovac
EDBT1
2012 ASTERIX: An Open Source System for "Big Data" Management and Analysis
abstract
At UC Irvine, we are building a next generation parallel database system, called ASTERIX, as our approach to addressing today's "Big Data" management challenges. ASTERIX aims to combine time-tested principles from parallel database systems with those of the Web-scale computing community, such as fault tolerance for long running jobs. In this demo, we present a whirlwind tour of ASTERIX, highlighting a few of its key features. We will demonstrate examples of our data definition language to model semi-structured data, and examples of interesting queries using our declarative query language. In particular, we will show the capabilities of ASTERIX for answering geo-spatial queries and fuzzy queries, as well as ASTERIX' data feed construct for continuously ingesting data.
Sattam Alsubaiee, Yasser Altowim, Hotham Altwaijry, Alexander Behm, Vinayak R. Borkar, Yingyi Bu, Michael J. Carey 0001, Raman Grover, Zachary Heilbron, Young-Seok Kim, Chen Li 0001, Nicola Onose, Pouria Pirzadeh, Rares Vernica
Proc. VLDB Endow.14
2012 On the optimization of schedules for MapReduce workloads in the presence of shared scans
Joel L. Wolf, Andrey Balmin, Deepak Rajan, Kirsten Hildrum, Rohit Khandekar, Sujay S. Parekh, Kun-Lung Wu, Rares Vernica
VLDB J.8
2011 Hyracks: A flexible and extensible foundation for data-intensive computing
abstract
Hyracks is a new partitioned-parallel software platform designed to run data-intensive computations on large shared-nothing clusters of computers. Hyracks allows users to express a computation as a DAG of data operators and connectors. Operators operate on partitions of input data and produce partitions of output data, while connectors repartition operators' outputs to make the newly produced partitions available at the consuming operators. We describe the Hyracks end user model, for authors of dataflow jobs, and the extension model for users who wish to augment Hyracks' built-in library with new operator and/or connector types. We also describe our initial Hyracks implementation. Since Hyracks is in roughly the same space as the open source Hadoop platform, we compare Hyracks with Hadoop experimentally for several different kinds of use cases. The initial results demonstrate that Hyracks has significant promise as a next-generation platform for data-intensive applications.
Vinayak R. Borkar, Michael J. Carey 0001, Raman Grover, Nicola Onose, Rares Vernica
ICDE5
2011 ASTERIX: towards a scalable, semistructured data platform for evolving-world models
Alexander Behm, Vinayak R. Borkar, Michael J. Carey 0001, Raman Grover, Chen Li 0001, Nicola Onose, Rares Vernica, Alin Deutsch, Yannis Papakonstantinou, Vassilis J. Tsotras
Distributed Parallel Databases7
2010 Efficient parallel set-similarity joins using MapReduce
abstract
In this paper we study how to efficiently perform set-similarity joins in parallel using the popular MapReduce framework. We propose a 3-stage approach for end-to-end set-similarity joins. We take as input a set of records and output a set of joined records based on a set-similarity condition. We efficiently partition the data across nodes in order to balance the workload and minimize the need for replication. We study both self-join and R-S join cases, and show how to carefully control the amount of data kept in main memory on each node. We also propose solutions for the case where, even if we use the most fine-grained partitioning, the data still does not fit in the main memory of a node. We report results from extensive experiments on real datasets, synthetically increased in size, to evaluate the speedup and scaleup properties of the proposed algorithms using Hadoop.
Rares Vernica, Michael J. Carey 0001, Chen Li 0001
SIGMOD Conference1
2008 Entity categorization over large document collections
abstract
Extracting entities (such as people, movies) from documents and identifying the categories (such as painter, writer) they belong to enable structured querying and data analysis over unstructured document collections. In this paper, we focus on the problem of categorizing extracted entities. Most prior approaches developed for this task only analyzed the local document context within which entities occur. In this paper, we significantly improve the accuracy of entity categorization by (i) considering an entity's context across multiple documents containing it, and (ii) exploiting existing large lists of related entities (e.g., lists of actors, directors, books). These approaches introduce computational challenges because (a) the context of entities has to be aggregated across several documents and (b) the lists of related entities may be very large. We develop techniques to address these challenges. We present a thorough experimental study on real data sets that demonstrates the increase in accuracy and the scalability of our approaches.
Venkatesh Ganti, Arnd Christian König, Rares Vernica
KDD3
2008 SEPIA: estimating selectivities of approximate string predicates in large Databases
Chen Li 0001, Rares Vernica
VLDB J.3
2006 Relaxing Join and Selection Queries
Nick Koudas, Chen Li 0001, Anthony K. H. Tung, Rares Vernica
VLDB4