Ramakrishna Varadarajan

dblp:63/2265 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 8 first-authorArtificial intelligence and machine learning · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Information retrieval · 37% Query processing and optimization · 24% Database system architecture and tuning · 18%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval models
0.222010
Using Proximity Search to Estimate Authority Flow · IEEE Trans. Knowl. Data Eng. 2010
Explaining and Reformulating Authority Flow Queries · ICDE 2008
Database system architecture and tuning › database design
physical database design
0.212014
DBDesigner: A customizable physical design tool for Vertica Analytic Database · ICDE 2014
Query processing and optimization › materialization
late materialization
0.212013
Materialization strategies in the Vertica analytic database: Lessons learned · ICDE 2013
Data models and query languages › datalog › datalog query optimization
sideways information passing
0.212013
Materialization strategies in the Vertica analytic database: Lessons learned · ICDE 2013
Information retrieval
web search
0.122008
Beyond Single-Page Web Search Results · IEEE Trans. Knowl. Data Eng. 2008
Searching the web using composed pages · SIGIR 2006
Spatial and temporal data management › spatial query processing
proximity search
0.112010
Using Proximity Search to Estimate Authority Flow · IEEE Trans. Knowl. Data Eng. 2010
Information retrieval
ranking
0.112010
Using Proximity Search to Estimate Authority Flow · IEEE Trans. Knowl. Data Eng. 2010
Information retrieval › text summarization
query-focused summarization
0.112008
Beyond Single-Page Web Search Results · IEEE Trans. Knowl. Data Eng. 2008
Information retrieval
query reformulation
0.112008
Explaining and Reformulating Authority Flow Queries · ICDE 2008
Query processing and optimization
query result explanation
0.112008
Explaining and Reformulating Authority Flow Queries · ICDE 2008
Information retrieval › text summarization
web page summarization
0.112008
Beyond Single-Page Web Search Results · IEEE Trans. Knowl. Data Eng. 2008
Indexing and storage engines
column store
0.012013
Materialization strategies in the Vertica analytic database: Lessons learned · ICDE 2013
Query processing and optimization
analytical query processing
0.012012
The Vertica Analytic Database: C-Store 7 Years Later · Proc. VLDB Endow. 2012

Methods — techniques the papers use, named apart from their topics

optimizer cost estimation · 0.2cost-benefit model · 0.2experimental comparison · 0.2approximation · 0.1user feedback · 0.1personalized ranking · 0.1hyperlink analysis · 0.1heuristic algorithm · 0.1
YearPublicationVenuePosition
2014 DBDesigner: A customizable physical design tool for Vertica Analytic Database
abstract
In this paper, we present Vertica's customizable physical design tool, called the DBDesigner (DBD), that produces designs optimized for various scenarios and applications. For a given workload and space budget, DBD automatically recommends a physical design that optimizes query performance, storage footprint, fault tolerance and recovery to meet different customer requirements. Vertica is a distributed, massively parallel columnar database that physically organizes data into projections. Projections are attribute subsets from one or more tables with tuples sorted by one or more attributes, that are replicated or segmented (distributed) on cluster nodes. The key challenges involved in projection design are picking appropriate column sets, sort orders, cluster data distributions and column encodings. To achieve the desired trade-off between query performance and storage footprint, DBD operates under three different design policies: (a) load-optimized, (b) query-optimized or (c) balanced. These policies indirectly control the number of projections proposed and queries optimized to achieve the desired balance. To cater to query workloads that evolve over time, DBD also operates in a comprehensive and incremental design mode. In addition, DBD lets users override specific features of projection design based on their intimate knowledge about the data and query workloads. We present the complete physical design algorithm, describing in detail how projection candidates are efficiently explored and evaluated using optimizer's cost and benefit model. Our experimental results show that DBD produces good physical designs that satisfy a variety of customer use cases.
Ramakrishna Varadarajan, Vivek Bharathan, Ariel Cary, Jaimin Dave, Sreenath Bodagala
ICDE1
2013 Materialization strategies in the Vertica analytic database: Lessons learned
abstract
Column store databases allow for various tuple reconstruction strategies (also called materialization strategies). Early materialization is easy to implement but generally performs worse than late materialization. Late materialization is more complex to implement, and usually performs much better than early materialization, although there are situations where it is worse. We identify these situations, which essentially revolve around joins where neither input fits in memory (also called spilling joins). Sideways information passing techniques provide a viable solution to get the best of both worlds. We demonstrate how early materialization combined with sideways information passing allows us to get the benefits of late materialization, without the bookkeeping complexity or worse performance for spilling joins. It also provides some other benefits to query processing in Vertica due to positive interaction with compression and sort orders of the data. In this paper, we report our experiences with late and early materialization, highlight their strengths and weaknesses, and present the details of our sideways information passing implementation. We show experimental results of comparing these materialization strategies, which highlight the significant performance improvements provided by our implementation of sideways information passing (up to 72% on some TPC-H queries).
Lakshmikant Shrinivas, Sreenath Bodagala, Ramakrishna Varadarajan, Ariel Cary, Vivek Bharathan, Chuck Bear
ICDE3
2013 Comparing top-k XML lists
Ramakrishna Varadarajan, Fernando Farfán, Vagelis Hristidis
Inf. Syst.1
2012 WYSIWYE: An Algebra for Expressing Spatial and Textual Rules for Information Extraction
Vijil Chenthamarakshan, Ramakrishna Varadarajan, Prasad Deshpande, Raghu Krishnapuram, Knut Stolze
WAIM2
2012 The Vertica Analytic Database: C-Store 7 Years Later
abstract
This paper describes the system architecture of the Vertica Analytic Database (Vertica), a commercialization of the design of the C-Store research prototype. Vertica demonstrates a modern commercial RDBMS system that presents a classical relational interface while at the same time achieving the high performance expected from modern "web scale" analytic systems by making appropriate architectural choices. Vertica is also an instructive lesson in how academic systems research can be directly commercialized into a successful product.
Andrew Lamb, Matt Fuller, Ramakrishna Varadarajan, Nga Tran 0001, Ben Vandiver, Lyric Doshi, Chuck Bear
Proc. VLDB Endow.3
2010 Using Proximity Search to Estimate Authority Flow
abstract
Authority flow and proximity search have been used extensively in measuring the association between entities in data graphs, ranging from the web to relational and XML databases. These two ranking factors have been used and studied separately in the past. In addition to their semantic differences, a key advantage of proximity search is the existence of efficient execution algorithms. In contrast, due to the complexity of calculating the authority flow, current systems only use precomputed authority flows in runtime. This limitation prohibits authority flow to be used more effectively as a ranking factor. In this paper, we present a comparative analysis of the two ranking factors. We present an efficient approximation of authority flow based on proximity search. We analytically estimate the approximation error and how this affects the ranking of the results of a query.
Vagelis Hristidis, Yannis Papakonstantinou, Ramakrishna Varadarajan
IEEE Trans. Knowl. Data Eng.3
2009 Flexible and efficient querying and ranking on hyperlinked data sources
abstract
There has been an explosion of hyperlinked data in many domains, e.g., the biological Web. Expressive query languages and effective ranking techniques are required to convert this data into browsable knowledge. We propose the Graph Information Discovery (GID) framework to support sophisticated user queries on a rich web of annotated and hyperlinked data entries, where query answers need to be ranked in terms of some customized ranking criteria, e.g., PageRank or ObjectRank. GID has a data model that includes a schema graph and a data graph, and an intuitive query interface. The GID framework allows users to easily formulate queries consisting of sequences of hard filters (selection predicates) and soft filters (ranking criteria); it can also be combined with other specialized graph query languages to enhance their ranking capabilities. GID queries have a well-defined semantics and are implemented by a set of physical operators, each of which produces a ranked result graph. We discuss rewriting opportunities to provide an efficient evaluation of GID queries. Soft filters are a key feature of GID and they are implemented using authority flow ranking techniques; these are query dependent rankings and are expensive to compute at runtime. We present approximate optimization techniques for GID soft filter queries based on the properties of random walks, and using novel path-length-bound and graph-sampling approximation techniques. We experimentally validate our optimization techniques on large biological and bibliographic datasets. Our techniques can produce high quality (Top K) answers with a savings of up to an order of magnitude, in comparison to the evaluation time for the exact solution.
Ramakrishna Varadarajan, Vagelis Hristidis, Louiqa Raschid, Maria-Esther Vidal, Luis-Daniel Ibáñez, Héctor Rodríguez-Drumond
EDBT1
2008 Explaining and Reformulating Authority Flow Queries
abstract
Authority flow is an effective ranking mechanism for answering queries on a broad class of data. Systems have been developed to apply this principle on the Web (PageRank and topic sensitive PageRank), bibliographic databases (ObjectRank), and biological databases (Hubs of Knowledge project). However, these systems have the following drawbacks: (a) There is no way to explain to the user why a particular result received its current score; (b) The authority flow rates, which have been shown to dramatically affect the results' quality in ObjectRank, have to be set manually by a domain expert; (c) There is no query reformulation methodology to refine the query results according to the user's preferences. In this work, we address these shortcomings by introducing a framework and algorithms to explain query results and reformulate authority flow queries based on the user's feedback. The query reformulation process can be used to learn the user's preferences and automatically adjust the authority flow rates to facilitate personalized authority flow searching. We experimentally evaluate our algorithms in terms of performance and quality.
Ramakrishna Varadarajan, Vagelis Hristidis, Louiqa Raschid
ICDE1
2008 Beyond Single-Page Web Search Results
abstract
Given a user keyword query, current Web search engines return a list of individual Web pages ranked by their "goodness" with respect to the query. Thus, the basic unit for search and retrieval is an individual page, even though information on a topic is often spread across multiple pages. This degrades the quality of search results, especially for long or uncorrelated (multitopic) queries (in which individual keywords rarely occur together in the same document), where a single page is unlikely to satisfy the user's information need. We propose a technique that, given a keyword query, on the fly generates new pages, called composed pages, which contain all query keywords. The composed pages are generated by extracting and stitching together relevant pieces from hyperlinked Web pages and retaining links to the original Web pages. To rank the composed pages, we consider both the hyperlink structure of the original pages and the associations between the keywords within each page. Furthermore, we present and experimentally evaluate heuristic algorithms to efficiently generate the top composed pages. The quality of our method is compared to current approaches by using user surveys. Finally, we also show how our techniques can be used to perform query-specific summarization of Web pages.
Ramakrishna Varadarajan, Vagelis Hristidis, Tao Li 0001
IEEE Trans. Knowl. Data Eng.1
2006 A system for query-specific document summarization
abstract
There has been a great amount of work on query-independent summarization of documents. However, due to the success of Web search engines query-specific document summarization (query result snippets) has become an important problem, which has received little attention. We present a method to create query-specific summaries by identifying the most query-relevant fragments and combining them using the semantic associations within the document. In particular, we first add structure to the documents in the preprocessing stage and convert them to document graphs. Then, the best summaries are computed by calculating the top spanning trees on the document graphs. We present and experimentally evaluate efficient algorithms that support computing summaries in interactive time. Furthermore, the quality of our summarization method is compared to current approaches using a user survey.
Ramakrishna Varadarajan, Vagelis Hristidis
CIKM1
2006 Searching the web using composed pages
abstract
No abstract available.
Ramakrishna Varadarajan, Vagelis Hristidis, Tao Li 0001
SIGIR1
2005 Structure-based query-specific document summarization
abstract
Summarization of text documents is increasingly important with the amount of data available on the Internet. The large majority of current approaches view documents as linear sequences of words and create query-independent summaries. However, ignoring the structure of the document degrades the quality of summaries. Furthermore, the popularity of web search engines requires query-specific summaries. We present a method to create query-specific summaries by adding structure to documents by extracting associations between their fragments.
Ramakrishna Varadarajan, Vagelis Hristidis
CIKM1