Venky Harinarayan

dblp:71/3835 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Knowledge graphs · 31% Web and social media mining · 21% Information retrieval · 20%
Artificial intelligence
1 paper
Information extraction and text analysis · 100%

Topics — the 16 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge graphs › knowledge graph management
knowledge base curation
0.212013
Building, maintaining, and using knowledge bases: a report from the trenches · SIGMOD Conference 2013
Information retrieval
query understanding
0.212013
Building, maintaining, and using knowledge bases: a report from the trenches · SIGMOD Conference 2013
Recommender systems › e-commerce recommendation
gift recommendation
0.112012
Anatomy of a gift recommendation engine powered by social media · SIGMOD Conference 2012
Information retrieval › web search › web information retrieval
social media retrieval
0.012013
Entity Extraction, Linking, Classification, and Tagging for Social Media: A Wikipedia-Based Approach · Proc. VLDB Endow. 2013
Web and social media mining
social media analysis
0.012012
Anatomy of a gift recommendation engine powered by social media · SIGMOD Conference 2012
Distributed and cloud data management
federated database
0.011998
Virtual Database Technology · ICDE 1998
Data integration and cleaning › data integration system
virtual database
0.011998
Virtual Database Technology · ICDE 1998
Database system architecture and tuning › database design › physical database design
index selection
0.011997
Index Selection for OLAP · ICDE 1997
Query processing and optimization › OLAP
OLAP query optimization
0.011997
Index Selection for OLAP · ICDE 1997
Query processing and optimization › OLAP
data cube
0.011996
Implementing Data Cubes Efficiently · SIGMOD Conference 1996
Web and social media mining › social network analysis › influence maximization
greedy algorithm
0.011996
Implementing Data Cubes Efficiently · SIGMOD Conference 1996
Query processing and optimization › materialized view
materialized view selection
0.011996
Implementing Data Cubes Efficiently · SIGMOD Conference 1996
Query processing and optimization
aggregate query processing
0.011995
Aggregate-Query Processing in Data Warehousing Environments · VLDB 1995
Data integration and cleaning
data warehouse
0.011995
Aggregate-Query Processing in Data Warehousing Environments · VLDB 1995
Data integration and cleaning
heterogeneous data sources
0.011998
Virtual Database Technology · ICDE 1998
Indexing and storage engines › synopsis structure
summary tables
0.011997
Index Selection for OLAP · ICDE 1997

Methods — techniques the papers use, named apart from their topics

knowledge base linking · 0.3context and social signals · 0.3case study · 0.2social media signal extraction · 0.1greedy algorithm · 0.0data virtualization · 0.0performance bounds · 0.0lattice framework · 0.0
YearPublicationVenuePosition
2013 Building, maintaining, and using knowledge bases: a report from the trenches
abstract
A knowledge base (KB) contains a set of concepts, instances, and relationships. Over the past decade, numerous KBs have been built, and used to power a growing array of applications. Despite this flurry of activities, however, surprisingly little has been published about the end-to-end process of building, maintaining, and using such KBs in industry. In this paper we describe such a process. In particular, we describe how we build, update, and curate a large KB at Kosmix, a Bay Area startup, and later at WalmartLabs, a development and research lab of Walmart. We discuss how we use this KB to power a range of applications, including query understanding, Deep Web search, in-context advertising, event monitoring in social media, product search, social gifting, and social mining. Finally, we discuss how the KB team is organized, and the lessons learned. Our goal with this paper is to provide a real-world case study, and to contribute to the emerging direction of building, maintaining, and using knowledge bases for data management applications.
Omkar Deshpande, Digvijay S. Lamba, Michel Tourn, Sanjib Das, Sri Subramaniam, Anand Rajaraman, Venky Harinarayan, AnHai Doan
SIGMOD Conference7
2013 Entity Extraction, Linking, Classification, and Tagging for Social Media: A Wikipedia-Based Approach
abstract
Many applications that process social data, such as tweets, must extract entities from tweets (e.g., "Obama" and "Hawaii" in "Obama went to Hawaii"), link them to entities in a knowledge base (e.g., Wikipedia), classify tweets into a set of predefined topics, and assign descriptive tags to tweets. Few solutions exist today to solve these problems for social data, and they are limited in important ways. Further, even though several industrial systems such as OpenCalais have been deployed to solve these problems for text data, little if any has been published about them, and it is unclear if any of the systems has been tailored for social media. In this paper we describe in depth an end-to-end industrial system that solves these problems for social data. The system has been developed and used heavily in the past three years, first at Kosmix, a startup, and later at WalmartLabs. We show how our system uses a Wikipedia-based global "real-time" knowledge base that is well suited for social data, how we interleave the tasks in a synergistic fashion, how we generate and use contexts and social signals to improve task accuracy, and how we scale the system to the entire Twitter firehose. We describe experiments that show that our system outperforms current approaches. Finally we describe applications of the system at Kosmix and WalmartLabs, and lessons learned.
Rohit Kumar 0006, Digvijay S. Lamba, Nikesh Garera, Mitul Tiwari, Xiaoyong Chai, Sanjib Das, Sri Subramaniam, Anand Rajaraman, Venky Harinarayan, AnHai Doan
Proc. VLDB Endow.9
2012 Anatomy of a gift recommendation engine powered by social media
abstract
More and more people conduct their shopping online [1], especially during the holiday season [2]. Shopping online offers a lot of convenience, including the luxury of shopping from home, the ease of research, better prices, and in many cases access to unique products not available in stores.
Yannis Pavlidis, Madhusudan Mathihalli, Indrani Chakravarty, Arvind Batra, Ron Benson, Ravi Raj, Robert Yau, Mike McKiernan, Venky Harinarayan, Anand Rajaraman
SIGMOD Conference9
1998 Virtual Database Technology
abstract
Virtual database (VDB) technology makes external data behave as an extension of an enterprise's relational database (RDBMS) system. VDB technology enables the rapid deployment of applications with at least one of the following characteristics: large numbers of data sources; data sources that are autonomous (i.e. there is no centralized control); or data sources that can have a mixture of structured and unstructured data. The World Wide Web and most intranets have all of these characteristics and can thus benefit from VDB technology.
Ashish Gupta 0001, Venky Harinarayan, Anand Rajaraman
ICDE2
1997 Index Selection for OLAP
abstract
On-line analytical processing (OLAP) is a recent and important application of database systems. Typically, OLAP data is presented as a multidimensional "data cube." OLAP queries are complex and can take many hours or even days to run, if executed directly on the raw data. The most common method of reducing execution time is to precompute some of the queries into summary tables (subcubes of the data cube) and then to build indexes on these summary tables. In most commercial OLAP systems today, the summary tables that are to be precomputed are picked first, followed by the selection of the appropriate indexes on them. A trial-and-error approach is used to divide the space available between the summary tables and the indexes. This two-step process can perform very poorly. Since both summary tables and indexes consume the same resource-space-their selection should be done together for the most efficient use of space. The authors give algorithms that automate the selection of summary tables and indexes. In particular, they present a family of algorithms of increasing time complexities, and prove strong performance bounds for them. The algorithms with higher complexities have better performance bounds. However, the increase in the performance bound is diminishing, and they show that an algorithm of moderate complexity can perform fairly close to the optimal.
Himanshu Gupta 0001, Venky Harinarayan, Anand Rajaraman, Jeffrey D. Ullman
ICDE2
1996 Implementing Data Cubes Efficiently
abstract
Decision support applications involve complex queries on very large databases. Since response times should be small, query optimization is critical. Users typically view the data as multidimensional data cubes. Each cell of the data cube is a view consisting of an aggregation of interest, like total sales. The values of many of these cells are dependent on the values of other cells in the data cube..A common and powerful query optimization technique is to materialize some or all of these cells rather than compute them from raw data each time. Commercial systems differ mainly in their approach to materializing the data cube. In this paper, we investigate the issue of which cells (views) to materialize when it is too expensive to materialize all views. A lattice framework is used to express dependencies among views. We present greedy algorithms that work off this lattice and determine a good set of views to materialize. The greedy algorithm performs within a small constant factor of optimal under a variety of models. We then consider the most common case of the hypercube lattice and examine the choice of materialized views for hypercubes in detail, giving some good tradeoffs between the space used and the average time to answer a query. 1
Venky Harinarayan, Anand Rajaraman, Jeffrey D. Ullman
SIGMOD Conference1
1995 Optimization Using Tuple Subsumption
Venky Harinarayan, Ashish Gupta 0001
ICDT1
1995 Aggregate-Query Processing in Data Warehousing Environments
Ashish Gupta 0001, Venky Harinarayan, Dallan Quass
VLDB2