VLDB 2026 Research / reviewers in the wild / expert
Venky Harinarayan
dblp:71/3835
· DBLP profile ↗
8ranked-venue papers
2as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
7 papers |
Knowledge graphs · 31% Web and social media mining · 21% Information retrieval · 20% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge graphs › knowledge graph management
knowledge base curation |
0.2 | 1 | 2013 | Building, maintaining, and using knowledge bases: a report from the trenches · SIGMOD Conference 2013 |
Information retrieval
query understanding |
0.2 | 1 | 2013 | Building, maintaining, and using knowledge bases: a report from the trenches · SIGMOD Conference 2013 |
Recommender systems › e-commerce recommendation
gift recommendation |
0.1 | 1 | 2012 | Anatomy of a gift recommendation engine powered by social media · SIGMOD Conference 2012 |
Information retrieval › web search › web information retrieval
social media retrieval |
0.0 | 1 | 2013 | Entity Extraction, Linking, Classification, and Tagging for Social Media: A Wikipedia-Based Approach · Proc. VLDB Endow. 2013 |
Web and social media mining
social media analysis |
0.0 | 1 | 2012 | Anatomy of a gift recommendation engine powered by social media · SIGMOD Conference 2012 |
Distributed and cloud data management
federated database |
0.0 | 1 | 1998 | Virtual Database Technology · ICDE 1998 |
Data integration and cleaning › data integration system
virtual database |
0.0 | 1 | 1998 | Virtual Database Technology · ICDE 1998 |
Database system architecture and tuning › database design › physical database design
index selection |
0.0 | 1 | 1997 | Index Selection for OLAP · ICDE 1997 |
Query processing and optimization › OLAP
OLAP query optimization |
0.0 | 1 | 1997 | Index Selection for OLAP · ICDE 1997 |
Query processing and optimization › OLAP
data cube |
0.0 | 1 | 1996 | Implementing Data Cubes Efficiently · SIGMOD Conference 1996 |
Web and social media mining › social network analysis › influence maximization
greedy algorithm |
0.0 | 1 | 1996 | Implementing Data Cubes Efficiently · SIGMOD Conference 1996 |
Query processing and optimization › materialized view
materialized view selection |
0.0 | 1 | 1996 | Implementing Data Cubes Efficiently · SIGMOD Conference 1996 |
Query processing and optimization
aggregate query processing |
0.0 | 1 | 1995 | Aggregate-Query Processing in Data Warehousing Environments · VLDB 1995 |
Data integration and cleaning
data warehouse |
0.0 | 1 | 1995 | Aggregate-Query Processing in Data Warehousing Environments · VLDB 1995 |
Data integration and cleaning
heterogeneous data sources |
0.0 | 1 | 1998 | Virtual Database Technology · ICDE 1998 |
Indexing and storage engines › synopsis structure
summary tables |
0.0 | 1 | 1997 | Index Selection for OLAP · ICDE 1997 |
Methods — techniques the papers use, named apart from their topics
knowledge base linking · 0.3context and social signals · 0.3case study · 0.2social media signal extraction · 0.1greedy algorithm · 0.0data virtualization · 0.0performance bounds · 0.0lattice framework · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Building, maintaining, and using knowledge bases: a report from the trenchesabstractA knowledge base (KB) contains a set of concepts, instances, and relationships. Over the past decade, numerous KBs have been built, and used to power a growing array of applications. Despite this flurry of activities, however, surprisingly little has been published about the end-to-end process of building, maintaining, and using such KBs in industry. In this paper we describe such a process. In particular, we describe how we build, update, and curate a large KB at Kosmix, a Bay Area startup, and later at WalmartLabs, a development and research lab of Walmart. We discuss how we use this KB to power a range of applications, including query understanding, Deep Web search, in-context advertising, event monitoring in social media, product search, social gifting, and social mining. Finally, we discuss how the KB team is organized, and the lessons learned. Our goal with this paper is to provide a real-world case study, and to contribute to the emerging direction of building, maintaining, and using knowledge bases for data management applications. Omkar Deshpande, Digvijay S. Lamba, Michel Tourn, Sanjib Das, Sri Subramaniam, Anand Rajaraman, Venky Harinarayan, AnHai Doan |
SIGMOD Conference | 7 |
| 2013 | Entity Extraction, Linking, Classification, and Tagging for Social Media: A Wikipedia-Based ApproachabstractMany applications that process social data, such as tweets, must extract entities from tweets (e.g., "Obama" and "Hawaii" in "Obama went to Hawaii"), link them to entities in a knowledge base (e.g., Wikipedia), classify tweets into a set of predefined topics, and assign descriptive tags to tweets. Few solutions exist today to solve these problems for social data, and they are limited in important ways. Further, even though several industrial systems such as OpenCalais have been deployed to solve these problems for text data, little if any has been published about them, and it is unclear if any of the systems has been tailored for social media. In this paper we describe in depth an end-to-end industrial system that solves these problems for social data. The system has been developed and used heavily in the past three years, first at Kosmix, a startup, and later at WalmartLabs. We show how our system uses a Wikipedia-based global "real-time" knowledge base that is well suited for social data, how we interleave the tasks in a synergistic fashion, how we generate and use contexts and social signals to improve task accuracy, and how we scale the system to the entire Twitter firehose. We describe experiments that show that our system outperforms current approaches. Finally we describe applications of the system at Kosmix and WalmartLabs, and lessons learned. Rohit Kumar 0006, Digvijay S. Lamba, Nikesh Garera, Mitul Tiwari, Xiaoyong Chai, Sanjib Das, Sri Subramaniam, Anand Rajaraman, Venky Harinarayan, AnHai Doan |
Proc. VLDB Endow. | 9 |
| 2012 | Anatomy of a gift recommendation engine powered by social mediaabstractMore and more people conduct their shopping online [1], especially during the holiday season [2]. Shopping online offers a lot of convenience, including the luxury of shopping from home, the ease of research, better prices, and in many cases access to unique products not available in stores. Yannis Pavlidis, Madhusudan Mathihalli, Indrani Chakravarty, Arvind Batra, Ron Benson, Ravi Raj, Robert Yau, Mike McKiernan, Venky Harinarayan, Anand Rajaraman |
SIGMOD Conference | 9 |
| 1998 | Virtual Database TechnologyabstractVirtual database (VDB) technology makes external data behave as an extension of an enterprise's relational database (RDBMS) system. VDB technology enables the rapid deployment of applications with at least one of the following characteristics: large numbers of data sources; data sources that are autonomous (i.e. there is no centralized control); or data sources that can have a mixture of structured and unstructured data. The World Wide Web and most intranets have all of these characteristics and can thus benefit from VDB technology. Ashish Gupta 0001, Venky Harinarayan, Anand Rajaraman |
ICDE | 2 |
| 1997 | Index Selection for OLAPabstractOn-line analytical processing (OLAP) is a recent and important application of database systems. Typically, OLAP data is presented as a multidimensional "data cube." OLAP queries are complex and can take many hours or even days to run, if executed directly on the raw data. The most common method of reducing execution time is to precompute some of the queries into summary tables (subcubes of the data cube) and then to build indexes on these summary tables. In most commercial OLAP systems today, the summary tables that are to be precomputed are picked first, followed by the selection of the appropriate indexes on them. A trial-and-error approach is used to divide the space available between the summary tables and the indexes. This two-step process can perform very poorly. Since both summary tables and indexes consume the same resource-space-their selection should be done together for the most efficient use of space. The authors give algorithms that automate the selection of summary tables and indexes. In particular, they present a family of algorithms of increasing time complexities, and prove strong performance bounds for them. The algorithms with higher complexities have better performance bounds. However, the increase in the performance bound is diminishing, and they show that an algorithm of moderate complexity can perform fairly close to the optimal. Himanshu Gupta 0001, Venky Harinarayan, Anand Rajaraman, Jeffrey D. Ullman |
ICDE | 2 |
| 1996 | Implementing Data Cubes EfficientlyabstractDecision support applications involve complex queries on very large databases. Since response times should be small, query optimization is critical. Users typically view the data as multidimensional data cubes. Each cell of the data cube is a view consisting of an aggregation of interest, like total sales. The values of many of these cells are dependent on the values of other cells in the data cube..A common and powerful query optimization technique is to materialize some or all of these cells rather than compute them from raw data each time. Commercial systems differ mainly in their approach to materializing the data cube. In this paper, we investigate the issue of which cells (views) to materialize when it is too expensive to materialize all views. A lattice framework is used to express dependencies among views. We present greedy algorithms that work off this lattice and determine a good set of views to materialize. The greedy algorithm performs within a small constant factor of optimal under a variety of models. We then consider the most common case of the hypercube lattice and examine the choice of materialized views for hypercubes in detail, giving some good tradeoffs between the space used and the average time to answer a query. 1 Venky Harinarayan, Anand Rajaraman, Jeffrey D. Ullman |
SIGMOD Conference | 1 |
| 1995 | Optimization Using Tuple Subsumption
Venky Harinarayan, Ashish Gupta 0001 |
ICDT | 1 |
| 1995 | Aggregate-Query Processing in Data Warehousing Environments
Ashish Gupta 0001, Venky Harinarayan, Dallan Quass |
VLDB | 2 |