VLDB 2026 Research / reviewers in the wild / expert
Anand Rajaraman
dblp:r/ARajaraman
· DBLP profile ↗
29ranked-venue papers
7as first author
0since 2021 · last 2016
0000-0003-4484-6721ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 25 · 6 first-authorTheory of computation · 4 · 1 first-authorArtificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
20 papers |
Information retrieval · 26% Web and social media mining · 18% Knowledge graphs · 14% | |
| Artificial intelligence
4 papers |
Probabilistic and Bayesian machine learning · 35% Information extraction and text analysis · 26% Question answering and dialogue systems · 15% |
Topics — the 30 heaviest of 46, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process |
0.2 | 1 | 2015 | SEISMIC: A Self-Exciting Point Process Model for Predicting Tweet Popularity · KDD 2015 |
Web and social media mining › information diffusion
information diffusion prediction |
0.2 | 1 | 2015 | SEISMIC: A Self-Exciting Point Process Model for Predicting Tweet Popularity · KDD 2015 |
Knowledge graphs › knowledge graph management
knowledge base curation |
0.2 | 1 | 2013 | Building, maintaining, and using knowledge bases: a report from the trenches · SIGMOD Conference 2013 |
Information retrieval
query understanding |
0.2 | 1 | 2013 | Building, maintaining, and using knowledge bases: a report from the trenches · SIGMOD Conference 2013 |
Recommender systems › e-commerce recommendation
gift recommendation |
0.1 | 1 | 2012 | Anatomy of a gift recommendation engine powered by social media · SIGMOD Conference 2012 |
Data mining › text mining › information extraction
concept mining |
0.1 | 1 | 2010 | Towards The Web of Concepts: Extracting Concepts from Large Datasets · Proc. VLDB Endow. 2010 |
Data mining › pattern mining › association rule mining
market basket analysis |
0.1 | 1 | 2010 | Towards The Web of Concepts: Extracting Concepts from Large Datasets · Proc. VLDB Endow. 2010 |
Natural language and speech › Question answering and dialogue systems
question answering over structured data |
0.1 | 1 | 2009 | Answering Web Questions Using Structured Data - Dream or Reality? · Proc. VLDB Endow. 2009 |
Information retrieval › distributed information retrieval
federated search |
0.1 | 1 | 2009 | Kosmix: High-Performance Topic Exploration using the Deep Web · Proc. VLDB Endow. 2009 |
Information retrieval › web search › web information retrieval
hidden web |
0.1 | 1 | 2009 | Kosmix: High-Performance Topic Exploration using the Deep Web · Proc. VLDB Endow. 2009 |
Query processing and optimization › query rewriting
query transformation |
0.1 | 1 | 2009 | Kosmix: High-Performance Topic Exploration using the Deep Web · Proc. VLDB Endow. 2009 |
Information retrieval › interactive information retrieval › exploratory search
topic exploration |
0.1 | 1 | 2009 | Kosmix: High-Performance Topic Exploration using the Deep Web · Proc. VLDB Endow. 2009 |
Information retrieval
web search |
0.1 | 1 | 2009 | Answering Web Questions Using Structured Data - Dream or Reality? · Proc. VLDB Endow. 2009 |
Computer vision › Image recognition and object detection
visual recognition |
0.1 | 1 | 2016 | Data-Driven Disruption: The View from Silicon Valley · Proc. VLDB Endow. 2016 |
Cloud and datacenter computing
big data platform |
0.1 | 1 | 2016 | Data-Driven Disruption: The View from Silicon Valley · Proc. VLDB Endow. 2016 |
Information retrieval › web search › web information retrieval
social media retrieval |
0.0 | 1 | 2013 | Entity Extraction, Linking, Classification, and Tagging for Social Media: A Wikipedia-Based Approach · Proc. VLDB Endow. 2013 |
Distributed and cloud data management
mapreduce |
0.0 | 1 | 2012 | Muppet: MapReduce-Style Processing of Fast Data · Proc. VLDB Endow. 2012 |
Web and social media mining
social media analysis |
0.0 | 1 | 2012 | Anatomy of a gift recommendation engine powered by social media · SIGMOD Conference 2012 |
Information retrieval
query log analysis |
0.0 | 1 | 2010 | Towards The Web of Concepts: Extracting Concepts from Large Datasets · Proc. VLDB Endow. 2010 |
Data integration and cleaning › data extraction
web data extraction |
0.0 | 1 | 2001 | Querying Websites Using Compact Skeletons · PODS 2001 |
Data integration and cleaning › data extraction › web data extraction
wrapper generation |
0.0 | 1 | 2001 | Querying Websites Using Compact Skeletons · PODS 2001 |
Query processing and optimization › query rewriting
query answering using views |
0.0 | 2 | 1996 | Answering Queries Using Limited External Processors · PODS 1996 Answering Queries Using Templates with Binding Patterns · PODS 1995 |
Indexing and storage engines
caching |
0.0 | 1 | 2009 | Kosmix: High-Performance Topic Exploration using the Deep Web · Proc. VLDB Endow. 2009 |
Data integration and cleaning
heterogeneous data sources |
0.0 | 2 | 1998 | Querying Heterogeneous Information Sources Using Source Descriptions · VLDB 1996 Virtual Database Technology · ICDE 1998 |
Distributed and cloud data management
federated database |
0.0 | 1 | 1998 | Virtual Database Technology · ICDE 1998 |
Data integration and cleaning › data integration system
virtual database |
0.0 | 1 | 1998 | Virtual Database Technology · ICDE 1998 |
Database system architecture and tuning › database design › physical database design
index selection |
0.0 | 1 | 1997 | Index Selection for OLAP · ICDE 1997 |
Query processing and optimization › OLAP
OLAP query optimization |
0.0 | 1 | 1997 | Index Selection for OLAP · ICDE 1997 |
Data mining › time series analysis
change point detection |
0.0 | 1 | 1996 | Change Detection in Hierarchically Structured Information · SIGMOD Conference 1996 |
Query processing and optimization › OLAP
data cube |
0.0 | 1 | 1996 | Implementing Data Cubes Efficiently · SIGMOD Conference 1996 |
Methods — techniques the papers use, named apart from their topics
self-exciting point process · 0.4knowledge base linking · 0.3context and social signals · 0.3case study · 0.2social media signal extraction · 0.1mapupdate · 0.1statistical measures of support and confidence · 0.1federated search · 0.1greedy algorithm · 0.0skeleton-based extraction · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Data-Driven Disruption: The View from Silicon ValleyabstractWe live in an era where software is transforming industries, the sciences, and society as a whole. This exciting phenomenon has been described by the phrase "software is eating the world." It is becoming increasingly apparent that data is the fuel powering software's conquests. Data is the new disruptor. It's hard to believe that the first decade of the Big Data era is already behind us. Silicon Valley has been at the forefront of developing and applying data-driven approaches to create disruption at many levels: infrastructure (e.g., Hadoop and Spark), capabilities (e.g., image recognition and machine translation), and killer apps (e.g., self-driving cars and messaging bots). In this talk, we first look back on the past decade and share learnings from the frontlines of data-driven disruption. Looking ahead, we then describe challenges and opportunities for the next decade. Since this has also been a personal journey, we will use examples drawn from personal experience to illustrate each point. Anand Rajaraman |
Proc. VLDB Endow. | 1 |
| 2015 | SEISMIC: A Self-Exciting Point Process Model for Predicting Tweet PopularityabstractSocial networking websites allow users to create and share content. Big information cascades of post resharing can form as users of these sites reshare others' posts with their friends and followers. One of the central challenges in understanding such cascading behaviors is in forecasting information outbreaks, where a single post becomes widely popular by being reshared by many users. In this paper, we focus on predicting the final number of reshares of a given post. We build on the theory of self-exciting point processes to develop a statistical model that allows us to make accurate predictions. Our model requires no training or expensive feature engineering. It results in a simple and efficiently computable formula that allows us to answer questions, in real-time, such as: Given a post's resharing history so far, what is our current estimate of its final number of reshares? Is the post resharing cascade past the initial stage of explosive growth? And, which posts will be the most reshared in the future? Qingyuan Zhao, Murat A. Erdogdu, Hera Y. He, Anand Rajaraman, Jure Leskovec |
KDD | 4 |
| 2014 | Recall estimation for rare topic retrieval from large corpusesabstractThe problem of finding documents pertaining to a particular topic finds application in a variety of scenarios. Indeed, the demand for topically pertinent documents has led to myriad companies offering services to find and deliver them (perhaps along with sentiment analysis or clustering) to customers for any topics of interest. The methodologies used to uncover relevant documents range from manually curated keyword filters to trained classification models. Any serious topical analysis requires a sound understanding of key metrics behind the retrieval process, two of the most important being precision and recall. While precision can be easily and inexpensively measured by sampling from classified documents and utilizing (paid) human computation to mark incorrectly classified instances, it is not as straightforward to use the same approach for measuring recall. With most topics occurring relatively sparsely, an unbiased sampling approach becomes prohibitively expensive. In this paper, we introduce a recall measurement procedure requiring only relatively few human judgements. The technique makes use of pairs of sufficiently independent classifiers and the paper provides a detailed discussion of how such classifier pairs can be constructed in practice, with a focus on social media classifiers. We report the performance of the proposed method with simple keyword filters as well as with classifiers of progressive levels of complexity and show that under reasonable conditions, recall can be estimated to within 0.10 absolute error and 15% relative error, and often closer with a reduction of cost by a factor of as much as 1000x as compared with unbiased sampling. Praveen Bommannavar, Alek Kolcz, Anand Rajaraman |
IEEE BigData | 3 |
| 2014 | Anchor-Points Algorithms for Hamming and Edit Distances Using MapReduce
Foto N. Afrati, Anish Das Sarma, Anand Rajaraman, Pokey Rule, Semih Salihoglu, Jeffrey D. Ullman |
ICDT | 3 |
| 2013 | Building, maintaining, and using knowledge bases: a report from the trenchesabstractA knowledge base (KB) contains a set of concepts, instances, and relationships. Over the past decade, numerous KBs have been built, and used to power a growing array of applications. Despite this flurry of activities, however, surprisingly little has been published about the end-to-end process of building, maintaining, and using such KBs in industry. In this paper we describe such a process. In particular, we describe how we build, update, and curate a large KB at Kosmix, a Bay Area startup, and later at WalmartLabs, a development and research lab of Walmart. We discuss how we use this KB to power a range of applications, including query understanding, Deep Web search, in-context advertising, event monitoring in social media, product search, social gifting, and social mining. Finally, we discuss how the KB team is organized, and the lessons learned. Our goal with this paper is to provide a real-world case study, and to contribute to the emerging direction of building, maintaining, and using knowledge bases for data management applications. Omkar Deshpande, Digvijay S. Lamba, Michel Tourn, Sanjib Das, Sri Subramaniam, Anand Rajaraman, Venky Harinarayan, AnHai Doan |
SIGMOD Conference | 6 |
| 2013 | Entity Extraction, Linking, Classification, and Tagging for Social Media: A Wikipedia-Based ApproachabstractMany applications that process social data, such as tweets, must extract entities from tweets (e.g., "Obama" and "Hawaii" in "Obama went to Hawaii"), link them to entities in a knowledge base (e.g., Wikipedia), classify tweets into a set of predefined topics, and assign descriptive tags to tweets. Few solutions exist today to solve these problems for social data, and they are limited in important ways. Further, even though several industrial systems such as OpenCalais have been deployed to solve these problems for text data, little if any has been published about them, and it is unclear if any of the systems has been tailored for social media. In this paper we describe in depth an end-to-end industrial system that solves these problems for social data. The system has been developed and used heavily in the past three years, first at Kosmix, a startup, and later at WalmartLabs. We show how our system uses a Wikipedia-based global "real-time" knowledge base that is well suited for social data, how we interleave the tasks in a synergistic fashion, how we generate and use contexts and social signals to improve task accuracy, and how we scale the system to the entire Twitter firehose. We describe experiments that show that our system outperforms current approaches. Finally we describe applications of the system at Kosmix and WalmartLabs, and lessons learned. Rohit Kumar 0006, Digvijay S. Lamba, Nikesh Garera, Mitul Tiwari, Xiaoyong Chai, Sanjib Das, Sri Subramaniam, Anand Rajaraman, Venky Harinarayan, AnHai Doan |
Proc. VLDB Endow. | 8 |
| 2012 | Anatomy of a gift recommendation engine powered by social mediaabstractMore and more people conduct their shopping online [1], especially during the holiday season [2]. Shopping online offers a lot of convenience, including the luxury of shopping from home, the ease of research, better prices, and in many cases access to unique products not available in stores. Yannis Pavlidis, Madhusudan Mathihalli, Indrani Chakravarty, Arvind Batra, Ron Benson, Ravi Raj, Robert Yau, Mike McKiernan, Venky Harinarayan, Anand Rajaraman |
SIGMOD Conference | 10 |
| 2012 | Muppet: MapReduce-Style Processing of Fast DataabstractMapReduce has emerged as a popular method to process big data. In the past few years, however, not just big data, but fast data has also exploded in volume and availability. Examples of such data include sensor data streams, the Twitter Firehose, and Facebook updates. Numerous applications must process fast data. Can we provide a MapReduce-style framework so that developers can quickly write such applications and execute them over a cluster of machines, to achieve low latency and high scalability? In this paper we report on our investigation of this question, as carried out at Kosmix and WalmartLabs. We describe MapUpdate, a framework like MapReduce, but specifically developed for fast data. We describe Muppet, our implementation of MapUpdate. Throughout the description we highlight the key challenges, argue why MapReduce is not well suited to address them, and briefly describe our current solutions. Finally, we describe our experience and lessons learned with Muppet, which has been used extensively at Kosmix and WalmartLabs to power a broad range of applications in social media and e-commerce. Wang Lam, STS Prasad, Anand Rajaraman, Zoheb Vacheri, AnHai Doan |
Proc. VLDB Endow. | 4 |
| 2010 | Towards The Web of Concepts: Extracting Concepts from Large DatasetsabstractConcepts are sequences of words that represent real or imaginary entities or ideas that users are interested in. As a first step towards building a web of concepts that will form the backbone of the next generation of search technology, we develop a novel technique to extract concepts from large datasets. We approach the problem of concept extraction from corpora as a market-basket problem, adapting statistical measures of support and confidence. We evaluate our concept extraction algorithm on datasets containing data from a large number of users (e.g., the AOL query log data set), and we show that a high-precision concept set can be extracted. Aditya G. Parameswaran, Hector Garcia-Molina, Anand Rajaraman |
Proc. VLDB Endow. | 3 |
| 2009 | Answering Web Questions Using Structured Data - Dream or Reality?abstractThe question of which role structured data can play in Web search has been raised from the early days of the Web. On the one hand, structured data can be used to answer factual queries. On the other, large amounts of structured data can be used to better organize web-content and therefore to improve search on a wide range of queries. Anand Rajaraman, Sunita Sarawagi, William Tunstall-Pedoe, Gerhard Weikum, Alon Y. Halevy |
Proc. VLDB Endow. | 2 |
| 2009 | Kosmix: High-Performance Topic Exploration using the Deep WebabstractKosmix lies at the intersection of two important trends: topic exploration and the Deep Web. Topic exploration is a new approach to information discovery on the web that satisfies certain use cases not served well by conventional web search. The Deep Web, an inhospitable region for web crawlers, is emerging as a significant information resource. We describe the anatomy of Kosmix, the first general-purpose topic exploration engine to harness the Deep Web using a federated search approach. We focus in particular on the Kosmix approach to query tranformation and caching, which is essential to ensure reasonable performance. Anand Rajaraman |
Proc. VLDB Endow. | 1 |
| 2006 | Data Integration: The Teenage Years
Alon Y. Halevy, Anand Rajaraman, Joann J. Ordille |
VLDB | 2 |
| 2003 | Querying websites using compact skeletons
Anand Rajaraman, Jeffrey D. Ullman |
J. Comput. Syst. Sci. | 1 |
| 2001 | Querying Websites Using Compact SkeletonsabstractSeveral commercial applications, such as online comparison shopping and process automation, require integrating information that is scattered across multiple websites or XML documents. Much research has been devoted to this problem, resulting in several research prototypes and commercial implementations. Such systems rely on wrappers that provide relational or other structured interfaces to websites. Traditionally, wrappers have been constructed by hand on a per-website basis, constraining the scalability of the system. Anand Rajaraman, Jeffrey D. Ullman |
PODS | 1 |
| 2000 | Conjunctive query containment revisited
Chandra Chekuri, Anand Rajaraman |
Theor. Comput. Sci. | 2 |
| 1999 | E-Commerce Database Issues and ExperienceabstractNo abstract available. Anand Rajaraman |
SIGMOD Conference | 1 |
| 1999 | Answering Queries Using Limited External Query Processors
Alon Y. Halevy, Anand Rajaraman, Jeffrey D. Ullman |
J. Comput. Syst. Sci. | 2 |
| 1998 | Virtual Database TechnologyabstractVirtual database (VDB) technology makes external data behave as an extension of an enterprise's relational database (RDBMS) system. VDB technology enables the rapid deployment of applications with at least one of the following characteristics: large numbers of data sources; data sources that are autonomous (i.e. there is no centralized control); or data sources that can have a mixture of structured and unstructured data. The World Wide Web and most intranets have all of these characteristics and can thus benefit from VDB technology. Ashish Gupta 0001, Venky Harinarayan, Anand Rajaraman |
ICDE | 3 |
| 1997 | Index Selection for OLAPabstractOn-line analytical processing (OLAP) is a recent and important application of database systems. Typically, OLAP data is presented as a multidimensional "data cube." OLAP queries are complex and can take many hours or even days to run, if executed directly on the raw data. The most common method of reducing execution time is to precompute some of the queries into summary tables (subcubes of the data cube) and then to build indexes on these summary tables. In most commercial OLAP systems today, the summary tables that are to be precomputed are picked first, followed by the selection of the appropriate indexes on them. A trial-and-error approach is used to divide the space available between the summary tables and the indexes. This two-step process can perform very poorly. Since both summary tables and indexes consume the same resource-space-their selection should be done together for the most efficient use of space. The authors give algorithms that automate the selection of summary tables and indexes. In particular, they present a family of algorithms of increasing time complexities, and prove strong performance bounds for them. The algorithms with higher complexities have better performance bounds. However, the increase in the performance bound is diminishing, and they show that an algorithm of moderate complexity can perform fairly close to the optimal. Himanshu Gupta 0001, Venky Harinarayan, Anand Rajaraman, Jeffrey D. Ullman |
ICDE | 3 |
| 1997 | Conjunctive Query Containment Revisited
Chandra Chekuri, Anand Rajaraman |
ICDT | 2 |
| 1997 | The TSIMMIS Approach to Mediation: Data Models and Languages
Hector Garcia-Molina, Yannis Papakonstantinou, Dallan Quass, Anand Rajaraman, Yehoshua Sagiv, Jeffrey D. Ullman, Vasilis Vassalos, Jennifer Widom |
J. Intell. Inf. Syst. | 4 |
| 1996 | Answering Queries Using Limited External ProcessorsabstractWhen answering queries using external information sources, their contents can be described by views.To answer a query, we must rewrite it using the set of views presented by the Alon Y. Halevy, Anand Rajaraman, Jeffrey D. Ullman |
PODS | 2 |
| 1996 | Integrating Information by Outerjoins and Full DisjunctionsabstractOur motivationis the piecing together of tidbits of information found on the "web" into a usable information structure.The problem is related to that of computing the natural outerjoin of many relations in a way that preserves all possible connections among facts. Anand Rajaraman, Jeffrey D. Ullman |
PODS | 1 |
| 1996 | Change Detection in Hierarchically Structured InformationabstractDetecting and representing changes to data is important for active databases, data warehousing, view maintenance, and version and configuration management. Most previous work in change management has dealt with flat-file and relational data; we focus on hierarchically structured data. Since in many cases changes must be computed from old and new versions of the data, we define the hierarchical change detection problem as the problem of finding a "minimum-cost edit script" that transforms one data tree to another, and we present efficient algorithms for computing such an edit script. Our algorithms make use of some key domain characteristics to achieve substantially better performance than previous, generalpurpose algorithms. We study the performance of our algorithms both analytically and empirically, and we describe the application of our techniques to hierarchically structured documents. 1 Introduction We study the problem of detecting and representing changes to hierarchically stru... Sudarshan S. Chawathe, Anand Rajaraman, Hector Garcia-Molina, Jennifer Widom |
SIGMOD Conference | 2 |
| 1996 | Implementing Data Cubes EfficientlyabstractDecision support applications involve complex queries on very large databases. Since response times should be small, query optimization is critical. Users typically view the data as multidimensional data cubes. Each cell of the data cube is a view consisting of an aggregation of interest, like total sales. The values of many of these cells are dependent on the values of other cells in the data cube..A common and powerful query optimization technique is to materialize some or all of these cells rather than compute them from raw data each time. Commercial systems differ mainly in their approach to materializing the data cube. In this paper, we investigate the issue of which cells (views) to materialize when it is too expensive to materialize all views. A lattice framework is used to express dependencies among views. We present greedy algorithms that work off this lattice and determine a good set of views to materialize. The greedy algorithm performs within a small constant factor of optimal under a variety of models. We then consider the most common case of the hypercube lattice and examine the choice of materialized views for hypercubes in detail, giving some good tradeoffs between the space used and the average time to answer a query. 1 Venky Harinarayan, Anand Rajaraman, Jeffrey D. Ullman |
SIGMOD Conference | 2 |
| 1996 | LORE: A Lightweight Object REpository for Semistructured DataabstractNo abstract available. Dallan Quass, Jennifer Widom, Roy Goldman, Kevin Haas, Qingshan Luo, Jason McHugh, Svetlozar Nestorov, Anand Rajaraman, Hugo Rivero, Serge Abiteboul, Jeffrey D. Ullman, Janet L. Wiener |
SIGMOD Conference | 8 |
| 1996 | Querying Heterogeneous Information Sources Using Source Descriptions
Alon Y. Halevy, Anand Rajaraman, Joann J. Ordille |
VLDB | 2 |
| 1995 | Answering Queries Using Templates with Binding Patterns
Anand Rajaraman, Yehoshua Sagiv, Jeffrey D. Ullman |
PODS | 1 |
| 1993 | Connected Domination and Steiner Set on Asteroidal Triple-Free Graphs
Hari Balakrishnan, Anand Rajaraman, C. Pandu Rangan |
WADS | 2 |