Anand Rajaraman

dblp:r/ARajaraman · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
0since 2021 · last 2016
0000-0003-4484-6721ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 25 · 6 first-authorTheory of computation · 4 · 1 first-authorArtificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
20 papers
Information retrieval · 26% Web and social media mining · 18% Knowledge graphs · 14%
Artificial intelligence
4 papers
Probabilistic and Bayesian machine learning · 35% Information extraction and text analysis · 26% Question answering and dialogue systems · 15%

Topics — the 30 heaviest of 46, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process
0.212015
SEISMIC: A Self-Exciting Point Process Model for Predicting Tweet Popularity · KDD 2015
Web and social media mining › information diffusion
information diffusion prediction
0.212015
SEISMIC: A Self-Exciting Point Process Model for Predicting Tweet Popularity · KDD 2015
Knowledge graphs › knowledge graph management
knowledge base curation
0.212013
Building, maintaining, and using knowledge bases: a report from the trenches · SIGMOD Conference 2013
Information retrieval
query understanding
0.212013
Building, maintaining, and using knowledge bases: a report from the trenches · SIGMOD Conference 2013
Recommender systems › e-commerce recommendation
gift recommendation
0.112012
Anatomy of a gift recommendation engine powered by social media · SIGMOD Conference 2012
Data mining › text mining › information extraction
concept mining
0.112010
Towards The Web of Concepts: Extracting Concepts from Large Datasets · Proc. VLDB Endow. 2010
Data mining › pattern mining › association rule mining
market basket analysis
0.112010
Towards The Web of Concepts: Extracting Concepts from Large Datasets · Proc. VLDB Endow. 2010
Natural language and speech › Question answering and dialogue systems
question answering over structured data
0.112009
Answering Web Questions Using Structured Data - Dream or Reality? · Proc. VLDB Endow. 2009
Information retrieval › distributed information retrieval
federated search
0.112009
Kosmix: High-Performance Topic Exploration using the Deep Web · Proc. VLDB Endow. 2009
Information retrieval › web search › web information retrieval
hidden web
0.112009
Kosmix: High-Performance Topic Exploration using the Deep Web · Proc. VLDB Endow. 2009
Query processing and optimization › query rewriting
query transformation
0.112009
Kosmix: High-Performance Topic Exploration using the Deep Web · Proc. VLDB Endow. 2009
Information retrieval › interactive information retrieval › exploratory search
topic exploration
0.112009
Kosmix: High-Performance Topic Exploration using the Deep Web · Proc. VLDB Endow. 2009
Information retrieval
web search
0.112009
Answering Web Questions Using Structured Data - Dream or Reality? · Proc. VLDB Endow. 2009
Computer vision › Image recognition and object detection
visual recognition
0.112016
Data-Driven Disruption: The View from Silicon Valley · Proc. VLDB Endow. 2016
Cloud and datacenter computing
big data platform
0.112016
Data-Driven Disruption: The View from Silicon Valley · Proc. VLDB Endow. 2016
Information retrieval › web search › web information retrieval
social media retrieval
0.012013
Entity Extraction, Linking, Classification, and Tagging for Social Media: A Wikipedia-Based Approach · Proc. VLDB Endow. 2013
Distributed and cloud data management
mapreduce
0.012012
Muppet: MapReduce-Style Processing of Fast Data · Proc. VLDB Endow. 2012
Web and social media mining
social media analysis
0.012012
Anatomy of a gift recommendation engine powered by social media · SIGMOD Conference 2012
Information retrieval
query log analysis
0.012010
Towards The Web of Concepts: Extracting Concepts from Large Datasets · Proc. VLDB Endow. 2010
Data integration and cleaning › data extraction
web data extraction
0.012001
Querying Websites Using Compact Skeletons · PODS 2001
Data integration and cleaning › data extraction › web data extraction
wrapper generation
0.012001
Querying Websites Using Compact Skeletons · PODS 2001
Query processing and optimization › query rewriting
query answering using views
0.021996
Answering Queries Using Limited External Processors · PODS 1996
Answering Queries Using Templates with Binding Patterns · PODS 1995
Indexing and storage engines
caching
0.012009
Kosmix: High-Performance Topic Exploration using the Deep Web · Proc. VLDB Endow. 2009
Data integration and cleaning
heterogeneous data sources
0.021998
Querying Heterogeneous Information Sources Using Source Descriptions · VLDB 1996
Virtual Database Technology · ICDE 1998
Distributed and cloud data management
federated database
0.011998
Virtual Database Technology · ICDE 1998
Data integration and cleaning › data integration system
virtual database
0.011998
Virtual Database Technology · ICDE 1998
Database system architecture and tuning › database design › physical database design
index selection
0.011997
Index Selection for OLAP · ICDE 1997
Query processing and optimization › OLAP
OLAP query optimization
0.011997
Index Selection for OLAP · ICDE 1997
Data mining › time series analysis
change point detection
0.011996
Change Detection in Hierarchically Structured Information · SIGMOD Conference 1996
Query processing and optimization › OLAP
data cube
0.011996
Implementing Data Cubes Efficiently · SIGMOD Conference 1996

Methods — techniques the papers use, named apart from their topics

self-exciting point process · 0.4knowledge base linking · 0.3context and social signals · 0.3case study · 0.2social media signal extraction · 0.1mapupdate · 0.1statistical measures of support and confidence · 0.1federated search · 0.1greedy algorithm · 0.0skeleton-based extraction · 0.0
YearPublicationVenuePosition
2016 Data-Driven Disruption: The View from Silicon Valley
abstract
We live in an era where software is transforming industries, the sciences, and society as a whole. This exciting phenomenon has been described by the phrase "software is eating the world." It is becoming increasingly apparent that data is the fuel powering software's conquests. Data is the new disruptor. It's hard to believe that the first decade of the Big Data era is already behind us. Silicon Valley has been at the forefront of developing and applying data-driven approaches to create disruption at many levels: infrastructure (e.g., Hadoop and Spark), capabilities (e.g., image recognition and machine translation), and killer apps (e.g., self-driving cars and messaging bots). In this talk, we first look back on the past decade and share learnings from the frontlines of data-driven disruption. Looking ahead, we then describe challenges and opportunities for the next decade. Since this has also been a personal journey, we will use examples drawn from personal experience to illustrate each point.
Anand Rajaraman
Proc. VLDB Endow.1
2015 SEISMIC: A Self-Exciting Point Process Model for Predicting Tweet Popularity
abstract
Social networking websites allow users to create and share content. Big information cascades of post resharing can form as users of these sites reshare others' posts with their friends and followers. One of the central challenges in understanding such cascading behaviors is in forecasting information outbreaks, where a single post becomes widely popular by being reshared by many users. In this paper, we focus on predicting the final number of reshares of a given post. We build on the theory of self-exciting point processes to develop a statistical model that allows us to make accurate predictions. Our model requires no training or expensive feature engineering. It results in a simple and efficiently computable formula that allows us to answer questions, in real-time, such as: Given a post's resharing history so far, what is our current estimate of its final number of reshares? Is the post resharing cascade past the initial stage of explosive growth? And, which posts will be the most reshared in the future?
Qingyuan Zhao, Murat A. Erdogdu, Hera Y. He, Anand Rajaraman, Jure Leskovec
KDD4
2014 Recall estimation for rare topic retrieval from large corpuses
abstract
The problem of finding documents pertaining to a particular topic finds application in a variety of scenarios. Indeed, the demand for topically pertinent documents has led to myriad companies offering services to find and deliver them (perhaps along with sentiment analysis or clustering) to customers for any topics of interest. The methodologies used to uncover relevant documents range from manually curated keyword filters to trained classification models. Any serious topical analysis requires a sound understanding of key metrics behind the retrieval process, two of the most important being precision and recall. While precision can be easily and inexpensively measured by sampling from classified documents and utilizing (paid) human computation to mark incorrectly classified instances, it is not as straightforward to use the same approach for measuring recall. With most topics occurring relatively sparsely, an unbiased sampling approach becomes prohibitively expensive. In this paper, we introduce a recall measurement procedure requiring only relatively few human judgements. The technique makes use of pairs of sufficiently independent classifiers and the paper provides a detailed discussion of how such classifier pairs can be constructed in practice, with a focus on social media classifiers. We report the performance of the proposed method with simple keyword filters as well as with classifiers of progressive levels of complexity and show that under reasonable conditions, recall can be estimated to within 0.10 absolute error and 15% relative error, and often closer with a reduction of cost by a factor of as much as 1000x as compared with unbiased sampling.
Praveen Bommannavar, Alek Kolcz, Anand Rajaraman
IEEE BigData3
2014 Anchor-Points Algorithms for Hamming and Edit Distances Using MapReduce
Foto N. Afrati, Anish Das Sarma, Anand Rajaraman, Pokey Rule, Semih Salihoglu, Jeffrey D. Ullman
ICDT3
2013 Building, maintaining, and using knowledge bases: a report from the trenches
abstract
A knowledge base (KB) contains a set of concepts, instances, and relationships. Over the past decade, numerous KBs have been built, and used to power a growing array of applications. Despite this flurry of activities, however, surprisingly little has been published about the end-to-end process of building, maintaining, and using such KBs in industry. In this paper we describe such a process. In particular, we describe how we build, update, and curate a large KB at Kosmix, a Bay Area startup, and later at WalmartLabs, a development and research lab of Walmart. We discuss how we use this KB to power a range of applications, including query understanding, Deep Web search, in-context advertising, event monitoring in social media, product search, social gifting, and social mining. Finally, we discuss how the KB team is organized, and the lessons learned. Our goal with this paper is to provide a real-world case study, and to contribute to the emerging direction of building, maintaining, and using knowledge bases for data management applications.
Omkar Deshpande, Digvijay S. Lamba, Michel Tourn, Sanjib Das, Sri Subramaniam, Anand Rajaraman, Venky Harinarayan, AnHai Doan
SIGMOD Conference6
2013 Entity Extraction, Linking, Classification, and Tagging for Social Media: A Wikipedia-Based Approach
abstract
Many applications that process social data, such as tweets, must extract entities from tweets (e.g., "Obama" and "Hawaii" in "Obama went to Hawaii"), link them to entities in a knowledge base (e.g., Wikipedia), classify tweets into a set of predefined topics, and assign descriptive tags to tweets. Few solutions exist today to solve these problems for social data, and they are limited in important ways. Further, even though several industrial systems such as OpenCalais have been deployed to solve these problems for text data, little if any has been published about them, and it is unclear if any of the systems has been tailored for social media. In this paper we describe in depth an end-to-end industrial system that solves these problems for social data. The system has been developed and used heavily in the past three years, first at Kosmix, a startup, and later at WalmartLabs. We show how our system uses a Wikipedia-based global "real-time" knowledge base that is well suited for social data, how we interleave the tasks in a synergistic fashion, how we generate and use contexts and social signals to improve task accuracy, and how we scale the system to the entire Twitter firehose. We describe experiments that show that our system outperforms current approaches. Finally we describe applications of the system at Kosmix and WalmartLabs, and lessons learned.
Rohit Kumar 0006, Digvijay S. Lamba, Nikesh Garera, Mitul Tiwari, Xiaoyong Chai, Sanjib Das, Sri Subramaniam, Anand Rajaraman, Venky Harinarayan, AnHai Doan
Proc. VLDB Endow.8
2012 Anatomy of a gift recommendation engine powered by social media
abstract
More and more people conduct their shopping online [1], especially during the holiday season [2]. Shopping online offers a lot of convenience, including the luxury of shopping from home, the ease of research, better prices, and in many cases access to unique products not available in stores.
Yannis Pavlidis, Madhusudan Mathihalli, Indrani Chakravarty, Arvind Batra, Ron Benson, Ravi Raj, Robert Yau, Mike McKiernan, Venky Harinarayan, Anand Rajaraman
SIGMOD Conference10
2012 Muppet: MapReduce-Style Processing of Fast Data
abstract
MapReduce has emerged as a popular method to process big data. In the past few years, however, not just big data, but fast data has also exploded in volume and availability. Examples of such data include sensor data streams, the Twitter Firehose, and Facebook updates. Numerous applications must process fast data. Can we provide a MapReduce-style framework so that developers can quickly write such applications and execute them over a cluster of machines, to achieve low latency and high scalability? In this paper we report on our investigation of this question, as carried out at Kosmix and WalmartLabs. We describe MapUpdate, a framework like MapReduce, but specifically developed for fast data. We describe Muppet, our implementation of MapUpdate. Throughout the description we highlight the key challenges, argue why MapReduce is not well suited to address them, and briefly describe our current solutions. Finally, we describe our experience and lessons learned with Muppet, which has been used extensively at Kosmix and WalmartLabs to power a broad range of applications in social media and e-commerce.
Wang Lam, STS Prasad, Anand Rajaraman, Zoheb Vacheri, AnHai Doan
Proc. VLDB Endow.4
2010 Towards The Web of Concepts: Extracting Concepts from Large Datasets
abstract
Concepts are sequences of words that represent real or imaginary entities or ideas that users are interested in. As a first step towards building a web of concepts that will form the backbone of the next generation of search technology, we develop a novel technique to extract concepts from large datasets. We approach the problem of concept extraction from corpora as a market-basket problem, adapting statistical measures of support and confidence. We evaluate our concept extraction algorithm on datasets containing data from a large number of users (e.g., the AOL query log data set), and we show that a high-precision concept set can be extracted.
Aditya G. Parameswaran, Hector Garcia-Molina, Anand Rajaraman
Proc. VLDB Endow.3
2009 Answering Web Questions Using Structured Data - Dream or Reality?
abstract
The question of which role structured data can play in Web search has been raised from the early days of the Web. On the one hand, structured data can be used to answer factual queries. On the other, large amounts of structured data can be used to better organize web-content and therefore to improve search on a wide range of queries.
Anand Rajaraman, Sunita Sarawagi, William Tunstall-Pedoe, Gerhard Weikum, Alon Y. Halevy
Proc. VLDB Endow.2
2009 Kosmix: High-Performance Topic Exploration using the Deep Web
abstract
Kosmix lies at the intersection of two important trends: topic exploration and the Deep Web. Topic exploration is a new approach to information discovery on the web that satisfies certain use cases not served well by conventional web search. The Deep Web, an inhospitable region for web crawlers, is emerging as a significant information resource. We describe the anatomy of Kosmix, the first general-purpose topic exploration engine to harness the Deep Web using a federated search approach. We focus in particular on the Kosmix approach to query tranformation and caching, which is essential to ensure reasonable performance.
Anand Rajaraman
Proc. VLDB Endow.1
2006 Data Integration: The Teenage Years
Alon Y. Halevy, Anand Rajaraman, Joann J. Ordille
VLDB2
2003 Querying websites using compact skeletons
Anand Rajaraman, Jeffrey D. Ullman
J. Comput. Syst. Sci.1
2001 Querying Websites Using Compact Skeletons
abstract
Several commercial applications, such as online comparison shopping and process automation, require integrating information that is scattered across multiple websites or XML documents. Much research has been devoted to this problem, resulting in several research prototypes and commercial implementations. Such systems rely on wrappers that provide relational or other structured interfaces to websites. Traditionally, wrappers have been constructed by hand on a per-website basis, constraining the scalability of the system.
Anand Rajaraman, Jeffrey D. Ullman
PODS1
2000 Conjunctive query containment revisited
Chandra Chekuri, Anand Rajaraman
Theor. Comput. Sci.2
1999 E-Commerce Database Issues and Experience
abstract
No abstract available.
Anand Rajaraman
SIGMOD Conference1
1999 Answering Queries Using Limited External Query Processors
Alon Y. Halevy, Anand Rajaraman, Jeffrey D. Ullman
J. Comput. Syst. Sci.2
1998 Virtual Database Technology
abstract
Virtual database (VDB) technology makes external data behave as an extension of an enterprise's relational database (RDBMS) system. VDB technology enables the rapid deployment of applications with at least one of the following characteristics: large numbers of data sources; data sources that are autonomous (i.e. there is no centralized control); or data sources that can have a mixture of structured and unstructured data. The World Wide Web and most intranets have all of these characteristics and can thus benefit from VDB technology.
Ashish Gupta 0001, Venky Harinarayan, Anand Rajaraman
ICDE3
1997 Index Selection for OLAP
abstract
On-line analytical processing (OLAP) is a recent and important application of database systems. Typically, OLAP data is presented as a multidimensional "data cube." OLAP queries are complex and can take many hours or even days to run, if executed directly on the raw data. The most common method of reducing execution time is to precompute some of the queries into summary tables (subcubes of the data cube) and then to build indexes on these summary tables. In most commercial OLAP systems today, the summary tables that are to be precomputed are picked first, followed by the selection of the appropriate indexes on them. A trial-and-error approach is used to divide the space available between the summary tables and the indexes. This two-step process can perform very poorly. Since both summary tables and indexes consume the same resource-space-their selection should be done together for the most efficient use of space. The authors give algorithms that automate the selection of summary tables and indexes. In particular, they present a family of algorithms of increasing time complexities, and prove strong performance bounds for them. The algorithms with higher complexities have better performance bounds. However, the increase in the performance bound is diminishing, and they show that an algorithm of moderate complexity can perform fairly close to the optimal.
Himanshu Gupta 0001, Venky Harinarayan, Anand Rajaraman, Jeffrey D. Ullman
ICDE3
1997 Conjunctive Query Containment Revisited
Chandra Chekuri, Anand Rajaraman
ICDT2
1997 The TSIMMIS Approach to Mediation: Data Models and Languages
Hector Garcia-Molina, Yannis Papakonstantinou, Dallan Quass, Anand Rajaraman, Yehoshua Sagiv, Jeffrey D. Ullman, Vasilis Vassalos, Jennifer Widom
J. Intell. Inf. Syst.4
1996 Answering Queries Using Limited External Processors
abstract
When answering queries using external information sources, their contents can be described by views.To answer a query, we must rewrite it using the set of views presented by the
Alon Y. Halevy, Anand Rajaraman, Jeffrey D. Ullman
PODS2
1996 Integrating Information by Outerjoins and Full Disjunctions
abstract
Our motivationis the piecing together of tidbits of information found on the "web" into a usable information structure.The problem is related to that of computing the natural outerjoin of many relations in a way that preserves all possible connections among facts.
Anand Rajaraman, Jeffrey D. Ullman
PODS1
1996 Change Detection in Hierarchically Structured Information
abstract
Detecting and representing changes to data is important for active databases, data warehousing, view maintenance, and version and configuration management. Most previous work in change management has dealt with flat-file and relational data; we focus on hierarchically structured data. Since in many cases changes must be computed from old and new versions of the data, we define the hierarchical change detection problem as the problem of finding a "minimum-cost edit script" that transforms one data tree to another, and we present efficient algorithms for computing such an edit script. Our algorithms make use of some key domain characteristics to achieve substantially better performance than previous, generalpurpose algorithms. We study the performance of our algorithms both analytically and empirically, and we describe the application of our techniques to hierarchically structured documents. 1 Introduction We study the problem of detecting and representing changes to hierarchically stru...
Sudarshan S. Chawathe, Anand Rajaraman, Hector Garcia-Molina, Jennifer Widom
SIGMOD Conference2
1996 Implementing Data Cubes Efficiently
abstract
Decision support applications involve complex queries on very large databases. Since response times should be small, query optimization is critical. Users typically view the data as multidimensional data cubes. Each cell of the data cube is a view consisting of an aggregation of interest, like total sales. The values of many of these cells are dependent on the values of other cells in the data cube..A common and powerful query optimization technique is to materialize some or all of these cells rather than compute them from raw data each time. Commercial systems differ mainly in their approach to materializing the data cube. In this paper, we investigate the issue of which cells (views) to materialize when it is too expensive to materialize all views. A lattice framework is used to express dependencies among views. We present greedy algorithms that work off this lattice and determine a good set of views to materialize. The greedy algorithm performs within a small constant factor of optimal under a variety of models. We then consider the most common case of the hypercube lattice and examine the choice of materialized views for hypercubes in detail, giving some good tradeoffs between the space used and the average time to answer a query. 1
Venky Harinarayan, Anand Rajaraman, Jeffrey D. Ullman
SIGMOD Conference2
1996 LORE: A Lightweight Object REpository for Semistructured Data
abstract
No abstract available.
Dallan Quass, Jennifer Widom, Roy Goldman, Kevin Haas, Qingshan Luo, Jason McHugh, Svetlozar Nestorov, Anand Rajaraman, Hugo Rivero, Serge Abiteboul, Jeffrey D. Ullman, Janet L. Wiener
SIGMOD Conference8
1996 Querying Heterogeneous Information Sources Using Source Descriptions
Alon Y. Halevy, Anand Rajaraman, Joann J. Ordille
VLDB2
1995 Answering Queries Using Templates with Binding Patterns
Anand Rajaraman, Yehoshua Sagiv, Jeffrey D. Ullman
PODS1
1993 Connected Domination and Steiner Set on Asteroidal Triple-Free Graphs
Hari Balakrishnan, Anand Rajaraman, C. Pandu Rangan
WADS2