Alpa Jain

dblp:68/6659 · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 8 first-authorArtificial intelligence and machine learning · 9 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Query processing and optimization · 42% Data mining · 20% Information retrieval · 16%
Artificial intelligence
7 papers
Information extraction and text analysis · 70% Knowledge representation and reasoning · 30%
Software engineering, system software, and programming languages
1 paper
Debugging and program repair · 100%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
fact extraction
0.122006
Names and Similarities on the Web: Fact Extraction in the Fast Lane · ACL 2006
Organizing and Searching the World Wide Web of Facts - Step One: The One-Million Fact Extraction Challenge · AAAI 2006
Natural language and speech › Information extraction and text analysis
relation extraction
0.112011
Dynamic relationship and event discovery · WSDM 2011
Information retrieval
query suggestion
0.112011
Synthesizing high utility suggestions for rare web search queries · SIGIR 2011
Data mining › temporal data mining
temporal event detection
0.112011
Dynamic relationship and event discovery · WSDM 2011
Natural language and speech › Information extraction and text analysis
pattern learning
0.112010
I4E: interactive investigation of iterative information extraction · SIGMOD Conference 2010
Debugging and program repair
fault localization
0.112010
I4E: interactive investigation of iterative information extraction · SIGMOD Conference 2010
Query processing and optimization
join processing
0.112009
Join Optimization of Information Extraction Output: Quality Matters! · ICDE 2009
Query processing and optimization
query quality optimization
0.112009
A quality-aware optimizer for information extraction · ACM Trans. Database Syst. 2009
Query processing and optimization
top-k query processing
0.112009
Exploring a Few Good Tuples from Text Databases · ICDE 2009
Query processing and optimization › query optimization
cost-based optimization
0.112008
Optimizing SQL Queries over Text Databases · ICDE 2008
Natural language and speech › Information extraction and text analysis
open information extraction
0.112006
Organizing and Searching the World Wide Web of Facts - Step One: The One-Million Fact Extraction Challenge · AAAI 2006
Data mining › text mining
fact extraction
0.112006
Names and Similarities on the Web: Fact Extraction in the Fast Lane · ACL 2006
Web and social media mining › web mining
web text mining
0.112006
Names and Similarities on the Web: Fact Extraction in the Fast Lane · ACL 2006
Information retrieval
query logs
0.012011
Synthesizing high utility suggestions for rare web search queries · SIGIR 2011
Graph algorithms and graph theory › graph theory
dynamic graphs
0.012011
Dynamic relationship and event discovery · WSDM 2011
Visualization and visual analytics
interactive data analysis
0.012010
I4E: interactive investigation of iterative information extraction · SIGMOD Conference 2010
Data mining › text mining
information extraction
0.012009
A quality-aware optimizer for information extraction · ACM Trans. Database Syst. 2009
Data models and query languages › query language
query language semantics
0.012009
Exploring a Few Good Tuples from Text Databases · ICDE 2009
Query processing and optimization › SQL query processing
select-project-join query
0.012008
Optimizing SQL Queries over Text Databases · ICDE 2008
Information retrieval
web search
0.012006
Organizing and Searching the World Wide Web of Facts - Step One: The One-Million Fact Extraction Challenge · AAAI 2006

Methods — techniques the papers use, named apart from their topics

pattern-based extraction · 0.6temporal constraint clustering · 0.4information extraction · 0.3quality-aware join optimization · 0.2analytical cost model · 0.2textual source combination · 0.1query-level operations · 0.1ranked retrieval · 0.1maximum likelihood estimation · 0.1confidence scoring · 0.1cost-based optimization · 0.1distributional similarity · 0.1
YearPublicationVenuePosition
2013 Building Rich User Search Queries Profiles
Elif Aktolga, Alpa Jain, Emre Velipasaoglu
UMAP2
2011 Assisting web search users by destination reachability
abstract
Search engine users are increasingly performing complex tasks based on the simple keyword-in document-out paradigm. To assist users in accomplishing their tasks effectively, search engines provide query recommendations based on the user's current query. These are suggestions for follow-up queries given the user-provided query. A large number of techniques have been proposed in the past on mining such query recommendations which include past user sessions (e.g., sequence of queries within a specified window of time) to identify most frequently occurring pairs, using click-through graphs (e.g., a bipartite graph of queries and the urls on which users clicked) and rank these suggestions using some form of frequency counts from the past query logs. Given the limited number of queries that are offered (typically 5) it is important to effectively rank them. In this paper, we present a novel approach to ranking query recommendations which not only consider relevance to the original query but also take into account efficiency of a query at accomplishing a user search task at hand. We formalize the notion of query efficiency and show how our objective function effectively captures this as determined by a human study and eliminates biases introduced by click-through based metrics. To compute this objective function, we present a pseudosupervised learning technique where no explicit human experts are required to label samples. In addition, our techniques effectively characterize preferred url destinations and project each query into a higher dimension space where each sub-spaces represents user intent using these characteristics. Finally, we present an extensive evaluation of our proposed methods against production systems and show our method to increase task completion efficiency by 15%.
Alpa Jain, Larry Lai
CIKM2
2011 Building a generic debugger for information extraction pipelines
abstract
Complex information extraction (IE) pipelines are becoming an integral component of most text processing frameworks. We introduce a first system to help IE users analyze extraction pipeline semantics and operator transformations interactively while debugging. This allows the effort to be proportional to the need, and to focus on the portions of the pipeline under the greatest suspicion. We present a generic debugger for running post-execution analysis of any IE pipeline consisting of arbitrary types of operators. For this, we propose an effective provenance model for IE pipelines which captures a variety of operator types, ranging from those for which full to no specifications are available. We have evaluated our proposed algorithms and provenance model on large-scale real-world extraction pipelines.
Anish Das Sarma, Alpa Jain, Philip Bohannon
CIKM2
2011 Synthesizing high utility suggestions for rare web search queries
abstract
Search engines are continuously looking into methods to alleviate users' effort in finding desired information. For this, all major search engines employ query suggestions methods to facilitate effective query formulation and reformulation. Providing high quality query suggestions is a critical task for search engines and so far most research efforts have focused on tapping various information available in search query logs to identify potential suggestions. By relying on this single source of information, suggestion providing systems often restrict themselves to only previously observed query sessions. Therefore, a critical challenge faced by query suggestions provision mechanism is that of coverage, i.e., the number of unique queries for which users are provided with suggestions, while keeping the suggestion quality high. To address this problem, we propose a novel way of generating suggestions for user search queries by moving beyond the dependency on search query logs and providing synthetic suggestions for web search queries. The key challenges in providing synthetic suggestions include identifying important concepts in a query and systematically exploring related concepts while ensuring that the resulting suggestions are relevant to the user query and of high utility. We present an end-to-end system to generate synthetic suggestions that builds upon novel query-level operations and combines information available from various textual sources. We evaluate our suggestion system over a large-scale real-world dataset of query logs and show that our methods increase the coverage of query-suggestion pairs by up to 39% without compromising the quality or the utility of the suggestions.
Alpa Jain, Umut Ozertem, Emre Velipasaoglu
SIGIR1
2011 Dynamic relationship and event discovery
abstract
This paper studies the problem of dynamic relationship and event discovery. A large body of previous work on relation extraction focuses on discovering predefined and static relationships between entities. In contrast, we aim to identify temporally defined (e.g., co-bursting) relationships that are not predefined by an existing schema, and we identify the underlying time constrained events that lead to these relationships. The key challenges in identifying such events include discovering and verifying dynamic connections among entities, and consolidating binary dynamic connections into events consisting of a set of entities that are connected at a given time period. We formalize this problem and introduce an efficient end-to-end pipeline as a solution. In particular, we introduce two formal notions, global temporal constraint cluster and local temporal constraint cluster, for detecting dynamic events. We further design efficient algorithms for discovering such events from a large graph of dynamic relationships. Finally, detailed experiments on real data show the
Anish Das Sarma, Alpa Jain, Cong Yu 0001
WSDM2
2010 Organizing query completions for web search
abstract
All state-of-the-art web search engines implement an auto-completion mechanism - an assistive technology enabling users to effectively formulate their search queries by predicting the next characters or words that they are likely to type. Query completions (or suggestions) are typically mined from past user interactions with the search engine, e.g., from query logs, clickthrough patterns, or query reformulations; they are ranked by some measure of query popularity, e.g., query frequency or clickthrough rate. Current query suggestion tools largely assume that the set of suggestions provided to the users is homogeneous, corresponding to a single real-world interpretation of the query. In this paper, we hypothesize that, in some cases, users would benefit from an alternative presentation of the suggestions, one where suggestions are not only ordered by likelihood but also organized by high-level user intent. Rich search suggestion interaction frameworks that reduce the user effort in identifying the set of relevant suggestions open new and promising directions towards improving user experience. Along these lines, we propose clustering the set of suggestions presented to a search engine user, and assigning an appropriate label to each subset of suggestions to help users quickly identify useful ones. For this, we present a variety of unsupervised clustering techniques for search suggestions, based on the information available to a large-scale web search engine. We evaluate our novel search suggestion presentation techniques on a real-world dataset of query logs. Based on a set of user studies, we show that by extending the existing assistance layer to effectively group suggestions and label them - while accounting for the query popularity - we substantially increase the user's satisfaction.
Alpa Jain, Gilad Mishne
CIKM1
2010 FactRank: Random Walks on a Web of Facts
Alpa Jain, Patrick Pantel
COLING1
2010 Open Entity Extraction from Web Search Query Logs
Alpa Jain, Marco Pennacchiotti
COLING1
2010 I4E: interactive investigation of iterative information extraction
abstract
Information extraction systems are increasingly being used to mine structured information from unstructured text documents. A commonly used unsupervised technique is to build iterative information extraction (IIE) systems that learn task-specific rules, called patterns, to generate the desired tuples. Oftentimes, output from an information extraction system may contain unexpected results which may be due to an incorrect pattern, incorrect tuple, or both. In such scenarios, users and developers of the extraction system could greatly benefit from an investigation tool that can quickly help them reason about and repair the output.
Anish Das Sarma, Alpa Jain, Divesh Srivastava
SIGMOD Conference2
2009 Identifying comparable entities on the web
abstract
Web search engines are often presented with user queries that involve comparisons of real-world entities. Thus far, this interaction has typically been captured by users submitting appropriately designed keyword queries for which they are presented a list of relevant documents. Richer interactions that explicitly allow for a comparative analysis of entities represent a new potential direction to improve the search experience. With this in mind, we present an initial step of mining comparable entities from sources of information available to a large-scale Web search engine, namely, search query logs and documents from a Web crawl. Our mining methods generate a diverse set of comparables consisting of entities from a broad class of categories, such as medicines, appliances, electronics, and vacation destinations.
Alpa Jain, Patrick Pantel
CIKM1
2009 Join Optimization of Information Extraction Output: Quality Matters!
abstract
Information extraction (IE) systems are trained to extract specific relations from text databases. Real-world applications often require that the output of multiple IE systems be joined to produce the data of interest. To optimize the execution of a join of multiple extracted relations, it is not sufficient to consider only execution time. In fact, the quality of the join output is of critical importance: unlike in the relational world, different join execution plans can produce join results of widely different quality whenever IE systems are involved. In this paper, we develop a principled approach to understand, estimate, and incorporate output quality into the join optimization process over extracted relations. We argue that the output quality is affected by (a) the configuration of the IE systems used to process documents, (b) the document retrieval strategies used to retrieve documents, and (c) the actual join algorithm used. Our analysis considers several alternatives for these factors, and predicts the output quality - and, of course, the execution time - of the alternate execution plans. We establish the accuracy of our analytical models, as well as study the effectiveness of a quality-aware join optimizer, with a large-scale experimental evaluation over real-world text collections and state-of-the-art IE systems.
Alpa Jain, Panagiotis G. Ipeirotis, AnHai Doan, Luis Gravano
ICDE1
2009 Exploring a Few Good Tuples from Text Databases
abstract
Information extraction from text databases is a useful paradigm to populate relational tables and unlock the considerable value hidden in plain-text documents. However, information extraction can be expensive, due to various complex text processing steps necessary in uncovering the hidden data. There are a large number of text databases available, and not every text database is necessarily relevant to every relation. Hence, it is important to be able to quickly explore the utility of running an extractor for a specific relation over a given text database before carrying out the expensive extraction task. In this paper, we present a novel exploration methodology of finding a few good tuples for a relation that can be extracted from a database which allows for judging the relevance of the database for the relation. Specifically, we propose the notion of a good (k, lscr) query as one that can return any k tuples for a relation among the top-lscr fraction of tuples ranked by their aggregated confidence scores, provided by the extractor; if these tuples have high scores, the database can be determined as relevant to the relation. We formalize the access model for information extraction, and investigate efficient query processing algorithms for good (k, lscr) queries, which do not rely on any prior knowledge about the extraction task or the database. We demonstrate the viability of our algorithms using a detailed experimental study with real text databases.
Alpa Jain, Divesh Srivastava
ICDE1
2009 A quality-aware optimizer for information extraction
abstract
A large amount of structured information is buried in unstructured text. Information extraction systems can extract structured relations from the documents and enable sophisticated, SQL-like queries over unstructured text. Information extraction systems are not perfect and their output has imperfect precision and recall (i.e., contains spurious tuples and misses good tuples). Typically, an extraction system has a set of parameters that can be used as “knobs” to tune the system to be either precision- or recall-oriented. Furthermore, the choice of documents processed by the extraction system also affects the quality of the extracted relation. So far, estimating the output quality of an information extraction task has been an ad hoc procedure, based mainly on heuristics. In this article, we show how to use Receiver Operating Characteristic (ROC) curves to estimate the extraction quality in a statistically robust way and show how to use ROC analysis to select the extraction parameters in a principled manner. Furthermore, we present analytic models that reveal how different document retrieval strategies affect the quality of the extracted relation. Finally, we present our maximum likelihood approach for estimating, on the fly, the parameters required by our analytic models to predict the runtime and the output quality of each execution plan. Our experimental evaluation demonstrates that our optimization approach predicts accurately the output quality and selects the fastest execution plan that satisfies the output quality restrictions.
Alpa Jain, Panagiotis G. Ipeirotis
ACM Trans. Database Syst.1
2008 Optimizing SQL Queries over Text Databases
abstract
Text documents often embed data that is structured in nature, and we can expose this structured data using information extraction technology. By processing a text database with information extraction systems, we can materialize a variety of structured "relations," over which we can then issue regular SQL queries. A key challenge to process SQL queries in this text-based scenario is efficiency: information extraction is time-consuming, so query processing strategies should minimize the number of documents that they process. Another key challenge is result quality: in the traditional relational world, all correct execution strategies for a SQL query produce the same (correct) result; in contrast, a SQL query execution over a text database might produce answers that are not fully accurate or complete, for a number of reasons. To address these challenges, we study a family of select-project-join SQL queries over text databases, and characterize query processing strategies on their efficiency and - critically - on their result quality as well. We optimize the execution of SQL queries over text databases in a principled, cost-based manner, incorporating this tradeoff between efficiency and result quality in a user-specific fashion. Our large-scale experiments- over real data sets and multiple information extraction systems - show that our SQL query processing approach consistently picks appropriate execution strategies for the desired balance between efficiency and result quality.
Alpa Jain, AnHai Doan, Luis Gravano
ICDE1
2007 SQL Queries Over Unstructured Text Databases
abstract
Text documents often embed data that is structured in nature. By processing a text database with information extraction systems, we can define a variety of structured "relations" over which we can then issue SQL queries. Processing SQL queries in this text-based scenario presents multiple challenges. One key challenge is efficiency: information extraction is a time-consuming process, so query processing strategies should pick efficient extraction systems whenever possible, and also minimize the number of documents that they process. Another key challenge is result quality: extraction systems might output erroneous information or miss information that they should capture; also, efficiency-related query processing decisions (e.g., to avoid processing large numbers of useless documents) may compromise result completeness. To address these challenges, we characterize SQL query processing strategies in terms of their efficiency and result quality, and discuss the (user-specific) tradeoff between these two properties.
Alpa Jain, AnHai Doan, Luis Gravano
ICDE1
2006 Organizing and Searching the World Wide Web of Facts - Step One: The One-Million Fact Extraction Challenge
Marius Pasca, Dekang Lin, Jeffrey Bigham, Andrei Lifchits, Alpa Jain
AAAI5
2006 Names and Similarities on the Web: Fact Extraction in the Fast Lane
abstract
In a new approach to large-scale extraction of facts from unstructured text, distributional similarities become an integral part of both the iterative acquisition of high-coverage contextual extraction patterns, and the validation and ranking of candidate facts. The evaluation measures the quality and coverage of facts extracted from one hundred million Web documents, starting from ten seed facts and using no additional knowledge, lexicons or complex tools.
Marius Pasca, Dekang Lin, Jeffrey Bigham, Andrei Lifchits, Alpa Jain
ACL5