VLDB 2026 Research / reviewers in the wild / expert
Indrajit Bhattacharya
dblp:b/IndrajitBhattacharya
· DBLP profile ↗
43ranked-venue papers
11as first author
11since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 7 first-author · 9 since 2021Databases, data management, data science and information retrieval · 21 · 7 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 2Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | RetinaQA: A Robust Knowledge Base Question Answering Model for both Answerable and Unanswerable QuestionsabstractAn essential requirement for a real-world Knowledge Base Question Answering (KBQA) system is the ability to detect answerability of questions when generating logical forms.However, state-of-the-art KBQA models assume all questions to be answerable.Recent research has found that such models, when superficially adapted to detect answerability, struggle to satisfactorily identify the different categories of unanswerable questions, and simultaneously preserve good performance for answerable questions.Towards addressing this issue, we propose RetinaQA, a new KBQA model that unifies two key ideas in a single KBQA architecture: (a) discrimination over candidate logical forms, rather than generating these, for handling schema-related unanswerability, and (b) sketch-filling-based construction of candidate logical forms for handling data-related unaswerability.Our results show that RetinaQA significantly outperforms adaptations of state-of-the-art KBQA models in handling both answerable and unanswerable questions and demonstrates robustness across all categories of unanswerability.Notably, Reti-naQA also sets a new state-of-the-art for answerable KBQA, surpassing existing models.We release our code-base 1 for further research. Prayushi Faldu, Indrajit Bhattacharya, Mausam |
ACL (1) | 2 |
| 2024 | Few-shot Transfer Learning for Knowledge Base Question Answering: Fusing Supervised Models with In-Context LearningabstractMayur Patidar, Riya Sawhney, Avinash Singh, Biswajit Chatterjee, Mausam ., Indrajit Bhattacharya. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Mayur Patidar, Riya Sawhney, Avinash Kumar Singh, Biswajit Chatterjee, Mausam, Indrajit Bhattacharya |
ACL (1) | 6 |
| 2024 | FUGAN: A GAN Based Facial Reconstructor for Accurate Unveiling of Hidden Faces
Mrinmoy Sadhukhan, Indrajit Bhattacharya, Paramartha Dutta |
ICPR (6) | 2 |
| 2023 | Do I have the Knowledge to Answer? Investigating Answerability of Knowledge Base QuestionsabstractMayur Patidar, Prayushi Faldu, Avinash Singh, Lovekesh Vig, Indrajit Bhattacharya, Mausam -. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Mayur Patidar, Prayushi Faldu, Avinash Kumar Singh, Lovekesh Vig, Indrajit Bhattacharya, Mausam |
ACL (1) | 5 |
| 2022 | A Weak Supervision Approach for Predicting Difficulty of Technical Interview QuestionsabstractPredicting difficulty of questions is crucial for technical interviews. However, such questions are long-form and more open-ended than factoid and multiple choice questions explored so far for question difficulty prediction. Existing models also require large volumes of candidate response data for training. We study weak-supervision and use unsupervised algorithms for both question generation and difficulty prediction. We create a dataset of interview questions with difficulty scores for deep learning and use it to evaluate SOTA models for question difficulty prediction trained using weak supervision. Our analysis brings out the task’s difficulty as well as the promise of weak supervision for it. Arpita Kundu, Subhasish Ghosh, Pratik Saini, Tapas Nayak, Indrajit Bhattacharya |
COLING | 5 |
| 2021 | Joint Learning of Representations for Web-tables, Entities and Types using Graph Convolutional NetworkabstractExisting approaches for table annotation with entities and types either capture the structure of table using graphical models, or learn embeddings of table entries without accounting for the complete syntactic structure.We propose TabGCN, which uses Graph Convolutional Networks to capture the complete structure of tables, knowledge graph and the training annotations, and jointly learns embeddings for table elements as well as the entities and types.To account for knowledge incompleteness, TabGCN's embeddings can be used to discover new entities and types.Using experiments on 5 benchmark datasets, we show that TabGCN significantly outperforms multiple state-of-the-art baselines for table annotation, while showing promising performance on downstream table-related applications. Aniket Pramanick, Indrajit Bhattacharya |
EACL | 2 |
| 2021 | Complex Question Answering on knowledge graphs using machine translation and multi-task learningabstractSaurabh Srivastava, Mayur Patidar, Sudip Chowdhury, Puneet Agarwal, Indrajit Bhattacharya, Gautam Shroff. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Mayur Patidar, Sudip Chowdhury, Puneet Agarwal, Indrajit Bhattacharya, Gautam Shroff |
EACL | 5 |
| 2021 | Learning-based Assistant for Data Migration of Enterprise Information SystemsabstractData migration from source to target information system is a critical step for modernizing information systems. Central to data migration is data transform that transforms the source system data into target system. In this paper we present a tool that assists the experts in creating the data transformation specification by (a) suggesting candidate field matches between the source and target data models using machine learning and knowledge representation, and (b) rules for the data transformation using program synthesis. It takes the expert’s feedback for the identified matches and synthesized rules and proposes new matches and transformation rules. We have executed our tool on real-life industrial data. Our schema matching recall at 5 is 0.76, while for the rule generator recall at 2 is 0.81. Sayandeep Mitra, Debayan Mukherjee, Atreya Bandyopadhyay, Rajdip Chowdhury, Raveendra Kumar Medicherla, Indrajit Bhattacharya, Ravindra Naik |
ASE | 6 |
| 2021 | Generating An Optimal Interview Question Plan Using A Knowledge Graph And Integer Linear ProgrammingabstractSoham Datta, Prabir Mallick, Sangameshwar Patil, Indrajit Bhattacharya, Girish Palshikar. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Soham Datta, Prabir Mallick, Sangameshwar Patil, Indrajit Bhattacharya, Girish Keshav Palshikar |
NAACL-HLT | 4 |
| 2021 | Analyzing Topic Transitions in Text-Based Social Cascades Using Dual-Network Hawkes Process
Jayesh Choudhari, Srikanta J. Bedathur, Indrajit Bhattacharya, Anirban Dasgupta 0001 |
PAKDD (1) | 3 |
| 2021 | Semi-structured Document Annotation Using Entity and Relation Types
Arpita Kundu, Subhasis Ghosh, Indrajit Bhattacharya |
ECML/PKDD (3) | 3 |
| 2020 | Discovering Knowledge Graph Schema from Short Natural Language Text via DialogabstractWe study the problem of schema discovery for knowledge graphs.We propose a solution where an agent engages in multi-turn dialog with an expert for this purpose.Each minidialog focuses on a short natural language statement, and looks to elicit the expert's desired schema-based interpretation of that statement, taking into account possible augmentations to the schema.The overall schema evolves by performing dialog over a collection of such statements.We take into account the probability that the expert does not respond to a query, and model this probability as a function of the complexity of the query.For such mini-dialogs with response uncertainty, we propose a dialog strategy that looks to elicit the schema over as short a dialog as possible.By combining the notion of uncertainty sampling from active learning with generalized binary search, the strategy asks the query with the highest expected reduction of entropy.We show that this significantly reduces dialog complexity while engaging the expert in meaningful dialog. Subhasis Ghosh, Arpita Kundu, Aniket Pramanick, Indrajit Bhattacharya |
SIGdial | 4 |
| 2019 | Enabling Human-Like Task Identification From Natural ConversationabstractA robot as a coworker or a cohabitant is becoming mainstream day-by-day with the development of low-cost sophisticated hardware. However, an accompanying software stack that can aid the usability of the robotic hardware remains the bottleneck of the process, especially if the robot is not dedicated to a single job. Programming a multi-purpose robot requires an on the fly mission scheduling capability that involves task identification and plan generation. The problem dimension increases if the robot accepts tasks from a human in natural language. Though recent advances in NLP and planner development can solve a variety of complex problems, their amalgamation for a dynamic robotic task handler is used in a limited scope. Specifically, the problem of formulating a planning problem from natural language instructions is not studied in details. In this work, we provide a non-trivial method to combine an NLP engine and a planner such that a robot can successfully identify tasks and all the relevant parameters and generate an accurate plan for the task. Additionally, some mechanism is required to resolve the ambiguity or missing pieces of information in natural language instruction. Thus, we also develop a dialogue strategy that aims to gather additional information with minimal question-answer iterations and only when it is necessary. This work makes a significant stride towards enabling a human-like task understanding capability in a robot. Pradip Pramanick, Chayan Sarkar, P. Balamuralidhar, Ajay Kattepur, Indrajit Bhattacharya, Arpan Pal 0001 |
IROS | 5 |
| 2019 | Your instruction may be crisp, but not clear to me!abstractThe number of robots deployed in our daily surroundings is ever-increasing. Even in the industrial setup, the use of coworker robots is increasing rapidly. These cohabitant robots perform various tasks as instructed by co-located human beings. Thus, a natural interaction mechanism plays a big role in the usability and acceptability of the robot, especially by a non-expert user. The recent development in natural language processing (NLP) has paved the way for chatbots to generate an automatic response for users’ query. A robot can be equipped with such a dialogue system. However, the goal of human-robot interaction is not focused on generating a response to queries, but it often involves performing some tasks in the physical world. Thus, a system is required that can detect user intended task from the natural instruction along with the set of pre- and post-conditions. In this work, we develop a dialogue engine for a robot that can classify and map a task instruction to the robot’s capability. If there is some ambiguity in the instructions or some required information is missing, which is often the case in natural conversation, it asks an appropriate question(s) to resolve it. The goal is to generate minimal and pin-pointed queries for the user to resolve an ambiguity. We evaluate our system for a telepresence scenario where a remote user instructs the robot for various tasks. Our study based on 12 individuals shows that the proposed dialogue strategy can help a novice user to effectively interact with a robot, leading to satisfactory user experience. Pradip Pramanick, Chayan Sarkar, Indrajit Bhattacharya |
RO-MAN | 3 |
| 2018 | Discovering Topical Interactions in Text-Based Cascades Using Hidden Markov Hawkes ProcessesabstractSocial media conversations unfold based on complex interactions between users, topics and time. While recent models have been proposed to capture network strengths between users, users' topical preferences and temporal patterns between posting and response times, interaction patterns between topics has not been studied. We propose the Hidden Markov Hawkes Process (HMHP) that incorporates topical Markov Chains within Hawkes processes to jointly model topical interactions along with user-user and user-topic patterns. We propose a Gibbs sampling algorithm for HMHP that jointly infers the network strengths, diffusion paths, the topics of the posts as well as the topic-topic interactions. We show using experiments on real and semi-synthetic data that HMHP is able to generalize better and recover the network strengths, topics and diffusion paths more accurately than state-of-the-art baselines. More interestingly, HMHP finds insightful interactions between topics in real tweets which no existing model is able to do. Jayesh Choudhari, Anirban Dasgupta 0001, Indrajit Bhattacharya, Srikanta J. Bedathur |
ICDM | 3 |
| 2018 | A multi-criteria evaluation approach in navigation technique for micro-jet for damage & need assessment in disaster response scenarios
Tamal Mondal, Indrajit Bhattacharya, Prithviraj Pramanik, Naiwrita Boral, Jaydeep Roy, Subhanjan Saha, Sujoy Saha |
Knowl. Based Syst. | 2 |
| 2018 | CTMR-collaborative time-stamp based multicast routing for delay tolerant networks in post disaster scenario
J. K. Mandal 0001, Indrajit Bhattacharya, Tamal Mondal, Sourav Sanu Shaw |
Peer-to-Peer Netw. Appl. | 3 |
| 2017 | Stance Classification of Context-Dependent ClaimsabstractRoy Bar-Haim, Indrajit Bhattacharya, Francesco Dinuzzo, Amrita Saha, Noam Slonim. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Roy Bar-Haim, Indrajit Bhattacharya, Francesco Dinuzzo, Amrita Saha, Noam Slonim |
EACL (1) | 2 |
| 2017 | Mitigating selfish, blackhole and wormhole attacks in DTN in a secure, cooperative wayabstractIn this work, security threats posed by selfish, blackhole and wormhole attacks in a delay tolerant network (DTN) are addressed. A cooperative approach has been proposed as a solution where malicious behaviour of a node is tested whenever another node tries to send a message to it. Performance parameters, like received and sent messages to neighbouring nodes have been used along with calculated faith values (CFV) to derive latest CFVs. The process is repeated for each communication instance. In addition, a simple yet efficient credibility-based cryptographic encryption technique is applied to encrypt message contents that works by appending the signatures and CFV values of each encountered nodes. The method has been tested, by implementing it over the spray and wait DTN protocol, in the ONE simulator. The proposed strategy is able to avoid the three types of malicious nodes effectively and increases the trustworthiness of communication process greatly. Amit Kr. Gupta, J. K. Mandal 0001, Indrajit Bhattacharya |
Int. J. Inf. Comput. Secur. | 3 |
| 2016 | An Optimization Approach for Matching Textual Domain Models with Existing CodeabstractWe address the task of mapping a given textual domain model with the source code of an application which is in the same domain but was developed independently of the domain model. The key novelty of our approach is to use mathematical optimization to find a mapping between the elements in the two sides that maximizes the instances of clusters of related elements on each side being mapped to clusters of similarly related elements on the other side. We describe experiments wherein we apply our approach to the task of matching two real, open-source applications to corresponding industry-standard domain models. In comparison with previous approaches that leverage relationships, but are formulated as heuristics rather than as a principled optimization problem, our approach gives up to 40% higher precision given a desired level of recall. Tejas Patil, Raghavan Komondoor, Deepak D'Souza, Indrajit Bhattacharya |
ICSME | 4 |
| 2016 | DirMove: direction of movement based routing in DTN architecture for post-disaster scenario
Indrajit Bhattacharya, Parthasarathi Banerjee, J. K. Mandal 0001, Animesh Mukherjee 0001 |
Wirel. Networks | 2 |
| 2015 | Online Topic-based Social Influence Analysis for the Wimbledon ChampionshipsabstractVarious industries are turning to social media to identify key influencers on topics of interest. Following this trend, the All England Lawn Tennis and Croquet Club (AELTC) is keen to analyze the `social pulse' around the famous Wimbledon Championships. IBM developed and deployed social influence analysis capability for AELTC during the 2014 edition of the Championship. The design and implementation of influence analysis technology in the real world involves several challenges. In this paper, we define various functional and usability criteria that social influence scores should satisfy, and propose a multi-dimensional definition of influence that satisfies these criteria. We highlight the need to identify both all-time influencers and recent influencers, and track user influences over multiple time-scales for this purpose. We also stress the importance of aspect-specific influence analysis, and investigate an approach that uses an aspect hierarchy that annotates tweets with topics or aspects before analyzing them for influence. We also describe interesting insights discovered by our tool and the lessons that we learnt from this engagement. Varun Embar, Indrajit Bhattacharya, Vinayaka Pandit, Roman Vaculín |
KDD | 2 |
| 2014 | A bayesian framework for estimating properties of network diffusionsabstractThe analysis of network connections, diffusion processes and cascades requires evaluating properties of the diffusion network. Properties of interest often involve variables that are not explicitly observed in real world diffusions. Connection strengths in the network and diffusion paths of infections over the network are examples of such hidden variables. These hidden variables therefore need to be estimated for these properties to be evaluated. In this paper, we propose and study this novel problem in a Bayesian framework by capturing the posterior distribution of these hidden variables given the observed cascades, and computing the expectation of these properties under this posterior distribution. We identify and characterize interesting network diffusion properties whose expectations can be computed exactly and efficiently, either wholly or in part. For properties that are not `nice' in this sense, we propose a Gibbs Sampling framework for Monte Carlo integration. In detailed experiments using various network diffusion properties over multiple synthetic and real datasets, we demonstrate that the proposed approach is significantly more accurate than a frequentist plug-in baseline. We also propose a map-reduce implementation of our framework and demonstrate that this can analyze cascades with millions of infections in minutes. Varun Embar, Rama Kumar Pasumarthi, Indrajit Bhattacharya |
KDD | 3 |
| 2013 | Nested Hierarchical Dirichlet Process for Nonparametric Entity-Topic Analysis
Priyanka Agrawal, Lavanya Sita Tekumalla, Indrajit Bhattacharya |
ECML/PKDD (2) | 3 |
| 2013 | A Layered Dirichlet Process for Hierarchical Segmentation of Sequential Grouped Data
Adway Mitra, Ranganath B. N., Indrajit Bhattacharya |
ECML/PKDD (2) | 3 |
| 2012 | Dynamic Multi-relational Chinese Restaurant Process for Analyzing Influences on Users in Social MediaabstractWe study the problem of analyzing influence of various factors affecting individual messages posted in social media. The problem is challenging because of various types of influences propagating through the social media network that act simultaneously on any user. Additionally, the topic composition of the influencing factors and the susceptibility of users to these influences evolve over time. This problem has not been studied before, and off-the-shelf models are unsuitable for this purpose. To capture the complex interplay of these various factors, we propose a new non-parametric model called the Dynamic Multi-Relational Chinese Restaurant Process. This accounts for the user network for data generation and also allows the parameters to evolve over time. Designing inference algorithms for this model suited for large scale social-media data is another challenge. To this end, we propose a scalable and multi-threaded inference algorithm based on online Gibbs Sampling. Extensive evaluations on large-scale Twitter and Face book data show that the extracted topics when applied to authorship and commenting prediction outperform state-of-the-art baselines. More importantly, our model produces valuable insights on topic trends and user personality trends beyond the capability of existing approaches. Himabindu Lakkaraju, Indrajit Bhattacharya, Chiranjib Bhattacharyya |
ICDM | 2 |
| 2012 | Cross-Guided Clustering: Transfer of Relevant Supervision across TasksabstractLack of supervision in clustering algorithms often leads to clusters that are not useful or interesting to human reviewers. We investigate if supervision can be automatically transferred for clustering a target task, by providing a relevant supervised partitioning of a dataset from a different source task. The target clustering is made more meaningful for the human user by trading-off intrinsic clustering goodness on the target task for alignment with relevant supervised partitions in the source task, wherever possible. We propose a cross-guided clustering algorithm that builds on traditional k-means by aligning the target clusters with source partitions. The alignment process makes use of a cross-task similarity measure that discovers hidden relationships across tasks. When the source and target tasks correspond to different domains with potentially different vocabularies, we propose a projection approach using pivot vocabularies for the cross-domain similarity measure. Using multiple real-world and synthetic datasets, we show that our approach improves clustering accuracy significantly over traditional k-means and state-of-the-art semi-supervised clustering baselines, over a wide range of data characteristics and parameter settings. Indrajit Bhattacharya, Shantanu Godbole, Sachindra Joshi, Ashish Verma 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2011 | Learning Dirichlet Processes from Partially Observed GroupsabstractMotivated by the task of vernacular news analysis using known news topics from national news-papers, we study the task of topic analysis, where given source datasets with observed topics, data items from a target dataset need to be assigned either to observed source topics or to new ones. Using Hierarchical Dirichlet Processes for addressing this task imposes unnecessary and often inappropriate generative assumptions on the observed source topics. In this paper, we explore Dirichlet Processes with partially observed groups (POG-DP). POG-DP avoids modeling the given source topics. Instead, it directly models the conditional distribution of the target data as a mixture of a Dirichlet Process and the posterior distribution of a Hierarchical Dirichlet Process with known groups and topics. This introduces coupling between selection probabilities of all topics within a source, leading to effective identification of source topics. We further improve on this with a Combinatorial Dirichlet Process with partially observed groups (POG-CDP) that captures finer grained coupling between related topics by choosing intersections between sources. We evaluate our models in three different real-world applications. Using extensive experimentation, we compare against several baselines to show that our model performs significantly better in all three applications. Avinava Dubey, Indrajit Bhattacharya, Mrinal Kanti Das, Tanveer A. Faruquie, Chiranjib Bhattacharyya |
ICDM | 2 |
| 2011 | Exploiting Coherence for the Simultaneous Discovery of Latent Facets and associated SentimentsabstractFacet-based sentiment analysis involves discovering the latent facets, sentiments and their associations. Traditional facet-based sentiment analysis algorithms typically perform the various tasks in sequence, and fail to take advantage of the mutual reinforcement of the tasks. Additionally, inferring sentiment levels typically requires domain knowledge or human intervention. In this paper, we propose a series of probabilistic models that jointly discover latent facets and sentiment topics, and also order the sentiment topics with respect to a multi-point scale, in a language and domain independent manner. This is achieved by simultaneously capturing both short-range syntactic structure and long range semantic dependencies between the sentiment and facet words. The models further incorporate coherence in reviews, where reviewers dwell on one facet or sentiment level before moving on, for more accurate facet and sentiment discovery. For reviews which are supplemented with ratings, our models automatically order the latent sentiment topics, without requiring seed-words or domain-knowledge. To the best of our knowledge, our work is the first attempt to combine the notions of syntactic and semantic dependencies in the domain of review mining. Further, the concept of facet and sentiment coherence has not been explored earlier either. Extensive experimental results on real world review data show that the proposed models outperform various state of the art baselines for facet-based sentiment analysis. Himabindu Lakkaraju, Chiranjib Bhattacharyya, Indrajit Bhattacharya, Srujana Merugu |
SDM | 3 |
| 2010 | Building re-usable dictionary repositories for real-world text miningabstractText mining, though still a nascent industry, has been growing quickly along with the awareness of the importance of unstructured data in business analytics, customer retention and extension, social media, and legal applications. There has been a recent increase in the number of commercial text mining product and service offerings, but successful or wide-spread deployments are rare, mainly due to a dependence on the expertise and skill of practitioners. Accordingly, there is a growing need for re-usable repositories for text mining. In this paper, we focus on dictionary-based text mining and its role in enabling practitioners in understanding and analyzing large text datasets. We motivate and define the problem of exploratory dictionary construction for capturing concepts of interest, and propose a framework for efficient construction, tuning, and re-use of these dictionaries across datasets. The construction framework offers a range of interaction modes to the user to quickly build concept dictionaries over large datasets. We also show how to adapt one or more dictionaries across domains and tasks, thereby enabling reuse of knowledge and effort in industrial practice. We present results and case studies on real-life CRM analytics datasets, where such repositories and tooling significantly cut down practitioner time and effort for dictionary-based text mining. Shantanu Godbole, Indrajit Bhattacharya, Ashish Verma 0001 |
CIKM | 2 |
| 2010 | A Cluster-Level Semi-supervision Model for Interactive Clustering
Avinava Dubey, Indrajit Bhattacharya, Shantanu Godbole |
ECML/PKDD (1) | 2 |
| 2009 | Cross-Guided Clustering: Transfer of Relevant Supervision across Domains for Improved ClusteringabstractLack of supervision in clustering algorithms often leads to clusters that are not useful or interesting to human reviewers. We investigate if supervision can be automatically transferred to a clustering task in a target domain, by providing a relevant supervised partitioning of a dataset from a different source domain. The target clustering is made more meaningful for the human user by trading off intrinsic clustering goodness on the target dataset for alignment with relevant supervised partitions in the source dataset, wherever possible. We propose a cross-guided clustering algorithm that builds on traditional k-means by aligning the target clusters with source partitions. The alignment process makes use of a cross-domain similarity measure that discovers hidden relationships across domains with potentially different vocabularies. Using multiple real-world datasets, we show that our approach improves clustering accuracy significantly over traditional k-means. Indrajit Bhattacharya, Shantanu Godbole, Sachindra Joshi, Ashish Verma 0001 |
ICDM | 1 |
| 2009 | Enabling analysts in managed services for CRM analyticsabstractData analytics tools and frameworks abound, yet rapid deployment of analytics solutions that deliver actionable insights from business data remains a challenge. The primary reason is that on-field practitioners are required to be both technically proficient and knowledgeable about the business. The recent abundance of unstructured business data has thrown up new opportunities for analytics, but has also multiplied the deployment challenge, since interpretation of concepts derived from textual sources require a deep understanding of the business. In such a scenario, a managed service for analytics comes up as the best alternative. A managed analytics service is centered around a business analyst who acts as a liaison between the business and the technology. This calls for new tools that assist the analyst to be efficient in the tasks that she needs to execute. Also, the analytics needs to be repeatable, in that the delivered insights should not depend heavily on the expertise of specific analysts. These factors lead us to identify new areas that open up for KDD research in terms of 'time-to-insight' and repeatability for these analysts. We present our analytics framework in the form of a managed service offering for CRM analytics. We describe different analyst-centric tools using a case study from real-life engagements and demonstrate their effectiveness. Indrajit Bhattacharya, Shantanu Godbole, Ashish Verma 0001, Jeff Achtermann, Kevin English |
KDD | 1 |
| 2008 | Structured entity identification and document categorization: two tasks with one joint modelabstractTraditionally, research in identifying structured entities in documents has proceeded independently of document categorization research. In this paper, we observe that these two tasks have much to gain from each other. Apart from direct references to entities in a database, such as names of person entities, documents often also contain words that are correlated with discriminative entity attributes, such age-group and income-level of persons. This happens naturally in many enterprise domains such as CRM, Banking, etc. Then, entity identification, which is typically vulnerable against noise and incompleteness in direct references to entities in documents, can benefit from document categorization with respect to such attributes. In return, entity identification enables documents to be categorized according to different label-sets arising from entity attributes without requiring any supervision. In this paper, we propose a probabilistic generative model for joint entity identification and document categorization. We show how the parameters of the model can be estimated using an EM algorithm in an unsupervised fashion. Using extensive experiments over real and semi-synthetic data, we demonstrate that the two tasks can benefit immensely from each other when performed jointly using the proposed model. Indrajit Bhattacharya, Shantanu Godbole, Sachindra Joshi |
KDD | 1 |
| 2007 | Online Collective Entity Resolution
Indrajit Bhattacharya, Lise Getoor |
AAAI | 1 |
| 2007 | Query-time Entity ResolutionabstractEntity resolution is the problem of reconciling database references corresponding to the same real-world entities. Given the abundance of publicly available databases that have unresolved entities, we motivate the problem of query-time entity resolution quick and accurate resolution for answering queries over such `unclean' databases at query-time. Since collective entity resolution approaches --- where related references are resolved jointly --- have been shown to be more accurate than independent attribute-based resolution for off-line entity resolution, we focus on developing new algorithms for collective resolution for answering entity resolution queries at query-time. For this purpose, we first formally show that, for collective resolution, precision and recall for individual entities follow a geometric progression as neighbors at increasing distances are considered. Unfolding this progression leads naturally to a two stage `expand and resolve' query processing strategy. In this strategy, we first extract the related records for a query using two novel expansion operators, and then resolve the extracted records collectively. We then show how the same strategy can be adapted for query-time entity resolution by identifying and resolving only those database references that are the most helpful for processing the query. We validate our approach on two large real-world publication databases where we show the usefulness of collective resolution and at the same time demonstrate the need for adaptive strategies for query processing. We then show how the same queries can be answered in real-time using our adaptive approach while preserving the gains of collective resolution. In addition to experiments on real datasets, we use synthetically generated data to empirically demonstrate the validity of the performance trends predicted by our analysis of collective entity resolution over a wide range of structural characteristics in the data. Indrajit Bhattacharya, Lise Getoor |
J. Artif. Intell. Res. | 1 |
| 2007 | Collective entity resolution in relational dataabstractMany databases contain uncertain and imprecise references to real-world entities. The absence of identifiers for the underlying entities often results in a database which contains multiple references to the same entity. This can lead not only to data redundancy, but also inaccuracies in query processing and knowledge extraction. These problems can be alleviated through the use of entity resolution . Entity resolution involves discovering the underlying entities and mapping each database reference to these entities. Traditionally, entities are resolved using pairwise similarity over the attributes of references. However, there is often additional relational information in the data. Specifically, references to different entities may cooccur. In these cases, collective entity resolution, in which entities for cooccurring references are determined jointly rather than independently, can improve entity resolution accuracy. We propose a novel relational clustering algorithm that uses both attribute and relational information for determining the underlying domain entities, and we give an efficient implementation. We investigate the impact that different relational similarity measures have on entity resolution quality. We evaluate our collective entity resolution algorithm on multiple real-world databases. We show that it improves entity resolution performance over both attribute-based baselines and over algorithms that consider relational information but do not resolve entities collectively. In addition, we perform detailed experiments on synthetically generated data to identify data characteristics that favor collective relational resolution over purely attribute-based algorithms. Indrajit Bhattacharya, Lise Getoor |
ACM Trans. Knowl. Discov. Data | 1 |
| 2006 | Query-time entity resolutionabstractThe goal of entity resolution is to reconcile database references corresponding to the same real-world entities. Given the abundance of publicly available databases where entities are not resolved, we motivate the problem of quickly processing queries that require resolved entities from such 'unclean' databases. We propose a two-stage collective resolution strategy for processing queries. We then show how it can be performed on-the-fly by adaptively extracting and resolving those database references that are the most helpful for resolving the query. We validate our approach on two large real-world publication databases where we show the usefulness of collective resolution and at the same time demonstrate the need for adaptive strategies for query processing. We then show how the same queries can be answered in real time using our adaptive approach while preserving the gains of collective resolution. Indrajit Bhattacharya, Lise Getoor, Louis Licamele |
KDD | 1 |
| 2006 | A Latent Dirichlet Model for Unsupervised Entity ResolutionabstractEntity resolution has received considerable attention in recent years. Given many references to underlying entities, the goal is to predict which references correspond to the same entity. We show how to extend the Latent Dirichlet Allocation model for this task and propose a probabilistic model for collective entity resolution for relational domains where references are connected to each other. Our approach differs from other recently proposed entity resolution approaches in that it is a) generative, b) does not make pair-wise decisions and c) captures relations between entities through a hidden group variable. We propose a novel sampling algorithm for collective entity resolution which is unsupervised and also takes entity relations into account. Additionally, we do not assume the domain of entities to be known and show how to infer the number of entities from the data. We demonstrate the utility and practicality of our relational entity resolution approach for author resolution in two real-world bibliographic datasets. In addition, we present preliminary results on characterizing conditions under which relational information is useful. Indrajit Bhattacharya, Lise Getoor |
SDM | 1 |
| 2005 | Similarity Searching in Peer-to-Peer DatabasesabstractWe consider the problem of handling similarity queries in peer-to-peer databases. We propose an indexing and searching mechanism which, given a query object, returns the set of objects in the database that are semantically related to the query. We propose an indexing scheme which clusters data such that semantically related objects are partitioned into a small set of clusters, allowing for a simple and efficient similarity search strategy. Our indexing scheme also decouples object and node locations. Our adaptive replication and randomized lookup schemes exploit this feature and ensure that the number of copies of an object is proportional to its popularity and all replicas are equally likely to serve a given query, thus achieving perfect load balancing. The techniques developed in this work are oblivious to the underlying DHT topology and can be implemented on a variety of structured overlays such as CAN, CHORD, Pastry, and Tapestry. We also present DHT-independent analytical guarantees for the performance of our algorithms in terms of search accuracy, cost, and load-balance; the experimental results from our simulations confirm the insights derived from these analytical models Indrajit Bhattacharya, Srinivas R. Kashyap, Srinivasan Parthasarathy 0002 |
ICDCS | 1 |
| 2004 | Unsupervised Sense Disambiguation Using Bilingual Probabilistic ModelsabstractWe describe two probabilistic models for unsupervised word-sense disambiguation using parallel corpora. The first model, which we call the Sense model, builds on the work of Diab and Resnik (2002) that uses both parallel text and a sense inventory for the target language, and recasts their approach in a probabilistic framework. The second model, which we call the Concept model, is a hierarchical model that uses a concept latent variable to relate different language specific sense labels. We show that both models improve performance on the word sense disambiguation task over previous unsupervised approaches, with the Concept model showing the largest improvement. Furthermore, in learning the Concept model, as a by-product, we learn a sense inventory for the parallel language. Indrajit Bhattacharya, Lise Getoor, Yoshua Bengio |
ACL | 1 |
| 2002 | Quantified Computation Tree Logic
Anindya C. Patthak, Indrajit Bhattacharya, Anirban Dasgupta 0001, Pallab Dasgupta, P. P. Chakrabarti 0001 |
Inf. Process. Lett. | 2 |
| 2002 | An Object-Oriented Fuzzy Data Model for Similarity Detection in Image DatabasesabstractWe introduce a fuzzy set theoretic approach for dealing with uncertainty in images in the context of spatial and topological relations existing among the objects in the image. We propose an object-oriented graph theoretic model for representing an image and this model allows us to assess the similarity between images using the concept of (fuzzy) graph matching. Sufficient flexibility has been provided in the similarity algorithm so that different features of an image may be independently focused upon. Arun K. Majumdar, Indrajit Bhattacharya, Amit K. Saha |
IEEE Trans. Knowl. Data Eng. | 2 |