Sachindra Joshi

dblp:96/2418 · DBLP profile ↗
← Back
55ranked-venue papers
5as first author
12since 2021 · last 2024
0009-0009-8089-9493ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 17 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8 · 3 since 2021Theory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
abstract
Distribution matching methods for language model alignment such as Generation with Distributional Control (GDC) and Distributional Policy Gradient (DPG) have not received the same level of attention in reinforcement learning from human feedback (RLHF) as contrastive methods such as Sequence Likelihood Calibration (SLiC), Direct Preference Optimization (DPO) and its variants. We identify high variance of the gradient estimate as the primary reason for the lack of success of these methods and propose a self-normalized baseline to reduce the variance. We further generalize the target distribution in DPG, GDC and DPO by using Bayes' rule to define the reward-conditioned posterior. The resulting approach, referred to as BRAIn - Bayesian Reward-conditioned Amortized Inference acts as a bridge between distribution matching methods and DPO and significantly outperforms prior art in summarization and Antropic HH tasks.
Gaurav Pandey 0001, Yatin Nandwani, Tahira Naseem, Guangxuan Xu, Dinesh Raghu, Sachindra Joshi, Asim Munawar, Ramón Fernandez Astudillo
ICML7
2023 Pointwise Mutual Information Based Metric and Decoding Strategy for Faithful Generation in Document Grounded Dialogs
abstract
A major concern in using deep learning based generative models for document-grounded dialogs is the potential generation of responses that are not faithful to the underlying document.Existing automated metrics used for evaluating the faithfulness of response with respect to the grounding document measure the degree of similarity between the generated response and the document's content.However, these automated metrics are far from being well aligned with human judgments.Therefore, to improve the measurement of faithfulness, we propose a new metric that utilizes (Conditional) Point-wise Mutual Information (PMI) between the generated response and the source document, conditioned on the dialogue.PMI quantifies the extent to which the document influences the generated response -with a higher PMI indicating a more faithful response.We build upon this idea to create a new decoding technique that incorporates PMI into the response generation process to predict more faithful responses.Our experiments on the BEGIN benchmark demonstrate an improved correlation of our metric with human evaluation.We also show that our decoding technique is effective in generating more faithful responses when compared to standard decoding techniques on a set of publicly available document-grounded dialog datasets.
Yatin Nandwani, Dinesh Raghu, Sachindra Joshi, Luis A. Lastras
EMNLP4
2022 Structural Constraints and Natural Language Inference for End-to-End Flowchart Grounded Dialog Response Generation
abstract
Flowchart grounded dialog systems converse with users by following a given flowchart and a corpus of FAQs.The existing state-of-the-art approach (Raghu et al., 2021) for learning such a dialog system, named FLONET, has two main limitations.(1) It uses a Retrieval Augmented Generation (RAG) framework which represents a flowchart as a bag of nodes.By doing so, it loses the connectivity structure between nodes which can aid in better response generation.(2) Typically dialogs progress with the agent asking polar (Y/N) questions, but users often respond indirectly without the explicit use of polar words.In such cases, it fails to understand the correct polarity of the answer.To overcome these issues, we propose Structure-Aware FLONET (SA-FLONET) which infuses structural constraints derived from the connectivity structure of flowcharts into the RAG framework.It uses natural language inference to better predict the polarity of indirect Y/N answers.We find that SA-FLONET outperforms FLONET, with a success rate improvement of 68% and 123% in flowchart grounded response generation and zero-shot flowchart grounded response generation tasks respectively.
Dinesh Raghu, Suraj Joshi, Sachindra Joshi, Mausam
EMNLP3
2022 Learning as Conversation: Dialogue Systems Reinforced for Information Acquisition
abstract
Pengshan Cai, Hui Wan, Fei Liu, Mo Yu, Hong Yu, Sachindra Joshi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Pengshan Cai, Hui Wan 0001, Fei Liu 0004, Mo Yu, Hong Yu 0001, Sachindra Joshi
NAACL-HLT6
2022 DG2: Data Augmentation Through Document Grounded Dialogue Generation
abstract
Collecting data for training dialog systems can be extremely expensive due to the involvement of human participants and the need for extensive annotation.Especially in documentgrounded dialog systems, human experts need to carefully read the unstructured documents to answer the users' questions.As a result, existing document-grounded dialog datasets are relatively small-scale and obstruct the effective training of dialogue systems.In this paper, we propose an automatic data augmentation technique grounded on documents through a generative dialogue model.The dialogue model consists of a user bot and agent bot that can synthesize diverse dialogues given an input document, which are then used to train a downstream model.When supplementing the original dataset, our method achieves significant improvement over traditional data augmentation methods.We also achieve competitive performance in the low-resource setting.
Qingyang Wu, Song Feng 0002, Derek Chen, Sachindra Joshi, Luis A. Lastras
SIGDIAL4
2021 Bootstrapping Dialog Models from Human to Human Conversation Logs
abstract
State-of-the-art commercial dialog platforms provide powerful tools to build a conversational agent. These platforms provide complete control to the dialog designer to model user-agent interactions. However, a dialog designer needs to rely on domain experts to manually build the dialog model -- by creating dialog flow nodes and modeling user intents. This process is laborious, time consuming and expensive and does not allow the designer to exploit human to human conversation logs effectively. In this work, we present a research prototype that can ingest human-to-human conversation logs between an end-user and an agent, and suggest user-intents and agent-responses, given a conversation context. We utilize human to human conversation logs to build two emulators: user and agent. An agent emulator models an agent response given the conversation context so far, and a user emulator outputs possible user responses. Our system is able to recommend conversational intents as well as conversation flow using emulators based on real-world data, thus making the process of designing a bot more efficient. To the best our knowledge this is the first system that enables data-driven dialog model creation by emulating users and agents.
Pankaj Dhoolia, Danish Contractor, Sachindra Joshi
AAAI4
2021 Doc2Bot: Document grounded Bot Framework
abstract
Conversational agents, or chatbots, are widely used to provide customer care and other informational support. Currently, the development of chatbots using standard frameworks requires a lot of manual crafting by subject matter experts (SMEs). On the other hand, while learning-based approaches to dialog have made significant advancements, they require training with a large volume of dialog data, which chatbot developers typically do not have access to. To tackle these challenges, we introduce DOC2BOT, a system that supports the automated construction of chatbots by digesting various forms of documents such as business manuals, HowTos, and customer support pages that organizations own. In addition to that, DOC2BOT provides a user-friendly experience to SMEs, and to minimize their effort by supporting intuitive interactions and streamlining their workflow.
Kshitij Fadnis, Pankaj Dhoolia, Qingzi Vera Liao, Steven Ross, Nathaniel Mills, Sachindra Joshi, Luis A. Lastras
AAAI7
2021 MultiDoc2Dial: Modeling Dialogues Grounded in Multiple Documents
abstract
We propose MultiDoc2Dial, a new task and dataset on modeling goal-oriented dialogues grounded in multiple documents.Most previous works treat document-grounded dialogue modeling as a machine reading comprehension task based on a single given document or passage.In this work, we aim to address more realistic scenarios where a goaloriented information-seeking conversation involves multiple topics, and hence is grounded on different documents.To facilitate such a task, we introduce a new dataset that contains dialogues grounded in multiple documents from four different domains.We also explore modeling the dialogue-based and documentbased context in the dataset.We present strong baseline approaches and various experimental results, aiming to support further research efforts on such a task. Social Security CreditsYou must earn at least 40 Social Security credits to qualify for social security benefits. Number of Credit Needed for Disability BenefitsTo be eligible for disability benefits, you must meet a recent work test and a duration work test.Number of Credit Needed for Retirement Benefits If you are born after 1928, you will need 40 credits to qualify for retirement benefits.30 years or older -In general, you must have at least 20 credits in the 10-year period immediately before you become disabled.U1: I need help with SSDI.I heard that it could benefit my relatives too.I am in my 50s.A2: Yes SSDI pays benefits to you and
Song Feng 0002, Siva Sankalp Patel, Hui Wan 0001, Sachindra Joshi
EMNLP (1)4
2021 End-to-End Learning of Flowchart Grounded Task-Oriented Dialogs
abstract
We propose a novel problem within end-toend learning of task oriented dialogs (TOD), in which the dialog system mimics a troubleshooting agent who helps a user by diagnosing their problem (e.g., car not starting).Such dialogs are grounded in domain-specific flowcharts, which the agent is supposed to follow during the conversation.Our task exposes novel technical challenges for neural TOD, such as grounding an utterance to the flowchart without explicit annotation, referring to additional manual pages when user asks a clarification question, and ability to follow unseen flowcharts at test time.We release a dataset (FLODIAL) consisting of 2,738 dialogs grounded on 12 different troubleshooting flowcharts.We also design a neural model, FLONET, which uses a retrieval-augmented generation architecture to train the dialog agent.Our experiments find that FLONET can do zero-shot transfer to unseen flowcharts, and sets a strong baseline for future research.
Dinesh Raghu, Shantanu Agarwal, Sachindra Joshi, Mausam
EMNLP (1)3
2021 Integrating Dialog History into End-to-End Spoken Language Understanding Systems
abstract
End-to-end spoken language understanding (SLU) systems that process human-human or human-computer interactions are often context independent and process each turn of a conversation independently. Spoken conversations on the other hand, are very much context dependent, and dialog history contains useful information that can improve the processing of each conversational turn. In this paper, we investigate the importance of dialog history and how it can be effectively integrated into end-to-end SLU systems. While processing a spoken utterance, our proposed RNN transducer (RNN-T) based SLU model has access to its dialog history in the form of decoded transcripts and SLU labels of previous turns. We encode the dialog history as BERT embeddings, and use them as an additional input to the SLU model along with the speech features for the current utterance. We evaluate our approach on a recently released spoken dialog data set, the HarperValleyBank corpus. We observe significant improvements: 8% for dialog action and 30% for caller intent recognition tasks, in comparison to a competitive context independent end-to-end baseline system.
Jatin Ganhotra, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Sachindra Joshi, George Saon, Zoltán Tüske, Brian Kingsbury
Interspeech4
2021 Explaining Neural Network Predictions on Sentence Pairs via Learning Word-Group Masks
abstract
Hanjie Chen, Song Feng, Jatin Ganhotra, Hui Wan, Chulaka Gunasekara, Sachindra Joshi, Yangfeng Ji. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Song Feng 0002, Jatin Ganhotra, Hui Wan 0001, R. Chulaka Gunasekara, Sachindra Joshi, Yangfeng Ji
NAACL-HLT6
2021 Does Structure Matter? Encoding Documents for Machine Reading Comprehension
abstract
Hui Wan, Song Feng, Chulaka Gunasekara, Siva Sankalp Patel, Sachindra Joshi, Luis Lastras. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Hui Wan 0001, Song Feng 0002, R. Chulaka Gunasekara, Siva Sankalp Patel, Sachindra Joshi, Luis A. Lastras
NAACL-HLT5
2020 Mask & Focus: Conversation Modelling by Learning Concepts
Gaurav Pandey 0001, Dinesh Raghu, Sachindra Joshi
AAAI3
2020 Unsupervised Learning of Interpretable Dialog Models
abstract
Recently several deep learning based models have been proposed for end-to-end learning of dialogs. While these models can be trained from data without the need for any additional annotations, it is hard to interpret them. On the other hand, there exist traditional state based dialog systems, where the states of the dialog are discrete and hence easy to interpret. However these states need to be handcrafted and annotated in the data. To achieve the best of both worlds, we propose Latent State Tracking Network (LSTN) using which we learn an interpretable model in unsupervised manner. The model defines a discrete latent variable at each turn of the conversation which can take a finite set of values. Since these discrete variables are not present in the training data, we use EM algorithm to train our model in unsupervised manner. In the experiments, we show that LSTN can help achieve interpretability in dialog models without much decrease in performance compared to end-to-end approaches.
Dhiraj Madan, Dinesh Raghu, Gaurav Pandey 0001, Sachindra Joshi
ECAI4
2020 doc2dial: A Goal-Oriented Document-Grounded Dialogue Dataset
abstract
We introduce doc2dial, a new dataset of goal-oriented dialogues that are grounded in the associated documents.Inspired by how the authors compose documents for guiding end users, we first construct dialogue flows based on the content elements that corresponds to higher-level relations across text sections as well as lower-level relations between discourse units within a section.Then we present these dialogue flows to crowd contributors to create conversational utterances.The dataset includes over 4500 annotated conversations with an average of 14 turns that are grounded in over 450 documents from four domains.Compared to the prior document-grounded dialogue datasets, this dataset covers a variety of dialogue scenes in information-seeking conversations.For evaluating the versatility of the dataset, we introduce multiple dialogue modeling tasks and present baseline approaches.
Song Feng 0002, Hui Wan 0001, R. Chulaka Gunasekara, Siva Sankalp Patel, Sachindra Joshi, Luis A. Lastras
EMNLP (1)5
2020 Conversational Document Prediction to Assist Customer Care Agents
abstract
Jatin Ganhotra, Haggai Roitman, Doron Cohen, Nathaniel Mills, Chulaka Gunasekara, Yosi Mass, Sachindra Joshi, Luis Lastras, David Konopnicki. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Jatin Ganhotra, Haggai Roitman, Doron Cohen 0001, Nathaniel Mills, R. Chulaka Gunasekara, Yosi Mass, Sachindra Joshi, Luis A. Lastras, David Konopnicki
EMNLP (1)7
2020 Neural Conversational QA: Learning to Reason vs Exploiting Patterns
abstract
Neural Conversational QA tasks like ShARC require systems to answer questions based on the contents of a given passage.On studying recent state-of-the-art models on the ShARC QA task, we found indications that the models learn spurious clues/patterns in the dataset.Furthermore, we show that a heuristic-based program designed to exploit these patterns can have performance comparable to that of the neural models.In this paper we share our findings about four types of patterns found in the ShARC corpus and describe how neural models exploit them.Motivated by the aforementioned findings, we create and share a modified dataset that has fewer spurious patterns, consequently allowing models to learn better.
Nikhil Verma, Dhiraj Madan, Danish Contractor, Sachindra Joshi
EMNLP (1)6
2019 A Practical Dialogue-Act-Driven Conversation Model for Multi-Turn Response Selection
abstract
Harshit Kumar, Arvind Agarwal, Sachindra Joshi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Arvind Agarwal, Sachindra Joshi
EMNLP/IJCNLP (1)3
2018 Dialogue Act Sequence Labeling Using Hierarchical Encoder With CRF
abstract
Dialogue Act recognition associate dialogue acts (i.e., semantic labels) to utterances in a conversation. The problem of associating semantic labels to utterances can be treated as a sequence labeling problem. In this work, we build a hierarchical recurrent neural network using bidirectional LSTM as a base unit and the conditional random field (CRF) as the top layer to classify each utterance into its corresponding dialogue act. The hierarchical network learns representations at multiple levels, i.e., word level, utterance level, and conversation level. The conversation level representations are input to the CRF layer, which takes into account not only all previous utterances but also their dialogue acts, thus modeling the dependency among both, labels and utterances, an important consideration of natural dialogue. We validate our approach on two different benchmark data sets, Switchboard and Meeting Recorder Dialogue Act, and show performance improvement over the state-of-the-art methods by 2.2% and 4.1% absolute points, respectively. It is worth noting that the inter-annotator agreement on Switchboard data set is 84%, and our method is able to achieve the accuracy of about 79% despite being trained on the noisy data.
Arvind Agarwal, Riddhiman Dasgupta, Sachindra Joshi
AAAI4
2018 Exemplar Encoder-Decoder for Neural Conversation Generation
abstract
In this paper we present the Exemplar Encoder-Decoder network (EED), a novel conversation model that learns to utilize similar examples from training data to generate responses.Similar conversation examples (context-response pairs) from training data are retrieved using a traditional TF-IDF based retrieval model.The retrieved responses are used to create exemplar vectors that are used by the decoder to generate the response.The contribution of each retrieved response is weighed by the similarity of corresponding context with the input context.We present detailed experiments on two large data sets and find that our method outperforms state of the art sequence to sequence generative models on several recently proposed evaluation metrics.We also observe that the responses generated by the proposed EED model are more informative and diverse compared to existing state-of-the-art method.
Gaurav Pandey 0001, Danish Contractor, Sachindra Joshi
ACL (1)4
2018 Dialogue-act-driven Conversation Model : An Experimental Study
abstract
The utility of additional semantic information for the task of next utterance selection in an automated dialogue system is the focus of study in this paper. In particular, we show that additional information available in the form of dialogue acts –when used along with context given in the form of dialogue history– improves the performance irrespective of the underlying model being generative or discriminative. In order to show the model agnostic behavior of dialogue acts, we experiment with several well-known models such as sequence-to-sequence encoder-decoder model, hierarchical encoder-decoder model, and Siamese-based models with and without hierarchy; and show that in all models, incorporating dialogue acts improves the performance by a significant margin. We, furthermore, propose a novel way of encoding dialogue act information, and use it along with hierarchical encoder to build a model that can use the sequential dialogue act information in a natural way. Our proposed model achieves an MRR of about 84.8% for the task of next utterance selection on a newly introduced Daily Dialogue dataset, and outperform the baseline models. We also provide a detailed analysis of results including key insights that explain the improvement in MRR because of dialog act information.
Arvind Agarwal, Sachindra Joshi
COLING3
2017 Generating Natural Language Question-Answer Pairs from a Knowledge Graph Using a RNN Based Question Generation Model
abstract
Sathish Reddy, Dinesh Raghu, Mitesh M. Khapra, Sachindra Joshi. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Sathish Reddy, Dinesh Raghu, Mitesh M. Khapra, Sachindra Joshi
EACL (1)4
2017 Finding Dominant User Utterances And System Responses in Conversations
abstract
There are several dialog frameworks which allow manual specification of intents and rule based dialog flow. The rule based framework provides good control to dialog designers at the expense of being more time consuming and laborious. The job of a dialog designer can be reduced if we could identify pairs of user intents and corresponding responses automatically from prior conversations between users and agents. In this paper we propose an approach to find these frequent user utterances (which serve as examples for intents) and corresponding agent responses. We propose a novel SimCluster algorithm that extends standard K-means algorithm to simultaneously cluster user utterances and agent utterances by taking their adjacency information into account. The method also aligns these clusters to provide pairs of intents and response groups. We compare our results with those produced by using simple Kmeans clustering on a real dataset and observe upto 10% absolute improvement in F1-scores. Through our experiments on synthetic dataset, we show that our algorithm gains more advantage over K-means algorithm when the data has large variance.
Dhiraj Madan, Sachindra Joshi
IJCNLP(1)2
2017 Incomplete Follow-up Question Resolution using Retrieval based Sequence to Sequence Learning
abstract
Intelligent personal assistants (IPAs) and interactive question answering (IQA) systems frequently encounter incomplete follow-up questions. The incomplete follow-up questions only make sense when seen in conjunction with the conversation context: the previous question and answer. Thus, IQA and IPA systems need to utilize the conversation context in order to handle the incomplete follow-up questions and generate an appropriate response. In this work, we present a retrieval based sequence to sequence learning system that can generate the complete (or intended) question for an incomplete follow-up question (given the conversation context). We can train our system using only a small labeled dataset (with only a few thousand conversations), by decomposing the original problem into two simpler and independent problems. The first problem focuses solely on selecting the candidate complete questions from a library of question templates (built offline using the small labeled conversations dataset). In the second problem, we re-rank the selected candidate questions using a neural language model (trained on millions of unlabelled questions independently). Our system can achieve a BLEU score of 42.91, as compared to 29.11 using an existing generation based approach. We further demonstrate the utility of our system as a plug-in module to an existing QA pipeline. Our system when added as a plug-in module, enables Siri to achieve an improvement of 131.57% in answering incomplete follow-up questions.
Sachindra Joshi
SIGIR2
2016 Non-sentential Question Resolution using Sequence to Sequence Learning
abstract
An interactive Question Answering (QA) system frequently encounters non-sentential (incomplete) questions. These non-sentential questions may not make sense to the system when a user asks them without the context of conversation. The system thus needs to take into account the conversation context to process the question. In this work, we present a recurrent neural network (RNN) based encoder decoder network that can generate a complete (intended) question, given an incomplete question and conversation context. RNN encoder decoder networks have been show to work well when trained on a parallel corpus with millions of sentences, however it is extremely hard to obtain conversation data of this magnitude. We therefore propose to decompose the original problem into two separate simplified problems where each problem focuses on an abstraction. Specifically, we train a semantic sequence model to learn semantic patterns, and a syntactic sequence model to learn linguistic patterns. We further combine syntactic and semantic sequence models to generate an ensemble model. Our model achieves a BLEU score of 30.15 as compared to 18.54 using a standard RNN encoder decoder model.
Sachindra Joshi
COLING2
2015 A statistical approach for Non-Sentential Utterance Resolution for Interactive QA System
abstract
Non-Sentential Utterances (NSUs) are short utterances that do not have the form of a full sentence but nevertheless convey a complete sentential meaning in the context of a conversation.NSUs are frequently used to ask follow up questions during interactions with question answer (QA) systems resulting into in-correct answers being presented to their users.Most of the current methods for resolving such NSUs have adopted rule or grammar based approach and have limited applicability.In this paper, we present a data driven statistical method for resolving such NSUs.Our method is based on the observation that humans identify keyword appearing in an NSU and place them in the context of conversation to construct a meaningful sentence.We adapt the keyword to question (K2Q) framework to generate natural language questions using keywords appearing in an NSU and its context.The resulting questions are ranked using different scoring methods in a statistical framework.Our evaluation on a data-set collected using mTurk shows that the proposed method perform significantly better than the previous work that has largely been rule based.
Dinesh Raghu, Sathish Indurthi, Jitendra Ajmera, Sachindra Joshi
SIGDIAL Conference4
2014 Domain Cartridge: Unsupervised Framework for Shallow Domain Ontology Construction from Corpus
abstract
In this work we propose an unsupervised framework to construct a shallow domain ontology from corpus. It is essential for Information Retrieval systems, Question-Answering systems, Dialogue etc. to identify important concepts in the domain and the relationship between them. We identify important domain terms of which multi-words form an important component. We show that the incorporation of multi-words improves parser performance, resulting in better parser output, which improves the performance of an existing Question-Answering system by upto 7%. On manually annotated smartphone dataset, the proposed system identifies 40:87% of the domain terms, compared to 22% recall obtained using WordNet, 43:77% by Yago and 53:74% by BabelNet respectively. However, it does not use any manually annotated resource like the compared systems. Thereafter, we propose a framework to construct a shallow ontology from the discovered domain terms by identifying four domain relations namely, Synonyms ('similar-to'), Type-Of ('is-a'), Action-On ('methods') and Feature-Of ('attributes'), where we achieve significant performance improvement over WordNet, BabelNet and Yago without using any mode of supervision or manual annotation.
Subhabrata Mukherjee, Jitendra Ajmera, Sachindra Joshi
CIKM3
2014 Automatic generation of question answer pairs from noisy case logs
abstract
In a customer support scenario, a lot of valuable information is recorded in the form of `case logs'. Case logs are primarily written for future references or manual inspections and therefore are written in a hasty manner and are very noisy. In this paper, we propose techniques that exploit these case logs to mine real customer concerns or problems and then map them to well written knowledge articles for that enterprise. This mapping results into generation of question-answer (QA) pairs. These QA pairs can be used for a variety of applications such as dynamically updating the frequently-asked-questions (FAQs), updating the knowledge repository etc. In this paper we show the utility of these discovered QA pairs as training data for a question-answering system. Our approach for mining the case logs is based on a composite model consisting of two generative models, viz, hidden Markov model (HMM) and latent Dirichlet allocation (LDA) model. The LDA model explains the long-range dependencies across words due to their semantic similarity and HMM models the sequential patterns present in these case logs. Such processing results in crisp `problem statement' segments which are indicative of the real customer concerns. Our experiments show that this approach finds crisp problem-statements in 56% of the cases and outperforms other alternate methods for segmentation such as HMM, LDA and conditional random field (CRF). After finding these crisp problem-statements, appropriate answers are looked up from an existing knowledge repository index forming candidate QA pairs. We show that considering only the problemstatement segments for which the answers can be found further improves the segmentation performance to 82%. Finally, we show that when these QA pairs are used as training data, the performance of a question-answering system can be improved significantly.
Jitendra Ajmera, Sachindra Joshi, Ashish Verma 0001, Amol Mittal
ICDE2
2014 Author-Specific Sentiment Aggregation for Polarity Prediction of Reviews
Subhabrata Mukherjee, Sachindra Joshi
LREC2
2014 Joint Author Sentiment Topic Model
abstract
Traditional works in sentiment analysis and aspect rating prediction do not take author preferences and writing style into account during rating prediction of reviews. In this work, we introduce Joint Author Sentiment Topic Model (JAST), a generative process of writing a review by an author. Authors have different topic preferences, ‘emotional’ attachment to topics, writing style based on the distribution of semantic (topic) and syntactic (background) words and their tendency to switch topics. JAST uses Latent Dirichlet Allocation to learn the distribution of author-specific topic preferences and emotional attachment to topics. It uses a Hidden Markov Model to capture short range syntactic and long range semantic dependencies in reviews to capture coherence in author writing style. JAST jointly discovers the topics in a review, author preferences for the topics, topic ratings as well as the overall review rating from the point of view of an author. To the best of our knowledge, this is the first work in Natural Language Processing to bring all these dimensions together to have an author-specific generative model of a review.
Subhabrata Mukherjee, Gaurab Basu, Sachindra Joshi
SDM3
2013 Sentiment Aggregation using ConceptNet Ontology
Subhabrata Mukherjee, Sachindra Joshi
IJCNLP2
2012 Finding Influential Authors in Brand-Page Communities
Hemant Purohit, Jitendra Ajmera, Sachindra Joshi, Ashish Verma 0001, Amit P. Sheth
ICWSM3
2012 Spoken Document Clustering Using Word Confusion Networks
Shajith Ikbal, Sachindra Joshi, Ashish Verma 0001, Om Deshmukh
INTERSPEECH2
2012 Using Sequential Unconstrained Minimization Techniques to simplify SVM solvers
Sachindra Joshi, Jayadeva, Ganesh Ramakrishnan, Suresh Chandra 0001
Neurocomputing1
2012 Data and task parallelism in ILP using MapReduce
Ashwin Srinivasan 0001, Tanveer A. Faruquie, Sachindra Joshi
Mach. Learn.3
2012 Cross-Guided Clustering: Transfer of Relevant Supervision across Tasks
abstract
Lack of supervision in clustering algorithms often leads to clusters that are not useful or interesting to human reviewers. We investigate if supervision can be automatically transferred for clustering a target task, by providing a relevant supervised partitioning of a dataset from a different source task. The target clustering is made more meaningful for the human user by trading-off intrinsic clustering goodness on the target task for alignment with relevant supervised partitions in the source task, wherever possible. We propose a cross-guided clustering algorithm that builds on traditional k-means by aligning the target clusters with source partitions. The alignment process makes use of a cross-task similarity measure that discovers hidden relationships across tasks. When the source and target tasks correspond to different domains with potentially different vocabularies, we propose a projection approach using pivot vocabularies for the cross-domain similarity measure. Using multiple real-world and synthetic datasets, we show that our approach improves clustering accuracy significantly over traditional k-means and state-of-the-art semi-supervised clustering baselines, over a wide range of data characteristics and parameter settings.
Indrajit Bhattacharya, Shantanu Godbole, Sachindra Joshi, Ashish Verma 0001
ACM Trans. Knowl. Discov. Data3
2011 Labeling Unlabeled Data using Cross-Language Guided Clustering
Sachindra Joshi, Danish Contractor, Sumit Negi
IJCNLP1
2011 Auto-Grouping Emails For Faster E-Discovery
Sachindra Joshi, Danish Contractor, Kenney Ng, Prasad Deshpande, Thomas Hampp
Proc. VLDB Endow.1
2009 Cross-Guided Clustering: Transfer of Relevant Supervision across Domains for Improved Clustering
abstract
Lack of supervision in clustering algorithms often leads to clusters that are not useful or interesting to human reviewers. We investigate if supervision can be automatically transferred to a clustering task in a target domain, by providing a relevant supervised partitioning of a dataset from a different source domain. The target clustering is made more meaningful for the human user by trading off intrinsic clustering goodness on the target dataset for alignment with relevant supervised partitions in the source dataset, wherever possible. We propose a cross-guided clustering algorithm that builds on traditional k-means by aligning the target clusters with source partitions. The alignment process makes use of a cross-domain similarity measure that discovers hidden relationships across domains with potentially different vocabularies. Using multiple real-world datasets, we show that our approach improves clustering accuracy significantly over traditional k-means.
Indrajit Bhattacharya, Shantanu Godbole, Sachindra Joshi, Ashish Verma 0001
ICDM3
2009 Automatically Extracting Dialog Models from Conversation Transcripts
abstract
There is a growing need for task-oriented natural language dialog systems that can interact with a user to accomplish a given objective. Recent work on building task-oriented dialog systems have emphasized the need for acquiring task-specific knowledge from un-annotated conversational data. In our work we acquire task-specific knowledge by defining sub-task as the key unit of a task-oriented conversation. We propose an unsupervised, apriori like algorithm that extracts the sub-tasks and their valid orderings from un-annotated human-human conversations. Modeling dialogues as a combination of sub-tasks and their valid orderings easily captures the variability in conversations. It also provides us the ability to map our dialogue model to AIML constructs and therefore use off-the-shelf AIML interpreters to build task-oriented chat-bots. We conduct experiments on real world data sets to establish the effectiveness of the sub-task extraction process. We codify the extracted sub-tasks in an AIML knowledge base and build a chatbot using this knowledge base. We also show the usefulness of the chatbot in automatically handling customer requests by performing a user evaluation study.
Sumit Negi, Sachindra Joshi, Anup Chalamalla, L. Venkata Subramaniam
ICDM2
2009 An investigation into feature construction to assist word sense disambiguation
Lucia Specia, Ashwin Srinivasan 0001, Sachindra Joshi, Ganesh Ramakrishnan, Maria das Graças Volpe Nunes
Mach. Learn.3
2008 RAD: A Scalable Framework for Annotator Development
abstract
Developments in semantic search technology have motivated the need for efficient and scalable entity annotation techniques. We demonstrate RAD: a tool for Rapid Annotator Development on a document collection. RAD builds on a recent approach (Ramakrishnan et al., 2006) that translates entity annotation rules into equivalent operations on the inverted index of the collection, to directly generate an annotation index (which can be used in search applications). To make the framework scalable, we use an industrial strength indexer, Lucene (http://lucene.apache.org) and introduce some modifications to its API. The index also serves as a suitable representation for making quick comparisons with an indexed ground truth of annotations on the same collection to evaluate precision and recall of the annotations. RAD achieves at least an order of magnitude speedup over the standard approach of annotating a document-at-a-time as adopted by GATE (Cunnignham et al., 2002). The speedup factor increases with increase in the size of the collection, making RAD scalable. We cache intermediate results from the index operations, enabling quick update of the annotation index as well as speedy evaluation when rules are modified. This makes RAD suitable for rapid and interactive development of annotators.
Sanjeet Khaitan, Ganesh Ramakrishnan, Sachindra Joshi, Anup Chalamalla
ICDE3
2008 Learning Decision Lists with Known Rules for Text Mining
Venkatesan Chakravarthy, Sachindra Joshi, Ganesh Ramakrishnan, Shantanu Godbole, Sreeram Balakrishnan
IJCNLP2
2008 Feature Construction Using Theory-Guided Sampling and Randomised Search
Sachindra Joshi, Ganesh Ramakrishnan, Ashwin Srinivasan 0001
ILP1
2008 Automatic pronunciation evaluation and classification
Om Deshmukh, Sachindra Joshi, Ashish Verma 0001
INTERSPEECH2
2008 Structured entity identification and document categorization: two tasks with one joint model
abstract
Traditionally, research in identifying structured entities in documents has proceeded independently of document categorization research. In this paper, we observe that these two tasks have much to gain from each other. Apart from direct references to entities in a database, such as names of person entities, documents often also contain words that are correlated with discriminative entity attributes, such age-group and income-level of persons. This happens naturally in many enterprise domains such as CRM, Banking, etc. Then, entity identification, which is typically vulnerable against noise and incompleteness in direct references to entities in documents, can benefit from document categorization with respect to such attributes. In return, entity identification enables documents to be categorized according to different label-sets arising from entity attributes without requiring any supervision. In this paper, we propose a probabilistic generative model for joint entity identification and document categorization. We show how the parameters of the model can be estimated using an EM algorithm in an unsupervised fashion. Using extensive experiments over real and semi-synthetic data, we demonstrate that the two tasks can benefit immensely from each other when performed jointly using the proposed model.
Indrajit Bhattacharya, Shantanu Godbole, Sachindra Joshi
KDD3
2007 Using ILP to Construct Features for Information Extraction from Semi-structured Text
Ganesh Ramakrishnan, Sachindra Joshi, Sreeram Balakrishnan, Ashwin Srinivasan 0001
ILP2
2006 Search result summarization and disambiguation via contextual dimensions
abstract
Topic hierarchies are a popular method of summarizing the results obtained in response to a query in various search applications. However, topic hierarchies are rigid when they are pre-defined and somewhat unintuitive when they are dynamically generated by statistical techniques. In this paper, we propose an alternative approach to query disambiguation and result summarization by placing the results in set of contextual dimensions which can be viewed as facets. For the generic search scenario, we illustrate our approach by using three types of contextual dimensions, namely, concepts, features, and specializations. We use NLP techniques and a data mining algorithm to select distinct contexts.
Krishna Prasad Chitrapura, Sachindra Joshi, Raghu Krishnapuram
CIKM2
2006 Entity Annotation based on Inverse Index Operations
Ganesh Ramakrishnan, Sreeram Balakrishnan, Sachindra Joshi
EMNLP3
2006 Automatic Sales Lead Generation from Web Data
abstract
Speed to market is critical to companies that are driven by sales in a competitive market. The earlier a potential customer can be approached in the decision making process of a purchase, the higher are the chances of converting that prospect into a customer. Traditional methods to identify sales leads such as company surveys and direct marketing are manual, expensive and not scalable. Over the past decade the World Wide Web has grown into an information-mesh, with most important facts being reported through Web sites. Several news papers, press releases, trade journals, business magazines and other related sources are on-line. These sources could be used to identify prospective buyers automatically. In this paper, we present a system called ETAP (Electronic Trigger Alert Program) that extracts trigger events from Web data that help in identifying prospective buyers. Trigger events are events of corporate relevance and indicative of the propensity of companies to purchase new products associated with these events. Examples of trigger events are change in management, revenue growth and mergers & acquisitions. The unstructured nature of information makes the extraction task of trigger events difficult. We pose the problem of trigger events extraction as a classification problem and develop methods for learning trigger event classifiers using existing classification methods. We present methods to automatically generate the training data required to learn the classifiers. We also propose a method of feature abstraction that uses named entity recognition to solve the problem of data sparsity. We score and rank the trigger events extracted from ETAP for easy browsing. Our experiments show the effectiveness of the method and thus establish the feasibility of automatic sales lead generation using the Web data.
Ganesh Ramakrishnan, Sachindra Joshi, Sumit Negi, Raghu Krishnapuram, Sreeram Balakrishnan
ICDE2
2004 EShopMonitor: A Web Content Monitoring Tool
abstract
Data presented on commerce sites runs into thousands of pages, and is typically delivered from multiple back-end sources. This makes it difficult to identify incorrect, anomalous, or interesting data such as $9.99 air fares, missing links, drastic changes in prices and addition of new products or promotions. We describe a system that monitors Web sites automatically and generates various types of reports so that the content of the site can be monitored and the quality maintained. The solution designed and implemented by us consists of a site crawler that crawls dynamic pages, an information miner that learns to extract useful information from the pages based on examples provided by the user, and a reporter that can be configured by the user to answer specific queries. The tool can also be used for identifying price trends and new products or promotions at competitor sites. A pilot run of the tool has been successfully completed at the ibm.com site.
Neeraj Agrawal, Rema Ananthanarayanan, Sachindra Joshi, Raghu Krishnapuram, Sumit Negi
ICDE4
2003 Classification of Text Documents Based on Minimum System Entropy
Raghu Krishnapuram, Krishna Prasad Chitrapura, Sachindra Joshi
ICML3
2003 A bag of paths model for measuring structural similarity in Web documents
abstract
Structural information (such as layout and look-and-feel) has been extensively used in the literatuce for extraction of interesting or relevant data, efficient storage, and query optimization. Traditionally, tree models (such as DOM trees) have been used to represent structural information, especially in the case of HTML and XML documents. However, computation of structural similarity between documents based on the tree model is computationally expensive. In this paper, we propose an alternative scheme for representing the structural information of documents based on the paths contained in the corresponding tree model. Since the model includes partial information about parents, children and siblings, it allows us to define a new family of meaningful (and at the same time computationally simple) structural similarity measures. Our experimental results based on the SIGMOD XML data set as well as HTML document collections from ibm.com, dell.com, and amazon.com show that the representation is powerful enough to produce good clusters of structurally similar pages.
Sachindra Joshi, Neeraj Agrawal, Raghu Krishnapuram, Sumit Negi
KDD1
2003 A matrix density based algorithm to hierarchically co-cluster documents and words
abstract
This paper proposes an algorithm to hierarchically cluster documents. Each cluster is actually a cluster of documents and an associated cluster of words, thus a document-word co-cluster. Note that, the vector model for documents creates the document-word matrix, of which every co-cluster is a submatrix. One would intuitively expect a submatrix made up of high values to be a good document cluster, with the corresponding word cluster containing its most distinctive features. Our algorithm looks to exploit this. We have defined matrix density, and our algorithm basically uses matrix density considerations in its working.The algorithm is a partitional-agglomerative algorithm. The partitioning step involves the identification of dense submatrices so that the respective row sets partition the row set of the complete matrix. The hierarchical agglomerative step involves merging the most similar submatrices until we are down to the required number of clusters (if we want a flat clustering) or until we have just the single complete matrix left (if we are interested in a hierarchical arrangement of documents). It also generates apt labels for each cluster or hierarchy node. The similarity measure between clusters that we use here for the merging cleverly uses the fact that the clusters here are co-clusters, and is a key point of difference from existing agglomerative algorithms. We will refer to the proposed algorithm as RPSA (Rowset Partitioning and Submatrix Agglomeration). We have compared it as a clustering algorithm with Spherical K-Means and Spectral Graph Partitioning. We have also evaluated some hierarchies generated by the algorithm.
Bhushan Mandhani, Sachindra Joshi, Krishna Kummamuru
WWW2
2001 Mining Generalised Disjunctive Association Rules
abstract
This paper introduces generalised disjunctive association rules such as "People who buy bread also buy butter jam", and "People who buy either raincoats or umbrellas also buy flashlights". A generalised disjunctive association rule allows the disjunction of conjuncts, "People who buy jackets also buy bow ties or neckties and tiepins". Such rules capture contextual inter-relationships among items.Given a context (antecedent), there may be a large number of generalised disjunctive association rules that satisfy the minsupp and minconf constraints. It is computationally expensive to find all such rules. We present algorithm thrifty traverse which borrows concepts such as subsumption from propositional logic to mine a subset of such rules in a computationally feasible way. We experimented with our algorithm on US census data as well as transaction data from a grocery superstore to demonstrate its computational feasibility, utility and scalability.
Amit Anil Nanavati, Krishna Prasad Chitrapura, Sachindra Joshi, Raghu Krishnapuram
CIKM3