VLDB 2026 Research / reviewers in the wild / expert
Deepak P 0001
dblp:33/1882 · also Deepak Padmanabhan 0001, P. Deepak 0001
· DBLP profile ↗
66ranked-venue papers
21as first author
17since 2021 · last 2024
0000-0002-1336-2356ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 41 · 13 first-author · 8 since 2021Artificial intelligence and machine learning · 28 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Fairer Centroids in K-means ClusteringabstractThere has been much recent interest in developing fair clustering algorithms that seek to do justice to the representation of groups defined along sensitive attributes such as race and sex. Within the centroid clustering paradigm, these algorithms are seen to generate clusterings where different groups are disadvantaged within different clusters with respect to their representativity, i.e., distance to centroid. In view of this deficiency, we propose a novel notion of cluster-level centroid fairness that targets the representativity unfairness borne by groups within each cluster, along with a metric to quantify the same. Towards operationalising this notion, we draw on ideas from political philosophy aligned with consideration for the worst-off group to develop Fair-Centroid; a new clustering method that focusses on enhancing the representativity of the worst-off group within each cluster. Our method uses an iterative optimisation paradigm wherein an initial cluster assignment is refined by reassigning objects to clusters such that the worst-off group in each cluster is benefitted. We compare our notion with a related fairness notion and show through extensive empirical evaluations on real-world datasets that our method significantly enhances cluster-level centroid fairness at low impact on cluster coherence. Stanley Simoes, Deepak P 0001, Muiris MacCarthaigh |
AAAI | 2 |
| 2024 | REDAffectiveLM: leveraging affect enriched embedding and transformer-based neural language model for readers' emotion detection
Anoop Kadan, Deepak P 0001, Manjary P. Gangan, Savitha Sam Abraham, V. L. Lajish |
Knowl. Inf. Syst. | 2 |
| 2023 | Group Fairness in Case-Based Reasoning
Shania Mitra, Ditty Mathew, Deepak P 0001, Sutanu Chakraborti |
ICCBR | 3 |
| 2023 | The Case for Circularities in Case-Based Reasoning
Adwait P. Parsodkar, Deepak P 0001, Sutanu Chakraborti |
ICCBR | 2 |
| 2023 | FiSH: fair spatial hot spotsabstractAbstract Pervasiveness of tracking devices and enhanced availability of spatially located data has deepened interest in using them for various policy interventions, through computational data analysis tasks such as spatial hot spot detection. In this paper, we consider, for the first time to our best knowledge, fairness in detecting spatial hot spots. We motivate the need for ensuring fairness through statistical parity over the collective population covered across chosen hot spots. We then characterize the task of identifying a diverse set of solutions in the noteworthiness-fairness trade-off spectrum, to empower the user to choose a trade-off justified by the policy domain. Being a novel task formulation, we also develop a suite of evaluation metrics for fair hot spots, motivated by the need to evaluate pertinent aspects of the task. We illustrate the computational infeasibility of identifying fair hot spots using naive and/or direct approaches and devise a method, codenamed FiSH, for efficiently identifying high-quality, fair and diverse sets of spatial hot spots. FiSH traverses the tree-structured search space using heuristics that guide it towards identifying noteworthy and fair sets of spatial hot spots. Through an extensive empirical analysis over a real-world dataset from the domain of human development, we illustrate that FiSH generates high-quality solutions at fast response times. Towards assessing the relevance of FiSH in real-world context, we also provide a detailed discussion of how it could fit within the current practice of hot spots policing, as read within the historical context of the evolution of the practice. Deepak P 0001, Sowmya S. Sundaram |
Data Min. Knowl. Discov. | 1 |
| 2023 | On Efficient Large Maximal Biplex DiscoveryabstractCohesive subgraph discovery is an important problem in bipartite graph mining. In this paper, we focus on one kind of cohesive structure, called k-biplex, where each vertex of one side is disconnected from at most k vertices of the other side. We consider the large maximal k-biplex enumeration problem which is to list all those maximal k-biplexes with the number of vertices at each side at least a non-negative integer . This formulation, we observe, has various applications and targets to find non-redundant results by excluding non-maximal ones. Existing approaches suffer from massive redundant computations and can only run on small and moderate datasets. Towards improving scalability, we propose an efficient tree-based algorithm with two advanced strategies and powerful pruning techniques. Experimental results on real and synthetic datasets show the superiority of our algorithm over existing approaches. Kaiqiang Yu, Cheng Long 0001, Deepak P 0001, Tanmoy Chakraborty 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Deep Extreme Mixture Model for Time Series ForecastingabstractTime Series Forecasting (TSF) has been a topic of extensive research, which has many real world applications such as weather prediction, stock market value prediction, traffic control etc. Many machine learning models have been developed to address TSF, yet, predicting extreme values remains a challenge to be effectively addressed. Extreme events occur rarely, but tend to cause a huge impact, which makes extreme event prediction important. Assuming light tailed distributions, such as Gaussian distribution, on time series data does not do justice to the modeling of extreme points. To tackle this issue, we develop a novel approach towards improving attention to extreme event prediction. Within our work, we model time series data distribution, as a mixture of Gaussian distribution and Generalized Pareto distribution (GPD). In particular, we develop a novel Deep eXtreme Mixture Model (DXtreMM) for univariate time series forecasting, which addresses extreme events in time series. The model consists of two modules: 1) Variational Disentangled Auto-encoder (VD-AE) based classifier and 2) Multi Layer Perceptron (MLP) based forecaster units combined with Generalized Pareto Distribution (GPD) estimators for lower and upper extreme values separately. VD-AE Classifier model predicts the possibility of occurrence of an extreme event given a time segment, and forecaster module predicts the exact value. Through extensive set of experiments on real-world datasets we have shown that our model performs well for extreme events and is comparable with the existing baseline methods for normal time step forecasting. Abilasha S, Sahely Bhadra, Ahmed Zaheer Dadarkar, Deepak P 0001 |
CIKM | 4 |
| 2022 | Never Judge a Case by Its (Unreliable) Neighbors: Estimating Case Reliability for CBR
Adwait P. Parsodkar, Deepak P 0001, Sutanu Chakraborti |
ICCBR | 2 |
| 2022 | On Efficient Large Maximal Biplex Discovery (Extended abstract)abstractCohesive subgraph discovery is an important problem in bipartite graph mining. In this paper, we focus on one kind of cohesive structure, called$k$-biplex, where each vertex of one side is disconnected from at most$k$vertices of the other side. We consider the large maximal$k$-biplex enumeration problem which is to list all those maximal$k$-biplexes with the number of vertices at each side at least a non-negative integer$\theta$. This formulation aims to find non-redundant results by excluding non-maximal ones and has various applications. Existing approaches suffer from massive redundant computations and can only run on small and moderate datasets. Towards improving scalability, we propose an efficient tree-based algorithm with two advanced strategies and powerful pruning techniques. Experimental results show the superiority of our algorithm over existing approaches. Kaiqiang Yu, Cheng Long 0001, Deepak P 0001, Tanmoy Chakraborty 0002 |
ICDE | 3 |
| 2022 | Span Detection for Kinematics Word Problems
Savitha Sam Abraham, Deepak P 0001, Sowmya S. Sundaram |
ICONIP (6) | 2 |
| 2022 | Warping resilient scalable anomaly detection in time series
Abilasha S, Sahely Bhadra, Deepak P 0001, Anish Mathew |
Neurocomputing | 3 |
| 2021 | Cross-modal Data Linkage for Common Entity Identification
Pragya Prakash, Jay Rawal, Snehal Gupta, Deepak P 0001, Mukesh K. Mohania |
ADMA | 4 |
| 2021 | Revisiting Fast and Slow Thinking in Case-Based Reasoning
Srashti Kaurav, Devi Ganesan, Deepak P 0001, Sutanu Chakraborti |
ICCBR | 3 |
| 2021 | Towards Richer Realizations of Holographic CBR
Renganathan Subramanian, Devi Ganesan, Deepak P 0001, Sutanu Chakraborti |
ICCBR | 3 |
| 2021 | Unsupervised Keyword Combination Query Generation from Online Health Related Content for Evidence-Based Fact CheckingabstractFalse information in the domain of online health related articles is of great concern, which can be witnessed in the current pandemic situation of Covid-19. It is markedly different from fake news in the political context as health information should be evaluated against the most recent and reliable medical resources such as scholarly repositories. However, one of the challenges with such an approach is the retrieval of the pertinent resources. In this work, we formulate a new unsupervised task of generating queries using keywords extracted from a health-related article which can be further applied to retrieve relevant authoritative and reliable medical content from scholarly repositories to assess the article’s veracity. We propose a three-step approach for it and illustrate that our method is able to generate effective queries. We also curate a new dataset to aid the evaluation for this task which will be made available upon request. Pritam Deka, Anna Jurek-Loughrey, Deepak P 0001 |
iiWAS | 3 |
| 2021 | FairLOF: Fairness in Outlier DetectionabstractAbstract An outlier detection method may be considered fair over specified sensitive attributes if the results of outlier detection are not skewed toward particular groups defined on such sensitive attributes. In this paper, we consider the task of fair outlier detection. Our focus is on the task of fair outlier detection over multiple multi-valued sensitive attributes (e.g., gender, race, religion, nationality and marital status, among others), one that has broad applications across modern data scenarios. We propose a fair outlier detection method,FairLOF, that is inspired by the popularLOFformulation for neighborhood-based outlier detection. We outline ways in which unfairness could be induced withinLOFand develop three heuristic principles to enhance fairness, which form the basis of theFairLOFmethod. Being a novel task, we develop an evaluation framework for fair outlier detection, and use that to benchmarkFairLOFon quality and fairness of results. Through an extensive empirical evaluation over real-world datasets, we illustrate thatFairLOFis able to achieve significant improvements in fairness at sometimes marginal degradations on result quality as measured against the fairness-agnosticLOFmethod. We also show that a generalization of our method, namedFairLOF-Flex, is able to open possibilities of further deepening fairness in outlier detection beyond what is offered byFairLOF. Deepak P 0001, Savitha Sam Abraham |
Data Sci. Eng. | 1 |
| 2021 | Emotion-aware polarity lexicons for Twitter sentiment analysisabstractAbstract Theoretical frameworks in psychology map the relationships between emotions and sentiments. In this paper, we study the role of such mapping for computational emotion detection from text (e.g., social media) with an aim to understand the usefulness of an emotion‐rich corpus of documents (e.g., tweets) to learn polarity lexicons for sentiment analysis. We propose two different methods that leverage a corpus of emotion‐labelled tweets to learn word‐polarity lexicons. The proposed methods model the emotion corpus using a generative unigram mixture model, combined with the emotion‐sentiment mapping proposed in psychology for automated generation of word‐polarity lexicons that capture emotion‐rich vocabulary. We comparatively evaluate the quality of the proposed mixture model in learning emotion‐aware sentiment lexicons with those generated using supervised latent dirichlet allocation (sLDA) and word‐document‐frequency (WDF) statistics. Sentiment analysis experiments on benchmark Twitter data sets confirm the quality of our proposed lexicons. Further, a comparative analysis with sLDA, WDF‐based emotion‐aware lexicons, and standard sentiment lexicons that are agnostic to emotion knowledge suggests that the proposed lexicons lead to a significantly better performance in both sentiment classification and sentiment intensity prediction tasks. Anil Bandhakavi, Nirmalie Wiratunga, Stewart Massie, Deepak P 0001 |
Expert Syst. J. Knowl. Eng. | 4 |
| 2020 | Distributed Representations for Arithmetic Word ProblemsabstractWe consider the task of learning distributed representations for arithmetic word problems. We outline the characteristics of the domain of arithmetic word problems that make generic text embedding methods inadequate, necessitating a specialized representation learning method to facilitate the task of retrieval across a wide range of use cases within online learning platforms. Our contribution is two-fold; first, we propose several 'operators' that distil knowledge of the domain of arithmetic word problems and schemas into word problem transformations. Second, we propose a novel neural architecture that combines LSTMs with graph convolutional networks to leverage word problems and their operator-transformed versions to learn distributed representations for word problems. While our target is to ensure that the distributed representations are schema-aligned, we do not make use of schema labels in the learning process, thus yielding an unsupervised representation learning method. Through an evaluation on retrieval over a publicly available corpus of word problems, we illustrate that our framework is able to consistently improve upon contemporary generic text embeddings in terms of schema-alignment. Sowmya S. Sundaram, Deepak P 0001, Savitha Sam Abraham |
AAAI | 2 |
| 2020 | Fairness in Unsupervised LearningabstractData in digital form is expanding at an exponential rate, far outpacing any chance of getting any significant fraction labelled manually. This has resulted in heightened research emphasis on unsupervised learning, learning in the absence of labels. In fact, unsupervised learning has been often dubbed as the next frontier of AI. Unsupervised learning is the most plausible model to analyze the bulk of passively collected data that spans across various domains; e.g., social media footprints, safety/surveilance cameras, IoT devices, sensors, smartphone apps, medical wearables, traffic sensing devices and public wi-fi access. While fairness in supervised learning, such as classification tasks, has inspired a large amount of research in the past few years, work on fair unsupervised learning has been relatively slow in picking up. This tutorial targets to provide an overview of: (i) fairness issues in unsupervised learning drawing abundantly from political philosophy, (ii) current research in fair unsupervised learning, and (iii) new directions to extend the state-of-the-art in fair unsupervised learning. While we intend to broadly cover all tasks in unsupervised learning, our focus will be on clustering, retrieval and representation learning. In a unique departure from conventional data science tutorials, we will place significant emphasis on presenting and debating pertinent literature from ethics and philosophy. Overall, this half-day tutorial brings a strong emphasis on ensuring strong interdisciplinarity. Deepak P 0001, Joemon M. Jose, Sanil V |
CIKM | 1 |
| 2020 | Fairness in Clustering with Multiple Sensitive Attributes
Savitha Sam Abraham, Deepak P 0001, Sowmya S. Sundaram |
EDBT | 2 |
| 2020 | Local connectivity in centroid clusteringabstractClustering is a fundamental task in unsupervised learning, one that targets to group a dataset into clusters of similar objects. There has been recent interest in embedding normative considerations around fairness within clustering formulations. In this paper, we propose 'local connectivity' as a crucial factor in assessing membership desert in centroid clustering. We use local connectivity to refer to the support offered by the local neighborhood of an object towards supporting its membership to the cluster in question. We motivate the need to consider local connectivity of objects in cluster assignment, and provide ways to quantify local connectivity in a given clustering. We then exploit concepts from density-based clustering and devise LOFKM, a clustering method that seeks to deepen local connectivity in clustering outputs, while staying within the framework of centroid clustering. Through an empirical evaluation over real-world datasets, we illustrate that LOFKM achieves notable improvements in local connectivity at reasonable costs to clustering quality, illustrating the effectiveness of the method. Deepak P 0001 |
IDEAS | 1 |
| 2020 | Emotion cognizance improves health fake news identificationabstractIdentifying fake news is increasingly being recognized as an important computational task with high potential social impact. Misinformation is routinely injected into almost every domain of news including politics, health, science, business, etc., among which, the fake news in the health domain poses serious risk and harm to health and well-being in modern societies. In this paper, we consider the utility of the affective character of news articles for fake news identification in the health domain and present evidence that emotion cognizant representations are significantly more suited for the task. We outline a simple technique that works by leveraging emotion intensity lexicons to develop emotion-amplified text representations and evaluate the utility of such a representation for identifying fake news relating to health in various supervised and unsupervised scenarios. The consistent and notable empirical gains that we observe over a range of technique types and parameter settings establish the utility of the emotional information in news articles, an often overlooked aspect, for the task of misinformation identification in the health domain. Anoop Kadan, Deepak P 0001, V. L. Lajish |
IDEAS | 2 |
| 2020 | ReSCo-CC: Unsupervised Identification of Key Disinformation SentencesabstractDisinformation is often presented in long textual articles, especially when it relates to domains such as health, often seen in relation to COVID-19. These articles are typically observed to have a number of trustworthy sentences among which core disinformation sentences are scattered. In this paper, we propose a novel unsupervised task of identifying sentences containing key disinformation within a document that is known to be untrustworthy. We design a three-phase statistical NLP solution for the task which starts with embedding sentences within a bespoke feature space designed for the task. Sentences represented using those features are then clustered, following which the key sentences are identified through proximity scoring. We also curate a new dataset with sentence level disinformation scorings to aid evaluation for this task; the dataset is being made publicly available to facilitate further research. Based on a comprehensive empirical evaluation against techniques from related tasks such as claim detection and summarization, as well as against simplified variants of our proposed approach, we illustrate that our method is able to identify core disinformation effectively. Soumya Suvra Ghosal, Deepak P 0001, Anna Jurek-Loughrey |
iiWAS | 2 |
| 2020 | Modeling Implicit Communities from Geo-Tagged Event Traces Using Spatio-Temporal Point Processes
Ankita Likhyani, P. K. Srijith, Deepak P 0001, Srikanta J. Bedathur |
WISE (1) | 4 |
| 2020 | Fair Outlier Detection
Deepak P 0001, Savitha Sam Abraham |
WISE (2) | 1 |
| 2019 | Location-Specific Influence Quantification in Location-Based Social NetworksabstractLocation-based social networks (LBSNs) such as Foursquare offer a platform for users to share and be aware of each other’s physical movements. As a result of such a sharing of check-in information with each other, users can be influenced to visit (or check-in) at the locations visited by their friends. Quantifying such influences in these LBSNs is useful in various settings such as location promotion, personalized recommendations, mobility pattern prediction, and so forth. In this article, we develop a model to quantify the influence specific to a location between a pair of users. Specifically, we develop a framework called LoCaTe , that combines (a) a user mobility model based on kernel density estimates; (b) a model of the semantics of the location using topic models; and (c) a user correlation model that uses an exponential distribution. We further develop LoCaTe+ , an advanced model within the same framework where user correlation is quantified using a Mutually Exciting Hawkes Process. We show the applicability of LoCaTe and LoCaTe+ for location promotion and location recommendation tasks using LBSNs. Our models are validated using a long-term crawl of Foursquare data collected between January 2015 and February 2016, as well as other publicly available LBSN datasets. Our experiments demonstrate the efficacy of the LoCaTe framework in capturing location-specific influence between users. We also show that our models improve over state-of-the-art models for the task of location promotion as well as location recommendation. Ankita Likhyani, Srikanta J. Bedathur, Deepak P 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2018 | Content and Context: Two-Pronged Bootstrapped Learning for Regex-Formatted Entity ExtractionabstractRegular expressions are an important building block of rule-based information extraction systems. Regexes can encode rules to recognize instances of simple entities which can then feed into the identification of more complex cross-entity relationships. Manually crafting a regex that recognizes all possible instances of an entity is difficult since an entity can manifest in a variety of different forms. Thus, the problem of automatically generalizing manually crafted seed regexes to improve the recall of IE systems has attracted research attention. In this paper, we propose a bootstrapped approach to improve the recall for extraction of regex-formatted entities, with the only source of supervision being the seed regex. Our approach starts from a manually authored high precision seed regex for the entity of interest, and uses the matches of the seed regex and the context around these matches to identify more instances of the entity. These are then used to identify a set of diverse, high recall regexes that are representative of this entity. Through an empirical evaluation over multiple real world document corpora, we illustrate the effectiveness of our approach. Stanley Simoes, Deepak P 0001, Munu Sairamesh, Deepak Khemani, Sameep Mehta |
AAAI | 2 |
| 2018 | Fast Identification of Interesting Spatial Regions with Applications in Human Development Research
Carl Duffy, Deepak P 0001, Cheng Long 0001, M. Satish Kumar, Amit Thorat, Amaresh Dubey |
DEXA (2) | 2 |
| 2018 | It Pays to Be Certain: Unsupervised Record Linkage via Ambiguity Minimization
Anna Jurek-Loughrey, Deepak P 0001 |
PAKDD (3) | 2 |
| 2018 | Leveraging semantic resources in diversified query expansionabstractA search query, being a very concise grounding of user intent, could potentially have many possible interpretations. Search engines hedge their bets by diversifying top results to cover multiple such possibilities so that the user is likely to be satisfied, whatever be her intended interpretation. Diversified Query Expansion is the problem of diversifying query expansion suggestions, so that the user can specialize the query to better suit her intent, even before perusing search results. In this paper, we consider the usage of semantic resources and tools to arrive at improved methods for diversified query expansion. In particular, we develop two methods, those that leverage Wikipedia and pre-learnt distributional word embeddings respectively. Both the approaches operate on a common three-phase framework; that of first taking a set of informative terms from the search results of the initial query, then building a graph, following by using a diversity-conscious node ranking to prioritize candidate terms for diversified query expansion. Our methods differ in the second phase, with the first method Select-Link-Rank (SLR) linking terms with Wikipedia entities to accomplish graph construction; on the other hand, our second method, Select-Embed-Rank (SER), constructs the graph using similarities between distributional word embeddings. Through an empirical analysis and user study, we show that SLR ourperforms state-of-the-art diversified query expansion methods, thus establishing that Wikipedia is an effective resource to aid diversified query expansion. Our empirical analysis also illustrates that SER outperforms the baselines convincingly, asserting that it is the best available method for those cases where SLR is not applicable; these include narrow-focus search systems where a relevant knowledge base is unavailable. Our SLR method is also seen to outperform a state-of-the-art method in the task of diversified entity ranking. Adit Krishnan, Deepak P 0001, Sayan Ranu, Sameep Mehta |
World Wide Web | 2 |
| 2017 | LTRo: Learning to Route Queries in Clustered P2P IR
Rami Suleiman Alkhawaldeh, Deepak P 0001, Joemon M. Jose, Fajie Yuan |
ECIR | 2 |
| 2017 | Latent Space Embedding for Retrieval in Question-Answer ArchivesabstractCommunity-driven Question Answering (CQA) systems such as Yahoo!Answers have become valuable sources of reusable information.CQA retrieval enables usage of historical CQA archives to solve new questions posed by users.This task has received much recent attention, with methods building upon literature from translation models, topic models, and deep learning.In this paper, we devise a CQA retrieval technique, LASER-QA, that embeds question-answer pairs within a unified latent space preserving the local neighborhood structure of question and answer spaces.The idea is that such a space mirrors semantic similarity among questions as well as answers, thereby enabling high quality retrieval.Through an empirical analysis on various real-world QA datasets, we illustrate the improved effectiveness of LASER-QA over state-of-theart methods. Deepak P 0001, Dinesh Garg, Shirish K. Shevade |
EMNLP | 1 |
| 2017 | LoCaTe: Influence Quantification for Location Promotion in Location-based Social NetworksabstractLocation-based social networks (LBSNs) such as Foursquare offer a platform for users to share and be aware of each other’s physical movements. As a result of such a sharing of check-in information with each other, users can be influenced to visit (or check-in) at the locations visited by their friends. Quantifying such influences in these LBSNs is useful in various settings such as location promotion, personalized recommendations, mobility pattern prediction etc. In this paper, we focus on the problem of location promotion and develop a model to quantify the influence specific to a location between a pair of users. Specifically, we develop a joint model called LoCaTe, consisting of (i) user mobility model estimated using kernel density estimates; (ii) a model of the semantics of the location using topic models; and (iii) a model of time-gap between check-ins using exponential distribution. We validate our model on a long-term crawl of Foursquare data collected between Jan 2015 Feb 2016, as well as on publicly available LBSN datasets. Our experiments demonstrate that LoCaTe significantly outperforms state-of-the-art models for the same task. Ankita Likhyani, Srikanta J. Bedathur, Deepak P 0001 |
IJCAI | 3 |
| 2017 | Lexicon based feature extraction for emotion text classification
Anil Bandhakavi, Nirmalie Wiratunga, Deepak P 0001, Stewart Massie |
Pattern Recognit. Lett. | 3 |
| 2016 | Evaluating Document Retrieval Methods for Resource Selection in Clustered P2P IRabstractResource Selection (or Query Routing) is an important step in P2P IR. Though analogous to document retrieval in the sense of choosing a relevant subset of resources, resource selection methods have evolved independently from those for document retrieval. Among the reasons for such divergence is that document retrieval targets scenarios where underlying resources are semantically homogeneous, whereas peers would manage diverse content. We observe that semantic heterogeneity is mitigated in the clustered 2-tier P2P IR architecture resource selection layer by way of usage of clustering, and posit that this necessitates a re-look at the applicability of document retrieval methods for resource selection within such a framework. This paper empirically benchmarks document retrieval models against the state-of-the-art resource selection models for the problem of resource selection in the clustered P2P IR architecture, using classical IR evaluation metrics. Our benchmarking study illustrates that document retrieval models significantly outperform other methods for the task of resource selection in the clustered P2P IR architecture. This indicates that clustered P2P IR framework can exploit advancements in document retrieval methods to deliver corresponding improvements in resource selection, indicating potential convergence of these fields for the clustered P2P IR architecture. Rami Suleiman Alkhawaldeh, Joemon M. Jose, Deepak P 0001 |
CIKM | 3 |
| 2016 | Leveraging Stratification in Twitter SamplingabstractWith Tweet volumes reaching 500 million a day, sampling is inevitable for any application using Twitter data. Realizing this, data providers such as Twitter, Gnip and Boardreader license sampled data streams priced in accordance with the sample size. Big Data applications working with sampled data would be interested in working with a large enough sample that is representative of the universal dataset. Previous work focusing on the representativeness issue has considered ensuring that global occurrence rates of key terms, be reliably estimated from the sample. Present technology allows sample size estimation in accordance with probabilistic bounds on occurrence rates for the case of uniform random sampling. In this paper, we consider the problem of further improving sample size estimates by leveraging stratification in Twitter data. We analyze our estimates through an extensive study using simulations and real-world data, establishing the superiority of our method over uniform random sampling. Our work provides the technical know-how for data providers to expand their portfolio to include stratified sampled datasets, whereas applications are benefited by being able to monitor more topics/events at the same data and computing cost. Vikas Joshi, Deepak P 0001, L. Venkata Subramaniam |
ECAI | 2 |
| 2016 | MixKMeans: Clustering Question-Answer ArchivesabstractCommunity-driven Question Answering (CQA) systems that crowdsource experiential information in the form of questions and answers and have accumulated valuable reusable knowledge.Clustering of QA datasets from CQA systems provides a means of organizing the content to ease tasks such as manual curation and tagging.In this paper, we present a clustering method that exploits the two-part question-answer structure in QA datasets to improve clustering quality.Our method, MixKMeans, composes question and answer space similarities in a way that the space on which the match is higher is allowed to dominate.This construction is motivated by our observation that semantic similarity between question-answer data (QAs) could get localized in either space.We empirically evaluate our method on a variety of real-world labeled datasets.Our results indicate that our method significantly outperforms stateof-the-art clustering methods for the task of clustering question-answer archives. Deepak P 0001 |
EMNLP | 1 |
| 2016 | Clustering-Based Query Routing in Cooperative Semi-Structured Peer to Peer NetworksabstractWe consider the problem of resource selection in clustered Peer-to-Peer Information Retrieval (P2P IR) networks with cooperative peers. The clustered P2P IR framework presents a significant departure from general P2P IR architectures by employing clustering to ensure content coherence between resources at the resource selection layer, without disturbing document allocation. We propose that such a property could be leveraged in resource selection by adapting well-studied and popular inverted lists for centralized document retrieval. Accordingly, we propose the Inverted PeerCluster Index (IPI), an approach that adapts the inverted lists, in a straightforward manner, for resource selection in clustered P2P IR. IPI also encompasses a strikingly simple peer-specific scoring mechanism that exploits the said index for resource selection. Through an extensive empirical analysis on P2P IR testbeds, we establish that IPI competes well with the sophisticated state-of-the-art methods in virtually every parameter of interest for the resource selection task, in the context of clustered P2P IR. Rami Suleiman Alkhawaldeh, Joemon M. Jose, Deepak P 0001 |
ICTAI | 3 |
| 2016 | Select, Link and Rank: Diversified Query Expansion and Entity Ranking Using Wikipedia
Adit Krishnan, Deepak P 0001, Sayan Ranu, Sameep Mehta |
WISE (1) | 2 |
| 2015 | Entity Linking for Web Search Queries
Deepak P 0001, Sayan Ranu, Prithu Banerjee, Sameep Mehta |
ECIR | 1 |
| 2015 | Indexing and matching trajectories under inconsistent sampling ratesabstractQuantifying the similarity between two trajectories is a fundamental operation in analysis of spatio-temporal databases. While a number of distance functions exist, the recent shift in the dynamics of the trajectory generation procedure violates one of their core assumptions; a consistent and uniform sampling rate. In this paper, we formulate a robust distance function called Edit Distance with Projections (EDwP) to match trajectories under inconsistent and variable sampling rates through dynamic interpolation. This is achieved by deploying the idea of projections that goes beyond matching only the sampled points while aligning trajectories. To enable efficient trajectory retrievals using EDwP, we design an index structure called TrajTree. TrajTree derives its pruning power by employing the unique combination of bounding boxes with Lipschitz embedding. Extensive experiments on real trajectory databases demonstrate EDwP to be up to 5 times more accurate than the state-of-the-art distance functions. Additionally, TrajTree increases the efficiency of trajectory retrievals by up to an order of magnitude over existing techniques. Sayan Ranu, Deepak P 0001, Aditya Telang, Prasad Deshpande, Sriram Raghavan |
ICDE | 2 |
| 2014 | Unsupervised Solution Post Identification from Discussion ForumsabstractDiscussion forums have evolved into a dependable source of knowledge to solve common problems.However, only a minority of the posts in discussion forums are solution posts.Identifying solution posts from discussion forums, hence, is an important research problem.In this paper, we present a technique for unsupervised solution post identification leveraging a so far unexplored textual feature, that of lexical correlations between problems and solutions.We use translation models and language models to exploit lexical correlations and solution post character respectively.Our technique is designed to not rely much on structural features such as post metadata since such features are often not uniformly available across forums.Our clustering-based iterative solution identification approach based on the EM-formulation performs favorably in an empirical evaluation, beating the only unsupervised solution identification technique from literature by a very large margin.We also show that our unsupervised technique is competitive against methods that require supervision, outperforming one such technique comfortably. Deepak P 0001, Karthik Visweswariah |
ACL (1) | 1 |
| 2014 | Fast Mining of Interesting Phrases from Subsets of Text CorporaabstractWe address the problem of mining interesting phrases from subsets of a text corpus where the subset is specified using a set of features such as keywords that form a query. Previous algorithms for the problem have proposed solutions that involve sifting through a phrase dictionary based index or a document-based index where the solution is linear in either the phrase dictionary size or the size of the document subset. We propose the usage of an independence assumption between query keywords given the top correlated phrases, wherein the pre-processing could be reduced to discovering phrases from among the top phrases per each feature in the query. We then outline an indexing mechanism where per-keyword phrase lists are stored either in disk or memory, so that popular aggregation algorithms such as No Random Access and Sort-merge Join may be adapted to do the scoring at real-time to identify the top interesting phrases. Though such an approach is expected to be approximate, we empirically illustrate that very high accuracies (of over 90%) are achieved against the results of exact algorithms. Due to the simplified list-aggregation, we are also able to provide response times that are orders of magnitude better than state-of-the-art algorithms. Interestingly, our disk-based approach outperforms the in-memory baselines by up to hundred times and sometimes more, confirming the superiority of the proposed method. Deepak P 0001, Atreyee Dey, Debapriyo Majumdar |
EDBT | 1 |
| 2014 | Detecting localized homogeneous anomalies over spatio-temporal data
Aditya Telang, Deepak P 0001, Salil Joshi 0001, Prasad Deshpande, Ranjana Rajendran |
Data Min. Knowl. Discov. | 2 |
| 2013 | Query Suggestions for Textual Problem Solution Repositories
Deepak P 0001, Sutanu Chakraborti, Deepak Khemani |
ECIR | 1 |
| 2012 | Two-part segmentation of text documentsabstractWe consider the problem of segmenting text documents that have a two-part structure such as a problem part and a solution part. Documents of this genre include incident reports that typically involve description of events relating to a problem followed by those pertaining to the solution that was tried. Segmenting such documents into the component two parts would render them usable in knowledge reuse frameworks such as Case-Based Reasoning. This segmentation problem presents a hard case for traditional text segmentation due to the lexical inter-relatedness of the segments. We develop a two-part segmentation technique that can harness a corpus of similar documents to model the behavior of the two segments and their inter-relatedness using language models and translation models respectively. In particular, we use separate language models for the problem and solution segment types, whereas the inter-relatedness between segment types is modeled using an IBM Model 1 translation model. We model documents as being generated starting from the problem part that comprises of words sampled from the problem language model, followed by the solution part whose words are sampled either from the solution language model or from a translation model conditioned on the words already chosen in the problem part. We show, through an extensive set of experiments on real-world data, that our approach outperforms the state-of-the-art text segmentation algorithms in the accuracy of segmentation, and that such improved accuracy translates well to improved usability in Case-based Reasoning systems. We also analyze the robustness of our technique to varying amounts and types of noise and empirically illustrate that our technique is quite noise tolerant, and degrades gracefully with increasing amounts of noise. Deepak P 0001, Karthik Visweswariah, Nirmalie Wiratunga, Sadiq Sani |
CIKM | 1 |
| 2012 | Retrieving similar discussion forum threads: a structure based approachabstractOnline forums are becoming a popular way of finding useful information on the web. Search over forums for existing discussion threads so far is limited to keyword-based search due to the minimal effort required on part of the users. However, it is often not possible to capture all the relevant context in a complex query using a small number of keywords. Example-based search that retrieves similar discussion threads given one exemplary thread is an alternate approach that can help the user provide richer context and vastly improve forum search results. In this paper, we address the problem of finding similar threads to a given thread. Towards this, we propose a novel methodology to estimate similarity between discussion threads. Our method exploits the thread structure to decompose threads in to set of weighted overlapping components. It then estimates pairwise thread similarities by quantifying how well the information in the threads are mutually contained within each other using lexical similarities between their underlying components. We compare our proposed methods on real datasets against state-of-the-art thread retrieval mechanisms wherein we illustrate that our techniques outperform others by large margins on popular retrieval evaluation measures such as NDCG, MAP, [email protected] and MRR. In particular, consistent improvements of up to 10% are observed on all evaluation measures. Amit Singh 0003, Deepak P 0001, Dinesh Raghu |
SIGIR | 2 |
| 2012 | Finding Relevant Tweets
Deepak P 0001, Sutanu Chakraborti |
WAIM | 1 |
| 2012 | Improving Recall of Regular Expressions for Information Extraction
Karin Murthy, Deepak P 0001, Prasad Deshpande |
WISE | 2 |
| 2012 | Interpretable and reconfigurable clustering of document datasets by deriving word-based rules
Vipin Balachandran, Deepak P 0001, Deepak Khemani |
Knowl. Inf. Syst. | 2 |
| 2012 | Exploiting Evidence from Unstructured Data to Enhance Master Data ManagementabstractMaster data management (MDM) integrates data from multiple structured data sources and builds a consolidated 360-degree view of business entities such as customers and products. Today's MDM systems are not prepared to integrate information from unstructured data sources, such as news reports, emails, call-center transcripts, and chat logs. However, those unstructured data sources may contain valuable information about the same entities known to MDM from the structured data sources. Integrating information from unstructured data into MDM is challenging as textual references to existing MDM entities are often incomplete and imprecise and the additional entity information extracted from text should not impact the trustworthiness of MDM data. In this paper, we present an architecture for making MDM text-aware and showcase its implementation as IBM Info-Sphere MDM Extension for Unstructured Text Correlation, an add-on to IBM InfoSphere Master Data Management Standard Edition. We highlight how MDM benefits from additional evidence found in documents when doing entity resolution and relationship discovery. We experimentally demonstrate the feasibility of integrating information from unstructured data sources into MDM. Karin Murthy, Prasad Deshpande, Atreyee Dey, Ramanujam Halasipuram, Mukesh K. Mohania, Deepak P 0001, Jennifer Reed, Scott Schumacher |
Proc. VLDB Endow. | 6 |
| 2011 | More or better: on trade-offs in compacting textual problem solution repositoriesabstractIn this paper, we look into the problem of filtering problem solution repositories (from sources such as community-driven question answering systems) to render them more suitable for usage in knowledge reuse systems. We explore harnessing the fuzzy nature of usability of a solution to a problem, for such compaction. Fuzzy usabilities lead to several challenges; notably, the trade-off between choosing generic or better solutions. We develop an approach that can heed to a user specification of the trade-off between these criteria and introduce several quality measures based on fuzzy usability estimates to ascertain the quality of a problem-solution repository for usage in a Case Based Reasoning system. We establish, through a detailed empirical analysis, that our approach outperforms state-of-the-art approaches on virtually all quality measures. Deepak P 0001, Sutanu Chakraborti, Deepak Khemani |
CIKM | 1 |
| 2011 | Efficient reverse skyline retrieval with arbitrary non-metric similarity measuresabstractA Reverse Skyline query returns all objects whose skyline contains the query object. In this paper, we consider Reverse Skyline query processing where the distance between attribute values are not necessarily metric. We outline real world cases that motivate Reverse Skyline processing in such scenarios. We consider various optimizations to develop efficient algorithms for Reverse Skyline processing. Firstly, we consider block-based processing of objects to optimize on IO costs. We then explore pre-processing to re-arrange objects on disk to speed-up computational and IO costs. We then present our main contribution, which is a method of using group-level reasoning and early pruning to micro-optimize processing by reducing attribute level comparisons. An extensive empirical evaluation with real-world datasets and synthetic data of varying characteristics shows that our optimization techniques are indeed very effective in dramatically speeding Reverse Skyline processing, both in terms of computational costs and IO costs. Prasad Deshpande, Deepak P 0001 |
EDBT | 2 |
| 2011 | Fast Rule Mining Over Multi-Dimensional WindowsabstractAssociation rule mining is an indispensable tool for discovering insights from large databases and data warehouses.The data in a warehouse being multi-dimensional, it is often useful to mine rules over subsets of data defined by selections over the dimensions.Such interactive rule mining over multi-dimensional query windows is difficult since rule mining is computationally expensive.Current methods using pre-computation of frequent itemsets require counting of some itemsets by revisiting the transaction database at query time, which is very expensive.We develop a method (RMW) that identifies the minimal set of itemsets to compute and store for each cell, so that rule mining over any query window may be performed without going back to the transaction database.We give formal proofs that the set of itemsets chosen by RMW is sufficient to answer any query and also prove that it is the optimal set to be computed for 1 dimensional queries.We demonstrate through an extensive empirical evaluation that RMW achieves extremely fast query response time compared to existing methods, with only moderate overhead in pre-computation and storage. Mahashweta Das, Deepak P 0001, Prasad Deshpande, Ramakrishnan Kannan |
SDM | 2 |
| 2010 | Efficient RkNN Retrieval with Arbitrary Non-Metric Similarity MeasuresabstractA R k NN query returns all objects whose nearest k neighbors contain the query object. In this paper, we consider R k NN query processing in the case where the distances between attribute values are not necessarily metric. Dissimilarities between objects could then be a monotonic aggregate of dissimilarities between their values, such aggregation functions being specified at query time. We outline real world cases that motivate R k NN processing in such scenarios. We consider the AL-Tree index and its applicability in R k NN query processing. We develop an approach that exploits the group level reasoning enabled by the AL-Tree in R k NN processing. We evaluate our approach against a Naive approach that performs sequential scans on contiguous data and an improved block-based approach that we provide. We use real-world datasets and synthetic data with varying characteristics for our experiments. This extensive empirical evaluation shows that our approach is better than existing methods in terms of computational and disk access costs, leading to significantly better response times. Deepak P 0001, Prasad Deshpande |
Proc. VLDB Endow. | 1 |
| 2009 | Interpretable and reconfigurable clustering of document datasets by deriving word-based rulesabstractClusters of text documents output by clustering algorithms are often hard to interpret. We describe motivating real-world scenarios that necessitate reconfigurability and high interpretability of clusters and outline the problem of generating clusterings with interpretable and reconfigurable cluster models. We develop a clustering algorithm toward the outlined goal of building interpretable and reconfigurable cluster models; it works by generating rules with disjunctions and conditions on the frequencies of words, to decide on the membership of a document to a cluster. Each cluster is comprised of precisely the set of documents that satisfy the corresponding rule. We show that our approach outperforms the unsupervised decision tree approach by huge margins. We show that the purity and f-measure losses to achieve interpretability are as little as 5% and 3% respectively using our approach. Vipin Balachandran, Deepak P 0001, Deepak Khemani |
CIKM | 2 |
| 2009 | Efficient skyline retrieval with arbitrary similarity measuresabstractA skyline query returns a set of objects that are not dominated by other objects. An object is said to dominate another if it is closer to the query than the latter on all factors under consideration. In this paper, we consider the case where the similarity measures may be arbitrary and do not necessarily come from a metric space. We first explore middleware algorithms, analyze how skyline retrieval for non-metric spaces can be done on the middleware backend, and lay down a necessary and sufficient stopping condition for middleware-based skyline algorithms. We develop the Balanced Access Algorithm, which is provably more IO-friendly than the state-of-the-art algorithm for skyline query processing on middleware and show that BAA outperforms the latter by orders of magnitude. We also show that without prior knowledge about data distributions, it is unlikely to have a middleware algorithm that is more IO-friendly than BAA. In fact, we empirically show that BAA is very close to the absolute lower bound of IO costs for middleware algorithms. Further, we explore the non-middleware setting and devise an online algorithm for skyline retrieval which uses a recently proposed value space index over non-metric spaces (AL-Tree [10]). The AL-Tree based algorithm is able to prune subspaces and efficiently maintain candidate sets leading to better performance. We compare our algorithms to existing ones which can work with arbitrary similarity measures and show that our approaches are better in terms of computational and disk access costs leading to significantly better response times. Deepak P 0001, Prasad Deshpande, Debapriyo Majumdar, Raghu Krishnapuram |
EDBT | 1 |
| 2009 | CAESAR: A Context-Aware, Social Recommender System for Low-End Mobile DevicesabstractMobile-enabled social networks applications are becoming increasingly popular. Most of the current social network applications have been designed for high-end mobile devices, and they rely upon features such as GPS, capabilities of the world wide web, and rich media support. However, a significant fraction of mobile user base, especially in the developing world, own low-end devices that are only capable of voice and short text messages (SMS). In this context, a natural question is whether one can design meaningful social network-based applications that can work well with these simple devices, and if so, what the real challenges are. Towards answering these questions, this paper presents a social network-based recommender system that has been explicitly designed to work even with devices that just support phone calls and SMS. Our design of the social network based recommender system incorporates three features that complement each other to derive highly targeted ads. First, we analyze information such as customer's address books to estimate the level of social affinity among various users. This social affinity information is used to identify the recommendations to be sent to an individual user. Second, we combine the social affinity information with the spatio-temporal context of users and historical responses of the user to further refine the set of recommendations and to decide when a recommendation would be sent. Third, social affinity computation and spatio-temporal contextual association are continuously tuned through user feedback. We outline the challenges in building such a system, and outline approaches to deal with such challenges. Lakshmish Ramaswamy, Deepak P 0001, Ramana Polavarapu, Kutila Gunasekera, Dinesh Garg, Karthik Visweswariah, Shivkumar Kalyanaraman |
Mobile Data Management | 2 |
| 2008 | Efficient online top-K retrieval with arbitrary similarity measuresabstractThe top-k retrieval problem requires finding k objects most similar to a given query object. Similarities between objects are most often computed as aggregated similarities of their attribute values. We consider the case where the similarities between attribute values are arbitrary (non-metric), due to which standard space partitioning indexes cannot be used. Among the most popular techniques that can handle arbitrary similarity measures is the family of threshold algorithms. These were designed as middleware algorithms that assume that similarity lists for each attribute are available and focus on efficiently merging these lists to arrive at the results. In this paper, we explore multi-dimensional indexing of non-metric spaces that can lead to efficient pruning of the search space utilizing inter-attribute relationships, during top-k computation. We propose an indexing structure, the AL-Tree and an algorithm to do top-k retrieval using it in an online fashion. The ALTree exploits the fact that many real world attributes come from a small value space. We show that our algorithm performs much better than the threshold based algorithms in terms of computational cost due to efficient pruning of the search space. Further, it out-performs them in terms of IOs by upto an order of magnitude in case of dense datasets. Prasad Deshpande, Deepak P 0001, Krishna Kummamuru |
EDBT | 2 |
| 2008 | Intelligent user assistance for cost effective usage of mobile phoneabstractCost is the governing factor which defines the penetration and adoption of mobile phones in lower strata (lower income group) of the society particularly in developing countries like India. In this paper we describe an enhancement to the mobile phone design which interacts with the user to facilitate cost-conscious usage of the mobile phone. In particular, we propose an intelligent component in the mobile phone which tracks mobile usage pattern and informs the user of deviations in usage and suggests means of using the device cost effectively. Deepak P 0001, Anuradha Bhamidipaty, Swati Challa |
IUI | 1 |
| 2008 | Unsupervised Segmentation of Conversational TranscriptsabstractContact centers provide dialog based support to organizations to address various customer related issues. We have observed that the calls received at contact centers mostly follow well defined patterns. Such call flows not only specify how an agent should proceed in a call, handle objections, persuade customers, follow compliance issues, etc but also help to structure the operational process of call handling. Automatically identifying such patterns in terms of distinct segments from a collection of transcripts of conversations would improve productivity of agents as well as track compliance to guidelines. Call transcripts from call centers typically tend to be noisy owing to the noise arising from agent/caller distractions, and errors introduced by the speech recognition engine. Such noise makes classical text segmentation algorithms such as TextTiling, which work on each transcript in isolation, very inappropriate. But such noise effects become statistically insignificant over a corpus of similar calls. In this paper, we propose an algorithm to segment conversational transcripts in an unsupervised way utilizing corpus level information of similar call transcripts. We show that our approach outperforms the classical TextTiling algorithm and also describe ways to improve the segmentation using limited supervision. We discuss various ways of evaluating such an algorithm. We apply the proposed algorithm to a corpus of transcripts of calls from a car reservation call center and evaluate it using various evaluation measures. We apply segmentation to the problem of automatically checking the compliance of agents and show that our segmentation algorithm considerably improves the precision. Krishna Kummamuru, Deepak P 0001, Shourya Roy, L. Venkata Subramaniam |
SDM | 2 |
| 2007 | SymAB : Symbol-Based Address Book for the Semi-literate Mobile User
Anuradha Bhamidipaty, Deepak P 0001 |
INTERACT (1) | 2 |
| 2007 | Optimizing on Mobile Usage Cost for the Lower Income Group: Insights and Recommendations
Deepak P 0001, Anuradha Bhamidipaty |
INTERACT (1) | 1 |
| 2007 | Mining conversational text for procedures with applications in contact centers
Deepak P 0001, Krishna Kummamuru |
Int. J. Document Anal. Recognit. | 1 |
| 2006 | Building Clusters of Related Words: An Unsupervised Approach
Deepak P 0001, Delip Rao, Deepak Khemani |
PRICAI | 1 |
| 2004 | Context Disambiguation in Web Search ResultsabstractIt is a common experience while Web searching that one gets to see pages that are not of interest. Partly these are due to a word or words in the search query having different contexts, the user obviously expecting to find pages related to the context of interest. This paper proposes a method for disambiguating contexts in Web search results. Deepak P 0001, Jyothi John, Sandeep Parameswaran |
ICWS | 1 |