VLDB 2026 Research / reviewers in the wild / expert
Pawan Goyal 0002
dblp:77/2307-2
· DBLP profile ↗
32ranked-venue papers in the field
3as first author
8since 2021 · last 2026
0000-0002-9414-8166ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 23 (1 first)Data Mining & Knowledge Discovery · 6Database Systems & Data Management · 2 (2 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion
Rahul Mehta 0008, Kavin R. V, Indrajit Pal, Tushar Abhishek, Pawan Goyal 0002, Manish Gupta 0001 |
SIGIR | 5 |
| 2024 | Instruction-Guided Bullet Point Summarization of Long Financial Earnings Call TranscriptsabstractWhile automatic summarization techniques have made significant advancements, their primary focus has been on summarizing short news articles or documents that have clear structural patterns like scientific articles or government reports. There has not been much exploration into developing efficient methods for summarizing financial documents, which often contain complex facts and figures. Here, we study the problem of bullet point summarization of long Earning Call Transcripts (ECTs) using the recently released ECTSum dataset. We leverage an unsupervised question-based extractive module followed by a parameter efficient instruction-tuned abstractive module to solve this task. Our proposed model FLANFinBPS achieves new state-of-the-art performances outperforming the strongest baseline with 14.88% average ROUGE score gain, and is capable of generating factually consistent bullet point summaries that capture the important facts discussed in the ECTs. We make the codebase publicly available at https://github.com/subhendukhatuya/FLAN-FinBPS. Subhendu Khatuya, Koushiki Sinha, Niloy Ganguly, Saptarshi Ghosh 0001, Pawan Goyal 0002 |
SIGIR | 5 |
| 2024 | Legal Statute Identification: A Case Study using State-of-the-Art Datasets and MethodsabstractLegal Statute Identification (LSI) involves identifying the relevant statutes (articles of law) given the facts (evidence) of a legal case. There are several key challenges in LSI, such as (i)~usage of label (statute) semantics which can be complicated and confusing; (ii)~the input text (i.e., the facts) are very long and noisy; (iii)~the label distribution usually follows a long tail, making predictions for the rare labels challenging. Although multiple methods have been proposed to address these challenges, there has not been any comprehensive study to establish the effects of these factors on different models/approaches. In this work, we reproduce several LSI models on two popular LSI datasets and study the effect of the above-mentioned challenges. We conduct thorough experiments with transformer-based encoders such as BERT and Longformer. We further try out different combinations of these encoders with approaches devised specifically for LSI, which essentially use different mechanisms to model the statute texts to enhance fact representations. Our experiments yield several interesting insights into how the above-mentioned challenges are addressed by different models, the interplay of different encoding and statute text handling measures, and how the nature of the LSI datasets affects the model performances. Finally, we also analyze the explanability capabilities of different approaches using human-annotated rationales. Shounak Paul, Rajas Bhatt, Pawan Goyal 0002, Saptarshi Ghosh 0001 |
SIGIR | 3 |
| 2023 | Legal IR and NLP: The History, Challenges, and State-of-the-Art
Debasis Ganguly, Jack G. Conrad, Kripabandhu Ghosh, Saptarshi Ghosh 0001, Pawan Goyal 0002, Paheli Bhattacharya, Shubham Kumar Nigam, Shounak Paul |
ECIR (3) | 5 |
| 2022 | MTLTS: A Multi-Task Framework To Obtain Trustworthy Summaries From Crisis-Related MicroblogsabstractOccurrences of catastrophes such as natural or man-made disasters trigger the spread of rumours over social media at a rapid pace. Presenting a trustworthy and summarized account of the unfolding event in near real-time to the consumers of such potentially unreliable information thus becomes an important task. In this work, we propose MTLTS, the first end-to-end solution for the task that jointly determines the credibility and summary-worthiness of tweets. Our credibility verifier is designed to recursively learn the structural properties of a Twitter conversation cascade, along with the stances of replies towards the source tweet. We then take a hierarchical multi-task learning approach, where the verifier is trained at a lower layer, and the summarizer is trained at a deeper layer where it utilizes the verifier predictions to determine the salience of a tweet. Different from existing disaster-specific summarizers, we model tweet summarization as a supervised task. Such an approach can automatically learn summary-worthy features, and can therefore generalize well across domains. When trained on the PHEME dataset [29], not only do we outperform the strongest baselines for the auxiliary task of verification/rumour detection, we also achieve 21 - 35% gains in the verified ratio of summary tweets, and 16 - 20% gains in ROUGE1-F1 scores over the existing state-of-the-art solutions for the primary task of trustworthy summarization. Rajdeep Mukherjee, Uppada Vishnu, Hari Chandana Peruri, Sourangshu Bhattacharya, Koustav Rudra, Pawan Goyal 0002, Niloy Ganguly |
WSDM | 6 |
| 2021 | Which acts model happiness?: an exploratory analysis on Twitter and GoodreadsabstractModeling and analysis of affective and inner states is gaining prominence in research. Articulating the entire spectrum, ranging from recipes of long-term happiness to factors leading to depression, we frame a model of happiness states of people comprising of three states: G (lasting happiness), P (flickering) and I (frustration), respectively. The definitions of these states are based on psychology literature. We used a XgBoost Classifier to categorize 54,066 Twitter users based on their tweets and analysed the results including what kind of friends each category of users have (for 120 users obtained after thresholding 213 manually labelled users). Analysing XgBoost classification we could re-confirm characteristics mentioned in the definition of the three states (G, P, I) and find out more traits/characteristics beyond the definition as well. We observed that G users are more people-oriented. G and P users are more work-oriented than I users. G users are elder in age to P or I users. I users were found to be more religious than P owing to shelter-seeking traits. Qualitative analysis shows that G group suggests long-term vision, selfless and positive qualities, religious mindset and positive demeanour as expected. I group suggests negative feelings and activities and sensual words as expected. P group has traces of both G and I. P group contains dominating, strong words and extreme negative reactions. We found 21,115 users having Twitter and Goodreads handles to study what kind of books users of each category read. Reading patterns of G constitute of academic/technical, religion, inspirational/self-help and romance. Those of P users are fantasy/fiction, sports, LGBT/BDSM/Erotica and horror/violence/betrayal. I users tend to read fantasy/fiction, death and indiscriminately any arbitrary topic. G and P users make friends in the same category whereas I users tend to have friends in P category, but not among themselves. Mayank Bhasin, Pawan Goyal 0002 |
ASONAM | 3 |
| 2021 | Joint Autoregressive and Graph Models for Software and Developer Social Networks
Rima Hazra, Hardik Aggarwal, Pawan Goyal 0002, Animesh Mukherjee 0001, Soumen Chakrabarti |
ECIR (1) | 3 |
| 2021 | Reproducibility, Replicability and Beyond: Assessing Production Readiness of Aspect Based Sentiment Analysis in the Wild
Rajdeep Mukherjee, Shreyas Shetty, Subrata Chattopadhyay, Subhadeep Maji, Samik Datta, Pawan Goyal 0002 |
ECIR (2) | 6 |
| 2020 | Read what you need: Controllable Aspect-based Opinion Summarization of Tourist ReviewsabstractManually extracting relevant aspects and opinions from large volumes of user-generated text is a time-consuming process. Summaries, on the other hand, help readers with limited time budgets to quickly consume the key ideas from the data. State-of-the-art approaches for multi-document summarization, however, do not consider user preferences while generating summaries. In this work, we argue the need and propose a solution for generating personalized aspect-based opinion summaries from large collections of online tourist reviews. We let our readers decide and control several attributes of the summary such as the length and specific aspects of interest among others. Specifically, we take an unsupervised approach to extract coherent aspects from tourist reviews posted onTripAdvisor. We then propose an Integer Linear Programming (ILP) based extractive technique to select an informative subset of opinions around the identified aspects while respecting the user-specified values for various control parameters. Finally, we evaluate and compare our summaries using crowdsourcing and ROUGE-based metrics and obtain competitive results. Rajdeep Mukherjee, Hari Chandana Peruri, Uppada Vishnu, Pawan Goyal 0002, Sourangshu Bhattacharya, Niloy Ganguly |
SIGIR | 4 |
| 2020 | Network measures: A new paradigm towards reliable novel word sense detection
Abhik Jana, Animesh Mukherjee 0001, Pawan Goyal 0002 |
Inf. Process. Manag. | 3 |
| 2019 | Fully Contextualized Biomedical NER
Ashim Gupta, Pawan Goyal 0002, Sudeshna Sarkar, Mahanandeeshwar Gattu |
ECIR (2) | 2 |
| 2019 | DeepTagRec: A Content-cum-User Based Tag Recommendation Framework for Stack Overflow
Suman Kalyan Maity, Abhishek Panigrahi, Sayan Ghosh 0002, Arundhati Banerjee, Pawan Goyal 0002, Animesh Mukherjee 0001 |
ECIR (2) | 5 |
| 2019 | Misleading Metadata Detection on YouTube
Priyank Palod, Ayush Patwari, Sudhanshu Bahety, Saurabh Bagchi, Pawan Goyal 0002 |
ECIR (2) | 5 |
| 2019 | Automated Early Leaderboard Generation from Comparative Tables
Mayank Singh 0001, Rajdeep Sarkar, Atharva Vyas, Pawan Goyal 0002, Animesh Mukherjee 0001, Soumen Chakrabarti |
ECIR (1) | 4 |
| 2019 | Thou Shalt Not Hate: Countering Online Hate Speech
Binny Mathew, Punyajoy Saha, Hardik Tharad, Subham Rajgaria, Prajwal Singhania, Suman Kalyan Maity, Pawan Goyal 0002, Animesh Mukherjee 0001 |
ICWSM | 7 |
| 2019 | Addressing Vocabulary Gap in E-commerce SearchabstractE-commerce customers express their purchase intents in several ways, some of which may use a different vocabulary than that of the product catalog. For example, the intent for "women maternity gown" is often expressed with the query, "ladies pregnancy dress". Search engines typically suffer from poor performance on such queries because of low overlap between query terms and specifications of the desired products. Past work has referred to these queries as vocabulary gap queries. In our experiments, we show that our technique significantly outperforms strong baselines and also show its real-world effectiveness with an online A/B experiment. Subhadeep Maji, Manish Bansal, Kalyani Roy, Mohit Kumar 0008, Pawan Goyal 0002 |
SIGIR | 6 |
| 2018 | Automated Assistance in E-commerce: An Approach Based on Category-Sensitive Retrieval
Anirban Majumder, Abhay Pande, Kondalarao Vonteru, Abhishek Gangwar, Subhadeep Maji, Pankaj Bhatia, Pawan Goyal 0002 |
ECIR | 7 |
| 2018 | Identifying Sub-events and Summarizing Disaster-Related Information from MicroblogsabstractIn recent times, humanitarian organizations increasingly rely on social media to search for information useful for disaster response. These organizations have varying information needs ranging from general situational awareness (i.e., to understand a bigger picture) to focused information needs e.g., about infrastructure damage, urgent needs of affected people. This research proposes a novel approach to help crisis responders fulfill their information needs at different levels of granularities. Specifically, the proposed approach presents simple algorithms to identify sub-events and generate summaries of big volume of messages around those events using an Integer Linear Programming (ILP) technique. Extensive evaluation on a large set of real world Twitter dataset shows (a). our algorithm can identify important sub-events with high recall (b). the summarization scheme shows (6---30%) higher accuracy of our system compared to many other state-of-the-art techniques. The simplicity of the algorithms ensures that the entire task is done in real time which is needed for practical deployment of the system. Koustav Rudra, Pawan Goyal 0002, Niloy Ganguly, Prasenjit Mitra 0001, Muhammad Imran 0002 |
SIGIR | 2 |
| 2018 | Extracting and Summarizing Situational Information from the Twitter Social Media during DisastersabstractMicroblogging sites like Twitter have become important sources of real-time information during disaster events. A large amount of valuable situational information is posted in these sites during disasters; however, the information is dispersed among hundreds of thousands of tweets containing sentiments and opinions of the masses. To effectively utilize microblogging sites during disaster events, it is necessary to not only extract the situational information from the large amounts of sentiments and opinions, but also to summarize the large amounts of situational information posted in real-time. During disasters in countries like India, a sizable number of tweets are posted in local resource-poor languages besides the normal English-language tweets. For instance, in the Indian subcontinent, a large number of tweets are posted in Hindi/Devanagari (the national language of India), and some of the information contained in such non-English tweets is not available (or available at a later point of time) through English tweets. In this work, we develop a novel classification-summarization framework which handles tweets in both English and Hindi—we first extract tweets containing situational information, and then summarize this information. Our proposed methodology is developed based on the understanding of how several concepts evolve in Twitter during disaster. This understanding helps us achieve superior performance compared to the state-of-the-art tweet classifiers and summarization approaches on English tweets. Additionally, to our knowledge, this is the first attempt to extract situational information from non-English tweets. Koustav Rudra, Niloy Ganguly, Pawan Goyal 0002, Saptarshi Ghosh 0001 |
ACM Trans. Web | 3 |
| 2017 | Extracting Social Lists from TwitterabstractSocial list queries like 'valentines day gift ideas', 'best anniversary messages for your parents', etc. are quite popular on web search engines. Users expect instant answers comprising of a list of relevant items (social list) for such a query. Surprisingly, current search engines do not provide any crisp instant answers for queries in this critical query segment. To the best of our knowledge, we propose the first system that tackles such queries. Although such social factors are heavily discussed on online social networks like Twitter, extracting such lists from tweets is quite challenging. How to discover such lists from tweets? We present a system that identifies these 'social lists' from a large number of Twitter hashtags using a high recall classifier trained using novel task-specific features with good accuracy. Further, we briefly discuss how list items can be extracted from related tweets. Experiments over a dataset of ~4M tweets show that our recall-optimized system can obtain up to 75.5% precision at 95.3% recall. Ankan Mullick, Pawan Goyal 0002, Niloy Ganguly, Manish Gupta 0001 |
ASONAM | 2 |
| 2017 | Extracting Entities of Interest from Comparative Product ReviewsabstractThis paper presents a deep learning based approach to extract product comparison information out of user reviews on various e-commerce websites. Any comparative product review has three major entities of information: the names of the products being compared, the user opinion (predicate) and the feature or aspect under comparison. All these informing entities are dependent on each other and bound by the rules of the language, in the review. We observe that their inter-dependencies can be captured well using LSTMs. We evaluate our system on existing manually labeled datasets and observe out-performance over the existing Semantic Role Labeling (SRL) framework popular for this task. Jatin Arora 0001, Sumit Agrawal, Pawan Goyal 0002, Sayan D. Pathak |
CIKM | 3 |
| 2017 | Relay-Linking Models for Prominence and Obsolescence in Evolving NetworksabstractThe rate at which nodes in evolving social networks acquire links (friends, citations) shows complex temporal dynamics. Preferential attachment and link copying models, while enabling elegant analysis, only capture rich-gets-richer effects, not aging and decline. Recent aging models are complex and heavily parameterized; most involve estimating 1-3 parameters per node. These parameters are intrinsic: they explain decline in terms of events in the past of the same node, and do not explain, using the network, where the linking attention might go instead. We argue that traditional characterization of linking dynamics are insufficient to judge the faithfulness of models. We propose a new temporal sketch of an evolving graph, and introduce several new characterizations of a network's temporal dynamics. Then we propose a new family of frugal aging models with no per-node parameters and only two global parameters. Our model is based on a surprising inversion or undoing of triangle completion, where an old node relays a citation to a younger follower in its immediate vicinity. Despite very few parameters, the new family of models shows remarkably better fit with real data. Before concluding, we analyze temporal signatures for various research communities yielding further insights into their comparative dynamics. To facilitate reproducible research, we shall soon make all the codes and the processed dataset available in the public domain. Mayank Singh 0001, Rajdeep Sarkar, Pawan Goyal 0002, Animesh Mukherjee 0001, Soumen Chakrabarti |
KDD | 3 |
| 2017 | Using re-ranking to boost deep learning based community question retrievalabstractThe current study presents a two-stage question retrieval approach which, in the first phase, retrieves similar questions for a given query using a deep learning based approach and in the second phase, re-ranks initially retrieved questions on the basis of inter-question similarities. The suggested deep learning based approach is trained using several surface features of texts and the associated weights are pre-trained using a deep generative model for better initialization. The proposed retrieval model outperforms standard baseline question retrieval approaches. The proposed re-ranking approach performs inference over a similarity graph constructed with the initially retrieved questions and re-ranks the questions based on their similarity with other relevant questions. Suggested re-ranking approach significantly improves the precision for the retrieval task. Krishnendu Ghosh, Plaban Kumar Bhowmick, Pawan Goyal 0002 |
WI | 3 |
| 2016 | PEQ: An Explainable, Specification-based, Aspect-oriented Product Comparator for E-commerceabstractWhile purchasing a product, consumers often rely on specifications as well as online reviews of the product for decision-making. While comparing, one often has in mind a specific aspect or a set of aspects which are of interest to them. Previous work has used comparative sentences, where two entities are compared directly in a single sentence by the review author, towards the comparison task. In this paper, we extend the existing model by incorporating the feature specifications of the products, which are easily available, and learn the importance to be associated with each of them. To test the validity of these product ranking measures, we comprehensively test it on a digital camera dataset from Amazon.com and the results show good empirical outperformance over the state-of-the-art baselines. Abhishek Sikchi, Pawan Goyal 0002, Samik Datta |
CIKM | 2 |
| 2016 | FeRoSA: A Faceted Recommendation System for Scientific Articles
Tanmoy Chakraborty 0002, Amrith Krishna, Mayank Singh 0001, Niloy Ganguly, Pawan Goyal 0002, Animesh Mukherjee 0001 |
PAKDD (2) | 5 |
| 2015 | Extracting Situational Information from Microblogs during Disaster Events: a Classification-Summarization ApproachabstractMicroblogging sites like Twitter have become important sources of real-time information during disaster events. A significant amount of valuable situational information is available in these sites; however, this information is immersed among hundreds of thousands of tweets, mostly containing sentiments and opinion of the masses, that are posted during such events. To effectively utilize microblogging sites during disaster events, it is necessary to (i) extract the situational information from among the large amounts of sentiment and opinion, and (ii) summarize the situational information, to help decision-making processes when time is critical. In this paper, we develop a novel framework which first classifies tweets to extract situational information, and then summarizes the information. The proposed framework takes into consideration the typicalities pertaining to disaster events where (i) the same tweet often contains a mixture of situational and non-situational information, and (ii) certain numerical information, such as number of casualties, vary rapidly with time, and thus achieves superior performance compared to state-of-the-art tweet summarization approaches. Koustav Rudra, Subham Ghosh, Niloy Ganguly, Pawan Goyal 0002, Saptarshi Ghosh 0001 |
CIKM | 4 |
| 2015 | The Role Of Citation Context In Predicting Long-Term Citation Profiles: An Experimental Study Based On A Massive Bibliographic Text DatasetabstractThe impact and significance of a scientific publication is measured mostly by the number of citations it accumulates over the years. Early prediction of the citation profile of research articles is a significant as well as challenging problem. In this paper, we argue that features gathered from the citation contexts of the research papers can be very relevant for citation prediction. Analyzing a massive dataset of nearly 1.5 million computer science articles and more than 26 million citation contexts, we show that average countX (number of times a paper is cited within the same article) and average citeWords (number of words within the citation context) discriminate between various citation ranges as well as citation categories. We use these features in a stratified learning framework for future citation prediction. Experimental results show that the proposed model significantly outperforms the existing citation prediction models by a margin of 8-10% on an average under various experimental settings. Specifically, the features derived from the citation context help in predicting long-term citation behavior. Mayank Singh 0001, Vikas Patidar, Suhansanu Kumar, Tanmoy Chakraborty 0002, Animesh Mukherjee 0001, Pawan Goyal 0002 |
CIKM | 6 |
| 2015 | A Stratified Learning Approach for Predicting the Popularity of Twitter Idioms
Suman Kalyan Maity, Pawan Goyal 0002, Animesh Mukherjee 0001 |
ICWSM | 3 |
| 2015 | On the Formation of Circles in Co-authorship NetworksabstractThe availability of an overwhelmingly large amount of bibliographic information including citation and co-authorship data makes it imperative to have a systematic approach that will enable an author to organize her own personal academic network profitably. An effective method could be to have one's co-authorship network arranged into a set of ``circles'', which has been a recent practice for organizing relationships (e.g., friendship) in many online social networks. In this paper, we propose an unsupervised approach to automatically detect circles in an ego network such that each circle represents a densely knit community of researchers. Our model is an unsupervised method which combines a variety of node features and node similarity measures. The model is built from a rich co-authorship network data of more than 8 hundred thousand authors. In the first level of evaluation, our model achieves 13.33% improvement in terms of overlapping modularity compared to the best among four state-of-the-art community detection methods. Further, we conduct a task-based evaluation -- two basic frameworks for collaboration prediction are considered with the circle information (obtained from our model) included in the feature set. Experimental results show that including the circle information detected by our model improves the prediction performance by 9.87% and 15.25% on average in terms of AUC (Area under the ROC) and [email protected] (Precision at Top 20) respectively compared to the case, where the circle information is not present. Tanmoy Chakraborty 0002, Sikhar Patranabis, Pawan Goyal 0002, Animesh Mukherjee 0001 |
KDD | 3 |
| 2013 | A novel neighborhood based document smoothing model for information retrieval
Pawan Goyal 0002, Laxmidhar Behera, T. Martin McGinnity |
Inf. Retr. | 1 |
| 2013 | A Context-Based Word Indexing Model for Document SummarizationabstractExisting models for document summarization mostly use the similarity between sentences in the document to extract the most salient sentences. The documents as well as the sentences are indexed using traditional term indexing measures, which do not take the context into consideration. Therefore, the sentence similarity values remain independent of the context. In this paper, we propose a context sensitive document indexing model based on the Bernoulli model of randomness. The Bernoulli model of randomness has been used to find the probability of the cooccurrences of two terms in a large corpus. A new approach using the lexical association between terms to give a context sensitive weight to the document terms has been proposed. The resulting indexing weights are used to compute the sentence similarity matrix. The proposed sentence similarity measure has been used with the baseline graph-based ranking models for sentence extraction. Experiments have been conducted over the benchmark DUC data sets and it has been shown that the proposed Bernoulli-based sentence similarity model provides consistent improvements over the baseline IntraLink and UniformLink methods [1]. Pawan Goyal 0002, Laxmidhar Behera, T. Martin McGinnity |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | Query Representation through Lexical Association for Information RetrievalabstractA user query for information retrieval (IR) applications may not contain the most appropriate terms (words) as actually intended by the user. This is usually referred to as the term mismatch problem and is a crucial research issue in IR. Using the notion of relevance, we provide a comprehensive theoretical analysis of a parametric query vector, which is assumed to represent the information needs of the user. A lexical association function has been derived analytically using the system relevance criteria. The derivation is further justified using an empirical evidence from the user relevance criteria. Such analytical derivation as presented in this paper provides a proper mathematical framework to the query expansion techniques, which have largely been heuristic in the existing literature. By using the generalized retrieval framework, the proposed query representation model is equally applicable to the vector space model (VSM), Okapi best matching 25 (Okapi BM25), and Language Model (LM). Experiments over various data sets from TREC show that the proposed query representation gives statistically significant improvements over the baseline Okapi BM25 and LM as well as other well-known global query expansion techniques. Empirical results along with the theoretical foundations of the query representation confirm that the proposed model extends the state of the art in global query expansion. Pawan Goyal 0002, Laxmidhar Behera, T. Martin McGinnity |
IEEE Trans. Knowl. Data Eng. | 1 |