EDBT 2026 Demo / reviewers in the wild / expert
Tirthankar Ghosal
dblp:215/3590
· DBLP profile ↗
28ranked-venue papers
7as first author
24since 2021 · last 2026
0000-0002-2358-522XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 15 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Operational Safety via Agentic Dialogue Hazard Identification AnalysisabstractOperational safety in high-stakes domains such as industrial process control, autonomous, and safety-critical systems demand reliable hazard identification. While large language models (LLMs) have shown promise in automating safety analysis tasks, single-turn, monolithic inference is brittle: it lacks the self-correction, deliberation, and contextual refinement that safety engineers apply iteratively. In this paper, we introduce HAZDIAL, a framework that investigates whether structured agentic dialogue (multi-agent, multi-turn interactions) improves the quality of NLP-based hazard identification over single-pass baselines. We systematically compare two dialogue modalities: adversarial debate and constructive discussion, and propose an genetic algorithm-based agentic interaction optimization. We evaluate all configurations against a curated golden dataset using standard classification metrics (accuracy, precision, recall, F1) and a novel dialogue metrics. This work advances the intersection of dialogue systems, multi-agent reasoning, and AI safety, providing empirical evidence for dialogue-driven hazard analysis. Sanjay Das, Ran Elgedawy, Ethan Seefried, Ryan Burchfield, Tirthankar Ghosal |
SIGDIAL | 5 |
| 2025 | PROV-AGENT: Unified Provenance for Tracking AI Agent Interactions in Agentic WorkflowsabstractLarge Language Models (LLMs) and other foundation models are increasingly used as the core of AI agents. In agentic workflows, these agents plan tasks, interact with humans and peers, and influence scientific outcomes across federated and heterogeneous environments. However, agents can hallucinate or reason incorrectly, propagating errors when one agent’s output becomes another’s input. Thus, assuring that agents’ actions are transparent, traceable, reproducible, and reliable is critical to assess hallucination risks and mitigate their workflow impacts. While provenance techniques have long supported these principles, existing methods fail to capture and relate agent-centric metadata such as prompts, responses, and decisions with the broader workflow context and downstream outcomes. In this paper, we introduce PROV-AGENT, a provenance model that extends W3C PROV and leverages the Model Context Protocol (MCP) and data observability to integrate agent interactions into end-to-end workflow provenance. Our contributions include: (1) a provenance model tailored for agentic workflows, (2) a near real-time, open-source system for capturing agentic provenance, and (3) a cross-facility evaluation spanning edge, cloud, and HPC environments, demonstrating support for critical provenance queries and agent reliability analysis. Renan Souza 0001, Amal Gueroudji, Stephen DeWitt, Daniel Rosendo, Tirthankar Ghosal, Robert B. Ross, Prasanna Balaprakash, Rafael Ferreira da Silva |
eScience | 5 |
| 2025 | Can Large Language Models Unlock Novel Scientific Research Ideas?abstractThe widespread adoption of Large Language Models (LLMs) and publicly available Chat-GPT have marked a significant turning point in the integration of Artificial Intelligence (AI) into people's everyday lives.This study examines the ability of Large Language Models (LLMs) to generate future research ideas from scientific papers.Unlike tasks such as summarization or translation, idea generation lacks a clearly defined reference set or structure, making manual evaluation the default standard.However, human evaluation in this setting is extremely challenging -it requires substantial domain expertise, contextual understanding of the paper, and awareness of the current research landscape.This makes it time-consuming, costly, and fundamentally non-scalable, particularly as new LLMs are being released at a rapid pace.Currently, there is no automated evaluation metric specifically designed for this task.To address this gap, we propose two automated evaluation metrics: Idea Alignment Score (IAScore) and Idea Distinctness Index.We further conducted human evaluation to assess the novelty, relevance, and feasibility of the generated future research ideas.This investigation offers insights into the evolving role of LLMs in idea generation, highlighting both its capability and limitations.Our work contributes to the ongoing efforts in evaluating and utilizing language models for generating future research ideas.We make our datasets and codes publicly available 1 ."Innovation is seeing what everybody has seen and thinking what nobody has thought" -Dr. Sandeep Kumar 0009, Tirthankar Ghosal, Vinayak Goyal, Asif Ekbal |
EMNLP | 2 |
| 2025 | chatHPC: Empowering HPC users with large language models
Junqi Yin, Jesse Hines, Emily J. Herron, Tirthankar Ghosal, Suzanne Prentice, Vanessa Lama, Feiyi Wang |
J. Supercomput. | 4 |
| 2024 | Longform Multimodal Lay Summarization of Scientific Papers: Towards Automatically Generating Science Blogs from Research ArticlesabstractScience communication, in layperson’s terms, is essential to reach the general population and also maximize the impact of underlying scientific research. Hence, good science blogs and journalistic reviews of research articles are so well-read and critical to conveying science. Scientific blogging goes beyond traditional research summaries, offering experts a platform to articulate findings in layperson’s terms. It bridges the gap between intricate research and its comprehension by the general public, policymakers, and other researchers. Amid the rapid expansion of scientific data and the accelerating pace of research, credible science blogs serve as vital artifacts for evidence-based information to the general non-expert audience. However, writing a scientific blog or even a short lay summary requires significant time and effort. Here, we are intrigued what if the process of writing a scientific blog based on a given paper could be semi-automated to produce the first draft? In this paper, we introduce a novel task of Artificial Intelligence (AI)-based science blog generation from a research article. We leverage the idea that presentations and science blogs share a symbiotic relationship in their aim to clarify and elucidate complex scientific concepts. Both rely on visuals, such as figures, to aid comprehension. With this motivation, we create a new dataset of science blogs using the presentation transcript and the corresponding slides. We create a dataset containing a paper’s presentation transcript and figures annotated from nearly 3000 papers. We then propose a multimodal attention model to generate a blog text and select the most relevant figures to explain a research article in layperson’s terms, essentially a science blog. Our experimental results with respect to both automatic and human evaluation metrics show the effectiveness of our proposed approach and the usefulness of our proposed dataset. Sandeep Kumar 0009, Guneet Singh Kohli, Tirthankar Ghosal, Asif Ekbal |
LREC/COLING | 3 |
| 2024 | 'Quis custodiet ipsos custodes?' Who will watch the watchmen? On Detecting AI-generated peer-reviewsabstractThe integrity of the peer-review process is vital for maintaining scientific rigor and trust within the academic community.With the steady increase in the usage of large language models (LLMs) like ChatGPT in academic writing, there is a growing concern that AI-generated texts could compromise scientific publishing, including peer-reviews.Previous works have focused on generic AI-generated text detection or have presented an approach for estimating the fraction of peer-reviews that can be AIgenerated.Our focus here is to solve a realworld problem by assisting the editor or chair in determining whether a review is written by ChatGPT or not.To address this, we introduce the Term Frequency (TF) model, which posits that AI often repeats tokens, and the Review Regeneration (RR) model, which is based on the idea that ChatGPT generates similar outputs upon re-prompting.We stress test these detectors against token attack and paraphrasing.Finally, we propose an effective defensive strategy to reduce the effect of paraphrasing on our models.Our findings suggest both our proposed methods perform better than the other AI text detectors.Our RR model is more robust, although our TF model performs better than the RR model without any attacks.We make our code, dataset, and model public 12 . Sandeep Kumar 0009, Mohit Sahu, Vardhan Gacche, Tirthankar Ghosal, Asif Ekbal |
EMNLP | 4 |
| 2024 | Emotion aided multi-task framework for video embedded misinformation detection
Rina Kumari, Vipin Gupta, Nischal Ashok, Tirthankar Ghosal, Asif Ekbal |
Multim. Tools Appl. | 4 |
| 2023 | When Reviewers Lock Horns: Finding Disagreements in Scientific Peer ReviewsabstractTo this date, the efficacy of the scientific publishing enterprise fundamentally rests on the strength of the peer review process.The journal editor or the conference chair primarily relies on the expert reviewers' assessment, identify points of agreement and disagreement and try to reach a consensus to make a fair and informed decision on whether to accept or reject a paper.However, with the escalating number of submissions requiring review, especially in toptier Artificial Intelligence (AI) conferences, the editor/chair, among many other works, invests a significant, sometimes stressful effort to mitigate reviewer disagreements.Here in this work, we introduce a novel task of automatically identifying contradictions among reviewers on a given article.To this end, we introduce Con-traSciView, a comprehensive review-pair contradiction dataset on around 8.5k papers (with around 28k review pairs containing nearly 50k review pair comments) from the open reviewbased ICLR and NeurIPS conferences.We further propose a baseline model that detects contradictory statements from the review pairs.To the best of our knowledge, we make the first attempt to identify disagreements among peer reviewers automatically.We make our dataset and code public for further investigations 1 . Sandeep Kumar 0009, Tirthankar Ghosal, Asif Ekbal |
EMNLP | 2 |
| 2023 | Identifying multimodal misinformation leveraging novelty detection and emotion recognition
Rina Kumari, Nischal Ashok, Pawan Kumar Agrawal, Tirthankar Ghosal, Asif Ekbal |
J. Intell. Inf. Syst. | 4 |
| 2022 | How Confident Was Your Reviewer? Estimating Reviewer Confidence from Peer Review Texts
Prabhat Kumar Bharti, Tirthankar Ghosal, Mayank Agrawal, Asif Ekbal |
DAS | 2 |
| 2022 | BetterPR: A Dataset for Estimating the Constructiveness of Peer Review Comments
Prabhat Kumar Bharti, Tirthankar Ghosal, Mayank Agarwal, Asif Ekbal |
TPDL | 2 |
| 2022 | Investigations on Meta Review Generation from Peer Review Texts Leveraging Relevant Sub-tasks in the Peer Review Pipeline
Asheesh Kumar, Tirthankar Ghosal, Saprativa Bhattacharjee, Asif Ekbal |
TPDL | 2 |
| 2022 | ELITR Minuting Corpus: A Novel Dataset for Automatic Minuting from Multi-Party Meetings in English and CzechabstractTaking minutes is an essential component of every meeting, although the goals, style, and procedure of this activity (“minuting” for short) can vary. Minuting is a rather unstructured writing activity and is affected by who is taking the minutes and for whom the intended minutes are. With the rise of online meetings, automatic minuting would be an important benefit for the meeting participants as well as for those who might have missed the meeting. However, automatically generating meeting minutes is a challenging problem due to a variety of factors including the quality of automatic speech recorders (ASRs), availability of public meeting data, subjective knowledge of the minuter, etc. In this work, we present the first of its kind dataset on Automatic Minuting. We develop a dataset of English and Czech technical project meetings which consists of transcripts generated from ASRs, manually corrected, and minuted by several annotators. Our dataset, AutoMin, consists of 113 (English) and 53 (Czech) meetings, covering more than 160 hours of meeting content. Upon acceptance, we will publicly release (aaa.bbb.ccc) the dataset as a set of meeting transcripts and minutes, excluding the recordings for privacy reasons. A unique feature of our dataset is that most meetings are equipped with more than one minute, each created independently. Our corpus thus allows studying differences in what people find important while taking the minutes. We also provide baseline experiments for the community to explore this novel problem further. To the best of our knowledge AutoMin is probably the first resource on minuting in English and also in a language other than English (Czech). Anna Nedoluzhko, Muskaan Singh, Marie Hledíková, Tirthankar Ghosal, Ondrej Bojar |
LREC | 4 |
| 2022 | Novelty Detection in Community Question Answering Forums
Tirthankar Ghosal, Vignesh Edithal, Tanik Saikh, Saprativa Bhattacharjee, Asif Ekbal, Pushpak Bhattacharyya |
PACLIC | 1 |
| 2022 | Automatic Minuting: A Pipeline Method for Generating Minutes from Multi-Party Meeting Proceedings
Kartik Shinde, Tirthankar Ghosal, Muskaan Singh, Ondrej Bojar |
PACLIC | 2 |
| 2022 | Novelty Detection: A Perspective from Natural Language ProcessingabstractAbstract The quest for new information is an inborn human trait and has always been quintessential for human survival and progress. Novelty drives curiosity, which in turn drives innovation. In Natural Language Processing (NLP), Novelty Detection refers to finding text that has some new information to offer with respect to whatever is earlier seen or known. With the exponential growth of information all across the Web, there is an accompanying menace of redundancy. A considerable portion of the Web contents are duplicates, and we need efficient mechanisms to retain new information and filter out redundant information. However, detecting redundancy at the semantic level and identifying novel text is not straightforward because the text may have less lexical overlap yet convey the same information. On top of that, non-novel/redundant information in a document may have assimilated from multiple source documents, not just one. The problem surmounts when the subject of the discourse is documents, and numerous prior documents need to be processed to ascertain the novelty/non-novelty of the current one in concern. In this work, we build upon our earlier investigations for document-level novelty detection and present a comprehensive account of our efforts toward the problem. We explore the role of pre-trained Textual Entailment (TE) models to deal with multiple source contexts and present the outcome of our current investigations. We argue that a multipremise entailment task is one close approximation toward identifying semantic-level non-novelty. Our recent approach either performs comparably or achieves significant improvement over the latest reported results on several datasets and across several related tasks (paraphrasing, plagiarism, rewrite). We critically analyze our performance with respect to the existing state of the art and show the superiority and promise of our approach for future investigations. We also present our enhanced dataset TAP-DLND 2.0 and several baselines to the community for further research on document-level novelty detection. Tirthankar Ghosal, Tanik Saikh, Tameesh Biswas, Asif Ekbal, Pushpak Bhattacharyya |
Comput. Linguistics | 1 |
| 2022 | What the fake? Probing misinformation detection standing on the shoulder of novelty and emotion
Rina Kumari, Nischal Ashok, Tirthankar Ghosal, Asif Ekbal |
Inf. Process. Manag. | 3 |
| 2021 | Attend to Your Review: A Deep Neural Network to Extract Aspects from Peer Reviews
Rajeev Verma, Kartik Shinde, Hardik Arora, Tirthankar Ghosal |
ICONIP (6) | 4 |
| 2021 | A Multitask Learning Approach for Fake News Detection: Novelty, Emotion, and Sentiment Lend a Helping HandabstractThe recent explosion in false information on social media has led to intensive research on automatic fake news detection models and fact-checkers. Fake news and misinformation, due to its peculiarity and rapid dissemination, have posed many interesting challenges to the Natural Language Processing (NLP) and Machine Learning (ML) community. Admissible literature shows that novel information includes the element of surprise, which is the principal characteristic for the amplification and virality of misinformation. Novel and emotional information attracts immediate attention in the reader. Emotion is the presentation of a certain feeling or sentiment. Sentiment helps an individual to convey his emotion through expression and hence the two are co-related. Thus, Novelty of the news item and thereafter detecting the Emotional state and Sentiment of the reader appear to be three key ingredients, tightly coupled with misinformation. In this paper we propose a deep multitask learning model that jointly performs novelty detection, emotion recognition, sentiment prediction, and misinformation detection. Our proposed model achieves the state-of-the-art(SOTA) performance for fake news detection on three benchmark datasets, viz. ByteDance, Fake News Challenge(FNC), and Covid-Stance with 11.55%, 1.58%, and 21.76% improvement in accuracy, respectively. The proposed approach also shows the efficacy over the single-task framework with an accuracy gain of 11.53, 28.62, and 14.31 percentage points for the above three datasets. The source code is available at https://github.com/Nish-19/Multitask-Fake-News-NES. Rina Kumari, Nischal Ashok, Tirthankar Ghosal, Asif Ekbal |
IJCNN | 3 |
| 2021 | ARGUABLY @ AI Debater-NLPCC 2021 Task 3: Argument Pair Extraction from Peer Review and Rebuttals
Guneet Singh Kohli, Prabsimran Kaur, Muskaan Singh, Tirthankar Ghosal, Prashant Singh Rana |
NLPCC (2) | 4 |
| 2021 | A Neuro-Symbolic Approach for Question Answering on Research Articles
Komal Gupta, Tirthankar Ghosal, Asif Ekbal |
PACLIC | 2 |
| 2021 | An Empirical Performance Analysis of State-of-the-Art Summarization Models for Automatic Minuting
Muskaan Singh, Tirthankar Ghosal, Ondrej Bojar |
PACLIC | 2 |
| 2021 | Misinformation detection using multitask learning with mutual learning for novelty detection and emotion recognition
Rina Kumari, Nischal Ashok, Tirthankar Ghosal, Asif Ekbal |
Inf. Process. Manag. | 3 |
| 2021 | Is your document novel? Let attention guide you. An attention-based model for document-level novelty detectionabstractAbstract Detecting, whether a document contains sufficient new information to be deemed as novel , is of immense significance in this age of data duplication. Existing techniques for document-level novelty detection mostly perform at the lexical level and are unable to address the semantic-level redundancy. These techniques usually rely on handcrafted features extracted from the documents in a rule-based or traditional feature-based machine learning setup. Here, we present an effective approach based on neural attention mechanism to detect document-level novelty without any manual feature engineering. We contend that the simple alignment of texts between the source and target document(s) could identify the state of novelty of a target document. Our deep neural architecture elicits inference knowledge from a large-scale natural language inference dataset, which proves crucial to the novelty detection task. Our approach is effective and outperforms the standard baselines and recent work on document-level novelty detection by a margin of $\sim$ 3% in terms of accuracy. Tirthankar Ghosal, Vignesh Edithal, Asif Ekbal, Pushpak Bhattacharyya, Srinivasa Satya Sameer Kumar Chivukula, George Tsatsaronis 0001 |
Nat. Lang. Eng. | 1 |
| 2019 | DeepSentiPeer: Harnessing Sentiment in Review Texts to Recommend Peer Review DecisionsabstractAutomatically validating a research artefact is one of the frontiers in Artificial Intelligence (AI) that directly brings it close to competing with human intellect and intuition.Although criticized sometimes, the existing peer review system still stands as the benchmark of research validation.The present-day peer review process is not straightforward and demands profound domain knowledge, expertise, and intelligence of human reviewer(s), which is somewhat elusive with the current state of AI.However, the peer review texts, which contains rich sentiment information of the reviewer, reflecting his/her overall attitude towards the research in the paper, could be a valuable entity to predict the acceptance or rejection of the manuscript under consideration.Here in this work, we investigate the role of reviewers sentiments embedded within peer review texts to predict the peer review outcome.Our proposed deep neural architecture takes into account three channels of information: the paper, the corresponding reviews, and the review polarity to predict the overall recommendation score as well as the final decision.We achieve significant performance improvement over the baselines (∼ 29% error reduction) proposed in a recently released dataset of peer reviews.An AI of this kind could assist the editors/program chairs as an additional layer of confidence in the final decision making, especially when non-responding/missing reviewers are frequent in present day peer review. Tirthankar Ghosal, Rajeev Verma, Asif Ekbal, Pushpak Bhattacharyya |
ACL (1) | 1 |
| 2019 | To Comprehend the New: On Measuring the Freshness of a DocumentabstractDetecting the novelty or freshness of an entire document is essential in this age of data duplication and semantic-level redundancy all across the web. Current techniques for the problem mostly root on handcrafted similarity and divergence based measures to classify a document as novel or non-novel. However, document-level novelty detection is relatively less explored in literature if compared to its sentence-level counterpart. In this work, we present a deep neural architecture to automatically predict the amount of new information contained in a document in the form of a novelty score. Along with, we offer a dataset of more than 7500 documents, annotated at the sentence-level to facilitate further research. Our approach which learns the notion of novelty and redundancy only from the data achieves significant performance improvement over the existing methods and adopted baselines (~17% error reduction). Also, our approach complies with the Two-Stage theory of human recall essential to comprehend new information. Tirthankar Ghosal, Asif Ekbal, Pushpak Bhattacharyya |
IJCNN | 1 |
| 2018 | Novelty Goes Deep. A Deep Neural Solution To Document Level Novelty DetectionabstractThe rapid growth of documents across the web has necessitated finding means of discarding redundant documents and retaining novel ones. Capturing redundancy is challenging as it may involve investigating at a deep semantic level. Techniques for detecting such semantic redundancy at the document level are scarce. In this work we propose a deep Convolutional Neural Networks (CNN) based model to classify a document as novel or redundant with respect to a set of relevant documents already seen by the system. The system is simple and do not require any manual feature engineering. Our novel scheme encodes relevant and relative information from both source and target texts to generate an intermediate representation which we coin as the Relative Document Vector (RDV). The proposed method outperforms the existing state-of-the-art on a document-level novelty detection dataset by a margin of ∼5% in terms of accuracy. We further demonstrate the effectiveness of our approach on a standard paraphrase detection dataset where paraphrased passages closely resemble to semantically redundant documents. Tirthankar Ghosal, Vignesh Edithal, Asif Ekbal, Pushpak Bhattacharyya, George Tsatsaronis 0001, Srinivasa Satya Sameer Kumar Chivukula |
COLING | 1 |
| 2018 | TAP-DLND 1.0 : A Corpus for Document Level Novelty Detection
Tirthankar Ghosal, Amitra Salam, Swati Tiwari, Asif Ekbal, Pushpak Bhattacharyya |
LREC | 1 |