Debdoot Mukherjee

dblp:29/2804 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
7since 2021 · last 2026
0009-0009-4424-4449ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A Unified Geospatial Clustering Framework to Identify Varying Density Clusters in E-Commerce Logistics
abstract
In e-commerce logistics, accurate geospatial clustering is essential for optimizing resource allocation, manpower planning, and delivery network design. However, existing density-based clustering approaches, particularly their reliance on heuristic parameter tuning, have been underexplored in datasets with significant density variations, limiting robustness and scalability. This study presents an unsupervised framework that extends DBSCAN by leveraging Gaussian Mixture Models (GMM). First, we propose a method that systematically identifies suitable clustering scales through statistical modeling. Second, the approach iteratively applies DBSCAN to extract clusters from dense to sparse regions, overcoming single-parameter limitations. Finally, we validate the method through large-scale offline experiments using data from over 200 last-mile dispatch centers (LMDC). The results demonstrate the framework’s effectiveness in identifying heterogeneous geographic demand patterns and supporting workforce planning and operational benchmarking. This framework provides a scalable solution to a critical challenge in e-commerce logistics, offering a valuable reference for strategic and operational decision-making.
Arpit Tiwari, Bhavuk Singhal, Anshu Aditya, Aryan Tiwari, Shubham Jain 0012, Debashis Mukherjee, Debdoot Mukherjee
AAAI7
2026 TRUST: Transaction Risk via Unified Sequence and Topology
abstract
Abuse detection in e-commerce platforms is critical for preventing operational losses, particularly for transaction types vulnerable to abuse such as Return-to-Origin (RTO) in Cash-on-Delivery (COD) workflows. Detecting such abuse accurate, real-time decisions to intercept malicious orders before placement, imposing stringent sub-second latency requirements on deployed systems. In this work, we present TRUST, a deployed, production-scale abuse detection system based on a unified architecture of heterogeneous Graph Neural Networks (GNNs) and Transformer-based sequence encoders. This design enables joint reasoning over multi-relational entity interactions and temporal behavioural signals, allowing the model to combine complementary information for effective abuse detection when either modality is sparse or absent. TRUST processes millions of transactions daily with an average inference latency of ~25 ms, achieving a ~9.6% absolute precision improvement over a strong XGBoost baseline in live RTO detection. We report systematic ablation studies across both graph and sequence stages, evaluating GNN variants, sampling strategies, sequence lengths, and positional encoding schemes to guide architectural choices. Deployed end-to-end in a high-throughput environment, TRUST demonstrates that GNN–Transformer cascades can deliver state-of-the-art accuracy, scalability, and operational reliability in real-world abuse detection, offering a reproducible blueprint for similar industry-scale applications.
Rithvik Y., Bhavuk Singhal, Shubham Jain 0012, Akshat Garg, Karan Tanwar, Anshu Aditya, Debashis Mukherjee, Debdoot Mukherjee
AAAI8
2025 GeoIndia V2: A Unified Graph and Language Model for Context-Aware Geocoding
abstract
Geocoding in India presents unique challenges due to the unstructured, multilingual and diverse nature of its address systems. While recent advances in geospatial AI have explored the combination of spatial and semantic cues, existing methods often fall short in effectively integrating both dimensions for robust address resolution. In this work, we propose GeoIndia-V2, an enhanced version of GeoIndia [21], that unifies geospatial and semantic modeling through a novel fusion framework. Our unified model combines the Graphormer architecture [27] and a Pre-trained Transformer based Language Model (PTLM) that is trained from scratch on proprietary Indian address data, using our proposed Key Modulated Cross-Attention (KMCA) mechanism. KMCA enables deep cross-modal interaction between geospatial topology and linguistic structure and allows the model to reason contextually across both geographic and textual dimention, effectively handling the semantic intricacies of Indian addresses-including colloquial usage, inconsistent formatting, and multilinguality. We leverage last-mile e-commerce delivery data to construct a fine-grained graph of neighbourhood connectivity, enabling Graphormer to capture rich spatial relationships. Unlike prior methods that rely on self-loops, we generate graphs dynamically at inference time to exploit Graphormer's topological strength. Additionally, we introduce a generative decoding strategy for predicting hierarchical H3 cells. https://www.uber.com/en-IN/blog/h3/, moving beyond conventional bit-wise classification approaches. To the best of our knowledge, this is the first method to explicitly fuse graph-based geospatial learning with language-driven semantic modeling via cross-attention in the Indian geocoding context. Our approach significantly outperforms existing solutions and marks a substantial advancement toward building scalable real-world geocoding systems for complex address ecosystems like India.
Arpit Tiwari, Bhavuk Singhal, Anshu Aditya, Shubham Jain 0012, Debashis Mukherjee, Debdoot Mukherjee
CIKM6
2024 Study of Abuse Detection in Continuous Speech for Indian Languages
abstract
The presence of abusive content on social media platforms poses a significant challenge in maintaining a positive online environment for millions of users. While automatic abuse detection has seen extensive use in the text domain, audio abuse detection remains relatively unexplored. This paper addresses the modeling challenges inherent in identifying offensive content within real-life audio recordings, particularly in the multilingual context of two Indian languages. We introduce a cascaded model that combines an automatic speech recognition system with textual keyword spotting and compare it with an end-to-end model utilizing audio-level feature embeddings and neural classifiers. Our findings, based on the ADIMA dataset, demonstrate that both methods achieve similar discriminability, but the cascaded linguistic approach stands out for its enhanced explainability and the potential to leverage well-established Natural Language Processing techniques for abuse detection.
Rini A. Sharon, Debdoot Mukherjee
ICASSP2
2022 3MASSIV: Multilingual, Multimodal and Multi-Aspect dataset of Social Media Short Videos
abstract
We present 3MASSIV, a multilingual, multimodal and multi-aspect, expertly-annotated dataset of diverse short videos extracted from short-video social media platform - Moj. 3MASSIV comprises of 50k short videos (20 seconds average duration) and 100K unlabeled videos in 11 different languages and captures popular short video trends like pranks, fails, romance, comedy expressed via unique audio-visual formats like self-shot videos, reaction videos, lip-synching, self-sung songs, etc. 3MASSIV presents an opportunity for multimodal and multilingual semantic understanding on these unique videos by annotating them for concepts, affective states, media types, and audio language. We present a thorough analysis of 3MASSIV and highlight the variety and unique aspects of our dataset compared to other contemporary popular datasets with strong baselines. We also show how the social media content in 3MASSIV is dynamic and temporal in nature, which can be used for semantic understanding tasks and cross-lingual analysis.
Vikram Gupta, Trisha Mittal, Puneet Mathur, Mayank Maheshwari, Aniket Bera, Debdoot Mukherjee, Dinesh Manocha
CVPR7
2022 ADIMA: Abuse Detection In Multilingual Audio
abstract
Abusive content detection in spoken text can be addressed by performing Automatic Speech Recognition (ASR) and leveraging advancements in natural language processing. However, ASR models introduce latency and often perform sub-optimally for abusive words as they are underrepresented in training corpora and not spoken clearly or completely. Exploration of this problem entirely in the audio domain has largely been limited by the lack of audio datasets. Building on these challenges, we propose ADIMA, a novel, linguistically diverse, ethically sourced, expert annotated and well- balanced multilingual abuse detection audio dataset comprising of 11,775 audio samples in 10 Indic languages spanning 65 hours and spoken by 6,446 unique users. Through quantitative experiments across monolingual and cross-lingual zeroshot settings, we take the first step in democratizing audio based content moderation in Indic languages and set forth our dataset to pave future work. Dataset and code are available at: https://github.com/ShareChatAI/Adima
Vikram Gupta, Rini A. Sharon, Ramit Sawhney, Debdoot Mukherjee
ICASSP4
2022 Multilingual and Multimodal Abuse Detection
abstract
The presence of abusive content on social media platforms is undesirable as it severely impedes healthy and safe social media interactions.While automatic abuse detection has been widely explored in textual domain, audio abuse detection still remains unexplored.In this paper, we attempt abuse detection in conversational audio from a multimodal perspective in a multilingual social media setting.Our key hypothesis is that along with the modelling of audio, incorporating discriminative information from other modalities can be highly beneficial for this task.Our proposed method, MADA, explicitly focuses on two modalities other than the audio itself, namely, the underlying emotions expressed in the abusive audio and the semantic information encapsulated in the corresponding textual form.Observations prove that MADA demonstrates gains over audio-only approaches on the ADIMA dataset.We test the proposed approach on 10 different languages and observe consistent gains in the range 0.6%-5.2%by leveraging multiple modalities.We also perform extensive ablation experiments for studying the contributions of every modality and observe the best results while leveraging all the modalities together.Additionally, we perform experiments to empirically confirm that there is a strong correlation between underlying emotions and abusive behaviour.
Rini A. Sharon, Heet Shah, Debdoot Mukherjee, Vikram Gupta
INTERSPEECH3
2020 Understanding Chat Messages for Sticker Recommendation in Messaging Apps
abstract
Stickers are popularly used in messaging apps such as Hike to visually express a nuanced range of thoughts and utterances to convey exaggerated emotions. However, discovering the right sticker from a large and ever expanding pool of stickers while chatting can be cumbersome. In this paper, we describe a system for recommending stickers in real time as the user is typing based on the context of the conversation. We decompose the sticker recommendation (SR) problem into two steps. First, we predict the message that the user is likely to send in the chat. Second, we substitute the predicted message with an appropriate sticker. Majority of Hike's messages are in the form of text which is transliterated from users' native language to the Roman script. This leads to numerous orthographic variations of the same message and makes accurate message prediction challenging. To address this issue, we learn dense representations of chat messages employing character level convolution network in an unsupervised manner. We use them to cluster the messages that have the same meaning. In the subsequent steps, we predict the message cluster instead of the message. Our approach does not depend on human labelled data (except for validation), leading to fully automatic updation and tuning pipeline for the underlying models. We also propose a novel hybrid message prediction model, which can run with low latency on low-end phones that have severe computational limitations. Our described system has been deployed for more than 6 months and is being used by millions of users along with hundreds of thousands of expressive stickers.
Abhishek Laddha, Mohamed Hanoosh, Debdoot Mukherjee, Parth Patwa, Ankur Narang
AAAI3
2020 Learning Multigraph Node Embeddings Using Guided Lévy Flights
Aman Roy, Vinayak Kumar, Debdoot Mukherjee, Tanmoy Chakraborty 0002
PAKDD (1)3
2019 Heterogeneous Edge Embedding for Friend Recommendation
Janu Verma, Srishti Gupta 0003, Debdoot Mukherjee, Tanmoy Chakraborty 0002
ECIR (2)3
2019 Contextual Typeahead Sticker Suggestions on Hike Messenger
abstract
In this demonstration, we present Hike's sticker recommendation system, which helps users choose the right sticker to substitute the next message that they intend to send in a chat. We describe how the system addresses the issue of numerous orthographic variations for chat messages and operates under 20 milliseconds with low CPU and memory footprint on device.
Mohamed Hanoosh, Abhishek Laddha, Debdoot Mukherjee
IJCAI3
2013 A Case Based Approach to Serve Information Needs in Knowledge Intensive Processes
Debdoot Mukherjee, Jeanette Blomberg, Rama Akkiraju, Dinesh Raghu, Monika Gupta 0002, Sugata Ghosal, Taiga Nakamura
ICSOC1
2013 Bug resolution catalysts: identifying essential non-committers from bug repositories
abstract
Bugs are inevitable in software projects. Resolving bugs is the primary activity in software maintenance. Developers, who fix bugs through code changes, are naturally important participants in bug resolution. However, there are other participants in these projects who do not perform any code commits. They can be reporters reporting bugs; people having a deep technical know-how of the software and providing valuable insights on how to solve the bug; bug-tossers who re-assign the bugs to the right set of developers. Even though all of them act on the bugs by tossing and commenting, not all of them may be crucial for bug resolution. In this paper, we formally define essential non-committers and try to identify these bug resolution catalysts. We empirically study 98304 bug reports across 11 open source and 5 commercial software projects for validating the existence of such catalysts. We propose a network analysis based approach to construct a Minimal Essential Graph that identifies such people in a project. Finally, we suggest ways of leveraging this information for bug triaging and bug report summarization.
Senthil Mani, Seema Nagar, Debdoot Mukherjee, Ramasuri Narayanam, Vibha Sinha, Amit Anil Nanavati
MSR3
2013 Which work-item updates need your response?
abstract
Work-item notifications alert the team collaborating on a work-item about any update to the work-item (e.g., addition of comments, change in status). However, as software professionals get involved with multiple tasks in project(s), they are inundated by too many notifications from the work-item tool. Users are upset that they often miss the notifications that solicit their response in the crowd of mostly useless ones. We investigate the severity of this problem by studying the work-item repositories of two large collaborative projects and conducting a user study with one of the project teams. We find that, on an average, only 1 out of every 5 notifications that are received by the users require a response from them. We propose TWINY - a machine learning based approach to predict whether a notification will prompt any action from its recipient. Such a prediction can help to suitably mark up notifications and to decide whether a notification needs to be sent out immediately or be bundled in a message digest. We conduct empirical studies to evaluate the efficacy of different classification techniques in this setting. We find that incremental learning algorithms are ideally suited, and ensemble methods appear to give the best results in terms of prediction accuracy.
Debdoot Mukherjee, Malika Garg
MSR1
2011 Serving Information Needs in Business Process Consulting
Monika Gupta 0002, Debdoot Mukherjee, Senthil Mani, Vibha Sinha, Saurabh Sinha 0001
BPM2
2011 Using MATCON to generate CASE tools that guide deployment of pre-packaged applications
abstract
The complex process of adapting pre-packaged applications, such as Oracle or SAP, to an organization's needs is full of challenges. Although detailed, structured, and well-documented methods govern this process, the consulting team implementing the method must spend a huge amount of manual effort to make sure the guidelines of the method are followed as intended by the method author. MATCON breaks down the method content, documents, templates, and work products into reusable objects, and enables them to be cataloged and indexed so these objects can be easily found and reused on subsequent projects. By using models and meta-modeling the reusable methods, we automatically produce a CASE tool to apply these methods, thereby guiding consultants through this complex process. The resulting tool helps consultants create the method deliverables for the initial phases of large customization projects. Our MATCON output, referred to as Consultant Assistant, has shown significant savings in training costs, a 20 - 30% improvement in productivity, and positive results in large Oracle and SAP implementations.
Elad Fein, Natalia Razinkov, Shlomit Shachor, Pietro Mazzoleni, SweeFen Goh, Richard Goodwin, Manisha Bhandar, Shyh-Kwei Chen, Juhnyoung Lee, Vibha Sinha, Senthil Mani, Debdoot Mukherjee, Biplav Srivastava, Pankaj Dhoolia
ICSE12
2011 A Service Model for Development and Test Clouds
Debdoot Mukherjee, Monika Gupta 0002, Vibha Sinha, Nianjun Zhou
ICSOC1
2010 From Informal Process Diagrams to Formal Process Models
Debdoot Mukherjee, Pankaj Dhoolia, Saurabh Sinha 0001, Aubrey J. Rembert, Mangala Gowri Nanda
BPM1
2009 Efficient Testing of Service-Oriented Applications Using Semantic Service Stubs
abstract
Service-oriented applications can be expensive to test because services are hosted remotely, are potentially shared among many users, and may have costs associated with their invocation. In this paper, we present an approach for reducing the costs of testing such applications. The key observation underlying our approach is that certain aspects of an application can be tested using locally deployed semantic service stubs, instead of actual remote services.A semantic service stub incorporates some of the service functionality, such as verifying preconditions and generating output messages based on post conditions. We illustrate how semantic stubs can enable the client test suite to be partitioned into subsets, some of which need not be executed using remote services. We also present a case study that demonstrates the feasibility of the approach, and potential cost savings for testing. The main benefits of our approach are that it can (1) reduce the number of test cases that need to be run to invoke remote services, (2) ensure that certain aspects of application functionality are well-tested before service integration occurs.
Senthil Mani, Vibha Sinha, Saurabh Sinha 0001, Pankaj Dhoolia, Debdoot Mukherjee, Soham Chakraborty 0001
ICWS5
2008 Determining QoS of WS-BPEL Compositions
Debdoot Mukherjee, Pankaj Jalote, Mangala Gowri Nanda
ICSOC1