Subhadeep Maji

dblp:216/7323 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 84% Recommender systems · 16%
Artificial intelligence
1 paper
Knowledge representation and reasoning · 61% Information extraction and text analysis · 30% Language models and text generation · 9%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
e-commerce search
0.822020
Predicting Session Length for Product Search on E-commerce Platform · SIGIR 2020
Addressing Vocabulary Gap in E-commerce Search · SIGIR 2019
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
first-order logic constraints
0.412020
Logic Constrained Pointer Networks for Interpretable Textual Similarity · IJCAI 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
logical constraints
0.412020
Logic Constrained Pointer Networks for Interpretable Textual Similarity · IJCAI 2020
Natural language and speech › Information extraction and text analysis › text similarity › semantic similarity
semantic textual similarity
0.412020
Logic Constrained Pointer Networks for Interpretable Textual Similarity · IJCAI 2020
Information retrieval › user behavior
search session analysis
0.412020
Predicting Session Length for Product Search on E-commerce Platform · SIGIR 2020
Recommender systems › user modeling
user intent modeling
0.412020
Predicting Session Length for Product Search on E-commerce Platform · SIGIR 2020
Information retrieval
web search
0.412020
Predicting Session Length for Product Search on E-commerce Platform · SIGIR 2020
Information retrieval › e-commerce search
query-product matching
0.412019
Addressing Vocabulary Gap in E-commerce Search · SIGIR 2019
Information retrieval
query understanding
0.412019
Addressing Vocabulary Gap in E-commerce Search · SIGIR 2019
Information retrieval
vocabulary mismatch
0.412019
Addressing Vocabulary Gap in E-commerce Search · SIGIR 2019
Information retrieval › evaluation
retrieval effectiveness
0.112019
Addressing Vocabulary Gap in E-commerce Search · SIGIR 2019

Methods — techniques the papers use, named apart from their topics

weibull distribution · 0.4sentinel gating · 0.4recurrent neural network · 0.4pointer network · 0.4BERT · 0.4a/b experiments · 0.4
YearPublicationVenuePosition
2025 Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
abstract
The use of human speech to train LLMs poses privacy concerns due to these models' ability to generate samples that closely resemble artifacts in the training data. We propose a speaker privacy-preserving representation learning method through the Universal Speech Codec (USC), a computationally efficient codec that disentangles speech into: (i) privacy-preserving semantically rich representations, capturing content and speech paralinguistics, and (ii) residual acoustic and speaker representations that enable high-fidelity reconstruction. Evaluations show that USC's semantic representation preserves content, prosody, and sentiment, while removing identifiable traits. Additionally, we present an evaluation methodology for measuring privacy-preserving properties. We compare USC against other speech codecs and demonstrate its effectiveness on privacy-preserving representation learning, showcasing the trade-offs between speaker anonymization and paralinguistics retention.1
Biel Tura Vecino, Subhadeep Maji, Aravind Varier, Antonio Bonafonte, Ivan Valles, Michael Owen, Constantinos Papayiannis, Leif Rädel, Grant P. Strimel, Oluwaseyi Feyisetan, Roberto Barra-Chicote, Ariya Rastrow, Volker Leutnant, Trevor Wood
INTERSPEECH2
2023 Unsupervised Domain Adaptation With Global and Local Graph Neural Networks Under Limited Supervision and Its Application to Disaster Response
abstract
Identification and categorization of social media posts generated during disasters are crucial to reduce the suffering of the affected people. However, the lack of labeled data is a significant bottleneck in learning an effective categorization system for a disaster. This motivates us to study the problem as unsupervised domain adaptation (UDA) between a previous disaster with labeled data (source) and a current disaster (target). However, if the amount of labeled data available is limited, it restricts the learning capabilities of the model. To handle this challenge, we use limited labeled data along with abundantly available unlabeled data, generated during a source disaster to propose a novel two-part graph neural network (GNN). The first part extracts domain-agnostic global information by constructing a token-level graph across domains and the second part preserves local instance-level semantics. In our experiments, we show that the proposed method outperforms state-of-the-art techniques by 2.74% weighted F1 score on average on two standard public datasets in the area of disaster management. We also report experimental results for granular actionable multilabel classification datasets in disaster domain for the first time, on which we outperform BERT by 3.00% on average w.r.t. weighted F1. Additionally, we show that our approach can retain performance when minimal labeled data are available.
Samujjwal Ghosh, Subhadeep Maji, Maunendra Sankar Desarkar
IEEE Trans. Comput. Soc. Syst.2
2021 Reproducibility, Replicability and Beyond: Assessing Production Readiness of Aspect Based Sentiment Analysis in the Wild
Rajdeep Mukherjee, Shreyas Shetty, Subrata Chattopadhyay, Subhadeep Maji, Samik Datta, Pawan Goyal 0002
ECIR (2)4
2020 A Regularised Intent Model for Discovering Multiple Intents in E-Commerce Tail Queries
Subhadeep Maji, Priyank Patel, Bharat Thakarar, Krishna Azad Tripathi
ECIR (1)1
2020 Logic Constrained Pointer Networks for Interpretable Textual Similarity
abstract
Systematically discovering semantic relationships in text is an important and extensively studied area in Natural Language Processing, with various tasks such as entailment, semantic similarity, etc. Decomposability of sentence-level scores via subsequence alignments has been proposed as a way to make models more interpretable. We study the problem of aligning components of sentences leading to an interpretable model for semantic textual similarity. In this paper, we introduce a novel pointer network based model with a sentinel gating function to align constituent chunks, which are represented using BERT. We improve this base model with a loss function to equally penalize misalignments in both sentences, ensuring the alignments are bidirectional. Finally, to guide the network with structured external knowledge, we introduce first-order logic constraints based on ConceptNet and syntactic knowledge. The model achieves an F1 score of 97.73 and 96.32 on the benchmark SemEval datasets for the chunk alignment task, showing large improvements over the existing solutions. Source code is available at https://github.com/manishb89/interpretable_sentence_similarity
Subhadeep Maji, Manish Bansal, Kalyani Roy, Pawan Goyal 0002
IJCAI1
2020 Predicting Session Length for Product Search on E-commerce Platform
abstract
Estimation of session duration for an e-commerce search engine is important for various downstream applications, including user satisfaction prediction, personalization, and diversification of search results. It has been shown in previous studies that search session length has a strong correlation with user's explore vs specific purchase intent. Based on previous work [14], we hypothesize that early prediction of session length distribution can be used to control the degree of explore vs exploit (loosely related to diversification v/s personalization) for Search Engine Result Pages (SERPs) in the user's session to follow. In this work, we try to early predict the user's session length, which will enable the control on explore v/s exploit of the search results. Towards this end, based on previous work and strong empirical evidence, we hypothesize session lengths are Weibull distributed and propose its parameters being modeled by a Recurrent Neural Network over actions in user's search sessions. Through experimentation, we demonstrate that our method performs better as compared to strong baselines for the same.
Shashank Gupta 0001, Subhadeep Maji
SIGIR2
2019 Addressing Vocabulary Gap in E-commerce Search
abstract
E-commerce customers express their purchase intents in several ways, some of which may use a different vocabulary than that of the product catalog. For example, the intent for "women maternity gown" is often expressed with the query, "ladies pregnancy dress". Search engines typically suffer from poor performance on such queries because of low overlap between query terms and specifications of the desired products. Past work has referred to these queries as vocabulary gap queries. In our experiments, we show that our technique significantly outperforms strong baselines and also show its real-world effectiveness with an online A/B experiment.
Subhadeep Maji, Manish Bansal, Kalyani Roy, Mohit Kumar 0008, Pawan Goyal 0002
SIGIR1
2018 Automated Assistance in E-commerce: An Approach Based on Category-Sensitive Retrieval
Anirban Majumder, Abhay Pande, Kondalarao Vonteru, Abhishek Gangwar, Subhadeep Maji, Pankaj Bhatia, Pawan Goyal 0002
ECIR5