Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Samik Datta

dblp:93/9023 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 50% Language models and text generation · 25% Learning paradigms · 25%
Databases, data mining, and information retrieval
3 papers
Recommender systems · 48% Web and social media mining · 27% Data mining · 14%
Theoretical computer science
3 papers
Graph algorithms and graph theory · 63% Mathematical optimization · 21% Approximation and online algorithms · 10%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process › temporal point process
marked temporal point process
0.912025
ExPERT: Modeling Human Behavior Under External Stimuli Aware Personalized MTPP · AAAI 2025
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
temporal point process
0.912025
ExPERT: Modeling Human Behavior Under External Stimuli Aware Personalized MTPP · AAAI 2025
Natural language and speech › Language models and text generation
text representation
0.912025
ExPERT: Modeling Human Behavior Under External Stimuli Aware Personalized MTPP · AAAI 2025
Machine learning › Learning paradigms › semi-supervised learning
graph-based semi-supervised learning
0.412020
Understanding the Success of Graph-based Semi-Supervised Learning using Partially Labelled Stochastic Block Model · IJCAI 2020
Machine learning › Learning paradigms › semi-supervised learning › graph-based semi-supervised learning
label propagation
0.412020
Understanding the Success of Graph-based Semi-Supervised Learning using Partially Labelled Stochastic Block Model · IJCAI 2020
Graph algorithms and graph theory › graph clustering › community detection
stochastic block model
0.412020
Understanding the Success of Graph-based Semi-Supervised Learning using Partially Labelled Stochastic Block Model · IJCAI 2020
Data mining › predictive modeling
event prediction
0.312025
ExPERT: Modeling Human Behavior Under External Stimuli Aware Personalized MTPP · AAAI 2025
Web and social media mining › user behavior analysis
user behavior modeling
0.312025
ExPERT: Modeling Human Behavior Under External Stimuli Aware Personalized MTPP · AAAI 2025
Recommender systems › conversion rate prediction
delayed feedback
0.212023
Modelling Delayed Redemption with Importance Sampling and Pre-Redemption Engagement · KDD 2023
Machine learning and data management
importance sampling
0.212023
Modelling Delayed Redemption with Importance Sampling and Pre-Redemption Engagement · KDD 2023
Computational social science and digital humanities
social network analysis
0.112012
Capacitated team formation problem on social networks · KDD 2012
Computational social science and digital humanities › social computing
team formation
0.112012
Capacitated team formation problem on social networks · KDD 2012
Mathematical optimization
combinatorial optimization
0.112012
Capacitated team formation problem on social networks · KDD 2012
Web and social media mining › social network analysis
influence maximization
0.112010
Viral Marketing for Multiple Products · ICDM 2010
Web and social media mining
social network analysis
0.112010
Viral Marketing for Multiple Products · ICDM 2010
Approximation and online algorithms
approximation algorithms
0.012010
Viral Marketing for Multiple Products · ICDM 2010
Approximation and online algorithms › approximation algorithms › combinatorial approximation algorithms
greedy approximation
0.012010
Viral Marketing for Multiple Products · ICDM 2010

Methods — techniques the papers use, named apart from their topics

transformer hawkes process · 1.7language model embeddings · 0.9language model embedding · 0.9importance sampling · 0.7engagement modeling · 0.7social network analysis · 0.3combinatorial optimization · 0.3influence estimation heuristics · 0.2greedy hill climbing · 0.2
YearPublicationVenuePosition
2025 ExPERT: Modeling Human Behavior Under External Stimuli Aware Personalized MTPP
abstract
Marked Temporal Point Process (MTPP) -- the de-facto sequence model for continuous-time event sequences -- historically employed for modeling human-generated action sequences, lack awareness of external stimuli. In this study, we propose a novel framework developed over Transformer Hawkes Process (THP) to incorporate external stimuli in a domain-agnostic manner. Furthermore, we integrate personalization into our framework by employing language model-based representations of user and event descriptions, which is essential for modeling human-generated action sequences. Towards evaluating the efficacy, we put together a comprehensive benchmark comprising 5 datasets (2 novel additions, and 3 repurposed from existing open datasets) harvested from several domains, spanning education, e-commerce, online payment, and discussion forum. On average, we achieve 9.35% gain in type-prediction accuracy and 7.38% reduction in time-prediction RMSE across all datasets over SOTA MTPP baselines. We demonstrate the superior performance of our proposed model through extensive ablations and showcasing its ability to capture complex combinations of external stimuli in a synthetic set up.
Subhendu Khatuya, Ritvik Vij, Paramita Koley, Samik Datta, Niloy Ganguly
AAAI4
2024 Integrated Two Variant Deep Learners for Aspect-Based Sentiment Analysis: An Improved Meta-Heuristic-Based Model
abstract
Several traditional methods were tested on standard datasets to evaluate client emotions transmitted via internet portals. Customers, on the other hand, continue to have difficulty obtaining aspect-oriented viewpoints voiced by other customers, and the accuracy of the current model is insufficient. The suggested Aspect-Based Sentimental Analysis (ABSA) starts with pre-processing, which includes “stop word and punctuation removal, lower case conversion, and stemming.” Aspect extraction, which entails dividing the nouns and adjectives, as well as verbs and adverbs, is the following step. The weighted polarity features from the “Vader sentiment intensity analyzer, as well as the word2vector and Term Frequency-Inverse Document Frequency (TF-IDF)” are concatenated. OIDL stands for Optimized Integrated Deep Learning, which combines two types of deep learners. The first is the combination of concatenated features with “Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN),” while the second is the combination of concatenated features with RNN. The Improved Coyote Optimization Algorithm (ICOA) improves both deep learners, and the conclusion of sentiment analysis result is considered both models. Thus, the suggested model surpasses standard methodologies regarding precision and accuracy, according to the results of the experiments.
Samik Datta, Satyajit Chakrabarti
Cybern. Syst.1
2023 Modelling Delayed Redemption with Importance Sampling and Pre-Redemption Engagement
abstract
Rewards-based programs are popular within e-commerce online stores, with the goal of providing serendipitous incentives to delight customers. These rewards (or incentives) could be in the form of cashback, free-shipping or discount coupons on purchases within specific categories. The success of such programs relies on their ability to identify relevant rewards for customers, from a wide variety of incentives available on the online store. Estimating the likelihood of a customer redeeming an incentive is challenging due to 1) data sparsity: relatively rare occurrence of coupon redemptions as compared to issuances, and 2) delayed feedback: customers taking time to redeem, resulting in inaccurate model refresh, compounded by data drift due to new customers and coupons.
Samik Datta, Anshuman Mourya, Anirban Majumder, Vineet Chaoji
KDD1
2021 Reproducibility, Replicability and Beyond: Assessing Production Readiness of Aspect Based Sentiment Analysis in the Wild
Rajdeep Mukherjee, Shreyas Shetty, Subrata Chattopadhyay, Subhadeep Maji, Samik Datta, Pawan Goyal 0002
ECIR (2)5
2021 Graph-based semi-supervised learning through the lens of safety
abstract
Graph-based semi-supervised learning (G-SSL) algorithms have witnessed rapid development and widespread usage across a variety of applications in recent years. However, the theoretical characterisation of the efficacy of such algorithms has remained an under-explored area. We introduce a novel algorithm for G-SSL, CSX, whose objective function extends those of Label Propagation and Expander, two popular G-SSL algorithms. We provide data-dependent generalisation error bounds for all three aforementioned algorithms when they are applied to graphs drawn from a partially labelled extension of a versatile latent space graph generative model. The bounds we obtain enable us to characterise the predictive performance as measured by accuracy in terms of homophily and label quantity. Building on this we develop a key notion of GLM-safety which enables us to compare G-SSL algorithms on the basis of the range of graphs on which they obtain a guaranteed accuracy. We show that the proposed algorithm CSX has a better GLM-safety profile than Label Propagation and Expander while achieving comparable or better accuracy on synthetic as well as real-world benchmark networks.
Shreyas Sheshadri, Avirup Saha, Priyank Patel, Samik Datta, Niloy Ganguly
UAI4
2020 Understanding the Success of Graph-based Semi-Supervised Learning using Partially Labelled Stochastic Block Model
abstract
With the proliferation of learning scenarios with an abundance of instances, but limited amount of high-quality labels, semi-supervised learning algorithms came to prominence. Graph-based semi-supervised learning (G-SSL) algorithms, of which Label Propagation (LP) is a prominent example, are particularly well-suited for these problems. The premise of LP is the existence of homophily in the graph, but beyond that nothing is known about the efficacy of LP. In particular, there is no characterisation that connects the structural constraints, volume and quality of the labels to the accuracy of LP. In this work, we draw upon the notion of recovery from the literature on community detection, and provide guarantees on accuracy for partially-labelled graphs generated from the Partially-Labelled Stochastic Block Model (PLSBM). Extensive experiments performed on synthetic data verify the theoretical findings.
Avirup Saha, Shreyas Sheshadri, Samik Datta, Niloy Ganguly, Disha Makhija, Priyank Patel
IJCAI3
2017 SOPER: Discovering the Influence of Fashion and the Many Faces of User from Session Logs using Stick Breaking Process
abstract
Recommending lifestyle articles is of immediate interest to the e-commerce industry and is beginning to attract research attention. Often followed strategies, such as recommending popular items are inadequate for this vertical because of two reasons. Firstly, users have their own personal preference over items, referred to as personal styles, which lead to the long-tail phenomenon. Secondly, each user displays multiple personas, each persona has a preference over items which could be dictated by a particular occasion, e.g. dressing for a party would be different from dressing to go to office. Recommendation in this vertical is crucially dependent on discovering styles for each of the multiple personas. There is no literature which addresses this problem.
Lucky Dhakad, Mrinal Kanti Das, Chiranjib Bhattacharyya, Samik Datta, Mihir Kale, Vivek Mehta
CIKM4
2016 PEQ: An Explainable, Specification-based, Aspect-oriented Product Comparator for E-commerce
abstract
While purchasing a product, consumers often rely on specifications as well as online reviews of the product for decision-making. While comparing, one often has in mind a specific aspect or a set of aspects which are of interest to them. Previous work has used comparative sentences, where two entities are compared directly in a single sentence by the review author, towards the comparison task. In this paper, we extend the existing model by incorporating the feature specifications of the products, which are easily available, and learn the importance to be associated with each of them. To test the validity of these product ranking measures, we comprehensively test it on a digital camera dataset from Amazon.com and the results show good empirical outperformance over the state-of-the-art baselines.
Abhishek Sikchi, Pawan Goyal 0002, Samik Datta
CIKM3
2016 TweetGrep: Weakly Supervised Joint Retrieval and Sentiment Analysis of Topical Tweets
Satarupa Guha, Tanmoy Chakraborty 0002, Samik Datta, Vasudeva Varma
ICWSM3
2012 Capacitated team formation problem on social networks
abstract
In a team formation problem, one is required to find a group of users that can match the requirements of a collaborative task. Example of such collaborative tasks abound, ranging from software product development to various participatory sensing tasks in knowledge creation. Due to the nature of the task, team members are often required to work on a co-operative basis. Previous studies [1, 2] have indicated that co-operation becomes effective in presence of social connections. Therefore, effective team selection requires the team members to be socially close as well as a division of the task among team members so that no user is overloaded by the assignment. In this work, we investigate how such teams can be formed on a social network.
Anirban Majumder, Samik Datta, K. V. M. Naidu
KDD2
2010 Viral Marketing for Multiple Products
abstract
Viral Marketing, the idea of exploiting social interactions of users to propagate awareness for products, has gained considerable focus in recent years. One of the key issues in this area is to select the best seeds that maximize the influence propagated in the social network. In this paper, we define the seed selection problem (called t-Influence Maximization, or t-IM) for multiple products. Specifically, given the social network and t products along with their seed requirements, we want to select seeds for each product that maximize the overall influence. As the seeds are typically sent promotional messages, to avoid spamming users, we put a hard constraint on the number of products for which any single user can be selected as a seed. In this paper, we design two efficient techniques for the t-IM problem, called Greedy and FairGreedy. The Greedy algorithm uses simple greedy hill climbing, but still results in a 1/3-approximation to the optimum. Our second technique, FairGreedy, allocates seeds with not only high overall influence (close to Greedy in practice), but also ensures fairness across the influence of different products. We also design efficient heuristics for estimating the influence of the selected seeds, that are crucial for running the seed selection on large social network graphs. Finally, using extensive simulations on real-life social graphs, we show the effectiveness and scalability of our techniques compared to existing and naive strategies.
Samik Datta, Anirban Majumder, Nisheeth Shrivastava
ICDM1