Aditya Pal

dblp:70/7564 · DBLP profile ↗
← Back
29ranked-venue papers
17as first author
2since 2021 · last 2025
0000-0001-8079-0935ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 13 first-authorHuman-computer interaction and ubiquitous computing · 11 · 5 first-authorArtificial intelligence and machine learning · 9 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
12 papers
Recommender systems · 37% Information retrieval · 23% Data mining · 20%
Artificial intelligence
4 papers
Graph learning · 57% Probabilistic and Bayesian machine learning · 22% Information extraction and text analysis · 14%
Human-computer interaction and pervasive computing
5 papers
Collaborative and social computing · 100%

Topics — the 30 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
clustering
0.622020
PinnerSage: Multi-Modal User Embedding Framework for Recommendations at Pinterest · KDD 2020
Identifying topical authorities in microblogs · WSDM 2011
Collaborative and social computing
online communities
0.532014
Selecting an effective niche: an ecological view of the success of online communities · CHI 2014
Goals and perceived success of online enterprise communities: what is important to leaders & members? · CHI 2014
Socializing volunteers in an online community: a field experiment · CSCW 2012
Information retrieval › search engines
expert finding
0.532015
Discovering Experts across Multiple Domains · SIGIR 2015
Exploring Question Selection Bias to Identify Experts and Potential Experts in Community Question Answering · ACM Trans. Inf. Syst. 2012
Identifying topical authorities in microblogs · WSDM 2011
Machine learning › Graph learning › graph neural network
graph convolutional network
0.412020
MultiSage: Empowering GCN with Contextualized Multi-Embeddings on Web-Scale Multipartite Networks · KDD 2020
Machine learning › Graph learning
graph neural network
0.412020
MultiSage: Empowering GCN with Contextualized Multi-Embeddings on Web-Scale Multipartite Networks · KDD 2020
Recommender systems
graph-based recommendation
0.412020
MultiSage: Empowering GCN with Contextualized Multi-Embeddings on Web-Scale Multipartite Networks · KDD 2020
Data mining › clustering
hierarchical clustering
0.412020
PinnerSage: Multi-Modal User Embedding Framework for Recommendations at Pinterest · KDD 2020
Recommender systems › representation learning for recommendation
user embedding
0.412020
PinnerSage: Multi-Modal User Embedding Framework for Recommendations at Pinterest · KDD 2020
Recommender systems
sequential recommendation
0.412019
Hierarchical Temporal Convolutional Networks for Dynamic Recommender Systems · WWW 2019
Information retrieval
question answering
0.322015
Metrics and Algorithms for Routing Questions to User Communities · ACM Trans. Inf. Syst. 2015
Question temporality: identification and uses · CSCW 2012
Recommender systems
cold-start recommendation
0.212016
Discovery of Topical Authorities in Instagram · WWW 2016
Web and social media mining
social network analysis
0.212016
Discovery of Topical Authorities in Instagram · WWW 2016
Recommender systems
user recommendation
0.212016
Discovery of Topical Authorities in Instagram · WWW 2016
Natural language and speech › Information extraction and text analysis
emotion recognition
0.212015
Detecting Emotions in Social Media: A Constrained Optimization Approach · IJCAI 2015
Recommender systems › social recommendation
community recommendation
0.212015
Metrics and Algorithms for Routing Questions to User Communities · ACM Trans. Inf. Syst. 2015
Information retrieval › question answering
question routing
0.212015
Metrics and Algorithms for Routing Questions to User Communities · ACM Trans. Inf. Syst. 2015
Collaborative and social computing › social media
enterprise social media
0.212015
Inferring Employee Engagement from Social Media · CHI 2015
Mathematical optimization
constrained optimization
0.212015
Detecting Emotions in Social Media: A Constrained Optimization Approach · IJCAI 2015
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.212013
Discovering Hierarchical Structure for Sources and Entities · AAAI 2013
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
indian buffet process
0.212013
Discovering Hierarchical Structure for Sources and Entities · AAAI 2013
Information retrieval › question answering
community question answering
0.112012
Exploring Question Selection Bias to Identify Experts and Potential Experts in Community Question Answering · ACM Trans. Inf. Syst. 2012
Data integration and cleaning
data fusion
0.112012
Information integration over time in unreliable and uncertain environments · WWW 2012
Data integration and cleaning › truth discovery
source reliability estimation
0.112012
Information integration over time in unreliable and uncertain environments · WWW 2012
Data integration and cleaning
truth discovery
0.112012
Information integration over time in unreliable and uncertain environments · WWW 2012
Web and social media mining › social media analysis
microblog analysis
0.112011
Identifying topical authorities in microblogs · WSDM 2011
Data mining › clustering
probabilistic clustering
0.112011
Identifying topical authorities in microblogs · WSDM 2011
Machine learning › Deep learning architectures and training › convolutional neural network › convolutional neural network architecture
temporal convolutional network
0.112019
Hierarchical Temporal Convolutional Networks for Dynamic Recommender Systems · WWW 2019
Web and social media mining › online social networks
instagram
0.112016
Discovery of Topical Authorities in Instagram · WWW 2016
Information retrieval › ranking › search ranking
proximity ranking
0.112015
Metrics and Algorithms for Routing Questions to User Communities · ACM Trans. Inf. Syst. 2015
Information retrieval
ranking
0.112015
Metrics and Algorithms for Routing Questions to User Communities · ACM Trans. Inf. Syst. 2015

Methods — techniques the papers use, named apart from their topics

random walk · 0.9multi-GPU training · 0.9attention mechanism · 0.9recurrent neural network · 0.8queue-based mini-batch generation · 0.8hierarchical temporal convolutional networks · 0.8data caching · 0.8predictive modeling · 0.7dictionary-based linguistic analysis · 0.7constrained optimization · 0.4ward hierarchical clustering · 0.4medoids · 0.4label propagation · 0.2survey · 0.2quantitative analysis · 0.2qualitative analysis · 0.2interviews · 0.2indian buffet process · 0.2
YearPublicationVenuePosition
2025 SigNet-TAM: Indian Sign Language Recognition with ResNet50-BiLSTM and Temporal Attention Mechanism
Abraham Vellaiparambil Jose, Aditya Pal, Vaibhav Prajapati, Rangachary Kommanduri, Mrinmoy Ghorai, Himangshu Sarma
CGI (2)2
2025 Alphanumeric Fingerspelling: A New Large-Scale Dataset and Comparative Analysis of Methods for Indian Sign Language
Aditya Pal, Abraham Vellaiparambil Jose, Vaibhav Prajapati, K. Rangachary, Mrinmoy Ghorai, Himangshu Sarma
CGI (3)1
2020 PinnerSage: Multi-Modal User Embedding Framework for Recommendations at Pinterest
abstract
Latent user representations are widely adopted in the tech industry for powering personalized recommender systems. Most prior work infers a single high dimensional embedding to represent a user, which is a good starting point but falls short in delivering a full understanding of the user's interests. In this work, we introduce PinnerSage, an end-to-end recommender system that represents each user via multi-modal embeddings and leverages this rich representation of users to provides high quality personalized recommendations. PinnerSage achieves this by clustering users' actions into conceptually coherent clusters with the help of a hierarchical clustering method (Ward) and summarizes the clusters via representative pins (Medoids) for efficiency and interpretability. PinnerSage is deployed in production at Pinterest and we outline the several design decisions that makes it run seamlessly at a very large scale. We conduct several offline and online A/B experiments to show that our method significantly outperforms single embedding methods.
Aditya Pal, Pong Eksombatchai, Charles Rosenberg 0001, Jure Leskovec
KDD1
2020 MultiSage: Empowering GCN with Contextualized Multi-Embeddings on Web-Scale Multipartite Networks
abstract
Graph convolutional networks (GCNs) are a powerful class of graph neural networks. Trained in a semi-supervised end-to-end fashion, GCNs can learn to integrate node features and graph structures to generate high-quality embeddings that can be used for various downstream tasks like search and recommendation. However, existing GCNs mostly work on homogeneous graphs and consider a single embedding for each node, which do not sufficiently model the multi-facet nature and complex interaction of nodes in real-world networks. Here, we present a contextualized GCN engine by modeling the multipartite networks of target nodes and their intermediatecontext nodes that specify the contexts of their interactions. Towards the neighborhood aggregation process, we devise a contextual masking operation at the feature level and a contextual attention mechanism at the node level to achieve interaction contextualization by treating neighboring target nodes based on intermediate context nodes. Consequently, we compute multiple embeddings for target nodes that capture their diverse facets and different interactions during graph convolution, which is useful for fine-grained downstream applications. To enable efficient web-scale training, we build a parallel random walk engine to pre-sample contextualized neighbors, and a Hadoop2-based data provider pipeline to pre-join training data, dynamically reduce multi-GPU training time, and avoid high memory cost. Extensive experiments on the bipartite Pinterest graph and tripartite OAG graph corroborate the advantage of the proposed system.
Carl Yang 0001, Aditya Pal, Andrew Zhai, Nikil Pancha, Jiawei Han 0001, Charles Rosenberg 0001, Jure Leskovec
KDD2
2019 Hierarchical Temporal Convolutional Networks for Dynamic Recommender Systems
abstract
Recommender systems that can learn from cross-session data to dynamically predict the next item a user will choose are crucial for online platforms. However, existing approaches often use out-of-the-box sequence models which are limited by speed and memory consumption, are often infeasible for production environments, and usually do not incorporate cross-session information, which is crucial for effective recommendations. Here we propose Hierarchical Temporal Convolutional Networks (HierTCN), a hierarchical deep learning architecture that makes dynamic recommendations based on users' sequential multi-session interactions with items. HierTCN is designed for web-scale systems with billions of items and hundreds of millions of users. It consists of two levels of models: The high-level model uses Recurrent Neural Networks (RNN) to aggregate users' evolving long-term interests across different sessions, while the low-level model is implemented with Temporal Convolutional Networks (TCN), utilizing both the long-term interests and the short-term interactions within sessions to predict the next interaction. We conduct extensive experiments on a public XING dataset and a large-scale Pinterest dataset that contains 6 million users with 1.6 billion interactions. We show that HierTCN is 2.5x faster than RNN-based models and uses 90% less data memory compared to TCN-based models. We further develop an effective data caching scheme and a queue-based mini-batch generator, enabling our model to be trained within 24 hours on a single GPU. Our model consistently outperforms state-of-the-art dynamic recommendation methods, with up to 18% improvement in recall and 10% in mean reciprocal rank.
Jiaxuan You, Yichen Wang 0001, Aditya Pal, Pong Eksombatchai, Charles Rosenberg 0001, Jure Leskovec
WWW3
2018 Label Propagation with Neural Networks
abstract
Label Propagation (LP) is a popular transductive learning method for very large datasets, in part due to its simplicity and ability to parallelize. However, it has limited ability to handle node features, and its accuracy can be sensitive to the number of iterations. We propose an algorithm called LPNN that solves these problems by a loose-coupling of LP with a feature-based classifier. We experimentally establish the effectiveness of LPNN.
Aditya Pal, Deepayan Chakrabarti
CIKM1
2016 Predicting Attitude and Actions of Twitter Users
abstract
In this paper, we present computational models to predict Twitter users' attitude towards a specific brand through their personal and social characteristics. We also predict their likelihood of taking different actions based on their attitudes. In order to operationalize our research on users' attitude and actions, we collected ground-truth data through surveys of Twitter users. We have conducted experiments using two real world datasets to validate the effectiveness of our attitude and action prediction framework. Finally, we show how our models can be integrated with a visual analytics system for customer intervention.
Jalal Mahmud, Geli Fei, Anbang Xu, Aditya Pal, Michelle X. Zhou
IUI4
2016 Discovery of Topical Authorities in Instagram
abstract
Instagram has more than 400 million monthly active accounts who share more than 80 million pictures and videos daily. This large volume of user-generated content is the application's notable strength, but also makes the problem of finding the authoritative users for a given topic challenging. Discovering topical authorities can be useful for providing relevant recommendations to the users. In addition, it can aid in building a catalog of topics and top topical authorities in order to engage new users, and hence provide a solution to the cold-start problem. In this paper, we present a novel approach that we call the Authority Learning Framework (ALF) to find topical authorities in Instagram. ALF is based on the self-described interests of the follower base of popular accounts. We infer regular users' interests from their self-reported biographies that are publicly available and use Wikipedia pages to ground these interests as fine-grained, disambiguated concepts. We propose a generalized label propagation algorithm to propagate the interests over the follower graph to the popular accounts. We show that even if biography-based interests are sparse at an individual user level they provide strong signals to infer the topical authorities and let us obtain a high precision authority list per topic. Our experiments demonstrate that ALF performs significantly better at user recommendation task compared to fine-tuned and competitive methods, via controlled experiments, in-the-wild tests, and over an expert-curated list of topical authorities.
Aditya Pal, Amac Herdagdelen, Sourav Chatterji, Sumit Taank, Deepayan Chakrabarti
WWW1
2015 Inferring Employee Engagement from Social Media
abstract
Employees increasingly are expressing ideas and feelings through enterprise social media. Recent work in CHI and CSCW has applied linguistic analysis towards understanding employee experiences. In this paper, we apply dictionary based linguistic analysis to measure 'Employee Engagement'. Employee engagement is a measure of employee willingness to apply discretionary effort towards organizational goals, and plays an important role in organizational outcomes such as financial or operational results. Organizations typically use surveys to measure engagement. This paper describes an approach to model employee engagement based on word choice in social media. This method can potentially complement surveys, thus providing more real-time insights into engagement and allowing organizations to address engagement issues faster. Our results predicting engagement scores on a survey by combining demographics with social media text demonstrate that social media text has significant predictive power compared to demographic data alone. We also find that engagement may be a state than a stable trait since social media posts closer to the administration of the survey had the most predictive power. We further identify the minimum number of social media posts required per employee for the best prediction.
N. Sadat Shami, Michael J. Muller, Aditya Pal, Mikhil Masli, Werner Geyer
CHI3
2015 Detecting Emotions in Social Media: A Constrained Optimization Approach
Yichen Wang 0001, Aditya Pal
IJCAI2
2015 Discovering Experts across Multiple Domains
abstract
Researchers have focused on finding experts in individual domains, such as emails, forums, question answering, blogs, and microblogs. In this paper, we propose an algorithm for finding experts across these different domains. To do this, we propose an expertise framework that aims at extracting key expertise features and building an unified scoring model based on SVM ranking algorithm. We evaluate our model on a real World dataset and show that it is significantly better than the prior state-of-art.
Aditya Pal
SIGIR1
2015 Metrics and Algorithms for Routing Questions to User Communities
abstract
An online community consists of a group of users who share a common interest, background, or experience, and their collective goal is to contribute toward the welfare of the community members. Several websites allow their users to create and manage niche communities, such as Yahoo! Groups, Facebook Groups, Google+ Circles, and WebMD Forums. These community services also exist within enterprises, such as IBM Connections. Question answering within these communities enables their members to exchange knowledge and information with other community members. However, the onus of finding the right community for question asking lies with an individual user. The overwhelming number of communities necessitates the need for a good question routing strategy so that new questions get routed to an appropriately focused community and thus get resolved in a reasonable time frame. In this article, we consider the novel problem of routing a question to the right community and propose a framework for selecting and ranking the relevant communities for a question. We propose several novel features for modeling the three main entities of the system: questions, users, and communities. We propose features such as language attributes, inclination to respond, user familiarity, and difficulty of a question; based on these features, we propose similarity metrics between the routed question and the system entities. We introduce aCutoff-Aggregation(CA) algorithm that aggregates the entity similarity within a community to compute that community's relevance. We introduce twok-nearest-neighbor (knn) algorithms that are a natural instantiation of theCAalgorithm, which are computationally efficient and evaluate several ranking algorithms over the aggregate similarity scores computed by the twoknnalgorithms. We propose clustering techniques to speed up our recommendation framework and show how pipelining can improve the model performance. We demonstrate the effectiveness of our framework on two large real-world datasets.
Aditya Pal
ACM Trans. Inf. Syst.1
2014 Goals and perceived success of online enterprise communities: what is important to leaders & members?
abstract
Online communities are successful only if they achieve their goals, but there has been little direct study of goals. We analyze novel data characterizing the goals of enterprise online communities, assessing the importance of goals for leaders, how goals influence member perceptions of community value, and how goals relate to success measures proposed in the literature. We find that most communities have multiple goals and common goals are learning, reuse of resources, collaboration, networking, influencing change, and innovation. Leaders and members agree that all of these goals are important, but their perceptions of success on goals do not align with each other, or with commonly used behavioral success measures. We conclude that simple behavioral measures and leader perceptions are not good success metrics, and propose alternatives based on specific goals members and leaders judge most important.
Tara Matthews, Jilin Chen, Steve Whittaker 0001, Aditya Pal, Haiyi Zhu, Hernan Badenes, Barton A. Smith
CHI4
2014 Selecting an effective niche: an ecological view of the success of online communities
abstract
Online communities serve various important functions, but many fail to thrive. Research on community success has traditionally focused on internal factors. In contrast, we take an ecological view to understand how the success of a community is influenced by other communities. We measured a community's relationship with other communities - its "niche" - through four dimensions: topic overlap, shared members, content linking, and shared offline organizational affiliation. We used a mixed-method approach, combining the quantitative analysis of 9495 online enterprise communities and interviews with community members. Our results show that too little or too much overlap in topic with other communities causes a community's activity to suffer. We also show that this main result is moderated in predictable ways by whether the community shares members with, links to content in, or shares an organizational affiliation with other communities. These findings provide new insight on community success, guiding online community designers on how to effectively position their community in relation to others.
Haiyi Zhu, Jilin Chen, Tara Matthews, Aditya Pal, Hernan Badenes, Robert E. Kraut
CHI4
2014 System U: automatically deriving personality traits from social media for people recommendation
abstract
This paper presents a system, System U, which automatically derives people's personality traits from social media and recommends people for different tasks. The system leverages linguistic signals appearing in a person's social media activities to compute the personality portraits including Big Five personality, fundamental needs and basic human values. This system and technology can be used in a wide variety of personalized applications, such as recommending people to answer questions.
Hernan Badenes, Mateo N. Bengualid, Jilin Chen, Liang Gou, Eben M. Haber, Jalal Mahmud, Jeffrey Nichols 0001, Aditya Pal, Jerald Schoudt, Barton A. Smith, Ying Xuan, Huahai Yang, Michelle X. Zhou
RecSys8
2013 Discovering Hierarchical Structure for Sources and Entities
abstract
In this paper, we consider the problem of jointly learning hierarchies over a set of sources and entities based on their containment relationship. We model the concept of hierarchy using a set of latent binary features and propose a generative model that assigns those latent features to sources and entities in order to maximize the probability of the observed containment. To avoid fixing the number of features beforehand, we consider a non-parametric approach based on the Indian Buffet Process. The hierarchies produced by our algorithm can be used for completing missing associations and discovering structural bindings in the data. Using simulated and real datasets we provide empirical evidence of the effectiveness of the proposed approach in comparison to the existing hierarchy agnostic approaches.
Aditya Pal, Nilesh N. Dalvi, Kedar Bellare
AAAI1
2013 Routing questions for collaborative answering in community question answering
abstract
Community Question Answering (CQA) service enables its users to exchange knowledge in the form of questions and answers. By allowing the users to contribute knowledge, CQA not only satisfies the question askers but also provides valuable references to other users with similar queries. Due to a large volume of questions, not all questions get fully answered. As a result, it can be useful to route a question to a potential answerer. In this paper, we present a question routing scheme which takes into account the answering, commenting and voting propensities of the users. Unlike prior work which focuses on routing a question to the most desirable expert, we focus on routing it to a group of users - who would be willing to collaborate and provide useful answers to that question. Through empirical evidence, we show that more answers and comments are desirable for improving the lasting value of a question-answer thread. As a result, our focus is on routing a question to a team of compatible users. We propose a recommendation model that takes into account the compatibility, topical expertise and availability of the users. Our experiments over a large real-world dataset shows the effectiveness of our approach over several baseline models.
Shuo Chang, Aditya Pal
ASONAM2
2013 Question routing to user communities
abstract
An online community consists of a group of users who share a common interest, background, or experience and their collective goal is to contribute towards the welfare of the community members. Question answering is an important feature that enables community members to exchange knowledge within the community boundary. The overwhelming number of communities necessitates the need for a good question routing strategy so that new questions gets routed to the appropriately focused community and thus get resolved. In this paper, we consider the novel problem of routing questions to the right community and propose a framework to select the right set of communities for a question. We begin by using several prior proposed features for users and add some additional features, namely language attributes and inclination to respond, for community modeling. Then we introduce two k nearest neighbor based aggregation algorithms for computing community scores. We show how these scores can be combined to recommend communities and test the effectiveness of the recommendations over a large real world dataset.
Aditya Pal, Fei Wang 0001, Michelle X. Zhou, Jeffrey Nichols 0001, Barton A. Smith
CIKM1
2012 Socializing volunteers in an online community: a field experiment
abstract
Although many off-line organizations give their employees training, mentorship, a cohort and other socialization experiences that improve their retention and productivity, online production communities rarely do this. This paper describes the planning, execution and evaluation of a socialization regime for an online technical support community. In a two-phase project, we first automatically identified from participants' early behavior, those with high potential to become core members. We then designed, delivered and experimentally evaluated socialization experiences intended to build commitment and competence among these potential core members. We were able to identify potential core members with high accuracy from only two weeks of behavior. A year later, those classified as potential core members participated in the community ten times more actively than those not identified. In an evaluation experiment, some potential core members were randomly assigned to receive socialization experiences, while others were not. A year later, those who had participated in the socialization regime contributed more answers in the community compared to those in the control condition. The socialization experiences, however, undercut their sense of connection to the community and the quality of their contributions. We discuss what was effective and what could be improved in designing socialization experiences for online groups.
Rosta Farzan, Robert E. Kraut, Aditya Pal, Joseph A. Konstan
CSCW3
2012 Question temporality: identification and uses
abstract
In this paper, we introduce the concept of question temporality as a measure of the usefulness of the answers provided on the questions asked in the Question Answering sites (QA). We define question temporality based on when the answers provided on the questions would expire. We use classification methods to show that the question temporality can be assessed automatically. Our regression analysis highlights features that predict temporality of the questions. Our research can be instructive for interface designers to design temporality-aware interfaces and influence selection of questions and answers for display.
Aditya Pal, James Margatan, Joseph A. Konstan
CSCW1
2012 Evolution of Experts in Question Answering Communities
Aditya Pal, Shuo Chang, Joseph A. Konstan
ICWSM1
2012 Tracking Spatio-Temporal Diffusion in Climate Data
abstract
A forest canopy forms a critical platform for complex interactions between the vegetation and the atmosphere boundary layer and is considered as a crucial piece for environmental scientists in their understanding of the ecosystem and its response to the climate change. Microfronts represent a class of these interactions characterized by a moving mass of air that introduce fluctuations in ambient temperature and humidity on small spatial and temporal scales. In this paper, we present a joint spatio-temporal hidden markov model that simultaneously incorporates neighborhood dependencies in space and time. We show that our approach can trace the diffusion of microfronts more effectively than several baseline methods over a sensor data from Brazilian rainforest and a synthetically generated dataset.
Jaya Kawale, Aditya Pal, Rob Fatland
SDM2
2012 Information integration over time in unreliable and uncertain environments
abstract
Often an interesting true value such as a stock price, sports score, or current temperature is only available via the observations of noisy and potentially conflicting sources. Several techniques have been proposed to reconcile these conflicts by computing a weighted consensus based on source reliabilities, but these techniques focus on static values. When the real-world entity evolves over time, the noisy sources can delay, or even miss, reporting some of the real-world updates. This temporal aspect introduces two key challenges for consensus-based approaches: (i) due to delays, the mapping between a source's noisy observation and the real-world update it observes is unknown, and (ii) missed updates may translate to missing values for the consensus problem, even if the mapping is known. To overcome these challenges, we propose a formal approach that models the history of updates of the real-world entity as a hidden semi-Markovian process (HSMM). The noisy sources are modeled as observations of the hidden state, but the mapping between a hidden state (i.e. real-world update) and the observation (i.e. source value) is unknown. We propose algorithms based on Gibbs Sampling and EM to jointly infer both the history of real-world updates as well as the unknown mapping between them and the source values. We demonstrate using experiments on real-world datasets how our history-based techniques improve upon history-agnostic consensus-based approaches.
Aditya Pal, Vibhor Rastogi, Ashwin Machanavajjhala, Philip Bohannon
WWW1
2012 Exploring Question Selection Bias to Identify Experts and Potential Experts in Community Question Answering
abstract
Community Question Answering (CQA) services enable their users to exchange knowledge in the form of questions and answers. These communities thrive as a result of a small number of highly active users, typically calledexperts, who provide a large number of high-quality useful answers. Expert identification techniques enable community managers to take measures to retain the experts in the community. There is further value in identifying the experts during the first few weeks of their participation as it would allow measures to nurture and retain them. In this article we address two problems: (a) How to identify current experts in CQA? and (b) How to identify users who have potential of becoming experts in future (potential experts)? In particular, we propose a probabilistic model that captures the selection preferences of users based on the questions they choose for answering. The probabilistic model allows us to run machine learning methods for identifying experts and potential experts. Our results over several popular CQA datasets indicate that experts differ considerably from ordinary users in their selection preferences; enabling us to predict experts with higher accuracy over several baseline models. We show that selection preferences can be combined with baseline measures to improve the predictive performance even further.
Aditya Pal, F. Maxwell Harper, Joseph A. Konstan
ACM Trans. Inf. Syst.1
2011 What's in a @name? How Name Value Biases Judgment of Microblog Authors
Aditya Pal, Scott Counts
ICWSM1
2011 Connecting Mutually Influencing Bloggers
Aditya Pal, Jaya Kawale
ICWSM1
2011 Early Detection of Potential Experts in Question Answering Communities
Aditya Pal, Rosta Farzan, Joseph A. Konstan, Robert E. Kraut
UMAP1
2011 Identifying topical authorities in microblogs
abstract
Content in microblogging systems such as Twitter is produced by tens to hundreds of millions of users. This diversity is a notable strength, but also presents the challenge of finding the most interesting and authoritative authors for any given topic. To address this, we first propose a set of features for characterizing social media authors, including both nodal and topical metrics. We then show how probabilistic clustering over this feature space, followed by a within-cluster ranking procedure, can yield a final list of top authors for a given topic. We present results across several topics, along with results from a user study confirming that our method finds authors who are significantly more interesting and authoritative than those resulting from several baseline conditions. Additionally our algorithm is computationally feasible in near real-time scenarios making it an attractive alternative for capturing the rapidly changing dynamics of microblogs.
Aditya Pal, Scott Counts
WSDM1
2010 Expert identification in community question answering: exploring question selection bias
abstract
Community Question Answering (CQA) services enables users to ask and answer questions. In these communities, there are typically a small number of experts amongst the large population of users. We study which questions a user select for answering and show that experts prefer answering questions where they have a higher chance of making a valuable contribution. We term this preferential selection as question selection bias and propose a mathematical model to estimate it. Our results show that using Gaussian classification models we can effectively distinguish experts from ordinary users over their selection biases. In order to estimate these biases, only a small amount of data per user is required, which makes an early identification of expertise a possibility. Further, our study of bias evolution reveals that they do not show significant changes over time indicating that they emanates from the intrinsic characteristics of users.
Aditya Pal, Joseph A. Konstan
CIKM1