Mossaab Bagdouri

dblp:88/10915 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 7 first-authorArtificial intelligence and machine learning · 4 · 4 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 86% Data mining · 14%
Artificial intelligence
3 papers
Question answering and dialogue systems · 95% Trustworthy machine learning · 5%
Human-computer interaction and pervasive computing
2 papers
Collaborative and social computing · 91% Health and well-being technologies · 9%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 87% Smart cities and intelligent transportation · 13%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
community question answering
0.522017
Building Bridges across Social Platforms: Answering Twitter Questions with Yahoo! Answers · SIGIR 2017
Cross-Platform Question Routing for Better Question Answering · SIGIR 2015
Natural language and speech › Question answering and dialogue systems › community question answering
question routing
0.522017
Building Bridges across Social Platforms: Answering Twitter Questions with Yahoo! Answers · SIGIR 2017
Cross-Platform Question Routing for Better Question Answering · SIGIR 2015
Information retrieval
evaluation
0.532016
Pearson Rank: A Head-Weighted Gap-Sensitive Score-Based Correlation Coefficient · SIGIR 2016
Sequential testing in classifier evaluation yields biased estimates of effectiveness · SIGIR 2013
Cross-Platform Question Routing for Better Question Answering · SIGIR 2015
Information retrieval › evaluation
rank correlation
0.212016
Pearson Rank: A Head-Weighted Gap-Sensitive Score-Based Correlation Coefficient · SIGIR 2016
Data mining › predictive modeling › classification
classifier evaluation
0.212013
Sequential testing in classifier evaluation yields biased estimates of effectiveness · SIGIR 2013
Computational social science and digital humanities › social computing
crisis informatics
0.112012
"Beacons of hope" in decentralized coordination: learning from on-the-ground medical twitterers during the 2010 Haiti earthquake · CSCW 2012
Computational social science and digital humanities
social media analysis
0.112012
"Beacons of hope" in decentralized coordination: learning from on-the-ground medical twitterers during the 2010 Haiti earthquake · CSCW 2012
Collaborative and social computing › online communities
collective identity
0.112012
Blogs as a collective war diary · CSCW 2012
Collaborative and social computing
computer-supported cooperative work
0.112012
"Beacons of hope" in decentralized coordination: learning from on-the-ground medical twitterers during the 2010 Haiti earthquake · CSCW 2012
Collaborative and social computing › social media
social media analysis
0.112012
Blogs as a collective war diary · CSCW 2012
Information retrieval › ranking
learning to rank
0.112017
Building Bridges across Social Platforms: Answering Twitter Questions with Yahoo! Answers · SIGIR 2017
Information retrieval
ranking
0.112017
Building Bridges across Social Platforms: Answering Twitter Questions with Yahoo! Answers · SIGIR 2017
Information retrieval › evaluation › test collection
test collection reusability
0.112016
Pearson Rank: A Head-Weighted Gap-Sensitive Score-Based Correlation Coefficient · SIGIR 2016
Machine learning › Trustworthy machine learning › fairness and bias
evaluation bias
0.012013
Sequential testing in classifier evaluation yields biased estimates of effectiveness · SIGIR 2013
Information retrieval › evaluation › effectiveness metrics
effectiveness measure estimation
0.012013
Sequential testing in classifier evaluation yields biased estimates of effectiveness · SIGIR 2013
Smart cities and intelligent transportation › disaster management
disaster response
0.012012
Blogs as a collective war diary · CSCW 2012

Methods — techniques the papers use, named apart from their topics

question transformation · 0.6learning to rank · 0.6text classification · 0.3sequential testing · 0.3topic modeling · 0.3social media analysis · 0.3pronoun analysis · 0.3inductive analysis · 0.3pearson correlation · 0.2kendall's tau · 0.2
YearPublicationVenuePosition
2017 Building Bridges across Social Platforms: Answering Twitter Questions with Yahoo! Answers
abstract
This paper investigates techniques for answering microblog questions by searching in a large community question answering website. Some question transformations are considered, some proprieties of the answering platform are examined, how to select among the various available configurations in a learning-to-rank framework is studied.
Mossaab Bagdouri, Douglas W. Oard
SIGIR1
2016 Journalists and Twitter: A Multidimensional Quantitative Description of Usage Patterns
Mossaab Bagdouri
ICWSM1
2016 Pearson Rank: A Head-Weighted Gap-Sensitive Score-Based Correlation Coefficient
abstract
One way of evaluating the reusability of a test collection is to determine whether removing the unique contributions of some system would alter the preference order between that system and others. Rank correlation measures such as Kendall's tau are often used for this purpose. Rank correlation measures are appropriate for ordinal measures in which only preference order is important, but many evaluation measures produce system scores in which both the preference order and the magnitude of the score difference are important. Such measures are referred to as interval. Pearson's rho offers one way in which correlation can be computed over results from an interval measure such that smaller errors in the gap size are preferred. When seeking to improve over existing systems, we care the most about comparisons among the best systems. For that purpose we prefer head-weighed measures such as tau_AP, which is designed for ordinal data. No present head weighted measure fully leverages the information present in interval effectiveness measures. This paper introduces such a measure, referred to as Pearson Rank.
Ning Gao 0006, Mossaab Bagdouri, Douglas W. Oard
SIGIR2
2015 Profession-Based Person Search in Microblogs: Using Seed Sets to Find Journalists
abstract
We introduce the problem of searching for professionals in microblogging platforms. We describe a study of how a group of professional journalists with some common characteristics (e.g., works in a specific language, belongs to certain region, or specializes in a particular media) can be found. Starting from seed sets of different sizes, social network features and profile content features are used to find additional journalists. The results show that combining the social network features of the reciprocated mentions and a bidirectional friend/follower graph provides a signal stronger than either of them taken independently, that both social network and profile content features are useful, and that profile content features are able to find larger numbers of less prominent journalists. We apply our methods to find the Twitter accounts of British and Arab journalists.
Mossaab Bagdouri, Douglas W. Oard
CIKM1
2015 On Predicting Deletions of Microblog Posts
abstract
Among the many classification tasks on Twitter content, predicting whether a tweet will be deleted has to date received relatively little attention. Deletions occur for a variety of reasons, which can make the classification task challenging. Moreover, deletion prediction might serve different goals, the characteristics of which should be reflected in the evaluation design. This paper addresses the problem of deletion prediction by analyzing the distribution of deleted tweets, presenting a new evaluation framework, exploring tweet-based and user-based features, and reporting prediction scores.
Mossaab Bagdouri, Douglas W. Oard
CIKM1
2015 Cross-Platform Question Routing for Better Question Answering
abstract
The last two decades have seen an increasing interest in the task of question answering (QA). Earlier approaches focused on automated retrieval and extraction models. Recent developments have more focus on community driven QA. This work addresses this task through cross-platform question routing. We study question types as well as the answers that can be gathered from different platforms. After developing new evaluation measures, we optimize for various constraints of the user needs. We consider models that work for the general public, before adapting them to some special demographics (Arab journalists).
Mossaab Bagdouri
SIGIR1
2014 CLIR for Informal Content in Arabic Forum Posts
abstract
The field of Cross-Language Information Retrieval (CLIR) addresses the problem of finding documents in some language that are relevant to a question posed in a different language. Retrieving answers to questions written using formal vocabulary from collections of informal documents, as with many types of social media, is a largely unexplored subfield of CLIR. Because formal and informal content are often intermingled, CLIR systems that excel at finding formal content may tend to select formal over informal content. To measure this effect, a test collection annotated for both relevance and informality is needed. This paper describes the development of a small test collection for this task, with questions posed in formal English and the documents consisting of intermixed formal and informal Arabic. Experiments with this collection show that dialect classification can help to recognize informal content, thus improving precision. At the same time, the results indicate that neither dialect-tuned morphological analysis nor a lightweight CLIR approach that minimizes propagation of translation errors yet yield a reliable improvement in recall for informal content when compared to a straightforward document translation architecture.
Mossaab Bagdouri, Douglas W. Oard, Vittorio Castelli
CIKM1
2013 Towards minimizing the annotation cost of certified text classification
abstract
The common practice of testing a sequence of text classifiers learned on a growing training set, and stopping when a target value of estimated effectiveness is first met, introduces a sequential testing bias. In settings where the effectiveness of a text classifier must be certified (perhaps to a court of law), this bias may be unacceptable. The choice of when to stop training is made even more complex when, as is common, the annotation of training and test data must be paid for from a common budget: each new labeled training example is a lost test example. Drawing on ideas from statistical power analysis, we present a framework for joint minimization of training and test annotation that maintains the statistical validity of effectiveness estimates, and yields a natural definition of an optimal allocation of annotations to training and test data. We identify the development of allocation policies that can approximate this optimum as a central question for research. We then develop simulation-based power analysis methods for van Rijsbergen's F-measure, and incorporate them in four baseline allocation policies which we study empirically. In support of our studies, we develop a new analytic approximation of confidence intervals for the F-measure that is of independent interest.
Mossaab Bagdouri, William Webber, David D. Lewis, Douglas W. Oard
CIKM1
2013 Sequential testing in classifier evaluation yields biased estimates of effectiveness
abstract
It is common to develop and validate classifiers through a process of repeated testing, with nested training and/or test sets of increasing size. We demonstrate in this paper that such repeated testing leads to biased estimates of classifier effectiveness. Experiments on a range of text classification tasks under three sequential testing frameworks show all three lead to optimistic estimates of effectiveness. We calculate empirical adjustments to unbias estimates on our data set, and identify directions for research that could lead to general techniques for avoiding bias while reducing labeling costs.
William Webber, Mossaab Bagdouri, David D. Lewis, Douglas W. Oard
SIGIR2
2012 Blogs as a collective war diary
abstract
Disaster-related research in human-centered computing has typically focused on the shorter-term, emergency period of a disaster event, whereas effects of some crises are long-term, lasting years. Social media archived on the Internet provides researchers the opportunity to examine societal reactions to a disaster over time. In this paper we examine how blogs written during a protracted conflict might reflect a collective view of the event. The sheer amount of data originating from the Internet about a significant event poses a challenge to researchers; we employ topic modeling and pronoun analysis as methods to analyze such large-scale data. First, we discovered that blog war topics temporally tracked the actual, measurable violence in the society suggesting that blog content can be an indicator of the health or state of the affected population. We also found that people exhibited a collective identity when they blogged about war, as evidenced by a higher use of first-person plural pronouns compared to blogging on other topics. Blogging about daily life decreased as violence in the society increased; when violence waned, there was a resurgence of daily life topics, potentially illustrating how a society returns to normalcy.
Gloria Mark, Mossaab Bagdouri, Leysia Palen, James H. Martin, Ban Al-Ani, Kenneth M. Anderson
CSCW2
2012 "Beacons of hope" in decentralized coordination: learning from on-the-ground medical twitterers during the 2010 Haiti earthquake
abstract
We examine the public, social media communications of 110 emergency medical response teams and organizations in the immediate aftermath of the January 12, 2010 Haiti earthquake. We found the teams through an inductive analysis of Twitter communications acquired over the three-week emergency period from 89,114 Twitterers. We then analyzed the teams' Twitter streams, as well as all digital media they generated and pointed to in their streams - blog posts, photographs, videos, status updates and field reports - to understand the medical coordination challenges they faced from pre-deployment readiness to on-the-ground action. Here we identify opportunities for improving coordination in a decentralized and distributed environment where staffing, disease trajectories, and other circumstances rapidly change. We extrapolate from these findings to theorize about how "beaconing" behavior is a sign of latent potential for coordination upon which mechanisms of coordination can capitalize.
Aleksandra Sarcevic, Leysia Palen, Joanne I. White, Kate Starbird, Mossaab Bagdouri, Kenneth M. Anderson
CSCW5