Amac Herdagdelen

dblp:02/4543 · also Amaç Herdagdelen · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Recommender systems · 52% Web and social media mining · 34% Information retrieval · 15%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems
cold-start recommendation
0.212016
Discovery of Topical Authorities in Instagram · WWW 2016
Web and social media mining
social network analysis
0.212016
Discovery of Topical Authorities in Instagram · WWW 2016
Recommender systems
user recommendation
0.212016
Discovery of Topical Authorities in Instagram · WWW 2016
Information retrieval
query reformulation
0.112010
Generalized syntactic and semantic models of query reformulation · SIGIR 2010
Web and social media mining › online social networks
instagram
0.112016
Discovery of Topical Authorities in Instagram · WWW 2016

Methods — techniques the papers use, named apart from their topics

wikipedia grounding · 0.2label propagation · 0.2probabilistic term rewrite · 0.1levenshtein distance · 0.1generative model · 0.1
YearPublicationVenuePosition
2023 The Geography of Facebook Groups in the United States
abstract
We present a de-identified and aggregated dataset based on geographical patterns of Facebook Groups usage and demonstrate its association with measures of social capital. The dataset is aggregated at United States county level. Established spatial measures of social capital are known to vary across US counties. Their availability and recency depends on running costly surveys. We examine to what extent a dataset based on usage patterns of Facebook Groups, which can be generated at regular intervals, could be used as a partial proxy by capturing local online associations. We identify four main latent factors that distinguish Facebook group engagement by county, obtained by exploratory factor analysis. The first captures small and private groups, dense with friendship connections. The second captures very local and small groups. The third captures non-local, large, public groups, with more age mixing. The fourth captures partially local groups of medium to large size. Only two of these factors, the first and third, correlate with offline community level social capital measures, while the second and fourth do not. Together and individually, the factors are predictive of offline social capital measures, even controlling for various demographic attributes of the counties. To our knowledge this is the first systematic test of the association between offline regional social capital and patterns of online community engagement in the same regions. By making the dataset available to the research community, we hope to contribute to the ongoing studies in social capital.
Amac Herdagdelen, Lada A. Adamic, Bogdan State
ICWSM1
2020 What Makes People Feel Close to Online Groups? The Roles of Group Attributes and Group Types
Robert E. Kraut, John M. Levine, Marisol Martinez-Escobar, Amac Herdagdelen
ICWSM4
2016 Discovery of Topical Authorities in Instagram
abstract
Instagram has more than 400 million monthly active accounts who share more than 80 million pictures and videos daily. This large volume of user-generated content is the application's notable strength, but also makes the problem of finding the authoritative users for a given topic challenging. Discovering topical authorities can be useful for providing relevant recommendations to the users. In addition, it can aid in building a catalog of topics and top topical authorities in order to engage new users, and hence provide a solution to the cold-start problem. In this paper, we present a novel approach that we call the Authority Learning Framework (ALF) to find topical authorities in Instagram. ALF is based on the self-described interests of the follower base of popular accounts. We infer regular users' interests from their self-reported biographies that are publicly available and use Wikipedia pages to ground these interests as fine-grained, disambiguated concepts. We propose a generalized label propagation algorithm to propagate the interests over the follower graph to the popular accounts. We show that even if biography-based interests are sparse at an individual user level they provide strong signals to infer the topical authorities and let us obtain a high precision authority list per topic. Our experiments demonstrate that ALF performs significantly better at user recommendation task compared to fine-tuned and competitive methods, via controlled experiments, in-the-wild tests, and over an expert-curated list of topical authorities.
Aditya Pal, Amac Herdagdelen, Sourav Chatterji, Sumit Taank, Deepayan Chakrabarti
WWW2
2012 Bootstrapping a Game with a Purpose for Commonsense Collection
abstract
Text mining has been very successful in extracting huge amounts of commonsense knowledge from data, but the extracted knowledge tends to be extremely noisy. Manual construction of knowledge repositories, on the other hand, tends to produce high-quality data in very small amounts. We propose an architecture to combine the best of both worlds: A game with a purpose that induces humans to clean up data automatically extracted by text mining. First, a text miner trained on a set of known commonsense facts harvests many more candidate facts from corpora. Then, a simple slot-machine-with-a-purpose game presents these candidate facts to the players for verification by playing. As a result, a new dataset of high precision commonsense knowledge is created. This combined architecture is able to produce significantly better commonsense facts than the state-of-the-art text miner alone. Furthermore, we report that bootstrapping (i.e., training the text miner on the output of the game) improves the subsequent performance of the text miner.
Amac Herdagdelen, Marco Baroni
ACM Trans. Intell. Syst. Technol.1
2011 Stereotypical gender actions can be extracted from web text
abstract
We extracted gender-specific actions from text corpora and Twitter, and compared them with stereotypical expectations of people. We used Open Mind Common Sense (OMCS), a common sense knowledge repository, to focus on actions that are pertinent to common sense and daily life of humans. We use the gender information of Twitter users and web-corpus-based pronoun/name gender heuristics to compute the gender bias of the actions. With high recall, we obtained a Spearman correlation of 0.47 between corpus-based predictions and a human gold standard, and an area under the ROC curve of 0.76 when predicting the polarity of the gold standard. We conclude that it is feasible to use natural text (and a Twitter-derived corpus in particular) in order to augment common sense repositories with the stereotypical gender expectations of actions. We also present a dataset of 441 common sense actions with human judges' ratings on whether the action is typically/slightly masculine/feminine (or neutral), and another larger dataset of 21,442 actions automatically rated by the methods we investigate in this study.
Amac Herdagdelen, Marco Baroni
J. Assoc. Inf. Sci. Technol.1
2010 Learning Dense Models of Query Similarity from User Click Logs
Fabio De Bona, Stefan Riezler, Keith B. Hall, Massimiliano Ciaramita, Amac Herdagdelen, Maria Holmqvist
HLT-NAACL5
2010 Generalized syntactic and semantic models of query reformulation
abstract
We present a novel approach to query reformulation which combines syntactic and semantic information by means of generalized Levenshtein distance algorithms where the substitution operation costs are based on probabilistic term rewrite functions. We investigate unsupervised, compact and efficient models, and provide empirical evidence of their effectiveness. We further explore a generative model of query reformulation and supervised combination methods providing improved performance at variable computational costs. Among other desirable properties, our similarity measures incorporate information-theoretic interpretations of taxonomic relations such as specification and generalization.
Amac Herdagdelen, Massimiliano Ciaramita, Daniel Mahler, Maria Holmqvist, Keith B. Hall, Stefan Riezler, Enrique Alfonseca
SIGIR1