VLDB 2026 Research / reviewers in the wild / expert
Ingmar Weber
dblp:w/IngmarWeber
· DBLP profile ↗
80ranked-venue papers in the field
15as first author
10since 2021 · last 2026
0000-0003-4169-2579ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 59 (9 first)Data Mining & Knowledge Discovery · 12 (5 first)Database Systems & Data Management · 5 (1 first)Other / Interdisciplinary · 2Big Data, Cloud & Distributed Data Systems · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bowling with ChatGPT: On the Evolving User Interactions with Conversational AI SystemsabstractRecent studies have discussed how users are increasingly using conversational AI systems, powered by LLMs, for information seeking, decision support, and even emotional support. However, these macro-level observations offer limited insight into how the purpose of these interactions shifts over time, how users frame their interactions with the system, and how steering dynamics unfold in these human-AI interactions. To examine these evolving dynamics, we gathered and analyzed a unique dataset InVivoGPT: consisting of 825K ChatGPT interactions, donated by 300 users through their GDPR data rights. Our analyses reveal three key findings. First, participants increasingly turn to ChatGPT for a broader range of purposes, including substantial growth in sensitive domains such as health and mental health. Second, interactions become more socially framed: the system anthropomorphizes itself at rising rates, participants more frequently treat it as a companion, and personal data disclosure becomes both more common and more diverse. Third, conversational steering becomes more prominent, especially after the release of GPT-4o, with conversations where the participants followed a model-initiated suggestion quadrupling over the period of our dataset. Overall, our results show that conversational AI systems are shifting from functional tools to social partners, raising important questions about their design and governance. Sai Keerthana Karnam, Abhisek Dash, Krishna P. Gummadi, Animesh Mukherjee 0001, Ingmar Weber, Savvas Zannettou |
WWW | 5 |
| 2025 | From Satellites to Social Media: What Data Tells us About Society
Ingmar Weber |
DATA | 1 |
| 2025 | Ethnic Diversity and Spatial Dynamics in Google Maps Reviews: Insights from a German Border CityabstractLocation-based online services, such as Google Maps, provide a valuable lens for examining social dynamics in both online (virtual) and offline (physical) spaces. In particular, online reviews offer insights into how cultural or ethnic differences shape mobility patterns, place preferences, and public expression. This study analyzes the spatial behaviour of select ethnic groups in a German border city by integrating Google Maps reviews with demographic data. Using comparative analysis, we identify patterns in place engagement and disparities in urban space usage across different population groups. The findings highlight how socio-demographic factors influence the frequency and types of places visited, revealing gaps in urban accessibility. These insights demonstrate the potential of Google Maps data for understanding socio-spatial dynamics and inform strategies for more inclusive and data-driven urban planning. Ethel Elikem Mensah, Vikram Kamath Cannanure, Ingmar Weber |
ICWSM | 3 |
| 2025 | Studying Behavioral Addiction by Combining Surveys and Digital Traces: A Case Study of TikTokabstractOpaque algorithms disseminate and mediate the content that users consume on online social media platforms. This algorithmic mediation serves users with contents of their liking, on the other hand, it may cause several inadvertent risks to society at scale. While some of these risks, e.g., filter bubbles or dissemination of hateful content, are well studied in the community, behavioral addiction, designated by the Digital Services Act (DSA) as a potential systemic risk, has been understudied. In this work, we aim to study if one can effectively diagnose behavioral addiction using digital data traces from social media platforms. Focusing on the TikTok short-format video platform as a case study, we employ a novel mixed methodology of combining survey responses with data donations of behavioral traces. We survey 1590 TikTok users and stratify them into three addiction groups (i.e., less/moderately/highly likely addicted). Then, we obtain data donations from 107 surveyed participants. By analyzing users' data we find that, among others, highly likely addicted users spend more time watching TikTok videos and keep coming back to TikTok throughout the day, indicating a compulsion to use the platform. Finally, by using basic user engagement features, we train classifier models to identify highly likely addicted users with F1 >= 0.55. The performance of the classifier models suggests predicting addictive users solely based on their usage is rather difficult. Sepehr Mousavi, Abhisek Dash, Krishna P. Gummadi, Ingmar Weber |
ICWSM | 5 |
| 2024 | Exploring Global Gender Gaps in the Blockchain Domain: Insights from LinkedIn Advertising DataabstractBlockchain technology has gained widespread attention through Bitcoin, but the blockchain domain is still striving to increase gender diversity and widely assess skills gaps by gender. There is limited awareness of women’s participation in blockchain, prompting this study to assess and explore global gender gaps in interests, skills, and professions within the field. By analyzing gender-disaggregated data from LinkedIn’s advertisement platform, we reveal that women are significantly underrepresented in blockchain compared to men, with the gender gap being even more pronounced than in the broader IT sector. This study delves into the volume, velocity, variety, veracity, and value that LinkedIn Ad data offers to assess gender gaps in the blockchain domain at a global level. Reham Al Tamime, Markus Strohmaier, Ingmar Weber |
IEEE Big Data | 3 |
| 2024 | Analyzing Mentions of Death in COVID-19 TweetsabstractMany researchers have analyzed the potential of using tweets for epidemiology in general and for nowcasting COVID-19 trends in specific. Here, we focus on a subset of tweets that mention a personal, COVID-related death. We show that focusing on this set improves the correlation with official death statistics in six countries, while also picking up on mortality trends specific to different age groups and socio-economic groups. Furthermore, qualitative analysis reveals how politicized many of the mentioned deaths are. To help others reproduce and build on our work, we release a dataset of annotated tweets for academic research. Divya Mani Adhikari, Muhammad Imran 0002, Umair Qazi, Ingmar Weber |
ICWSM | 4 |
| 2024 | Gender Gaps in Online Social Connectivity, Promotion and Relocation Reports on LinkedInabstractOnline professional social networking platforms provide opportunities to expand networks strategically for job opportunities and career advancement. A large body of research shows that women’s offline networks are less advantageous than men’s. How online platforms such as LinkedIn may reflect or reproduce gendered networking behaviours, or how online social connectivity may affect outcomes differentially by gender is not well understood. This paper analyses aggregate, anonymised data from almost 10 million LinkedIn users in the UK and US information technology (IT) sector collected from the site’s advertising platform to explore how being connected to Big Tech companies (‘social connectivity’) varies by gender, and how gender, age, seniority and social connectivity shape the propensity to report job promotions or relocations. Consistent with previous studies, we find there are fewer women compared to men on LinkedIn in IT. Furthermore, female users are less likely to be connected to Big Tech companies than men. However, when we further analyse recent promotion or relocation reports, we find women are more likely than men to have reported a recent promotion at work, suggesting high-achieving women may be self-selecting onto LinkedIn. Even among this positively selected group, though, we find men are more likely to report a recent relocation. Social connectivity emerges as a significant predictor of promotion and relocation reports, with an interaction effect between gender and social connectivity indicating the payoffs to social connectivity for promotion and relocation reports are larger for women. This suggests that online networking has the potential for larger impacts for women, who experience greater disadvantage in traditional networking contexts, and calls for further research to understand differential impacts of online networking for socially disadvantaged groups. Ghazal Kalhor, Hannah Gardner, Ingmar Weber, Ridhi Kashyap |
ICWSM | 3 |
| 2023 | Partisan US News Media Representations of Syrian RefugeesabstractWe investigate how representations of Syrian refugees (2011-2021) differ across US partisan news outlets. We analyze 47,388 articles from the online US media about Syrian refugees to detail differences in reporting between left- and right-leaning media. We use various NLP techniques to understand these differences. Our polarization and question answering results indicated that left-leaning media tended to represent refugees as child victims, welcome in the US, and right-leaning media cast refugees as Islamic terrorists. We noted similar results with our sentiment and offensive speech scores over time, which detail possibly unfavorable representations of refugees in right-leaning media. A strength of our work is how the different techniques we have applied validate each other. Based on our results, we provide several recommendations. Stakeholders may utilize our findings to intervene around refugee representations, and design communications campaigns that improve the way society sees refugees and possibly aid refugee outcomes. Marzieh Babaeianjelodar, Yiwen Shi, Kamila Janmohamed, Rupak Sarkar, Ingmar Weber, Thomas Davidson, Munmun De Choudhury, Jonathan Huang, Shweta Yadav 0001, Ashiqur R. KhudaBukhsh, Chris T. Bauch, Preslav Nakov, Orestis Papakyriakopoulos, Koustuv Saha, Kaveh Khoshnood, Navin Kumar 0004 |
ICWSM | 6 |
| 2023 | Host-Centric Social Connectedness of Migrants in Europe on FacebookabstractExtant literature has explored the social integration process of migrants settling in host communities. However, this literature typically takes a migrant-centric view, implicitly putting the burden of a successful integration on the migrant, and trying to identify the factors that lead to integration along various dimensions. In this paper, we flip this point of view by studying the attributes of natives that govern their propensity to form social ties with migrants.We do so by using anonymous and aggregate social network data provided by Facebook’s advertising platform. More specifically, we look at factors that influence the propensity for a likely-to-be non-Muslim Facebook user to have at least one social connection to a Facebook user who celebrates Ramadan. Given that, in the European context, following Islam is predominantly tied to a migration background, this gives us a lens into cross-cultural native-migrant connectivity. Our study considers demographic attributes of the host population, such as age, gender, and education level, as well as spatial variation across 30 European cities. Our findings suggest that young, educated, and male Facebook users are relatively more likely to build cross-cultural ties, compared to older, less educated, and female Facebook users. We also observe heterogeneity across the analyzed cities. Aparup Khatua, Emilio Zagheni, Ingmar Weber |
ICWSM | 3 |
| 2023 | Gender Pay Gap in Sports on a Fan-Request Celebrity Video SiteabstractThe internet is often thought of as a democratizer, enabling equality in aspects such as pay, as well as a tool introducing novel communication and monetization opportunities. In this study we examine athletes on Cameo, a website that enables bi-directional fan-celebrity interactions, questioning whether the well-documented gender pay gaps in sports persist in this digital setting. Traditional studies into gender pay gaps in sports are mostly in a centralized setting where an organization decides the pay for the players, while Cameo facilitates grass-roots fan engagement where fans pay for video messages from their preferred athletes. The results showed that even on such a platform gender pay gaps persist, both in terms of cost-per-message, and in the number of requests, proxied by number of ratings. For instance, we find that female athletes have a median pay of 30$ per-video, while the same statistic is 40$ for men. The results also contribute to the study of parasocial relationships and personalized fan engagements over a distance. Something that has become more relevant during the ongoing COVID-19 pandemic, where in-person fan engagement has often been limited. Nazanin Sabri, Stephen Reysen, Ingmar Weber |
WWW | 3 |
| 2020 | The Relative Value of Facebook Advertising Data for Poverty Mapping
Masoomali Fatehkia, Benjamin Coles, Ferda Ofli, Ingmar Weber |
ICWSM | 4 |
| 2020 | Facebook Ads as a Demographic Tool to Measure the Urban-Rural DivideabstractIn the global move toward urbanization, making sure the people remaining in rural areas are not left behind in terms of development and policy considerations is a priority for governments worldwide. However, it is increasingly challenging to track important statistics concerning this sparse, geographically dispersed population, resulting in a lack of reliable, up-to-date data. In this study, we examine the usefulness of the Facebook Advertising platform, which offers a digital “census” of over two billions of its users, in measuring potential rural-urban inequalities. We focus on Italy, a country where about 30% of the population lives in rural areas. First, we show that the population statistics that Facebook produces suffer from instability across time and incomplete coverage of sparsely populated municipalities. To overcome such limitation, we propose an alternative methodology for estimating Facebook Ads audiences that nearly triples the coverage of the rural municipalities from 19% to 55% and makes feasible fine-grained sub-population analysis. Using official national census data, we evaluate our approach and confirm known significant urban-rural divides in terms of educational attainment and income. Extending the analysis to Facebook-specific user “interests” and behaviors, we provide further insights on the divide, for instance, finding that rural areas show a higher interest in gambling. Notably, we find that the most predictive features of income in rural areas differ from those for urban centres, suggesting researchers need to consider a broader range of attributes when examining rural wellbeing. The findings of this study illustrate the necessity of improving existing tools and methodologies to include under-represented populations in digital demographic studies – the failure to do so could result in misleading observations, conclusions, and most importantly, policies. Daniele Rama, Yelena Mejova, Michele Tizzoni, Kyriaki Kalimeri, Ingmar Weber |
WWW | 5 |
| 2019 | Rock, Rap, or Reggaeton?: Assessing Mexican Immigrants' Cultural Assimilation Using Facebook Data, abstractThe degree to which Mexican immigrants in the U.S. are assimilating culturally has been widely debated. To examine this question, we focus on musical taste, a key symbolic resource that signals the social positions of individuals. We adapt an assimilation metric from earlier work to analyze self-reported musical interests among immigrants in Facebook. We use the relative levels of interest in musical genres, where a similarity to the host population in musical preferences is treated as evidence of cultural assimilation. Contrary to skeptics of Mexican assimilation, we find significant cultural convergence even among first-generation immigrants, which problematizes their use as assimilative “benchmarks” in the literature. Further, 2nd generation Mexican Americans show high cultural convergence vis-à-vis both Anglos and African-Americans, with the exception of those who speak Spanish. Rather than conforming to a single assimilation path, our findings reveal how Mexican immigrants defy simple unilinear theoretical expectations and illuminate their uniquely heterogeneous character. René D. Flores, Timothy Riffe, Ingmar Weber, Emilio Zagheni |
WWW | 4 |
| 2018 | Professional Gender Gaps Across US Cities
Karri Haranko, Emilio Zagheni, Venkata Rama Kiran Garimella, Ingmar Weber |
ICWSM | 4 |
| 2018 | Mater Certa Est, Pater Numquam: What Can Facebook Advertising Data Tell Us about Male Fertility Rates?
Francesco Rampazzo, Emilio Zagheni, Ingmar Weber, Maria Rita Testa, Francesco C. Billari |
ICWSM | 3 |
| 2018 | A Social Media Based Examination of the Effects of Counseling Recommendations after Student Deaths on College Campuses
Koustuv Saha, Ingmar Weber, Munmun De Choudhury |
ICWSM | 2 |
| 2017 | Visualizing Health Awareness in the Middle East
Matheus Araújo 0001, Yelena Mejova, Michaël Aupetit 0001, Ingmar Weber |
ICWSM | 4 |
| 2017 | Automated Hate Speech Detection and the Problem of Offensive Language
Thomas Davidson, Dana Warmsley, Michael W. Macy, Ingmar Weber |
ICWSM | 4 |
| 2017 | A Long-Term Analysis of Polarization on Twitter
Venkata Rama Kiran Garimella, Ingmar Weber |
ICWSM | 2 |
| 2017 | Face-to-BMI: Using Computer Vision to Infer Body Mass Index on Social Media
Enes Kocabey, Mustafa Camurcu, Ferda Ofli, Yusuf Aytar, Antonio Torralba 0001, Ingmar Weber |
ICWSM | 7 |
| 2017 | Fertility and Its Meaning: Evidence from Search Behavior
Jussi Ojala, Emilio Zagheni, Francesco C. Billari, Ingmar Weber |
ICWSM | 4 |
| 2017 | Is Saki #delicious?: The Food Perception Gap on Instagram and Its Relation to HealthabstractFood is an integral part of our life and what and how much we eat crucially affects our health. Our food choices largely depend on how we perceive certain characteristics of food, such as whether it is healthy, delicious or if it qualifies as a salad. But these perceptions differ from person to person and one person's "single lettuce leaf" might be another person's "side salad". Studying how food is perceived in relation to what it actually is typically involves a laboratory setup. Here we propose to use recent advances in image recognition to tackle this problem. Concretely, we use data for 1.9 million images from Instagram from the US to look at systematic differences in how a machine would objectively label an image compared to how a human subjectively does. We show that this difference, which we call the "perception gap", relates to a number of health outcomes observed at the county level. To the best of our knowledge, this is the first time that image recognition is being used to study the "misalignment" of how people describe food images vs. what they actually depict. Ferda Ofli, Yusuf Aytar, Ingmar Weber, Raggi al Hammouri, Antonio Torralba 0001 |
WWW | 3 |
| 2016 | From migration corridors to clusters: The value of Google+ data for migration studiesabstractRecently, there have been considerable efforts to use online data to investigate international migration. These efforts show that Web data are valuable for estimating migration rates and are relatively easy to obtain. However, existing studies have only investigated flows of people along migration corridors, i.e. between pairs of countries. In our work, we use data about “places lived” from millions of Google+ users in order to study migration `clusters', i.e. groups of countries in which individuals have lived sequentially. For the first time, we consider information about more than two countries people have lived in. We argue that these data are very valuable because this type of information is not available in traditional demographic sources which record country-to-country migration flows independent of each other. We show that migration clusters of country triads cannot be identified using information about bilateral flows alone. To demonstrate the additional insights that can be gained by using data about migration clusters, we first develop a model that tries to predict the prevalence of a given triad using only data about its constituent pairs. We then inspect the groups of three countries which are more or less prominent, compared to what we would expect based on bilateral flows alone. Next, we identify a set of features such as a shared language or colonial ties that explain which triple of country pairs are more or less likely to be clustered when looking at country triples. Then we select and contrast a few cases of clusters that provide some qualitative information about what our data set shows. The type of data that we use is potentially available for a number of social media services. We hope that this first study about migration clusters will stimulate the use of Web data for the development of new theories of international migration that could not be tested appropriately before. Johnnatan Messias, Fabrício Benevenuto, Ingmar Weber, Emilio Zagheni |
ASONAM | 3 |
| 2016 | #greysanatomy vs. #yankees: Demographics and Hashtag Use on Twitter
Jisun An, Ingmar Weber |
ICWSM | 2 |
| 2016 | Social Media Participation in an Activist Movement for Racial Equality
Munmun De Choudhury, Shagun Jhaver, Benjamin Sugar, Ingmar Weber |
ICWSM | 4 |
| 2016 | The Road to Popularity: The Dilution of Growing Audience on Twitter
Przemyslaw A. Grabowicz, Mahmoudreza Babaei, Juhi Kulshrestha, Ingmar Weber |
ICWSM | 4 |
| 2016 | You Are What Apps You Use: Demographic Prediction Based on User's Apps
Eric Malmi, Ingmar Weber |
ICWSM | 2 |
| 2016 | Analyzing the Targets of Hate in Online Social Media
Leandro Araújo, Mainack Mondal, Denzil Correa, Fabrício Benevenuto, Ingmar Weber |
ICWSM | 5 |
| 2016 | A Large-Scale Study of Online Shopping BehaviorabstractThe continuous growth of e-commerce has stimulated great interest in generating theories and models for online consumer behavior. While studies on online consumer behavior are widespread, research on relating Internet browsing activities to online shopping behavior are scarce. This paper provides an exploratory analysis on the relationship between online browsing habits and consumers' pre-shopping effort, as one of the indicators of shopping behavior. The data used in this study was extracted from 88,637 users with more than half a million shopping instances from two large online retailers, Amazon and Walmart. Our findings provide insights for scholars to form hypotheses and design models or theories to explain online consumer behavior. Practitioners may also use the results of this study to make strategic decisions. Soroosh Nalchigar, Ingmar Weber, Parisa Lak, Ayse Basar Bener |
IDEAS | 2 |
| 2015 | Understanding Musical Diversity via Online Social Media
Ingmar Weber, Mor Naaman, Sarah Vieweg |
ICWSM | 2 |
| 2015 | Algorithms and criteria for diversification of news article comments
Giorgos Giannopoulos, Marios Koniaris, Ingmar Weber, Alejandro Jaimes, Timos K. Sellis |
J. Intell. Inf. Syst. | 3 |
| 2014 | Gender Asymmetries in Reality and Fiction: The Bechdel Test of Social Media
David García 0001, Ingmar Weber, Venkata Rama Kiran Garimella |
ICWSM | 2 |
| 2014 | Get Back! You Don't Know Me Like That: The Social Mediation of Fact Checking Interventions in Twitter Conversations
Aniko Hannak, Drew Margolin, Brian Keegan, Ingmar Weber |
ICWSM | 4 |
| 2014 | Using Co-Following for Personalized Out-of-Context Twitter Friend Recommendation
Ingmar Weber, Venkata Rama Kiran Garimella |
ICWSM | 1 |
| 2014 | Visualizing User-Defined, Discriminative Geo-Temporal Twitter Activity
Ingmar Weber, Venkata Rama Kiran Garimella |
ICWSM | 1 |
| 2014 | Who watches (and shares) what on youtube? and when?: using twitter to understand youtube viewershipabstractBy combining multiple social media datasets, it is possible to gain insight into each dataset that goes beyond what could be obtained with either individually. In this paper we combine user-centric data from Twitter with video-centric data from YouTube to build a rich picture of who watches and shares what on YouTube. We study 87K Twitter users, 5.6 million YouTube videos and 15 million video sharing events from user-, video- and sharing-event-centric perspectives. We show that features of Twitter users correlate with YouTube features and sharing-related features. For example, urban users are quicker to share than rural users. We find a superlinear relationship between initial Twitter shares and the final amounts of views. We discover that Twitter activity metrics play more role in video popularity than mere amount of followers. We also reveal the existence of correlated behavior concerning the time between video creation and sharing within certain timescales, showing the time onset for a coherent response, and the time limit after which collective responses are extremely unlikely. Response times depend on the category of the video, suggesting Twitter video sharing is highly dependent on the video content. To the best of our knowledge, this is the first large-scale study combining YouTube and Twitter data, and it reveals novel, detailed insights into who watches (and shares) what on YouTube, and when. Adiya Abisheva, Venkata Rama Kiran Garimella, David García 0001, Ingmar Weber |
WSDM | 4 |
| 2014 | Query recommendation in the information domain of childrenabstractChildren represent an increasing group of web users. Some of the key problems that hamper their search experience is their limited vocabulary, their difficulty in using the right keywords, and the inappropriateness of their general‐purpose query suggestions. In this work, we propose a method that uses tags from social media to suggest queries related to children's topics. Concretely, we propose a simple yet effective approach to bias a random walk defined on a bipartite graph of web resources and tags through keywords that are more commonly used to describe resources for children. We evaluate our method using a large query log sample of queries submitted by children. We show that our method outperforms by a large margin the query suggestions of modern search engines and state‐of‐the art query suggestions based on random walks. We improve further the quality of the ranking by combining the score of the random walk with topical and language modeling features to emphasize even more the child‐related aspects of the query suggestions. Sergio Duarte Torres, Djoerd Hiemstra, Ingmar Weber, Pavel Serdyukov |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2014 | Analysis of Search and Browsing Behavior of Young Users on the WebabstractThe Internet is increasingly used by young children for all kinds of purposes. Nonetheless, there are not many resources especially designed for children on the Internet and most of the content online is designed for grown-up users. This situation is problematic if we consider the large differences between young users and adults since their topic interests, computer skills, and language capabilities evolve rapidly during childhood. There is little research aimed at exploring and measuring the difficulties that children encounter on the Internet when searching for information and browsing for content. In the first part of this work, we employed query logs from a commercial search engine to quantify the difficulties children of different ages encounter on the Internet and to characterize the topics that they search for. We employed query metrics (e.g., the fraction of queries posed in natural language), session metrics (e.g., the fraction of abandoned sessions), and click activity (e.g., the fraction of ad clicks). The search logs were also used to retrace stages of child development. Concretely, we looked for changes in interests (e.g., the distribution of topics searched) and language development (e.g., the readability of the content accessed and the vocabulary size). In the second part of this work, we employed toolbar logs from a commercial search engine to characterize the browsing behavior of young users, particularly to understand the activities on the Internet that trigger search. We quantified the proportion of browsing and search activity in the toolbar sessions and we estimated the likelihood of a user to carry out search on the Web vertical and multimedia verticals (i.e., videos and images) given that the previous event is another search event or a browsing event. We observed that these metrics clearly demonstrate an increased level of confusion and unsuccessful search sessions among children. We also found a clear relation between the reading level of the clicked pages and characteristics of the users such as age and educational attainment. In terms of browsing behavior, children were found to start their activities on the Internet with a search engine (instead of directly browsing content) more often than adults. We also observed a significantly larger amount of browsing activity for the case of teenager users. Interestingly we also found that if children visit knowledge-related Web sites (i.e., information-dense pages such as Wikipedia articles), they subsequently do more Web searches than adults. Additionally, children and especially teenagers were found to have a greater tendency to engage in multimedia search, which calls to improve the aggregation of multimedia results into the current search result pages. Sergio Duarte Torres, Ingmar Weber, Djoerd Hiemstra |
ACM Trans. Web | 2 |
| 2013 | #Egypt: visualizing Islamist vs. secular tension on TwitterabstractWe present a demo that shows Twitter hashtag usage in Egypt from the angle of Islamist vs. secular polarization. The demo not only provides insights about current events in Egypt, but also reveals differences in attitudes of the two political camps with respect to major events abroad. The demo is publicly accessible at http://sc1.qcri.org/twitter/egypt/weekly/. Ingmar Weber, Venkata Rama Kiran Garimella |
ASONAM | 1 |
| 2013 | Secular vs. Islamist polarization in Egypt on TwitterabstractWe use public data from Twitter, both in English and Arabic, to study the phenomenon of secular vs. Islamist polarization in Twitter. Starting with a set of prominent seed Twitter users from both camps, we follow retweeting edges to obtain an extended network of users with inferred political orientation. We present an in-depth description of the members of the two camps, both in terms of behavior on Twitter and in terms of offline characteristics such as gender. Through the identification of partisan users, we compute a valence on the secular vs. Islamist axis for hashtags and use this information both to analyze topical interests and to quantify how polarized society as a whole is at a given point in time. For the last 12 months, large values on this "polarization barometer" coincided with periods of violence. Tweets are furthermore annotated using hand-crafted dictionaries to quantify the usage of (i) religious terms, (ii) derogatory terms referring to other religions, and (ii) references to charitable acts. The combination of all the information allows us to test and quantify a number of stereo-typical hypotheses such as (i) that religiosity and political Islamism are correlated, (ii) that political Islamism and negative views on other religions are linked, (iii) that religiosity goes hand in hand with charitable giving, and (iv) that the followers of the Egyptian Muslim Brotherhood are more tightly connected and expressing themselves "in unison" than the secular opposition. Whereas a lot of existing literature on the Arab Spring and the Egyptian Revolution is largely of qualitative and descriptive nature, our contribution lies in providing a quantitative and data-driven analysis of online communication in this dynamic and politically charged part of the world. Ingmar Weber, Venkata Rama Kiran Garimella, Alaa Batayneh |
ASONAM | 1 |
| 2013 | PLEAD 2013: politics, elections and dataabstractWhat is the role of the internet in politics general and during campaigns in particular? And what is the role of large amounts of user data in all of this? Ingmar Weber, Ana-Maria Popescu, Marco Pennacchiotti |
CIKM | 1 |
| 2013 | Detecting Friday Night Party Photos: Semantics for Tag Recommendation
Philip J. McParlane, Yelena Mejova, Ingmar Weber |
ECIR | 3 |
| 2013 | Political Hashtag Trends
Ingmar Weber, Venkata Rama Kiran Garimella, Asmelash Teka Hadgu |
ECIR | 1 |
| 2013 | From Republicans to Teenagers - Group Membership and Search (GRUMPS)
Ingmar Weber, Djoerd Hiemstra, Pavel Serdyukov |
ECIR | 1 |
| 2013 | Learning to question: leveraging user preferences for shopping adviceabstractWe present ShoppingAdvisor, a novel recommender system that helps users in shopping for technical products. ShoppingAdvisor leverages both user preferences and technical product attributes in order to generate its suggestions. The system elicits user preferences via a tree-shaped flowchart, where each node is a question to the user. At each node, ShoppingAdvisor suggests a ranking of products matching the preferences of the user, and that gets progressively refined along the path from the tree's root to one of its leafs. Mahashweta Das, Gianmarco De Francisci Morales, Aristides Gionis, Ingmar Weber |
KDD | 4 |
| 2013 | From machu_picchu to "rafting the urubamba river": anticipating information needs via the entity-query graphabstractWe study the problem of anticipating user search needs, based on their browsing activity. Given the current web page p that a user is visiting we want to recommend a small and diverse set of search queries that are relevant to the content of p, but also non-obvious and serendipitous. Ilaria Bordino, Gianmarco De Francisci Morales, Ingmar Weber, Francesco Bonchi |
WSDM | 3 |
| 2013 | Studying inter-national mobility through IP geolocationabstractThe increasing ubiquity of Internet use has opened up new avenues in the study of human mobility. Easily-obtainable geolocation data resulting from repeated logins to the same website offer the possibility of observing long-term patterns of mobility for a large number of individuals. We use data on the geographic locations from where over 100 million anonymized users log into Yahoo!~services to generate the first global map of short- and medium-term mobility flows. We develop a protocol to identify anonymized users who, over a one-year period, had spent more than 3 months in a different country from their stated country of residence ("migrants"), and users who spent less than a month in another country ("tourists"). We compute aggregate estimates of migration propensities between countries, as inferred from a user's location over the observed period. Geolocation data allow us to characterize also the pendularity of migration flows -- i.e., the extent to which migrants travel back and forth between their countries of origin and destination. We use data regarding visa regimes, colonial ties, geographic location and economic development to predict migration and tourism flows. Our analysis shows the persistence of traditional migration patterns as well as the emergence of new routes. Migrations tend to be more pendular between countries that are close to each other. We observe particularly high levels of pendularity within the European Economic Area, even after we control for distance and visa regimes. The dataset, methodology and results presented have important implications for the travel industry, as well as for several disciplines in social sciences, including geography, demography and the sociology of networks. Bogdan State, Ingmar Weber, Emilio Zagheni |
WSDM | 2 |
| 2013 | Data-driven political scienceabstractThe tutorial will summarize the state-of-the art in the growing area of computational political science. Like many others, this research domain is being revolutionized by the availability of open, big data and the increasing reach and importance of social media. The surging interest on the part of the academic community is matched by intense efforts on the part of political campaigns to use online data in order to learn how to best disseminate information and reach the right potential donors or voters. In this context, a tutorial can summarize existing methods in a fascinating, high-interest area and allow participants with diverse backgrounds to get inspiration from the methods and problems studied. The tutorial will feature seminal research concerning (i) political polarization, (ii) election prediction and polling, and (iii) political campaigning and influence propagation. The goal is not only to familiarize attendees with ideas from related conferences such as WWW, ICWSM or CIKM, but also to present ideas and quantitative methods closer to political science such as Poole's and Rosenthal's NOMINATE score for a politician's political orientation. Ingmar Weber, Ana-Maria Popescu, Marco Pennacchiotti |
WSDM | 1 |
| 2013 | Sponsored search, market equilibria, and the Hungarian Method
Paul Dütting, Monika Henzinger, Ingmar Weber |
Inf. Process. Lett. | 3 |
| 2013 | Piggybacking on Social NetworksabstractThe popularity of social-networking sites has increased rapidly over the last decade. A basic functionalities of social-networking sites is to present users with streams of events shared by their friends. At a systems level, materialized per-user views are a common way to assemble and deliver such event streams on-line and with low latency. Access to the data stores, which keep the user views, is a major bottleneck of social-networking systems. We propose to improve the throughput of these systems by using social piggybacking, which consists of processing the requests of two friends by querying and updating the view of a third common friend. By using one such hub view, the system can serve requests of the first friend without querying or updating the view of the second. We show that, given a social graph, social piggybacking can minimize the overall number of requests, but computing the optimal set of hubs is an NP-hard problem. We propose anO(logn) approximation algorithm and a heuristic to solve the problem, and evaluate them using the full Twitter and Flickr social graphs, which have up to billions of edges. Compared to existing approaches, using social piggybacking results in similar throughput in systems with few servers, but enables substantial throughput improvements as the size of the system grows, reaching up to a 2-factor increase. We also evaluate our algorithms on a real social networking system prototype and we show that the actual increase in throughput corresponds nicely to the gain anticipated by our cost function. Aristides Gionis, Flavio Paiva Junqueira, Vincent Leroy 0001, Marco Serafini, Ingmar Weber |
Proc. VLDB Endow. | 5 |
| 2013 | A Comprehensive Study of Techniques for URL-Based Web Page Language ClassificationabstractGiven only the URL of a Web page, can we identify its language? In this article we examine this question. URL-based language classification is useful when the content of the Web page is not available or downloading the content is a waste of bandwidth and time. We built URL-based language classifiers for English, German, French, Spanish, and Italian by applying a variety of algorithms and features. As algorithms we used machine learning algorithms which are widely applied for text classification and state-of-art algorithms for language identification of text. As features we used words, various sized n-grams, and custom-made features (our novel feature set). We compared our approaches with two baseline methods, namely classification by country code top-level domains and classification by IP addresses of the hosting Web servers. We trained and tested our classifiers in a 10-fold cross-validation setup on a dataset obtained from the Open Directory Project and from querying a commercial search engine. We obtained the lowest F1-measure for English (94) and the highest F1-measure for German (98) with the best performing classifiers. We also evaluated the performance of our methods: (i) on a set of Web pages written in Adobe Flash and (ii) as part of a language-focused crawler. In the first case, the content of the Web page is hard to extract and in the second page downloading pages of the “wrong” language constitutes a waste of bandwidth. In both settings the best classifiers have a high accuracy with an F1-measure between 95 (for English) and 98 (for Italian) for the Adobe Flash pages and a precision between 90 (for Italian) and 97 (for French) for the language-focused crawler. Eda Baykan, Monika Henzinger, Ingmar Weber |
ACM Trans. Web | 3 |
| 2012 | Maximizing revenue from strategic recommendations under decaying trustabstractSuppose your sole interest in recommending a product to me is to maximize the amount paid to you by the seller for a sequence of recommendations. How should you recommend optimally if I become more inclined to ignore you with each irrelevant recommendation you make? Finding an answer to this question is a key challenge in all forms of marketing that rely on and explore social ties; ranging from personal recommendations to viral marketing. Paul Dütting, Monika Henzinger, Ingmar Weber |
CIKM | 3 |
| 2012 | Query recommendation for childrenabstractOne of the biggest problems that children experience while searching the web occurs during the query formulation process. Children have been found to struggle formulating queries based on keywords given their limited vocabulary and their difficulty to choose the right keywords. Sergio Duarte Torres, Djoerd Hiemstra, Ingmar Weber, Pavel Serdyukov |
CIKM | 3 |
| 2012 | PLEAD 2012: politics, elections and dataabstractWhat is the role of the internet in politics general and during campaigns in particular? And what is the role of large amounts of user data in all of this? In the 2008 U.S. presidential campaign the Democrats were far more successful than the Republicans in utilizing online media for mobilization, co-ordination and fundraising. For the first time, social media and the Internet played a fundamental role in political campaigns. However, technical research in this area has been surprisingly limited and fragmented. The goal of this workshop is to bring together, for the first time, researchers working at the intersection of social network analysis, computational social science and political science, to share and discuss their ideas in a common forum; and to inspire further developments in this growing, fascinating field. The workshop has Filippo Menczer as keynote speaker, it includes technical presentations of accepted papers and concludes with a panel discussion where scientists and media experts from different fields can interact and share views. Ingmar Weber, Ana-Maria Popescu, Marco Pennacchiotti |
CIKM | 1 |
| 2012 | Political search trendsabstractWe present Political Search Trends, a browser based web search analysis tool that (i) assigns a political leaning to web search queries, (ii) detects trending political queries in a given week, and (iii) links search queries to fact-checked statements. In terms of methodology, it showcases the power of analyzing queries leading to clicks on selected, annotated web sites of interest. Ingmar Weber, Venkata Rama Kiran Garimella, Erik Borra |
SIGIR | 1 |
| 2012 | Diversifying User Comments on News Articles
Giorgos Giannopoulos, Ingmar Weber, Alejandro Jaimes, Timos K. Sellis |
WISE | 2 |
| 2012 | A large-scale sentiment analysis for Yahoo! answersabstractSentiment extraction from online web documents has recently been an active research topic due to its potential use in commercial applications. By sentiment analysis, we refer to the problem of assigning a quantitative positive/negative mood to a short bit of text. Most studies in this area are limited to the identification of sentiments and do not investigate the interplay between sentiments and other factors. In this work, we use a sentiment extraction tool to investigate the influence of factors such as gender, age, education level, the topic at hand, or even the time of the day on sentiments in the context of a large online question answering site. We start our analysis by looking at direct correlations, e.g., we observe more positive sentiments on weekends, very neutral ones in the Science & Mathematics topic, a trend for younger people to express stronger sentiments, or people in military bases to ask the most neutral questions. We then extend this basic analysis by investigating how properties of the (asker, answerer) pair affect the sentiment present in the answer. Among other things, we observe a dependence on the pairing of some inferred attributes estimated by a user's ZIP code. We also show that the best answers differ in their sentiments from other answers, e.g., in the Business & Finance topic, best answers tend to have a more neutral sentiment than other answers. Finally, we report results for the task of predicting the attitude that a question will provoke in answers. We believe that understanding factors influencing the mood of users is not only interesting from a sociological point of view, but also has applications in advertising, recommendation, and search. Onur Küçüktunç, Berkant Barla Cambazoglu, Ingmar Weber, Hakan Ferhatosmanoglu |
WSDM | 3 |
| 2012 | Answers, not links: extracting tips from yahoo! answers to address how-to web queriesabstractWe investigate the problem of mining "tips" from Yahoo! Answers and displaying those tips in response to related web queries. Here, a "tip" is a short, concrete and self-contained bit of non-obvious advice such as "To zest a lime if you don't have a zester : use a cheese grater." Ingmar Weber, Antti Ukkonen, Aristides Gionis |
WSDM | 1 |
| 2011 | What and how children search on the webabstractThe Internet has become an important part of the daily life of children as a source of information and leisure activities. Nonetheless, given that most of the content available on the web is aimed at the general public, children are constantly exposed to inappropriate content, either because the language goes beyond their reading skills, their attention span differs from grown-ups or simple because the content is not targeted at children as is the case of ads and adult content. In this work we employed a large query log sample from a commercial web search engine to identify the struggles and search behavior of children of the age of 6 to young adults of the age of 18. Concretely we hypothesized that the large and complex volume of information to which children are exposed leads to ill-defined searches and to disorientation during the search process. For this purpose, we quantified their search difficulties based on query metrics (e.g. fraction of queries posed in natural language), session metrics (e.g. fraction of abandoned sessions) and click activity (e.g. fraction of ad clicks). We also used the search logs to retrace stages of child development. Concretely we looked for changes in the user interests (e.g. distribution of topics searched), language development (e.g. readability of the content accessed) and cognitive development (e.g. sentiment expressed in the queries) among children and adults. We observed that these metrics clearly demonstrate an increased level of confusion and unsuccessful search sessions among children. We also found a clear relation between the reading level of the clicked pages and the demographics characteristics of the users such as age and average educational attainment of the zone in which the user is located. Sergio Duarte Torres, Ingmar Weber |
CIKM | 2 |
| 2011 | Who uses web search for what: and howabstractWe analyze a large query log of 2.3 million anonymous registered users from a web-scale U.S. search engine in order to jointly analyze their on-line behavior in terms of who they might be (demographics), what they search for (query topics), and how they search (session analysis). We examine basic demographics from registration information provided by the users, augmented with U.S. census data, analyze basic session statistics, classify queries into types (navigational, informational, transactional) based on click entropy, classify queries into topic categories, and cluster users based on the queries they issued. We then examine the resulting clusters in terms of demographics and search behavior. Our analysis of the data suggests that there are important differences in search behavior across different demographic groups in terms of the topics they search for, and how they search (e.g., white conservatives are those likely to have voted republican, mostly white males, who search for business, home, and gardening related topics; Baby Boomers tend to be primarily interested in Finance and a large fraction of their sessions consist of simple navigational queries related to online banking, etc.). Finally, we examine regional search differences, which seem to correlate with differences in local industries (e.g., gambling related queries are highest in Las Vegas and lowest in Salt Lake City; searches related to actors are about three times higher in L.A. than in any other region). Ingmar Weber, Alejandro Jaimes |
WSDM | 1 |
| 2011 | An expressive mechanism for auctions on the webabstractAuctions are widely used on the Web. Applications range from internet advertising to platforms such as eBay. In most of these applications the auctions in use are single/multi-item auctions with unit demand. The main drawback of standard mechanisms for this type of auctions, such as VCG and GSP, is the limited expressiveness that they offer to the bidders. The General Auction Mechanism (GAM) of [1] is taking a first step towards addressing the problem of limited expressiveness by computing a bidder optimal, envy free outcome for linear utility functions with identical slopes and a single discontinuity per bidder-item pair. We show that in many practical situations this does not suffice to adequately model the preferences of the bidders, and we overcome this problem by presenting the first mechanism for piece-wise linear utility functions with non-identical slopes and multiple discontinuities. Our mechanism runs in polynomial time. Like GAM it is incentive compatible for inputs that fulfill a certain non-degeneracy requirement, but our requirement is more general than the requirement of GAM. For discontinuous utility functions that are non-degenerate as well as for continuous utility functions the outcome of our mechanism is a competitive equilibrium. We also show how our mechanism can be used to compute approximately bidder optimal, envy free outcomes for a general class of continuous utility functions via piece-wise linear approximation. Finally, we prove hardness results for even more expressive settings. Paul Dütting, Monika Henzinger, Ingmar Weber |
WWW | 3 |
| 2011 | Offline file assignments for online load balancing
Paul Dütting, Monika Henzinger, Ingmar Weber |
Inf. Process. Lett. | 3 |
| 2011 | A Comprehensive Study of Features and Algorithms for URL-Based Topic ClassificationabstractGiven only the URL of a Web page, can we identify its topic? We study this problem in detail by exploring a large number of different feature sets and algorithms on several datasets. We also show that the inherent overlap between topics and the sparsity of the information in URLs makes this a very challenging problem. Web page classification without a page’s content is desirable when the content is not available at all, when a classification is needed before obtaining the content, or when classification speed is of utmost importance. For our experiments we used five different corpora comprising a total of about 3 million (URL, classification) pairs. We evaluated several techniques for feature generation and classification algorithms. The individual binary classifiers were then combined via boosting into metabinary classifiers. We achieve typical F-measure values between 80 and 85, and a typical precision of around 86. The precision can be pushed further over 90 while maintaining a typical level of recall between 30 and 40. Eda Baykan, Monika Henzinger, Ludmila Marian, Ingmar Weber |
ACM Trans. Web | 4 |
| 2011 | Camera Brand Congruence and Camera Model Propagation in the Flickr Social GraphabstractGiven that my friends on Flickr use cameras of brand X, am I more likely to also use a camera of brand X? Given that one of these friends changes her brand, am I likely to do the same? Do new camera models pop up uniformly in the friendship graph? Or do early adopters then “convert” their friends? Which factors influence the conversion probability of a user? These are the kind of questions addressed in this work. Direct applications involve personalized advertising in social networks. For our study, we crawled a complete connected component of the Flickr friendship graph with a total of 67M edges and 3.9M users. 1.2M of these users had at least one public photograph with valid model metadata, which allowed us to assign camera brands and models to users and time slots. Similarly, we used, where provided in a user’s profile, information about a user’s geographic location and the groups joined on Flickr. Concerning brand congruence, our main findings are the following. First, a pair of friends on Flickr has a higher probability of being congruent, that is, using the same brand, compared to two random users (27% vs. 19%). Second, the degree of congruence goes up for pairs of friends (i) in the same country (29%), (ii) who both only have very few friends (30%), and (iii) with a very high cliqueness (38%). Third, given that a user changes her camera model between March-May 2007 and March-May 2008, high cliqueness friends are more likely than random users to do the same (54% vs. 48%). Fourth, users using high-end cameras are far more loyal to their brand than users using point-and-shoot cameras, with a probability of staying with the same brand of 60% vs 33%, given that a new camera is bought. Fifth, these “expert” users’ brand congruence reaches 66% for high cliqueness friends. All these differences are statistically significant at 1%. As for the propagation of new models in the friendship graph, we observe the following. First, the growth of connected components of users converted to a particular, new camera model differs distinctly from random growth. Second, the decline of dissemination of a particular model is close to random decline. This illustrates that users influence their friends to change to a particular new model, rather than from a particular old model. Third, having many converted friends increases the probability of the user to convert herself. Here differences between friends from the same or from different countries are more pronounced for point-and-shoot than for digital single-lens reflex users. Fourth, there was again a distinct difference between arbitrary friends and high cliqueness friends in terms of prediction quality for conversion. Adish Singla, Ingmar Weber |
ACM Trans. Web | 2 |
| 2010 | Demographic information flowsabstractIn advertising and content relevancy prediction it is important to understand whether, over time, information that reaches one demographic group spreads to others. In this paper we analyze the query log of a large U.S. web search engine to determine whether the same queries are performed by different demographic groups at different times, particularly when there are query bursts. We obtain aggregate demographic features from user-provided registration information (gender, birth year, ZIP code), U.S. census data, and election results. Given certain queries, we examine trends (from high to low and vice versa) and changes in the statistical spread of the demographic features of users that issue the queries over time periods that include query bursts. Our analysis shows that for certain types of queries (movies and news) distinct demographic groups perform searches at different times, suggesting that information related to such queries flows between them. Queries of movie titles, for instance, tend to be issued first by young and then by older users, where a sudden jump in age occurs upon the movie's release. To the best of our knowledge, this is the first time this problem has been studied using search query logs. Ingmar Weber, Alejandro Jaimes |
CIKM | 1 |
| 2010 | The demographics of web searchabstractHow does the web search behavior of "rich" and "poor" people differ? Do men and women tend to click on difffferent results for the same query? What are some queries almost exclusively issued by African Americans? These are some of the questions we address in this study. Ingmar Weber, Carlos Castillo 0001 |
SIGIR | 1 |
| 2010 | How much is your personal recommendation worth?abstractSuppose you buy a new laptop and, simply because you like it so much, you recommend it to friends, encouraging them to purchase it as well. What would be an adequate price for the vendor of the laptop to pay for your recommendation? Paul Dütting, Monika Henzinger, Ingmar Weber |
WWW | 3 |
| 2010 | Tagging and navigabilityabstractWe consider the problem of optimal tagging for navigational purposes in one's own collection. What is the best that a forgetful user can hope for in terms of ease of retrieving a labeled object? We prove that the number of tags has to increase logarithmically in the collection size to maintain a manageable result set. Using Flickr data we then show that users do indeed apply more and more tags as their collection grows and that this is not due to a global increase in tagging activity. However, as the additional terms applied are not statistically independent, users of large collections still have to deal with larger and larger result sets, even when more tags are used as search terms. We pose optimal tag suggestion for navigational purposes as an open problem. Adish Singla, Ingmar Weber |
WWW | 2 |
| 2009 | Camera brand congruence in the Flickr social graphabstractGiven that my friends on Flickr use cameras of brand X, am I more likely to also use a camera of brand X? Given that one of these friends changes her brand, am I likely to do the same? These are the kind of questions addressed in this work. Direct applications involve personalized advertising in social networks. Adish Singla, Ingmar Weber |
WSDM | 2 |
| 2009 | Purely URL-based topic classificationabstractGiven only the URL of a web page, can we identify its topic? This is the question that we examine in this paper. Usually, web pages are classified using their content, but a URL-only classifier is preferable, (i) when speed is crucial, (ii) to enable content filtering before an (objection-able) web page is downloaded, (iii) when a page's content is hidden in images, (iv) to annotate hyperlinks in a personalized web browser, without fetching the target page, and (v) when a focused crawler wants to infer the topic of a target page before devoting bandwidth to download it. We apply a machine learning approach to the topic identification task and evaluate its performance in extensive experiments on categorized web pages from the Open Directory Project (ODP). When training separate binary classifiers for each topic, we achieve typical F-measure values between 80 and 85, and a typical precision of around 85. We also ran experiments on a small data set of university web pages. For the task of classifying these pages into faculty, student, course and project pages, our methods improve over previous approaches by 13.8 points of F-measure. Eda Baykan, Monika Henzinger, Ludmila Marian, Ingmar Weber |
WWW | 4 |
| 2009 | Rethinking email message and people searchabstractWe show how a number of novel email search features can be implemented without any kind of natural language processing (NLP) or advanced data mining. Our approach inspects the email headers of all messages a user has ever sent or received and it creates simple per-contact summaries, including simple information about the message exchange history, the domain of the sender or even the sender's gender. With these summaries advanced questions/tasks such as "Who do I still need to reply to?" or "Find 'fun' messages sent by friends." become possible. As a proof of concept, we implemented a Mozilla-Thunderbird extension, adding powerful people search to the popular email client. Sebastian Michel 0001, Ingmar Weber |
WWW | 2 |
| 2008 | Personalized, interactive tag recommendation for flickrabstractWe study the problem of personalized, interactive tag recom-mendation for Flickr: While a user enters/selects new tags for a particular picture, the system suggests related tags to her, based on the tags that she or other people have used in the past along with (some of) the tags already entered. The suggested tags are dynamically updated with every ad-ditional tag entered/selected. We describe a new algorithm, called Hybrid, which can be applied to this problem, and show that it outperforms previous algorithms. It has only a single tunable parameter, which we found to be very robust. Apart from this new algorithm and its detailed analysis, our main contributions are (i) a clean methodology which leads to conservative performance estimates, (ii) showing how classical classification algorithms can be applied to this problem, (iii) introducing a new cost measure, which cap-tures the effort of the whole tagging process, (iv) clearly identifying, when purely local schemes (using only a user’s tagging history) can or cannot be improved by global schemes (using everybody’s tagging history). Nikhil Garg 0005, Ingmar Weber |
RecSys | 2 |
| 2008 | Personalized tag suggestion for flickrabstractWe present a system for personalized tag suggestion for Flickr: While the user is entering/selecting new tags for a particular picture, the system is suggesting related tags to her, based on the tags that she or other people have used in the past along with (some of) the tags already entered. The suggested tags are dynamically updated with every ad-ditional tag entered/selected. We describe three algorithms which can be applied to this problem. In experiments, our best-performing method yields an improvement in precision of 10-15 % over a baseline method very similar to the sys-tem currently used by Flickr. Our system is accessible at Nikhil Garg 0005, Ingmar Weber |
WWW | 2 |
| 2008 | Output-sensitive autocompletion search
Hannah Bast, Christian Worm Mortensen, Ingmar Weber |
Inf. Retr. | 3 |
| 2008 | Web page language identification based on URLsabstractGiven only the URL of a web page, can we identify its language? This is the question that we examine in this paper. Such a language classifier is, for example, useful for crawlers of web search engines, which frequently try to satisfy certain language quotas. To determine the language of uncrawled web pages, they have to download the page, which might be wasteful, if the page is not in the desired language. With URL-based language classifiers these redundant downloads can be avoided. We apply a variety of machine learning algorithms to the language identification task and evaluate their performance in extensive experiments for five languages: English, French, German, Spanish and Italian. Our best methods achieve an F-measure, averaged over all languages, of around .90 for both a random sample of 1,260 web page from a large web crawl and for 25k pages from the ODP directory. For 5k pages of web search engine results we even achieve an F-measure of .96. The achieved recall for these collections is .93, .88 and .95 respectively. Two independent human evaluators performed considerably worse on the task, with an F-measure of .75 and a typical recall of a mere .67. Using only country-code top-level domains, such as .de or .fr yields a good precision, but a typical recall of below .60 and an F-measure of around .68. Eda Baykan, Monika Henzinger, Ingmar Weber |
Proc. VLDB Endow. | 3 |
| 2007 | The CompleteSearch Engine: Interactive, Efficient, and Towards IR& DB Integration
Hannah Bast, Ingmar Weber |
CIDR | 2 |
| 2007 | Efficient interactive query expansion with complete searchabstractWe present an efficient realization of the following interactive search engine \nfeature: as the user is typing the query, words that are related to the last \nquery word and that would lead to good hits are suggested, as well as selected \nsuch hits. The realization has three parts: (i) building clusters of related \nterms, (ii) adding this information as artificial words to the index such that \n(iii) the described feature reduces to an instance of prefix search and \ncompletion. An efficient solution for the latter is provided by the \nCompleteSearch engine, with which we have integrated the proposed feature. For \nbuilding the clusters of related terms we propose a variant of latent semantic \nindexing that, unlike standard approaches, is completely transparent to the \nuser. By experiments on two large test-collections, we demonstrate that the \nfeature is provided at only a slight increase in query processing time and \nindex size. Hannah Bast, Debapriyo Majumdar, Ingmar Weber |
CIKM | 3 |
| 2007 | ESTER: efficient search on text, entities, and relationsabstractWe present ESTER, a modular and highly efficient system for combined full-text and ontology search. ESTER builds on a query engine that supports two basic operations: prefix search and join. Both of these can be implemented very efficiently with a compact index, yet in combination provide powerful querying capabilities. We show how ESTER can answer basic SPARQL graph-pattern queries on the ontology by reducing them to a small number of these two basic operations. ESTER further supports a natural blend of such semantic queries with ordinary full-text queries. Moreover, the prefix search operation allows for a fully interactive and proactive user interface, which after every keystroke suggests to the user possible semantic interpretations of his or her query, and speculatively executes the most likely of these interpretations. As a proof of concept, we applied ESTER to the English Wikipedia, which contains about 3 million documents, combined with the recent YAGO ontology, which contains about 2.5 million facts. For a variety of complex queries, ESTER achieves worst-case query processing times of a fraction of a second, on a single machine, with an index size of about 4 GB. Hannah Bast, Alexandru Chitea, Fabian M. Suchanek, Ingmar Weber |
SIGIR | 4 |
| 2006 | Type less, find more: fast autocompletion search with a succinct indexabstractWe consider the following full-text search autocompletion feature. Imagine a user of a search engine typing a query. Then with every letter being typed, we would like an instant display of completions of the last query word which would lead to good hits. At the same time, the best hits for any of these completions should be displayed. Known indexing data structures that apply to this problem either incur large processing times for a substantial class of queries, or they use a lot of space. We present a new indexing data structure that uses no more space than a state-of-the-art compressed inverted index, but with 10 times faster query processing times. Even on the large TREC Terabyte collection, which comprises over 25 million documents, we achieve, on a single machine and with the index on disk, average response times of one tenth of a second. We have built a full-fledged, interactive search engine that realizes the proposed autocompletion feature combined with support for proximity search, semi-structured (XML) text, subword and phrase completion, and semantic tags. Hannah Bast, Ingmar Weber |
SIGIR | 2 |
| 2006 | Output-Sensitive Autocompletion SearchabstractWe consider the following autocompletion search scenario: imagine a user of a search engine typing a query; then with every keystroke display those completions of the last query word that would lead to the best hits, and also display the best such hits. The following problem is at the core of this feature: for a fixed document collection, given a set D of documents, and an alphabetical range W of words, compute the set of all word-in-document pairs ( w , d ) from the collection such that w ∈ W and d ∈ D . We present a new data structure with the help of which such autocompletion queries can be processed, on the average, in time linear in the input plus output size, independent of the size of the underlying document collection. At the same time, our data structure uses no more space than an inverted index. Actual query processing times on a large test collection correlate almost perfectly with our theoretical bound. Hannah Bast, Christian Worm Mortensen, Ingmar Weber |
SPIRE | 3 |