VLDB 2026 Research / reviewers in the wild / expert
Martha A. Larson
dblp:75/107
· DBLP profile ↗
132ranked-venue papers
12as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 75 · 7 first-author · 12 since 2021Databases, data management, data science and information retrieval · 49 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 32 · 3 first-author · 7 since 2021Security and privacy · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revealing the Impact of Visual Text Style on Attribute-based Descriptions Produced by Large Visual Language ModelsabstractWhen the visual style of text is considered, a wide variety can be observed in font, color, and size. However, when a word is read, its meaning is independent of the style in which it has been written or rendered. In this paper, we investigate whether, and how, the style in which a word is visualized in an image impacts the description that a Large Visual Language Model (LVLM) provides for the concept to which that word refers. Specifically, we investigate how functional text styles (readability-oriented, e.g., black sans-serif) versus decorative styles (display-oriented, e.g., colored cursive/script) affect LVLMs’ descriptions of a concept in terms of the attributes of that concept. Our experiments study the situation in which the LVLM is able to correctly identify the concept referred to by a visual text, i.e., by a word or words rendered as an image, and in which the visual text style should not influence the attribute-based description that the LVLM produces. Our experimental results reveal that even when the concept is correctly identified, text style influences the model’s attribute-based descriptions of the concept. Our findings demonstrate non-trivial style leakage from text style into semantic inference and motivate style-aware evaluation and mitigation for LVLM-based multimedia systems. Martha A. Larson, Zhengyu Zhao 0001 |
ICMR | 2 |
| 2026 | Frequency Is What You Need: Considering Word Frequency When Text Masking Benefits Vision-Language Model Pre-trainingabstractVision Language Models (VLMs) can be trained more efficiently if training sets can be reduced in size. Recent work has shown the benefits of masking text during VLM training using a variety of strategies (truncation, random masking, block masking and syntax masking) and has reported syntax masking as the top performer. In this paper, we analyze the impact of different text masking strategies on the word frequency in the training data, and show that this impact is connected to model success. This finding motivates Contrastive Language-Image Pre-training with Word Frequency Masking (CLIPF), our proposed masking approach, which directly leverages word frequency. Extensive experiments demonstrate the advantages of CLIPF over syntax masking and other existing approaches, particularly when the number of input tokens decreases. We show that not only CLIPF, but also other existing masking strategies, outperform syntax masking when enough epochs are used during training, a finding of practical importance for selecting a text masking method for VLM training. Our code is available online. Mingliang Liang, Martha A. Larson |
WACV | 2 |
| 2025 | Enhancing Vision-Language Model Pre-Training with Image-Text Pair Pruning Based on Word FrequencyabstractWe propose Word-Frequency-based Image-Text Pair Pruning (WFPP), a novel data pruning method that reduces the amount of data needed for training a Visual-Language Model (VLM) while preserving performance. WFPP removes image-text pairs across the entire training dataset according to a score calculated on the basis of word frequency. The effect of WFPP is to balance word frequency in the training data, i.e., reduce the dominance of highly frequent words, while retaining the contribution of less frequent words. WFPP is competitive with MetaCLIP, currently the leading text-image pair pruning approach, but has the advantage of not relying on metadata. Our experiments demonstrate that applying WFPP when training a CLIP model can boost performance on a wide range of downstream tasks. Mingliang Liang, Martha A. Larson |
CBMI | 2 |
| 2025 | Dual-Objective Adversarial Disentanglement for Protecting Speech Data used for Diagnosing Parkinson's DiseaseabstractRecently, the challenge of protecting privacy-sensitive information in speech data has received growing attention. While many studies have explored the protection of speaker identity, the protection of individual speaker attributes, such as gender, has not been thoroughly investigated. In this paper, we propose a dual-objective approach to adversarial disentanglement that protects the gender attribute of the speaker in speech data used for the diagnosis of Parkinson's disease (PD). The approach combines an adversarial Gradient Reversal Layer (GRL) objective with a utility objective. Experiments on the PC-GITA and NeuroVoz PD speech datasets show that our approach can block the ability of a classifier to infer the gender of the speaker, while preserving the utility for diagnostic purposes. Our work contributes to speech privacy, but also to the understanding of gender for PD diagnosis. Mehtab Ur Rahman, Martha A. Larson, Louis ten Bosch, Cristian Tejedor García |
CBMI | 2 |
| 2025 | Evaluating the Usefulness of Non-Diagnostic Speech Data for Developing Parkinson's Disease ClassifiersabstractContains fulltext : 322329.pdf (Publisher’s version ) (Open Access) Terry Yi Zhong, Esther Janse, Cristian Tejedor García, Louis ten Bosch, Martha A. Larson |
INTERSPEECH | 5 |
| 2025 | Typographic Attacks in a Multi-Image SettingabstractXiaomeng Wang, Zhengyu Zhao, Martha Larson. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zhengyu Zhao 0001, Martha A. Larson |
NAACL (Long Papers) | 3 |
| 2025 | Resisting Bag-Based Attribute Profiling by Adding Adversarial Items to Existing Media ProfilesabstractBag-based classification is a supervised machine learning method that makes a prediction based on a bag of items. Unfortunately, it can be misused as an attribute profiling attack, where the attacker’s objective is to infer a privacy-sensitive attribute of a target user from that user’s shared social media profile, i.e., a bag of images or other media. Despite this threat, existing studies on profiling attacks are limited to the item-level perspective, i.e., attack and defense of a single item. In this work, we move obfuscation defenses against attribute profiling beyond the existing single-item research to study the multi-item, bag-based case, which is more practically relevant because it considers the full attack surface. Defense against bag-based profiling is difficult, because, in general, content shared on social media can never be completely deleted. For this reason, we study defenses that involve extensions, referred to aspivoting additions, to existing profiles, which aim to change (i.e., pivot) the output of the bag-based classifier without removing items contained in the original profile. We propose three different pivoting additions: Adversarial Noise (AdvN), Adversarially Perturbed Items (AdvPI), and Natural Items (NatI). We experimentally demonstrate the ability of these pivoting additions to compromise the performance of three deep bag-based classifiers, representing late-, intermediate- and early-fusion approaches. Overall, our work provides an introduction to the risk of bag-based profiling and a systematic study of defenses. Zhuoran Liu 0001, Zhengyu Zhao 0001, Martha A. Larson |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Mutant Texts: A Technique for Uncovering Unexpected Inconsistencies in Large-Scale Vision-Language Models
Mingliang Liang, Zhouran Liu, Martha A. Larson |
MMM (4) | 3 |
| 2023 | Exploring the Importance of Sign Language Phonology for a Deep Neural NetworkabstractWe conduct an initial investigation to gain insight into whether a deep neural network learns phonological aspects of sign language when classifying video recordings of isolated signs from a continuous signing scenario.We train a series of neural networks to distinguish pairs of signs in Dutch Sign Language, controlling the phonological difference between the signs in each pair.Our results suggest that the intrinsic dimension of the final hidden layer of a network is surprisingly insensitive to the phonological difference between the signs in a pair.However, the ability of the network to discriminate two signs shows a clear trend towards increasing with increasing phonological distinctiveness. Related WorkIsolated sign language recognition.Sign language recognition (SLR) is the problem of recognizing and identifying a particular sign in a video clip.In this paper, we study isolated SLR, also known as word-level SLR, which 1 https://github.com/JavierMartnz/MindTheLinguisticGap Javier Martinez Rodriguez, Martha A. Larson, Louis ten Bosch |
ESANN | 2 |
| 2023 | Beyond Neural-on-Neural Approaches to Speaker Gender ProtectionabstractRecent research has proposed approaches that modify speech to defend against gender inference attacks. The goal of these protection algorithms is to control the availability of information about a speaker’s gender, a privacy-sensitive attribute. Currently, the common practice for developing and testing gender protection algorithms is "neural-on-neural", i.e., perturbations are generated and tested with a neural network. In this paper, we propose to go beyond this practice to strengthen the study of gender protection. First, we demonstrate the importance of testing gender inference attacks that are based on speech features historically developed by speech scientists, alongside the conventionally used neural classifiers. Next, we argue that researchers should use speech features to gain insight into how protective modifications change the speech signal. Finally, we point out that gender-protection algorithms should be compared with novel "vocal adversaries", human-executed voice adaptations, in order to improve interpretability and enable before-the-mic protection. Loes van Bemmel, Zhuoran Liu 0001, Nik Vaessen, Martha A. Larson |
ICASSP | 4 |
| 2023 | Image Shortcut Squeezing: Countering Perturbative Availability Poisons with CompressionabstractPerturbative availability poisoning (PAP) adds small changes to images to prevent their use for model training. Current research adopts the belief that practical and effective approaches to countering such poisons do not exist. In this paper, we argue that it is time to abandon this belief. We present extensive experiments showing that 12 state-of-the-art PAP methods are vulnerable to Image Shortcut Squeezing (ISS), which is based on simple compression. For example, on average, ISS restores the CIFAR-10 model accuracy to 81.73%, surpassing the previous best preprocessing-based countermeasures by 37.97% absolute. ISS also (slightly) outperforms adversarial training and has higher generalizability to unseen perturbation norms and also higher efficiency. Our investigation reveals that the property of PAP perturbations depends on the type of surrogate model used for poison generation, and it explains why a specific ISS compression yields the best performance for a specific type of PAP perturbation. We further test stronger, adaptive poisoning, and show it falls short of being an ideal defense against ISS. Overall, our results demonstrate the importance of considering various (simple) countermeasures to ensure the meaningfulness of analysis carried out during the development of availability poisons. Zhuoran Liu 0001, Zhengyu Zhao 0001, Martha A. Larson |
ICML | 3 |
| 2023 | Exploring Privacy-Preserving Techniques on Synthetic Data as a Defense Against Model Inversion Attacks
Manel Slokom, Peter-Paul de Wolf, Martha A. Larson |
ISC | 3 |
| 2023 | Textual Concept Expansion with Commonsense Knowledge to Improve Dual-Stream Image-Text Matching
Mingliang Liang, Zhuoran Liu 0001, Martha A. Larson |
MMM (1) | 3 |
| 2023 | The Importance of Image Interpretation: Patterns of Semantic Misclassification in Real-World Adversarial Images
Zhengyu Zhao 0001, Nga Dang, Martha A. Larson |
MMM (2) | 3 |
| 2023 | Adversarial Image Color Transformations in Explicit Color Filter SpaceabstractDeep Neural Networks have been shown to be vulnerable to adversarial images. Conventional attacks strive for indistinguishable adversarial images with strictly restricted perturbations. Recently, researchers have moved to explore distinguishable yet non-suspicious adversarial images and demonstrated that color transformation attacks are effective. In this work, we propose Adversarial Color Filter (AdvCF), a novel color transformation attack that is optimized with gradient information in the parameter space of a simple color filter. In particular, our color filter space is explicitly specified so that we are able to provide a systematic analysis of model robustness against adversarial color transformations, from both the attack and defense perspectives. In contrast, existing color transformation attacks do not offer the opportunity for systematic analysis due to the lack of such an explicit space. We further demonstrate the effectiveness of our AdvCF in fooling image classifiers and also compare it with other color transformation attacks regarding their robustness to defenses and image acceptability through an extensive user study. We also highlight the human-interpretability of AdvCF and show its superiority over the state-of-the-art human-interpretable color transformation attack on both image acceptability and efficiency. Additional results provide interesting new insights into model robustness against AdvCF in another three visual tasks. Zhengyu Zhao 0001, Zhuoran Liu 0001, Martha A. Larson |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Color the Word: Leveraging Web Images for Machine Translation of Untranslatable Words
Yana van de Sande, Martha A. Larson |
MMM (1) | 2 |
| 2022 | NewsImages: addressing the depiction gap with an online news dataset for text-image rematchingabstractWe present NewsImages, a dataset of online news items, and the related NewsImages rematching task. The goal of NewsImages is to provide researchers with a means of studying the depiction gap, which we define to be the difference between what an image literally depicts and the way in which it is connected to the text that it accompanies. Online news is a domain in which the image-text connection is known to be indirect: The news article does not describe what is literally depicted in the image. We validate NewsImages with experiments that show the dataset's and the task's use for studying occurring connections between image and text, as well as addressing the depiction gap, which include sparse data, diversity of content, and importance of background knowledge. Andreas Lommatzsch, Benjamin Kille, Özlem Özgöbek, Yuxiao Zhou 0002, Jelena Tesic, Cláudio Bartolomeu, David Semedo, Lidia Pivovarova, Mingliang Liang, Martha A. Larson |
MMSys | 10 |
| 2022 | When Machine Learning Models Leak: An Exploration of Synthetic Training Data
Manel Slokom, Peter-Paul de Wolf, Martha A. Larson |
PSD | 3 |
| 2021 | Social Signals and Multimedia: Past, Present, FutureabstractThe rising popularity of Artificial Intelligence (AI) has brought considerable public interest as well faster and more direct transfer of research ideas into practice. One of the aspects of AI that still trails behind considerably is the role of machines in interpreting, enhancing, modeling, generating, and influencing social behavior. Such behavior is captured as social signals, usually by sensors recording multiple modalities, making it classic multimedia data. Such behavior can also be generated by an AI system when interacting with humans. Using AI techniques in combination with multimedia data can be used to pursue multiple goals, two of which are high-lighted here. First, supporting people during social interactions and helping them to fulfil their social needs either actively or passively.Second, improving our understanding of how people collaborate, build relationships, and process self identity. Despite the rise of fields such as Social Signal Processing, a similar panel organised at ACM Multimedia 2014, and an area on social and emotional signal sat the ACM MM since 2014, we argue that we have yet to truly fulfil the potential of the combining social signals and multimedia. This panel asks where we have come far enough and what remaining challenges there are in light of recent global events. Hayley Hung, Cathal Gurrin, Martha A. Larson, Hatice Gunes, Fabien Ringeval, Elisabeth André, Louis-Philippe Morency |
ACM Multimedia | 3 |
| 2021 | Screen Gleaning: A Screen Reading TEMPEST Attack on Mobile Devices Exploiting an Electromagnetic Side Channel
Zhuoran Liu 0001, Niels Samwel, Leo Weissbart, Zhengyu Zhao 0001, Dirk Lauret, Lejla Batina, Martha A. Larson |
NDSS | 7 |
| 2021 | On Success and Simplicity: A Second Look at Transferable Targeted AttacksabstractAchieving transferability of targeted attacks is reputed to be remarkably difficult. The current state of the art has resorted to resource-intensive solutions that necessitate training model(s) for each target class with additional data. In our investigation, we find, however, that simple transferable attacks which require neither model training nor additional data can achieve surprisingly strong targeted transferability. This insight has been overlooked until now, mainly because the widespread practice of attacking with only few iterations has largely limited the attack convergence to optimal targeted transferability. In particular, we, for the first time, identify that a very simple logit loss can largely surpass the commonly adopted cross-entropy loss, and yield even better results than the resource-intensive state of the art. Our analysis spans a variety of transfer scenarios, especially including three new, realistic scenarios: an ensemble transfer scenario with little model similarity, a worse-case scenario with low-ranked target classes, and also a real-world attack on the Google Cloud Vision API. Results in these new transfer scenarios demonstrate that the commonly adopted, easy scenarios cannot fully reveal the actual strength of different attacks and may cause misleading comparative results. We also show the usefulness of the simple logit loss for generating targeted universal adversarial perturbations in a data-free manner. Overall, the aim of our analysis is to inspire a more meaningful evaluation on targeted transferability. Code is available at https://github.com/ZhengyuZhao/Targeted-Tansfer. Zhengyu Zhao 0001, Zhuoran Liu 0001, Martha A. Larson |
NeurIPS | 3 |
| 2021 | Pivoting Image-based Profiles Toward Privacy: Inhibiting Malicious Profiling with Adversarial AdditionsabstractUsers build up profiles online consisting of items that they have shared or interacted with. In this work, we look at profiles that consist of images. We address the issue of privacy-sensitive information being automatically inferred from these user profiles, against users’ will and best interest. We introduce the concept of a privacy pivot, which is a strategic change that users can make in their sharing that will inhibit malicious profiling. Importantly, the pivot helps put privacy control into the hands of the users. Further, it does not require users to delete any of the existing images in their profiles, nor does it require a radical change in their sharing intentions, i.e., what they would like to communicate with their profile. Previous work has investigated adversarial images for privacy protection, but has focused on individual images. Here, we move further to study image sets comprising image profiles. We define a conceptual formulation of the challenge of the privacy pivot in the form of an “Anti-Profiling Model”. Within this model, we propose a basic pivot solution that uses adversarial additions to effectively inhibit the predictions of profilers using set-based image classification. Zhuoran Liu 0001, Zhengyu Zhao 0001, Martha A. Larson |
UMAP | 3 |
| 2021 | Adversarial Item Promotion: Vulnerabilities at the Core of Top-N Recommenders that Use Images to Address Cold StartabstractE-commerce platforms provide their customers with ranked lists of recommended items matching the customers’ preferences. Merchants on e-commerce platforms would like their items to appear as high as possible in the top-N of these ranked lists. In this paper, we demonstrate how unscrupulous merchants can create item images that artificially promote their products, improving their rankings. Recommender systems that use images to address the cold start problem are vulnerable to this security risk. We describe a new type of attack, Adversarial Item Promotion (AIP), that strikes directly at the core of Top-N recommenders: the ranking mechanism itself. Existing work on adversarial images in recommender systems investigates the implications of conventional attacks, which target deep learning classifiers. In contrast, our AIP attacks are embedding attacks that seek to push features representations in a way that fools the ranker (not a classifier) and directly leads to item promotion. We introduce three AIP attacks insider attack, expert attack, and semantic attack, which are defined with respect to three successively more realistic attack models. Our experiments evaluate the danger of these attacks when mounted against three representative visually-aware recommender algorithms in a framework that uses images to address cold start. We also evaluate potential defenses, including adversarial training and find that common, currently-existing, techniques do not eliminate the danger of AIP attacks. In sum, we show that using images to address cold start opens recommender systems to potential threats with clear practical implications. Zhuoran Liu 0001, Martha A. Larson |
WWW | 2 |
| 2021 | Towards user-oriented privacy for recommender system data: A personalization-based approach to gender obfuscation for user profilesabstractIn this paper, we propose a new privacy solution for the data used to train a recommender system, i.e., the user–item matrix. The user–item matrix contains implicit information, which can be inferred using a classifier, leading to potential privacy violations. Our solution, called Personalized Blurring (PerBlur), is a simple, yet effective, approach to adding and removing items from users’ profiles in order to generate an obfuscated user–item matrix. The novelty of PerBlur is personalization of the choice of items used for obfuscation to the individual user profiles. PerBlur is formulated within a user-oriented paradigm of recommender system data privacy that aims at making privacy solutions understandable, unobtrusive, and useful for the user. When obfuscated data is used for training, a recommender system algorithm is able to reach performance comparable to what is attained when it is trained on the original, unobfuscated data. At the same time, a classifier can no longer reliably use the obfuscated data to predict the gender of users, indicating that implicit gender information has been removed. In addition to introducing PerBlur, we make several key contributions. First, we propose an evaluation protocol that creates a fair environment to compare between different obfuscation conditions. Second, we carry out experiments that show that gender obfuscation impacts the fairness and diversity of recommender system results. In sum, our work establishes that a simple, transparent approach to gender obfuscation can protect user privacy while at the same time improving recommendation results for users by maintaining fairness and enhancing diversity. Manel Slokom, Alan Hanjalic, Martha A. Larson |
Inf. Process. Manag. | 3 |
| 2020 | Adversarial Color Enhancement: Generating Unrestricted Adversarial Images by Optimizing a Color Filter
Zhengyu Zhao 0001, Zhuoran Liu 0001, Martha A. Larson |
BMVC | 3 |
| 2020 | Towards Large Yet Imperceptible Adversarial Image Perturbations With Perceptual Color DistanceabstractThe success of image perturbations that are designed to fool image classifier is assessed in terms of both adversarial effect and visual imperceptibility. The conventional assumption on imperceptibility is that perturbations should strive for tight Lp-norm bounds in RGB space. In this work, we drop this assumption by pursuing an approach that exploits human color perception, and more specifically, minimizing perturbation size with respect to perceptual color distance. Our first approach, Perceptual Color distance C&W (PerC-C&W), extends the widely-used C&W approach and produces larger RGB perturbations. PerC-C&W is able to maintain adversarial strength, while contributing to imperceptibility. Our second approach, Perceptual Color distance Alternating Loss (PerC-AL), achieves the same outcome, but does so more efficiently by alternating between the classification loss and perceptual color difference when updating perturbations. Experimental evaluation shows PerC approaches outperform conventional Lp approaches in terms of robustness and transferability, and also demonstrates that the PerC distance can provide added value on top of existing structure-based methods to creating image perturbations. Zhengyu Zhao 0001, Zhuoran Liu 0001, Martha A. Larson |
CVPR | 3 |
| 2020 | The Connection between the Text and Images of News Articles: New Insights for Multimedia AnalysisabstractWe report on a case study of text and images that reveals the inadequacy of simplistic assumptions about their connection and interplay. The context of our work is a larger effort to create automatic systems that can extract event information from online news articles about flooding disasters. We carry out a manual analysis of 1000 articles containing a keyword related to flooding. The analysis reveals that the articles in our data set cluster into seven categories related to different topical aspects of flooding, and that the images accompanying the articles cluster into five categories related to the content they depict. The results demonstrate that flood-related news articles do not consistently report on a single, currently unfolding flooding event and we should also not assume that a flood-related image will directly relate to a flooding-event described in the corresponding article. In particular, spatiotemporal distance is important. We validate the manual analysis with an automatic classifier demonstrating the technical feasibility of multimedia analysis approaches that admit more realistic relationships between text and images. In sum, our case study confirms that closer attention to the connection between text and images has the potential to improve the collection of multimodal information from news articles. Nelleke Oostdijk, Hans van Halteren, Mustafa Erkan Basar, Martha A. Larson |
LREC | 4 |
| 2020 | The World has Changed - The World Needs to Change. What Multimedia has to Offer for Our Common Digital FutureabstractNot only the current coronavirus is holding the world in breath. Beyond this current health crisis the world is facing several global challenges from climate change and environmental damage, access to clean water and food, socio-economic inequalities to name a few. The United have very well framed these global challenges in their 17 Sustainability Goals for a future in prosperity and equal opportunities for all, to be achieved by 2030. There is no one simple solution, no one easy cure in sight to address these pressing challenges of our days. Rather a collective approach of all of us is needed which in sum will be contributing to these. Obviously, the field of multimedia has contributed to many tools and applications that are so much in demand these days to stay connected while keeping the distance. But there is much more we can offer to our common digital future. Our future health system, global access to education, decent work, and reducing inequalities are just some of these goals where we our field can contribute. In this panel we will discuss which path we could follow. Susanne Boll, Hari Sundaram, Svetha Venkatesh, Martha A. Larson, Mohan Kankanhalli |
ACM Multimedia | 4 |
| 2019 | Who's Afraid of Adversarial Queries?: The Impact of Image Modifications on Content-based Image RetrievalabstractAn adversarial query is an image that has been modified to disrupt content-based image retrieval (CBIR), while appearing nearly untouched to the human eye. This paper presents an analysis of adversarial queries for CBIR based on neural, local, and global features. We introduce an innovative neural image perturbation approach, called Perturbations for Image Retrieval Error (PIRE), that is capable of blocking neural-feature-based CBIR. PIRE differs significantly from existing approaches that create images adversarial with respect to CNN classifiers because it is unsupervised, i.e., it needs no labeled data from the data set to which it is applied. Our experimental analysis demonstrates the surprising effectiveness of PIRE in blocking CBIR, and also covers aspects of PIRE that must be taken into account in practical settings, including saving images, image quality and leaking adversarial queries into the background collection. Our experiments also compare PIRE (a neural approach) with existing keypoint removal and injection approaches (which modify local features). Finally, we discuss the challenges that face multimedia researchers in the future study of adversarial queries. Zhuoran Liu 0001, Zhengyu Zhao 0001, Martha A. Larson |
ICMR | 3 |
| 2019 | Reproducible Experiments on Adaptive Discriminative Region Discovery for Scene RecognitionabstractThis companion paper supports the replication of scene image recognition experiments using Adaptive Discriminative Region Discovery (Adi-Red), an approach presented at ACM Multimedia 2018. We provide a set of artifacts that allow the replication of the experiments using a Python implementation. All the experiments are covered in a single shell script, which requires the installation of an environment, following our instructions, or using ReproZip.The data sets (images and labels) are automatically downloaded, and the train-test splits used in the experiments are created. The first experiment is from the original paper, and the second supports exploration of the resolution of the scale-specific input image, an interesting additional parameter. For both experiments, five other parameters can be adjusted: the threshold used to select the number of discriminative patches, the number of scales used, the type of patch selection (Adi-Red, dense or random), the architecture and pre-training data set of the pre-trained CNN feature extractor. The final output includes four tables (original Table 1, Table 2 and Table 4, and a table for the resolution experiment) and two plots (original Figure 3 and Figure 4). Zhengyu Zhao 0001, Zhuoran Liu 0001, Martha A. Larson, Ahmet Iscen, Naoko Nitta |
ACM Multimedia | 3 |
| 2019 | The Representation of Speech in Deep Neural Networks
Odette Scharenborg, Nikki van der Gouw, Martha A. Larson, Elena Marchiori |
MMM (2) | 3 |
| 2019 | Remembering winter was coming - Character-oriented video summaries of TV series
Xavier Bost, Serigne Gueye, Vincent Labatut, Martha A. Larson, Georges Linarès, Damien Malinas, Raphaël Roth |
Multim. Tools Appl. | 4 |
| 2019 | Top-N Recommendation with Multi-Channel Positive Feedback using Factorization MachinesabstractUser interactions can be considered to constitute different feedback channels, for example, view, click, like or follow, that provide implicit information on users’ preferences. Each implicit feedback channel typically carries a unary, positive-only signal that can be exploited by collaborative filtering models to generate lists of personalized recommendations. This article investigates how a learning-to-rank recommender system can best take advantage of implicit feedback signals from multiple channels. We focus on Factorization Machines (FMs) with Bayesian Personalized Ranking (BPR), a pairwise learning-to-rank method, that allows us to experiment with different forms of exploitation. We perform extensive experiments on three datasets with multiple types of feedback to arrive at a series of insights. We compare conventional, direct integration of feedback types with our proposed method, which exploits multiple feedback channels during the sampling process of training. We refer to our method as multi-channel sampling. Our results show that multi-channel sampling outperforms conventional integration, and that sampling with the relative “level” of feedback is always superior to a level-blind sampling approach. We evaluate our method experimentally on three datasets in different domains and observe that with our multi-channel sampler the accuracy of recommendations can be improved considerably compared to the state-of-the-art models. Further experiments reveal that the appropriate sampling method depends on particular properties of datasets such as popularity skewness. Babak Loni, Roberto Pagano, Martha A. Larson, Alan Hanjalic |
ACM Trans. Inf. Syst. | 3 |
| 2018 | The Conversation Continues: the Effect of Lyrics and Music Complexity of Background Music on Spoken-Word RecognitionabstractBackground music in social interaction settings can hinder conversation.Yet, little is known of how specific properties of music impact speech processing.This paper addresses this knowledge gap by investigating the effect of the 1) complexity of the background music, and 2) the presence versus absence of sung lyrics on spoken-word recognition in background music.To answer these questions, a word identification experiment was run in which Dutch participants listened to Dutch CVC words embedded in stretches of background music in four conditions: low/high complexity and with lyrics/music-only, and at three SNRs.Music stretches with and without lyrics were sampled from the same song in order to control for factors beyond the complexity of the music and the presence of lyrics.The results showed a clear negative impact of more complex music and the presence of lyrics in background music on spoken-word recognition.The results open a path for future work, and suggest that social spaces (e.g., restaurants, cafés and bars) should make careful choices of music to promote conversation. Odette Scharenborg, Martha A. Larson |
INTERSPEECH | 2 |
| 2018 | From Volcano to Toyshop: Adaptive Discriminative Region Discovery for Scene RecognitionabstractAs deep learning approaches to scene recognition emerge, they have continued to leverage discriminative regions at multiple scales, building on practices established by conventional image classification research. However, approaches remain largely generic, and do not carefully consider the special properties of scenes. In this paper, inspired by the intuitive differences between scenes and objects, we propose Adi-Red, an adaptive approach to discriminative region discovery for scene recognition. Adi-Red uses a CNN classifier, which was pre-trained using only image-level scene labels, to discover discriminative image regions directly. These regions are then used as a source of features to perform scene recognition. The use of the CNN classifier makes it possible to adapt the number of discriminative regions per image using a simple, yet elegant, threshold, at relatively low computational cost. Experimental results on the scene recognition benchmark dataset SUN397 demonstrate the ability of Adi-Red to outperform the state of the art. Additional experimental analysis on the Places dataset reveals the advantages of Adi-Red, and highlight how they are specific to scenes. We attribute the effectiveness of Adi-Red to the ability of adaptive region discovery to avoid introducing noise, while also not missing out on important information. Zhengyu Zhao 0001, Martha A. Larson |
ACM Multimedia | 2 |
| 2018 | Geo-Distinctive Visual Element Matching for Location Estimation of ImagesabstractWe propose an image representation and matching approach that substantially improves visual-based location estimation for images. The main novelty of the approach, called distinctive visual element matching (DVEM), is its use of representations that are specific to the query image whose location is being predicted. These representations are based on visual element clouds, which robustly capture the connection between the query and visual evidence from candidate locations. We then maximize the influence of visual elements that are geo-distinctive because they do not occur in images taken at many other locations. We carry out experiments and analysis for both geo-constrained and geo-unconstrained location estimation cases using two large-scale, publicly available datasets: the San Francisco Landmark dataset with 1.06 million street-view images and the MediaEval'15 Placing Task dataset with 5.6 million geo-tagged images from Flickr. We present examples that illustrate the highly transparent mechanics of the approach, which are based on commonsense observations about the visual patterns in image collections. Our results show that the proposed method delivers a considerable performance improvement compared to the state-of-the-art. Xinchao Li, Martha A. Larson, Alan Hanjalic |
IEEE Trans. Multim. | 2 |
| 2018 | Detecting Socially Significant Music Events Using Temporally Noisy LabelsabstractIn this paper, we focus on event detection over the timeline of a music track. Such technology is motivated by the need for innovative applications such as searching, nonlinear access, and recommendation. Event detection over the timeline requires time-code level labels in order to train machine learning models. We use timed comments from SoundCloud, a modern social music sharing platform, to obtain these labels. While in this way the need for tedious and time-consuming manual labeling can be reduced, the challenge is that timed comments are subject to additional temporal noise, as they occur in the temporal neighborhood of the actual events. We investigate the utility of such noisy timed comments as training labels through a case study, in which we investigate three types of events in electronic dance music (EDM): drop, build, and break. These socially significant events play a key role in an EDM track's unfolding and are popular in social media circles. These events are interesting for detection, and here we leverage the timed comments generated in the course of the online social activity around them. We propose a two-stage learning method that relies on noisy timed comments and, given a music track, marks the events on the timeline. In the experiments, we focus, in particular, on investigating to which extent noisy timed comments can replace manually acquired expert labels. The conclusions we draw during this study provide useful insights that motivate further research in the field of event detection. Karthik Yadati, Martha A. Larson, Cynthia C. S. Liem, Alan Hanjalic |
IEEE Trans. Multim. | 2 |
| 2017 | The Geo-Privacy Bonus of Popular Photo EnhancementsabstractToday's geo-location estimation approaches are able to infer the location of a target image using its visual content alone. These approaches typically exploit visual matching techniques, applied to a large collection of background images with known geo-locations. Users who are unaware that visual analysis and retrieval approaches can compromise their geo-privacy, unwittingly open themselves to risks of crime or other unintended consequences. This paper lays the groundwork for a new approach to geo-privacy of social images: Instead of requiring a change of user behavior, we start by investigating users' existing photo-sharing practices. We carry out a series of experiments using a large collection of social images (8.5M) to systematically analyze how photo editing practices impact the performance of geo-location estimation. We find that standard image enhancements, including filters and cropping, already serve as natural geo-privacy protectors. In our experiments, up to 19% of images whose location would otherwise be automatically predictable were unlocalizeable after enhancement. We conclude that it would be wrong to assume that geo-visual privacy is a lost cause in today's world of rapidly maturing machine learning. Instead, protecting users against the unwanted effects of pixel-based inference is a viable research field. A starting point is understanding the geo-privacy bonus of already established user behavior. Jaeyoung Choi 0002, Martha A. Larson, Xinchao Li, Gerald Friedland, Alan Hanjalic |
ICMR | 2 |
| 2017 | On the Automatic Identification of Music for Common ActivitiesabstractIn this paper, we address the challenge of identifying music suitable to accompany typical daily activities. We first derive a list of common activities by analyzing social media data. Then, an automatic approach is proposed to find music for these activities. Our approach is inspired by our experimentally acquired findings (a) that genre and instrument information, i.e., as appearing in the textual metadata, are not sufficient to distinguish music appropriate for different types of activities, and (b) that existing content-based approaches in the music information retrieval community do not overcome this insufficiency. The main contributions of our work are (a) our analysis of the properties of activity-related music that inspire our use of novel high-level features, e.g., drop-like events, and (b) our approach's novel method of extracting and combining low-level features, and, in particular, the joint optimization of the time window for feature aggregation and the number of features to be used. The effectiveness of the approach method is demonstrated in a comprehensive experimental study including failure analysis. Karthik Yadati, Cynthia C. S. Liem, Martha A. Larson, Alan Hanjalic |
ICMR | 3 |
| 2017 | Multimodal Video-to-Video Linking: Turning to the Crowd for Insight and Evaluation
Maria Eskevich, Martha A. Larson, Robin Aly, Serwah Sabetghadam, Gareth J. F. Jones, Roeland Ordelman, Benoit Huet |
MMM (2) | 2 |
| 2017 | CitRec 2017: International Workshop on Recommender Systems for CitizensabstractThe "International Workshop on Recommender Systems for Citizens" (CitRec) is focused on a novel type of recommender systems both in terms of ownership and purpose: recommender systems run by citizens and serving society as a whole. Jie Yang 0028, Zhu Sun 0001, Alessandro Bozzon, Jie Zhang 0002, Martha A. Larson |
RecSys | 5 |
| 2017 | A Stream-based Resource for Multi-Dimensional Evaluation of Recommender AlgorithmsabstractRecommender System research has evolved to focus on developing algorithms capable of high performance in online systems. This development calls for a new evaluation infrastructure that supports multi-dimensional evaluation of recommender systems. Today's researchers should analyze algorithms with respect to a variety of aspects including predictive performance and scalability. Researchers need to subject algorithms to realistic conditions in online A/B tests. We introduce two resources supporting such evaluation methodologies: the new data set of stream recommendation interactions released for CLEF NewsREEL 2017, and the new Open Recommendation Platform (ORP). The data set allows researchers to study a stream recommendation problem closely by "replaying" it locally, and ORP makes it possible to take this evaluation "live" in a living lab scenario. Specifically, ORP allows researchers to deploy their algorithms in a live stream to carry out A/B tests. To our knowledge, NewsREEL is the first online news recommender system resource to be put at the disposal of the research community. In order to encourage others to develop comparable resources for a wide range of domains, we present a list of practical lessons learned in the development of the dataset and ORP. Benjamin Kille, Andreas Lommatzsch, Frank Hopfgartner, Martha A. Larson, Arjen P. de Vries |
SIGIR | 4 |
| 2016 | ChaLearn Joint Contest on Multimedia Challenges Beyond Visual Analysis: An overviewabstractThis paper provides an overview of the Joint Contest on Multimedia Challenges Beyond Visual Analysis. We organized an academic competition that focused on four problems that require effective processing of multimodal information in order to be solved. Two tracks were devoted to gesture spotting and recognition from RGB-D video, two fundamental problems for human computer interaction. Another track was devoted to a second round of the first impressions challenge of which the goal was to develop methods to recognize personality traits from short video clips. For this second round we adopted a novel collaborative-competitive (i.e., coopetition) setting. The fourth track was dedicated to the problem of video recommendation for improving user experience. The challenge was open for about 45 days, and received outstanding participation: almost 200 participants registered to the contest, and 20 teams sent predictions in the final stage. The main goals of the challenge were fulfilled: the state of the art was advanced considerably in the four tracks, with novel solutions to the proposed problems (mostly relying on deep learning). However, further research is still required. The data of the four tracks will be available to allow researchers to keep making progress in the four tracks. Hugo Jair Escalante, Víctor Ponce-López, Jun Wan 0001, Michael Riegler 0001, Albert Clapés, Sergio Escalera, Isabelle Guyon, Xavier Baró, Pål Halvorsen, Henning Müller, Martha A. Larson |
ICPR | 12 |
| 2016 | Experiences with Shared Resources for Research and Education in Speech and Language Processingabstract\n Contains fulltext :\n 161889.pdf (Publisher’s version ) (Open Access)\n Rebecca Bates 0001, Eric Fosler-Lussier, Florian Metze, Martha A. Larson, Gina-Anne Levow, Emily Mower Provost |
INTERSPEECH | 4 |
| 2016 | Right inflight?: a dataset for exploring the automatic prediction of movies suitable for a watching situationabstractIn this paper, we present the dataset Right Inflight developed to support the exploration of the match between video content and the situation in which that content is watched. Specifically, we look at videos that are suitable to be watched on an airplane, where the main assumption is that that viewers watch movies with the intent of relaxing themselves and letting time pass quickly, despite the inconvenience and discomfort of flight. The aim of the dataset is to support the development of recommender systems, as well as computer vision and multimedia retrieval algorithms capable of automatically predicting which videos are suitable for inflight consumption. Our ultimate goal is to promote a deeper understanding of how people experience video content, and of how technology can support people in finding or selecting video content that supports them in regulating their internal states in certain situations. Right Inflight consists of 318 human-annotated movies, for which we provide links to trailers, a set of pre-computed low-level visual, audio and text features as well as user ratings. The annotation was performed by crowdsourcing workers, who were asked to judge the appropriateness of movies for inflight consumption. Michael Riegler 0001, Martha A. Larson, Concetto Spampinato, Pål Halvorsen, Mathias Lux, Jonas Markussen, Konstantin Pogorelov, Carsten Griwodz, Håkon Kvale Stensland |
MMSys | 2 |
| 2016 | RecSys Challenge 2016: Job RecommendationsabstractThe 2016 ACM Recommender Systems Challenge focused on the problem of job recommendations. Given a large dataset from XING that consisted of anonymized user profiles, job postings, and interactions between them, the participating teams had to predict postings that a user will interact with. The challenge ran for four months with 366 registered teams. 119 of those teams actively participated and submitted together 4,232 solutions yielding in an impressive neck-and-neck race that was decided within the last days of the challenge. Fabian Abel, András A. Benczúr, Daniel Kohlsdorf, Martha A. Larson, Róbert Pálovics |
RecSys | 4 |
| 2016 | Bayesian Personalized Ranking with Multi-Channel User FeedbackabstractPairwise learning-to-rank algorithms have been shown to allow recommender systems to leverage unary user feedback. We propose Multi-feedback Bayesian Personalized Ranking (MF-BPR), a pairwise method that exploits different types of feedback with an extended sampling method. The feedback types are drawn from different "channels", in which users interact with items (e.g., clicks, likes, listens, follows, and purchases). We build on the insight that different kinds of feedback, e.g., a click versus a like, reflect different levels of commitment or preference. Our approach differs from previous work in that it exploits multiple sources of feedback simultaneously during the training process. The novelty of MF-BPR is an extended sampling method that equates feedback sources with "levels" that reflect the expected contribution of the signal. We demonstrate the effectiveness of our approach with a series of experiments carried out on three datasets containing multiple types of feedback. Our experimental results demonstrate that with a right sampling method, MF-BPR outperforms BPR in terms of accuracy. We find that the advantage of MF-BPR lies in its ability to leverage level information when sampling negative items. Babak Loni, Roberto Pagano, Martha A. Larson, Alan Hanjalic |
RecSys | 3 |
| 2016 | Algorithms Aside: Recommendation As The Lens Of LifeabstractIn this position paper, we take the experimental approach of putting algorithms aside, and reflect on what recommenders would be for people if they were not tied to technology. By looking at some of the shortcomings that current recommenders have fallen into and discussing their limitations from a human point of view, we ask the question: if freed from all limitations, what should, and what could, RecSys be? We then turn to the idea that life itself is the best recommender system, and that people themselves are the query. By looking at how life brings people in contact with options that suit their needs or match their preferences, we hope to shed further light on what current RecSys could be doing better. Finally, we look at the forms that RecSys could take in the future. By formulating our vision beyond the reach of usual considerations and current limitations, including business models, algorithms, data sets, and evaluation methodologies, we attempt to arrive at fresh conclusions that may inspire the next steps taken by the community of researchers working on RecSys. Tamas Motajcsek, Jean-Yves Le Moine, Martha A. Larson, Daniel Kohlsdorf, Andreas Lommatzsch, Domonkos Tikk, Omar Alonso, Paolo Cremonesi, Andrew M. Demetriou, Kristaps Dobrajs, Franca Garzotto, Ayse Göker, Frank Hopfgartner, Davide Malagoli, Thuy Ngoc Nguyen 0001, Jasminko Novak, Francesco Ricci 0001, Mario Scriminaci, Marko Tkalcic, Anna Zacchi |
RecSys | 3 |
| 2016 | The Contextual Turn: from Context-Aware to Context-Driven Recommender SystemsabstractA critical change has occurred in the status of context in recommender systems. In the past, context has been considered 'additional evidence'. This past picture is at odds with many present application domains, where user and item information is scarce. Such domains face continuous cold start conditions and must exploit session rather than user information. In this paper, we describe the `Contextual Turn?: the move towards context-driven recommendation algorithms for which context is critical, rather than additional. We cover application domains, algorithms that promise to address the challenges of context-driven recommendation, and the steps that the community has taken to tackle context-driven problems. Our goal is to point out the commonalities of context-driven problems, and urge the community to address the overarching challenges that context-driven recommendation poses. Roberto Pagano, Paolo Cremonesi, Martha A. Larson, Balázs Hidasi, Domonkos Tikk, Alexandros Karatzoglou, Massimo Quadrana |
RecSys | 3 |
| 2016 | Introduction to the Special Issue on Crowd in Intelligent SystemsabstractNo abstract available. Kuan-Ta Chen, Omar Alonso, Martha A. Larson, Irwin King |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2015 | Pairwise geometric matching for large-scale object retrievalabstractSpatial verification is a key step in boosting the performance of object-based image retrieval. It serves to eliminate unreliable correspondences between salient points in a given pair of images, and is typically performed by analyzing the consistency of spatial transformations between the image regions involved in individual correspondences. In this paper, we consider the pairwise geometric relations between correspondences and propose a strategy to incorporate these relations at significantly reduced computational cost, which makes it suitable for large-scale object retrieval. In addition, we combine the information on geometric relations from both the individual correspondences and pairs of correspondences to further improve the verification accuracy. Experimental results on three reference datasets show that the proposed approach results in a substantial performance improvement compared to the existing methods, without making concessions regarding computational efficiency. Xinchao Li, Martha A. Larson, Alan Hanjalic |
CVPR | 2 |
| 2015 | Evento 360: Social Event Discovery from Web-scale Multimedia CollectionabstractWe present Evento 360 (URL: http://evento360.info), an online interactive social event browser, which allows the user to explore events detected within a web-scale multimedia corpus. The system addresses five key aspects of social multimedia event detection and summarization: multimodality, scale, diversity of representations, noise of multimedia items, and missing metadata. The detection algorithm uses unsupervised clustering approach that exploits temporal, spatial and textual metadata. For each detected event cluster, to choose the best subset of photos that meet both relevance and diversity criteria, the system uses hierarchical clustering that exploits both visual and audio information. Evento 360's user interface provides a search feature that is not limited to a certain set of events, but rather can handle an arbitrary event query. It allows the user to retrieve and explore relevant events. The system scales well and is effective in producing high-quality summaries of the detected events. Jaeyoung Choi 0002, Eungchan Kim, Martha A. Larson, Gerald Friedland, Alan Hanjalic |
ACM Multimedia | 3 |
| 2015 | Overview of the 2015 Workshop on Speech, Language and Audio in MultimediaabstractThe Workshop on Speech, Language and Audio in Multimedia (SLAM) positions itself at at the crossroad of multiple scientific fields (music and audio processing, speech processing, natural language processing and multimedia) to discuss and stimulate research results, projects, datasets and benchmarks initiatives where audio, speech and language are applied to multimedia data. While the first two editions were collocated with major speech events, SLAM'15 is deeply rooted in the multimedia community, opening up to computer vision and multimodal fusion. To this end, the workshop emphasizes video hyperlinking as an showcase where computer vision meets speech and language. Such techniques provide a powerful illustration of how multimedia technologies incorporating speech, language and audio can make multimedia content collections better accessible, and thereby more useful, to users. Guillaume Gravier, Gareth J. F. Jones, Martha A. Larson, Roeland Ordelman |
ACM Multimedia | 3 |
| 2015 | Overview of ACM RecSys CrowdRec 2015 Workshop: Crowdsourcing and Human Computation for Recommender Systems
Martha A. Larson, Domonkos Tikk, Roberto Turrin |
RecSys | 1 |
| 2015 | Recurrent neural network language model adaptation with curriculum learning
Yangyang Shi, Martha A. Larson, Catholijn M. Jonker |
Comput. Speech Lang. | 2 |
| 2015 | Integrating meta-information into recurrent neural network language models
Yangyang Shi, Martha A. Larson, Joris Pelemans, Catholijn M. Jonker, Patrick Wambacq, Pascal Wiggers, Kris Demuynck |
Speech Commun. | 2 |
| 2015 | Uploader Intent for Online Video: Typology, Inference, and ApplicationsabstractWe investigate automatic inference of uploader intent for online video, i.e., prediction of the reason for which a user has uploaded a particular video to the Internet. Users upload video for specific reasons, but rarely state these reasons explicitly in the video metadata. Information about the reasons motivating uploaders has the potential ultimately to benefit a wide range of application areas, including video production, video-based advertising , and video search. In this paper, we apply a combination of social-Web mining and crowdsourcing to arrive at a typology that characterizes the uploader intent of a broad range of videos. We then use a set of multimodal features, including visual semantic features, found to be indicative of uploader intent in order to classify videos automatically into uploader intent classes. We evaluate our approach on a dataset containing ca. 3K crowdsourcing-annotated videos and demonstrate its usefulness in prediction tasks relevant to common application areas. Christoph Kofler, Subhabrata Bhattacharya, Martha A. Larson, Tao Chen 0015, Alan Hanjalic, Shih-Fu Chang |
IEEE Trans. Multim. | 3 |
| 2015 | Global-Scale Location Prediction for Social Images Using Geo-Visual RankingabstractWe propose an automatic method that addresses the challenge of predicting the geo-location of social images using only the visual content of those images. Our method is able to generate a geo-location prediction for an image globally . In this respect, it contrasts with other existing approaches, specifically with those that generate predictions restricted to specific cities, landmarks, or an otherwise pre-defined set of locations. The essence and the main novelty of our ranking-based method is that for a given query image a geo-location is recommended based on the evidence collected from images that are not only geographically close to this geo-location, but also have sufficient visual similarity to the query image within the considered image collection. Our method is evaluated experimentally on a public dataset of 8.8 million geo-tagged images from Flickr, released by the MediaEval 2013 evaluation benchmark. Experiments show that the proposed method delivers a substantial performance improvement compared to the existing related approaches, particularly for queries with high numbers of neighbors . In addition, a detailed analysis of the method's performance reveals the impact of different visual feature extraction and image matching strategies, as well as the densities and types of images found at different locations, on the prediction accuracy. Xinchao Li, Martha A. Larson, Alan Hanjalic |
IEEE Trans. Multim. | 2 |
| 2015 | Exploiting the Deep-Link Commentsphere to Support Non-Linear Video AccessabstractIn this paper, we investigate the usefulness of deep links for improving video search results. Deep links are time-coded comments with which viewers express their reactions to the content at specific time-points of a video that they find noteworthy. The rationale underlying our work is that deep links can open up an interesting new perspective on the relevance of a video, namely focusing on individual video segments, in addition to the existing ones that typically concern a video as a whole. In this perspective, deep-link comments provide non-linear access to videos via their time-codes, which can match alternate dimensions of user needs that extend beyond topical and affective relevance. We explore the different types of deep-link comments and develop a viewer expressive reaction variety (VERV) typology that captures how viewers deep-link on YouTube. We validate this typology through a user study on Amazon Mechanical Turk to show that it is a typology human annotators can agree upon. We then demonstrate, through experiments, that deep-link comments can automatically be classified into VERV categories and show the potential of our proposed usage of deep-link comments for video search through a user study. Raynor Vliegendhart, Martha A. Larson, Babak Loni, Alan Hanjalic |
IEEE Trans. Multim. | 2 |
| 2014 | CARS2: Learning Context-aware Representations for Context-aware RecommendationsabstractRich contextual information is typically available in many recommendation domains allowing recommender systems to model the subtle effects of context on preferences. Most contextual models assume that the context shares the same latent space with the users and items. In this work we propose CARS2, a novel approach for learning context-aware representations for context-aware recommendations. We show that the context-aware representations can be learned using an appropriate model that aims to represent the type of interactions between context variables, users and items. We adapt the CARS2 algorithms to explicit feedback data by using a quadratic loss function for rating prediction, and to implicit feedback data by using a pairwise and a listwise ranking loss functions for top-N recommendations. By using stochastic gradient descent for parameter estimation we ensure scalability. Experimental evaluation shows that our CARS2 models achieve competitive recommendation performance, compared to several state-of-the-art approaches. Yue Shi 0002, Alexandros Karatzoglou, Linas Baltrunas, Martha A. Larson, Alan Hanjalic |
CIKM | 4 |
| 2014 | Cross-Domain Collaborative Filtering with Factorization Machines
Babak Loni, Yue Shi 0002, Martha A. Larson, Alan Hanjalic |
ECIR | 3 |
| 2014 | SocialZap: Catch-up on Interesting Television Fragments Discovered from Social MediaabstractIn this paper we present SocialZap, a multimedia search engine that finds the most interesting fragments, zap points, in a television broadcast based on microblog posts and socially tagged photos. The main novelty of SocialZap is the fully-automatic transfer of the learned viewers interest from textual posts to the visual channel, without the need for any manual effort in the process. Once SocialZap finds the zap points, users can easily browse through a television broadcast and directly watch the interesting fragments. Thus, SocialZap adds social experience to watching television. Svetlana Kordumova, Christoph Kofler, Dennis C. Koelma, Bouke Huurnink, Bauke Freiburg, Joris Kleinveld, Manuel van Rijn, Marco van Deursen, Martha A. Larson, Cees Snoek |
ICMR | 9 |
| 2014 | VideoJot: A Multifunctional Video Annotation ToolabstractVideos are becoming more and more a tool of communication. There are how-to videos, people are discussing actions of others based on their recorded performance, e.g., in soccer, or they simply record videos of great moments and show them to friends and family. In this paper we focus on very specific how-to videos and present a novel, web based annotation tool, that combines (i) zoom, (ii) drawing, and (iii) temporal social bookmarking in video streams. Moreover, we present a short study on the usefulness of the tool to communicate general concepts of a specific video game based on a captured game session. Michael Riegler 0001, Mathias Lux, Vincent Charvillat, Axel Carlier, Raynor Vliegendhart, Martha A. Larson |
ICMR | 6 |
| 2014 | How 'How' Reflects What's What: Content-based Exploitation of How Users Frame Social ImagesabstractIn this paper, we introduce the concept of intentional framing, defined as the sum of the choices that a photographer makes on how to portray the subject matter of an image. We carry out analysis experiments that demonstrate the existence of a correspondence between image similarity that is calculated automatically on the basis of global feature representations, and image similarity that is perceived by humans at the level of intentional frames. Intentional framing has profound implications: The existence of a fundamental image-interpretation principle that explains the importance of global representations in capturing human-perceived image semantics reaches beyond currently dominant assumptions in multimedia research. The ability of fast global-feature approaches to compete with more `sophisticated' approaches, which are computationally more complex, is demonstrated using a simple search method (SimSea) to classify a large (2M) collection of social images by tag class. In short, intentional framing provides a principled connection between human interpretations of images and lightweight, fast image processing methods. Moving forward, it is critical that the community explicitly exploits such approaches, as the social image collections that we tackle, continue to grow larger. Michael Riegler 0001, Martha A. Larson, Mathias Lux, Christoph Kofler |
ACM Multimedia | 2 |
| 2014 | Fashion 10000: an enriched social image dataset for fashion and clothingabstractIn this work, we present a new social image dataset related to the fashion and clothing domain. The dataset contains more than 32000 images, their context and social metadata. Furthermore the dataset is enriched with several types of annotations collected from the Amazon Mechanical Turk (AMT) crowdsourcing platform, which can serve as ground truth for various content analysis algorithms. This dataset has been successfully used at the Crowdsourcing task of the 2013 MediaEval Multimedia Benchmarking initiative. The dataset contributes to several research areas such as Crowdsourcing, multimedia content and context analysis as well as hybrid human/automatic approaches. In this paper, the dataset is described in detail and the dataset collection strategy, statistics, applications of dataset and its contribution to MediaEval 2013 is discussed. Babak Loni, Lei Yen Cheung, Michael Riegler 0001, Alessandro Bozzon, Luke R. Gottlieb, Martha A. Larson |
MMSys | 6 |
| 2014 | Overview of ACM RecSys CrowdRec 2014 workshop: crowdsourcing and human computation for recommender systemsabstractThe CrowdRec workshop brings together the recommender system community for discussion and exchange of ideas. Its goal is to allow the potential of human computation and crowdsourcing to be exploited fully and sustainably, leading to the development of improved recommendation and information filtering technologies. Currently, the complete range of possible intelligent contributions that recommender systems could elicit from users is under-explored, and its full extent is unknown. Critical questions addressed in the workshop include how to: formulate crowdtasks, match tasks with crowdmembers, ensure the quality of crowd input, and integrate feedback from the crowd in an optimal manner to improve recommendation. Further, crowdsourcing can also be exploited for system design and system evaluation. Martha A. Larson, Paolo Cremonesi, Alexandros Karatzoglou |
RecSys | 1 |
| 2014 | 'Free lunch' enhancement for collaborative filtering with factorization machinesabstractThe advantage of Factorization Machines over other factorization models is their ability to easily integrate and efficiently exploit auxiliary information to improve Collaborative Filtering. Until now, this auxiliary information has been drawn from external knowledge sources beyond the user-item matrix. In this paper, we demonstrate that Factorization Machines can exploit additional representations of information inherent in the user-item matrix to improve recommendation performance. We refer to our approach as 'Free Lunch' enhancement since it leverages clusters that are based on information that is present in the user-item matrix, but not otherwise directly exploited during matrix factorization. Borrowing clustering concepts from codebook sharing, our approach can also make use of 'Free Lunch' information inherent in a user-item matrix from a auxiliary domain that is different from the target domain of the recommender. Our approach improves performance both in the joint case, in which the auxiliary and target domains share users, and in the disjoint case, in which they do not. Although 'Free Lunch' enhancement does not apply equally well to any given domain or domain combination, our overall conclusion is that Factorization Machines present an opportunity to exploit information that is ubiquitously present, but commonly under-appreciated by Collaborative Filtering algorithms. Babak Loni, Alan Said, Martha A. Larson, Alan Hanjalic |
RecSys | 3 |
| 2014 | Intent-Aware Video Search Result OptimizationabstractVideo search engines are relatively successful at returning search results that users find to be on topic. These results do not, however, completely satisfy the user's information need unless they also fulfill the user's intent, i.e., the immediate goal a user seeks to accomplish with video search. Satisfying a user's information need to its full extent poses a particular challenge to video search engines because user intent is often not explicitly reflected in the query. In this paper, we propose a multimodal approach that addresses this challenge by refining the results lists returned by a mainstream video search engine in order to optimally capture user intent. Our approach is based on the insight that the results lists returned by video search engines do contain videos that satisfy user's intent, but that videos with the highest potential for satisfaction are often buried within or scattered over the results list. The proposed approach consists of three steps. In the first step, it analyzes the initial results list to determine the intent distribution pattern. On the basis of this pattern, in the second step, it refines the video search results list such that the top of the list better reveals intent. The third step further improves this refinement by visual reranking, exploiting intent-sensitive lightweight visual features extracted from thumbnails. Extensive evaluation of the approach includes a user study carried out on a crowdsourcing platform and a system-oriented evaluation. Evaluation results demonstrate that our approach leads to a substantial improvement of the information need satisfaction at users. Christoph Kofler, Martha A. Larson, Alan Hanjalic |
IEEE Trans. Multim. | 2 |
| 2014 | Predicting Failing Queries in Video SearchabstractThe ability to predict when a video search query is not likely to deliver satisfying search results is expected to enable more effective search results optimizations and improved search experience for users. In this paper, we propose a novel context-aware query failure prediction approach that predicts whether a particular query submitted in a user's search session is likely to fail. The approach builds on the well-known concept of query performance prediction introduced in conventional text-based Web search to estimate the query's retrieval performance, but extends this concept with two novel characteristics, user indicators and engine indicators. User indicators are derived from transaction logs, capture the patterns of user interactions with the video search engine, and exploit the context in which a particular query was submitted. Engine indicators are derived from the search results list and measure the consistency of visual search results at the level of visual concepts and textual metadata associated with videos. Extensive evaluation of the approach on a test set containing over one million video search queries shows its effectiveness and demonstrates a significant improvement over traditional and state-of-the-art baseline approaches. Christoph Kofler, Linjun Yang, Martha A. Larson, Tao Mei 0001, Alan Hanjalic, Shipeng Li 0001 |
IEEE Trans. Multim. | 3 |
| 2014 | Corpus Development for Affective Video IndexingabstractAffective video indexing is the area of research that develops techniques to automatically generate descriptions of video content that encode the emotional reactions which the video content evokes in viewers. This paper provides a set of corpus development guidelines based on state-of-the-art practice intended to support researchers in this field. Affective descriptions can be used for video search and browsing systems offering users affective perspectives. The paper is motivated by the observation that affective video indexing has yet to fully profit from the standard corpora (data sets) that have benefited conventional forms of video indexing. Affective video indexing faces unique challenges, since viewer-reported affective reactions are difficult to assess. Moreover affect assessment efforts must be carefully designed in order to both cover the types of affective responses that video content evokes in viewers and also capture the stable and consistent aspects of these responses. We first present background information on affect and multimedia and related work on affective multimedia indexing, including existing corpora. Three dimensions emerge as critical for affective video corpora, and form the basis for our proposed guidelines: the context of viewer response, personal variation among viewers, and the effectiveness and efficiency of corpus creation. Finally, we present examples of three recent corpora and discuss how these corpora make progressive steps towards fulfilling the guidelines. Mohammad Soleymani 0001, Martha A. Larson, Thierry Pun, Alan Hanjalic |
IEEE Trans. Multim. | 2 |
| 2013 | K-component recurrent neural network language models using curriculum learningabstractConventional n-gram language models are known for their limited ability to capture long-distance dependencies and their brittleness with respect to within-domain variations. In this paper, we propose a k-component recurrent neural network language model using curriculum learning (CL-KRNNLM) to address within-domain variations. Based on a Dutch-language corpus, we investigate three methods of curriculum learning that exploit dedicated component models for specific sub-domains. Under an oracle situation in which context information is known during testing, we experimentally test three hypotheses. The first is that domain-dedicated models perform better than general models on their specific domains. The second is that curriculum learning can be used to train recurrent neural network language models (RNNLMs) from general patterns to specific patterns. The third is that curriculum learning, used as an implicit weighting method to adjust the relative contributions of general and specific patterns, outperforms conventional linear interpolation. Under the condition that context information is unknown during testing, the CL-KRNNLM also achieves improvement over conventional RNNLM by 13% relative in terms of word prediction accuracy. Finally, the CL-KRNNLM is tested in an additional experiment involving N-best rescoring on a standard data set. Here, the context domains are created by clustering the training data using Latent Dirichlet Allocation and k-means clustering. Yangyang Shi, Martha A. Larson, Catholijn M. Jonker |
ASRU | 2 |
| 2013 | GAPfm: optimal top-n recommendations for graded relevance domainsabstractRecommender systems are frequently used in domains in which users express their preferences in the form of graded judgments, such as ratings. Current ranking techniques are based on one of two sub-optimal approaches: either they optimize for a binary metric such as Average Precision, which discards information on relevance levels, or they optimize for Normalized Discounted Cumulative Gain (NDCG), which ignores the dependence of an item's contribution on the relevance of more highly ranked items. We address the shortcomings of existing approaches by proposing GAPfm, the Graded Average Precision factor model, which is a latent factor model for top-N recommendation in domains with graded relevance data. The model optimizes the Graded Average Precision metric that has been proposed recently for assessing the quality of ranked results lists for graded relevance. GAPfm's advantages are twofold: it maintains full information about graded relevance and also addresses the limitations of models that optimize NDCG. Experimental results show that GAPfm achieves substantial improvements on the top-N recommendation task, compared to several state-of-the-art approaches. Yue Shi 0002, Alexandros Karatzoglou, Linas Baltrunas, Martha A. Larson, Alan Hanjalic |
CIKM | 4 |
| 2013 | I want to be Sachin Tendulkar!: a spoken english cricket game for rural studentsabstractWe present a mobile phone based cricket game for improving the spoken English pronunciation of school children designed and tested within a specific socio-cultural context in rural India, the Mewat district of Haryana State. The development of the game concept was informed by a field study, which identified the cultural restrictions respected by the community as well community interests and motivations. The game was accessible using a low-end mobile phone and evaluated with a group of 63 students from classes 4 and 5 of a rural school. The results suggest that the cricket game can be effectively used to engage students and to improve their spoken English skills in the given setting. Further, the game serves as an informative example of how factors impacting the acceptability and appropriation of a technology in a particular setting can be taken into account from the very beginning of the design process. Martha A. Larson, Nitendra Rajput, Abhigyan Singh |
CSCW | 1 |
| 2013 | CLiMF: Collaborative Less-Is-More Filtering
Yue Shi 0002, Alexandros Karatzoglou, Linas Baltrunas, Martha A. Larson, Nuria Oliver, Alan Hanjalic |
IJCAI | 4 |
| 2013 | Speed up of recurrent neural network language models with sentence independent subsampling stochastic gradient descent
Yangyang Shi, Mei-Yuh Hwang, Kaisheng Yao, Martha A. Larson |
INTERSPEECH | 4 |
| 2013 | Exploiting the succeeding words in recurrent neural network language models
Yangyang Shi, Martha A. Larson, Pascal Wiggers, Catholijn M. Jonker |
INTERSPEECH | 2 |
| 2013 | Multimedia information seeking through search and hyperlinkingabstractSearching for relevant webpages and following hyperlinks to related content is a widely accepted and effective approach to information seeking on the textual web. Existing work on multimedia information retrieval has focused on search for individual relevant items or on content linking without specific attention to search results. We describe our research exploring integrated multimodal search and hyperlinking for multimedia data. Our investigation is based on the MediaEval 2012 Search and Hyperlinking task. This includes a known-item search task using the Blip10000 internet video collection, where automatically created hyperlinks link each relevant item to related items within the collection. The search test queries and link assessment for this task was generated using the Amazon Mechanical Turk crowdsourcing platform. Our investigation examines a range of alternative methods which seek to address the challenges of search and hyperlinking using multimodal approaches. The results of our experiments are used to propose a research agenda for developing effective techniques for search and hyperlinking of multimedia content. Maria Eskevich, Gareth J. F. Jones, Robin Aly, Roeland Ordelman, Danish Nadeem, Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot, Tom De Nies, Pedro Debevere, Rik Van de Walle, Petra Galuscáková, Pavel Pecina, Martha A. Larson |
ICMR | 15 |
| 2013 | Geo-visual ranking for location prediction of social imagesabstractPredicting geographic location using exclusively the visual content of images holds the promise of greatly benefiting users' access to media collections. In this paper, we present a visual-content-based approach that predicts where in the world a social image was taken. We employ a ranking method that assigns a query photo the geo-location of its most likely geo-visual neighbor in the social image collection. The novelty of the approach is that ranking makes use not only of the photos themselves, but also their geo-visual neighbors. In contrast to other approaches, we do not restrict the locations we predict to landmarks or specific cities. The approach is evaluated on a set of 3 million geo-tagged photos from Flickr, released by MediaEval 2012. Experiments show that the proposed system delivers a substantive performance improvement compared with previously proposed, related visual content-based approaches. The discussion illustrates how photo densities, geo-visual redundancy and uploader patterns characteristic of social image collections impacts the performance. Xinchao Li, Martha A. Larson, Alan Hanjalic |
ICMR | 2 |
| 2013 | ACM multimedia 2013 workshop on crowdsourcing for multimediaabstractThe topic "Crowdsourcing for Multimedia" encompasses the full range of techniques that combine human intelligence and a large number of individual contributors to advance the state of the art in multimedia research. The ACM Multimedia 2013 Workshop on Crowdsourcing for Multimedia (CrowdMM 2013) provided a forum for presenting new crowdsourcing techniques, exchanging innovative crowdsourcing ideas, and discussing crowdsourcing best practices for multimedia. The workshop program consisted of presented papers, a keynote speech and a panel discussion. A special feature of this year's workshop was the "Crowdsourcing for Multimedia Ideas Competition", the results of which were presented at the workshop. Kuan-Ta Chen, Wei-Ta Chu, Martha A. Larson |
ACM Multimedia | 3 |
| 2013 | Crowdsourcing for multimedia researchabstractCrowdsourcing techniques make use of intelligent contributions of large number of human crowdmembers. This tutorial introduces researchers to the applications of crowdsourcing to multimedia analysis with the aim of allowing them to understand the potentials and limitations of crowdsourcing tools and techniques. We emphasize the fact that crowdsourcing represents a further development along a pre-existing continuum of techniques, and discuss the added advantages that new developments offer. We provide a basic overview of human computation, with an emphasis on example cases in which crowdsourcing has been applied to generate data sets, to improve automatic multimedia content analysis, and to elicit user needs or multimedia system requirements. Different techniques and considerations in using human computation methods to acquire high-quality data and annotations are discussed and demonstrated. Mohammad Soleymani 0001, Martha A. Larson |
ACM Multimedia | 2 |
| 2013 | How do we deep-link?: leveraging user-contributed time-links for non-linear video accessabstractThis paper studies a new way of accessing videos in a non-linear fashion. Existing non-linear access methods allow users to jump into videos at points that depict specific visual concepts or that are likely to elicit affective reactions. We believe that deep-link comments, which occur unprompted on social video sharing platforms, offer a new opportunity beyond existing methods. With deep-link comments, viewers express themselves about a particular moment in a video by including a time-code. Deep-link comments are special because they reflect viewer perceptions of noteworthiness, that include, but extend beyond depicted conceptual content and induced affective reactions. Based on deep-link comments collected from YouTube, we develop a Viewer Expressive Reaction Variety (VERV) taxonomy that captures how viewers deep-link. We validate the taxonomy with a user study on a crowdsourcing platform and discuss how it extends conventional relevance criteria. We carry out experiments which show that deep-link comments can be automatically filtered and sorted into VERV categories. Raynor Vliegendhart, Babak Loni, Martha A. Larson, Alan Hanjalic |
ACM Multimedia | 3 |
| 2013 | Fashion-focused creative commons social datasetabstractIn this work, we present a fashion-focused Creative Commons dataset, which is designed to contain a mix of general images as well as a large component of images that are focused on fashion (i.e., relevant to particular clothing items or fashion accessories). The dataset contains 4810 images and related metadata. Furthermore, a ground truth on image's tags is presented. Ground truth generation for large-scale datasets is a necessary but expensive task. Traditional expert based approaches have become an expensive and non-scalable solution. For this reason, we turn to crowdsourcing techniques in order to collect ground truth labels; in particular we make use of the commercial crowdsourcing platform, Amazon Mechanical Turk (AMT). Two different groups of annotators (i.e., trusted annotators known to the authors and crowdsourcing workers on AMT) participated in the ground truth creation. Annotation agreement between the two groups is analyzed. Applications of the dataset in different contexts are discussed. This dataset contributes to research areas such as crowdsourcing for multimedia, multimedia content analysis, and design of systems that can elicit fashion preferences from users. Babak Loni, María Menéndez-Blanco, Mihai Georgescu, Luca Galli, Claudio Massari, Ismail Sengör Altingövde, Davide Martinenghi, Mark S. Melenhorst, Raynor Vliegendhart, Martha A. Larson |
MMSys | 10 |
| 2013 | Blip10000: a social video dataset containing SPUG content for tagging and retrievalabstractThe increasing amount of digital multimedia content available is inspiring potential new types of user interaction with video data. Users want to easily find the content by searching and browsing. For this reason, techniques are needed that allow automatic categorisation, searching the content and linking to related information. In this work, we present a dataset that contains comprehensive semi-professional user-generated (SPUG) content, including audiovisual content, user-contributed metadata, automatic speech recognition transcripts, automatic shot boundary files, and social information for multiple 'social levels'. We describe the principal characteristics of this dataset and present results that have been achieved on different tasks. Sebastian Schmiedeke, Isabelle Ferrané, Maria Eskevich, Christoph Kofler, Martha A. Larson, Yannick Estève, Lori Lamel, Gareth J. F. Jones, Thomas Sikora |
MMSys | 6 |
| 2013 | xCLiMF: optimizing expected reciprocal rank for data with multiple levels of relevanceabstractExtended Collaborative Less-is-More Filtering xCLiMF is a learning to rank model for collaborative filtering that is specifically designed for use with data where information on the level of relevance of the recommendations exists, e.g. through ratings. xCLiMF can be seen as a generalization of the Collaborative Less-is-More Filtering (CLiMF) method that was proposed for top-N recommendations using binary relevance (implicit feedback) data. The key contribution of the xCLiMF algorithm is that it builds a recommendation model by optimizing Expected Reciprocal Rank, an evaluation metric that generalizes reciprocal rank in order to incorporate user feedback with multiple levels of relevance. Experimental results on real-world datasets show the effectiveness of xCLiMF, and also demonstrate its advantage over CLiMF when more than two levels of relevance exist in the data. Yue Shi 0002, Alexandros Karatzoglou, Linas Baltrunas, Martha A. Larson, Alan Hanjalic |
RecSys | 4 |
| 2013 | Unifying rating-oriented and ranking-oriented collaborative filtering for improved recommendation
Yue Shi 0002, Martha A. Larson, Alan Hanjalic |
Inf. Sci. | 2 |
| 2013 | Mining contextual movie similarity with matrix factorization for context-aware recommendationabstractContext-aware recommendation seeks to improve recommendation performance by exploiting various information sources in addition to the conventional user-item matrix used by recommender systems. We propose a novel context-aware movie recommendation algorithm based on joint matrix factorization (JMF). We jointly factorize the user-item matrix containing general movie ratings and other contextual movie similarity matrices to integrate contextual information into the recommendation process. The algorithm was developed within the scope of the mood-aware recommendation task that was offered by the Moviepilot mood track of the 2010 context-aware movie recommendation (CAMRa) challenge. Although the algorithm could generalize to other types of contextual information, in this work, we focus on two: movie mood tags and movie plot keywords. Since the objective in this challenge track is to recommend movies for a user given a specified mood, we devise a novel mood-specific movie similarity measure for this purpose. We enhance the recommendation based on this measure by also deploying the second movie similarity measure proposed in this article that takes into account the movie plot keywords. We validate the effectiveness of the proposed JMF algorithm with respect to the recommendation performance by carrying out experiments on the Moviepilot challenge dataset. We demonstrate that exploiting contextual information in JMF leads to significant improvement over several state-of-the-art approaches that generate movie recommendations without using contextual information. We also demonstrate that our proposed mood-specific movie similarity is better suited for the task than the conventional mood-based movie similarity measures. Finally, we show that the enhancement provided by the movie similarity capturing the plot keywords is particularly helpful in improving the recommendation to those users who are significantly more active in rating the movies than other users. Yue Shi 0002, Martha A. Larson, Alan Hanjalic |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2013 | Nontrivial landmark recommendation using geotagged photosabstractOnline photo-sharing sites provide a wealth of information about user behavior and their potential is increasing as it becomes ever-more common for images to be associated with location information in the form of geotags. In this article, we propose a novel approach that exploits geotagged images from an online community for the purpose of personalized landmark recommendation. Under our formulation of the task, recommended landmarks should be relevant to user interests and additionally they should constitute nontrivial recommendations. In other words, recommendations of landmarks that are highly popular and frequently visited and can be easily discovered through other information sources such as travel guides should be avoided in favor of recommendations that relate to users' personal interests. We propose a collaborative filtering approach to the personalized landmark recommendation task within a matrix factorization framework. Our approach, WMF-CR, combines weighted matrix factorization and category-based regularization. The integrated weights emphasize the contribution of nontrivial landmarks in order to focus the recommendation model specifically on the generation of nontrivial recommendations. They support the judicious elimination of trivial landmarks from consideration without also discarding information valuable for recommendation. Category-based regularization addresses the sparse data problem, which is arguably even greater in the case of our landmark recommendation task than in other recommendation scenarios due to the limited amount of travel experience recorded in the online image set of any given user. We use category information extracted from Wikipedia in order to provide the system with a method to generalize the semantics of landmarks and allow the model to relate them not only on the basis of identity, but also on the basis of topical commonality. The proposed approach is computational scalable, that is, its complexity is linear with the number of observed preferences in the user-landmark preference matrix and the number of nonzero similarities in the category-based landmark similarity matrix. We evaluate the approach on a large collection of geotagged photos gathered from Flickr. Our experimental results demonstrate that WMF-CR outperforms several state-of-the-art baseline approaches in recommending nontrivial landmarks. Additionally, they demonstrate that the approach is well suited for addressing data sparseness and provides particular performance improvement in the case of users who have limited travel experience, that is, have visited only few cities or few landmarks. Yue Shi 0002, Pavel Serdyukov, Alan Hanjalic, Martha A. Larson |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2013 | Generating Visual Summaries of Geographic Areas Using Community-Contributed ImagesabstractIn this paper, we present a novel approach for automatic visual summarization of a geographic area that exploits user-contributed images and related explicit and implicit metadata collected from popular content-sharing websites. By means of this approach, we search for a limited number of representative but diverse images to represent the area within a certain radius around a specific location. Our approach is based on the random walk with restarts over a graph that models relations between images, visual features extracted from them, associated text, as well as the information on the uploader and commentators. In addition to introducing a novel edge weighting mechanism, we propose in this paper a simple but effective scheme for selecting the most representative and diverse set of images based on the information derived from the graph. We also present a novel evaluation protocol, which does not require input of human annotators, but only exploits the geographical coordinates accompanying the images in order to reflect conditions on image sets that must necessarily be fulfilled in order for users to find them representative and diverse. Experiments performed on a collection of Flickr images, captured around 207 locations in Paris, demonstrate the effectiveness of our approach. Stevan Rudinac, Alan Hanjalic, Martha A. Larson |
IEEE Trans. Multim. | 3 |
| 2013 | Learning Crowdsourced User Preferences for Visual Summarization of Image CollectionsabstractIn this paper we propose a novel approach to selecting images suitable for inclusion in the visual summaries. The approach is grounded in insights about how people summarize image collections. We utilize the Amazon Mechanical Turk crowdsourcing platform to obtain a large number of manually created visual summaries as well as information about criteria for image inclusion in the summary. Based on these large-scale user tests, we propose an automatic image selection approach, which jointly utilizes the analysis of image content, context, popularity, visual aesthetic appeal as well as the sentiment derived from the comments posted on the images. In our approach we do not describe images based on their properties only, but also in the context of semantically related images, which improves robustness and effectively enables propagation of sentiment, aesthetic appeal as well as various inherent attributes associated with a particular group of images. We discuss the phenomenon of a low inter-user agreement, which makes an automatic evaluation of visual summaries a challenging task and propose a solution inspired by the text summarization and machine translation communities. The experiments performed on a collection of geo-referenced Flickr images demonstrate the effectiveness of our image selection approach. Stevan Rudinac, Martha A. Larson, Alan Hanjalic |
IEEE Trans. Multim. | 2 |
| 2012 | Creating a Data Collection for Evaluating Rich Speech Retrieval
Maria Eskevich, Gareth J. F. Jones, Martha A. Larson, Roeland Ordelman |
LREC | 3 |
| 2012 | GeoMM'12: ACM international workshop on geotagging and its applications in multimediaabstractGeotagging is the process of adding geographical identification metadata to various media files such as photos, videos, websites, messages, and tweets. It is not limited to GPS sensor data but an extension of current multimedia files with a wide variety of location-specific information. The GeoMM'12 workshop presents research on recent research on geotagging within the context of multimedia analysis. This workshop aims to not only provide more cutting edge algorithms, but also motivate novel applications in this promising field. Liangliang Cao, Gerald Friedland, Martha A. Larson |
ACM Multimedia | 3 |
| 2012 | ACM multimedia 2012 workshop on crowdsourcing for multimediaabstractCrowdsourcing for multimedia involves exploiting both human intelligence and the combination of a large number of individual human contributions (i.e., the 'wisdom of the crowd') to develop techniques, systems and data sets that advance the state of the art. The ACM Multimedia 2012 Workshop on Crowdsourcing for Multimedia (CrowdMM 2012) provides a forum presenting crowdsourcing techniques for multimedia, as well as innovative ideas exemplifying how multimedia research can benefit from crowdsourcing. Through presented papers, invited talks and a panel, the workshop will promote interactive discussion on the scope and research potentials of crowdsourcing. The goal is to provide information to the multimedia research community on the principles of crowdsourcing and to inspire researchers to address the limitations of current studies by innovative use of human computation and collective intelligence. The workshop views crowdsourcing in the broad sense: it encompasses both unsolicited human contributions, e.g., tags assigned by users to images, and also solicited contributions, e.g., annotations gathered by making use of crowdsourcing platforms that micro-outsource tasks to a large pool of human workers. Kuan-Ta Chen, Wei-Ta Chu, Martha A. Larson, Wei Tsang Ooi |
ACM Multimedia | 3 |
| 2012 | Intent and its discontents: the user at the wheel of the online video search engineabstractWe embrace the position of the user in the driver's seat of the video search engine by proposing a principled framework for multimedia retrieval that moves beyond what users are searching for also to encompass why they search. This 'why' is understood as the reason, purpose or immediate goal behind a user information need, which we identify as the underlying user 'intent'. Breaking information needs down into a topical dimension representing 'what' and an intent dimension representing 'why' will allow online video search engines to provide users with more satisfying search results. Until now, research on intent has remained small scale, limited by the lack of a systematic method for arriving at possible dimensions of user intent that provide productive areas of inquiry for multimedia research. We demonstrate how mining information from user descriptions of video information needs on the social Web makes it possible to identify useful intent categories for online video search and carry out validation experiments showing that these categories display enough invariance to be successfully modeled by a video search engine. In a final experiment, we demonstrate the potential for these categories to improve video retrieval with a large user study confirming that users associate salient differences within topically homogenous video search engine results lists with these intent categories. This reveals the potential to refine video results list using user intent. Alan Hanjalic, Christoph Kofler, Martha A. Larson |
ACM Multimedia | 3 |
| 2012 | When video search goes wrong: predicting query failure using search engine logs and visual search resultsabstractThe recent increase in the volume and variety of video content available online presents growing challenges for video search. Users face increased difficulty in formulating effective queries and search engines must deploy highly effective algorithms to provide relevant results. Although lately much effort has been invested in optimizing video search engine results, relatively little attention has been given to predicting for which queries results optimization is most useful, i.e., predicting which queries will fail. Being able to predict when a video search query would fail is likely to make the video search result optimization more efficient and effective, improve the search experience for the user by providing support in the query formulation process and in this way boost the development of video search engines in general. While insight about a query's performance in general could be obtained using the well-known concept of query performance prediction (QPP), we propose a novel approach for predicting a failure of a video search query in the specific context of a search session. Our 'context-aware query failure' prediction approach uses a combination of 'user indicators' and 'engine indicators' to predict whether a particular query is likely to fail in the context of a particular search session. User indicators are derived from the search log and capture the patterns of query (re)formulation behavior and the click-through data of a user during a typical video search session. Engine indicators are derived from the video search results list and capture the visual variance of search results that would be offered to the user for the given query. We validate our approach experimentally on a test set containing 1+ million video search queries and show its effectiveness compared to a set of conventional QPP baselines. Our approach achieves a 13% relative improvement over the baseline. Christoph Kofler, Linjun Yang, Martha A. Larson, Tao Mei 0001, Alan Hanjalic, Shipeng Li 0001 |
ACM Multimedia | 3 |
| 2012 | LikeLines: collecting timecode-level feedback for web videos through user interactionsabstractConventional online video players do not make the inner structure of the video apparent, making it hard to jump straight to the interesting parts. Our LikeLines system provides its users with a navigable heat map of interesting regions for the videos they are watching. Its novelty lies in its combination of content analysis and both explicit and implicit user interactions. The system can be readily used and deployed to collect large amounts of interaction data needed for in-depth research on timecode-level feedback. Raynor Vliegendhart, Martha A. Larson, Alan Hanjalic |
ACM Multimedia | 2 |
| 2012 | CLiMF: learning to maximize reciprocal rank with collaborative less-is-more filteringabstractIn this paper we tackle the problem of recommendation in the scenarios with binary relevance data, when only a few (k) items are recommended to individual users. Past work on Collaborative Filtering (CF) has either not addressed the ranking problem for binary relevance datasets, or not specifically focused on improving top-k recommendations. To solve the problem we propose a new CF approach, Collaborative Less-is-More Filtering (CLiMF). In CLiMF the model parameters are learned by directly maximizing the Mean Reciprocal Rank (MRR), which is a well-known information retrieval metric for measuring the performance of top-k recommendations. We achieve linear computational complexity by introducing a lower bound of the smoothed reciprocal rank metric. Experiments on two social network datasets demonstrate the effectiveness and the scalability of CLiMF, and show that CLiMF significantly outperforms a naive baseline and two state-of-the-art CF methods. Yue Shi 0002, Alexandros Karatzoglou, Linas Baltrunas, Martha A. Larson, Nuria Oliver, Alan Hanjalic |
RecSys | 4 |
| 2012 | TFMAP: optimizing MAP for top-n context-aware recommendationabstractIn this paper, we tackle the problem of top-N context-aware recommendation for implicit feedback scenarios. We frame this challenge as a ranking problem in collaborative filtering (CF). Much of the past work on CF has not focused on evaluation metrics that lead to good top-N recommendation lists in designing recommendation models. In addition, previous work on context-aware recommendation has mainly focused on explicit feedback data, i.e., ratings. We propose TFMAP, a model that directly maximizes Mean Average Precision with the aim of creating an optimally ranked list of items for individual users under a given context. TFMAP uses tensor factorization to model implicit feedback data (e.g., purchases, clicks) with contextual information. Yue Shi 0002, Alexandros Karatzoglou, Linas Baltrunas, Martha A. Larson, Alan Hanjalic, Nuria Oliver |
SIGIR | 4 |
| 2012 | Adaptive diversification of recommendation results via latent factor portfolioabstractThis paper studies result diversification in collaborative filtering. We argue that the diversification level in a recommendation list should be adapted to the target users' individual situations and needs. Different users may have different ranges of interests -- the preference of a highly focused user might include only few topics, whereas that of the user with broad interests may encompass a wide range of topics. Thus, the recommended items should be diversified according to the interest range of the target user. Such an adaptation is also required due to the fact that the uncertainty of the estimated user preference model may vary significantly between users. To reduce the risk of the recommendation, we should take the difference of the uncertainty into account as well. Yue Shi 0002, Jun Wang 0012, Martha A. Larson, Alan Hanjalic |
SIGIR | 4 |
| 2012 | Special issue on searching speechabstractNo abstract available. Martha A. Larson, Franciska de Jong, Wessel Kraaij, Steve Renals |
ACM Trans. Inf. Syst. | 1 |
| 2011 | The where in the tweetabstractTwitter is a widely-used social networking service which enables its users to post text-based messages, so-called tweets. POI tags on tweets can show more human-readable high-level information about a place rather than just a pair of coordinates. In this paper, we attempt to predict the POI tag of a tweet based on its textual content and time of posting. Potential applications include accurate positioning when GPS devices fail and disambiguating places located near each other. We consider this task as a ranking problem, i.e., we try to rank a set of candidate POIs according to a tweet by using language and time models. To tackle the sparsity of tweets tagged with POIs, we use web pages retrieved by search engines as an additional source of evidence. From our experiments, we find that users indeed leak some information about their accurate locations in their tweets. Pavel Serdyukov, Arjen P. de Vries, Carsten Eickhoff, Martha A. Larson |
CIKM | 5 |
| 2011 | A peer's-eye view: network term clouds in a peer-to-peer system
Raynor Vliegendhart, Martha A. Larson, Christoph Kofler, Johan A. Pouwelse |
CIKM | 2 |
| 2011 | To Seek, Perchance to Fail: Expressions of User Needs in Internet Video Search
Christoph Kofler, Martha A. Larson, Alan Hanjalic |
ECIR | 2 |
| 2011 | Reranking Collaborative Filtering with Multiple Self-contained Modalities
Yue Shi 0002, Martha A. Larson, Alan Hanjalic |
ECIR | 2 |
| 2011 | How Far Are We in Trust-Aware Recommendation?
Yue Shi 0002, Martha A. Larson, Alan Hanjalic |
ECIR | 2 |
| 2011 | Personalized Landmark Recommendation Based on Geotags from Photo Sharing Sites
Yue Shi 0002, Pavel Serdyukov, Alan Hanjalic, Martha A. Larson |
ICWSM | 4 |
| 2011 | Automatic tagging and geotagging in video collections and communitiesabstractAutomatically generated tags and geotags hold great promise to improve access to video collections and online communities. We overview three tasks offered in the MediaEval 2010 benchmarking initiative, for each, describing its use scenario, definition and the data set released. For each task, a reference algorithm is presented that was used within MediaEval 2010 and comments are included on lessons learned. The Tagging Task, Professional involves automatically matching episodes in a collection of Dutch television with subject labels drawn from the keyword thesaurus used by the archive staff. The Tagging Task, Wild Wild Web involves automatically predicting the tags that are assigned by users to their online videos. Finally, the Placing Task requires automatically assigning geo-coordinates to videos. The specification of each task admits the use of the full range of available information including user-generated metadata, speech recognition transcripts, audio, and visual features. Martha A. Larson, Mohammad Soleymani 0001, Pavel Serdyukov, Stevan Rudinac, Christian Wartena, Vanessa Murdock 0001, Gerald Friedland, Roeland Ordelman, Gareth J. F. Jones |
ICMR | 1 |
| 2011 | Frontiers in multimedia searchabstractThis outline summarizes our tutorial "Frontiers in Multimedia Search", whose goal is to provide insights into the most recent developments in the field of multimedia retrieval and to identify the issues and bottlenecks that could determine the directions of research focus for the coming years. We present an overview of new algorithms and techniques, particularly concentrating on those innovative approaches that are informed by neighboring fields including information retrieval, speech and language processing and network analysis. We also discuss evaluation of new algorithms, in particular, making use of crowdsourcing for the development of the necessary data sets. Alan Hanjalic, Martha A. Larson |
ACM Multimedia | 2 |
| 2011 | Alice's worlds of wonder: exploiting tags to understand images in terms of size and scaleabstractThe 'Wonderlands of Size and Scale' system extends Flickr search functionality with a mechanism that exploits information implicit in image tag sets to filter images on the basis of characteristics related to size and scale. The innovative contribution of the 'Wonderlands' system is threefold: first, its use of an understanding of real-world physical entities depicted in images in terms of size and scale to filter images for display to the user; second, its application of our recently proposed 'Reading between the Tags' approach, which infers the real-world size of physical objects depicted in images by combining user-assigned image tags and natural language statistics mined from the Web; and, third, its trim realization of an engaging application that is implemented in a light-weight manner such that it can be executed as a live extension to the Flickr search engine. Results of a simple system-oriented evaluation support the conclusion that 'Wonderlands' reaches its aim of presenting users with images sorted with respect to the size- and scale-characteristics of their visually depicted content. Christoph Kofler, Martha A. Larson, Alan Hanjalic |
ACM Multimedia | 2 |
| 2011 | Reading between the tags to predict real-world size-class for visually depicted objects in imagesabstractMultimedia information retrieval stands to benefit from the availability of additional information about tags and how they relate to the content visually depicted in images. We propose a generic approach that contributes to improving the informativeness of image tags by combining generalizations about the distributional tendencies of physical objects in the real world and statistics of natural language use patterns that have been mined from the Web. The approach, which we refer to as 'Reading between the Tags,' provides for each tag associated with an image, first, a prediction concerning corporeality, i.e., whether or not the tag denotes a physical entity, and, then, concerning the real-world size of that entity, i.e., large, medium or small. Mining takes place using a set of Language Use Frames (LUFs) that are composed of natural language neighborhoods characteristic of tag classes. We validate our approach with a series of experiments on a set of images from the MIRFLICKR data set using ground truth created with standard crowdsourcing techniques. The main experiments demonstrate the effectiveness of our approach for size-class prediction. A further experiment shows that size-class prediction can be improved and made image-specific using general and relatively small sets of visual concepts. A final experiment confirms that the set of LUFs can also be chosen automatically via statistical feature selection. Martha A. Larson, Christoph Kofler, Alan Hanjalic |
ACM Multimedia | 1 |
| 2011 | Finding representative and diverse community contributed images to create visual summaries of geographic areasabstractThis paper presents an automatic approach that uses community-contributed images to create representative and diverse visual summaries of specific geographic areas. Complex relations between images, extracted visual features, text associated with the images as well as users and their social network are modeled using a multimodal graph. To compute affinities between nodes in the graph we rely on the proven concept of random walk with restarts. The novelty of our approach lies in its use of the multimodal graph to create a diverse, yet representative, image set. Further, we introduce an edge-weighting mechanism for the fusion of heterogeneous modalities. We evaluate our summaries with a new protocol that tests for representativeness and diversity using image geo-coordinates and is independent of the need for human evaluators. The experiments, performed on a set of Flickr images, demonstrate the effectiveness of our approach. Stevan Rudinac, Alan Hanjalic, Martha A. Larson |
ACM Multimedia | 3 |
| 2011 | Tags as Bridges between Domains: Improving Recommendation with Tag-Induced Cross-Domain Collaborative Filtering
Yue Shi 0002, Martha A. Larson, Alan Hanjalic |
UMAP | 2 |
| 2010 | Exploiting Result Consistency to Select Query Expansions for Spoken Content Retrieval
Stevan Rudinac, Martha A. Larson, Alan Hanjalic |
ECIR | 2 |
| 2010 | Towards affective state modeling in narrative and conversational settingsabstractWe carry out two studies on affective state modeling for communication settings that involve unilateral intent on the part of one participant (the evoker) to shift the affective state of another participant (the experiencer). The first investigates viewer response in a narrative setting using a corpus of docu-mentaries annotated with viewer-reported narrative peaks. The second investigates affective triggers in a conversational set-ting using a corpus of recorded interactions, annotated with continuous affective ratings, between a human interlocutor and an emotionally colored agent. In each case, we build a “one-sided ” model using indicators derived from the speech of one participant. Our classification experiments confirm the viabil-ity of our models and provide insight into useful features. Index Terms: affect, speech recognition, audio analysis, natural language communication Bart Jochems, Martha A. Larson, Roeland Ordelman, Ronald Poppe, Khiet P. Truong |
INTERSPEECH | 2 |
| 2010 | Contextual verification for open vocabulary spoken term detectionabstractIn spoken term detection, subword speech recognition is a viable means for addressing the out-of-vocabulary (OOV) problem at query time. Applying fuzzy error compensation techniques is needed for coping with inevitable recognition errors, but can lead to high false alarm rates especially for short queries. We propose two novel methods which reject false alarms based on the context of the hypothesized result and the distance to phonetically similar queries. Using the proposed methods, we obtain an increase in precision of 11% absolute at equal recall. Daniel Schneider 0003, Timo Mertens, Martha A. Larson, Joachim Köhler |
INTERSPEECH | 3 |
| 2010 | Advances in multimedia retrieval, part i: frontiers in multimedia searchabstractNo abstract available. Alan Hanjalic, Martha A. Larson |
ACM Multimedia | 2 |
| 2010 | Multimedia content with a speech track: ACM multimedia 2010 workshop on searching spontaneous conversational speechabstractNo abstract available. Martha A. Larson, Roeland Ordelman, Florian Metze, Wessel Kraaij, Franciska de Jong |
ACM Multimedia | 1 |
| 2010 | Exploiting noisy visual concept detection to improve spoken content based video retrievalabstractIn this paper, we present a technique for unsupervised construction of concept vectors, concept-based representations of complete video units, from the noisy shot-level output of a set of visual concept detectors. We deploy these vectors to improve spoken-content-based video retrieval using Query Expansion Selection (QES). Our QES approach analyzes results lists returned in response to several alternative query expansions, applying a coherence indicator calculated on top-ranked items to choose the appropriate expansion. The approach is data driven, does not require prior training and relies solely on the analysis of the collection being queried and the results lists produced for the given query text. The experiments, performed on two datasets, TRECVID 2007/2008 and TRECVID 2009, demonstrate the effectiveness of our approach and show that a small set of well-selected visual concept detectors is sufficient to improve retrieval performance. Stevan Rudinac, Martha A. Larson, Alan Hanjalic |
ACM Multimedia | 2 |
| 2010 | List-wise learning to rank with matrix factorization for collaborative filteringabstractA ranking approach, ListRank-MF, is proposed for collaborative filtering that combines a list-wise learning-to-rank algorithm with matrix factorization (MF). A ranked list of items is obtained by minimizing a loss function that represents the uncertainty between training lists and output lists produced by a MF ranking model. ListRank-MF enjoys the advantage of low complexity and is analytically shown to be linear with the number of observed ratings for a given user-item matrix. We also experimentally demonstrate the effectiveness of ListRank-MF by comparing its performance with that of item-based collaborative recommendation and a related state-of-the-art collaborative ranking approach (CoFiRank). Yue Shi 0002, Martha A. Larson, Alan Hanjalic |
RecSys | 2 |
| 2010 | Visual concept-based selection of query expansions for spoken content retrievalabstractIn this paper we present a novel approach to semantic-theme-based video retrieval that considers entire videos as retrieval units and exploits automatically detected visual concepts to improve the results of retrieval based on spoken content. We deploy a query prediction method that makes use of a coherence indicator calculated on top returned documents and taking into account the information about visual concepts presence in videos to make a choice between query expansion methods. The main contribution of our approach is in its ability to exploit noisy shot-level concept detection to improve semantic-theme-based video retrieval. Strikingly, improvement is possible using an extremely limited set of concepts. In the experiments performed on TRECVID 2007 and 2008 datasets our approach shows an interesting performance improvement compared to the best performing baseline. Stevan Rudinac, Martha A. Larson, Alan Hanjalic |
SIGIR | 2 |
| 2010 | Predicting podcast preference: An analysis framework and its applicationabstractAbstract Finding worthwhile podcasts can be difficult for listeners since podcasts are published in large numbers and vary widely with respect to quality and repute. Independently of their informational content, certain podcasts provide satisfying listening material while other podcasts have little or no appeal. In this paper we present PodCred, a framework for analyzing listener appeal, and we demonstrate its application to the task of automatically predicting the listening preferences of users. First, we describe the PodCred framework, which consists of an inventory of factors contributing to user perceptions of the credibility and quality of podcasts. The framework is designed to support automatic prediction of whether or not a particular podcast will enjoy listener preference. It consists of four categories of indicators related to the Podcast Content , the Podcaster , the Podcast Context , and the Technical Execution of the podcast. Three studies contributed to the development of the PodCred framework: a review of the literature on credibility for other media, a survey of prescriptive guidelines for podcasting, and a detailed data analysis. Next, we report on a validation exercise in which the PodCred framework is applied to a real‐world podcast preference prediction task. Our validation focuses on select framework indicators that show promise of being both discriminative and readily accessible. We translate these indicators into a set of easily extractable “surface” features and use them to implement a basic classification system. The experiments carried out to evaluate system use popularity levels in iTunes as ground truth and demonstrate that simple surface features derived from the PodCred framework are indeed useful for classifying podcasts. Manos Tsagkias, Martha A. Larson, Maarten de Rijke |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2009 | Investigating the Global Semantic Impact of Speech Recognition Error on Spoken Content Collections
Martha A. Larson, Manos Tsagkias, Jiyin He, Maarten de Rijke |
ECIR | 1 |
| 2009 | Exploiting Surface Features for the Prediction of Podcast Preference
Manos Tsagkias, Martha A. Larson, Maarten de Rijke |
ECIR | 2 |
| 2009 | Searching multimedia content with a spontaneous conversational speech trackabstractNo abstract available. Martha A. Larson, Roeland Ordelman, Franciska de Jong, Wessel Kraaij, Joachim Köhler |
ACM Multimedia | 1 |
| 2009 | Exploiting user similarity based on rated-item pools for improved user-based collaborative filteringabstractAn approach to user-based collaborative filtering is proposed that refines prediction of item ratings that is based on global user similarity by incorporating information derived from a more detailed user comparison made on the basis of Rated Item Pools (RIPs). The preference spectrum defined by items that a user has rated, and ranging from best-liked to most disliked items, is divided into item sets, or RIPs, which supply the basis for a fine-grained calculation of similarity between users. The RIP-based approach makes it possible for the model to take advantage of user tastes that are matched at one end of the spectrum, e.g., two users agree on favorites, without requiring complete correspondence of item ratings between user profiles. The approach improves rating prediction, as compared to a baseline that uses the global user similarity alone. It does not unduly inflate computational complexity or rely on external resources, common shortcomings of competing rating prediction methods. Cases in which the nearest neighbors are relatively dissimilar, known to be challenging for user-based collaborative filtering, demonstrate particularly substantial improvement. Performance is shown to be stable across the choice of neighborhood size, number of pools and relative pool size. Yue Shi 0002, Martha A. Larson, Alan Hanjalic |
RecSys | 2 |
| 2009 | An effective coherence measure to determine topical consistency in user-generated contentabstractWhen searching for blogs on a specific topic, information seekers prefer blogs that place a central focus on that topic over blogs whose mention of the topic is diffuse or incidental. In order to present users with better blog feed search results, we developed a measure of topical consistency that is able to capture whether or not a blog is topically focused. The measure, called the coherence score , is inspired by the genetics literature and captures the tightness of the clustering structure of a data set relative to a background collection. In a set of experiments on synthetic data, the coherence score is shown to provide a faithful reflection of topic clustering structure. The properties that make the coherence score more appropriate than lexical cohesion, a common measure of topical structure, are discussed. Retrieval experiments show that integrating the coherence score as a prior in a language modeling-based approach to blog feed search improves retrieval effectiveness. The coherence score must, however, be used judiciously in order to avoid boosting the ranking of irrelevant but topically focused blogs. To this end, we experiment with a series of weighting schemes that adjust the contribution of the coherence score according to the relevance of a blog to the user query. An appropriate weighting scheme is able to improve retrieval performance. Finally, we show that the coherence score can be reliably estimated with a sample exceeding 20 posts in size. Consistent with this finding, experiments show that the best retrieval performance is achieved if coherence scores are used when a blog contains more than 20 posts. Jiyin He, Wouter Weerkamp, Martha A. Larson, Maarten de Rijke |
Int. J. Document Anal. Recognit. | 3 |
| 2008 | Using Coherence-Based Measures to Predict Query Difficulty
Jiyin He, Martha A. Larson, Maarten de Rijke |
ECIR | 2 |
| 2008 | Term clouds as surrogates for user generated speechabstractUser generated spoken audio remains a challenge for Automatic Speech Recognition (ASR) technology and content-based audio surrogates derived from ASR-transcripts must be error robust. An investigation of the use of term clouds as surrogates for podcasts demonstrates that ASR term clouds closely approximate term clouds derived from human-generated transcripts across a range of cloud sizes. A user study confirms the conclusion that ASR-clouds are viable surrogates for depicting the content of podcasts. Manos Tsagkias, Martha A. Larson, Maarten de Rijke |
SIGIR | 2 |
| 2003 | Using syllable-based indexing features and language models to improve German spoken document retrievalabstractS.II/1217-II/1220 Martha A. Larson, Stefan Eickeler |
INTERSPEECH | 1 |
| 2002 | Exploring sub-word features and linear support vector machines for German spoken document classification
Martha A. Larson, Stefan Eickeler, Gerhard Paass, Edda Leopold, Jörg Kindermann |
INTERSPEECH | 1 |
| 2002 | Creation of an Annotated German Broadcast Speech Database for Spoken Document Retrieval
Stefan Eickeler, Martha A. Larson, Wolff Rüter, Joachim Köhler |
LREC | 2 |
| 2002 | SVM Classification Using Sequences of Phonemes and Syllables
Gerhard Paass, Edda Leopold, Martha A. Larson, Jörg Kindermann, Stefan Eickeler |
PKDD | 3 |
| 2000 | Compound splitting and lexical unit recombination for improved performance of a speech recognition system for German parliamentary speechesabstractS.945-948 Martha A. Larson, Daniel Willett, Joachim Köhler, Gerhard Rigoll |
INTERSPEECH | 1 |