EDBT 2026 Demo / reviewers in the wild / expert
Alejandro Jaimes
dblp:45/956 · also Alejandro Jaimes-Larrarte
· DBLP profile ↗
77ranked-venue papers
19as first author
10since 2021 · last 2025
0009-0003-1965-6237ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 13 first-author · 1 since 2021Artificial intelligence and machine learning · 27 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 27 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CEHA: A Dataset of Conflict Events in the Horn of AfricaabstractNatural Language Processing (NLP) of news articles can play an important role in understanding the dynamics and causes of violent conflict. Despite the availability of datasets categorizing various conflict events, the existing labels often do not cover all of the fine-grained violent conflict event types relevant to areas like the Horn of Africa. In this paper, we introduce a new benchmark dataset Conflict Events in the Horn of Africa region (CEHA) and propose a new task for identifying violent conflict events using online resources with this dataset. The dataset consists of 500 English event descriptions regarding conflict events in the Horn of Africa region with fine-grained event-type definitions that emphasize the cause of the conflict. This dataset categorizes the key types of conflict risk according to specific areas required by stakeholders in the Humanitarian-Peace-Development Nexus. Additionally, we conduct extensive experiments on two tasks supported by this dataset: Event-relevance Classification and Event-type Classification. Our baseline models demonstrate the challenging nature of these tasks and the usefulness of our dataset for model evaluations in low-resource settings. Di Lu 0003, Shihao Ran, Elizabeth M. Olson, Hemank Lamba, Aoife Cahill, Joel R. Tetreault, Alejandro Jaimes |
COLING | 8 |
| 2025 | Uchaguzi-2022: A Dataset of Citizen Reports on the 2022 Kenyan ElectionabstractOnline reporting platforms have enabled citizens around the world to collectively share their opinions and report in real time on events impacting their local communities. Systematically organizing (e.g., categorizing by attributes) and geotagging large amounts of crowdsourced information is crucial to ensuring that accurate and meaningful insights can be drawn from this data and used by policy makers to bring about positive change. These tasks, however, typically require extensive manual annotation efforts. In this paper we present Uchaguzi-2022, a dataset of 14k categorized and geotagged citizen reports related to the 2022 Kenyan General Election containing mentions of election-related issues such as official misconduct, vote count irregularities, and acts of violence. We use this dataset to investigate whether language models can assist in scalably categorizing and geotagging reports, thus highlighting its potential application in the AI for Social Good space. Roberto Mondini, Neema Kotonya, Robert L. Logan IV, Elizabeth M. Olson, Angela Oduor Lungati, Daniel Duke Odongo, Tim Ombasa, Hemank Lamba, Aoife Cahill, Joel R. Tetreault, Alejandro Jaimes |
COLING | 11 |
| 2023 | BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness MetricsabstractLiang Ma, Shuyang Cao, Robert L Logan IV, Di Lu, Shihao Ran, Ke Zhang, Joel Tetreault, Alejandro Jaimes. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shuyang Cao, Robert L. Logan IV, Di Lu 0003, Shihao Ran, Ke Zhang 0013, Joel R. Tetreault, Alejandro Jaimes |
ACL (1) | 8 |
| 2023 | Multimodal AI & LLMs for Peacekeeping and Emergency ResponseabstractWhen an emergency event, or an incident relevant for peacekeeping first occurs, getting the right information as quickly as possible is critical in saving lives. When an event is ongoing, information on what is happening can be critical in making decisions to keep people safe and take control of the particular situation unfolding. In both cases, first responders and peacekeepers have to quickly make decisions that include what resources to deploy and where. Fortunately, in most emergencies, people use social media to publicly share information. At the same time, sensor data is increasingly becoming available. But a platform to detect emergency situations and deliver the right information has to deal with ingesting thousands of noisy data points per second: sifting through and identifying relevant information, from different sources, in different formats, with varying levels of detail, in real time, so that relevant individuals and teams can be alerted at the right level and at the right time. In this talk I will describe the technical challenges in processing vast amounts of heterogeneous, noisy data in real time, highlighting the importance of interdisciplinary research and a human-centered approach to address problems in peacekeeping and emergency response. I will give specific examples specifically discussing how LLMs can be deployed at scale, including relevant future research directions in Multimedia. Alejandro Jaimes |
ACM Multimedia | 1 |
| 2022 | Temporal Event Reasoning Using Multi-source Auxiliary Learning Objectives
Xin Dong 0010, Tanay Kumar Saha, Ke Zhang 0013, Joel R. Tetreault, Alejandro Jaimes, Gerard de Melo |
ECIR (2) | 5 |
| 2022 | Mapping the Design Space of Human-AI Interaction in Text SummarizationabstractRuijia Cheng, Alison Smith-Renner, Ke Zhang, Joel Tetreault, Alejandro Jaimes-Larrarte. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ruijia Cheng, Alison Smith-Renner, Ke Zhang 0013, Joel R. Tetreault, Alejandro Jaimes |
NAACL-HLT | 5 |
| 2022 | An Exploration of Post-Editing Effectiveness in Text SummarizationabstractVivian Lai, Alison Smith-Renner, Ke Zhang, Ruijia Cheng, Wenjuan Zhang, Joel Tetreault, Alejandro Jaimes-Larrarte. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Vivian Lai, Alison Smith-Renner, Ke Zhang 0013, Ruijia Cheng, Joel R. Tetreault, Alejandro Jaimes |
NAACL-HLT | 7 |
| 2022 | AI & Public Data for Humanitarian and Emergency ResponseabstractWhen an emergency event, or an incident relevant for peacekeeping or humanitarian needs first occurs, getting the right information as quickly as possible is critical in saving lives. When an event is ongoing, information on what is happening can be critical in making decisions to keep people safe and take control of the particular situation unfolding. In both cases, first responders, peacekeepers, and others have to quickly make decisions that include what resources to deploy and where. Fortunately, in most emergencies, people use social media to publicly share information. At the same time, sensor data is increasingly becoming available. But a platform to detect emergency situations and deliver the right information has to deal with ingesting thousands of noisy data points per second: sifting through and identifying relevant information, from different sources, in different formats, with varying levels of detail, in real time, so that relevant individuals and teams can be alerted at the right level and at the right time. In this talk I will describe the technical challenges in processing vast amounts of heterogenous, noisy data in real time from the web and other sources, highlighting the importance of interdisciplinary research and a human-centered approach to address problems in humanitarian and emergency response. I will give specific examples and discuss relevant future research directions in Machine Learning, NLP, Information Retrieval, Computer Vision and other fields, highlighting the role of knowledge combined with Neural and other approaches. This talk will present an overview, and draw from some of our publications at CVPR, AAAI, EMNLP, and others. Alejandro Jaimes |
WSDM | 1 |
| 2021 | Journalistic Guidelines Aware News Image CaptioningabstractThe task of news article image captioning aims to generate descriptive and informative captions for news article images.Unlike conventional image captions that simply describe the content of the image in general terms, news image captions follow journalistic guidelines and rely heavily on named entities to describe the image content, often drawing context from the whole article they are associated with.In this work, we propose a new approach to this task, motivated by caption guidelines that journalists follow.Our approach, Journalistic Guidelines Aware News Image Captioning (JoGANIC), leverages the structure of captions to improve the generation quality and guide our representation design.Experimental results, including detailed ablation studies, on two large-scale publicly available datasets show that JoGANIC substantially outperforms state-of-the-art methods both on caption generation and named entity related metrics. Svebor Karaman, Joel R. Tetreault, Alejandro Jaimes |
EMNLP (1) | 4 |
| 2021 | Real-time Event Detection for Emergency Response TutorialabstractThe amount of public data being generated on a daily basis has grown exponentially in the last few years and continues to increase at incredible speed. Most of this data is unstructured and includes text in different formats, in different languages, from many different sources; images, video, audio, and data from sensors. A lot of that data contains information about events happening all over the world, many of which require emergency response. Detecting events in public data, in real time, is therefore critical in many applications: from getting information to first responders as quickly as possible, to creating situational awareness in such emergency situations, as getting the right information to the right places as quickly as possible is critical in saving lives. When an event is ongoing, information on what is happening can be critical in making decisions to keep people safe and take control of the particular situation unfolding. First responders have to quickly make decisions that include what resources to deploy and where. Fortunately, in most emergencies, people use social media to publicly share information. At the same time, sensor data is increasingly becoming available. In order to do this, efficient computational approaches must detect and deliver the right information to the right destination. This tutorial will cover techniques at the state-of-the art to detect events in real-time from large-scale heterogeneous sources. We will focus on NLP, Computer Vision, and Anomaly Detection techniques. We will give specific examples and discuss relevant future research directions in Machine Learning, NLP, Computer Vision and other fields relevant to real time event detection. We will also discuss applications of event detection. Alejandro Jaimes, Joel R. Tetreault |
KDD | 1 |
| 2020 | Unsupervised Detection of Sub-Events in Large Scale DisastersabstractSocial media plays a major role during and after major natural disasters (e.g., hurricanes, large-scale fires, etc.), as people “on the ground” post useful information on what is actually happening. Given the large amounts of posts, a major challenge is identifying the information that is useful and actionable. Emergency responders are largely interested in finding out what events are taking place so they can properly plan and deploy resources. In this paper we address the problem of automatically identifying important sub-events (within a large-scale emergency “event”, such as a hurricane). In particular, we present a novel, unsupervised learning framework to detect sub-events in Tweets for retrospective crisis analysis. We first extract noun-verb pairs and phrases from raw tweets as sub-event candidates. Then, we learn a semantic embedding of extracted noun-verb pairs and phrases, and rank them against a crisis-specific ontology. We filter out noisy and irrelevant information then cluster the noun-verb pairs and phrases so that the top-ranked ones describe the most important sub-events. Through quantitative experiments on two large crisis data sets (Hurricane Harvey and the 2015 Nepal Earthquake), we demonstrate the effectiveness of our approach over the state-of-the-art. Our qualitative evaluation shows better performance compared to our baseline. Chidubem Arachie, Manas Gaur, Sam Anzaroot, William Groves, Ke Zhang 0013, Alejandro Jaimes |
AAAI | 6 |
| 2020 | Multimodal Categorization of Crisis Events in Social MediaabstractRecent developments in image classification and natural language processing, coupled with the rapid growth in social media usage, have enabled fundamental advances in detecting breaking events around the world in real-time. Emergency response is one such area that stands to gain from these advances. By processing billions of texts and images a minute, events can be automatically detected to enable emergency response workers to better assess rapidly evolving situations and deploy resources accordingly. To date, most event detection techniques in this area have focused on image-only or text-only approaches, limiting detection performance and impacting the quality of information delivered to crisis response teams. In this paper, we present a new multimodal fusion method that leverages both images and texts as input. In particular, we introduce a cross-attention module that can filter uninformative and misleading components from weak modalities on a sample by sample basis. In addition, we employ a multimodal graph-based approach to stochastically transition between embeddings of different multimodal pairs during training to better regularize the learning process as well as dealing with limited training data by constructing new matched pairs from different samples. We show that our method outperforms the unimodal approaches and strong multimodal baselines by a large margin on three crisis-related tasks. Mahdi Abavisani, Liwei Wu 0001, Shengli Hu, Joel R. Tetreault, Alejandro Jaimes |
CVPR | 5 |
| 2016 | To Click or Not To Click: Automatic Selection of Beautiful Thumbnails from VideosabstractThumbnails play such an important role in online videos. As the most representative snapshot, they capture the essence of a video and provide the first impression to the viewers; ultimately, a great thumbnail makes a video more attractive to click and watch. We present an automatic thumbnail selection system that exploits two important characteristics commonly associated with meaningful and attractive thumbnails: high relevance to video content and superior visual aesthetic quality. Our system selects attractive thumbnails by analyzing various visual quality and aesthetic metrics of video frames, and performs a clustering analysis to determine the relevance to video content, thus making the resulting thumbnails more representative of the video. On the task of predicting thumbnails chosen by professional video editors, we demonstrate the effectiveness of our system against six baseline methods, using a real-world dataset of 1,118 videos collected from Yahoo Screen. In addition, we study what makes a frame a good thumbnail by analyzing the statistical relationship between thumbnail frames and non-thumbnail frames in terms of various image quality features. Our study suggests that the selection of a good thumbnail is highly correlated with objective visual quality metrics, such as the frame texture and sharpness, implying the possibility of building an automatic thumbnail selection system based on visual aesthetics. Yale Song, Miriam Redi, Jordi Vallmitjana, Alejandro Jaimes |
CIKM | 4 |
| 2016 | TGIF: A New Dataset and Benchmark on Animated GIF DescriptionabstractWith the recent popularity of animated GIFs on social media, there is need for ways to index them with rich meta-data. To advance research on animated GIF understanding, we collected a new dataset, Tumblr GIF (TGIF), with 100K animated GIFs from Tumblr and 120K natural language descriptions obtained via crowdsourcing. The motivation for this work is to develop a testbed for image sequence description systems, where the task is to generate natural language descriptions for animated GIFs or video clips. To ensure a high quality dataset, we developed a series of novel quality controls to validate free-form text input from crowd-workers. We show that there is unambiguous association between visual content and natural language descriptions in our dataset, making it an ideal benchmark for the visual content captioning task. We perform extensive statistical analyses to compare our dataset to existing image and video description datasets. Next, we provide baseline results on the animated GIF description task, using three representative techniques: nearest neighbor, statistical machine translation, and recurrent neural networks. Finally, we show that models fine-tuned from our animated GIF description dataset can be helpful for automatic movie description. Yuncheng Li, Yale Song, Liangliang Cao, Joel R. Tetreault, Larry Goldberg, Alejandro Jaimes, Jiebo Luo 0001 |
CVPR | 6 |
| 2016 | How to Compete Online for News Audience: Modeling Words that Attract ClicksabstractHeadlines are particularly important for online news outlets where there are many similar news stories competing for users' attention. Traditionally, journalists have followed rules-of-thumb and experience to master the art of crafting catchy headlines, but with the valuable resource of large-scale click-through data of online news articles, we can apply quantitative analysis and text mining techniques to acquire an in-depth understanding of headlines. In this paper, we conduct a large-scale analysis and modeling of 150K news articles published over a period of four months on the Yahoo home page. We define a simple method to measure click-value of individual words, and analyze how temporal trends and linguistic attributes affect click-through rate (CTR). We then propose a novel generative model, headline click-based topic model (HCTM), that extends latent Dirichlet allocation (LDA) to reveal the effect of topical context on the click-value of words in headlines. HCTM leverages clicks in aggregate on previously published headlines to identify words for headlines that will generate more clicks in the future. We show that by jointly taking topics and clicks into account we can detect changes in user interests within topics. We evaluate HCTM in two different experimental settings and compare its performance with ALDA (adapted LDA), LDA, and TextRank. The first task, full headline, is to retrieve full headline used for a news article given the body of news article. The second task, good headline, is to specifically identify words in the headline that have high click values for current news audience. For full headline task, our model performs on par with ALDA, a state-of-the art web-page summarization method that utilizes click-through information. For good headline task, which is of more practical importance to both individual journalists and online news outlets, our model significantly outperforms all other comparative methods. Joon Hee Kim, Amin Mantrach, Alejandro Jaimes, Alice Oh |
KDD | 3 |
| 2016 | Humor in Collective Discourse: Unsupervised Funniness Detection in the New Yorker Cartoon Caption Contest
Dragomir R. Radev, Amanda Stent, Joel R. Tetreault, Aasish Pappu, Aikaterini Iliakopoulou, Agustin Chanfreau, Paloma de Juan, Jordi Vallmitjana, Alejandro Jaimes, Rahul Jha, Robert Mankoff |
LREC | 9 |
| 2016 | Mouse Activity as an Indicator of Interestingness in VideoabstractAutomatic detection of interesting moments in video has many real-world applications such as video summarization and efficient online video browsing. In this paper, we present a lightweight and scalable solution to this problem based on user mouse activity while watching video. Unlike previous approaches that analyze video content to infer the interestingness, we leverage the implicit user feedback obtained from thousands of online video watching sessions. This makes our method computationally efficient and scalable to billions of videos. Most importantly, our approach can handle a variety of video genres because we make no assumption on what constitutes interestingness: we let the crowd tell us through their mouse activity. By analyzing 106,212 user sessions collected from a popular online video website, we show that mouse activity is highly indicative of interestingness, and that our approach has competitive performance to several state-of-the-art methods. Gloria Zen, Paloma de Juan, Yale Song, Alejandro Jaimes |
ICMR | 4 |
| 2016 | Predicting celebrity attendees at public events using stock photo metadata
Xin Shuai, Neil O'Hare, Luca Maria Aiello, Alejandro Jaimes |
Multim. Tools Appl. | 4 |
| 2015 | A Large-Scale Study of User Image Search Behavior on the WebabstractIn this study, we analyze user image search behavior from a large-scale Yahoo! Image Search query log, based on the hypothesis that behavior is dependent on query type. We categorize queries using two orthogonal taxonomies (subject-based and facet-based) and identify important query types at the intersection of these taxonomies. We study user search behavior on a large-scale set of search sessions for each query type, examining characteristics of sessions, query reformulation patterns, click patterns, and page view patterns. We identify important behavioral differences across query types, in particular showing that some query types are more exploratory, while others correspond to focused search. We also supplement our study with a survey to link the behavioral differences to users' intent. Our findings shed light on the importance of considering query categories to better understand user behavior on image search platforms. Jaimie Yejean Park, Neil O'Hare, Rossano Schifanella, Alejandro Jaimes, Chin-Wan Chung |
CHI | 4 |
| 2015 | Video co-summarization: Video summarization by visual co-occurrenceabstractWe present video co-summarization, a novel perspective to video summarization that exploits visual co-occurrence across multiple videos. Motivated by the observation that important visual concepts tend to appear repeatedly across videos of the same topic, we propose to summarize a video by finding shots that co-occur most frequently across videos collected using a topic keyword. The main technical challenge is dealing with the sparsity of co-occurring patterns, out of hundreds to possibly thousands of irrelevant shots in videos being considered. To deal with this challenge, we developed a Maximal Biclique Finding (MBF) algorithm that is optimized to find sparsely co-occurring patterns, discarding less co-occurring patterns even if they are dominant in one video. Our algorithm is parallelizable with closed-form updates, thus can easily scale up to handle a large number of videos simultaneously. We demonstrate the effectiveness of our approach on motion capture and self-compiled YouTube datasets. Our results suggest that summaries generated by visual co-occurrence tend to match more closely with human generated summaries, when compared to several popular unsupervised techniques. Wen-Sheng Chu, Yale Song, Alejandro Jaimes |
CVPR | 3 |
| 2015 | TVSum: Summarizing web videos using titlesabstractVideo summarization is a challenging problem in part because knowing which part of a video is important requires prior knowledge about its main topic. We present TVSum, an unsupervised video summarization framework that uses title-based image search results to find visually important shots. We observe that a video title is often carefully chosen to be maximally descriptive of its main topic, and hence images related to the title can serve as a proxy for important visual concepts of the main topic. However, because titles are free-formed, unconstrained, and often written ambiguously, images searched using the title can contain noise (images irrelevant to video content) and variance (images of different topics). To deal with this challenge, we developed a novel co-archetypal analysis technique that learns canonical visual concepts shared between video and images, but not in either alone, by finding a joint-factorial representation of two data sets. We introduce a new benchmark dataset, TVSum50, that contains 50 videos and their shot-level importance scores annotated via crowdsourcing. Experimental results on two datasets, SumMe and TVSum50, suggest our approach produces superior quality summaries compared to several recently proposed approaches. Yale Song, Jordi Vallmitjana, Amanda Stent, Alejandro Jaimes |
CVPR | 4 |
| 2015 | Construction and evaluation of ontological tag trees
Chetan Kumar Verma, Vijay Mahadevan, Nikhil Rasiwasia, Gaurav Aggarwal, Alejandro Jaimes, Sujit Dey |
Expert Syst. Appl. | 6 |
| 2015 | Algorithms and criteria for diversification of news article comments
Giorgos Giannopoulos, Marios Koniaris, Ingmar Weber, Alejandro Jaimes, Timos K. Sellis |
J. Intell. Inf. Syst. | 4 |
| 2014 | 6 Seconds of Sound and Vision: Creativity in Micro-videosabstractThe notion of creativity, as opposed to related concepts such as beauty or interestingness, has not been studied from the perspective of automatic analysis of multimedia content. Meanwhile, short online videos shared on social media platforms, or micro-videos, have arisen as a new medium for creative expression. In this paper we study creative micro-videos in an effort to understand the features that make a video creative, and to address the problem of automatic detection of creative content. Defining creative videos as those that are novel and have aesthetic value, we conduct a crowdsourcing experiment to create a dataset of over 3, 800 micro-videos labelled as creative and non-creative. We propose a set of computational features that we map to the components of our definition of creativity, and conduct an analysis to determine which of these features correlate most with creative video. Finally, we evaluate a supervised approach to automatically detect creative video, with promising results, showing that it is necessary to model both aesthetic value and novelty to achieve optimal classification accuracy. Miriam Redi, Neil O'Hare, Rossano Schifanella, Michele Trevisiol, Alejandro Jaimes |
CVPR | 5 |
| 2014 | Cold-start news recommendation with domain-dependent browse graphabstractOnline social networks and mash-up services create opportunities to connect different web services otherwise isolated. Specifically in the case of news, users are very much exposed to news articles while performing other activities, such as social networking or web searching. Browsing behavior aimed at the consumption of news, especially in relation to the visits coming from other domains, has been mainly overlooked in previous work. To address that, we build a BrowseGraph out of the collective browsing traces extracted from a large viewlog of Yahoo News (0.5B entries), and we define the ReferrerGraph as its subgraph induced by the sessions with the same referrer domain. The structural and temporal properties of the graph show that browsing behavior in news is highly dependent on the referrer URL of the session, in terms of type of content consumed and time of consumption. We build on this observation and propose a news recommender that addresses the cold-start problem: given a user landing on a page of the site for the first time, we aim to predict the page she will visit next. We compare 24 flavors of recommenders belonging to the families of content-based, popularity-based, and browsing-based models. We show that the browsing-based recommender that takes into account the referrer URL is the best performing, achieving a prediction accuracy of 48% in conditions of heavy data sparsity. Michele Trevisiol, Luca Maria Aiello, Rossano Schifanella, Alejandro Jaimes |
RecSys | 4 |
| 2014 | Random walks based modularity: application to semi-supervised learningabstractAlthough criticized for some of its limitations, modularity remains a standard measure for analyzing social networks. Quantifying the statistical surprise in the arrangement of the edges of the network has led to simple and powerful algorithms. However, relying solely on the distribution of edges instead of more complex structures such as paths limits the extent of modularity. Indeed, recent studies have shown restrictions of optimizing modularity, for instance its resolution limit. We introduce here a novel, formal and well-defined modularity measure based on random walks. We show how this modularity can be computed from paths induced by the graph instead of the traditionally used edges. We argue that by computing modularity on paths instead of edges, more informative features can be extracted from the network. We verify this hypothesis on a semi-supervised classification procedure of the nodes in the network, where we show that, under the same settings, the features of the random walk modularity help to classify better than the features of the usual modularity. Additionally, the proposed approach outperforms the classical label propagation procedure on two data sets of labeled social networks. Robin Devooght, Amin Mantrach, Ilkka Kivimäki, Hugues Bersini, Alejandro Jaimes, Marco Saerens |
WWW | 5 |
| 2014 | A time-based collective factorization for topic discovery and monitoring in newsabstractDiscovering and tracking topic shifts in news constitutes a new challenge for applications nowadays. Topics evolve,emerge and fade, making it more difficult for the journalist -or the press consumer- to decrypt the news. For instance, the current Syrian chemical crisis has been the starting point of the UN Russian initiative and also the revival of the US France alliance. A topical mapping representing how the topics evolve in time would be helpful to contextualize information. As far as we know, few topic tracking systems can provide such temporal topic connections. In this paper, we introduce a novel framework inspired from Collective Factorization for online topic discovery able to connect topics between different time-slots. The framework learns jointly the topics evolution and their time dependencies. It offers the user the ability to control, through one unique hyper-parameter, the tradeoff between the past accumulated knowledge and the current observed data. We show, on semi-synthetic datasets and on Yahoo News articles, that our method is competitive with state-of-the-art techniques while providing a simple way to monitor topics evolution (including emerging and disappearing topics). Carmen Vaca, Amin Mantrach, Alejandro Jaimes, Marco Saerens |
WWW | 3 |
| 2013 | Search behaviour on photo sharing platformsabstractThe behaviour, goals, and intentions of users while searching for images in large scale online collections are not well understood, with image search log analysis providing limited insights, in part because they tend only to have access to user search and result click information. In this paper we study user search behaviour in a large photo-sharing platform, analyzing all user actions during search sessions (i.e. including post result-click pageviews). Search accounts for a significant part of user interactions with such platforms, and we show differences between the queries issued on such platforms and those on general image search. We show that search behaviour is influenced by the query type, and also depends on the user. Finally, we analyse how users behave when they reformulate their queries, and develop URL class prediction models for image search, showing that query-specific models significantly outperform query-agnostic models. The insights provided in this paper are intended as a launching point for the design of better interfaces and ranking models for image search. Silviu Maniu, Neil O'Hare, Luca Maria Aiello, Luca Chiarandini, Alejandro Jaimes |
ICME | 5 |
| 2013 | Leveraging Browsing Patterns for Topic Discovery and Photostream Recommendation
Luca Chiarandini, Przemyslaw A. Grabowicz, Michele Trevisiol, Alejandro Jaimes |
ICWSM | 4 |
| 2013 | Cultural Dimensions in Twitter: Time, Individualism and Power
Ruth Olimpia Garcia Gavilanes, Daniele Quercia, Alejandro Jaimes |
ICWSM | 3 |
| 2013 | Automatic selection of social media responses to newsabstractSocial media responses to news have increasingly gained in importance as they can enhance a consumer's news reading experience, promote information sharing and aid journalists in assessing their readership's response to a story. Given that the number of responses to an online news article may be huge, a common challenge is that of selecting only the most interesting responses for display. This paper addresses this challenge by casting message selection as an optimization problem. We define an objective function which jointly models the messages' utility scores and their entropy. We propose a near-optimal solution to the underlying optimization problem, which leverages the submodularity property of the objective function. Our solution first learns the utility of individual messages in isolation and then produces a diverse selection of interesting messages by maximizing the defined objective function. The intuitions behind our work are that an interesting selection of messages contains diverse, informative, opinionated and popular messages referring to the news article, written mostly by users that have authority on the topic. Our intuitions are embodied by a rich set of content, social and user features capturing the aforementioned aspects. We evaluate our approach through both human and automatic experiments, and demonstrate it outperforms the state of the art. Additionally, we perform an in-depth analysis of the annotated ``interesting'' responses, shedding light on the subjectivity around the selection process and the perception of interestingness. Tadej Stajner, Bart Thomee, Ana-Maria Popescu, Marco Pennacchiotti, Alejandro Jaimes |
KDD | 5 |
| 2013 | Analyzing Favorite Behavior in Flickr
Marek Lipczak, Michele Trevisiol, Alejandro Jaimes |
MMM (1) | 3 |
| 2013 | Competition-based networks for expert findingabstractFinding experts in question answering platforms has important applications, such as question routing or identification of best answers. Addressing the problem of ranking users with respect to their expertise, we propose Competition-Based Expertise Networks (CBEN), a novel community expertise network structure based on the principle of competition among the answerers of a question. We evaluate our approach on a very large dataset from Yahoo! Answers using a variety of centrality measures. We show that it outperforms state-of-the-art network structures and, unlike previous methods, is able to consistly outperform simple metrics like best answer count. We also analyse question answering forums in Yahoo! Answers, and show that they can be characterised by factual or subjective information seeking behavior, social discussions and the conducting of polls or surveys. We find that the ability to identify experts greatly depends on the type of forum, which is directly reflected in the structural properties of the expertise networks. Çigdem Aslay, Neil O'Hare, Luca Maria Aiello, Alejandro Jaimes |
SIGIR | 4 |
| 2013 | Distinguishing topical and social groups based on common identity and bond theoryabstractSocial groups play a crucial role in social media platforms because they form the basis for user participation and engagement. Groups are created explicitly by members of the community, but also form organically as members interact. Due to their importance, they have been studied widely (e.g., community detection, evolution, activity, etc.). One of the key questions for understanding how such groups evolve is whether there are different types of groups and how they differ. In Sociology, theories have been proposed to help explain how such groups form. In particular, the common identity and common bond theory states that people join groups based on identity (i.e., interest in the topics discussed) or bond attachment (i.e., social relationships). The theory has been applied qualitatively to small groups to classify them as either topical or social. We use the identity and bond theory to define a set of features to classify groups into those two categories. Using a dataset from Flickr, we extract user-defined groups and automatically-detected groups, obtained from a community detection algorithm. We discuss the process of manual labeling of groups into social or topical and present results of predicting the group label based on the defined features. We directly validate the predictions of the theory showing that the metrics are able to forecast the group type with high accuracy. In addition, we present a comparison between declared and detected groups along topicality and sociality dimensions. Przemyslaw A. Grabowicz, Luca Maria Aiello, Víctor M. Eguíluz, Alejandro Jaimes |
WSDM | 4 |
| 2013 | Sensing Trending Topics in TwitterabstractOnline social and news media generate rich and timely information about real-world events of all kinds. However, the huge amount of data available, along with the breadth of the user base, requires a substantial effort of information filtering to successfully drill down to relevant topics and events. Trending topic detection is therefore a fundamental building block to monitor and summarize information originating from social sources. There are a wide variety of methods and variables and they greatly affect the quality of results. We compare six topic detection methods on three Twitter datasets related to major events, which differ in their time scale and topic churn rate. We observe how the nature of the event considered, the volume of activity over time, the sampling procedure and the pre-processing of the data all greatly affect the quality of detected topics, which also depends on the type of detection method used. We find that standard natural language processing techniques can perform well for social streams on very focused topics, but novel techniques designed to mine the temporal distribution of concepts are needed to handle more heterogeneous streams containing multiple stories evolving in parallel. One of the novel topic detection methods we propose, based on -grams cooccurrence and topic ranking, consistently achieves the best performance across all these conditions, thus being more reliable than other state-of-the-art techniques. Luca Maria Aiello, Georgios Petkos, Carlos J. Martín-Dancausa, David P. A. Corney, Symeon Papadopoulos, Ryan Skraba, Ayse Göker, Ioannis Kompatsiaris, Alejandro Jaimes |
IEEE Trans. Multim. | 9 |
| 2012 | Discovering Social Photo Navigation PatternsabstractIn general, user browsing behavior has been examined within specific tasks (e.g., search), or in the context of particular web sites or services ( e.g., in shopping sites). However, with the growth of social networks and the proliferation of many different types of web services ( e.g., news aggregators, blogs, forums, etc.), the web can be viewed as an ecosystem in which a user's actions in a particular web service may be influenced by the service she arrived from ( e.g., are users browsing patterns similar if they arrive at a website via search or via links in aggregators?). In particular, since photos in services like Flickr are used extensively throughout the web, it is common for visitors to the site to arrive via links in many different types of web sites. In this paper, we depart from the hypothesis that visitors to social sites such as Flickr behave differently depending on where they come from. For this purpose, we analyze a large sample of Flickr user logs to discover social photo navigation patterns. More specifically, we classify pages within Flickr into different categories ( e.g., "add a friend page", "single photo page," etc.), and by clustering sessions discover important differences in social photo navigation that manifest themselves depending on the type of site users visit before visiting Flickr. Our work examines photo navigation patterns in Flickr for the first time taking into account the referrer domain. Our analysis is useful in that it can contribute to a better understanding of how people use photo services like Flickr, and it can be used to inform the design of user modeling and recommendation algorithms, among others. Luca Chiarandini, Michele Trevisiol, Alejandro Jaimes |
ICME | 3 |
| 2012 | Lin-spiration: using a mixture of spiral and linear visualization layouts to explore time seriesabstractTime series data is pervasive in many domains and interactive visualization of such data is useful for a wide range of tasks including analysis and prediction. In spite of the importance of visualizing time series data and the fact that time series data is often easily interpretable, traditional approaches are either very simple and limited, or are aimed at domain experts. In this paper, we propose a novel interactive visualization paradigm for exploring and comparing multiple sets of time series data. In particular, we propose a focus+context approach, where a "focus" segment of a time series is zoomed into and visualized using a linear layout at one scale, while the remaining segments of the time series (i.e., the context) are visualized using spiral data layouts. Our paradigm allows the user to dynamically select and compare different sections of each time series independently, facilitating the exploration of time series data in a fun and engaging way. Eduardo Graells-Garrido, Alejandro Jaimes |
IUI | 2 |
| 2012 | A human-centered perspective on multimedia data science: tutorial overviewabstractThis tutorial focuses on the analysis of user behavior in multimedia through large-scale data analysis. This includes discovering and leveraging search and navigation patterns, understanding how elements of interaction impact behavior, and how we can use controlled experiments in combination with user studies and other techniques to gain insights into human behavior with a particular emphasis on multimedia, particularly in the context of social media. Alejandro Jaimes |
ACM Multimedia | 1 |
| 2012 | Predicting participants in public events using stock photosabstractPictures taken by journalists for distribution and for inclusion in stock photo collections are often enriched with metadata. One key aspect of such photos is that they focus largely on events and feature celebrities and other public figures. They may provide interesting insights into how such public figures are related to each other in terms of the events they attend, and in their social proximity in terms of how often they are photographed together. In this paper, we study a corpus of approximately 9 million stock photographs taken over a 10 year period and, using their metadata, we extract a social network from co-appearance of public figures in events depicted in the photographs. We exploit this latent social information and combine it with the rich image metadata to explore the possibility of predicting attendees at future events, showing promising performance for this task. Neil O'Hare, Luca Maria Aiello, Alejandro Jaimes |
ACM Multimedia | 3 |
| 2012 | PRiSMA: searching images in parallelabstractPRiSMA is an image search application for tablet and desktop devices intended to facilitate and promote the searching of images in parallel. With an intuitive user interface, users can branch their queries into multiple horizontal sliding strips to simultaneously explore different perspectives of large image collections (e.g., colors, geographical location or topic). Strips can be easily created, tailored, merged, and removed, allowing users to effectively perform multiple queries and manage the results in a dynamic and orderly fashion. With PRiSMA we aim to explore the potential and limitations of parallel image search from a user perspective. Pancho Tolchinsky, Luca Chiarandini, Alejandro Jaimes |
ACM Multimedia | 3 |
| 2012 | Image ranking based on user browsing behaviorabstractRanking of images is difficult because many factors determine their importance (e.g., popularity, quality, entertainment value, context, etc.). In social media platforms, ranking also depends on social interactions and on the visibility of the images both inside and outside those platforms. In this context, the application of standard ranking methods is not clearly understood, and neither are the subtleties associated with taking into account social interaction, internal, and external factors. In this paper, we use a large Flickr dataset and investigate these factors by performing an in-depth analysis of several ranking algorithms using both internal (i.e., within Flickr) and external (i.e., links from outside of Flickr) factors. We analyze rankings given by common metrics used in image retrieval (e.g., number of favorites), and compare them with metrics based on page views (e.g., time spent, number of views). In addition, we represent users' navigation by a graph and combine session models with some of these metrics, comparing with PageRank and BrowseRank. Our experiments show significant differences between the rankings, providing insights on the impact of social interactions, internal, and external factors in image ranking. Michele Trevisiol, Luca Chiarandini, Luca Maria Aiello, Alejandro Jaimes |
SIGIR | 4 |
| 2012 | Diversifying User Comments on News Articles
Giorgos Giannopoulos, Ingmar Weber, Alejandro Jaimes, Timos K. Sellis |
WISE | 3 |
| 2012 | Correlating financial time series with micro-blogging activityabstractWe study the problem of correlating micro-blogging activity with stock-market events, defined as changes in the price and traded volume of stocks. Specifically, we collect messages related to a number of companies, and we search for correlations between stock-market events for those companies and features extracted from the micro-blogging messages. The features we extract can be categorized in two groups. Features in the first group measure the overall activity in the micro-blogging platform, such as number of posts, number of re-posts, and so on. Features in the second group measure properties of an induced interaction graph, for instance, the number of connected components, statistics on the degree distribution, and other graph-based properties. Eduardo J. Ruiz, Vagelis Hristidis, Carlos Castillo 0001, Aristides Gionis, Alejandro Jaimes |
WSDM | 5 |
| 2012 | Societally connected multimedia across culturesabstractThe advance of the Internet in the past decade has radically changed the way people communicate and collaborate with each other. Physical distance is no more a barrier in online social networks, but cultural differences (at the individual, community, as well as societal levels) still govern human-human interactions and must be considered and leveraged in the online world. The rapid deployment of high-speed Internet allows humans to interact using a rich set of multimedia data such as texts, pictures, and videos. This position paper proposes to define a new research area called ‘connected multimedia’, which is the study of a collection of research issues of the super-area social media that receive little attention in the literature. By connected multimedia, we mean the study of the social and technical interactions among users, multimedia data, and devices across cultures and explicitly exploiting the cultural differences. We justify why it is necessary to bring attention to this new research area and what benefits of this new research area may bring to the broader scientific research community and the humanity. Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang, Noshir S. Contractor, Alex Hauptmann 0001, Alejandro Jaimes, Wanqing Li 0001, Alexander C. Loui, Tao Mei 0001, Nicu Sebe, Yonghong Tian 0001, Vincent S. Tseng, Qing Wang 0015, Changsheng Xu, Shiwen Yu |
J. Zhejiang Univ. Sci. C | 7 |
| 2011 | Do all birds tweet the same?: characterizing twitter around the worldabstractSocial media services have spread throughout the world in just a few years. They have become not only a new source of information, but also new mechanisms for societies world-wide to organize themselves and communicate. Therefore, social media has a very strong impact in many aspects -- at personal level, in business, and in politics, among many others. In spite of its fast adoption, little is known about social media usage in different countries, and whether patterns of behavior remain the same or not. To provide deep understanding of differences between countries can be useful in many ways, e.g.: to improve the design of social media systems (which features work best for which country?), and influence marketing and political campaigns. Moreover, this type of analysis can provide relevant insight into how societies might differ. In this paper we present a summary of a large-scale analysis of Twitter for an extended period of time. We analyze in detail various aspects of social media for the ten countries we identified as most active. We collected one year's worth of data and report differences and similarities in terms of activity, sentiment, use of languages, and network structure. To the best of our knowledge, this is the first on-line social network study of such characteristics. Barbara Poblete, Ruth Olimpia Garcia Gavilanes, Marcelo Mendoza, Alejandro Jaimes |
CIKM | 4 |
| 2011 | Who uses web search for what: and howabstractWe analyze a large query log of 2.3 million anonymous registered users from a web-scale U.S. search engine in order to jointly analyze their on-line behavior in terms of who they might be (demographics), what they search for (query topics), and how they search (session analysis). We examine basic demographics from registration information provided by the users, augmented with U.S. census data, analyze basic session statistics, classify queries into types (navigational, informational, transactional) based on click entropy, classify queries into topic categories, and cluster users based on the queries they issued. We then examine the resulting clusters in terms of demographics and search behavior. Our analysis of the data suggests that there are important differences in search behavior across different demographic groups in terms of the topics they search for, and how they search (e.g., white conservatives are those likely to have voted republican, mostly white males, who search for business, home, and gardening related topics; Baby Boomers tend to be primarily interested in Finance and a large fraction of their sessions consist of simple navigational queries related to online banking, etc.). Finally, we examine regional search differences, which seem to correlate with differences in local industries (e.g., gambling related queries are highest in Las Vegas and lowest in Salt Lake City; searches related to actors are about three times higher in L.A. than in any other region). Ingmar Weber, Alejandro Jaimes |
WSDM | 2 |
| 2011 | Social Network Analysis and Mining for Business ApplicationsabstractSocial network analysis has gained significant attention in recent years, largely due to the success of online social networking and media-sharing sites, and the consequent availability of a wealth of social network data. In spite of the growing interest, however, there is little understanding of the potential business applications of mining social networks. While there is a large body of research on different problems and methods for social network mining, there is a gap between the techniques developed by the research community and their deployment in real-world applications. Therefore the potential business impact of these techniques is still largely unexplored. In this article we use a business process classification framework to put the research topics in a business context and provide an overview of what we consider key problems and techniques in social network analysis and mining from the perspective of business applications. In particular, we discuss data acquisition and preparation, trust, expertise, community structure, network dynamics, and information propagation. In each case we present a brief overview of the problem, describe state-of-the art approaches, discuss business application examples, and map each of the topics to a business process classification framework. In addition, we provide insights on prospective business applications, challenges, and future research directions. The main contribution of this article is to provide a state-of-the-art overview of current techniques while providing a critical perspective on business applications of social network analysis and mining. Francesco Bonchi, Carlos Castillo 0001, Aristides Gionis, Alejandro Jaimes |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2010 | Demographic information flowsabstractIn advertising and content relevancy prediction it is important to understand whether, over time, information that reaches one demographic group spreads to others. In this paper we analyze the query log of a large U.S. web search engine to determine whether the same queries are performed by different demographic groups at different times, particularly when there are query bursts. We obtain aggregate demographic features from user-provided registration information (gender, birth year, ZIP code), U.S. census data, and election results. Given certain queries, we examine trends (from high to low and vice versa) and changes in the statistical spread of the demographic features of users that issue the queries over time periods that include query bursts. Our analysis shows that for certain types of queries (movies and news) distinct demographic groups perform searches at different times, suggesting that information related to such queries flows between them. Queries of movie titles, for instance, tend to be issued first by young and then by older users, where a sudden jump in age occurs upon the movie's release. To the best of our knowledge, this is the first time this problem has been studied using search query logs. Ingmar Weber, Alejandro Jaimes |
CIKM | 2 |
| 2010 | Human-centered multimedia systems: tutorial overviewabstractThis tutorial will focus on technical analysis and interaction techniques formulated from the perspective of key human factors in a user-centered approach to developing multimedia systems. The tutorial will take a holistic view on the research issues and applications of Human-Centered Systems, focusing on four main areas: (1) multimodal interaction: visual (body, gaze, gesture); (2) image indexing and retrieval: user behavior, context modeling, cultural issues, and machine learning for user-centric approaches; (3) multimedia data: conceptual analysis at different levels (feature, cognitive, and affective); and (4) sources of contextual information and case studies in multi-camera networks. Nicu Sebe, Alejandro Jaimes, Hamid K. Aghajan |
ACM Multimedia | 2 |
| 2010 | Sonify your face: facial expressions for sound generationabstractWe present a novel visual creativity tool that automatically recognizes facial expressions and tracks facial muscle movements in real time to produce sounds. The facial expression recognition module detects and tracks a face and outputs a feature vector of motions of specific locations in the face. The feature vector is used as input to a Bayesian network which classifies facial expressions into several categories (e.g., angry, disgusted, happy, etc.). The classification results are used along with the feature vector to generate a combination of sounds that change in real time depending on the person's facial expressions. We explain the artistic motivation behind the work, the basic components of our tool, and possible applications in the arts (performance, installation) and in the medical domain. Finally, we report on the experience of approximately 25 users of our system at a conference demonstration session, of 9 participants in a pilot study to assess the system's usability, and discuss our experience installing the work at an important digital arts festival (RE-NEW 2009). Roberto Valenti, Alejandro Jaimes, Nicu Sebe |
ACM Multimedia | 2 |
| 2010 | Will recommenders kill search?: recommender systems - an industry perspectiveabstractAt the 2010 annual ACM Conference on Recommender Systems (RecSys 2010) a panel addressed emerging topics regarding recommender systems as a whole and specifically their role in industry. This report summarizes answers from a distinguished group of industry leaders representing different industries in which recommender systems are highly relevant. Panel members discuss questions regarding the role of recommender systems in their own industry area, killer applications, opportunities, and future directions. Ido Guy, Alejandro Jaimes, Pau Agulló, Pat Moore, Palash Nandy, Chahab Nastar, Henrik Schinzel |
RecSys | 2 |
| 2010 | Workshop on the practical use of recommender systems algorithms & technologyabstractUser modeling, adaptation, and personalization techniques have hit the mainstream. The explosion of social network websites, on-line user-generated content platforms, and the tremendous growth in computational power of mobile devices are generating incredibly large amounts of user data, and an increasing desire of users to "personalize" (their desktop, e-mail, news site, phone). The potential value of personalization has become clear both as a commodity for the benefit or enjoyment of end-users, and as an enabler of new or better services -- a strategic opportunity to enhance and expand businesses. An exciting characteristic of recommender systems is that they draw the interest of industry and businesses while posing very interesting research and scientific challenges. Jérôme Picault, Dimitre Kostadinov, Pablo Castells, Alejandro Jaimes |
RecSys | 4 |
| 2009 | ClusTR: Exploring Multivariate Cluster Correlations and Topic Trends
Luigi Di Caro, Alejandro Jaimes |
ECML/PKDD (2) | 2 |
| 2009 | Integration of Context and Content for Multimedia Management: An Introduction to the Special IssueabstractThe 11 papers in this special issue focus on the integration of context and content for multimedia management. Jiebo Luo 0001, Alan Hanjalic, Qi Tian 0001, Alejandro Jaimes |
IEEE Trans. Multim. | 4 |
| 2009 | Introduction to the special section for the best papers of ACM multimedia 2008abstractintroduction Share on Introduction to the special section for the best papers of ACM multimedia 2008 Authors: K. Selçuk Candan Arizona State University, USA Arizona State University, USAView Profile , Alberto Del Bimbo Università degli Studi di Firenze, Italy Università degli Studi di Firenze, ItalyView Profile , Carsten Griwodz Simula Research Laboratory, Norway Simula Research Laboratory, NorwayView Profile , Alejandro Jaimes Telefonica Research, Spain Telefonica Research, SpainView Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 5Issue 3August 2009 Article No.: 18pp 1–3https://doi.org/10.1145/1556134.1556135Published:14 August 2009Publication History 0citation329DownloadsMetricsTotal Citations0Total Downloads329Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access K. Selçuk Candan, Alberto Del Bimbo, Carsten Griwodz, Alejandro Jaimes |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2008 | Facial expression recognition as a creative interfaceabstractWe present an audiovisual creativity tool that automatically recognizes facial expressions in real time, producing sounds in combination with images. The facial expression recognition component detects and tracks a face and outputs a feature vector of motions of specific locations in the face. The feature vector is used as input to a Bayesian network which classifies facial expressions into several categories (e.g., angry, disgusted, happy, etc.). The classification results are used along with the feature vector to generate a combination of sounds and images that change in real time depending on the person's facial expressions. We explain the basic components of our tool and several possible applications in the arts (performance, installation) and medical domains. Roberto Valenti, Alejandro Jaimes, Nicu Sebe |
IUI | 2 |
| 2008 | Graphical representation of meetings on mobile devicesabstractThe AMIDA Mobile Meeting Assistant is a system that allows remote participants to attend a meeting through a mobile device. The system improves the engagement in the meeting of the remote participants with respect to voice-only solutions thanks to the use of visual annotations and the capture of slides. The visual focus of attention of meeting participants and other annotations serve to reconstruct a 2D or a 3D representation of the meeting on a mobile device (smart phone). A first version of the system has been implemented, and feedback from a user study and from industrial partners shows that the Mobile Meeting Assistant's functionalities are positively appreciated, and sets priorities for future developments. Lukas Matena, Alejandro Jaimes, Andrei Popescu-Belis |
Mobile HCI | 2 |
| 2008 | 3rd international workshop on human-centered computing (HCC '08)abstractIn this workshop summary we describe the motivation for continued discussion in Human-Centered Computing, giving an outline of the articles presented at the workshop, its expected outcomes, and future activities. We emphasize the reasoning behind a non-traditional format for the workshop, which builds on the previous workshops on "Human-Centered Multimedia" held in conjunction with ACM Multimedia 2007 and 2006. Alejandro Jaimes, Daniela Nicklas 0001, Nicu Sebe |
ACM Multimedia | 1 |
| 2008 | Connecting artists and scientists in multimedia researchabstractHistorically, the ACM Multimedia Conference is split into a "technical" program and an "arts" program. These programs sometimes seem completely separate from one another, victims of a "semantic gap" between disciplines. The goal of this panel is to create a space in which scientists learn from artists, and arts from science. We need to discover new connections between modalities of research. In order to create the most exciting and powerful future forms of interactive multimedia systems, the ones that will create the most beneficial broader impact on humanity, we need to foster new collaborations between artists and scientists. This panel seeks to bridge the great divide of language and communities that has fragmented us, creating a new space for developing connections between the arts and sciences of multimedia research, as embodied through the artists and scientists of ACM Multimedia. The goal is to make this conference a premier site for catalyzing emergent connections. Andruid Kerne, Ron Wakkary, Frank Nack, Amanda Steggell, Alejandro Jaimes, K. Selçuk Candan, Alberto Del Bimbo, Pamela Jennings, Aleksandra Dulic |
ACM Multimedia | 5 |
| 2008 | A Study on the Granularity of User Modeling for Tag PredictionabstractOne of the characteristics of tag prediction mechanisms is that, typically, all user models are constructed with the same granularity. In this paper we hypothesize and empirically demonstrate that in order to increase tag prediction accuracy, the granularity of each user model has to be adapted to the level of usage of each particular user. We have constructed user models for tag prediction using association rules in Bibsonomy, a popular social bookmark and publication sharing system, at three granularity levels: (1) canonical, (2) stereotypical and (3) individual. Our experiments show that prediction accuracy improves if the level of granularity matches the level of participation of the user in the community (i.e., amount of tagging in Bibsonomy). Enrique Frías-Martínez, Manuel Cebrián, Alejandro Jaimes |
Web Intelligence | 3 |
| 2007 | Human-centered multimedia systems: tutorial overviewabstractThis tutorial will focus on technical analysis and interaction techniques formulated from the perspective of key human factors in a user-centered approach to developing multimedia systems. The tutorial will take a holistic view on the research issues and applications of Human-Centered Systems, focusing on four main areas: (1) multimodal interaction: visual (body, gaze, gesture); (2) image indexing and retrieval: user behavior, context modeling, cultural issues, and machine learning for user-centric approaches; (3) multimedia data: conceptual analysis at different levels (feature, cognitive, and affective); and (4) sources of contextual information and case studies in multi-camera networks. This full-day tutorial will consist of two parts: the first half will consist of presentations by the instructors, and the second part will consist of practical workgroup activities Alejandro Jaimes, Nicu Sebe |
ACM Multimedia | 1 |
| 2007 | Multimodal human-computer interaction: A survey
Alejandro Jaimes, Nicu Sebe |
Comput. Vis. Image Underst. | 1 |
| 2006 | Posture and activity silhouettes for self-reporting, interruption management, and attentive interfacesabstractIn this paper we present a novel system for monitoring a computer user's posture and activities in front of the computer (e.g., reading, speaking on the phone, etc.) for self-reporting. In our system, a camera and a microphone are placed in front of a computer work area (e.g., on top of the computer screen). The system can be used as a component in an attentive interface, or for giving the user real time feedback on the goodness of his current posture, and generating summaries of postures and activities over a specified period of time (e.g., hours, days, months, etc.). All elements of the system are highly customizable: the user decides what "good" postures are, what alarms and interruptions are triggered, if any, and what activity and posture summaries are generated. We present novel algorithms for posture measurement (using geometric features of the user's silhouette), and activity classification (using machine learning). Finally, we present experiments that show the feasibility of our approach. Alejandro Jaimes |
IUI | 1 |
| 2006 | Human-centered computing: a multimedia perspectiveabstractHuman-Centered Computing (HCC) is a set of methodologies that apply to any field that uses computers, in any form, in applications in which humans directly interact with devices or systems that use computer technologies. In this paper, we give an overview of HCC from a Multimedia perspective. We describe what we consider to be the three main areas of Human-Centered Multimedia (HCM): media production, analysis, and interaction. In addition, we identify the core characteristics of HCM, describe example applications, and propose a research agenda for HCM. Copyright 2006 ACM. Alejandro Jaimes, Nicu Sebe, Daniel Gatica-Perez |
ACM Multimedia | 1 |
| 2005 | Affective Meeting Video AnalysisabstractIn this paper we examine the affective content of meeting videos. First we asked five subjects to manually label three meeting videos using continuous response measurement (continuous-scale labeling in real-time) for energy and valence (the two dimensions of the human affect space). Then we automatically extracted audio-visual features to characterize the affective content of the videos. We compare the results of manual labeling and low-level automatic audio-visual feature extraction. Our analysis yields promising results, which suggest that affective meeting video analysis can lead to very interesting observations useful for automatic indexing. Alejandro Jaimes, Takeshi Nagamine, Kengo Omura, Nicu Sebe |
ICME | 1 |
| 2005 | Hotspot Components for Gesture-Based Interaction
Alejandro Jaimes |
INTERACT | 1 |
| 2005 | ACM multimedia interactive art program: an introduction to the presence/absence exhibitionabstractThe second ACM Multimedia Art program followed the successful formula used in ACM MM 2005, consisting of a session of long papers, a selection of posters and an art exhibition of multimedia works displayed at a gallery for a period encompassing the conference duration. "Presence/Absence" was selected as the central theme for the exhibition. In this paper, we discuss our motivations in organizing an art program at ACM MM, the exhibition theme, the works selected, and their potential impact in the technical community. Alejandro Jaimes, Andrew W. Senior, Wolfgang Muench |
ACM Multimedia | 1 |
| 2004 | ACM multimedia interactive art program: an introduction to the digital boundaries exhibitionabstractThe Digital Boundaries exhibition includes works that use multimedia to address issues of multiculturalism, identity, and awareness. By placing technology in new contexts to explore multimedia's impact on culture (and vice versa) we create a space for the discussion of new ideas and create an interdisciplinary impact by reinforcing a dialogue between the arts and multimedia communities. We discuss our motivation, the exhibition theme, the works selected, and their potential technical impact. Alejandro Jaimes, Pamela Jennings |
ACM Multimedia | 1 |
| 2004 | A visuospatial memory cue system for meeting video retrievalabstractWe present a system based on a new, memory-cue paradigm for retrieving meeting video scenes. The system graically represents important memory retrieval cues such as room layout, participant's faces and sitting positions, etc.. Queries are formulated dynamically: as the user graically manipulates the cues, the query results are shown. Our system (1) helps users easily express the cues they recall about a particular meeting; (2) helps users remember new cues for meeting video retrieval. We discuss the experiments that motivate this new approach, implementation, and future work. Takeshi Nagamine, Alejandro Jaimes, Kengo Omura, Kazutaka Hirata |
ACM Multimedia | 2 |
| 2004 | Interactive visualization of multi-stream meeting videos based on automatic visual content analysisabstractWe present a new approach to segment and visualize informally captured multi-stream meeting videos. We process the visual content in each stream individually by analyzing the differences between frames in each sequence to find change areas. These results are combined with face detection to determine visual activity in each of the streams. We then combine the activity scores from multiple streams and automatically generate a 3D representation of the video. Our representation allows the user to obtain an at-a-glance view of the video at different granularities of activity, view multiple streams simultaneously, and select particular points in time for viewing. We present experiments that suggest that low-level visual analysis can be effective for finding highlights that can be used for browsing multi-stream meeting videos. Alejandro Jaimes, Naofumi Yoshida, Kazumasa Murai, Kazutaka Hirata, Jun Miyazaki |
MMSP | 1 |
| 2004 | On the Image Content of a Web Segment: Chile as a Case Study
Alejandro Jaimes, Javier Ruiz-del-Solar, Rodrigo Verschae, Ricardo Baeza-Yates, Carlos Castillo 0001, D. Yaksic, Emilio Davis |
J. Web Eng. | 1 |
| 2003 | Interactive search fusion methods for video database retrievalabstractIn this paper, we investigate a new method for video database retrieval using interactive search fusion. Recent video analysis techniques have enabled the extraction of a variety of descriptors of features, concepts, clusters, classification results, speech and textual terms, MPEG-7 metadata, and so on. However, given an information need users are faced with a daunting task of trying to formulate queries over these multiple disparate data sources in order to retrieve the desired video content. In this paper, we explore a novel approach based on search fusion in which the user interactively builds a query by sequentially choosing among the descriptors and data sources and by selecting from various combining and score aggregation functions to fuse results of individual searches. For example, the system allows building of queries such as "retrieve video clips that have color of beach scenes, the detection of sky, and detection of water". In this paper we present the search fusion method and evaluate the performance on a large video database. John R. Smith, Alejandro Jaimes, Ching-Yung Lin, Milind R. Naphade, Apostol Natsev, Belle L. Tseng |
ICIP (1) | 2 |
| 2003 | Semi-automatic, data-driven construction of multimedia ontologiesabstractIn this paper we investigate semi-automatic construction of multimedia ontologies using a data-driven approach. We start with a collection of videos for which we wish to build an ontology (an explicit specification of a domain). Each video is pre-processed: scene cut detection, automatic speech recognition (ASR), and metadata extraction are performed. In addition we automatically index the videos based on visual content by extracting syntactic (e.g., color, texture, etc.) and semantic features (e.g., face, landscape, etc.). We then combine standard tools for ontology engineering and tools in content-based retrieval to semi-automatically build ontologies. In the first stage we process the text information available with the videos (ASR, metadata, and annotations, if any). Stop words (e.g., a, on, the) are eliminated and statistics (e.g., frequency, TFIDF, and entropy) are computed for all terms. Based on this data we manually select concepts and relationships to include in the ontology. Then we use content-based retrieval tools to assign multimedia entities (e.g., shots, videos, collections of videos) to concepts, properties, or relationships in the ontology, and to select multimedia entities as concepts, relationships, or properties in the ontology. We explore this methodology to construct multimedia ontologies from 24 hours of educational films from the 1940s-1960s used in the TREC video retrieval benchmark and discuss the problems encountered and future directions. Alejandro Jaimes, John R. Smith |
ICME | 1 |
| 2002 | Learning personalized video highlights from detailed MPEG-7 metadataabstractWe present a new framework for generating personalized video digests from detailed event metadata. In the new approach high level semantic features (e.g., number of offensive events) are extracted from an existing metadata signal using time windows (e.g., features within 16 sec. intervals). Personalized video digests are generated using a supervised learning algorithm which takes as input examples of important/unimportant events. Window-based features are extracted from the metadata and used to train the system and build a classifier that, given metadata for a new video, classifies segments into important and unimportant, according to a specific user, to generate personalized video digests. Our experimental results using soccer video suggest that extracting high level semantic information from existing metadata can be used effectively (80% precision and 85% recall using cross validation) in generating personalized video digests. Alejandro Jaimes, Tomio Echigo, Masayoshi Teraguchi, Fumiko Satoh |
ICIP (1) | 1 |
| 2002 | Duplicate detection in consumer photography and news videoabstractConsumers often make more than one photograph of the same scene, creating non-identical duplicates and near duplicates. In Kodak's consumer photography database, on average, 19% of the images, per roll, fall into this category. Automatic detection of duplicates, therefore, is extremely useful in applications that help users organize their image collections. We introduce the challenging problem of non-identical duplicate image detection in consumer photography, describe STELLA (a novel interactive personal image collection organization system), and give an overview of our novel framework for detecting duplicate and near duplicate consumer photographs and news videos. Alejandro Jaimes, Shih-Fu Chang, Alexander C. Loui |
ACM Multimedia | 1 |
| 2001 | A conceptual framework and empirical research for classifying visual descriptorsabstractAbstract This article presents exploratory research evaluating a conceptual structure for the description of visual content of images. The structure, which was developed from empirical research in several fields (e.g., Computer Science, Psychology, Information Studies, etc.), classifies visual attributes into a “Pyramid” containing four syntactic levels (type/technique, global distribution, local structure, composition), and six semantic levels (generic, specific, and abstract levels of both object and scene, respectively). Various experiments are presented, which address the Pyramid's ability to achieve several tasks: (1) classification of terms describing image attributes generated in a formal and an informal description task, (2) classification of terms that result from a structured approach to indexing, and (3) guidance in the indexing process. Several descriptions, generated by naive users and indexers, are used in experiments that include two image collections: a random Web sample, and a set of news images. To test descriptions generated in a structured setting, an Image Indexing Template (developed independently over several years of this project by one of the authors) was also used. The experiments performed suggest that the Pyramid is conceptually robust (i.e., can accommodate a full range of attributes), and that it can be used to organize visual content for retrieval, to guide the indexing process, and to classify descriptions obtained manually and automatically. Corinne Jörgensen, Alejandro Jaimes, Ana B. Benitez, Shih-Fu Chang |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2000 | Discovering Recurrent Visual Semantics in Consumer PhotographsabstractWe present techniques to semi-automatically discover recurrent visual semantics (RVS)-the repetitive appearance of visually similar elements such as objects and scenes-in consumer photographs. First, we introduce the detection of "bracketing" (very similar photographs) using an edge-correlation metric, which outperforms the color histogram. Then, we use color and novel composition features (based on automatic region segmentation) to perform scene-level clustering of images. We use a novel sequence-weighted technique, which uses the structure of standard film (only image sequence information), to perform hierarchical clustering. We show performance results of bracketing, explore clustering evaluation, and discuss STELLA, an interactive albuming and story telling application that uses these techniques to assist users in building digital albums. The STELLA system uses a new approach to album creation: instead of automatically creating albums, it provides an interactive environment that assists users in digital album creation. Alejandro Jaimes, Ana B. Benitez, Shih-Fu Chang, Alexander C. Loui |
ICIP | 1 |