EDBT 2026 Demo / reviewers in the wild / expert
Gerald Friedland
dblp:f/GeraldFriedland
· DBLP profile ↗
88ranked-venue papers
32as first author
3since 2021 · last 2025
0000-0002-9400-6539ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 75 · 25 first-author · 3 since 2021Artificial intelligence and machine learning · 19 · 4 first-authorDatabases, data management, data science and information retrieval · 5Human-computer interaction and ubiquitous computing · 3 · 1 first-authorComputer networks · 2 · 1 first-authorSecurity and privacy · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
26 papers |
Multimedia analysis and retrieval · 55% Audio and music processing · 40% Image and video processing · 5% | |
| Artificial intelligence
6 papers |
Speech recognition and synthesis · 57% Efficient and distributed learning · 41% Probabilistic and Bayesian machine learning · 2% | |
| Network and information security
6 papers |
Privacy and data protection · 96% Usable security · 4% | |
| Databases, data mining, and information retrieval
9 papers |
Machine learning and data management · 66% Web and social media mining · 20% Information retrieval · 8% | |
| Human-computer interaction and pervasive computing
3 papers |
Collaborative and social computing · 56% Ubiquitous computing and smart environments · 29% Human-AI interaction · 15% |
Topics — the 30 heaviest of 51, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
automated machine learning |
0.7 | 1 | 2023 | Efficient Multimedia Computing: Unleashing the Power of AutoML · ACM Multimedia 2023 |
Natural language and speech › Speech recognition and synthesis
speaker diarization |
0.6 | 5 | 2012 | Speaker Diarization: A Review of Recent Research · IEEE Trans. Speech Audio Process. 2012 The ICSI RT-09 Speaker Diarization System · IEEE Trans. Speech Audio Process. 2012 Estimating Dominance in Multi-Party Meetings Using Speaker Diarization · IEEE Trans. Speech Audio Process. 2011 |
Privacy and data protection › data confidentiality › content privacy
multimedia privacy |
0.6 | 3 | 2016 | Multimedia Privacy · ACM Multimedia 2016 Privacy concerns of sharing multimedia in social networks · ACM Multimedia 2013 Privacy concerns in multimedia and their solutions · ACM Multimedia 2012 |
Audio and music processing
audio representation learning |
0.5 | 2 | 2017 | DCAR: A Discriminative and Compact Audio Representation for Audio Processing · IEEE Trans. Multim. 2017 A Discriminative and Compact Audio Representation for Event Detection · ACM Multimedia 2016 |
Audio and music processing
sound event detection |
0.5 | 2 | 2017 | DCAR: A Discriminative and Compact Audio Representation for Audio Processing · IEEE Trans. Multim. 2017 A Discriminative and Compact Audio Representation for Event Detection · ACM Multimedia 2016 |
Multimedia analysis and retrieval
geotagging |
0.5 | 3 | 2014 | GeoMM 2014: the third ACM multimedia workshop ongeotagging and its applications in multimedia · ACM Multimedia 2014 Second ACM multimedia workshop on geotagging and its applications in multimedia (GeoMM 2013) · ACM Multimedia 2013 GeoMM'12: ACM international workshop on geotagging and its applications in multimedia · ACM Multimedia 2012 |
Audio and music processing › audio classification
acoustic scene classification |
0.3 | 1 | 2017 | DCAR: A Discriminative and Compact Audio Representation for Audio Processing · IEEE Trans. Multim. 2017 |
Multimedia analysis and retrieval
image retrieval |
0.3 | 1 | 2017 | Contextual Noise Reduction for Domain Adaptive Near-Duplicate Retrieval on Merchandize Images · IEEE Trans. Image Process. 2017 |
Multimedia analysis and retrieval › image retrieval
near-duplicate image retrieval |
0.3 | 1 | 2017 | Contextual Noise Reduction for Domain Adaptive Near-Duplicate Retrieval on Merchandize Images · IEEE Trans. Image Process. 2017 |
Image and video processing › feature representation
visual word quantization |
0.3 | 1 | 2017 | Contextual Noise Reduction for Domain Adaptive Near-Duplicate Retrieval on Merchandize Images · IEEE Trans. Image Process. 2017 |
Privacy and data protection › image privacy
bystander privacy |
0.3 | 1 | 2017 | Privacy Protection in Online Multimedia · ACM Multimedia 2017 |
Multimedia analysis and retrieval › multimodal summarization
event summarization |
0.2 | 1 | 2015 | Evento 360: Social Event Discovery from Web-scale Multimedia Collection · ACM Multimedia 2015 |
Multimedia analysis and retrieval › event detection
multimodal event detection |
0.2 | 1 | 2015 | Evento 360: Social Event Discovery from Web-scale Multimedia Collection · ACM Multimedia 2015 |
Multimedia analysis and retrieval › multimedia analysis
multimodal mining |
0.2 | 1 | 2015 | Multimedia COMMONS - Community-Organized Multimodal Mining: Opportunities for Novel Solutions (MMCommons Workshop 2015) · ACM Multimedia 2015 |
Multimedia analysis and retrieval
multimodal summarization |
0.2 | 1 | 2015 | Evento 360: Social Event Discovery from Web-scale Multimedia Collection · ACM Multimedia 2015 |
Collaborative and social computing
crowdsourcing |
0.2 | 1 | 2014 | Creating Experts From the Crowd: Techniques for Finding Workers for Difficult Tasks · IEEE Trans. Multim. 2014 |
Audio and music processing › speaker recognition
speaker identification |
0.2 | 2 | 2009 | Joke-o-mat: browsing sitcoms punchline by punchline · ACM Multimedia 2009 Live speaker identification in conversations · ACM Multimedia 2008 |
Privacy and data protection
social network privacy |
0.2 | 1 | 2013 | Privacy concerns of sharing multimedia in social networks · ACM Multimedia 2013 |
Audio and music processing › acoustic signal processing
acoustic feature analysis |
0.1 | 1 | 2012 | Name that room: room identification using acoustic features in a recording · ACM Multimedia 2012 |
Multimedia analysis and retrieval › audio-visual learning
audio-visual video analysis |
0.1 | 1 | 2012 | AMVA'12: ACM international workshop on audio and multimedia methods for large-scale video analysis · ACM Multimedia 2012 |
Multimedia analysis and retrieval › video analysis
large-scale video analytics |
0.1 | 1 | 2012 | AMVA'12: ACM international workshop on audio and multimedia methods for large-scale video analysis · ACM Multimedia 2012 |
Audio and music processing › room acoustics
reverberation |
0.1 | 1 | 2012 | Name that room: room identification using acoustic features in a recording · ACM Multimedia 2012 |
Multimedia analysis and retrieval › multimedia browsing
video browsing |
0.1 | 2 | 2010 | Joke-o-Mat HD: browsing sitcoms with human derived transcripts · ACM Multimedia 2010 Joke-o-mat: browsing sitcoms punchline by punchline · ACM Multimedia 2009 |
Ubiquitous computing and smart environments › mobile sensing
smartphone sensing |
0.1 | 1 | 2010 | Precise indoor localization using smart phones · ACM Multimedia 2010 |
Wireless sensing and localization › indoor localization
fingerprint-based localization |
0.1 | 1 | 2010 | Precise indoor localization using smart phones · ACM Multimedia 2010 |
Wireless sensing and localization
indoor localization |
0.1 | 1 | 2010 | Precise indoor localization using smart phones · ACM Multimedia 2010 |
Natural language and speech › Speech recognition and synthesis › speech analysis
prosodic features |
0.1 | 1 | 2009 | Prosodic and other Long-Term Features for Speaker Diarization · IEEE Trans. Speech Audio Process. 2009 |
Audio and music processing › speaker diarization
audio-visual speaker diarization |
0.1 | 1 | 2009 | Visual speaker localization aided by acoustic models · ACM Multimedia 2009 |
Audio and music processing
speaker diarization |
0.1 | 1 | 2009 | Visual speaker localization aided by acoustic models · ACM Multimedia 2009 |
Data mining › dimensionality reduction
manifold learning |
0.1 | 1 | 2016 | A Discriminative and Compact Audio Representation for Event Detection · ACM Multimedia 2016 |
Methods — techniques the papers use, named apart from their topics
AutoML · 1.3gaussian mixture model · 0.9grassmannian manifold optimization · 0.8unsupervised clustering · 0.4hierarchical clustering · 0.4visual words · 0.3context graph · 0.3anisotropic diffusion filter · 0.3visual cue analysis · 0.2wifi fingerprinting · 0.2statistical processing of radio signal strengths · 0.2annotator qualification · 0.2speaker modeling · 0.2temporal analysis · 0.2language model · 0.2human subject study · 0.2geotagging · 0.2geolocation analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From 2D to 3D: How Discrete Dependencies Enable Cross-Dimensional Inference in Neural Networks in Defiance of Euclidean Geometry
Gerald Friedland, Robert Mertens 0003 |
ISM | 1 |
| 2023 | Deep Layers Beware: Unraveling the Surprising Benefits of JPEG Compression for Image Classification Pre-processingabstractIn this paper, we explore the intriguing effects of JPEG compression as a pre-processing technique for image classification tasks. Building upon the findings of a previous study by Friedland et al., which demonstrated that substantial JPEG compression does not significantly degrade classification accuracy, we investigate the potential benefits and limitations of this approach when applied to various classifiers, such as AutoGluon-multimodal and EfficientNet. Our experiments not only confirm the original results but also reveal notable I/O benefits, with compressed images occupying as little as 14 % of the original dataset size while maintaining comparable accuracy.Despite these promising findings, we also document several investigations that did not yield beneficial outcomes. We found no evidence to suggest that JPEG compression leads to faster model convergence or allows smaller models to achieve the same accuracy. Additionally, our experiments showed that tabular classifiers could not match the performance of deep neural networks when trained on JPEG-compressed input, and that JPEG compression does not make classifiers more resilient to noise in input images.Together, our results provide a comprehensive evaluation of JPEG compression as a pre-processing technique for image classification. While the approach offers undeniable benefits in terms of data storage and accuracy preservation, it does not appear to yield advantages in terms of model convergence, model size, or robustness to noise. This study contributes valuable insights for researchers and practitioners working in multimedia signal processing and image recognition, paving the way for further exploration and optimization of multimedia compression techniques. Guruprasad Nayak, Gerald Friedland |
ISM | 2 |
| 2023 | Efficient Multimedia Computing: Unleashing the Power of AutoMLabstractAs the field of multimedia computing has grown rapidly, so has the need for larger datasets[5] and increased modeling capacity. Navigating this complex landscape often necessitates the use of sophisticated tools and cloud architectures, which all need to be addressed before the actual research commences. Recently, AutoML, an innovation previously exclusive to tabular data, has expanded to encompass multimedia data. This development has the potential to greatly streamline the research process, allowing researchers to shift their focus from model construction to the core content of their problems. In doing so, AutoML not only optimizes resource utilization but also boosts the reproducibility of results. The aim of this tutorial is to acquaint the multimedia community with AutoML technologies, underscoring their advantages and their practical applications in the field. Debanjan Datta, Gerald Friedland |
ACM Multimedia | 2 |
| 2020 | Reproducibility and Experimental Design for Machine Learning on Audio and Multimedia DataabstractThis tutorial provides an actionable perspective on the experimental design for machine learning experiments on multimedia data. The tutorial consists of lectures and hands-on exercises. The lectures provide an engineering introduction to machine learning design. By understanding the information flow and quantities in the scientific process, machine learners can be designed to be more efficient and their limits can be easier understood. The thought framework presented is derived from the traditional experimental sciences which require published results to be self-contained with regards to reproducibility. In the practical exercises, we will work on calculating and measuring quantities like Memory Equivalent Capacity or generalization ratio for different machine learners and data sets and discuss how these quantities relate to reproducible experimental design. Gerald Friedland |
ACM Multimedia | 1 |
| 2020 | DIME: An Online Tool for the Visual Comparison of Cross-modal Retrieval Models
Tony Zhao, Jaeyoung Choi 0002, Gerald Friedland |
MMM (2) | 3 |
| 2019 | Reproducibility and Experimental Design for Machine Learning on Audio and Multimedia DataabstractThis tutorial provides an actionable perspective on the experimental design for machine learning experiments on multimedia data. The tutorial consists of lectures and hands-on exercises. The lectures provide a theoretical introduction to machine learning design and signal processing. The thought framework presented is derived from the traditional experimental sciences which require published results to be self-contained with regards to reproducibility. In the practical exercises, we will work on calculating and measuring quantities like capacity or generalization ratio for different machine learners and data sets and discuss how these quantities relate to reproducible experimental design. Gerald Friedland |
ACM Multimedia | 1 |
| 2018 | Rethinking Summarization and Storytelling for Modern Social Multimedia
Stevan Rudinac, Tat-Seng Chua, Nicolás E. Díaz Ferreyra, Gerald Friedland, Tatjana Gornostaja, Benoit Huet, Rianne Kaptein, Krister Lindén, Marie-Francine Moens, Jaakko Peltonen, Miriam Redi, Markus Schedl, David A. Shamma, Alan F. Smeaton, Lexing Xie |
MMM (1) | 4 |
| 2017 | The Geo-Privacy Bonus of Popular Photo EnhancementsabstractToday's geo-location estimation approaches are able to infer the location of a target image using its visual content alone. These approaches typically exploit visual matching techniques, applied to a large collection of background images with known geo-locations. Users who are unaware that visual analysis and retrieval approaches can compromise their geo-privacy, unwittingly open themselves to risks of crime or other unintended consequences. This paper lays the groundwork for a new approach to geo-privacy of social images: Instead of requiring a change of user behavior, we start by investigating users' existing photo-sharing practices. We carry out a series of experiments using a large collection of social images (8.5M) to systematically analyze how photo editing practices impact the performance of geo-location estimation. We find that standard image enhancements, including filters and cropping, already serve as natural geo-privacy protectors. In our experiments, up to 19% of images whose location would otherwise be automatically predictable were unlocalizeable after enhancement. We conclude that it would be wrong to assume that geo-visual privacy is a lost cause in today's world of rapidly maturing machine learning. Instead, protecting users against the unwanted effects of pixel-based inference is a viable research field. A starting point is understanding the geo-privacy bonus of already established user behavior. Jaeyoung Choi 0002, Martha A. Larson, Xinchao Li, Gerald Friedland, Alan Hanjalic |
ICMR | 5 |
| 2017 | Privacy Protection in Online MultimediaabstractOnline multimedia has been growing rapidly due to ubiquitous mobile phones, widely deployed surveillance cameras, dashcams and mini-drones. When one takes photographs or videos at a public location, it is highly likely that some other people ("bystanders") also appear in the visual data. The data may be available online, such as shared by social media, and questions about privacy arise. This panel discusses the issues about privacy in online multimedia from legal, technological, and social aspects. Yung-Hsiang Lu, Andrea Cavallaro, Catherine Crump, Gerald Friedland, Keith Winstein |
ACM Multimedia | 4 |
| 2017 | Contextual Noise Reduction for Domain Adaptive Near-Duplicate Retrieval on Merchandize ImagesabstractIn this paper, we have proposed a novel method which utilizes the contextual relationship among visual words for reducing the Quantization errors in near-duplicate image retrieval (NDR). Instead of following the track of conventional NDR techniques which usually search new solutions by borrowing ideas from the text domain, we propose to model the problem back to image domain, which results in a more natural way of solution search. The idea of the proposed method is to construct a context graph that encapsulates the contextual relationship within an image and treat the graph as a pseudo-image, so that classical image filters can be adopted to reduce the mismapped visual words which are contextually inconsistent with others.With these contextual noises reduced, the method provides purified inputs to the subsequent processes in NDR, and improves the overall accuracy. More importantly, the purification further increases the sparsity of the image feature vectors, which thus speeds up the conventional methods by 1662% times and makes NDR practical to online applications on merchandize images where the requirement of response time is critical. The way of considering contextual noise reduction in image domain also makes the problem open to all sophisticated filters. Our study shows the classic anisotropic diffusion filter can be employed to address the cross-domain issue, resulting in the superiority of the method to conventional ones in both effectiveness and efficiency. Zhen-Qun Yang, Xiaoyong Wei, Zhang Yi 0001, Gerald Friedland |
IEEE Trans. Image Process. | 4 |
| 2017 | DCAR: A Discriminative and Compact Audio Representation for Audio ProcessingabstractThis paper presents a novel two-phase method for audio representation, discriminative and compact audio representation (DCAR), and evaluates its performance at detecting events and scenes in consumer-produced videos. In the first phase of DCAR, each audio track is modeled using a Gaussian mixture model (GMM) that includes several components to capture the variability within that track. The second phase takes into account both global structure and local structure. In this phase, the components are rendered more discriminative and compact by formulating an optimization problem on a Grassmannian manifold. The learned components can effectively represent the structure of audio. Our experiments used the YLI-MED and DCASE Acoustic Scenes datasets. The results show that variants on the proposed DCAR representation consistently outperform four popular audio representations (mv-vector, i-vector, GMM, and HEM-GMM). The advantage is significant for both easier and harder discrimination tasks; we discuss how these performance differences across tasks follow from how each type of model leverages (or does not leverage) the intrinsic structure of the data. Liping Jing, Bo Liu 0050, Jaeyoung Choi 0002, Adam Janin, Julia Bernd, Michael W. Mahoney, Gerald Friedland |
IEEE Trans. Multim. | 7 |
| 2016 | Multimedia PrivacyabstractThis tutorial brings together a number of recent advances at the nexus of multimedia analysis, online privacy, and social media mining. Our goal is to offer a multidisciplinary view of the emerging field of Multimedia Privacy: the study of privacy issues arising in the context of multimedia sharing in online platforms, and the pursuit of new approaches to mitigating those issues within multimedia computer science. Gerald Friedland, Symeon Papadopoulos, Julia Bernd, Ioannis Kompatsiaris |
ACM Multimedia | 1 |
| 2016 | A Discriminative and Compact Audio Representation for Event DetectionabstractThis paper presents a novel two-phase method for audio representation: Discriminative and Compact Audio Representation (DCAR). In the first phase, each audio track is modeled using a Gaussian mixture model (GMM) that includes several components to capture the variability within that track. The second phase takes into account both global structure and local structure. In this phase, the components are rendered more discriminative and compact by formulating an optimization problem on Grassmannian manifolds, which we found represents the structure of audio effectively. Experimental results on the YLI-MED dataset show that the proposed DCAR representation consistently outperforms state-of-the-art audio representations: i-vector, mv-vector, and GMM. Liping Jing, Bo Liu 0050, Jaeyoung Choi 0002, Adam Janin, Julia Bernd, Michael W. Mahoney, Gerald Friedland |
ACM Multimedia | 7 |
| 2016 | The Teaching Privacy CurriculumabstractA basic understanding of online privacy is essential to being an informed digital citizen, and therefore basic privacy education is becoming ever more necessary. Recently released high school and college computer science curricula acknowledge the significantly increased importance of fundamental knowledge about privacy, but do not yet provide concrete content in the area. To address this need, over the past two years, we have developed the Teaching Privacy Project (TPP) curriculum, http://teachingprivacy.org, which educates the general public about online privacy issues. We performed a pilot of our curriculum in a university course for non-CS majors and found that it was effective: weeks after last being exposed, students' privacy attitudes had shifted. In this paper, we describe our curriculum, our evaluation of it in the classroom, and our vision for future privacy education. Serge Egelman, Julia Bernd, Gerald Friedland, Dan Garcia 0001 |
SIGCSE | 3 |
| 2015 | An information-theoretic metric of fingerprint effectivenessabstractAudio fingerprinting refers to the process of extracting a robust, compact representation of audio which can be used to uniquely identify an audio segment. Works in the audio fingerprinting literature generally report results using system-level metrics. Because these systems are usually very complex, the overall system-level performance depends on many different factors. So, while these metrics are useful in understanding how well the entire system performs, they are not very useful in knowing how good or bad the fingerprint design is. In this work, we propose a metric of fingerprint effectiveness that decouples the effect of other system components such as the search mechanism or the nature of the database. The metric is simple, easy to compute, and has a clear interpretation from an information theory perspective. We demonstrate that the metric correlates directly with system-level metrics in assessing fingerprint effectiveness, and we show how it can be used in practice to diagnose the weaknesses in a fingerprint design. T. J. Tsai 0001, Gerald Friedland, Xavier Anguera Miró |
ICASSP | 2 |
| 2015 | Audio-Based Multimedia Event Detection with DNNs and Sparse SamplingabstractThis paper presents advances in analyzing audio content information to detect events in videos, such as a parade or a birthday party. We developed a set of tools for audio processing within the predominantly vision-focused deep neural network (DNN) framework Caffe. Using these tools, we show, for the first time, the potential of using only a DNN for audio-based multimedia event detection. Training DNNs for event detection using the entire audio track from each video causes a computational bottleneck. Here, we address this problem by developing a sparse audio frame-sampling method that improves event-detection speed and accuracy. We achieved a 10 percentage-point improvement in event classification accuracy, with a 200x reduction in the number of training input examples as compared to using the entire track. This reduction in input feature volume led to a 16x reduction in the size of the DNN architecture and a 300x reduction in training time. We applied our method using the recently released YLI-MED dataset and compared our results with a state-of-the-art system and with results reported in the literature for TRECVIDMED. Our results show much higher MAP scores compared to a baseline i-vector system - at a significantly reduced computational cost. The speed improvement is relevant for processing videos on a large scale, and could enable more effective deployment in mobile systems. Khalid Ashraf, Benjamin Elizalde, Forrest N. Iandola, Matthew W. Moskewicz, Julia Bernd, Gerald Friedland, Kurt Keutzer |
ICMR | 6 |
| 2015 | Evento 360: Social Event Discovery from Web-scale Multimedia CollectionabstractWe present Evento 360 (URL: http://evento360.info), an online interactive social event browser, which allows the user to explore events detected within a web-scale multimedia corpus. The system addresses five key aspects of social multimedia event detection and summarization: multimodality, scale, diversity of representations, noise of multimedia items, and missing metadata. The detection algorithm uses unsupervised clustering approach that exploits temporal, spatial and textual metadata. For each detected event cluster, to choose the best subset of photos that meet both relevance and diversity criteria, the system uses hierarchical clustering that exploits both visual and audio information. Evento 360's user interface provides a search feature that is not limited to a certain set of events, but rather can handle an arbitrary event query. It allows the user to retrieve and explore relevant events. The system scales well and is effective in producing high-quality summaries of the detected events. Jaeyoung Choi 0002, Eungchan Kim, Martha A. Larson, Gerald Friedland, Alan Hanjalic |
ACM Multimedia | 4 |
| 2015 | Multimedia COMMONS - Community-Organized Multimodal Mining: Opportunities for Novel Solutions (MMCommons Workshop 2015)abstractThe Multimedia COMMONS workshop laid the groundwork for developing a research community around the Multimedia Genome Project (MMGP), an initiative initially focused on annotation of---and research using---the 99.2 million images and nearly 800,000 videos in the Yahoo Flickr Creative Commons 100 Million dataset (YFCC100M). Current and potential users of the YFCC100M presented new research and systems that used this unprecedentedly large, unprecedentedly open-source dataset; discussed ideas for future data challenges and new benchmarking tasks that would not previously have been possible; and suggested priorities and plans for annotation and distribution based on community needs and interests. Gerald Friedland, Chong-Wah Ngo, David A. Shamma |
ACM Multimedia | 1 |
| 2015 | Teaching Privacy: What Every Student Needs to Know (Abstract Only)abstractAlthough frequent stories in the popular media have raised awareness about online privacy, most young people do not have a very good handle on what the specific issues are, nor the practical steps they can take to manage them. Teachers recognize that all their students' from future engineers to those totally bored by science need a realistic understanding of how online privacy works, so they can protect themselves online. In fact, the latest CS curricular recommendations include privacy but there is no comprehensive set of field-tested teaching materials. To address this, we are developing TROPE (Teachers' Resources for Online Privacy Education), a set of classroom-ready teaching materials (teachingprivacy.org). TROPE will provide educators with lesson modules, interactive demonstrations, and a teachers' guide, so they can readily integrate privacy into high school and college classes. Our goal for this workshop is twofold. First, we will introduce educators to TROPE and provide guidance on how they can cover privacy-related topics in their classrooms without being subject-matter experts. Second, we will solicit feedback and on-the-ground stories; by gaining a better understanding of specific problems faced by educators and students, we can increase TROPE's utility to teachers. We will provide teachers with up-to-date technical information about online privacy, including relevant highlights from our research; hands-on activities illustrating principles of online privacy; and an overview of the materials we are creating for TROPE. This will be an interactive workshop, driven by participants' questions, experiences, and interests. For CS educators at all levels; laptop or tablet recommended. Gerald Friedland, Serge Egelman, Dan Garcia 0001 |
SIGCSE | 1 |
| 2014 | Toward efficient, privacy-aware media classification on public databasesabstractThe ability to search databases by providing multimedia examples of voices, faces, or locations instead of textual descriptions can be tremendously useful. At the same time, uploading media for queries---especially media that contains sensitive content---means sharing private information with a potentially untrusted service provider. The growing field of privacy-preserving database searches attempts to resolve this tension. Within this scope of private searches, private media classification and retrieval is particularly challenging due to the inherent inexactness of recognition; to be useful, image or other media classification systems must identify approximate matches rather than just exact ones. This is difficult to reconcile with distortion-intolerant and resource-heavy privacy primitives, especially in web-scale databases. In this paper, we present an architecture for media classification on public databases that preserves client privacy while achieving asymptotic communication and computation costs that are sublinear in the size of the database. We demonstrate the usefulness of this architecture in the context of a privacy-preserving face recognition system. We observe order-of-magnitude speedups over state-of-the-art private face recognition systems. Giulia Fanti, Matthieu Finiasz, Gerald Friedland, Kannan Ramchandran |
ICMR | 3 |
| 2014 | GeoMM 2014: the third ACM multimedia workshop ongeotagging and its applications in multimediaabstractIt is our great pleasure to welcome you to the Third ACM Workshop on Geotagging and Its Applications in Multimedia -- GeoMM'14. This year's event continues the workshops in 2012 and 2013, with the goal of building a forum for the presentation and synthesis of vision and insight from leading experts and practitioners on the developing directions of geotagging research related to multimedia. Following the success in previous years, the GeoMM workshop serves as a venue for the premier research in geotagging and multimedia, and continues to attract submissions from a diverse set of researchers, who address newly arising problems within this emerging field. Five regular papers are presented in this workshop, covering a number of novel applications and new methodologies. An invited paper is also presented to introduce the related MediaEval 2014 Placing task, which consists of 5 million geotagged photos and 25,000 geotagged videos. We believe this workshop will benefit more and more research works in the broad research field. Liangliang Cao, Gerald Friedland, Lexing Xie |
ACM Multimedia | 2 |
| 2014 | Creating Experts From the Crowd: Techniques for Finding Workers for Difficult TasksabstractCrowdsourcing is currently used for a range of applications, either by exploiting unsolicited user-generated content, such as spontaneously annotated images, or by utilizing explicit crowdsourcing platforms such as Amazon Mechanical Turk to mass-outsource artificial-intelligence-type jobs. However, crowdsourcing is most often seen as the best option for tasks that do not require more of people than their uneducated intuition as a human being. This article describes our methods for identifying workers for crowdsourced tasks that are difficult for both machines and humans. It discusses the challenges we encountered in qualifying annotators and the steps we took to select the individuals most likely to do well at these tasks. Luke R. Gottlieb, Gerald Friedland, Jaeyoung Choi 0002, Pascal Kelm, Thomas Sikora |
IEEE Trans. Multim. | 2 |
| 2014 | Scalable multimedia content analysis on parallel platforms using pythonabstractIn this new era dominated by consumer-produced media there is a high demand for web-scalable solutions to multimedia content analysis. A compelling approach to making applications scalable is to explicitly map their computation onto parallel platforms. However, developing efficient parallel implementations and fully utilizing the available resources remains a challenge due to the increased code complexity, limited portability and required low-level knowledge of the underlying hardware. In this article, we present PyCASP, a Python-based framework that automatically maps computation onto parallel platforms from Python application code to a variety of parallel platforms. PyCASP is designed using a systematic, pattern-oriented approach to offer a single software development environment for multimedia content analysis applications. Using PyCASP, applications can be prototyped in a couple hundred lines of Python code and automatically scale to modern parallel processors. Applications written with PyCASP are portable to a variety of parallel platforms and efficiently scale from a single desktop Graphics Processing Unit (GPU) to an entire cluster with a small change to application code. To illustrate our approach, we present three multimedia content analysis applications that use our framework: a state-of-the-art speaker diarization application, a content-based music recommendation system based on the Million Song Dataset, and a video event detection system for consumer-produced videos. We show that across this wide range of applications, our approach achieves the goal of automatic portability and scalability while at the same time allowing easy prototyping in a high-level language and efficient performance of low-level optimized code. Ekaterina Gonina, Gerald Friedland, Eric Battenberg, Penporn Koanantakool, Michael B. Driscoll, Evangelos Georganas, Kurt Keutzer |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | Nowhere to hide: Exploring user-verification across Flickr accountsabstractThis work presents improved audio-based user-verification analysis and results on Flickr videos, using a subset of the MediaEval 2011 [1] data set. User-verification is a new task, where the goal is to determine if two pieces of media are uploaded by the same user. Our best results, with a 19.7% Equal Error Rate, and a 53.9% Miss Rate at 1% False Positive, are obtained using an i-vector [2] system. A frequency-matching system that requires 96% less computation time than the other systems is also explored, and may be better suited for processing large datasets from Flickr and other social networks. The results have significant privacy implications as they present a framework for exploiting users' tendencies to assume that different accounts remain as separate realms. Howard Lei, Jaeyoung Choi 0002, Gerald Friedland |
ICASSP | 3 |
| 2013 | Lost in segmentation: Three approaches for speech/non-speech detection in consumer-produced videosabstractTraditional speech/non-speech segmentation systems have been designed for specific acoustic conditions, such as broadcast news or meetings. However, little research has been done on consumer-produced audio. This type of media is constantly growing and has complex characteristics such as low quality recordings, environmental noise and overlapping sounds. This paper discusses an evaluation of three different approaches for speech/non-speech detection on consumer-produced audio. The approaches are state-of-the-art speech/non-speech detectors-one based on Gaussian Mixture Models (GMM), another on Support Vector Machines (SVM), and the last on Neural Networks (NN). Using the TRECVID MED 2012 database, we designed training/testing sets combinations to aid the understanding of what speech/non-speech detection on consumer-produced media entails and how traditional approaches to this detection performed in this domain. The results revealed that the cross-domain state-of-the-art GMM and SVM systems' tests underperformed a one-layer NN algorithm, which had 20% higher accuracy and computed audio 5 times faster. Benjamin Elizalde, Gerald Friedland |
ICME | 2 |
| 2013 | Exploring methods of improving speaker accuracy for speaker diarization
Mary Tai Knox, Nikki Mirghafori, Gerald Friedland |
INTERSPEECH | 3 |
| 2013 | An i-Vector Representation of Acoustic Environments for Audio-Based Video Event Detection on User Generated ContentabstractAudio-based video event detection (VED) on user-generated content (UGC) aims to find videos that show an observable event such as a wedding ceremony or birthday party rather than a sound, such as music, clapping or singing. The difficulty of video content analysis on UGC lies in the acoustic variability and lack of structure of the data. The UGC task has been explored mainly by computer vision, but can be benefited by the used of audio. The i-vector system is state-of-the-art in Speaker Verification, and is outperforming a conventional Gaussian Mixture Model (GMM)-based approach. The system compensates for undesired acoustic variability and extracts information from the acoustic environment, making it a meaningful choice for detection on UGC. This paper employs the i-vector-based system for audio-based VED on UGC and expands the understanding of the system on the task. It also includes a performance comparison with the conventional GMM-based and state-of-the-art Random Forest (RF)-based systems. The i-vector system aids audio-based event detection by addressing UGC audio characteristics. It outperforms the GMM-based system, and is competitive with the RF-based system in terms of the Missed Detection (MD) rate at 4% and 2.8% False Alarm (FA) rates, and complements the RF-based system by demonstrating slightly improvement in combination over the standalone systems. Benjamin Elizalde, Howard Lei, Gerald Friedland |
ISM | 3 |
| 2013 | Second ACM multimedia workshop on geotagging and its applications in multimedia (GeoMM 2013)abstractThe Workshop on Geotagging and Its Applications in Multimedia (GeoMM 2013) focuses on new applications and methods of geotagging and in geo-location support systems. As the location based multimedia becomes more and more popular in the era of Web and mobile applications, the increase in the use of geotagging and improvements in geo-location support systems open up a new dimension for the description, organization and manipulation of multimedia data. This new dimension radically expands the usefulness of multimedia data both for daily users of the Internet and social networking sites as well as for experts in particular application scenarios. The workshop serves as a venue for the premier research in geotagging and multimedia, and continues to attract submissions from a diverse set of researchers, who address newly arising problems within this emerging field. Liangliang Cao, Gerald Friedland, Pascal Kelm |
ACM Multimedia | 2 |
| 2013 | Human vs machine: establishing a human baseline for multimodal location estimationabstractOver the recent years, the problem of video location estimation (i.e., estimating the longitude/latitude coordinates of a video without GPS information) has been approached with diverse methods and ideas in the research community and significant improvements have been made. So far, however, systems have only been compared against each other and no systematic study on human performance has been conducted. Based on a human-subject study with 11,900 experiments, this article presents a human baseline for location estimation for different combinations of modalities (audio, audio/video, audio/video/text). Furthermore, this article compares state-of-the-art location estimation systems with the human baseline. Although the overall performance of humans' multimodal video location estimation is better than current machine learning approaches, the difference is quite small: For 41% of the test set, the machine's accuracy was superior to the humans. We present case studies and discuss why machines did better for some videos and not for others. Our analysis suggests new directions and priorities for future work on the improvement of location inference algorithms. Jaeyoung Choi 0002, Howard Lei, Venkatesan N. Ekambaram, Pascal Kelm, Luke R. Gottlieb, Thomas Sikora, Kannan Ramchandran, Gerald Friedland |
ACM Multimedia | 8 |
| 2013 | Privacy concerns of sharing multimedia in social networksabstractThis article summarizes the corresponding 3-hour tutorial at ACM Multimedia 2013. Gerald Friedland |
ACM Multimedia | 1 |
| 2013 | Exploiting innocuous activity for correlating users across sitesabstractWe study how potential attackers can identify accounts on different social network sites that all belong to the same user, exploiting only innocuous activity that inherently comes with posted content. We examine three specific features on Yelp, Flickr, and Twitter: the geo-location attached to a user's posts, the timestamp of posts, and the user's writing style as captured by language models. We show that among these three features the location of posts is the most powerful feature to identify accounts that belong to the same user in different sites. When we combine all three features, the accuracy of identifying Twitter accounts that belong to a set of Flickr users is comparable to that of existing attacks that exploit usernames. Our attack can identify 37% more accounts than using usernames when we instead correlate Yelp and Twitter. Our results have significant privacy implications as they present a novel class of attacks that exploit users' tendency to assume that, if they maintain different personas with different names, the accounts cannot be linked together; whereas we show that the posts themselves can provide enough information to correlate the accounts. Oana Goga, Howard Lei, Sree Hari Krishnan Parthasarathi, Gerald Friedland, Robin Sommer, Renata Teixeira |
WWW | 4 |
| 2013 | Narrative theme navigation for sitcoms supported by fan-generated scripts - Video navigation based on acoustic detection of actors and narrative elements
Gerald Friedland, Luke R. Gottlieb, Adam Janin |
Multim. Tools Appl. | 1 |
| 2013 | Editorial for automated media analysis and production for novel TV services
Gerald Friedland, Alberto Messina, Robbie De Sutter, Masanori Sano |
Multim. Tools Appl. | 1 |
| 2012 | How to put it into words - using random forests to extract symbol level descriptions from audio content for concept detectionabstractThis paper presents a system that uses symbolic representations of audio concepts as words for the descriptions of audio tracks, that enable it to go beyond the state of the art, which is audio event classification of a small number of audio classes in constrained settings, to large-scale classification in the wild. These audio words might be less meaningful for an annotator but they are descriptive for computer algorithms. We devise a random-forest vocabulary learning method with an audio word weighting scheme based on TF-IDF and TD-IDD, so as to combine the computational simplicity and accurate multi-class classification of the random forest with the data-driven discriminative power of the TF-IDF/TD-IDD methods. The proposed random forest clustering with text-retrieval methods significantly outperforms two state-of-the-art methods on the dry-run set and the full set of the TRECVID MED 2010 dataset. Po-Sen Huang, Robert Mertens 0003, Ajay Divakaran, Gerald Friedland, Mark Hasegawa-Johnson |
ICASSP | 4 |
| 2012 | Multimodal city-verification on flickr videos using acoustic and textual featuresabstractWe have performed city-verification of videos based on the videos' audio and metadata, using videos from the MediaEval Placing Task's video set, which contain consumer-produced videos “from-the-wild”. 18 cities were used as targets, for which acoustic and language models were trained, and against which test videos were scored. We have obtained the first known results for the city verification task, with an EER minimum of 21.8%, suggesting that ~80% of test videos, when tested against a correct target city, were identified as belonging to that city. This result is well above-chance, even as the videos contained very few city-specific audio and metadata features. We have also demonstrated the complementarity of audio and metadata for this task. Howard Lei, Jaeyoung Choi 0002, Gerald Friedland |
ICASSP | 3 |
| 2012 | Multimodal Location Estimation of Consumer Media: Dealing with Sparse Training DataabstractThis article describes a novel approach to the problem of associating geo-locations to consumer-produced multimedia data such as videos and photos that are publicly available on social networking websites such as Flickr. We specifically focus on the case where the available training data is sparse both in absolute numbers as well as geographic coverage when compared to the number of untagged query data. We develop a novel graphical model based framework for the problem of interest and pose the problem of geotagging as one of inference over this graph. The novelty of our algorithm lies in the fact that we jointly estimate the geo-locations of all the query videos, which helps obtain performance improvements over existing algorithms in the literature that process each query video independently. Our system enables the query videos to act as "virtual" training data that effectively bootstrap the geo-tagging process. The quality of the database improves with each additional query video in the system. Further, our modeling provides a generic theoretical framework that can be used to incorporate any other available textual, visual or audio features. We evaluate our algorithm on the MediaEval 2011 Placing Task data set and show that for fixed training data the system performance improves with an increasing number of unlabeled test data. The performance gains are shown to be over 10% as compared to existing algorithms in the literature. Jaeyoung Choi 0002, Gerald Friedland, Venkatesan N. Ekambaram, Kannan Ramchandran |
ICME | 2 |
| 2012 | Where did I go wrong?: Identifying troublesome segments for speaker diarization systemsabstractThe focus of this work is to identify types of segments that are difficult for speaker diarization systems. The diarization outputs of five state-of-the-art systems are analyzed on short/long seg-ments as well as segments surrounding speaker changepoints. We found that for all five systems as the duration of the segment decreased the diarization error rate (DER) increased. Also, seg-ments immediately preceding and following speaker change-points performed much worse than their respective counter-parts. In fact, at least 40 % of the DER for all five systems is attributed to time within 0.5 seconds of a speaker changepoint. We hope the results of this work motivate future improvements of speaker diarization systems. Index Terms: speaker diarization, error analysis, rich transcrip-tion Mary Tai Knox, Nikki Mirghafori, Gerald Friedland |
INTERSPEECH | 3 |
| 2012 | GeoMM'12: ACM international workshop on geotagging and its applications in multimediaabstractGeotagging is the process of adding geographical identification metadata to various media files such as photos, videos, websites, messages, and tweets. It is not limited to GPS sensor data but an extension of current multimedia files with a wide variety of location-specific information. The GeoMM'12 workshop presents research on recent research on geotagging within the context of multimedia analysis. This workshop aims to not only provide more cutting edge algorithms, but also motivate novel applications in this promising field. Liangliang Cao, Gerald Friedland, Martha A. Larson |
ACM Multimedia | 2 |
| 2012 | Privacy concerns in multimedia and their solutionsabstractThis article summarizes the corresponding 3-hour tutorial at ACM Multimedia 2012. Gerald Friedland |
ACM Multimedia | 1 |
| 2012 | AMVA'12: ACM international workshop on audio and multimedia methods for large-scale video analysisabstractMedia sharing sites on the Internet and the one-click upload capability of smartphones have led to a deluge of online multimedia content. Everyday, thousands of videos are uploaded into the web creating an ever-growing demand for methods to make them easier to retrieve, search, and index. While visual information is a very important part of a video, acoustic information often complements it. This is especially true for the analysis of consumer-produced, "unconstrained" videos from social media networks, such as YouTube uploads or Flickr content. The goal of the 1st ACM International Workshop on Audio and Multimedia Methods for Large-Scale Video Analysis (AMVA) is to bring together researchers and practitioners in this newly emerging field, and to foster discussion on future directions of the topic by providing a forum for focused exchanges on new ideas, developments, and results. The aim is to build a strong community and a venue that at some point can become its own conference. Gerald Friedland, Daniel P. W. Ellis, Florian Metze |
ACM Multimedia | 1 |
| 2012 | Name that room: room identification using acoustic features in a recordingabstractThis paper presents a system for identifying the room in an audio or video recording through the analysis of acoustical properties. The room identification system was tested using a corpus of 13440 reverberant audio samples. With no common content between the training and testing data, an accuracy of 61% for musical signals and 85% for speech signals was achieved. This approach could be applied in a variety of scenarios where knowledge about the acoustical environment is desired, such as location estimation, music recommendation, or emergency response systems. Nils Peters, Howard Lei, Gerald Friedland |
ACM Multimedia | 3 |
| 2012 | The ICSI RT-09 Speaker Diarization SystemabstractThe speaker diarization system developed at the International Computer Science Institute (ICSI) has played a prominent role in the speaker diarization community, and many researchers in the rich transcription community have adopted methods and techniques developed for the ICSI speaker diarization engine. Although there have been many related publications over the years, previous articles only presented changes and improvements rather than a description of the full system. Attempting to replicate the ICSI speaker diarization system as a complete entity would require an extensive literature review, and might ultimately fail due to component description version mismatches. This paper therefore presents the first full conceptual description of the ICSI speaker diarization system as presented to the National Institute of Standards Technology Rich Transcription 2009 (NIST RT-09) evaluation, which consists of online and offline subsystems, multi-stream and single-stream implementations, and audio and audio-visual approaches. Some of the components, such as the online system, have not been previously described. The paper also includes all necessary preprocessing steps, such as Wiener filtering, speech activity detection and beamforming. Gerald Friedland, Adam Janin, David Imseng, Xavier Anguera Miró, Luke R. Gottlieb, Marijn Huijbregts, Mary Tai Knox, Oriol Vinyals |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Speaker Diarization: A Review of Recent ResearchabstractSpeaker diarization is the task of determining “who spoke when?” in an audio or video recording that contains an unknown amount of speech and also an unknown number of speakers. Initially, it was proposed as a research topic related to automatic speech recognition, where speaker diarization serves as an upstream processing step. Over recent years, however, speaker diarization has become an important key technology for many tasks, such as navigation, retrieval, or higher level inference on audio data. Accordingly, many important improvements in accuracy and robustness have been reported in journals and conferences in the area. The application domains, from broadcast news, to lectures and meetings, vary greatly and pose different problems, such as having access to multiple microphones and multimodal information or overlapping speech. The most recent review of existing technology dates back to 2006 and focuses on the broadcast news domain. In this paper, we review the current state-of-the-art, focusing on research developed since 2006 that relates predominantly to speaker diarization for conference meetings. Finally, we present an analysis of speaker diarization performance as reported through the NIST Rich Transcription evaluations on meeting data and identify important areas for future research. Xavier Anguera Miró, Simon Bozonnet, Nicholas W. D. Evans, Corinne Fredouille, Gerald Friedland, Oriol Vinyals |
IEEE Trans. Speech Audio Process. | 5 |
| 2011 | Fast speaker diarization using a high-level scripting languageabstractMost current speaker diarization systems use agglomerative clustering of Gaussian Mixture Models (GMMs) to determine “who spoke when” in an audio recording. While state-of-the-art in accuracy, this method is computationally costly, mostly due to the GMM training, and thus limits the performance of current approaches to be roughly real-time. Increased sizes of current datasets require processing of hundreds of hours of data and thus make more efficient processing methods highly desirable. With the emergence of highly parallel multicore and manycore processors, such as graphics processing units (GPUs), one can re-implement GMM training to achieve faster than real-time performance by taking advantage of parallelism in the training computation. However, developing and maintaining the complex low-level GPU code is difficult and requires a deep understanding of the hardware architecture of the parallel processor. Furthermore, such low-level implementations are not readily reusable in other applications and not portable to other platforms, limiting programmer productivity. In this paper we present a speaker diarization system captured in under 50 lines of Python that achieves 50-250× faster than real-time performance by using a specialization framework to automatically map and execute computationally intensive GMM training on an NVIDIA GPU, without significant loss in accuracy. Ekaterina Gonina, Gerald Friedland, Henry Cook, Kurt Keutzer |
ASRU | 2 |
| 2011 | User verification: Matching the uploaders of videos across accountsabstractThis article presents an attempt to link the uploaders of videos based on the audio track of the videos. Using a subset of the MediaEval Placing Task's Flickr video set, which is labeled with the uploader's name, we conducted an experiment with a similar setup as a typical NIST speaker recognition evaluation run. Based on the assumption that the audio might be matched in various ways (speaker, channel, environmental noise, etc.), we trained one of ICSI's simplified speaker recognition systems on the audio tracks of the Flickr videos. Note that since the selection of videos is essentially random, the audio track can contain any sounds. We obtain an equal error rate of 36.7% on 312 videos with 11,550 trials. The result has implications for audio research, security applications, and raises privacy concerns. Howard Lei, Jaeyoung Choi 0002, Adam Janin, Gerald Friedland |
ICASSP | 4 |
| 2011 | Improved Overlapped Speech Handling for Speaker Diarization
Kofi Boakye, Oriol Vinyals, Gerald Friedland |
INTERSPEECH | 3 |
| 2011 | On the Applicability of Speaker Diarization to Audio Concept Detection for Multimedia RetrievalabstractRecently, audio concepts emerged as a useful building block in multimodal video retrieval systems. Information like "this file contains laughter", "this file contains engine sounds" or "this file contains slow music" can significantly improve purely visual based retrieval. The weak point of current approaches to audio concept detection is that they heavily rely on human annotators. In most approaches, audio material is manually inspected to identify relevant concepts. Then instances that contain examples of relevant concepts are selected -- again manually -- and used to train concept detectors. This approach comes with two major disadvantages: (1) it leads to rather abstract audio concepts that hardly cover the audio domain at hand and (2) the way human annotators identify audio concepts likely differs from the way a computer algorithm clusters audio data -- introducing additional noise in training data. This paper explores whether unsupervized audio segementation systems can be used to identify useful audio concepts by analyzing training data automatically and whether these audio concepts can be used for multimedia document classification and retrieval. A modified version of the ICSI (International Computer Science Institute) speaker diarization system finds segments in an audio track that have similar perceptual properties and groups these segments. This article provides an in-depth analysis on the statistic properties of similar acoustic segments identified by the diarization system in a predefined document set and the theoretical fitness of this approach to discern one document class from another. Robert Mertens 0003, Po-Sen Huang, Luke R. Gottlieb, Gerald Friedland, Ajay Divakaran |
ISM | 4 |
| 2011 | Automatic tagging and geotagging in video collections and communitiesabstractAutomatically generated tags and geotags hold great promise to improve access to video collections and online communities. We overview three tasks offered in the MediaEval 2010 benchmarking initiative, for each, describing its use scenario, definition and the data set released. For each task, a reference algorithm is presented that was used within MediaEval 2010 and comments are included on lessons learned. The Tagging Task, Professional involves automatically matching episodes in a collection of Dutch television with subject labels drawn from the keyword thesaurus used by the archive staff. The Tagging Task, Wild Wild Web involves automatically predicting the tags that are assigned by users to their online videos. Finally, the Placing Task requires automatically assigning geo-coordinates to videos. The specification of each task admits the use of the full range of available information including user-generated metadata, speech recognition transcripts, audio, and visual features. Martha A. Larson, Mohammad Soleymani 0001, Pavel Serdyukov, Stevan Rudinac, Christian Wartena, Vanessa Murdock 0001, Gerald Friedland, Roeland Ordelman, Gareth J. F. Jones |
ICMR | 7 |
| 2011 | Acoustic and multimodal processing for multimedia content analysisabstractThis article summarizes the corresponding 3-hour tutorial at ACM Multimedia 2011. Gerald Friedland |
ACM Multimedia | 1 |
| 2011 | Video2GPS: a demo of multimodal location estimation on flickr videosabstractThe following article describes our demo of an approach to determine the geo-coordinates of the recording place of Flickr videos based on both textual metadata and visual cues. The underlying system has been tested on the MediaEval 2010 Placing Task evaluation data, which consists of 5091 unfiltered test videos is able to classify 14% of the videos to within an accuracy of 10m. Gerald Friedland, Jaeyoung Choi 0002, Adam Janin |
ACM Multimedia | 1 |
| 2011 | Sherlock holmes' evil twin: on the impact of global inference for online privacyabstractUser-supplied content--in the form of photos, videos, and text--is a crucial ingredient to many web sites and services today. However, many users who provide content do not realize that their uploads may be leaking personal information in forms hard to intuitively grasp. Correlation of seemingly innocuous information can create inference chains that tell much more about individuals than they are aware of revealing. We contend that adversaries can systematically exploit such relationships by correlating information from different sources in what we term global inference attacks: assembling a comprehensive understanding from individual pieces found at a variety of locations, Sherlock-style. Not only are such attacks already technically viable given the capabilities that today's multimedia content analysis and correlation technologies readily provide, but we also find business models that provide adversaries with powerful incentives for pursuing them. Gerald Friedland, Gregor Maier, Robin Sommer, Nicholas Weaver |
NSPW | 1 |
| 2011 | Estimating Dominance in Multi-Party Meetings Using Speaker DiarizationabstractWith the increase in cheap commercially available sensors, recording meetings is becoming an increasingly practical option. With this trend comes the need to summarize the recorded data in semantically meaningful ways. Here, we investigate the task of automatically measuring dominance in small group meetings when only a single audio source is available. Past research has found that speaking length as a single feature, provides a very good estimate of dominance. For these tasks we use speaker segmentations generated by our automated faster than real-time speaker diarization algorithm, where the number of speakers is not known beforehand. From user-annotated data, we analyze how the inherent variability of the annotations affects the performance of our dominance estimation method. We primarily focus on examining of how the performance of the speaker diarization and our dominance tasks vary under different experimental conditions and computationally efficient strategies, and how this would impact on a practical implementation of such a system. Despite the use of a state-of-the-art speaker diarization algorithm, speaker segments can be noisy. On conducting experiments on almost 5 hours of audio-visual meeting data, our results show that the dominance estimation is robust to increasing diarization noise. Hayley Hung, Gerald Friedland, Daniel Gatica-Perez |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | An adaptive initialization method for speaker Diarization based on prosodic featuresabstractThe following article presents a novel, adaptive initialization scheme that can be applied to most state-of-the-art Speaker Diarization algorithms, i.e. algorithms that use agglomerative hierarchical clustering with Bayesian Information Criterion (BIC) and Gaussian Mixture Models (GMMs) of frame-based cepstral features (MFCCs). The initialization method is a combination of the recently proposed “adaptive seconds per Gaussian” (ASPG) method and a new pre-clustering and number of initial clusters estimation method based on prosodic features. The presented initialization method has two important advantages. First, the method requires no manual tuning and is robust against file length and speaker count variations. Second, the method outperforms our previously used initialization methods on all benchmark files that were presented in the 2006, 2007, and 2009 NIST Rich Transcription (RT) evaluations and results in a Diarization Error Rate (DER) improvement of up to 67% (relative). David Imseng, Gerald Friedland |
ICASSP | 2 |
| 2010 | Leveraging speaker diarization for meeting recognition from distant microphonesabstractWe investigate using state-of-the-art speaker diarization output for speech recognition purposes. While it seems obvious that speech recognition could benefit from the output of speaker diarization (“Who spoke when”) for effective feature normalization and model adaptation, such benefits have remained elusive in the very challenging domain of meeting recognition from distant microphones. In this study, we show that recognition gains are possible by careful post-processing of the diarization output. Still, recognition accuracy may suffer when the underlying diarization system performs worse than expected, even compared to far less sophisticated speaker-clustering techniques. We obtain a more accurate and robust overall system by combining recognition output with multiple speaker segmentations and clusterings. We evaluate our methods on data from the 2009 NIST Rich Transcription meeting recognition evaluation. Andreas Stolcke, Gerald Friedland, David Imseng |
ICASSP | 2 |
| 2010 | System output combination for improved speaker diarizationabstractInternational audience Simon Bozonnet, Nicholas W. D. Evans, Xavier Anguera Miró, Oriol Vinyals, Gerald Friedland, Corinne Fredouille |
INTERSPEECH | 5 |
| 2010 | Multimodal speaker diarization using oriented optical flow histogramsabstractSpeaker diarization is the task of partitioning an input stream into speaker homogeneous regions, or in other words, to determine ”who spoke when.” While approaches to this problem have traditionally relied entirely on the audio stream, the availability of accompanying video streams in recent diarization corpora has prompted the study of methods based on multimodal audio-visual features. In this work, we propose the use of robust video features based on oriented optical flow histograms. Using the state-of-the art ICSI diarization system, we show that, when combined with standard audio features, these features improve the diarization error rate by 14% percent over an audio-only baseline. Mary Tai Knox, Gerald Friedland |
INTERSPEECH | 2 |
| 2010 | A hybrid approach to online speaker diarizationabstractThis article presents a low-latency speaker diarization system (“who is speaking now?”) based on a hybrid approach that combines a traditional offline speaker diarization system (“who spoke when?”) with an online speaker identification system. The system fulfills all requirements of the diarization task, i.e. it does not need any a-priori information about the input, including no specific speaker models. After an initialization phase the approach allows a low-latency decision on the current speaker with an accuracy that is close to the underlying offline diarization system. The article describes the approach, evaluates the robustness of the system, and analyzes the latency/accuracy trade-off. Index Terms :S peaker Diarization, online, incremental, hybrid Carlos Vaquero, Oriol Vinyals, Gerald Friedland |
INTERSPEECH | 3 |
| 2010 | Discriminative training for hierarchical clustering in speaker diarizationabstractIn this paper, we propose a discriminative extension to agglomerative hierarchical clustering, a typical technique for speaker diarization, that fits seamlessly with most state-of-the art diarization algorithms. We propose to use maximum mutual information using bootstrapping i.e., initial predictions are used as input for retraining of models in an unsupervised fashion. This article describes this new approach, analyzes its behavior, and presents results on the official NIST Rich Transcription datasets. We show an absolute improvement of 4% DER with respect to the generative approach baseline. We also observe a strong correlation between the original error and the amount of improvement, that is, the better our predicted labels are, the more gain we obtain from discriminative training, which we interpret as a strong indication for the high potential of the extension. Oriol Vinyals, Gerald Friedland, Nelson Morgan |
INTERSPEECH | 2 |
| 2010 | Parallelizing Speaker-Attributed Speech Recognition for Meeting BrowsingabstractThe following article presents an application for browsing meeting recordings by speaker and keyword which we call the Meeting Diarist. The goal of the system is to enable browsing of the content with rich meta-data in a graphical user interface shortly after the end of meeting, even when the application runs on a contemporary laptop. We there-fore developed novel parallel methods for speaker diarization and multi-hypothesis speech recognition that are optimized to run on multicore and many core architectures. This paper presents the underlying parallel speaker diarization and speech recognition realizations, a comparison of results based on NIST RT07 evaluation data, and a description of the final application. Gerald Friedland, Jike Chong, Adam Janin |
ISM | 1 |
| 2010 | Multimodal location estimationabstractIn this article we define a multimedia content analysis problem, which we call multimodal location estimation: Given a video/image/audio file, the task is to determine where it was recorded. A single indication, such as a unique landmark, might already pinpoint a location precisely. In most cases, however, a combination of evidence from the visual and the acoustic domain will only narrow down the set of possible answers. Therefore, approaches to tackle this task should be inherently multimedia. While the task is hard, in fact sometimes unsolvable, training data can be leveraged from the Internet in large amounts. Moreover, even partially successful automatic estimation of location opens up new possibilities in video content matching, archiving, and organization. It could revolutionize law enforcement and computer-aided intelligence agency work, especially since both semi-automatic and fully automatic approaches would be possible. In this article, we describe our idea of growing multimodal location estimation as a research field in the multimedia community. Based on examples and scenarios, we propose a multimedia approach to leverage cues from the visual and the acoustic portions of a video as well as from given metadata. We also describe experiments to estimate the amount of available training data that could potentially be used as publicly available infrastructure for research in this field. Finally, we present an initial set of results based on acoustic and visual cues and discuss the massive challenges involved and some possible paths to solutions. Gerald Friedland, Oriol Vinyals, Trevor Darrell |
ACM Multimedia | 1 |
| 2010 | Joke-o-Mat HD: browsing sitcoms with human derived transcriptsabstractJoke-o-mat HD is a system that allows a user to navigate sitcoms (such as Seinfeld) by "narrative themes", including scenes, punchlines, and dialog segments. The themes can be filtered by the main actors and by keyword. For example, the user can select to see only punchlines by Kramer that contain the word "armoire". The system infers the narrative themes using segmentation of the audio track into laughter, actors, words, and music. The segmentation can be generated either by an expert annotator, via automatic methods, or by exploiting human derived (HD) "found" data such as fan-generated scripts and closed captions. We demonstrate browsing one episode of Seinfeld using all three methods of generating segmentations. Adam Janin, Luke R. Gottlieb, Gerald Friedland |
ACM Multimedia | 3 |
| 2010 | Precise indoor localization using smart phonesabstractWe present an indoor localization application leveraging the sensing capabilities of current state of the art smart phones. To the best of our knowledge, our application is the first one to be implemented in smart phones and integrating both offline and online phases of fingerprinting, delivering an accuracy of up to 1.5 meters. In particular, we have studied the possibilities offered by WiFi radio, cellular communications radio, accelerometer and magnetometer, already embedded in smart phones, with the intention to build a multimodal solution for localization. We have also implemented a new approach for the statistical processing of radio signal strengths, showing that it can outperform existing deterministic techniques. Eladio Martin, Oriol Vinyals, Gerald Friedland, Ruzena Bajcsy |
ACM Multimedia | 3 |
| 2010 | 3rd international workshop on automated information extraction in media productionabstractThe third Workshop on Automated Information Extraction in Media Production (AIEMPro10) aims at fostering exchange of ideas and of practices between leading experts in research and leading actors in the media community, in order to catalyze the migration towards new ways of producing media content, aided by large scale introduction of tools for automated multimedia analysis and understanding. Furthermore, the workshop helps researchers in better understanding what are some real-life key requirements which would enable their scientific developments come into wider adoption. Alberto Messina, Robbie De Sutter, Jean-Pierre Evain, Masanori Sano, Gerald Friedland |
ACM Multimedia | 5 |
| 2010 | Cybercasing the Joint: On the Privacy Implications of Geo-Tagging
Gerald Friedland, Robin Sommer |
HotSec | 1 |
| 2010 | Tuning-Robust Initialization Methods for Speaker DiarizationabstractThis paper investigates a typical speaker diarization system regarding its robustness against initialization parameter variation and presents a method to reduce manual tuning of these values significantly. The behavior of an agglomerative hierarchical clustering system is studied to determine which initialization parameters impact accuracy most. We show that the accuracy of typical systems is indeed very sensitive to the values chosen for the initialization parameters and factors such as the duration of speech in the recording. We then present a solution that reduces the sensitivity of the initialization values and therefore reduces the need for manual tuning significantly while at the same time increasing the accuracy of the system. For short meetings extracted from the previous (2006, 2007, and 2009) National Institute of Standards and Technology (NIST) Rich Transcription (RT) evaluation data, the decrease of the diarization error rate is up to 50% relative. The approach consists of a novel initialization parameter estimation method for speaker diarization that uses agglomerative clustering with Bayesian information criterion (BIC) and Gaussian mixture models (GMMs) of frame-based cepstral features (MFCCs). The estimation method balances the relationship between the optimal value of the seconds of speech data per Gaussian and the duration of the speech data and is combined with a novel nonuniform initialization method. This approach results in a system that performs better than the current ICSI baseline engine on datasets of the NIST RT evaluations of the years 2006, 2007, and 2009. David Imseng, Gerald Friedland |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Dialocalization: Acoustic speaker diarization and visual localization as joint optimization problemabstractThe following article presents a novel audio-visual approach for unsupervised speaker localization in both time and space and systematically analyzes its unique properties. Using recordings from a single, low-resolution room overview camera and a single far-field microphone, a state-of-the-art audio-only speaker diarization system (speaker localization in time) is extended so that both acoustic and visual models are estimated as part of a joint unsupervised optimization problem. The speaker diarization system first automatically determines the speech regions and estimates “who spoke when,” then, in a second step, the visual models are used to infer the location of the speakers in the video. We call this process “dialocalization.” The experiments were performed on real-world meetings using 4.5 hours of the publicly available AMI meeting corpus. The proposed system is able to exploit audio-visual integration to not only improve the accuracy of a state-of-the-art (audio-only) speaker diarization, but also adds visual speaker localization at little incremental engineering and computation costs. The combined algorithm has different properties, such as increased robustness, that cannot be observed in algorithms based on single modalities. The article describes the algorithm, presents benchmarking results, explains its properties, and systematically discusses the contributions of each modality. Gerald Friedland, Chuohao Yeo, Hayley Hung |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2009 | Robust Speaker Diarization for short speech recordingsabstractWe investigate a state-of-the-art speaker diarization system regarding its behavior on meetings that are much shorter (from 500 seconds down to 100 seconds) than those typically analyzed in speaker diarization benchmarks. First, the problems inherent to this task are analyzed. Then, we propose an approach that consists of a novel initialization parameter estimation method for typical state-of-the-art diarization approaches. The estimation method balances the relationship between the optimal value of the duration of speech data per Gaussian and the duration of the speech data, which is verified experimentally for the first time in this article. As a result, the diarization error rate for short meetings extracted from the 2006, 2007, and 2009 NIST RT evaluation data is decreased by up to 50% relative. David Imseng, Gerald Friedland |
ASRU | 2 |
| 2009 | Multi-modal speaker diarization of real-world meetings using compressed-domain video featuresabstractSpeaker diarization is originally defined as the task of determining ldquowho spoke whenrdquo given an audio track and no other prior knowledge of any kind. The following article shows a multi-modal approach where we improve a state-of-the-art speaker diarization system by combining standard acoustic features (MFCCs) with compressed domain video features. The approach is evaluated on over 4.5 hours of the publicly available AMI meetings dataset which contains challenges such as people standing up and walking out of the room. We show a consistent improvement of about 34% relative in speaker error rate (21% DER) compared to a state-of-the-art audio-only baseline. Gerald Friedland, Hayley Hung, Chuohao Yeo |
ICASSP | 1 |
| 2009 | Fusing short term and long term features for improved speaker diarizationabstractThe following article shows how a state-of-the-art speaker diarization system can be improved by combining traditional short-term features (MFCCs) with prosodic and other long-term features. First, we present a framework to study the speaker discriminability of 70 different long-term features. Then, we show how the top-ranked long-term features can be combined with short-term features to increase the accuracy of speaker diarization. The results were measured on standardized data sets (NIST RT) and show a consistent improvement of about 30% relative in diarization error rate compared to the best system presented at the NIST evaluation in 2007. This result was also verified on a wide set of meetings, which we call CombDev, that contains 21 meetings from previous evaluations. Since the prosodic and long-term features were selected using a diarization-independent speaker-discriminability study, we are confident that the same features are able to improve other systems that perform similar tasks Gerald Friedland, Oriol Vinyals, Christian Müller 0014 |
ICASSP | 1 |
| 2009 | Using Artistic Markers and Speaker Identification for Narrative-Theme Navigation of Seinfeld EpisodesabstractThis article describes a system to navigate Seinfeld episodes based on acoustic event detection and speaker identification of the audio track and subsequent inference of narrative themes based on genre-specific production rules. The system distinguishes laughter, music, and other noise as well as speech segments. Speech segments are then identified against pre-trained speaker models. Given this segmentation and the artistic production rules that underlie the "situation comedy" genre and Seinfeld in particular, the system enables a user to browse an episode by scene, punchline, and dialog segments. The themes can be filtered by the main actors, e.g. the user can choose to see only punchlines by Jerry and Kramer. Based on the length of the laughter, the top-5 punchlines are identified and presented to the user. The segmentation is then presented in an Applet-based graphical video browser that is intended to extend a typical YouTube videoplayer. Gerald Friedland, Luke R. Gottlieb, Adam Janin |
ISM | 1 |
| 2009 | Multimodal interfaces for automotive applications (MIAA)abstractThis paper summarizes the main objectives of the IUI workshop W2 on multimodal interfaces for automotive applications. Christian Müller 0014, Gerald Friedland |
IUI | 2 |
| 2009 | Joke-o-mat: browsing sitcoms punchline by punchlineabstractThis paper summarizes our contribution to the Yahoo! task of the ACM Multimedia Grand Challenge. This challenge asks for the robust automatic segmentation of videos according to "narrative themes". Based on the automatic segmentation methods presented in [1] and partly [2], we describe a system to navigate Seinfeld episodes based on automatic segmentation of the audio track only. The system distinguishes laughter, music, and other noise as well as speech segments. Speech segments are identified against pre-trained speaker models of the actors. Given this segmentation and the artistic production rules that underlie the genre situation comedy and Seinfeld in particular, the system enables a user to browse an episode by scene, by punchline, and by dialog segments. The themes can be filtered by the main actors, e.g. the user can select to see only punchlines by Jerry and Kramer. Based on the length of the laughter, the top 5 punchlines are also identified and presented to the user. Gerald Friedland, Luke R. Gottlieb, Adam Janin |
ACM Multimedia | 1 |
| 2009 | Visual speaker localization aided by acoustic modelsabstractThe following paper presents a novel audio-visual approach for unsupervised speaker locationing. Using recordings from a single, low-resolution room overview camera and a single far-field microphone, a state-of-the art audio-only speaker localization system (traditionally called speaker diarization) is extended so that both acoustic and visual models are estimated as part of a joint unsupervised optimization problem. The speaker diarization system first automatically determines the number of speakers and estimates "who spoke when", then, in a second step, the visual models are used to infer the location of the speakers in the video. The experiments were performed on real-world meetings using 4.5 hours of the publicly available AMI meeting corpus. The proposed system is able to exploit audio-visual integration to not only improve the accuracy of a state-of-the-art (audio-only) speaker diarization, but also adds visual speaker locationing at little incremental engineering and computation costs. Gerald Friedland, Chuohao Yeo, Hayley Hung |
ACM Multimedia | 1 |
| 2009 | Prosodic and other Long-Term Features for Speaker DiarizationabstractSpeaker diarization is defined as the task of determining ldquowho spoke whenrdquo given an audio track and no other prior knowledge of any kind. The following article shows how a state-of-the-art speaker diarization system can be improved by combining traditional short-term features (MFCCs) with prosodic and other long-term features. First, we present a framework to study the speaker discriminability of 70 different long-term features. Then, we show how the top-ranked long-term features can be combined with short-term features to increase the accuracy of speaker diarization. The results were measured on standardized datasets (NIST RT) and show a consistent improvement of about 30% relative in diarization error rate compared to the best system presented at the NIST evaluation in 2007. Gerald Friedland, Oriol Vinyals, Christian Müller 0014 |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Overlapped speech detection for improved speaker diarization in multiparty meetingsabstractState-of-the-art speaker diarization systems for meetings are now at a point where overlapped speech contributes significantly to the errors made by the system. However, little if no work has yet been done on detecting overlapped speech. We present our initial work toward developing an overlap detection system for improved meeting diarization. We investigate various features, with a focus on high-precision performance for use in the detector, and examine performance results on a subset of the AMI Meeting Corpus. For the high-quality signal case of a single mixed-headset channel signal, we demonstrate a relative improvement of about 7.4% DER over the baseline diarization system, while for the more challenging case of the single far-field channel signal relative improvement is 3.6%. We also outline steps towards improvement and moving beyond this initial phase. Kofi Boakye, B. Trueba-Hornero, Oriol Vinyals, Gerald Friedland |
ICASSP | 4 |
| 2008 | Estimating the dominant person in multi-party conversations using speaker diarization strategiesabstractIn this paper, we apply speaker diarization strategies from a single source to the task of estimating the dominant person in a group meeting. Previous work has shown that speaking length is strongly correlated with perceived dominance. Here we investigate this in more depth by considering two dominance tasks where there is full and majority agreement amongst ground-truth annotators. In addition, we investigate how 24 different speed-up and algorithmic strategies, and source types lead to interesting outcomes when applied to dominance estimation. We obtained the best performance of 77% using our slowest scheme and a single distant microphone (SDM). Within the top 3 out of 24 performing experiments in both dominance tasks, we show that we can use the furthest SDM, with no prior knowledge of the number of speakers and the fastest diarization scheme, which performs 1.3 times faster than real-time. Hayley Hung, Gerald Friedland, Daniel Gatica-Perez |
ICASSP | 3 |
| 2008 | Two's a crowd: improving speaker diarization by automatically identifying and excluding overlapped speechabstractWe present an update to our initial work [1] on overlapped speech detection for improving speaker diarization. Specifi-cally, we describe the addition of new features and feature warp-ing techniques that improve segmenter and, consequently, di-arization performance. We also demonstrate improved diariza-tion performance by additionally using overlap segment infor-mation in a new diarization pre-processing step which excludes overlap segments from speaker clustering. On a subset of the AMI Meeting Corpus we show that this overlap exclusion step nearly triples the relative improvement of diarization error rate as compared to overlap segment post-processing alone. Index Terms: speaker diarization, overlap detection 1. Kofi Boakye, Oriol Vinyals, Gerald Friedland |
INTERSPEECH | 3 |
| 2008 | Modulation spectrogram features for improved speaker diarizationabstractWe propose the use of modulation spectrogram features in speaker diarization. These features carry longer term characteristics of the acoustic signals than the widely used MFCCs, thus providing potential improvement by using both features in combination. Using the state-of-the-art ICSI speaker diarization system, an improvement of 20.77 % relative DER is obtained on the NIST Rich Transcription 2007 task with respect to the MFCC only system. Index Terms: modulation spectrogram, speaker diarization 1. Oriol Vinyals, Gerald Friedland |
INTERSPEECH | 2 |
| 2008 | A Hardware-Independent Fast Logarithm Approximation with Adjustable AccuracyabstractMany multimedia applications rely on the computation of logarithms, for example, when estimating log-likelihoods for Gaussian Mixture Models. Knowing of the demand to compute logarithms and other basic math functions rapidly, many hardware manufacturers provide libraries to perform calculations in hardware. Of course, these libraries are especially popular for the use in computer vision or audio analysis algorithms where a large amounts of data have to be processed. A downside of using specialized hardware though is that it increases the investment cost and the user is forced to use the same hardware, which is especially cumbersome when algorithms optimized for different specialized hardware are to be combined. This article presents the realization of a novel platform-independent, fast C-language implementation of the logarithm function. The idea behind the approach is to take advantage of the large amount of cache available in current processors. The logarithm implementation is compared to the current state of the art and we demonstrate the practical use of the algorithm in an actual speech analysis application. Oriol Vinyals, Gerald Friedland |
ISM | 2 |
| 2008 | Multimedia education: can we find unity in diversity?abstractThe field of multimedia is composed of a variety of research areas. This diversity makes multimedia such a special and interesting research field. However, the different vocabularies, methods, and cultures of the involved communities also introduce barriers that make it difficult to teach the field as a unified subject. The panel invites experts to discuss if and how we can teach the diversity in multimedia as a single subject and how we as researchers and educators can help to foster this goal. The following text provides background information to the topic and introduces the organizer's hypotheses to be discussed at the panel. Gerald Friedland, Wolfgang Hürst, Lars Knipping |
ACM Multimedia | 1 |
| 2008 | Live speaker identification in conversationsabstractThe following article describes our technical demonstration of an online speaker identification system for conversations. A laptop with an internal microphone is centrally placed in the table of a meeting room. The system is able to identify the current speaker independent of spoken text or language with a latency of about 1.5 seconds and an accuracy of about 85% (as evaluated against the NIST RT benchmark). A Java GUI shows the image of the current speaker along with a timeline containing past speakers. Speakers are added to the system's database using a one-minute training procedure. Gerald Friedland, Oriol Vinyals |
ACM Multimedia | 1 |
| 2007 | A fast-match approach for robust, faster than real-time speaker diarizationabstractDuring the past few years, speaker diarization has achieved satisfying accuracy in terms of speaker Diarization Error Rate (DER). The most successful approaches, based on agglomerative clustering, however, exhibit an inherent computational complexity which makes real-time processing, especially in combination with further processing steps, almost impossible. In this article we present a framework to speed up agglomerative clustering speaker diarization. The basic idea is to adopt a computationally cheap method to reduce the hypothesis space of the more expensive and accurate model selection via Bayesian Information Criterion (BIC). Two strategies based on the pitch-correlogram and the unscented-transform based approximation of KL-divergence are used independently as a fast-match approach to select the most likely clusters to merge. We performed the experiments using the existing ICSI speaker diarization system. The new system using KL-divergence fast-match strategy only performs 14% of total BIC comparisons needed in the baseline system, speeds up the system by 41% without affecting the speaker Diarization Error Rate (DER). The result is a robust and faster than real-time speaker diarization system. Oriol Vinyals, Gerald Friedland, Christian Müller 0014, Nikki Mirghafori, Chuck Wooters |
ASRU | 3 |
| 2007 | Using audio and video features to classify the most dominant person in a group meetingabstractThe automated extraction of semantically meaningful information from multi-modal data is becoming increasingly necessary due to the escalation of captured data for archival. A novel area of multi-modal data labelling, which has received relatively little attention, is the automatic estimation of the most dominant person in a group meeting. In this paper, we provide a framework for detecting dominance in group meetings using different audio and video cues. We show that by using a simple model for dominance estimation we can obtain promising results. Hayley Hung, Dinesh Babu Jayagopi, Chuohao Yeo, Gerald Friedland, Sileye O. Ba, Jean-Marc Odobez, Kannan Ramchandran, Nikki Mirghafori, Daniel Gatica-Perez |
ACM Multimedia | 4 |
| 2006 | A Practical Approach to Boundary Accurate Multi-Object Extraction from Still Images and VideosabstractThis paper presents recent improvements on a previous published practical approach to object segmentation from still images called SIOX: simple interactive object extraction. The basic method has been improved further and applied to a variety of applications. In the case of videos, this technique allows scale and rotation invariant real-time tracking of foreground objects including the classification of newly introduced objects. In the case of still-images, the method has been extended to cope with highly detailed textures with sub-pixel accuracy. The approach is robust in the presence of noise and can be applied to a variety of problems where objects have to be extracted, tracked and/or identified. The approach has been released as an open source framework. We discuss various applications of the algorithm and incorporated experiences and user feedback from several projects that have begun to integrate the algorithm Gerald Friedland, Kristian Jantz, Tobias Lenz, Fabian Wiesel, Raúl Rojas 0001 |
ISM | 1 |
| 2006 | Human-Centered Webcasting of Interactive-Whiteboard LecturesabstractIn our system for recording and transmitting lectures over the Internet, the board content is transmitted as vector graphics, producing thus a high quality image, while the video of the lecturer is sent as a separate stream. It is easy for the viewer to read the board but the lecturer appears in a separate window. As a result, two areas of the screen are competing for the viewer's attention, causing the widely known split attention effect. To eliminate this problem, the lecturer is extracted from the video stream and his or her image is pasted onto the board image at video stream rates. The lecturer can be dimmed from opaque to semitransparent, or even transparent. The article presents a detailed analysis of the underlying psychological problems and explains the multimedia techniques that are applied to achieve the solution Gerald Friedland, Raúl Rojas 0001 |
ISM | 1 |
| 2005 | SIOX: Simple Interactive Object Extraction in Still ImagesabstractThe following article presents an approach for interactive foreground extraction in still images that is currently being integrated into the GIMP. The presented approach has been derived from color signatures, a technique originating from image retrieval. The article explains the algorithm and presents some benchmark results to show the improvements in speed and accuracy compared to state of the art solutions. The article also describes how the algorithm can easily be adapted for video segmentation tasks. Gerald Friedland, Kristian Jantz, Raúl Rojas 0001 |
ISM | 1 |
| 2005 | The Virtual Technician: An Automatic Software Enhancer for Audio Recording in Lecture Halls
Gerald Friedland, Kristian Jantz, Lars Knipping, Raúl Rojas 0001 |
KES (1) | 1 |
| 2003 | Web Based Education as a Result of AI Supported Classroom Teaching
Gerald Friedland, Lars Knipping, Raúl Rojas 0001, Ernesto Tapia |
KES | 1 |