Viktor Rozgic

dblp:43/2871 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
abstract
Agentic search such as Deep Research systems-where agents autonomously browse the web, synthesize information, and return comprehensive citation-backed answers-represents a major shift in how users interact with web-scale information. While promising greater efficiency and cognitive offloading, the growing complexity and open-endedness of agentic search have outpaced existing evaluation benchmarks and methodologies, which largely assume short search horizons and static answers. In this paper, we introduce Mind2Web 2, a benchmark of 130 realistic, high-quality, and long-horizon tasks that require real-time web browsing and extensive information synthesis, constructed with over 1000 hours of human labor. To address the challenge of evaluating time-varying and complex answers, we propose a novel Agent-as-a-Judge framework. Our method constructs task-specific judge agents based on a tree-structured rubric design to automatically assess both answer correctness and source attribution. We conduct a comprehensive evaluation of ten frontier agentic search systems and human performance, along with a detailed error analysis to draw insights for future development. The best-performing system, OpenAI Deep Research, can already achieve 50-70% of human performance while spending half the time, highlighting its great potential. Altogether, Mind2Web 2 provides a rigorous foundation for developing and benchmarking the next generation of agentic search systems.
Boyu Gou, Zanming Huang, Yuting Ning, Yu Gu 0016, Michael Lin, Weijian Qi, Andrei Kopanev, Botao Yu, Bernal Jimenez Gutierrez, Yiheng Shu, Chan Hee Song, Jiaman Wu, Hanane Nour Moussa, Tianshu Zhang 0001, Yifei Li 0005, Tianci Xue, Zeyi Liao, Kai Zhang 0033, Boyuan Zheng 0001, Zhaowei Cai, Viktor Rozgic, Morteza Ziyadi, Huan Sun 0001, Yu Su 0001
NeurIPS23
2023 FedRPO: Federated Relaxed Pareto Optimization for Acoustic Event Classification
abstract
Performance and robustness of real-world Acoustic Event Classification (AEC) solutions depend on ability to train on diverse data from wide range of end-point devices and acoustic environments. Federated Learning (FL) provides a framework to leverage annotated and non-annotated AEC data from servers and client devices in a privacy preserving manner. In this work we propose a novel Federated Relaxed Pareto Optimization (FedRPO) method for semi-supervised FL with heterogeneous client data. In contrast to federated averaging class of FL algorithms (fedAvg) that perform unconstrained weighted aggregation across all data sources, FedRPO enables special treatment of data with high quality annotations vs. data with pseudo-labels of unknown, varying qualities. In particular, FedRPO computes the updates to the global model solving a constrained linear program, with explicit Pareto constraints to prevent performance degradation on annotated data, and controlled relaxation of the Pareto constraints on pseudo-labeled data to prevent learning of patterns in conflict with the annotated data. We show FedRPO significantly outperforms FedAvg on Amazon internal de-identified dataset on AEC tasks. On supervised learning, FedRPO improved precision by 32.5% over FedAvg when maintaining recall at 90%. Combined with FixMatch [1] for semi-supervised learning, FedRPO outperformed FedAvg on precision by 50.5% at 90% recall.
Meng Feng, Chieh-Chi Kao, Qingming Tang, Amit Solomon, Viktor Rozgic, Chao Wang 0018
ICASSP5
2023 Weight-Sharing Supernet for Searching Specialized Acoustic Event Classification Networks Across Device Constraints
abstract
Acoustic Event Classification (AEC) has been widely used in devices such as smart speakers and mobile phones for home safety or accessibility support [1]. As AEC models run on more and more devices with diverse computation resource constraints, it became increasingly expensive to develop models that are tuned to achieve optimal accuracy/computation trade-off for each given computation resource constraint. In this paper, we introduce a Once-For-All (OFA) Neural Architecture Search (NAS) framework for AEC. Specifically, we first train a weight-sharing supernet that supports different model architectures, followed by automatically searching for a model given specific computational resource constraints. Our experimental results showed that by just training once, the resulting model from NAS significantly outperforms both models trained individually from scratch and knowledge distillation (25.4% and 7.3% relative improvement). We also found that the benefit of weight-sharing supernet training of ultra-small models comes not only from searching but from optimization.
Guan-Ting Lin, Qingming Tang, Chieh-Chi Kao, Viktor Rozgic, Chao Wang 0018
ICASSP4
2023 Towards Paralinguistic-Only Speech Representations for End-to-End Speech Emotion Recognition
Georgios Ioannides, Michael Owen, Andrew Fletcher, Viktor Rozgic, Chao Wang 0018
INTERSPEECH4
2022 Federated Self-Supervised Learning for Acoustic Event Classification
abstract
Standard acoustic event classification (AEC) solutions require large-scale collection of data from client devices for model optimization. Federated learning (FL) is a compelling frame- work that decouples data collection and model training to enhance customer privacy. In this work, we investigate the feasibility of applying FL to improve AEC performance while no customer data can be directly uploaded to the server. We assume no pseudo labels can be inferred from on-device user inputs, aligning with the typical use cases of AEC. We adapt self-supervised learning to the FL framework for on-device continual learning of representations, and it results in improved performance of the downstream AEC classifiers with- out labeled/pseudo-labeled data available. Compared to the baseline w/o FL, the proposed method improves precision up to 20.3% relatively while maintaining the recall. Our work differs from prior work in FL in that our approach does not require user-generated learning targets, and the data we use is collected from our Beta program and is de-identified, to maximally simulate the production settings.
Meng Feng, Chieh-Chi Kao, Qingming Tang, Ming Sun 0007, Viktor Rozgic, Spyridon Matsoukas, Chao Wang 0018
ICASSP5
2022 Sentiment-Aware Automatic Speech Recognition Pre-Training for Enhanced Speech Emotion Recognition
abstract
We propose a novel multi-task pre-training method for Speech Emotion Recognition (SER). We pre-train SER model simultaneously on Automatic Speech Recognition (ASR) and sentiment classification tasks to make the acoustic ASR model more "emotion aware". We generate targets for the sentiment classification using text-to-sentiment model trained on publicly available data. Finally, we fine-tune the acoustic ASR on emotion annotated speech data. We evaluated the proposed approach on MSP-Podcast dataset, where we achieved the best reported concordance correlation coefficient (CCC) of 0.41 for valence prediction.
Ayoub Ghriss, Viktor Rozgic, Elizabeth Shriberg, Chao Wang 0018
ICASSP3
2022 Confidence Estimation for Speech Emotion Recognition Based on the Relationship Between Emotion Categories and Primitives
abstract
Confidence estimation for Speech Emotion Recognition (SER) is instrumental in improving the reliability in the behavior of downstream applications. In this work we propose (1) a novel confidence metric for SER based on the relationship between emotion primitives: arousal, valence, and dominance (AVD) and emotion categories (ECs), (2) EmoConfidNet - a DNN trained alongside the EC recognizer to predict the proposed confidence metric, and (3) a data filtering technique used to enhance the training of EmoConfidNet and the EC recognizer. For each training sample, we calculate distances from corresponding AVD annotation vectors to centroids of each EC in the AVD space, and define EC confidences as functions of the evaluated distances. EmoConfidNet is trained to predict confidence from the same acoustic representations used to train the EC recognizer. EmoConfidNet outperforms state-of-the-art confidence estimation methods on the MSP-Podcast and IEMOCAP datasets. For a fixed EC recognizer, after we reject the same number of low confidence predictions using EmoConfidNet, we achieve a higher F1 and unweighted average recall (UAR) than when rejecting using other methods.
Yang Li 0149, Constantinos Papayiannis, Viktor Rozgic, Elizabeth Shriberg, Chao Wang 0018
ICASSP3
2022 Improved Representation Learning For Acoustic Event Classification Using Tree-Structured Ontology
abstract
Acoustic events have a hierarchical structure analogous to a tree (or a directed acyclic graph). In this work, we propose a structure-aware semi-supervised learning framework for acoustic event classification (AEC). Our hypothesis is that the audio label structure contains useful information that is not available in audios and plain tags. We show that by organizing audio representations with a human-curated tree ontology, we can improve the quality of the learned audio representations for downstream AEC tasks. We use consistency training to use large amounts of unlabeled data for structured representation manifold learning. Experimental results indicate that our framework learns high quality representations which enable us to achieve comparable performance in discriminative tasks as fully supervised baselines. Moreover, our framework can better handle audios with unseen tags by confidently assigning a super-category (internal node like "animal" in Fig. 1) tag to the audio.
Arman Zharmagambetov, Qingming Tang, Chieh-Chi Kao, Ming Sun 0007, Viktor Rozgic, Jasha Droppo, Chao Wang 0018
ICASSP6
2021 Contrastive Unsupervised Learning for Speech Emotion Recognition
abstract
Speech emotion recognition (SER) is a key technology to enable more natural human-machine communication. However, SER has long suffered from a lack of public large-scale labeled datasets. To circumvent this problem, we investigate how unsupervised representation learning on unlabeled datasets can benefit SER. We show that the contrastive predictive coding (CPC) method can learn salient representations from unlabeled datasets, which improves emotion recognition performance. In our experiments, this method achieved state-of-the-art concordance correlation coefficient (CCC) performance for all emotion primitives (activation, valence, and dominance) on IEMOCAP. Additionally, on the MSP-Podcast dataset, our method obtained considerable performance improvements compared to baselines.
Joshua Levy, Andreas Stolcke, Viktor Rozgic, Spyridon Matsoukas, Constantinos Papayiannis, Daniel Bone, Chao Wang 0018
ICASSP5
2019 Multimodal and Multi-view Models for Emotion Recognition
abstract
Studies on emotion recognition (ER) show that combining lexical and acoustic information results in more robust and accurate models.The majority of the studies focus on settings where both modalities are available in training and evaluation.However, in practice, this is not always the case; getting ASR output may represent a bottleneck in a deployment pipeline due to computational complexity or privacyrelated constraints.To address this challenge, we study the problem of efficiently combining acoustic and lexical modalities during training while still providing a deployable acoustic model that does not require lexical inputs.We first experiment with multimodal models and two attention mechanisms to assess the extent of the benefits that lexical information can provide.Then, we frame the task as a multi-view learning problem to induce semantic information from a multimodal model into our acoustic-only network using a contrastive loss function.Our multimodal model outperforms the previous state of the art on the USC-IEMOCAP dataset reported on lexical and acoustic information.Additionally, our multi-view-trained acoustic network significantly surpasses models that have been exclusively trained with acoustic features.
Gustavo Aguilar, Viktor Rozgic, Chao Wang 0018
ACL (1)2
2019 Improving Emotion Classification through Variational Inference of Latent Variables
abstract
Conventional models for emotion recognition from speech signal are trained in supervised fashion using speech utterances with emotion labels. In this study we hypothesize that speech signal depends on multiple latent variables including the emotional state, age, gender, and speech content. We propose an Adversarial Autoencoder (AAE) to perform variational inference over the latent variables and reconstruct the input feature representations. Reconstruction of feature representations is used as an auxiliary task to aid the primary emotion recognition task. Experiments on the IEMOCAP dataset demonstrate that the auxiliary learning tasks improve emotion classification accuracy compared to a baseline supervised classifier. Further, we demonstrate that the proposed learning approach can be used for the end-to-end speech emotion recognition, as its applicable for models that operate on frame-level inputs.
Srinivas Parthasarathy, Viktor Rozgic, Ming Sun 0007, Chao Wang 0018
ICASSP2
2019 Semi-supervised Acoustic Event Detection Based on Tri-training
abstract
This paper presents our work of training acoustic event detection (AED) models using unlabeled dataset. Recent acoustic event detectors are based on large-scale neural networks, which are typically trained with huge amounts of labeled data. Labels for acoustic events are expensive to obtain, and relevant acoustic event audios can be limited, especially for rare events. In this paper we leverage an Internet-scale un-labeled dataset with potential domain shift to improve the detection of acoustic events. Based on the classic tri-training approach, our proposed method shows accuracy improvement over both the supervised training baseline, and semi-supervised self-training set-up, in all pre-defined acoustic event detection tasks. As our approach relies on ensemble models, we further show the improvements can be distilled to a single model via knowledge distillation, with the resulting single student model maintaining high accuracy of teacher ensemble models.
Bowen Shi 0002, Ming Sun 0007, Chieh-Chi Kao, Viktor Rozgic, Spyridon Matsoukas, Chao Wang 0018
ICASSP4
2019 Hierarchical Residual-pyramidal Model for Large Context Based Media Presence Detection
abstract
We study media presence detection, that is, learning to recognize if a sound segment (typically lasting for a few seconds) of a long recorded stream contains media (TV) sound. This problem is difficult because non-media sound sources can be quite diverse (e.g. human voicing, non-vocal sounds and non-human sounds), and the recorded sound can be a mixture of media and non-media sound.Different from speech recognition, where the recognizer needs to detect local phonetic variation, the key features used to distinguish media and non-media sounds are non-local features. Motivated by this, we propose a hierarchical model to learn representation of each pre-chunked segment within a long recorded stream jointly, and encourage every local representation to be not sensitive to variations within each segment. We also further explore the effects of techniques including stream based normalization and iteratively imputing missing labels of training dataset. Experimental results indicate that our proposed contextual based methods are effective for media presence detection.
Qingming Tang, Ming Sun 0007, Chieh-Chi Kao, Viktor Rozgic, Chao Wang 0018
ICASSP4
2019 Compression of Acoustic Event Detection Models with Quantized Distillation
abstract
Acoustic Event Detection (AED), aiming at detecting categories of events based on audio signals, has found application in many intelligent systems. Recently deep neural network significantly advances this field and reduces detection errors to a large scale. However how to efficiently execute deep models in AED has received much less attention. Meanwhile state-of-the-art AED models are based on large deep models, which are computational demanding and challenging to deploy on devices with constrained computational resources. In this paper, we present a simple yet effective compression approach which jointly leverages knowledge distillation and quantization to compress larger network (teacher model) into compact network (student model). Experimental results show proposed technique not only lowers error rate of original compact network by 15% through distillation but also further reduces its model size to a large extent (2% of teacher, 12% of full-precision student) through quantization.
Bowen Shi 0002, Ming Sun 0007, Chieh-Chi Kao, Viktor Rozgic, Spyridon Matsoukas, Chao Wang 0018
INTERSPEECH4
2018 Detecting Media Sound Presence in Acoustic Scenes
Constantinos Papayiannis, Justice Amoh, Viktor Rozgic, Shiva Sundaram, Chao Wang 0018
INTERSPEECH3
2017 Learning Discriminative Features via Label Consistent Neural Network
abstract
Deep Convolutional Neural Networks (CNN) enforce supervised information only at the output layer, and hidden layers are trained by back propagating the prediction error from the output layer without explicit supervision. We propose a supervised feature learning approach, Label Consistent Neural Network, which enforces direct supervision in late hidden layers in a novel way. We associate each neuron in a hidden layer with a particular class label and encourage it to be activated for input signals from the same class. More specifically, we introduce a label consistency regularization called "discriminative representation error" loss for late hidden layers and combine it with classification error loss to build our overall objective function. This label consistency constraint alleviates the common problem of gradient vanishing and tends to faster convergence, it also makes the features derived from late hidden layers discriminative enough for classification even using a simple k-NN classifier. Experimental results demonstrate that our approach achieves state-of-the-art performances on several public datasets for action and object category recognition.
Zhuolin Jiang, Yaming Wang, Larry Davis 0001, Walter Andrews, Viktor Rozgic
WACV5
2014 Multi-modal prediction of PTSD and stress indicators
abstract
Post-traumatic stress disorder (PTSD) is an anxiety disorder that affects a large population and that is currently diagnosed mostly through subject interviews and manual analysis of self-reported symptoms and of subject behavior. However, most PTSD cases are believed to go underdiagnosed and un-dertreated. We present a multi-modal system for computer-aided diagnosis of PTSD and stress that requires no clinician interview and relies principally in the elicitation of multimodal neurophysiological responses to audio-visual stimuli. We conduct a thorough evaluation of the discriminative power of the modalities involved (electro encephalography, galvanic skin-response, electrocardiography, head motion and speech), type of stimuli presented (audio, images, audio-and-images and video), and emotions evoked (positive, negative, and trauma-specific) between PTSD subjects and high and low-stress control groups. Our analysis indicates that the multi-modal prediction from the elicitation of trauma-specific emotions from images and audio is a promising approach to computer-aided diagnosis.
Viktor Rozgic, Amelio Vázquez Reina, Michael Crystal, Amit Srivastava, Veasna Tan, Chris Berka
ICASSP1
2014 Improving speech-based PTSD detection via multi-view learning
abstract
We demonstrate that by applying multi-view learning algorithms one can usefully leverage highly informative, highcost, psychophysiological data collected in a laboratory setting, to improve PTSD screening in the field, where only less-informative, low-cost, speech data are available. Cost metrics reflect resource requirements as well as subject receptivity to data collection. The speech-based representation involves distress indicator extraction from automatic speech recognition output, and a compact holistic audio representation based on the i-vector method. A prototype PTSD screening system was developed that benefits from highly informative EEG data yet, in the field, only relies on subjects' spoken commentary in response to open ended questions. Such a system can deliver screening with significantly increased engagement, to a broader population, leading to earlier intervention and improved outcomes. Using a recent dataset collected for multi-modal computer-aided diagnosis of PTSD, we demonstrate that the proposed method significantly improves speech-based PTSD detection, without requiring costly and aversive procedures at deployment.
Xiaodan Zhuang, Viktor Rozgic, Michael Crystal, Brian Marx
SLT2
2013 Robust EEG emotion classification using segment level decision fusion
abstract
In this paper we address single-trial binary classification of emotion dimensions (arousal, valence, dominance and liking) using electroencephalogram (EEG) signals that represent responses to audio-visual stimuli. We propose an innovative three step solution to this problem: (1) in contrast to the typical feature extraction on the response-level, we represent the EEG signal as a sequence of overlapping segments and extract feature vectors on the segment level; (2) transform segment level features to the response level features using projections based on a novel non-parametric nearest neighbor model; and (3) perform classification on the obtained response-level features. We demonstrate the efficacy of our approach by performing binary classification of emotion dimensions on DEAP (Dataset for Emotion Analysis using electroencephalogram, Physiological and Video Signals) and report state-of-the-art classification accuracies for all emotional dimensions.
Viktor Rozgic, Shiv Vitaladevuni, Rohit Prasad
ICASSP1
2012 Emotion Recognition using Acoustic and Lexical Features
Viktor Rozgic, Sankaranarayanan Ananthakrishnan, Shirin Saleem, Rohit Kumar 0001, Aravind Namandi Vembu, Rohit Prasad
INTERSPEECH1
2011 Estimation of ordinal approach-avoidance labels in dyadic interactions: Ordinal logistic regression approach
abstract
Behavioral Signal Processing aims at automating behavioral coding schemes such as those prevalent in psychology and mental health research. This paper describes a method to automatically quantify the approach-and-avoidance (AA) behavior, described by ordinal labels manually assigned by experts using either video-only or video-with-audio. We propose a novel ordinal regression (OR) algorithm and its hidden Markov model (HMM) extension for estimation of AA labels from visual motion capture based and acoustic features. The proposed algorithm transforms the OR to multiple binary classification problems, solves them by independent score-outputting classifiers and fits the cumulative logit logistic regression model with proportional odds (CLLRMP) to vectors of the classifier scores. The time series extension treats labels as states of the HMM with a likelihood function derived from the probabilistic CLLRMP output. We compare performances of the proposed algorithm applying the weighted binary SVMs in the second step (SVM-OLR), its time-series extension (HMM-SVM-OLR) and the baseline multi-class SVM. On the used dyadic interaction dataset the HMM-SVM-OLR achieves the highest estimation accuracies 71.6 % and 65.7 % for AA labels assigned respectively using video-only and video-with-audio.
Viktor Rozgic, Bo Xiao 0003, Athanasios Katsamanis, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan
ICASSP1
2011 Acoustic and Visual Cues of Turn-Taking Dynamics in Dyadic Interactions
abstract
In this paper we introduce an empirical study of multimodal cues of turn-taking dynamics in a social interaction context. We first identify pauses, gaps and overlapped speech segments in the dyadic conversation dataset. Second, we define two types of measurements, Mean Equalized Energy (MEE) and Animation Level (AL) on the audio and video channels, respectively. Then, we verify the hypothesis that the speaker with higher MEE or AL is more likely to take the floor after silence or overlapped speech. The results suggest that both the vocal and visual movement energy offer useful cues towards inferring the intention of the interlocutor to grab the floor. Index Terms: turn-taking, cues, equalized energy, motion vector
Bo Xiao 0003, Viktor Rozgic, Athanasios Katsamanis, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH2
2010 A new multichannel multi modal dyadic interaction database
abstract
In this work we present a new multi-modal database for analysis of participant behaviors in dyadic interactions. This database contains multiple channels with closeand far-field audio, a high definition camera array and motion capture data. Presence of the motion capture allows precise analysis of the body language low-level descriptors and its comparison with similar descriptors derived from video data. Data is manually labeled by multiple human annotators using psychologyinformed guides. This work also presents an initial analysis of approach-avoidance (A-A) behavior. Two sets of annotations are provided, one based on video only and the other obtained by using both the audio and video channels. Additionally, we describe the statistics of interaction descriptors and A-A labels on participants’ roles. Finally we provide an analysis of relations between various non-verbal features and approach/avoidance labels.
Viktor Rozgic, Bo Xiao 0003, Athanasios Katsamanis, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH1
2009 Optimal Allocation of Time-Resources for Multihypothesis Activity-Level Detection
Gautam Thatte, Viktor Rozgic, Ming Li 0026, Sabyasachi Ghosh, Urbashi Mitra, Shri Narayanan, Murali Annavaram, Donna Spruijt-Metz
DCOSS2
2008 Multimodal Speaker Segmentation in Presence of Overlapped Speech Segments
abstract
We propose a multimodal speaker segmentation algorithm with two main contributions: First, we suggest a hidden Markov model architecture that performs fusion of the three modalities: a multi-camera system for participant localization, a microphone array for speaker localization, and a speaker identification system; Second, we present a novel method for dealing with overlapped speech segments through a likelihood model of the microphone array observations that uses multiple local maxima of the Steered Power Response Generalized Cross Correlation Phase Transform (SPR-GCC-PHAT) function in the Joint Probabilistic Data Association (JPDA) framework. Results show that the proposed method outperforms standard speaker segmentation systems based on: (a) speaker identification and; (b) microphone array processing, for datasets with the significant portion (27.4%) of overlapped speech, and scores as high as 94.4% on the F-measure scale.
Viktor Rozgic, Kyu Jeong Han, Panayiotis G. Georgiou, Shri Narayanan
ISM1
2007 Information Theoretic Analysis of Direct Articulatory Measurements for Phonetic Discrimination
abstract
This paper focuses on the analysis of speech production signals (physical measurements from electromagnetic articulograph) from the perspective of phone discrimination. We explore two different signal representation schemes for the articulatory signals, one based on time-domain analysis and the other based on frequency domain. We quantify the amount of discrimination information offered by the speech production signals in identifying the phone labels through mutual information. Mutual information analyses establish that substantial discrimination information is present in the articulatory stream. Furthermore, phonological classification results with articulatory signals indicate higher accuracy compared to the acoustic signal.
Jorge F. Silva, Vivek Kumar Rangarajan Sridhar, Viktor Rozgic, Shri Narayanan
ICASSP (4)3
2007 Multimodal Meeting Monitoring: Improvements on Speaker Tracking and Segmentation through a Modified Mixture Particle Filter
abstract
In this paper we address improvements to our multimodal system for tracking of meeting participants and speaker segmentation with a focus on the microphone array modality. We propose an algorithm that uses Directions-of-Arrival estimated for each microphone pair as observations and performs tracking of an unknown number of acoustically-active meeting participants and subsequent speaker segmentation. We propose modified mixture particle filter (mMPF) for tracking of acoustic sources in the track-before-detection (TbD) framework. Trajectories of sound sources are reconstructed by the optimal assignment of posterior mixture components produced by mMPF in consecutive frames. Further, we propose a sequential optimal change-point detection algorithm which discovers speech segments in the reconstructed trajectories i.e., performs speaker segmentation. The algorithm is tested on a multi-participant meeting dataset both separately and as a part of the multimodal system. On the task of speaker detection in the multimodal setup we report significant improvement over our previous state of the art implementation.
Viktor Rozgic, Carlos Busso, Panayiotis G. Georgiou, Shri Narayanan
MMSP1