EDBT 2026 Demo / reviewers in the wild / expert
Patrick Cardinal
dblp:88/3832
· DBLP profile ↗
39ranked-venue papers
10as first author
12since 2021 · last 2026
0009-0000-9439-9910ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 26 · 9 first-author · 6 since 2021Security and privacy · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Weakly Supervised Learning for Facial Affective Behavior Analysis: A ReviewabstractRecent advances in deep learning (DL) and computational capacity have enabled facial affective behavior analysis (FABA) to progress from static images captured in controlled settings to fine-grained analysis of facial expressions in real world video data. However, training accurate DL models for FABA typically requires large-scale, expert-annotated datasets, which are costly to obtain and inherently noisy due to the ambiguity of labeling subtle facial expressions and action units (AUs). To mitigate these challenges, weakly supervised learning (WSL) has emerged as a promising paradigm for training models with weak annotations. In this paper, we present a structured taxonomy of WSL scenarios for FABA, organized according to the type of weak annotation and the specific affective task. Building on this taxonomy, we provide a critical synthesis of representative WSL methods for both classification (expression and AU recognition) and regression (expression and AU intensity estimation) tasks, focusing on their core methodological ideas, strengths, and limitations. Furthermore, we systematically summarize the comparative performance of WSL approaches along with widely adopted experimental setups and evaluation proto cols. Our critical assessment identifies key challenges and future research directions, including the need for efficient adaptation of foundation models and for the development of robust, scalable FABA systems suitable for real-world applications. Gnana Praveen Rajasekhar, Patrick Cardinal, Eric Granger |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Multi-Target Backdoor Attacks Against Speaker RecognitionabstractIn this work, we propose a multi-target backdoor attack against speaker identification using position-independent clicking sounds as triggers. Unlike previous single-target approaches, our method targets up to 50 speakers simultaneously, achieving success rates of up to 95.04%. To simulate more realistic attack conditions, we vary the signal-to-noise ratio between speech and trigger, demonstrating a trade-off between stealth and effectiveness. We further extend the attack to the speaker verification task by selecting the most similar training speaker—based on cosine similarity—as a proxy target. The attack is most effective when target and enrolled speaker pairs are highly similar, reaching success rates of up to 90% in such cases. Alexandrine Fortier, Sonal Joshi, Thomas Thebaud, Jesús Villalba 0001, Najim Dehak, Patrick Cardinal |
ASRU | 6 |
| 2023 | Recursive Joint Attention for Audio-Visual Fusion in Regression Based Emotion RecognitionabstractIn video-based emotion recognition (ER), it is important to effectively leverage the complementary relationship among audio (A) and visual (V) modalities, while retaining the intramodal characteristics of individual modalities. In this paper, a recursive joint attention model is proposed along with long short-term memory (LSTM) modules for the fusion of vocal and facial expressions in regression-based ER. Specifically, we investigated the possibility of exploiting the complementary nature of A and V modalities using a joint cross-attention model in a recursive fashion with LSTMs to capture the intramodal temporal dependencies within the same modalities as well as among the A-V feature representations. By integrating LSTMs with recursive joint cross-attention, our model can efficiently leverage both intra- and inter-modal relationships for the fusion of A and V modalities. The results of extensive experiments1performed on the challenging Affwild2 and Fatigue (private) datasets indicate that the proposed A-V fusion model can significantly outperform state-of-art-methods. Gnana Praveen Rajasekhar, Eric Granger, Patrick Cardinal |
ICASSP | 3 |
| 2022 | Towards Robust Speech-to-Text Adversarial AttackabstractThis paper introduces a novel adversarial algorithm for attacking the advanced speech-to-text transcription systems. Our proposed approach is based on developing an extension for the conventional distortion condition of the general adversarial optimization formulation using the Cramér integral probability metric. Minimizing over such a metric contributes to crafting signals very close to the subspace of legitimate speech recordings. That helps yield more robust adversarial signals against over-the-air playbacks without employing neither costly expectation over transformations nor static room impulse response simulations. Our approach considerably outperforms other targeted and non-targeted algorithms in terms of word error rate and sentence-level accuracy. Furthermore compared to seven other strong white and black-box adversarial attacks, our proposed approach is considerably more resilient against multiple consecutive over-the-air playbacks, corroborating its higher robustness in noisy environments. Mohammad Esmaeilpour, Patrick Cardinal, Alessandro L. Koerich |
ICASSP | 2 |
| 2022 | Named Entity Recognition for Audio De-IdentificationabstractData anonymization is often a task carried out by humans. Automating it would reduce the cost and time required to complete this task. This paper presents a pipeline to automate the anonymization of audio data in French. We propose a pipeline, which takes audio files with their transcriptions and removes the named entities (NEs) present in the audio. Our pipeline is made up of a forced aligner, which aligns words in an audio transcript with speech and a model that performs named entity recognition (NER). Then, the audio segments that correspond to NEs are substituted with silence to anonymize audio. We compared forced aligners and NER models to find the best ones for our scenario. We evaluated our pipeline on a small hand-annotated dataset, achieving an F1 score of 0.769. This result shows that automating this task is feasible. Guillaume Baril, Patrick Cardinal, Alessandro L. Koerich |
IJCNN | 2 |
| 2022 | Bi-discriminator GAN for tabular data synthesis
Mohammad Esmaeilpour, Nourhene Chaalia, Adel Abusitta 0001, François-Xavier Devailly, Wissem Maazoun, Patrick Cardinal |
Pattern Recognit. Lett. | 6 |
| 2022 | RSD-GAN: Regularized Sobolev Defense GAN Against Speech-to-Text Adversarial AttacksabstractThis letter introduces a new synthesis-based defense algorithm for counteracting with a varieties of adversarial attacks developed for challenging the performance of the cutting-edge speech-to-text transcription systems. Our algorithm implements a Sobolev-based GAN and proposes a novel regularizer for effectively controlling over the functionality of the entire generative model, particularly the discriminator network during training. Our achieved results upon carrying out numerous experiments on the victim DeepSpeech, Kaldi, and Lingvo speech transcription systems corroborate the remarkable performance of our defense approach against a comprehensive range of targeted and non-targeted adversarial attacks. Mohammad Esmaeilpour, Nourhene Chaalia, Patrick Cardinal |
IEEE Signal Process. Lett. | 3 |
| 2022 | Multidiscriminator Sobolev Defense-GAN Against Adversarial Attacks for End-to-End Speech SystemsabstractThis paper introduces a defense approach against end-to-end adversarial attacks developed for cutting-edge speech-to-text systems. The proposed defense algorithm has four steps. First, we use the short-time Fourier transform to represent speech signals with 2D spectrograms. Second, we iteratively find a safe vector using a spectrogram subspace projection operation. This operation minimizes the chordal distance adjustment between spectrograms with an additional regularization term. Third, we synthesize a spectrogram with such a safe vector using a novel GAN architecture trained with Sobolev integral probability metric. We impose an additional constraint on the generator network to improve the model’s performance in terms of stability and the total number of learned modes. Finally, we reconstruct the signal from the synthesized spectrogram and the Griffin-Lim phase approximation technique. We evaluate the proposed defense approach against six strong white and black-box adversarial attacks on DeepSpeech, Kaldi, and Lingvo models. The experimental results show that our algorithm outperforms other state-of-the-art defense algorithms in terms of accuracy and signal quality. Mohammad Esmaeilpour, Patrick Cardinal, Alessandro L. Koerich |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Cross Attentional Audio-Visual Fusion for Dimensional Emotion RecognitionabstractMultimodal analysis has recently drawn much interest in affective computing, since it can improve the overall accuracy of emotion recognition over isolated uni-modal approaches. The most effective techniques for multimodal emotion recognition efficiently leverage diverse and complimentary sources of information, such as facial, vocal, and physiological modalities, to provide comprehensive feature representations. In this paper, we focus on dimensional emotion recognition based on the fusion of facial and vocal modalities extracted from videos, where complex spatiotemporal relationships may be captured. Most of the existing fusion techniques rely on recurrent networks or conventional attention mechanisms that do not effectively leverage the complimentary nature of audiovisual (A-V) modalities. We introduce a cross-attentional fusion approach to extract the salient features across A - V modalities, allowing for accurate prediction of continuous values of valence and arousal. Our new cross-attentional A - V fusion model efficiently leverages the inter-modal relationships. In particular, it computes cross-attention weights to focus on the more contributive features across individual modalities, and thereby combine contributive feature representations, which are then fed to fully connected layers for the prediction of valence and arousal. The effectiveness of the proposed approach is validated experimentally on videos from the RECOLA and Fatigue (private) data-sets. Results indicate that our cross-attentional A - V fusion model is a cost-effective approach that outperforms state-of-the-art fusion approaches. Code is available: https://github.com/praveena2j/Cross-Attentional-AV-Fusion. Gnana Praveen Rajasekhar, Eric Granger, Patrick Cardinal |
FG | 3 |
| 2021 | Class-Conditional Defense GAN Against End-To-End Speech AttacksabstractIn this paper we propose a novel defense approach against end-to-end adversarial attacks developed to fool advanced speech-to-text systems such as DeepSpeech and Lingvo. Unlike conventional defense approaches, the proposed approach does not directly employ low-level transformations such as autoencoding a given input signal aiming at removing potential adversarial perturbation. Instead of that, we find an optimal input vector for a class conditional generative adversarial network through minimizing the relative chordal distance adjustment between a given test input and the generator network. Then, we reconstruct the 1D signal from the synthesized spectrogram and the original phase information derived from the given input signal. Hence, this reconstruction does not add any extra noise to the signal and according to our experimental results, our defense-GAN considerably outperforms conventional defense algorithms both in terms of word error rate and sentence level recognition accuracy. Mohammad Esmaeilpour, Patrick Cardinal, Alessandro L. Koerich |
ICASSP | 2 |
| 2021 | Deep domain adaptation with ordinal regression for pain assessment using weakly-labeled videos
Gnana Praveen Rajasekhar, Eric Granger, Patrick Cardinal |
Image Vis. Comput. | 3 |
| 2021 | Cyclic Defense GAN Against Speech Adversarial AttacksabstractThis paper proposes a new defense approach for counteracting state-of-the-art white and black-box adversarial attack algorithms. Our approach fits into the implicit reactive defense algorithm category since it does not directly manipulate the potentially malicious input signals. Instead, it reconstructs a similar signal with a synthesized spectrogram using a cyclic generative adversarial network. This cyclic framework helps to yield a stable generative model. Finally, we feed the reconstructed signal into the speech-to-text model for transcription. The conducted experiments on targeted and non-targeted adversarial attacks developed for attacking DeepSpeech, Kaldi, and Lingvo models demonstrate the proposed defense's effectiveness in adverse scenarios. Mohammad Esmaeilpour, Patrick Cardinal, Alessandro L. Koerich |
IEEE Signal Process. Lett. | 2 |
| 2020 | Deep Weakly Supervised Domain Adaptation for Pain Localization in VideosabstractAutomatic pain assessment has an important potential diagnostic value for populations that are incapable of articulating their pain experiences. As one of the dominating nonverbal channels for eliciting pain expression events, facial expressions has been widely investigated for estimating the pain intensity of individual. However, using state-of-the-art deep learning (DL) models in real-world pain estimation applications poses several challenges related to the subjective variations of facial expressions, operational capture conditions, and lack of representative training videos with labels. Given the cost of annotating intensity levels for every video frame, we propose a weakly-supervised domain adaptation (WSDA) technique that allows for training 3D CNNs for spatiotemporal pain intensity estimation using weakly labeled videos, where labels are provided on a periodic basis. In particular, WSDA integrates multiple instance learning into an adversarial deep domain adaptation framework to train an Inflated 3D-CNN (I3D) model such that it can accurately estimate pain intensities in the target operational domain. The training process relies on weak target loss, along with domain loss and source loss for domain adaptation of the I3D model. Experimental results obtained using labeled source domain RECOLA videos and weakly-labeled target domain UNBC-McMaster videos indicate that the proposed deep WSDA approach can achieve significantly higher level of sequence (bag)-level and frame (instance)-level pain localization accuracy than related state-of-the-art approaches. Gnana Praveen Rajasekhar, Eric Granger, Patrick Cardinal |
FG | 3 |
| 2020 | Detection of Adversarial Attacks and Characterization of Adversarial SubspaceabstractAdversarial attacks have always been a serious threat for any data-driven model. In this paper, we explore subspaces of adversarial examples in unitary vector domain, and we propose a novel detector for defending our models trained for environmental sound classification. We measure chordal distance between legitimate and malicious representation of sounds in unitary space of generalized Schur decomposition and show that their manifolds lie far from each other. Our front-end detector is a regularized logistic regression which discriminates eigenvalues of legitimate and adversarial spectrograms. The experimental results on three benchmarking datasets of environmental sounds represented by spectrograms reveal high detection rate of the proposed detector for eight types of adversarial attacks and it also outperforms other detection approaches. Mohammad Esmaeilpour, Patrick Cardinal, Alessandro L. Koerich |
ICASSP | 2 |
| 2020 | Adversarially Training for Audio ClassifiersabstractIn this paper, we investigate the potential effect of the adversarially training on the robustness of six advanced deep neural networks against a variety of targeted and non-targeted adversarial attacks. We firstly show that, the ResNet-56 model trained on the 2D representation of the discrete wavelet transform appended with the tonnetz chromagram outperforms other models in terms of recognition accuracy. Then we demonstrate the positive impact of adversarially training on this model as well as other deep architectures against six types of attack algorithms (white and black-box) with the cost of the reduced recognition accuracy and limited adversarial perturbation. We run our experiments on two benchmarking environmental sound datasets and show that without any imposed limitations on the budget allocations for the adversary, the fooling rate of the adversarially trained models can exceed 90%. In other words, adversarial attacks exist in any scales, but they might require higher adversarial perturbations compared to non-adversarially trained models. Raymel Alfonso Sallo, Mohammad Esmaeilpour, Patrick Cardinal |
ICPR | 3 |
| 2020 | A Robust Approach for Securing Audio Classification Against Adversarial AttacksabstractAdversarial audio attacks can be considered as a small perturbation unperceptive to human ears that is intentionally added to an audio signal and causes a machine learning model to make mistakes. This poses a security concern about the safety of machine learning models since the adversarial attacks can fool such models toward the wrong predictions. In this paper we first review some strong adversarial attacks that may affect both audio signals and their 2D representations and evaluate the resiliency of deep learning models and support vector machines (SVM) trained on 2D audio representations such as short time Fourier transform, discrete wavelet transform (DWT) and cross recurrent plot against several state-of-the-art adversarial attacks. Next, we propose a novel approach based on pre-processed DWT representation of audio signals and SVM to secure audio systems against adversarial attacks. The proposed architecture has several preprocessing modules for generating and enhancing spectrograms including dimension reduction and smoothing. We extract features from small patches of the spectrograms using the speeded up robust feature (SURF) algorithm which are further used to transform into cluster distance distribution using the K-Means++ algorithm. Finally, SURF-generated vectors are encoded by this codebook and the resulting codewords are used for training a SVM. All these steps yield to a novel approach for audio classification that provides a good tradeoff between accuracy and resilience. Experimental results on three environmental sound datasets show the competitive performance of the proposed approach compared to the deep neural networks both in terms of accuracy and robustness against strong adversarial attacks. Mohammad Esmaeilpour, Patrick Cardinal, Alessandro L. Koerich |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2019 | Emotion Recognition Using Fusion of Audio and Video FeaturesabstractIn this paper we propose a fusion approach to continuous emotion recognition that combines visual and auditory modalities in their representation spaces to predict the arousal and valence levels. The proposed approach employs a pre-trained convolution neural network and transfer learning to extract features from video frames that capture the emotional content. For the auditory content, a minimalistic set of parameters such as prosodic, excitation, vocal tract, and spectral descriptors are used as features. The fusion of these two modalities is carried out at a feature level, before training a single support vector regressor (SVR) or at a prediction level, after training one SVR for each modality. The proposed approach also includes preprocessing and post-processing techniques which contribute favorably to improving the concordance correlation coefficient (CCC). Experimental results for predicting spontaneous and natural emotions on the RECOLA dataset have shown that the proposed approach takes advantage of the complementary information of visual and auditory modalities and provides CCC of 0.749 and 0.565 for arousal and valence, respectively. The proposed approach outperforms the baseline system and several traditional approaches based on auditory and visual handcrafted features. Juan D. S. Ortega, Patrick Cardinal, Alessandro L. Koerich |
SMC | 2 |
| 2019 | End-to-end environmental sound classification using a 1D convolutional neural network
Sajjad Abdoli, Patrick Cardinal, Alessandro L. Koerich |
Expert Syst. Appl. | 2 |
| 2018 | Classification of Nonverbal Human Produced Audio Events: A Pilot StudyabstractThe accurate classification of nonverbal human produced audio events opens the door to numerous applications beyond health monitoring. Voluntary events, such as tongue clicking and teeth chattering, may lead to a novel way of silent interface command. Involuntary events, such as coughing and clearing the throat, may advance the current state-of-the-art in hearing health research. The challenge of such applications is the balance between the processing capabilities of a small intra-aural device and the accuracy of classification. In this pilot study, 10 nonverbal audio events are captured inside the ear canal blocked by an intra-aural device. The performance of three classifiers is investigated: Gaussian Mixture Model (GMM), Support Vector Machine and Multi-Layer Perceptron. Each classifier is trained using three different feature vector structures constructed using the mel-frequency cepstral (MFCC) coefficients and their derivatives. Fusion of the MFCCs with the auditory-inspired amplitude modulation features (AAMF) is also investigated. Classification is compared between binaural and monaural training sets as well as for noisy and clean conditions. The highest accuracy is achieved at 75.45% using the GMM classifier with the binaural MFCC+AAMF clean training set. Accuracy of 73.47% is achieved by training and testing the classifier with the binaural clean and noisy dataset. Rachel E. Bouserhal, Philippe Chabot, Milton Orlando Sarria-Paja, Patrick Cardinal, Jérémie Voix |
INTERSPEECH | 4 |
| 2016 | Automatic Dialect Detection in Arabic Broadcast SpeechabstractWe investigate different approaches for dialect identification in Arabic broadcast speech, using phonetic, lexical features obtained from a speech recognition system, and acoustic features using the i-vector framework. We studied both generative and discriminate classifiers, and we combined these features using a multi-class Support Vector Machine (SVM). We validated our results on an Arabic/English language identification task, with an accuracy of 100%. We used these features in a binary classifier to discriminate between Modern Standard Arabic (MSA) and Dialectal Arabic, with an accuracy of 100%. We further report results using the proposed method to discriminate between the five most widely used dialects of Arabic: namely Egyptian, Gulf, Levantine, North African, and MSA, with an accuracy of 52%. We discuss dialect identification errors in the context of dialect code-switching between Dialectal Arabic and MSA, and compare the error pattern between manually labeled data, and the output from our classifier. We also release the train and test data as standard corpus for dialect identification. Ahmed Ali 0002, Najim Dehak, Patrick Cardinal, Sameer Khurana, Sree Harsha Yella, James R. Glass, Peter Bell 0001, Steve Renals |
INTERSPEECH | 3 |
| 2016 | Native Language Detection Using the I-Vector Framework
Mohammed Senoussaoui, Patrick Cardinal, Najim Dehak, Alessandro L. Koerich |
INTERSPEECH | 2 |
| 2015 | Audio quotation marks for natural language understanding
Simon Boutin, Réal Tremblay, Patrick Cardinal, Doug Peters, Pierre Dumouchel |
INTERSPEECH | 3 |
| 2015 | Speaker adaptation using the i-vector technique for bottleneck featuresabstractDeep Neural Networks (DNN) have been largely used and successfully applied in the context of speaker independent Automatic Speech Recognition (ASR). However, these models are not easily adapted to model a specific speaker characteristic. Recently, one approach was proposed to address this issue, which consists of using the I-vector representation as input to the DNN. The I-vector is playing the role of providing information about the speaker as well as the environmental conditions for a given recording. This approach achieved a significant improvement in the context of a hybrid system of DNN combined with Hidden Markov Model (HMM). In this paper, we study the effect of speaker adaptation based on the I-vector framework in the context of stacked bottleneck features. These features, extracted from a second level of DNNs, are modelled by a classical Gaussian Mixture Model (GMM) ASR system. The proposed approach achieved an absolute WER improvement of 1.2% on an Arabic Broadcast news task. Index Terms: DNN, I-Vector, Bottleneck Features, Speech Recognition Patrick Cardinal, Najim Dehak, Yu Zhang 0033, James R. Glass |
INTERSPEECH | 1 |
| 2014 | Recent advances in ASR applied to an Arabic transcription system for Al-JazeeraabstractThis paper describes a detailed comparison of several state-of-the-art speech recognition techniques applied to a limited Ara-bic broadcast news dataset. The different approaches were all trained on 50 hours of transcribed audio from the Al-Jazeera news channel. The best results were obtained using i-vector-based speaker adaptation in a training scenario using the Min-imum Phone Error (MPE) criteria combined with sequential Deep Neural Network (DNN) training. We report results for two different types of test data: broadcast news reports, with a best word error rate (WER) of 17.86%, and a broadcast conver-sations with a best WER of 29.85%. The overall WER on this test set is 25.6%. Index Terms: Arabic, ASR system, Kaldi 1. Patrick Cardinal, Ahmed Ali 0002, Najim Dehak, Yu Zhang 0033, Tuka Al Hanai, James R. Glass, Stephan Vogel |
INTERSPEECH | 1 |
| 2014 | A complete KALDI recipe for building Arabic speech recognition systemsabstractIn this paper we present a recipe and language resources for training and testing Arabic speech recognition systems using the KALDI toolkit. We built a prototype broadcast news system using 200 hours GALE data that is publicly available through LDC. We describe in detail the decisions made in building the system: using the MADA toolkit for text normalization and vowelization; why we use 36 phonemes; how we generate pronunciations; how we build the language model. We report results using state-of-the-art modeling and decoding techniques. The scripts are released through KALDI and resources are made available on QCRI's language resources web portal. This is the first effort to share reproducible sizable training and testing results on MSA system. Ahmed Ali 0002, Patrick Cardinal, Najim Dehak, Stephan Vogel, James R. Glass |
SLT | 3 |
| 2013 | Large Vocabulary Speech Recognition on Parallel ArchitecturesabstractThe speed of modern processors has remained constant over the last few years but the integration capacity continues to follow Moore's law and thus, to be scalable, applications must be parallelized. The parallelization of the classical Viterbi beam search has been shown to be very difficult on multi-core processor architectures or massively threaded architectures such as Graphics Processing Unit (GPU). The problem with this approach is that active states are scattered in memory and thus, they cannot be efficiently transferred to the processor memory. This problem can be circumvented by using the A* search which uses a heuristic to significantly reduce the number of explored hypotheses. The main advantage of this algorithm is that the processing time is moved from the search in the recognition network to the computation of heuristic costs, which can be designed to take advantage of parallel architectures. Our parallel implementation of the A* decoder on a 4-core processor with a GPU led to a speed-up factor of 6.13 compared to the Viterbi beam search at its maximum capacity and an improvement of 4% absolute in accuracy at real-time. Patrick Cardinal, Pierre Dumouchel, Gilles Boulianne |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Using A* for the parallelization of speech recognition systemsabstractThe speed of modern processors has remained constant over the last few years but the integration capacity continues to follow Moore's law and thus, to be scalable, applications must be parallelized. This paper presents results in using the A* search algorithm in a large vocabulary speech recognition parallel system. This algorithm allows better parallelization over the Viterbi algorithm. First experiments with a “unigram approximation” heuristic resulted in approximatively 8.7 times less states being explored compared to our classical Viterbi decoder. The multi-thread implementation of the A* decoder led to a speed-up factor of 3 over its sequential counterpart. Patrick Cardinal, Gilles Boulianne, Pierre Dumouchel |
ICASSP | 1 |
| 2012 | CRIM's content-based audio copy detection system for TRECVID 2009
Vishwa Gupta, Gilles Boulianne, Patrick Cardinal |
Multim. Tools Appl. | 3 |
| 2010 | Content-based audio copy detection using nearest-neighbor mappingabstractWe report results on audio copy detection for TRECVID 2008 copy detection task. This task involves searching for transformed audio queries in over 200 hours of test audio. The queries were transformed in seven different ways, three of them involved mixing unrelated speech to the original query, making it a much more difficult task. We give results with a few promising algorithms and show that mapping each test frame to the nearest query frame results in robust audio copy detection. The minimal normalized detection cost rate (NDCR) for even the worst case transformations is less than 0.03. The algorithm provides easy parallel processing on a graphics processing unit, leading to a very fast search. Vishwa Gupta, Gilles Boulianne, Patrick Cardinal |
ICASSP | 3 |
| 2010 | Content-based advertisement detection
Patrick Cardinal, Vishwa Gupta, Gilles Boulianne |
INTERSPEECH | 1 |
| 2009 | Real-time correction of closed-captionsabstractAll in-text\treferences\tunderlined\tin\tblue\tare\tlinked\tto\tpublications\ton\tResearchGate, letting you\taccess\tand\tread\tthem\timmediately. Patrick Cardinal, Gilles Boulianne |
INTERSPEECH | 1 |
| 2009 | Using parallel architectures in speech recognitionabstractThe speed of modern processors has remained constant over the last few years and thus, to be scalable, applications must be parallelized.In addition to the main CPU, almost every computer is equipped with a Graphics Processors Unit (GPU) which is in essence a specialized parallel processor.This paper explores how performances of speech recognition systems can be enhanced by using GPU for the acoustic computations and multicore CPUs for the Viterbi search in a large vocabulary application.The multi-core implementation of our speech recognition system runs 1.3 times faster than the single-threaded CPU implementation.Addition of the GPU for dedicated acoustic computations increases the speed by a factor of 2.8, leading to a word accuracy improvement of 16.6% absolute at real-time, compared to the the single-threaded CPU implementation. Patrick Cardinal, Pierre Dumouchel, Gilles Boulianne |
INTERSPEECH | 1 |
| 2008 | GPU accelerated acoustic likelihood computationsabstractThis paper introduces the use of Graphics Processors Unit (GPU) for computing acoustic likelihoods in a speech recog-nition system. In addition to their high availability, GPUs pro-vide high computing performance at low cost. We have used a NVidia GeForce 8800GTX programmed with the CUDA (Com-pute Unified Device Architecture) which shows the GPU as a parallel coprocessor. The acoustic likelihoods are computed as dot products, operations for which GPUs are highly efficient. The implementation in our speech recognition system shows that GPU is 5x faster than the CPU SSE-based implementation. This improvement led to a speed up of 35 % on a large vocabu-lary task. Index Terms: Speech recognition, GPU 1. Patrick Cardinal, Pierre Dumouchel, Gilles Boulianne, Michel Comeau |
INTERSPEECH | 1 |
| 2007 | Real-Time Correction of Closed-Captions
Patrick Cardinal, Gilles Boulianne, Michel Comeau, Maryse Boisvert |
ACL | 1 |
| 2006 | Computer-assisted closed-captioning of live TV broadcasts in FrenchabstractGrowing needs for French closed-captioning of live TV broadcasts in Canada cannot be met only with stenography-based technology because of a chronic shortage of skilled stenographers. Using speech recognition for live closed-captioning, however, requires several specific problems to be solved, such as the need for low-latency real-time recognition, remote operation, automated model updates, and collaborative work. In this paper we describe our solutions to these problems and the implementation of a live captioning system based on the CRIM speech recognizer. We report results from field deployment in several projects. The oldest in operation has been broadcasting real-time closed-captions for more than 2 years. Index Terms: speech recognition, closed-captioning, model adaptation. Gilles Boulianne, Jean-Francois Beaumont, Maryse Boisvert, Julie Brousseau, Patrick Cardinal, Claude Chapdelaine, Michel Comeau, Pierre Ouellet, Frédéric Osterrath |
INTERSPEECH | 5 |
| 2005 | Segmentation of recordings based on partial transcriptionsabstractIn this paper, we present the approach we used to produce a training database from a set of recorded newscasts for which we had inaccurate transcriptions. These transcribed segments correspond to a set of prepared anchor texts and journalist stories, not necessarily in chronological order of their actual presentation. No segmental time boundary information is provided. Our main concern is thus to establish time marks that delimit the audio segments of the corresponding texts. To resolve this problem, we have developped a time marking procedure using our speech recognition engine. We obtain a segmentation accuracy of 80%. 1. Patrick Cardinal, Gilles Boulianne, Michel Comeau |
INTERSPEECH | 1 |
| 2003 | Automatic segmentation of film dialogues into phonemes and graphemesabstractIn film post-production, efficient methods for re-recording a dialogue or dubbing in a new language require a precisely time-aligned text, with individual letters time-coded to video frame resolution. Currently, this time alignment is performed by experts in a painstaking and slow process. To automate this process, we used CRIM’s largevocabulary HMM speech recognizer as a phoneme segmenter and measured its accuracy on typical film extracts in French and English. Our results reveal several characteristics of film dialogues, in addition to noise, that affect segmentation accuracy, such as speaking style or reverberant recordings. Despite these difficulties, an HMM-based segmenter trained on clean speech can still provide more than 89 % acceptable phoneme boundaries on typical film extracts. We also propose a method which provides the correspondence between aligned phonemes and graphemes of the text. The method does not use explicit rules, but rather computes an optimal string alignment according to an edit-distance metric. Together, HMM phoneme segmentation and phonemegrapheme correspondence meet the needs of film postproduction for a time-aligned text, and make it possible to automate a large part of the current post-synch process. 1. Gilles Boulianne, Jean-Francois Beaumont, Patrick Cardinal, Michel Comeau, Pierre Ouellet, Pierre Dumouchel |
INTERSPEECH | 3 |
| 2003 | Automated closed-captioning of live TV broadcast news in FrenchabstractThis paper describes the system currently under development at CRIM whose aim is to provide real-time closed captioning of live TV broadcast news in Canadian French. This project is done in collaboration with TVA Network, a national TV broadcaster and the RQST (a Québec association which promotes the use of subtitling). The automated closed-captioning system will use CRIM’s transducer-based large vocabulary French recognizer. The system will be totally integrated to the existing broadcaster’s equipment and working methods. First ”on-air” use will take place in February 2004. 1. Julie Brousseau, Jean-Francois Beaumont, Gilles Boulianne, Patrick Cardinal, Claude Chapdelaine, Michel Comeau, Frédéric Osterrath, Pierre Ouellet |
INTERSPEECH | 4 |
| 2002 | Disambiguation of Finite-State Transducers
N. Smaili, Patrick Cardinal, Gilles Boulianne, Pierre Dumouchel |
COLING | 2 |