EDBT 2026 Demo / reviewers in the wild / expert
Frank Rudzicz
dblp:36/6505
· DBLP profile ↗
81ranked-venue papers
11as first author
24since 2021 · last 2026
0000-0002-1139-3423ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 4 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorComputer networks · 2Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SphereEdit: Spherical Semantic Editing in Diffusion ModelsabstractDespite significant advances in diffusion models, achieving precise and composable image editing without task-specific training remains a challenge. Existing approaches often rely on iterative optimization or linear latent operations, which are slow, brittle, and prone to attribute entanglement (e.g., editing "lipstick" inadvertently alters skin tone). We introduce SphereEdit, a training-free framework that leverages the spherical geometry of diffusion embeddings and token aware cross-attention to enable interpretable, fine-grained control. We represent semantic attributes as unit vector directions in the denoiser’s prediction space and show that antipodal symmetry (‘old’ is approximately the negation of ‘young’) naturally supports bidirectional edits, while approximate orthogonality enables clean composition through spherical coefficient. At inference, these directions modulate cross-attention activations, producing spatially localized edits without optimization or fine-tuning. SphereEdit achieves sharper, more disentangled edits than prior baselines, while remaining plug-and-play and applicable across diverse image editing tasks. The code is available at https://github.com/sala-kon/SphereEdit Salamata Konate, Hassan Hamidi, Elham Dolatabadi, Frank Rudzicz, Laleh Seyyed-Kalantari |
WACV | 4 |
| 2026 | On the Limitations of Speaker DiarizationabstractABSTRACT Although speaker diarization has evolved to be more robust and more refined, including incorporating modern automatic speech recognition (ASR), current systems still suffer from several disruptive factors, like noise. We comprehensively evaluate the limitations of current diarization systems to uncover the underlying causes that hinder accuracy. Five open‐source diarization pipelines—both diarization‐only and joint ASR and diarization systems—are assessed on a set of heterogeneous benchmark data sets. We compare the performance of joint pipelines against those of diarization‐only systems, and analyse which audio characteristics hinder speaker discrimination, as well as the impact of using the speaker count as an input parameter. Our results indicate that diarization‐only and joint approaches are competitive with each other in unsupervised scenarios, and that providing the speaker count does not improve performance consistently. We also identify short audio duration and low speech‐to‐noise ratio (SNR) as the most impairing properties. We recommend using speech representation learning to further uncover underlying factors that affect diarization, pre‐processing techniques to remove noise, and performing hyper‐parameter tuning on, for example, the speech window length, and speech detection thresholds. Joana Amorim, João Pimentel 0004, Frank Rudzicz |
Expert Syst. J. Knowl. Eng. | 3 |
| 2025 | Trustworthy Medical Question Answering: An Evaluation-Centric SurveyabstractYinuo Wang, Baiyang Wang, Robert Mercer, Frank Rudzicz, Sudipta Singha Roy, Pengjie Ren, Zhumin Chen, Xindi Wang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Baiyang Wang, Robert E. Mercer, Frank Rudzicz, Sudipta Singha Roy, Pengjie Ren, Zhumin Chen, Xindi Wang 0001 |
EMNLP | 4 |
| 2025 | Filtered not Mixed: Filtering-Based Online Gating for Mixture of Large Language ModelsabstractWe propose MoE-F — a formalized mechanism for combining N pre-trained expert Large Language Models (LLMs) in online time-series prediction tasks by adaptively forecasting the best weighting of LLM predictions at every time step. Our mechanism leverages the conditional information in each expert's running performance to forecast the best combination of LLMs for predicting the time series in its next step. Diverging from static (learned) Mixture of Experts (MoE) methods, our approach employs time-adaptive stochastic filtering techniques to combine experts. By framing the expert selection problem as a finite state-space, continuous-time Hidden Markov model (HMM), we can leverage the Wohman-Shiryaev filter. Our approach first constructs N parallel filters corresponding to each of the N individual LLMs. Each filter proposes its best combination of LLMs, given the information that they have access to. Subsequently, the N filter outputs are optimally aggregated to maximize their robust predictive power, and this update is computed efficiently via a closed-form expression, thus generating our ensemble predictor.
Our contributions are:
- **(I)** the MoE-F algorithm — deployable as a plug-and-play filtering harness,
- **(II)** theoretical optimality guarantees of the proposed filtering-based gating algorithm (via optimality guarantees for its parallel Bayesian filtering and its robust aggregation steps), and
- **(III)** empirical evaluation and ablative results using state-of-the-art foundational and MoE LLMs on a real-world _Financial Market Movement_ task where MoE-F attains a remarkable 17% absolute and 48.5% relative F1 measure improvement over the next best performing individual LLM expert predicting short-horizon market movement based on streaming news. Further, we provide empirical evidence of substantial performance gains in applying MoE-F over specialized models in the _long-horizon time-series forecasting_ domain. Code available on github: https://github.com/raeidsaqur/moe-f Raeid Saqur, Anastasis Kratsios, Florian Krach, Yannick Limmer, Blanka Horvath, Frank Rudzicz |
ICLR | 6 |
| 2025 | How to Recover Long Audio Sequences Through Gradient Inversion Attack With Dynamic Segment-based Reconstruction
Xijie Zeng, Frank Rudzicz |
INTERSPEECH | 2 |
| 2025 | ACCORD: Closing the Commonsense Measurability GapabstractFrançois Roewer-Després, Jinyue Feng, Zining Zhu, Frank Rudzicz. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. François Roewer-Després, Jinyue Feng, Zining Zhu 0001, Frank Rudzicz |
NAACL (Long Papers) | 4 |
| 2024 | Auxiliary Knowledge-Induced Learning for Automatic Multi-Label Medical Document ClassificationabstractThe International Classification of Diseases (ICD) is an authoritative medical classification system of different diseases and conditions for clinical and management purposes. ICD indexing aims to assign a subset of ICD codes to a medical record. Since human coding is labour-intensive and error-prone, many studies employ machine learning techniques to automate the coding process. ICD coding is a challenging task, as it needs to assign multiple codes to each medical document from an extremely large hierarchically organized collection. In this paper, we propose a novel approach for ICD indexing that adopts three ideas: (1) we use a multi-level deep dilated residual convolution encoder to aggregate the information from the clinical notes and learn document representations across different lengths of the texts; (2) we formalize the task of ICD classification with auxiliary knowledge of the medical records, which incorporates not only the clinical texts but also different clinical code terminologies and drug prescriptions for better inferring the ICD codes; and (3) we introduce a graph convolutional network to leverage the co-occurrence patterns among ICD codes, aiming to enhance the quality of label representations. Experimental results show the proposed method achieves state-of-the-art performance on a number of measures. Xindi Wang 0001, Robert E. Mercer, Frank Rudzicz |
LREC/COLING | 3 |
| 2024 | Whister: Using Whisper's representations for Stuttering detection
Vrushank Changawala, Frank Rudzicz |
INTERSPEECH | 2 |
| 2024 | Self-Supervised Embeddings for Detecting Individual Symptoms of Depression
Sri Harsha Dumpala, Katerina Dikaios, Abraham Nunes, Frank Rudzicz, Rudolf Uher, Sageev Oore |
INTERSPEECH | 4 |
| 2024 | Developing Multi-Disorder Voice Protocols: A team science approach involving clinical expertise, bioethics, standards, and DEI
Anaïs Rameau, Satrajit Ghosh, Alexandros Sigaras, Olivier Elemento, Jean-Christophe Bélisle-Pipon, Vardit Ravitsky, Maria Powell, Alistair Johnson, David A. Dorr, Philip R. O. Payne, Micah Boyer, Stephanie Watts, Ruth Bahr, Frank Rudzicz, Jordan Lerner-Ellis, Shaheen Awan, Don Bolser, Yael Bensoussan |
INTERSPEECH | 14 |
| 2024 | Long-form evaluation of model editingabstractDomenic Rosati, Robie Gonzales, Jinkun Chen, Xuemin Yu, Yahya Kayani, Frank Rudzicz, Hassan Sajjad. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Domenic Rosati, Robie Gonzales, Jinkun Chen, Xuemin Yu, Melis Erkan, Yahya Kayani, Satya Deepika Chavatapalli, Frank Rudzicz, Hassan Sajjad 0001 |
NAACL-HLT | 8 |
| 2024 | Multi-stage Retrieve and Re-rank Model for Automatic Medical Coding RecommendationabstractXindi Wang, Robert Mercer, Frank Rudzicz. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xindi Wang 0001, Robert E. Mercer, Frank Rudzicz |
NAACL-HLT | 3 |
| 2024 | Representation Noising: A Defence Mechanism Against Harmful FinetuningabstractReleasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release of weights, weight stealing and fine-tuning APIs make closed models vulnerable to harmful fine-tuning attacks (HFAs). While safety measures like preventing jailbreaks and improving safety guardrails are important, such measures can easily be reversed through fine-tuning. In this work, we propose Representation Noising (\textsf{\small RepNoise}), a defence mechanism that operates even when attackers have access to the weights. \textsf{\small RepNoise} works by removing information about harmful representations such that it is difficult to recover them during fine-tuning. Importantly, our defence is also able to generalize across different subsets of harm that have not been seen during the defence process as long as they are drawn from the same distribution of the attack set. Our method does not degrade the general capability of LLMs and retains the ability to train the model on harmless tasks. We provide empirical evidence that the efficacy of our defence lies in its ``depth'': the degree to which information about harmful representations is removed across {\em all layers} of the LLM. We also find areas where \textsf{\small RepNoise} still remains ineffective and highlight how those limitations can inform future research. Domenic Rosati, Jan Wehner, Kai Williams, Lukasz Bartoszcze, Robie Gonzales, Carsten Maple, Subhabrata Majumdar, Hassan Sajjad 0001, Frank Rudzicz |
NeurIPS | 9 |
| 2023 | Investigating the Learning Behaviour of In-Context Learning: A Comparison with Supervised LearningabstractLarge language models (LLMs) have shown remarkable capacity for in-context learning (ICL), where learning a new task from just a few training examples is done without being explicitly pre-trained. However, despite the success of LLMs, there has been little understanding of how ICL learns the knowledge from the given prompts. In this paper, to make progress toward understanding the learning behaviour of ICL, we train the same LLMs with the same demonstration examples via ICL and supervised learning (SL), respectively, and investigate their performance under label perturbations (i.e., noisy labels and label imbalance) on a range of classification tasks. First, via extensive experiments, we find that gold labels have significant impacts on the downstream in-context performance, especially for large language models; however, imbalanced labels matter little to ICL across all model sizes. Second, when comparing with SL, we show empirically that ICL is less sensitive to label perturbations than SL, and ICL gradually attains comparable performance to SL as the model size increases. Xindi Wang 0001, Yufei Wang 0003, Can Xu 0002, Xiubo Geng, Chongyang Tao, Frank Rudzicz, Robert E. Mercer, Daxin Jiang |
ECAI | 7 |
| 2023 | A State-Vector Framework for Dataset EffectsabstractThe impressive success of recent deep neural network (DNN)-based systems is significantly influenced by the high-quality datasets used in training.However, the effects of the datasets, especially how they interact with each other, remain underexplored.We propose a statevector framework to enable rigorous studies in this direction.This framework uses idealized probing test results as the bases of a vector space.This framework allows us to quantify the effects of both standalone and interacting datasets.We show that the significant effects of some commonly-used language understanding datasets are characteristic and are concentrated on a few linguistic dimensions.Additionally, we observe some "spill-over" effects: the datasets could impact the models along dimensions that may seem unrelated to the intended tasks.Our state-vector framework paves the way for a systematic understanding of the dataset effects, a crucial component in responsible and robust model development. Esmat Sahak, Zining Zhu 0001, Frank Rudzicz |
EMNLP | 3 |
| 2022 | Language Modelling via Learning to RankabstractWe consider language modelling (LM) as a multi-label structured prediction task by re-framing training from solely predicting a single ground-truth word to ranking a set of words which could continue a given context. To avoid annotating top-k ranks, we generate them using pre-trained LMs: GPT-2, BERT, and Born-Again models. This leads to a rank-based form of knowledge distillation (KD). We also develop a method using N-grams to create a non-probabilistic teacher which generates the ranks without the need of a pre-trained LM. We confirm the hypotheses: that we can treat LMing as a ranking task and that we can do so without the use of a pre-trained LM. We show that rank-based KD generally gives a modest improvement to perplexity (PPL) -- though often with statistical significance -- when compared to Kullback–Leibler-based KD. Surprisingly, given the naivety of the method, the N-grams act as competitive teachers and achieve similar performance as using either BERT or a Born-Again model teachers. Unsurprisingly, GPT-2 always acts as the best teacher. Using it and a Transformer-XL student on Wiki-02, rank-based KD reduces a cross-entropy baseline from 65.27 to 55.94 and against a KL-based KD of 56.70. Arvid Frydenlund, Frank Rudzicz |
AAAI | 3 |
| 2022 | Neural reality of argument structure constructionsabstractIn lexicalist linguistic theories, argument structure is assumed to be predictable from the meaning of verbs.As a result, the verb is the primary determinant of the meaning of a clause.In contrast, construction grammarians propose that argument structure is encoded in constructions (or form-meaning pairs) that are distinct from verbs.Decades of psycholinguistic research have produced substantial empirical evidence in favor of the construction view.Here we adapt several psycholinguistic studies to probe for the existence of argument structure constructions (ASCs) in Transformerbased language models (LMs).First, using a sentence sorting experiment, we find that sentences sharing the same construction are closer in embedding space than sentences sharing the same verb.Furthermore, LMs increasingly prefer grouping by construction with more input data, mirroring the behaviour of non-native language learners.Second, in a "Jabberwocky" priming-based experiment, we find that LMs associate ASCs with meaning, even in semantically nonsensical sentences.Our work offers the first evidence for ASCs in LMs and highlights the potential to devise novel probing methods grounded in psycholinguistic research. Transitive DitransitiveCaused-motion Resultative Throw Anita threw the hammer.Chris threw Linda the pencil.Pat threw the keys onto the roof.Lyn threw the box apart. Zining Zhu 0001, Guillaume Thomas, Frank Rudzicz, Yang Xu 0023 |
ACL (1) | 4 |
| 2022 | KenMeSH: Knowledge-enhanced End-to-end Biomedical Text LabellingabstractCurrently, Medical Subject Headings (MeSH) are manually assigned to every biomedical article published and subsequently recorded in the PubMed database to facilitate retrieving relevant information.With the rapid growth of the PubMed database, large-scale biomedical document indexing becomes increasingly important.MeSH indexing is a challenging task for machine learning, as it needs to assign multiple labels to each article from an extremely large hierachically organized collection.To address this challenge, we propose KenMeSH, an end-to-end model that combines new text features and a dynamic Knowledge-enhanced mask attention that integrates document features with MeSH label hierarchy and journal correlation features to index MeSH terms.Experimental results show the proposed method achieves state-of-the-art performance on a number of measures. Xindi Wang 0001, Robert E. Mercer, Frank Rudzicz |
ACL (1) | 3 |
| 2022 | Predicting Fine-Tuning Performance with ProbingabstractLarge NLP models have recently shown impressive performance in language understanding tasks, typically evaluated by their finetuned performance.Alternatively, probing has received increasing attention as being a lightweight method for interpreting the intrinsic mechanisms of large NLP models.In probing, post-hoc classifiers are trained on "out-ofdomain" datasets that diagnose specific abilities.While probing the language models has led to insightful findings, they appear disjointed from the development of models.This paper explores the utility of probing deep NLP models to extract a proxy signal widely used in model development -the fine-tuning performance.We find that it is possible to use the accuracies of only three probing tests to predict the fine-tuning performance with errors 40% -80% smaller than baselines.We further discuss possible avenues where probing can empower the development of deep NLP models. Zining Zhu 0001, Soroosh Shahtalebi, Frank Rudzicz |
EMNLP | 3 |
| 2022 | A Remedy For Distributional Shifts Through Expected Domain TranslationabstractMachine learning models often fail to generalize to unseen domains due to the distributional shifts. A family of such shifts, “correlation shifts,” is caused by spurious correlations in the data. It is studied under the overarching topic of “domain generalization.” In this work, we employ multi-modal translation networks to tackle the correlation shifts that appear when data is sampled out-of-distribution. Learning a generative model from training domains enables us to translate each training sample under the special characteristics of other possible domains. We show that by training a predictor solely on the generated samples, the spurious correlations in training domains average out, and the invariant features corresponding to true correlations emerge. Our proposed technique, Expected Domain Translation (EDT), is benchmarked on the Colored MNIST dataset and drastically improves the state-of-the-art classification accuracy by 38% with train-domain validation model selection. Jean-Christophe Gagnon-Audet, Soroosh Shahtalebi, Frank Rudzicz, Irina Rish |
ICASSP | 3 |
| 2022 | MeSHup: Corpus for Full Text Biomedical Document IndexingabstractMedical Subject Heading (MeSH) indexing refers to the problem of assigning a given biomedical document with the most relevant labels from an extremely large set of MeSH terms. Currently, the vast number of biomedical articles in the PubMed database are manually annotated by human curators, which is time consuming and costly; therefore, a computational system that can assist the indexing is highly valuable. When developing supervised MeSH indexing systems, the availability of a large-scale annotated text corpus is desirable. A publicly available, large corpus that permits robust evaluation and comparison of various systems is important to the research community. We release a large scale annotated MeSH indexing corpus, MeSHup, which contains 1,342,667 full text articles, together with the associated MeSH labels and metadata, authors and publication venues that are collected from the MEDLINE database. We train an end-to-end model that combines features from documents and their associated labels on our corpus and report the new baseline. Xindi Wang 0001, Robert E. Mercer, Frank Rudzicz |
LREC | 3 |
| 2021 | How is BERT surprised? Layerwise detection of linguistic anomaliesabstractBai Li, Zining Zhu, Guillaume Thomas, Yang Xu, Frank Rudzicz. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zining Zhu 0001, Guillaume Thomas, Yang Xu 0023, Frank Rudzicz |
ACL/IJCNLP (1) | 5 |
| 2021 | Coughwatch: Real-World Cough Detection using SmartwatchesabstractContinuous monitoring of cough may provide insights into the health of individuals as well as the effectiveness of treatments. Smart-watches, in particular, are highly promising for such monitoring: they are inexpensive, unobtrusive, programmable, and have a variety of sensors. However, current mobile cough detection systems are not designed for smartwatches, and perform poorly when applied to real-world smartwatch data since they are often evaluated on data collected in the lab.In this work we propose CoughWatch, a lightweight cough detector for smartwatches that uses audio and movement data for in-the-wild cough detection. On our in-the-wild data, CoughWatch achieves a precision of 82% and recall of 55%, compared to 6% precision and 19% recall achieved by the current state-of-the-art approach. Furthermore, by incorporating gyroscope and accelerometer data, CoughWatch improves precision by up to 15.5 percentage points compared to an audio-only model. Daniyal Liaqat, Salaar Liaqat, Jun Lin Chen, Tina Sedaghat, Moshe Gabel, Frank Rudzicz, Eyal de Lara |
ICASSP | 6 |
| 2021 | Grad2Task: Improved Few-shot Text Classification Using Gradients for Task RepresentationabstractLarge pretrained language models (LMs) like BERT have improved performance in many disparate natural language processing (NLP) tasks. However, fine tuning such models requires a large number of training examples for each target task. Simultaneously, many realistic NLP problems are "few shot", without a sufficiently large training set. In this work, we propose a novel conditional neural process-based approach for few-shot text classification that learns to transfer from other diverse tasks with rich annotation. Our key idea is to represent each task using gradient information from a base model and to train an adaptation network that modulates a text classifier conditioned on the task representation. While previous task-aware few-shot learners represent tasks by input encoding, our novel task representation is more powerful, as the gradient captures input-output relationships of a task. Experimental results show that our approach outperforms traditional fine-tuning, sequential transfer learning, and state-of-the-art meta learning approaches on a collection of diverse few-shot tasks. We further conducted analysis and ablations to justify our design choices. Jixuan Wang, Kuan-Chieh Wang, Frank Rudzicz, Michael Brudno |
NeurIPS | 3 |
| 2020 | On Losses for Modern Language ModelsabstractBERT set many state-of-the-art results over varied NLU benchmarks by pre-training over two tasks: masked language modelling (MLM) and next sentence prediction (NSP), the latter of which has been highly criticized.In this paper, we 1) clarify NSP's effect on BERT pre-training, 2) explore fourteen possible auxiliary pre-training tasks, of which seven are novel to modern language models, and 3) investigate different ways to include multiple tasks into pre-training.We show that NSP is detrimental to training due to its context splitting and shallow semantic signal.We also identify six auxiliary pre-training tasks -sentence ordering, adjacent sentence prediction, TF prediction, TF-IDF prediction, a Fast-Sent variant, and a Quick Thoughts variant -that outperform a pure MLM baseline.Finally, we demonstrate that using multiple tasks in a multi-task pre-training framework provides better results than using any single auxiliary task.Using these methods, we outperform BERT Base on the GLUE benchmark using fewer than a quarter of the training tokens. Stephane Aroca-Ouellette, Frank Rudzicz |
EMNLP (1) | 2 |
| 2020 | Explainable Clinical Decision Support from TextabstractClinical prediction models often use structured variables and provide outcomes that are not readily interpretable by clinicians.Further, free-text medical notes may contain information not immediately available in structured variables.We propose a hierarchical CNNtransformer model with explicit attention as an interpretable, multi-task clinical language model, which achieves an AUROC of 0.75 and 0.78 on sepsis and mortality prediction on the English MIMIC-III dataset, respectively.We also explore the relationships between learned features from structured and unstructured variables using projection-weighted canonical correlation analysis.Finally, we outline a protocol to evaluate model usability in a clinical decision support context.From domain-expert evaluations, our model generates informative rationales that have promising real-life applications. Jinyue Feng, Chantal Shaib, Frank Rudzicz |
EMNLP (1) | 3 |
| 2020 | Word class flexibility: A deep contextualized approachabstractWord class flexibility refers to the phenomenon whereby a single word form is used across different grammatical categories.Extensive work in linguistic typology has sought to characterize word class flexibility across languages, but quantifying this phenomenon accurately and at scale has been fraught with difficulties.We propose a principled methodology to explore regularity in word class flexibility.Our method builds on recent work in contextualized word embeddings to quantify semantic shift between word classes (e.g., noun-to-verb, verb-to-noun), and we apply this method to 37 languages 1 .We find that contextualized embeddings not only capture human judgment of class variation within words in English, but also uncover shared tendencies in class flexibility across languages.Specifically, we find greater semantic variation when flexible lemmas are used in their dominant word class, supporting the view that word class flexibility is a directional process.Our work highlights the utility of deep contextualized models in linguistic typology. Guillaume Thomas, Yang Xu 0023, Frank Rudzicz |
EMNLP (1) | 4 |
| 2020 | An information theoretic view on selecting linguistic probesabstractThere is increasing interest in assessing the linguistic knowledge encoded in neural representations.A popular approach is to attach a diagnostic classifier -or "probe" -to perform supervised classification from internal representations.However, how to select a good probe is in debate.Hewitt and Liang (2019) showed that a high performance on diagnostic classification itself is insufficient, because it can be attributed to either "the representation being rich in knowledge", or "the probe learning the task", which Pimentel et al. (2020) challenged.We show this dichotomy is valid informationtheoretically.In addition, we find that the methods to construct and select good probes proposed by the two papers, control task (Hewitt and Liang, 2019) and control function (Pimentel et al., 2020), are equivalent -the errors of their approaches are identical (modulo irrelevant terms).Empirically, these two selection criteria lead to results that highly agree with each other. Zining Zhu 0001, Frank Rudzicz |
EMNLP (1) | 2 |
| 2020 | Speaker Diarization with Session-Level Speaker Embedding Refinement Using Graph Neural NetworksabstractDeep speaker embedding models have been commonly used as a building block for speaker diarization systems; however, the speaker embedding model is usually trained according to a global loss defined on the training data, which could be suboptimal for distinguishing speakers locally in a specific meeting session. In this work we present the first use of graph neural networks (GNNs) for the speaker diarization problem, utilizing a GNN to refine speaker embeddings locally using the structural information between speech segments inside each session. The speaker embeddings extracted by a pre-trained model are remapped into a new embedding space, in which the different speakers within a single session are better separated. The model is trained for linkage prediction in a supervised manner by minimizing the difference between the affinity matrix constructed by the refined embeddings and the ground-truth adjacency matrix. Spectral clustering is then applied on top of the refined embeddings. We show that the clustering performance of the refined speaker embeddings outperforms the original embeddings significantly on both simulated and real meeting data, and our system achieves the state-of-the-art result on the NIST SRE 2000 CALLHOME database. Jixuan Wang, Jian Wu 0027, Ranjani Ramamurthy, Frank Rudzicz, Michael Brudno |
ICASSP | 5 |
| 2020 | To BERT or not to BERT: Comparing Speech and Language-Based Approaches for Alzheimer's Disease DetectionabstractResearch related to automatically detecting Alzheimer's disease (AD) is important, given the high prevalence of AD and the high cost of traditional methods. Since AD significantly affects the content and acoustics of spontaneous speech, natural language processing and machine learning provide promising techniques for reliably detecting AD. We compare and contrast the performance of two such approaches for AD detection on the recent ADReSS challenge dataset: 1) using domain knowledge-based hand-crafted features that capture linguistic and acoustic phenomena, and 2) fine-tuning Bidirectional Encoder Representations from Transformer (BERT)-based sequence classification models. We also compare multiple feature-based regression models for a neuropsychological score task in the challenge. We observe that fine-tuned BERT models, given the relative importance of linguistics in cognitive impairment detection, outperform feature-based approaches on the AD detection task. Aparna Balagopalan, Benjamin Eyre, Frank Rudzicz, Jekaterina Novikova |
INTERSPEECH | 3 |
| 2020 | Speaker Attribution with Voice Profiles by Graph-Based Semi-Supervised LearningabstractSpeaker attribution is required in many real-world applications, such as meeting transcription, where speaker identity is assigned to each utterance according to speaker voice profiles. In this paper, we propose to solve the speaker attribution problem by using graph-based semi-supervised learning methods. A graph of speech segments is built for each session, on which segments from voice profiles are represented by labeled nodes while segments from test utterances are unlabeled nodes. The weight of edges between nodes is evaluated by the similarities between the pretrained speaker embeddings of speech segments. Speaker attribution then becomes a semi-supervised learning problem on graphs, on which two graph-based methods are applied: label propagation (LP) and graph neural networks (GNNs). The proposed approaches are able to utilize the structural information of the graph to improve speaker attribution performance. Experimental results on real meeting data show that the graph based approaches reduce speaker attribution error by up to 68% compared to a baseline speaker identification approach that processes each utterance independently. Jixuan Wang, Jian Wu 0027, Ranjani Ramamurthy, Frank Rudzicz, Michael Brudno |
INTERSPEECH | 5 |
| 2020 | Identification of Primary and Collateral Tracks in Stuttered SpeechabstractDisfluent speech has been previously addressed from two main perspectives: the clinical perspective focusing on diagnostic, and the Natural Language Processing (NLP) perspective aiming at modeling these events and detect them for downstream tasks. In addition, previous works often used different metrics depending on whether the input features are text or speech, making it difficult to compare the different contributions. Here, we introduce a new evaluation framework for disfluency detection inspired by the clinical and NLP perspective together with the theory of performance from (Clark, 1996) which distinguishes between primary and collateral tracks. We introduce a novel forced-aligned disfluency dataset from a corpus of semi-directed interviews, and present baseline results directly comparing the performance of text-based features (word and span information) and speech-based (acoustic-prosodic information). Finally, we introduce new audio features inspired by the word-based span features. We show experimentally that using these features outperformed the baselines for speech-based predictions on the present dataset. Rachid Riad, Anne-Catherine Bachoud-Lévi, Frank Rudzicz, Emmanuel Dupoux |
LREC | 3 |
| 2020 | Using word embeddings to improve the privacy of clinical notesabstractOBJECTIVE: In this work, we introduce a privacy technique for anonymizing clinical notes that guarantees all private health information is secured (including sensitive data, such as family history, that are not adequately covered by current techniques). MATERIALS AND METHODS: We employ a new "random replacement" paradigm (replacing each token in clinical notes with neighboring word vectors from the embedding space) to achieve 100% recall on the removal of sensitive information, unachievable with current "search-and-secure" paradigms. We demonstrate the utility of this paradigm on multiple corpora in a diverse set of classification tasks. RESULTS: We empirically evaluate the effect of our anonymization technique both on upstream and downstream natural language processing tasks to show that our perturbations, while increasing security (ie, achieving 100% recall on any dataset), do not greatly impact the results of end-to-end machine learning approaches. DISCUSSION: As long as current approaches utilize precision and recall to evaluate deidentification algorithms, there will remain a risk of overlooking sensitive information. Inspired by differential privacy, we sought to make it statistically infeasible to recreate the original data, although at the cost of readability. We hope that the work will serve as a catalyst to further research into alternative deidentification methods that can address current weaknesses. CONCLUSION: Our proposed technique can secure clinical texts at a low cost and extremely high recall with a readability trade-off while remaining useful for natural language processing classification tasks. We hope that our work can be used by risk-averse data holders to release clinical texts to researchers. Mohamed Abdalla 0001, Moustafa Abdalla, Frank Rudzicz, Graeme Hirst |
J. Am. Medical Informatics Assoc. | 3 |
| 2020 | A Conversational Robot for Older Adults with Alzheimer's DiseaseabstractAmid the rising cost of Alzheimer’s disease (AD), assistive health technologies can reduce care-giving burden by aiding in assessment, monitoring, and therapy. This article presents a pilot study testing the feasibility and effect of a conversational robot in a cognitive assessment task with older adults with AD. We examine the robot interactions through dialogue and miscommunication analysis, linguistic feature analysis, and the use of a qualitative analysis, in which we report key themes that were prevalent throughout the study. While conversations were typically better with human conversation partners (being longer, with greater engagement and less misunderstanding), we found that the robot was generally well liked by participants and that it was able to capture their interest in dialogue. Miscommunication due to issues of understanding and intelligibility did not seem to deter participants from their experience. Furthermore, in automatically extracting linguistic features, we examine how non-acoustic aspects of language change across participants with varying degrees of cognitive impairment, highlighting the robot’s potential as a monitoring tool. This pilot study is an exploration of how conversational robots can be used to support individuals with AD. Chloé Pou-Prom, Stefania Raimondo, Frank Rudzicz |
ACM Trans. Hum. Robot Interact. | 3 |
| 2019 | Centroid-based Deep Metric Learning for Speaker RecognitionabstractSpeaker embedding models that utilize neural networks to map utterances to a space where distances reflect similarity between speakers have driven recent progress in the speaker recognition task. However, there is still a significant performance gap between recognizing speakers in the training set and unseen speakers. The latter case corresponds to the few-shot learning task, where a trained model is evaluated on unseen classes. Here, we optimize a speaker embedding model with prototypical network loss (PNL), a state-of-the-art approach for the few-shot image classification task. The resulting embedding model outperforms the state-of-the-art triplet loss based models in both speaker verification and identification tasks, for both seen and unseen speakers. Jixuan Wang, Kuan-Chieh Wang, Marc T. Law, Frank Rudzicz, Michael Brudno |
ICASSP | 4 |
| 2018 | Learning multiview embeddings for assessing dementiaabstractAs the incidence of Alzheimer's Disease (AD) increases, early detection becomes crucial.Unfortunately, datasets for AD assessment are often sparse and incomplete.In this work, we leverage the multiview nature of a small AD dataset, DementiaBank, to learn an embedding that captures different modes of cognitive impairment.We apply generalized canonical correlation analysis (GCCA) to our dataset and demonstrate the added benefit of using multiview embeddings in two downstream tasks: identifying AD and predicting clinical scores.By including multiview embeddings, we obtain an F1 score of 0.82 in the classification task and a mean absolute error of 3.42 in the regression task.Furthermore, we show that multiview embeddings can be obtained from other datasets as well. Chloé Pou-Prom, Frank Rudzicz |
EMNLP | 2 |
| 2018 | Touch-Supported Voice Recording to Facilitate Forced Alignment of Text and Speech in an E-Reading InterfaceabstractReading a book together with a family member who has impaired vision or other difficulties reading is an important social bonding activity. However, for the person being read to, there is little support in making these experiences repeatable. While audio can easily be recorded, synchronizing it with the text for later playback requires the use of forced alignment algorithms, which do not perform well on amateur read-aloud speech. We propose a human-in-the-loop approach to augmenting such algorithms, in the form of touch metaphors during collocated read-aloud sessions using tablet e-readers. The metaphor is implemented as a finger-follows-text tracker. We explore how this could better handle the variability of amateur reading, which poses accuracy challenges for existing forced alignment techniques. Data collected from users reading aloud as assisted by touch metaphors show increases in the accuracy of forced alignment algorithms and reveal opportunities for how to better support reading aloud. Benett Axtell, Cosmin Munteanu, Carrie Demmans Epp, Yomna Aly, Frank Rudzicz |
IUI | 5 |
| 2018 | Speech in Smartwatch based AudioabstractNo abstract available. Daniyal Liaqat, Robert Wu 0002, Andrea Gershon, Hisham Alshaer, Frank Rudzicz, Eyal de Lara |
MobiSys | 5 |
| 2018 | Modified mean shift algorithmabstractThe mean shift (MS) algorithm is an iterative method introduced for locating modes of a probability density function. Although the MS algorithm has been widely used in many applications, the convergence of the algorithm has not yet been proven. In this study, the authors modify the MS algorithm in order to guarantee its convergence. The authors prove that the generated sequence using the proposed modified algorithm is a convergent sequence and the density estimate values along the generated sequence are monotonically increasing and convergent. In contrast to the MS algorithm, the proposed modified version does not require setting a stopping criterion a priori; instead, it guarantees the convergence after a finite number of iterations. The proposed modified version defines an upper bound for the number of iterations which is missing in the MS algorithm. The authors also present the matrix form of the proposed algorithm and show that, in contrast to the MS algorithm, the weight matrix is required to be computed once in the first iteration. The performance of the proposed modified version is compared with the MS algorithm and it was shown through the simulations that the proposed version can be used successfully to estimate cluster centres. Youness Aliyari Ghassabeh, Frank Rudzicz |
IET Image Process. | 2 |
| 2017 | On the impact of non-modal phonation on phonological featuresabstractDifferent modes of vibration of the vocal folds contribute significantly to the voice quality. The neutral mode phonation, often used in a modal voice, is one against which the other modes can be contrastively described, also called non-modal phonations. This paper investigates the impact of non-modal phonation on phonological posteriors, the probabilities of phonological features inferred from the speech signal using a deep learning approach. Five different non-modal phonations are considered: falsetto, creaky, harshness, tense and breathiness. The impact of such non-modal phonation on phonological features, the Sound Patterns of English (SPE), is investigated in both speech analysis and synthesis tasks. We found that breathy and tense phonation impact the SPE features less, creaky phonation impacts the features moderately, and harsh and falsetto phonation impact the phonological features the most. We also report invariant and the most different SPE features impacted by non-modal phonation. Milos Cernak, Elmar Nöth, Frank Rudzicz, Heidi Christensen, Juan Rafael Orozco-Arroyave, Raman Arora, Tobias Bocklet, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Juan Camilo Vásquez-Correa, Maria Yancheva, Alyssa Vann, Nikolai Vogler |
ICASSP | 3 |
| 2017 | Multi-view representation learning via gcca for multimodal analysis of Parkinson's diseaseabstractInformation from different bio-signals such as speech, handwriting, and gait have been used to monitor the state of Parkinson's disease (PD) patients, however, all the multimodal bio-signals may not always be available. We propose a method based on multi-view representation learning via generalized canonical correlation analysis (GCCA) for learning a representation of features extracted from handwriting and gait that can be used as a complement to speech-based features. Three different problems are addressed: classification of PD patients vs. healthy controls, prediction of the neurological state of PD patients according to the UPDRS score, and the prediction of a modified version of the Frenchay dysarthria assessment (m-FDA). According to the results, the proposed approach is suitable to improve the results in the addressed problems, specially in the prediction of the UPDRS, and m-FDA scores. Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Raman Arora, Elmar Nöth, Najim Dehak, Heidi Christensen, Frank Rudzicz, Tobias Bocklet, Milos Cernak, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Maria Yancheva, Alyssa Vann, Nikolai Vogler |
ICASSP | 7 |
| 2017 | Identifying and Avoiding Confusion in Dialogue with People with Alzheimer's DiseaseabstractAlzheimer's disease (AD) is an increasingly prevalent cognitive disorder in which memory, language, and executive function deteriorate, usually in that order. There is a growing need to support individuals with AD and other forms of dementia in their daily lives, and our goal is to do so through speech-based interaction. Given that 33% of conversations with people with middle-stage AD involve a breakdown in communication, it is vital that automated dialogue systems be able to identify those breakdowns and, if possible, avoid them. In this article, we discuss several linguistic features that are verbal indicators of confusion in AD (including vocabulary richness, parse tree structures, and acoustic cues) and apply several machine learning algorithms to identify dialogue-relevant confusion from speech with up to 82% accuracy. We also learn dialogue strategies to avoid confusion in the first place, which is accomplished using a partially observable Markov decision process and which obtains accuracies (up to 96.1%) that are significantly higher than several baselines. This work represents a major step towards automated dialogue systems for individuals with dementia. Hamid R. Chinaei, Leila Chan Currie, Andrew Danks, Hubert Lin, Tejas Mehta, Frank Rudzicz |
Comput. Linguistics | 6 |
| 2017 | Characterisation of voice quality of Parkinson's disease using differential phonological posterior features
Milos Cernak, Juan Rafael Orozco-Arroyave, Frank Rudzicz, Heidi Christensen, Juan Camilo Vásquez-Correa, Elmar Nöth |
Comput. Speech Lang. | 3 |
| 2016 | Vector-space topic models for detecting Alzheimer's diseaseabstractSemantic deficit is a symptom of language impairment in Alzheimer's disease (AD).We present a generalizable method for automatic generation of information content units (ICUs) for a picture used in a standard clinical task, achieving high recall, 96.8%, of human-supplied ICUs.We use the automatically generated topic model to extract semantic features, and train a random forest classifier to achieve an F-score of 0.74 in binary classification of controls versus people with AD using a set of only 12 features.This is comparable to results (0.72 F-score) with a set of 85 manual features.Adding semantic information to a set of standard lexicosyntactic and acoustic features improves F-score to 0.80.While control and dementia subjects discuss the same topics in the same contexts, controls are more informative per second of speech. Maria Yancheva, Frank Rudzicz |
ACL (1) | 2 |
| 2016 | CloudCAST - Remote Speech Technology for Speech ProfessionalsabstractInternational audience Phil D. Green, Ricard Marxer, Stuart P. Cunningham, Heidi Christensen, Frank Rudzicz, Maria Yancheva, André Coy, Massimiliano Malavasi, Lorenzo Desideri, Fabio Tamburini |
INTERSPEECH | 5 |
| 2016 | Speech Recognition in Alzheimer's Disease and in its Assessment
Luke Zhou, Kathleen C. Fraser, Frank Rudzicz |
INTERSPEECH | 3 |
| 2016 | Speech Production in Speech Technologies: Introduction to the CSL Special Issue
Karen Livescu, Frank Rudzicz, Eric Fosler-Lussier, Mark Hasegawa-Johnson, Jeff A. Bilmes |
Comput. Speech Lang. | 2 |
| 2016 | Principal differential analysis for detection of bilabial closure gestures from articulatory data
Farook Sattar, Frank Rudzicz |
Comput. Speech Lang. | 2 |
| 2016 | The mean shift algorithm and its relation to kernel regression
Youness Aliyari Ghassabeh, Frank Rudzicz |
Inf. Sci. | 2 |
| 2016 | Fast adaptive algorithms for optimal feature extraction from Gaussian data
Youness Aliyari Ghassabeh, Frank Rudzicz, Hamid Abrishami Moghaddam |
Pattern Recognit. Lett. | 2 |
| 2016 | Acoustic-articulatory relationships and inversion in sum-product and deep-belief networks
Frank Rudzicz, Arvid Frydenlund, Sean Robertson, Patricia Thaine |
Speech Commun. | 1 |
| 2016 | Manifold Learning for Multivariate Variable-Length Sequences With an Application to Similarity SearchabstractMultivariate variable-length sequence data are becoming ubiquitous with the technological advancement in mobile devices and sensor networks. Such data are difficult to compare, visualize, and analyze due to the nonmetric nature of data sequence similarity measures. In this paper, we propose a general manifold learning framework for arbitrary-length multivariate data sequences driven by similarity/distance (parameter) learning in both the original data sequence space and the learned manifold. Our proposed algorithm transforms the data sequences in a nonmetric data sequence space into feature vectors in a manifold that preserves the data sequence space structure. In particular, the feature vectors in the manifold representing similar data sequences remain close to one another and far from the feature points corresponding to dissimilar data sequences. To achieve this objective, we assume a semisupervised setting where we have knowledge about whether some of data sequences are similar or dissimilar, called the instance-level constraints. Using this information, one learns the similarity measure for the data sequence space and the distance measures for the manifold. Moreover, we describe an approach to handle the similarity search problem given user-defined instance level constraints in the learned manifold using a consensus voting scheme. Experimental results on both synthetic data and real tropical cyclone sequence data are presented to demonstrate the feasibility of our manifold learning framework and the robustness of performing similarity search in the learned manifold. Shen-Shyang Ho, Peng Dai 0002, Frank Rudzicz |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | EEG dimensionality reduction in automatic identification of synonymyabstractRecent work has demonstrated the feasibility of extracting semantic categories directly from cortical measures (e.g., electroencephalography, EEG) during receptive tasks. Here, we automatically classify speech stimuli as either synonymous or non-synonymous with a prior prime in a speech-receptive task given only EEG data with up to 86.84% accuracy. An analysis of variance reveals no significant difference among support vector machine and k-nearest neighbours classifiers, but a significant effect of the individual subject on accuracy. To perform classification, we reduce the highly-parameterized space by three successive techniques: a ranking based on t-test similarity, another based on principal components analysis (PCA), and a third on linear discriminant analysis. Emilio Parisotto, Youness Aliyari Ghassabeh, Siamak Freydoonnejad, Frank Rudzicz |
ICASSP | 4 |
| 2015 | Classifying phonological categories in imagined and articulated speechabstractThis paper presents a new dataset combining 3 modalities (EEG, facial, and audio) during imagined and vocalized phonemic and single-word prompts. We pre-process the EEG data, compute features for all 3 modalities, and perform binary classification of phonological categories using a combination of these modalities. For example, a deep-belief network obtains accuracies over 90% on identifying consonants, which is significantly more accurate than two baseline support vector machines. We also classify between the different states (resting, stimuli, active thinking) of the recording, achieving accuracies of 95%. These data may be used to learn multimodal relationships, and to develop silent-speech and brain-computer interfaces. Shunan Zhao, Frank Rudzicz |
ICASSP | 2 |
| 2015 | Lateralization in emotional speech perception following transcranial direct current stimulationabstractThe degree to which the perception of spoken emotion is lateralized in the brain remains a controversial topic. This work examines hemispheric differences in the perception of emotion in speech by applying tDCS, a neurostimulation protocol, to the T-RES speech emotion rating paradigm. We find several significant effects, including a strong interaction of prosody and neurostimulation for perceptual ratings when considering only lexical content, and that the perception of happiness does not appear to be affected by tDCS, but anger and (to a large extent) fear appear less intense after stimulation. Alex Francois-Nienaber, Jed A. Meltzer, Frank Rudzicz |
INTERSPEECH | 3 |
| 2015 | Automatic identification of received language in MEG
Emilio Parisotto, Youness Aliyari Ghassabeh, Matt J. MacDonald, Adelina Cozma, Elizabeth W. Pang, Frank Rudzicz |
INTERSPEECH | 6 |
| 2015 | Incremental algorithm for finding principal curvesabstractPrincipal curves are a non‐linear generalisation of principal components. They are smooth curves that pass through the middle of a data set to provide a new representation of those data to make tasks, such as visualisation and dimensionality reduction easier and more accurate. The subspace constrained mean shift (SCMS) algorithm is a recently proposed technique to find principal curves. The algorithm assumes that the complete data set is available in advance and that new data points cannot be added to the data set during the process. The algorithm finds the points on the principal curves by using the complete data set. In this paper, the authors investigate the situation where the entire data set is not available in advance and instead are sampled sequentially. They propose an incremental version of the SCMS algorithm that trains using a sequence of observations. Simulation results show the effectiveness of the proposed algorithm to find a principal curve using a stream of observations. Youness Aliyari Ghassabeh, Frank Rudzicz |
IET Signal Process. | 2 |
| 2015 | Sequential behavior prediction based on hybrid similarity and cross-user activity transfer
Peng Dai 0002, Shen-Shyang Ho, Frank Rudzicz |
Knowl. Based Syst. | 3 |
| 2015 | Fast incremental LDA feature extraction
Youness Aliyari Ghassabeh, Frank Rudzicz, Hamid Abrishami Moghaddam |
Pattern Recognit. | 2 |
| 2015 | 2D Psychoacoustic modeling of equivalent masking for automatic speech recognition
Peng Dai 0002, Frank Rudzicz, Ing Yann Soon, Alex Mihailidis, Huijun Ding |
Signal Process. | 2 |
| 2014 | Automatically identifying trouble-indicating speech behaviors in alzheimer's diseaseabstractAlzheimer's disease (AD) deteriorates executive, linguistic, and functional capacity and is rapidly becoming more prevalent. In particular, AD leads to an inability to follow simple dialogues. In this paper, we annotate two databases of dyad conversations, that include individuals with AD, with trouble indicating behaviors (TIBs). We then extract lexical/syntactic and acoustic features from all utterances and identify those that are most indicative of TIB (which include speech rate and utterance likelihoods in a standard language model) and classify utterances as having TIB or not with up to 79.5% accuracy. This will allow us to build automated dialogue systems and assessment tools that are sensitive to confusion in people with AD. Frank Rudzicz, Leila Chan Currie, Andrew Danks, Tejas Mehta, Shunan Zhao |
ASSETS | 1 |
| 2014 | Subject independent identification of breath sounds components using multiple classifiersabstractBreath sounds have been shown very valuable for diagnosis of obstructive sleep apnea. In this study, we present a subject independent method for automatic classification of breath and related sounds during sleep. An experienced operator manually labelled segments of breath sounds from 11 sleeping subjects as: inspiration, expiration, inspiratory snoring, expiratory snoring, wheezing, other noise, and non-audible. Ten features were extracted and fed into 3 different classifiers: näıve Bayes, Support Vector Machine, and Random Forest. Leave-one-out method was used in which data from each subject, in turn, is evaluated using models trained with all other subject. Mean accuracy for concurrent classification of all 7 classes reached 85.4%. Mean accuracy for separating data into 2 classes, snoring and non-snoring, reached 97.8%. To our knowledge, these are the highest accuracies achieved in automatic classification of all breath sounds components concurrently and for snoring, in a subject independent model. Hisham Alshaer, Aditya Pandya, T. Douglas Bradley, Frank Rudzicz |
ICASSP | 4 |
| 2014 | Automatic detection of expressed emotion in Parkinson's DiseaseabstractPatients with Parkinsons Disease (PD) frequently exhibit deficits in the production of emotional speech. In this paper, we examine the classification of emotional speech in patients with PD and the classification of PD speech. Participants were recorded speaking short statements with different emotional prosody which were classified with three methods (naïve Bayes, random forests, and support vector machines) using 209 unique auditory features. Feature sets were reduced using simple statistical testing. We achieve accuracies of 65.5% and 73.33% on classifying between the emotions and between PD vs. control, respectively. These results may assist in the future development of automated early detection systems for diagnosing patients with PD. Shunan Zhao, Frank Rudzicz, Leonardo G. Carvalho, Cesar Marquez-Chin, Steven R. Livingstone |
ICASSP | 2 |
| 2014 | Noisy Source Vector Quantization Using Kernel RegressionabstractThe problem of designing an optimal vector quantizer when there is access to the noise-free source has been well studied over the past five decades. However, in many real-world situations, the source output may be corrupted by some additive noise. In this case, we only have access to a noisy version of the data, but we expect a designed quantizer to minimize the distortion with respect to the clean (unavailable) data. It can be shown that the mean square distortion for an optimal noisy source vector quantization system can be decomposed into an optimum estimator, followed by an optimum source coder operating on the estimator output. We summarize this result first and then propose to use the kernel regression technique for estimating the clean data from the noisy version. The output of the kernel regression, as an estimate of the clean data, is quantized using the LBG vector quantizer. The proposed structure requires two sets of training data. The first set is used to train the kernel regression estimator. The second set is fed into the trained kernel regression system whose output is used to train the LBG vector quantizer. We show the effectiveness of the proposed structure through simulations with different numbers of code words. Youness Aliyari Ghassabeh, Frank Rudzicz |
IEEE Trans. Commun. | 2 |
| 2013 | Automatic detection of deception in child-produced speech using syntactic complexity features
Maria Yancheva, Frank Rudzicz |
ACL (1) | 2 |
| 2013 | Using text and acoustic features to diagnose progressive aphasia and its subtypesabstractThis paper presents experiments in automatically diagnosing primary progressive aphasia (PPA) and two of its subtypes, semantic dementia (SD) and progressive nonfluent aphasia (PNFA), from the acoustics of recorded narratives and textual analysis of the resultant transcripts. In order to train each of three types of classifier (naive Bayes, support vector machine, random forest), a large set of 81 available features must be reduced in size. Two methods of feature selection are therefore compared – one based on statistical significance and the other based on minimum-redundancy-maximum-relevance. After classifier optimization, PPA (or absence thereof) is correctly diagnosed across 87.4% of conditions, and the two subtypes of PPA are correctly classified 75.6% of the time. Kathleen C. Fraser, Frank Rudzicz, Elizabeth Rochon |
INTERSPEECH | 2 |
| 2013 | Adjusting dysarthric speech signals to be more intelligible
Frank Rudzicz |
Comput. Speech Lang. | 1 |
| 2012 | Sentence recognition from articulatory movements for silent speech interfacesabstractRecent research has demonstrated the potential of using an articulation-based silent speech interface for command-and-control systems. Such an interface converts articulation to words that can then drive a text-to-speech synthesizer. In this paper, we have proposed a novel near-time algorithm to recognize whole-sentences from continuous tongue and lip movements. Our goal is to assist persons who are aphonic or have a severe motor speech impairment to produce functional speech using their tongue and lips. Our algorithm was tested using a functional sentence data set collected from ten speakers (3012 utterances). The average accuracy was 94.89% with an average latency of 3.11 seconds for each sentence prediction. The results indicate the effectiveness of our approach and its potential for building a real-time articulation-based silent speech interface for clinical applications. Jun Wang 0037, Ashok Samal, Jordan R. Green, Frank Rudzicz |
ICASSP | 4 |
| 2012 | Whole-Word Recognition from Articulatory Movements for Silent Speech InterfacesabstractArticulation-based silent speech interfaces convert silently produced speech movements into audible words. These systems are still in their experimental stages, but have significant potential for facilitating oral communication in persons with laryngectomy or speech impairments. In this paper, we report the result of a novel, real-time algorithm that recognizes whole-words based on articulatory movements. This approach differs from prior work that has focused primarily on phoneme-level recognition based on articulatory features. On average, our algorithm missed 1.93 words in a sequence of twenty-five words with an average latency of 0.79 seconds for each word prediction using a data set of 5,500 isolated word samples collected from ten speakers. The results demonstrate the effectiveness of our approach and its potential for building a real-time articulation-based silent speech interface for health applications. Jun Wang 0037, Ashok Samal, Jordan R. Green, Frank Rudzicz |
INTERSPEECH | 4 |
| 2012 | Using articulatory likelihoods in the recognition of dysarthric speech
Frank Rudzicz |
Speech Commun. | 1 |
| 2011 | Adapting acoustic and lexical models to dysarthric speechabstractDysarthria is a motor speech disorder resulting from neurological damage to the part of the brain that controls the physical production of speech. It is, in part, characterized by pronunciation errors that include deletions, substitutions, insertions, and distortions of phonemes. These errors follow consistent intra-speaker patterns that we exploit through acoustic and lexical model adaptation to improve automatic speech recognition (ASR) on dysarthric speech. We show that acoustic model adaptation yields an average relative word error rate (WER) reduction of 36.99% and that pronunciation lexicon adaptation (PLA) further reduces the relative WER by an average of 8.29% on a large vocabulary task of over 1500 words for six speakers with severe to moderate dysarthria. PLA also shows an average relative WER reduction of 7.11% on speaker-dependent models evaluated using 5-fold cross-validation. Kinfe Tadesse Mengistu, Frank Rudzicz |
ICASSP | 2 |
| 2011 | Articulatory Knowledge in the Recognition of Dysarthric SpeechabstractDisabled speech is not compatible with modern generative and acoustic-only models of speech recognition (ASR). This work considers the use of theoretical and empirical knowledge of the vocal tract for atypical speech in labeling segmented and unsegmented sequences. These combined models are compared against discriminative models such as neural networks, support vector machines, and conditional random fields. Results show significant improvements in accuracy over the baseline through the use of production knowledge. Furthermore, although the statistics of vocal tract movement do not appear to be transferable between regular and disabled speakers, transforming the space of the former given knowledge of the latter before retraining gives high accuracy. This work may be applied within components of assistive software for speakers with dysarthria. Frank Rudzicz |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2010 | Correcting Errors in Speech Recognition with Articulatory Dynamics
Frank Rudzicz |
ACL | 1 |
| 2010 | Adaptive Kernel Canonical Correlation Analysis for Estimation of Task Dynamics from Acoustics
Frank Rudzicz |
ICASSP | 1 |
| 2010 | Identifying articulatory goals from kinematic data using principal differential analysisabstractArticulatory goals can be highly indicative of lexical intentions, but are rarely used in speech classification tasks. In this paper we show that principal differential analysis can be used to learn the behaviours of articulatory motions associated with certain high-level articulatory goals. This method accurately learns the parameters of second-order differential systems applied to data derived by electromagnetic articulography. On average, this approach is between 4.4% and 21.3% more accurate than an HMM and a neural network baseline. Michael Reimer, Frank Rudzicz |
INTERSPEECH | 2 |
| 2009 | Summarizing multiple spoken documents: finding evidence from untranscribed audio
Xiaodan Zhu 0001, Gerald Penn, Frank Rudzicz |
ACL/IJCNLP | 3 |
| 2009 | Applying discretized articulatory knowledge to dysarthric speechabstractThis paper applies two dynamic Bayes networks that include theoretical and measured kinematic features of the vocal tract, respectively, to the task of labeling phoneme sequences in unsegmented dysarthric speech. Speaker dependent and adaptive versions of these models are compared against two acoustic-only baselines, namely a hidden Markov model and a latent dynamic conditional random field. Both theoretical and kinematic models of the vocal tract perform admirably on speaker-dependent speech, and we show that the statistics of the latter are not necessarily transferable between speakers during adaptation. Frank Rudzicz |
ICASSP | 1 |
| 2009 | Phonological features in discriminative classification of dysarthric speechabstractIn an attempt to overcome problems associated with articulatory limitations and generative models, this work considers the use of phonological features in discriminative models for disabled speech. Specifically, we train feed-forward and recurrent neural networks, and radial basis and sequence-kernel support vector machines to abstractions of the vocal tract, and apply these models to phone recognition on dysarthric speech. The results show relative error reduction of between 1.5% and 10.9% with this approach against standard hidden Markov modeling, and increases in accuracy with speaker intelligibility across all classifiers. This work may be applied within components of assistive software for speakers with dysarthria. Frank Rudzicz |
ICASSP | 1 |
| 2007 | Comparing speaker-dependent and speaker-adaptive acoustic models for recognizing dysarthric speechabstractAcoustic modeling of dysarthric speech is complicated by its increased intra- and inter-speaker variability. The accuracy of speaker-dependent and speaker-adaptive models are compared for this task, with the latter prevailing across varying levels of speaker intelligibility. Frank Rudzicz |
ASSETS | 1 |
| 2006 | Clavius: Bi-Directional Parsing for Generic Multimodal Interaction
Frank Rudzicz |
ACL | 1 |
| 2004 | A framework for 3D visualisation and manipulation in an immersive space using an untethered bimanual gestural interfaceabstractImmersive Environments offer users the experience of being submerged in a virtual space, effectively transcending the boundary between the real and virtual world. We present a framework for visualization and manipulation of 3D virtual environments in which users need not resort to the awkward command vocabulary of traditional keyboard-and-mouse interaction. We have adapted the transparent toolglass paradigm as a gestural interface widget for a spatially immersive environment. To serve that purpose, we have implemented a bimanual gesture interpreter to recognize and translate a user's actions into commands for control of these widgets. In order to satisfy a primary design goal of keeping the user completely untethered, we use purely video-based tracking techniques. Yves Boussemart, François Rioux, Frank Rudzicz, Michael Wozniewski, Jeremy R. Cooperstock |
VRST | 3 |