Brendan Shillingford

dblp:126/9805 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Speech recognition and synthesis · 34% Trustworthy machine learning · 29% Deep learning architectures and training · 13%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
1 paper
Multimedia systems and quality of experience · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis › speech synthesis
text-to-speech
1.022022
More than Words: In-the-Wild Visually-Driven Prosody for Text-to-Speech · CVPR 2022
Sample Efficient Adaptive Text-to-Speech · ICLR (Poster) 2019
Bioinformatics and computational biology › computational neuroscience
neural decoding
0.912025
The Brain's Bitter Lesson: Scaling Speech Decoding With Self-Supervised Learning · ICML 2025
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
visual text-to-speech
0.612022
More than Words: In-the-Wild Visually-Driven Prosody for Text-to-Speech · CVPR 2022
Multimedia systems and quality of experience › multimedia synchronization
audio-visual synchronization
0.612022
More than Words: In-the-Wild Visually-Driven Prosody for Text-to-Speech · CVPR 2022
Machine learning › Trustworthy machine learning › language model interpretability
explanation consistency
0.412020
Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations · ACL 2020
Machine learning › Trustworthy machine learning
interpretability
0.412020
Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations · ACL 2020
Machine learning › Trustworthy machine learning › interpretability
natural language explanation
0.412020
Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations · ACL 2020
Natural language and speech › Language models and text generation
text generation
0.412020
Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations · ACL 2020
Natural language and speech › Machine translation
transliteration
0.312018
Recovering Missing Characters in Old Hawaiian Writing · EMNLP 2018
Machine learning › Deep learning architectures and training › recurrent neural network
gated recurrent network
0.312017
Cortical microcircuits as gated-recurrent neural networks · NIPS 2017
Machine learning › Deep learning architectures and training
recurrent neural network
0.312017
Cortical microcircuits as gated-recurrent neural networks · NIPS 2017
Machine learning › Transfer learning and domain adaptation › domain generalization
cross-subject generalization
0.312025
The Brain's Bitter Lesson: Scaling Speech Decoding With Self-Supervised Learning · ICML 2025
Natural language and speech › Language models and text generation › neural language model
recurrent neural network language model
0.112018
Recovering Missing Characters in Old Hawaiian Writing · EMNLP 2018

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 1.7magnetoencephalography · 1.7video-conditioned speech synthesis · 1.1sequence-to-sequence attacks · 0.4adversarial generation · 0.4meta-learning · 0.4recurrent neural network language model · 0.3finite state transducer · 0.3gated memory networks · 0.3LSTM · 0.3
YearPublicationVenuePosition
2025 The Brain's Bitter Lesson: Scaling Speech Decoding With Self-Supervised Learning
abstract
The past few years have seen remarkable progress in the decoding of speech from brain activity, primarily driven by large single-subject datasets. However, due to individual variation, such as anatomy, and differences in task design and scanning hardware, leveraging data across subjects and datasets remains challenging. In turn, the field has not benefited from the growing number of open neural data repositories to exploit large-scale deep learning. To address this, we develop neuroscience-informed self-supervised objectives, together with an architecture, for learning from heterogeneous brain recordings. Scaling to nearly **400 hours** of MEG data and **900 subjects**, our approach shows generalisation across participants, datasets, tasks, and even to *novel* subjects. It achieves **improvements of 15-27%** over state-of-the-art models and **matches *surgical* decoding performance with *non-invasive* data**. These advances unlock the potential for scaling speech decoding models beyond the current frontier.
Dulhan Jayalath, Gilad Landau, Brendan Shillingford, Mark W. Woolrich, Oiwi Parker Jones
ICML3
2022 More than Words: In-the-Wild Visually-Driven Prosody for Text-to-Speech
abstract
In this paper we present VDTTS, a Visually-Driven Text-to-Speech model. Motivated by dubbing, VDTTS takes ad-vantage of video frames as an additional input alongside text, and generates speech that matches the video signal. We demonstrate how this allows VDTTS to, unlike plain TTS models, generate speech that not only has prosodic variations like natural pauses and pitch, but is also synchronized to the input video. Experimentally, we show our model produces well-synchronized outputs, approaching the video-speech synchronization quality of the ground-truth, on several challenging benchmarks including “in-the-wild” content from VoxCeleb2. Supplementary demo videos demonstrating video-speech synchronization, robustness to speaker ID swapping, and prosody, presented at the project page.11Project page: http://google-research.github.io/lingvo-lab/vdtts
Michael Hassid, Michelle Tadmor Ramanovich, Brendan Shillingford, Miaosen Wang, Ye Jia, Tal Remez
CVPR3
2020 Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations
abstract
To increase trust in artificial intelligence systems, a promising research direction consists of designing neural models capable of generating natural language explanations for their predictions.In this work, we show that such models are nonetheless prone to generating mutually inconsistent explanations, such as "Because there is a dog in the image."and "Because there is no dog in the [same] image.",exposing flaws in either the decision-making process of the model or in the generation of the explanations.We introduce a simple yet effective adversarial framework for sanity checking models against the generation of inconsistent natural language explanations.Moreover, as part of the framework, we address the problem of adversarial attacks with full target sequences, a scenario that was not previously addressed in sequence-to-sequence attacks.Finally, we apply our framework on a state-of-the-art neural natural language inference model that provides natural language explanations for its predictions.Our framework shows that this model is capable of generating a significant number of inconsistent explanations.PREMISE: A guy in a red jacket is snowboarding in midair.
Oana-Maria Camburu, Brendan Shillingford, Pasquale Minervini, Thomas Lukasiewicz, Phil Blunsom
ACL2
2019 Recurrent Neural Network Transducer for Audio-Visual Speech Recognition
abstract
This work presents a large-scale audio-visual speech recognition system based on a recurrent neural network transducer (RNN-T) architecture. To support the development of such a system, we built a large audio-visual (A/V) dataset of segmented utterances extracted from YouTube public videos, leading to 31k hours of audio-visual training content. The performance of an audio-only, visual-only, and audio-visual system are compared on two large-vocabulary test sets: a set of utterance segments from public YouTube videos called YTDEV18 and the publicly available LRS3-TED set. To highlight the contribution of the visual modality, we also evaluated the performance of our system on the YTDEV18 set artificially corrupted with background noise and overlapping speech. To the best of our knowledge, our system significantly improves the state-of-the-art on the LRS3-TED set.
Takaki Makino, Hank Liao, Yannis M. Assael, Brendan Shillingford, Basilio Garcia, Otavio Braga, Olivier Siohan
ASRU4
2019 Sample Efficient Adaptive Text-to-Speech
Yutian Chen 0001, Yannis M. Assael, Brendan Shillingford, David Budden, Scott E. Reed, Heiga Zen, Luis C. Cobo, Andrew Trask, Ben Laurie, Caglar Gulcehre, Aäron van den Oord, Oriol Vinyals, Nando de Freitas
ICLR (Poster)3
2019 Large-Scale Visual Speech Recognition
abstract
This work presents a scalable solution to open-vocabulary visual speech recognition. To achieve this, we constructed the largest existing visual speech recognition dataset, consisting of pairs of text and video clips of faces speaking (3,886 hours of video). In tandem, we designed and trained an integrated lipreading system, consisting of a video processing pipeline that maps raw video to stable videos of lips and sequences of phonemes, a scalable deep neural network that maps the lip videos to sequences of phoneme distributions, and a production-level speech decoder that outputs sequences of words. The proposed system achieves a word error rate (WER) of 40.9% as measured on a held-out set. In comparison, professional lipreaders achieve either 86.4% or 92.9% WER on the same dataset when having access to additional types of contextual information. Our approach significantly improves on other lipreading approaches, including variants of LipNet and of Watch, Attend, and Spell (WAS), which are only capable of 89.8% and 76.8% WER respectively.
Brendan Shillingford, Yannis M. Assael, Matt Hoffman 0001, Thomas Paine, Cían Hughes, Utsav Prabhu, Hank Liao, Hasim Sak, Kanishka Rao, Lorrayne Bennett, Marie Mulville, Misha Denil, Ben Coppin, Ben Laurie, Andrew W. Senior, Nando de Freitas
INTERSPEECH1
2018 Recovering Missing Characters in Old Hawaiian Writing
abstract
In contrast to the older writing system of the 19th century, modern Hawaiian orthography employs characters for long vowels and glottal stops.These extra characters account for about one-third of the phonemes in Hawaiian, so including them makes a big difference to reading comprehension and pronunciation.However, transliterating between older and newer texts is a laborious task when performed manually.We introduce two related methods to help solve this transliteration problem automatically.One approach is implemented, endto-end, using finite state transducers (FSTs).The other is a hybrid deep learning approach, which approximately composes an FST with a recurrent neural network language model.
Brendan Shillingford, Oiwi Parker Jones
EMNLP1
2017 Cortical microcircuits as gated-recurrent neural networks
abstract
Cortical circuits exhibit intricate recurrent architectures that are remarkably similar across different brain areas. Such stereotyped structure suggests the existence of common computational principles. However, such principles have remained largely elusive. Inspired by gated-memory networks, namely long short-term memory networks (LSTMs), we introduce a recurrent neural network in which information is gated through inhibitory cells that are subtractive (subLSTM). We propose a natural mapping of subLSTMs onto known canonical excitatory-inhibitory cortical microcircuits. Our empirical evaluation across sequential image classification and language modelling tasks shows that subLSTM units can achieve similar performance to LSTM units. These results suggest that cortical circuits can be optimised to solve complex contextual problems and proposes a novel view on their computational function. Overall our work provides a step towards unifying recurrent networks as used in machine learning with their biological counterparts.
Rui Ponte Costa, Yannis M. Assael, Brendan Shillingford, Nando de Freitas, Tim P. Vogels
NIPS3
2013 "Dictionary Wars" (abstract only): an inverted, leaderboard-driven project for learning dictionary data structures
abstract
We present a highly reusable "inverted" project in which students learn asymptotic and practical behaviour of dictionary data structures--linked-lists, arrays, balanced trees, and hash tables--in an atmosphere of mild competition. Much like David Levine's Nifty Assignment "Sort Detective", rather than implementing the dictionaries, students' programs generate input to our (unlabeled) implementations, and students use timing data to label the implementations. Much like Bryant and O'Halloran's computer architecture labs, students also compete to "convince" a web-based, automated system that their input generators distinguish the dictionaries based on trend-line behaviour. Initial assessment results suggest the project makes substantially improves students' understanding of practical performance of various dictionary data structures, particularly hash tables. UBC has used the project in three terms, and we plan to use it at UBC and U Toronto in coming terms.
Kuba Karpierz, Joel Kitching, Brendan Shillingford, Elizabeth Ann Patitsas, Steven A. Wolfman
SIGCSE3