VLDB 2026 Research / reviewers in the wild / expert
Brendan Shillingford
dblp:126/9805
· DBLP profile ↗
9ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Speech recognition and synthesis · 34% Trustworthy machine learning · 29% Deep learning architectures and training · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% | |
| Computer graphics and multimedia
1 paper |
Multimedia systems and quality of experience · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis › speech synthesis
text-to-speech |
1.0 | 2 | 2022 | More than Words: In-the-Wild Visually-Driven Prosody for Text-to-Speech · CVPR 2022 Sample Efficient Adaptive Text-to-Speech · ICLR (Poster) 2019 |
Bioinformatics and computational biology › computational neuroscience
neural decoding |
0.9 | 1 | 2025 | The Brain's Bitter Lesson: Scaling Speech Decoding With Self-Supervised Learning · ICML 2025 |
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
visual text-to-speech |
0.6 | 1 | 2022 | More than Words: In-the-Wild Visually-Driven Prosody for Text-to-Speech · CVPR 2022 |
Multimedia systems and quality of experience › multimedia synchronization
audio-visual synchronization |
0.6 | 1 | 2022 | More than Words: In-the-Wild Visually-Driven Prosody for Text-to-Speech · CVPR 2022 |
Machine learning › Trustworthy machine learning › language model interpretability
explanation consistency |
0.4 | 1 | 2020 | Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations · ACL 2020 |
Machine learning › Trustworthy machine learning
interpretability |
0.4 | 1 | 2020 | Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations · ACL 2020 |
Machine learning › Trustworthy machine learning › interpretability
natural language explanation |
0.4 | 1 | 2020 | Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations · ACL 2020 |
Natural language and speech › Language models and text generation
text generation |
0.4 | 1 | 2020 | Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations · ACL 2020 |
Natural language and speech › Machine translation
transliteration |
0.3 | 1 | 2018 | Recovering Missing Characters in Old Hawaiian Writing · EMNLP 2018 |
Machine learning › Deep learning architectures and training › recurrent neural network
gated recurrent network |
0.3 | 1 | 2017 | Cortical microcircuits as gated-recurrent neural networks · NIPS 2017 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2017 | Cortical microcircuits as gated-recurrent neural networks · NIPS 2017 |
Machine learning › Transfer learning and domain adaptation › domain generalization
cross-subject generalization |
0.3 | 1 | 2025 | The Brain's Bitter Lesson: Scaling Speech Decoding With Self-Supervised Learning · ICML 2025 |
Natural language and speech › Language models and text generation › neural language model
recurrent neural network language model |
0.1 | 1 | 2018 | Recovering Missing Characters in Old Hawaiian Writing · EMNLP 2018 |
Methods — techniques the papers use, named apart from their topics
self-supervised learning · 1.7magnetoencephalography · 1.7video-conditioned speech synthesis · 1.1sequence-to-sequence attacks · 0.4adversarial generation · 0.4meta-learning · 0.4recurrent neural network language model · 0.3finite state transducer · 0.3gated memory networks · 0.3LSTM · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Brain's Bitter Lesson: Scaling Speech Decoding With Self-Supervised LearningabstractThe past few years have seen remarkable progress in the decoding of speech from brain activity, primarily driven by large single-subject datasets. However, due to individual variation, such as anatomy, and differences in task design and scanning hardware, leveraging data across subjects and datasets remains challenging. In turn, the field has not benefited from the growing number of open neural data repositories to exploit large-scale deep learning. To address this, we develop neuroscience-informed self-supervised objectives, together with an architecture, for learning from heterogeneous brain recordings. Scaling to nearly **400 hours** of MEG data and **900 subjects**, our approach shows generalisation across participants, datasets, tasks, and even to *novel* subjects. It achieves **improvements of 15-27%** over state-of-the-art models and **matches *surgical* decoding performance with *non-invasive* data**. These advances unlock the potential for scaling speech decoding models beyond the current frontier. Dulhan Jayalath, Gilad Landau, Brendan Shillingford, Mark W. Woolrich, Oiwi Parker Jones |
ICML | 3 |
| 2022 | More than Words: In-the-Wild Visually-Driven Prosody for Text-to-SpeechabstractIn this paper we present VDTTS, a Visually-Driven Text-to-Speech model. Motivated by dubbing, VDTTS takes ad-vantage of video frames as an additional input alongside text, and generates speech that matches the video signal. We demonstrate how this allows VDTTS to, unlike plain TTS models, generate speech that not only has prosodic variations like natural pauses and pitch, but is also synchronized to the input video. Experimentally, we show our model produces well-synchronized outputs, approaching the video-speech synchronization quality of the ground-truth, on several challenging benchmarks including “in-the-wild” content from VoxCeleb2. Supplementary demo videos demonstrating video-speech synchronization, robustness to speaker ID swapping, and prosody, presented at the project page.11Project page: http://google-research.github.io/lingvo-lab/vdtts Michael Hassid, Michelle Tadmor Ramanovich, Brendan Shillingford, Miaosen Wang, Ye Jia, Tal Remez |
CVPR | 3 |
| 2020 | Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language ExplanationsabstractTo increase trust in artificial intelligence systems, a promising research direction consists of designing neural models capable of generating natural language explanations for their predictions.In this work, we show that such models are nonetheless prone to generating mutually inconsistent explanations, such as "Because there is a dog in the image."and "Because there is no dog in the [same] image.",exposing flaws in either the decision-making process of the model or in the generation of the explanations.We introduce a simple yet effective adversarial framework for sanity checking models against the generation of inconsistent natural language explanations.Moreover, as part of the framework, we address the problem of adversarial attacks with full target sequences, a scenario that was not previously addressed in sequence-to-sequence attacks.Finally, we apply our framework on a state-of-the-art neural natural language inference model that provides natural language explanations for its predictions.Our framework shows that this model is capable of generating a significant number of inconsistent explanations.PREMISE: A guy in a red jacket is snowboarding in midair. Oana-Maria Camburu, Brendan Shillingford, Pasquale Minervini, Thomas Lukasiewicz, Phil Blunsom |
ACL | 2 |
| 2019 | Recurrent Neural Network Transducer for Audio-Visual Speech RecognitionabstractThis work presents a large-scale audio-visual speech recognition system based on a recurrent neural network transducer (RNN-T) architecture. To support the development of such a system, we built a large audio-visual (A/V) dataset of segmented utterances extracted from YouTube public videos, leading to 31k hours of audio-visual training content. The performance of an audio-only, visual-only, and audio-visual system are compared on two large-vocabulary test sets: a set of utterance segments from public YouTube videos called YTDEV18 and the publicly available LRS3-TED set. To highlight the contribution of the visual modality, we also evaluated the performance of our system on the YTDEV18 set artificially corrupted with background noise and overlapping speech. To the best of our knowledge, our system significantly improves the state-of-the-art on the LRS3-TED set. Takaki Makino, Hank Liao, Yannis M. Assael, Brendan Shillingford, Basilio Garcia, Otavio Braga, Olivier Siohan |
ASRU | 4 |
| 2019 | Sample Efficient Adaptive Text-to-Speech
Yutian Chen 0001, Yannis M. Assael, Brendan Shillingford, David Budden, Scott E. Reed, Heiga Zen, Luis C. Cobo, Andrew Trask, Ben Laurie, Caglar Gulcehre, Aäron van den Oord, Oriol Vinyals, Nando de Freitas |
ICLR (Poster) | 3 |
| 2019 | Large-Scale Visual Speech RecognitionabstractThis work presents a scalable solution to open-vocabulary visual speech recognition. To achieve this, we constructed the largest existing visual speech recognition dataset, consisting of pairs of text and video clips of faces speaking (3,886 hours of video). In tandem, we designed and trained an integrated lipreading system, consisting of a video processing pipeline that maps raw video to stable videos of lips and sequences of phonemes, a scalable deep neural network that maps the lip videos to sequences of phoneme distributions, and a production-level speech decoder that outputs sequences of words. The proposed system achieves a word error rate (WER) of 40.9% as measured on a held-out set. In comparison, professional lipreaders achieve either 86.4% or 92.9% WER on the same dataset when having access to additional types of contextual information. Our approach significantly improves on other lipreading approaches, including variants of LipNet and of Watch, Attend, and Spell (WAS), which are only capable of 89.8% and 76.8% WER respectively. Brendan Shillingford, Yannis M. Assael, Matt Hoffman 0001, Thomas Paine, Cían Hughes, Utsav Prabhu, Hank Liao, Hasim Sak, Kanishka Rao, Lorrayne Bennett, Marie Mulville, Misha Denil, Ben Coppin, Ben Laurie, Andrew W. Senior, Nando de Freitas |
INTERSPEECH | 1 |
| 2018 | Recovering Missing Characters in Old Hawaiian WritingabstractIn contrast to the older writing system of the 19th century, modern Hawaiian orthography employs characters for long vowels and glottal stops.These extra characters account for about one-third of the phonemes in Hawaiian, so including them makes a big difference to reading comprehension and pronunciation.However, transliterating between older and newer texts is a laborious task when performed manually.We introduce two related methods to help solve this transliteration problem automatically.One approach is implemented, endto-end, using finite state transducers (FSTs).The other is a hybrid deep learning approach, which approximately composes an FST with a recurrent neural network language model. Brendan Shillingford, Oiwi Parker Jones |
EMNLP | 1 |
| 2017 | Cortical microcircuits as gated-recurrent neural networksabstractCortical circuits exhibit intricate recurrent architectures that are remarkably similar across different brain areas. Such stereotyped structure suggests the existence of common computational principles. However, such principles have remained largely elusive. Inspired by gated-memory networks, namely long short-term memory networks (LSTMs), we introduce a recurrent neural network in which information is gated through inhibitory cells that are subtractive (subLSTM). We propose a natural mapping of subLSTMs onto known canonical excitatory-inhibitory cortical microcircuits. Our empirical evaluation across sequential image classification and language modelling tasks shows that subLSTM units can achieve similar performance to LSTM units. These results suggest that cortical circuits can be optimised to solve complex contextual problems and proposes a novel view on their computational function. Overall our work provides a step towards unifying recurrent networks as used in machine learning with their biological counterparts. Rui Ponte Costa, Yannis M. Assael, Brendan Shillingford, Nando de Freitas, Tim P. Vogels |
NIPS | 3 |
| 2013 | "Dictionary Wars" (abstract only): an inverted, leaderboard-driven project for learning dictionary data structuresabstractWe present a highly reusable "inverted" project in which students learn asymptotic and practical behaviour of dictionary data structures--linked-lists, arrays, balanced trees, and hash tables--in an atmosphere of mild competition. Much like David Levine's Nifty Assignment "Sort Detective", rather than implementing the dictionaries, students' programs generate input to our (unlabeled) implementations, and students use timing data to label the implementations. Much like Bryant and O'Halloran's computer architecture labs, students also compete to "convince" a web-based, automated system that their input generators distinguish the dictionaries based on trend-line behaviour. Initial assessment results suggest the project makes substantially improves students' understanding of practical performance of various dictionary data structures, particularly hash tables. UBC has used the project in three terms, and we plan to use it at UBC and U Toronto in coming terms. Kuba Karpierz, Joel Kitching, Brendan Shillingford, Elizabeth Ann Patitsas, Steven A. Wolfman |
SIGCSE | 3 |