VLDB 2026 Research / reviewers in the wild / expert
Layne Berry
dblp:330/4076
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Processor architecture and microarchitecture · 87% Energy-efficient computing · 13% | |
| Artificial intelligence
1 paper |
Vision and language · 62% Image recognition and object detection · 19% Knowledge representation and reasoning · 19% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › front-end
micro-op cache |
0.6 | 1 | 2022 | Speculative Code Compaction: Eliminating Dead Code via Speculative Microcode Transformations · MICRO 2022 |
Processor architecture and microarchitecture
microprogramming |
0.6 | 1 | 2022 | Speculative Code Compaction: Eliminating Dead Code via Speculative Microcode Transformations · MICRO 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.2 | 1 | 2022 | Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality · EMNLP 2022 |
Computer vision › Image recognition and object detection
image retrieval |
0.2 | 1 | 2022 | Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality · EMNLP 2022 |
Methods — techniques the papers use, named apart from their topics
speculative microcode transformation · 0.6probing tasks · 0.6manual annotation · 0.6data dependence prediction · 0.6data augmentation · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Temporally Streaming Audio-Visual Synchronization for Real-World VideosabstractWe introduce RealSync, a novel dataset designed to significantly enhance the training and evaluation of models for audio-visual synchronization (AV Sync) tasks. Sourced from high-quality YouTube channels, RealSync covers a wide range of content domains, providing an improved scale, diversity, and alignment with broadcast content compared to existing datasets. It features extended-length video samples, catering to the critical need for more comprehensive, real-world training and evaluation materials. Alongside this dataset, we present StreamSync, a model tailored for real-world AV Sync applications. StreamSync is designed to be backbone agnostic and incorporates a streaming mechanism that processes consecutive video segments dynamically, iteratively refining synchronization predictions. This innovative approach enables StreamSync to outperform existing models, offering superior synchronization accuracy with minimal computational cost per iteration. Together, our dataset and the StreamSync model establish a new benchmark for AVSync research, promising to drive the development of more robust and practical AVSync methods. https://github.com/jvoas655/StreamSync Jordan Voas, Wei-Cheng Tseng, Layne Berry, Xixi Hu 0001, Puyuan Peng, James Stuedemann, David F. Harwath |
WACV | 3 |
| 2024 | AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation ModelsabstractAudio-visual representation learning aims to develop systems with human-like perception by utilizing correlation between auditory and visual information. However, current models often focus on a limited set of tasks, and generalization abilities of learned representations are unclear. To this end, we propose the AV-SUPERB benchmark that enables general-purpose evaluation of unimodal audio/visual and bimodal fusion representations on 7 datasets covering 5 audio-visual tasks in speech and audio processing. We evaluate 5 recent self-supervised models and show that none of these models generalize to all tasks, emphasizing the need for future study on improving universal model performance. In addition, we show that representations may be improved with intermediate-task fine-tuning and audio event classification with AudioSet serves as a strong intermediate task. We release our benchmark with evaluation code1and a model submission platform2to encourage further research in audio-visual learning. Yuan Tseng, Layne Berry, I-Hsiang Chiu, Hsuan-Hao Lin, Max Liu, Puyuan Peng, Yi-Jen Shih, Hung-Yu Wang, Po-Yao Huang 0001, Chun-Mao Lai, Shang-Wen Li 0001, David F. Harwath, Yu Tsao 0001, Abdel-rahman Mohamed, Chi-Luen Feng, Hung-yi Lee |
ICASSP | 2 |
| 2023 | M-SpeechCLIP: Leveraging Large-Scale, Pre-Trained Models for Multilingual Speech to Image RetrievalabstractThis work investigates the use of large-scale, English-only pre-trained models (CLIP and HuBERT) for multilingual image-speech retrieval. For non-English image-speech retrieval, we outperform the current state-of-the-art performance by a wide margin both when training separate models for each language, and with a single model which processes speech in all three languages. We identify key differences in model behavior and performance between English and non-English settings, attributable to the English-only pre-training of CLIP and HuBERT, and investigate how fine-tuning the pre-trained models impacts these differences. Finally, we show that our models can be used for mono- and cross-lingual speech-text retrieval and cross-lingual speech-speech retrieval, despite never having seen any parallel speech-text or speech-speech data during training. Layne Berry, Yi-Jen Shih, Hsuan-Fu Wang, Heng-Jui Chang, Hung-yi Lee, David F. Harwath |
ICASSP | 1 |
| 2022 | Why is Winoground Hard? Investigating Failures in Visuolinguistic CompositionalityabstractRecent visuolinguistic pre-trained models show promising progress on various end tasks such as image retrieval and video captioning.Yet, they fail miserably on the recently proposed Winoground dataset (Thrush et al., 2022), which challenges models to match paired images and English captions, with items constructed to overlap lexically but differ in meaning (e.g., "there is a mug in some grass" vs. "there is some grass in a mug").By annotating the dataset using new fine-grained tags, we show that solving the Winoground task requires not just compositional language understanding, but a host of other abilities like commonsense reasoning or locating small, out-of-focus objects in low-resolution images.In this paper, we identify the dataset's main challenges through a suite of experiments on related tasks (probing task, image retrieval task), data augmentation, and manual inspection of the dataset.Our analysis suggests that a main challenge in visuolinguistic models may lie in fusing visual and textual representations, rather than in compositional language understanding.We release our annotation and code at https://github. com/ajd12342/why-winoground-hard. Anuj Diwan, Layne Berry, Eunsol Choi, David F. Harwath, Kyle Mahowald |
EMNLP | 2 |
| 2022 | Speculative Code Compaction: Eliminating Dead Code via Speculative Microcode TransformationsabstractThe computing landscape has been increasingly characterized by processor architectures with increasing core counts, while a majority of the software applications remain inherently sequential. Although state-of-the-art compilers feature sophisticated optimizations, a significant chunk of wasteful computation persists due to the presence of data-dependent operations and irregular control-flow patterns that are unpredictable at compile-time. This work presents speculative code compaction (SCC), a novel microarchitectural technique that significantly enhances the capabilities of the microcode engine to aggressively and speculatively eliminate dead code from hot code regions resident in the micro-op cache, and further generate a compact stream of micro-ops, based on dynamically predicted machine code invariants. SCC also extends existing micro-op cache designs to co-host multiple versions of unoptimized and speculatively optimized micro-op sequences, providing the fetch engine with significant flexibility to dynamically choose from and stream the appropriate set of micro-ops, as and when deemed profitable.SCC is a minimally-invasive technique that can be implemented at the processor front-end using a simple ALU and a register context table, and is yet able to substantially accelerate the performance of already compile-time optimized and machine-tuned code by an average of 6% (and as much as 30%), with an average of 12% (and as much as 24%) savings in energy consumption, while eliminating the need for profiling and offering increased adaptability to changing datasets and workload patterns. Logan Moody, Abdolrasoul Sharifi, Layne Berry, Joey Rudek, Jayesh Gaur, Jeff Parkhurst, Sreenivas Subramoney, Kevin Skadron, Ashish Venkat |
MICRO | 4 |
| 2022 | SpeechCLIP: Integrating Speech with Pre-Trained Vision and Language ModelabstractData-driven speech processing models usually perform well with a large amount of text supervision, but collecting transcribed speech data is costly. Therefore, we propose Speech-CLIP, a novel framework bridging speech and text through images to enhance speech models without transcriptions. We leverage state-of-the-art pre-trained HuBERT and CLIP, aligning them via paired images and spoken captions with minimal fine-tuning. SpeechCLIP outperforms prior state-of-the-art on image-speech retrieval and performs zero-shot speech-text retrieval without direct supervision from transcriptions. Moreover, SpeechCLIP can directly retrieve semantically related keywords from speech. Yi-Jen Shih, Hsuan-Fu Wang, Heng-Jui Chang, Layne Berry, Hung-yi Lee, David F. Harwath |
SLT | 4 |