EDBT 2026 Demo / reviewers in the wild / expert
Jake Sandakly
dblp:360/6696 · also Jacob Sandakly
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Audio and music processing · 85% Computational photography and imaging · 15% | |
| Artificial intelligence
2 papers |
Generative modeling · 72% 3D vision · 28% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › flow matching
conditional flow matching |
0.9 | 1 | 2025 | BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models · ICML 2025 |
Machine learning › Generative modeling
flow matching |
0.9 | 1 | 2025 | BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models · ICML 2025 |
Audio and music processing › spatial audio › binaural reproduction
binaural audio generation |
0.9 | 1 | 2025 | BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models · ICML 2025 |
Audio and music processing › spatial audio
binaural reproduction |
0.9 | 1 | 2025 | BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models · ICML 2025 |
Audio and music processing › acoustic rendering
novel-view acoustic synthesis |
0.9 | 1 | 2025 | SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding · CVPR 2025 |
Audio and music processing › spatial audio
spatial audio generation |
0.9 | 1 | 2025 | SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding · CVPR 2025 |
Audio and music processing
speech synthesis |
0.9 | 1 | 2025 | BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models · ICML 2025 |
Computer vision › 3D vision
human body modeling |
0.7 | 1 | 2023 | Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and Audio · NeurIPS 2023 |
Audio and music processing
spatial audio |
0.7 | 1 | 2023 | Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and Audio · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
streaming STFT/ISTFT · 1.7flow matching · 1.7causal u-net · 1.7spherical microphone array · 1.3multimodal dataset · 1.3panoramic RGB-D · 0.9cross-modal embedding · 0.9acoustic transfer function learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic BindingabstractWe introduce SoundVista, a method to generate the ambient sound of an arbitrary scene at novel viewpoints. Given a pre-acquired recording of the scene from sparsely distributed microphones, SoundVista can synthesize the sound of that scene from an unseen target viewpoint. The method learns the underlying acoustic transfer function that relates the signals acquired at the distributed microphones to the signal at the target viewpoint, using a limited number of known recordings. Unlike existing works, our method does not require constraints or prior knowledge of sound source details. Moreover, our method efficiently adapts to diverse room layouts, reference microphone configurations and unseen environments. To enable this, we introduce a visual-acoustic binding module that learns visual embeddings linked with local acoustic properties from panoramic RGB and depth data. We first leverage these embeddings to optimize the placement of reference microphones in any given scene. During synthesis, we leverage multiple embeddings extracted from reference locations to get adaptive weights for their contribution, conditioned on target viewpoint. We benchmark the task on both publicly available data and real-world settings. We demonstrate significant improvements over existing methods. Mingfei Chen, Israel D. Gebru, Ishwarya Ananthabhotla, Christian Richardt, Dejan Markovic, Jake Sandakly, Steven Krenn, Todd Keebler, Eli Shlizerman, Alexander Richard |
CVPR | 6 |
| 2025 | A2B: Neural Rendering of Ambisonic Recordings to BinauralabstractThis paper introduces a novel neural network model for rendering binaural audio directly from ambisonic recordings. We optimized the model end-to-end to learn a direct mapping between ambisonic and binaural signals. Our approach eliminates traditional processing steps that were required to mitigate artifacts due to spherical harmonic order truncation and spatial aliasing, as well as other complex filtering needed to compensate for near-field sound sources. To showcase the advantage of neural network-based rendering over traditional signal processing approaches, we introduce a new dataset that includes challenging near-field sound sources, including speech and background noises. We demonstrate that our model can produce binaural audio results that closely match the fidelity of ground truth binaural recordings. Our comprehensive validation shows that the proposed method outperforms existing methods on several error metrics as well as in subjective evaluations. Model code, demos and datasets are available on our project webpage. Israel D. Gebru, Todd Keebler, Jake Sandakly, Steven Krenn, Dejan Markovic, Julia Buffalini, Samuel Hassel, Alexander Richard |
ICASSP | 3 |
| 2025 | BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching ModelsabstractBinaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they struggle with rendering quality and streamable inference. Synthesizing high-quality binaural audio that is indistinguishable from real-world recordings requires precise modeling of binaural cues, room reverb, and ambient sounds. Additionally, real-world applications demand streaming inference. To address these challenges, we propose a flow matching based streaming binaural speech synthesis framework called BinauralFlow. We consider binaural rendering to be a generation problem rather than a regression problem and design a conditional flow matching model to render high-quality audio. Moreover, we design a causal U-Net architecture that estimates the current audio frame solely based on past information to tailor generative models for streaming inference. Finally, we introduce a continuous inference pipeline incorporating streaming STFT/ISTFT operations, a buffer bank, a midpoint solver, and an early skip schedule to improve rendering continuity and speed. Quantitative and qualitative evaluations demonstrate the superiority of our method over SOTA approaches. A perceptual study further reveals that our model is nearly indistinguishable from real-world recordings, with a 42% confusion rate. Susan Liang, Dejan Markovic, Israel D. Gebru, Steven Krenn, Todd Keebler, Jake Sandakly, Frank Yu, Samuel Hassel, Chenliang Xu, Alexander Richard |
ICML | 6 |
| 2023 | Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and AudioabstractWhile 3D human body modeling has received much attention in computer vision, modeling the acoustic equivalent, i.e. modeling 3D spatial audio produced by body motion and speech, has fallen short in the community. To close this gap, we present a model that can generate accurate 3D spatial audio for full human bodies. The system consumes, as input, audio signals from headset microphones and body pose, and produces, as output, a 3D sound field surrounding the transmitter's body, from which spatial audio can be rendered at any arbitrary position in the 3D space. We collect a first-of-its-kind multimodal dataset of human bodies, recorded with multiple cameras and a spherical array of 345 microphones. In an empirical evaluation, we demonstrate that our model can produce accurate body-induced sound fields when trained with a suitable loss. Dataset and code are available online. Xudong Xu, Dejan Markovic, Jake Sandakly, Todd Keebler, Steven Krenn, Alexander Richard |
NeurIPS | 3 |