VLDB 2026 Research / reviewers in the wild / expert
Carl Schissler
dblp:118/6128
· DBLP profile ↗
16ranked-venue papers
7as first author
3since 2021 · last 2022
0000-0002-8355-8856ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
13 papers |
Audio and music processing · 73% Rendering · 22% Virtual and augmented reality · 2% | |
| Artificial intelligence
4 papers |
Robot navigation and mapping · 46% 3D vision · 30% Representation and self-supervised learning · 23% |
Topics — the 19 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing
acoustic simulation |
2.5 | 8 | 2022 | SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning · NeurIPS 2022 Fast diffraction pathfinding for dynamic sound propagation · ACM Trans. Graph. 2021 Diffraction Kernels for Interactive Sound Propagation in Dynamic Environments · IEEE Trans. Vis. Comput. Graph. 2018 |
Audio and music processing › acoustic simulation
sound propagation |
2.1 | 7 | 2021 | Fast diffraction pathfinding for dynamic sound propagation · ACM Trans. Graph. 2021 Acoustic Classification and Optimization for Multi-Modal Rendering of Real-World Scenes · IEEE Trans. Vis. Comput. Graph. 2018 Diffraction Kernels for Interactive Sound Propagation in Dynamic Environments · IEEE Trans. Vis. Comput. Graph. 2018 |
Robotics › Robot navigation and mapping › mobile robot navigation › sensor-based navigation
audio-visual navigation |
1.0 | 2 | 2022 | SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning · NeurIPS 2022 SoundSpaces: Audio-Visual Navigation in 3D Environments · ECCV (6) 2020 |
Audio and music processing › acoustic simulation
geometric acoustics |
0.9 | 3 | 2021 | Fast diffraction pathfinding for dynamic sound propagation · ACM Trans. Graph. 2021 Interactive Sound Propagation and Rendering for Large Multi-Source Scenes · ACM Trans. Graph. 2017 Guided Multiview Ray Tracing for Fast Auralization · IEEE Trans. Vis. Comput. Graph. 2012 |
Audio and music processing
spatial audio |
0.9 | 3 | 2018 | Acoustic Classification and Optimization for Multi-Modal Rendering of Real-World Scenes · IEEE Trans. Vis. Comput. Graph. 2018 Efficient construction of the spatial room impulse response · VR 2017 Efficient HRTF-based Spatial Audio for Area and Volumetric Sources · IEEE Trans. Vis. Comput. Graph. 2016 |
Rendering › physically based rendering › wave optics
diffraction |
0.8 | 2 | 2021 | Fast diffraction pathfinding for dynamic sound propagation · ACM Trans. Graph. 2021 Diffraction Kernels for Interactive Sound Propagation in Dynamic Environments · IEEE Trans. Vis. Comput. Graph. 2018 |
Rendering
ray tracing |
0.8 | 5 | 2021 | Interactive sound propagation with bidirectional path tracing · ACM Trans. Graph. 2016 High-order diffraction and diffuse reflections for interactive sound propagation in large environments · ACM Trans. Graph. 2014 Fast diffraction pathfinding for dynamic sound propagation · ACM Trans. Graph. 2021 |
Robotics › Robot navigation and mapping
embodied navigation |
0.6 | 1 | 2022 | SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning · NeurIPS 2022 |
Computer vision › 3D vision
3d scene understanding |
0.5 | 1 | 2021 | Audio-Visual Floorplan Reconstruction · ICCV 2021 |
Computer vision › 3D vision › 3d scene reconstruction
floorplan reconstruction |
0.5 | 1 | 2021 | Audio-Visual Floorplan Reconstruction · ICCV 2021 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.4 | 1 | 2020 | VisualEchoes: Spatial Image Representation Learning Through Echolocation · ECCV (9) 2020 |
Machine learning › Representation and self-supervised learning
spatial representation learning |
0.4 | 1 | 2020 | VisualEchoes: Spatial Image Representation Learning Through Echolocation · ECCV (9) 2020 |
Rendering › ray tracing › path tracing
bidirectional path tracing |
0.4 | 2 | 2021 | Interactive sound propagation with bidirectional path tracing · ACM Trans. Graph. 2016 Fast diffraction pathfinding for dynamic sound propagation · ACM Trans. Graph. 2021 |
Audio and music processing
acoustic rendering |
0.4 | 2 | 2018 | Interactive Sound Propagation and Rendering for Large Multi-Source Scenes · ACM Trans. Graph. 2017 Acoustic Classification and Optimization for Multi-Modal Rendering of Real-World Scenes · IEEE Trans. Vis. Comput. Graph. 2018 |
Audio and music processing › spatial audio
head-related transfer function |
0.2 | 1 | 2016 | Efficient HRTF-based Spatial Audio for Area and Volumetric Sources · IEEE Trans. Vis. Comput. Graph. 2016 |
Rendering › illumination
diffuse reflection |
0.2 | 1 | 2014 | High-order diffraction and diffuse reflections for interactive sound propagation in large environments · ACM Trans. Graph. 2014 |
Multimedia analysis and retrieval
audio-visual learning |
0.2 | 1 | 2022 | SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning · NeurIPS 2022 |
Virtual and augmented reality
immersive audio |
0.2 | 2 | 2017 | Efficient construction of the spatial room impulse response · VR 2017 Efficient HRTF-based Spatial Audio for Area and Volumetric Sources · IEEE Trans. Vis. Comput. Graph. 2016 |
Computer animation and physical simulation
rigid body simulation |
0.1 | 1 | 2016 | SynCoPation: Interactive Synthesis-Coupled Sound Propagation · IEEE Trans. Vis. Comput. Graph. 2016 |
Methods — techniques the papers use, named apart from their topics
sim2real · 1.1geometry-based acoustic rendering · 1.1multimodal encoder-decoder · 1.0contrastive learning · 0.9uniform theory of diffraction · 0.5silhouette edge detection · 0.5a* pathfinding · 0.5ray tracing · 0.3diffraction kernel · 0.3acoustic impulse response measurement · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic LearningabstractWe introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from arbitrary microphone locations. Together with existing 3D visual assets, it supports an array of audio-visual research tasks, such as audio-visual navigation, mapping, source localization and separation, and acoustic matching. Compared to existing resources, SoundSpaces 2.0 has the advantages of allowing continuous spatial sampling, generalization to novel environments, and configurable microphone and material properties. To our knowledge, this is the first geometry-based acoustic simulation that offers high fidelity and realism while also being fast enough to use for embodied learning. We showcase the simulator's properties and benchmark its performance against real-world audio measurements. In addition, we demonstrate two downstream tasks---embodied navigation and far-field automatic speech recognition---and highlight sim2real performance for the latter. SoundSpaces 2.0 is publicly available to facilitate wider research for perceptual systems that can both see and hear. Changan Chen, Carl Schissler, Sanchit Garg, Philip Kobernik, Alexander Clegg, Paul Calamia, Dhruv Batra, Philip W. Robinson, Kristen Grauman |
NeurIPS | 2 |
| 2021 | Audio-Visual Floorplan ReconstructionabstractGiven only a few glimpses of an environment, how much can we infer about its entire floorplan? Existing methods can map only what is visible or immediately apparent from context, and thus require substantial movements through a space to fully map it. We explore how both audio and visual sensing together can provide rapid floorplan reconstruction from limited viewpoints. Audio not only helps sense geometry outside the camera’s field of view, but it also reveals the existence of distant freespace (e.g., a dog barking in another room) and suggests the presence of rooms not visible to the camera (e.g., a dishwasher humming in what must be the kitchen to the left). We introduce AV-Map, a novel multi-modal encoder-decoder framework that reasons jointly about audio and vision to reconstruct a floorplan from a short input video sequence. We train our model to predict both the interior structure of the environment and the associated rooms’ semantic labels. Our results on 85 large real-world environments show the impact: with just a few glimpses spanning 26% of an area, we can estimate the whole area with 66% accuracy—substantially better than the state of the art approach for extrapolating visual maps. Senthil Purushwalkam, Sebastià Vicenc Amengual Garí, Vamsi K. Ithapu, Carl Schissler, Philip W. Robinson, Abhinav Gupta 0001, Kristen Grauman |
ICCV | 4 |
| 2021 | Fast diffraction pathfinding for dynamic sound propagationabstractIn the context of geometric acoustic simulation, one of the more perceptually important yet difficult to simulate acoustic effects is diffraction, a phenomenon that allows sound to propagate around obstructions and corners. A significant bottleneck in real-time simulation of diffraction is the enumeration of high-order diffraction propagation paths in scenes with complex geometry (e.g. highly tessellated surfaces). To this end, we present a dynamic geometric diffraction approach that consists of an extensive mesh preprocessing pipeline and complementary runtime algorithm. The preprocessing module identifies a small subset of edges that are important for diffraction using a novel silhouette edge detection heuristic. It also extends these edges with planar diffraction geometry and precomputes a graph data structure encoding the visibility between the edges. The runtime module uses bidirectional path tracing against the diffraction geometry to probabilistically explore potential paths between sources and listeners, then evaluates the intensities for these paths using the Uniform Theory of Diffraction. It uses the edge visibility graph and the A* pathfinding algorithm to robustly and efficiently find additional high-order diffraction paths. We demonstrate how this technique can simulate 10th-order diffraction up to 568 times faster than the previous state of the art, and can efficiently handle large scenes with both high geometric complexity and high numbers of sources. Carl Schissler, Gregor Mückl, Paul Calamia |
ACM Trans. Graph. | 1 |
| 2020 | SoundSpaces: Audio-Visual Navigation in 3D Environments
Changan Chen, Unnat Jain, Carl Schissler, Sebastià Vicenc Amengual Garí, Ziad Al-Halah, Vamsi K. Ithapu, Philip W. Robinson, Kristen Grauman |
ECCV (6) | 3 |
| 2020 | VisualEchoes: Spatial Image Representation Learning Through Echolocation
Ruohan Gao, Changan Chen, Ziad Al-Halah, Carl Schissler, Kristen Grauman |
ECCV (9) | 4 |
| 2018 | Effects of virtual acoustics on target-word identification performance in multi-talker environmentsabstractMany virtual reality applications let multiple users communicate in a multi-talker environment, recreating the classic cocktail-party effect. While there is a vast body of research focusing on the perception and intelligibility of human speech in real-world scenarios with cocktail party effects, there is little work in accurately modeling and evaluating the effect in virtual environments. Given the goal of evaluating the impact of virtual acoustic simulation on the cocktail party effect, we conducted experiments to establish the signal-to-noise ratio (SNR) thresholds for target-word identification performance. Our evaluation was performed for sentences from the coordinate response measure corpus in presence of multi-talker babble. The thresholds were established under varying sound propagation and spatialization conditions. We used a state-of-the-art geometric acoustic system integrated into the Unity game engine to simulate varying conditions of reverberance (direct sound, direct sound & early reflections, direct sound and early reflections and late reverberation) and spatialization (mono, stereo, and binaural). Our results show that spatialization has the biggest effect on the ability of listeners to discern the target words in multi-talker virtual environments. Reverberance, on the other hand, slightly affects the target word discerning ability negatively. Atul Rungta, Nicholas Rewkowski, Carl Schissler, Philip W. Robinson, Ravish Mehra, Dinesh Manocha |
SAP | 3 |
| 2018 | Diffraction Kernels for Interactive Sound Propagation in Dynamic EnvironmentsabstractWe present a novel method to generate plausible diffraction effects for interactive sound propagation in dynamic scenes. Our approach precomputes a diffraction kernel for each dynamic object in the scene and combines them with interactive ray tracing algorithms at runtime. A diffraction kernel encapsulates the sound interaction behavior of individual objects in the free field and we present a new source placement algorithm to significantly accelerate the precomputation. Our overall propagation algorithm can handle highly-tessellated or smooth objects undergoing rigid motion. We have evaluated our algorithm's performance on different scenarios with multiple moving objects and demonstrate the benefits over prior interactive geometric sound propagation methods. We also performed a user study to evaluate the perceived smoothness of the diffracted field and found that the auditory perception using our approach is comparable to that of a wave-based sound propagation method. Atul Rungta, Carl Schissler, Nicholas Rewkowski, Ravish Mehra, Dinesh Manocha |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2018 | Acoustic Classification and Optimization for Multi-Modal Rendering of Real-World ScenesabstractWe present a novel algorithm to generate virtual acoustic effects in captured 3D models of real-world scenes for multimodal augmented reality. We leverage recent advances in 3D scene reconstruction in order to automatically compute acoustic material properties. Our technique consists of a two-step procedure that first applies a convolutional neural network (CNN) to estimate the acoustic material properties, including frequency-dependent absorption coefficients, that are used for interactive sound propagation. In the second step, an iterative optimization algorithm is used to adjust the materials determined by the CNN until a virtual acoustic simulation converges to measured acoustic impulse responses. We have applied our algorithm to many reconstructed real-world indoor scenes and evaluated its fidelity for augmented reality applications. Carl Schissler, Christian Loftin, Dinesh Manocha |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2017 | Efficient construction of the spatial room impulse responseabstractAn important component of the modeling of sound propagation for virtual reality (VR) is the spatialization of the room impulse response (RIR) for directional listeners. This involves convolution of the listener's head-related transfer function (HRTF) with the RIR to generate a spatial room impulse response (SRIR) which can be used to auralize the sound entering the listener's ear canals. Previous approaches tend to evaluate the HRTF for each sound propagation path, though this is too slow for interactive VR latency requirements. We present a new technique for computation of the SRIR that performs the convolution with the HRTF in the spherical harmonic (SH) domain for RIR partitions of a fixed length. The main contribution is a novel perceptually-driven metric that adaptively determines the lowest SH order required for each partition to result in no perceptible error in the SRIR. By using lower SH order for some partitions, our technique saves a significant amount of computation and is almost an order of magnitude faster than the previous approach. We compared the subjective impact of this new method to the previous one and observe a strong scene-dependent preference for our technique. As a result, our method is the first that can compute high-quality spatial sound for the entire impulse response fast enough to meet the audio latency requirements of interactive virtual reality applications. Carl Schissler, Peter Stirling, Ravish Mehra |
VR | 1 |
| 2017 | Interactive Sound Propagation and Rendering for Large Multi-Source ScenesabstractWe present an approach to generate plausible acoustic effects at interactive rates in large dynamic environments containing many sound sources. Our formulation combines listener-based backward ray tracing with sound source clustering and hybrid audio rendering to handle complex scenes. We present a new algorithm for dynamic late reverberation that performs high-order ray tracing from the listener against spherical sound sources. We achieve sublinear scaling with the number of sources by clustering distant sound sources and taking relative visibility into account. We also describe a hybrid convolution-based audio rendering technique that can process hundreds of thousands of sound paths at interactive rates. We demonstrate the performance on many indoor and outdoor scenes with up to 200 sound sources. In practice, our algorithm can compute more than 50 reflection orders at interactive rates on a multicore PC, and we observe a 5x speedup over prior geometric sound propagation algorithms. Carl Schissler, Dinesh Manocha |
ACM Trans. Graph. | 1 |
| 2016 | Adaptive impulse response modeling for interactive sound propagationabstractWe present novel techniques to accelerate the computation of impulse responses for interactive sound rendering. Our formulation is based on geometric acoustic algorithms that use ray tracing to compute the propagation paths from each source to the listener in large, dynamic scenes. In order to accelerate generation of realistic acoustic effects in multi-source scenes, we introduce two novel concepts: the impulse response cache and an adaptive frequency-driven ray tracing algorithm that exploits psychoacoustic characteristics of the impulse response length. As compared to prior approaches, we trace relatively fewer rays while maintaining high simulation fidelity for real-time applications. Furthermore, our approach can handle highly reverberant scenes and high-dynamic-range sources. We demonstrate its application in many scenarios and have observed a 5x speedup in computation time and about two orders of magnitude reduction in memory overhead compared to previous approaches. We also present the results of a preliminary user evaluation of our approach. Carl Schissler, Dinesh Manocha |
I3D | 1 |
| 2016 | Interactive sound propagation with bidirectional path tracingabstractWe introduce Bidirectional Sound Transport (BST) , a new algorithm that simulates sound propagation by bidirectional path tracing using multiple importance sampling. Our approach can handle multiple sources in large virtual environments with complex occlusion, and can produce plausible acoustic effects at an interactive rate on a desktop PC. We introduce a new metric based on the signal-to-noise ratio (SNR) of the energy response and use this metric to evaluate the performance of ray-tracing-based acoustic simulation methods. Our formulation exploits temporal coherence in terms of using the resulting sample distribution of the previous frame to guide the sample distribution of the current one. We show that our sample redistribution algorithm converges and better balances between early and late reflections. We evaluate our approach on different benchmarks and demonstrate significant speedup over prior geometric acoustic algorithms. Chunxiao Cao, Zhong Ren 0001, Carl Schissler, Dinesh Manocha, Kun Zhou 0001 |
ACM Trans. Graph. | 3 |
| 2016 | SynCoPation: Interactive Synthesis-Coupled Sound PropagationabstractRecent research in sound simulation has focused on either sound synthesis or sound propagation, and many standalone algorithms have been developed for each domain. We present a novel technique for coupling sound synthesis with sound propagation to automatically generate realistic aural content for virtual environments. Our approach can generate sounds from rigid-bodies based on the vibration modes and radiation coefficients represented by the single-point multipole expansion. We present a mode-adaptive propagation algorithm that uses a perceptual Hankel function approximation technique to achieve interactive runtime performance. The overall approach allows for high degrees of dynamism - it can support dynamic sources, dynamic listeners, and dynamic directivity simultaneously. We have integrated our system with the Unity game engine and demonstrate the effectiveness of this fully-automatic technique for audio content creation in complex indoor and outdoor scenes. We conducted a preliminary, online user-study to evaluate whether our Hankel function approximation causes any perceptible loss of audio quality. The results indicate that the subjects were unable to distinguish between the audio rendered using the approximate function and audio rendered using the full Hankel function in the Cathedral, Tuscany, and the Game benchmarks. Atul Rungta, Carl Schissler, Ravish Mehra, Chris Malloy, Ming C. Lin, Dinesh Manocha |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2016 | Efficient HRTF-based Spatial Audio for Area and Volumetric SourcesabstractWe present a novel spatial audio rendering technique to handle sound sources that can be represented by either an area or a volume in VR environments. As opposed to point-sampled sound sources, our approach projects the area-volumetric source to the spherical domain centered at the listener and represents this projection area compactly using the spherical harmonic (SH) basis functions. By representing the head-related transfer function (HRTF) in the same basis, we demonstrate that spatial audio which corresponds to an area-volumetric source can be efficiently computed as a dot product of the SH coefficients of the projection area and the HRTF. This results in an efficient technique whose computational complexity and memory requirements are independent of the complexity of the sound source. Our approach can support dynamic area-volumetric sound sources at interactive rates. We evaluate the performance of our technique in large complex VR environments and demonstrate significant improvement over the naive point-sampling technique. We also present results of a user evaluation, conducted to quantify the subjective preference of the user for our approach over the point-sampling approach in VR environments. Carl Schissler, Aaron Nicholls, Ravish Mehra |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2014 | High-order diffraction and diffuse reflections for interactive sound propagation in large environmentsabstractWe present novel algorithms for modeling interactive diffuse reflections and higher-order diffraction in large-scale virtual environments. Our formulation is based on ray-based sound propagation and is directly applicable to complex geometric datasets. We use an incremental approach that combines radiosity and path tracing techniques to iteratively compute diffuse reflections. We also present algorithms for wavelength-dependent simplification and visibility graph computation to accelerate higher-order diffraction at runtime. The overall system can generate plausible sound effects at interactive rates in large, dynamic scenes that have multiple sound sources. We highlight the performance in complex indoor and outdoor environments and observe an order of magnitude performance improvement over previous methods. Carl Schissler, Ravish Mehra, Dinesh Manocha |
ACM Trans. Graph. | 1 |
| 2012 | Guided Multiview Ray Tracing for Fast AuralizationabstractWe present a novel method for tuning geometric acoustic simulations based on ray tracing. Our formulation computes sound propagation paths from source to receiver and exploits the independence of visibility tests and validation tests to dynamically guide the simulation to high accuracy and performance. Our method makes no assumptions of scene layout and can account for moving sources, receivers, and geometry. We combine our guidance algorithm with a fast GPU sound propagation system for interactive simulation. Our implementation efficiently computes early specular paths and first order diffraction with a multiview tracing algorithm. We couple our propagation simulation with an audio output system supporting a high order interpolation scheme that accounts for attenuation, cross fading, and delay. The resulting system can render acoustic spaces composed of thousands of triangles interactively. Micah T. Taylor, Anish Chandak, Christian Lauterbach, Carl Schissler, Dinesh Manocha |
IEEE Trans. Vis. Comput. Graph. | 5 |