Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jonathan Swartz

dblp:89/5502 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
2since 2021 · last 2024
0000-0003-1959-6396ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
3D vision · 62% Deep learning architectures and training · 38%
Computer graphics and multimedia
2 papers
Computer animation and physical simulation · 99% Image and video coding · 1%

Topics — the 4 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › geometric deep learning
3d deep learning
0.812024
fVDB : A Deep-Learning Framework for Sparse, Large Scale, and High Performance Spatial Intelligence · ACM Trans. Graph. 2024
Computer animation and physical simulation
facial animation
0.812024
Near-realtime Facial Animation by Deep 3D Simulation Super-Resolution · ACM Trans. Graph. 2024
Computer vision › 3D vision
neural radiance field
0.212024
fVDB : A Deep-Learning Framework for Sparse, Large Scale, and High Performance Spatial Intelligence · ACM Trans. Graph. 2024
Computer vision › 3D vision › 3d reconstruction
point cloud reconstruction
0.212024
fVDB : A Deep-Learning Framework for Sparse, Large Scale, and High Performance Spatial Intelligence · ACM Trans. Graph. 2024

Methods — techniques the papers use, named apart from their topics

tensor-core convolution · 0.8super-resolution · 0.8sparse grid acceleration · 0.8neural network · 0.8jagged tensors · 0.8hierarchical DDA ray tracing · 0.8
YearPublicationVenuePosition
2024 Near-realtime Facial Animation by Deep 3D Simulation Super-Resolution
abstract
We present a neural network-based simulation super-resolution framework that can efficiently and realistically enhance a facial performance produced by a low-cost, real-time physics-based simulation to a level of detail that closely approximates that of a reference-quality off-line simulator with much higher resolution (27× element count in our examples) and accurate physical modeling. Our approach is rooted in our ability to construct a training set of paired frames, from the low- and high-resolution simulators respectively, that are in semantic correspondence with each other. We use face animation as an exemplar of such a simulation domain, where creating this semantic congruence is achieved by simply dialing in the same muscle actuation controls and skeletal pose in the two simulators. Our proposed neural network super-resolution framework generalizes from this training set to unseen expressions, compensates for modeling discrepancies between the two simulations due to limited resolution or cost-cutting approximations in the real-time variant, and does not require any semantic descriptors or parameters to be provided as input, other than the result of the real-time simulation. We evaluate the efficacy of our pipeline on a variety of expressive performances and provide comparisons and ablation experiments for plausible variations and alternatives to our proposed scheme. Our code is available at https://github.com/hjoonpark/3d-sim-super- res.git.
Hyojoon Park, Sangeetha Grama Srinivasan, Matthew Cong, Doyub Kim, Byungsoo Kim 0001, Jonathan Swartz, Ken Museth, Eftychios Sifakis
ACM Trans. Graph.6
2024 fVDB : A Deep-Learning Framework for Sparse, Large Scale, and High Performance Spatial Intelligence
abstract
We present f VDB, a novel GPU-optimized framework for deep learning on large-scale 3D data. f VDB provides a complete set of differentiable primitives to build deep learning architectures for common tasks in 3D learning such as convolution, pooling, attention, ray-tracing, meshing, etc. f VDB simultaneously provides a much larger feature set (primitives and operators) than established frameworks with no loss in efficiency: our operators match or exceed the performance of other frameworks with narrower scope. Furthermore, f VDB can process datasets with much larger footprint and spatial resolution than prior works, while providing a competitive memory footprint on small inputs. To achieve this combination of versatility and performance, f VDB relies on a single novel VDB index grid acceleration structure paired with several key innovations including GPU accelerated sparse grid construction, convolution using tensorcores, fast ray tracing kernels using a Hierarchical Digital Differential Analyzer algorithm (HDDA), and jagged tensors. Our framework is fully integrated with PyTorch enabling interoperability with existing pipelines, and we demonstrate its effectiveness on a number of representative tasks such as large-scale point-cloud segmentation, high resolution 3D generative modeling, unbounded scale Neural Radiance Fields, and large-scale point cloud reconstruction.
Francis Williams, Jonathan Swartz, Gergely Klár, Vijay Thakkar, Matthew Cong, Xuanchi Ren, Ruilong Li, Clement Fuji-Tsang, Sanja Fidler, Eftychios Sifakis, Ken Museth
ACM Trans. Graph.3
1995 A Resolution Independent Video Language
abstract
No abstract available.
Jonathan Swartz, Brian Christopher Smith
ACM Multimedia1
1990 Allophone clustering for continuous speech recognition
abstract
Two methods are presented for subword clustering. The first method is an agglomerative clustering algorithm. This method is completely data-driven and finds clusters without any external guidance. The second method uses decision trees for clustering. This method uses an expert-generated list of questions about contexts and recursively selects the most appropriate question to split the allophones. Preliminary results showed that when the training set has a good coverage of the allophonic variations in the test set, both method are capable of high-performance recognition. However, under vocabulary-independent conditions, the method using tree-based allophones outperformed agglomerative clustering because of its superior generalization capability.>
Kai-Fu Lee, Satoru Hayamizu, Hsiao-Wuen Hon, Cecil Huang, Jonathan Swartz, Robert Weide
ICASSP5