Nicholas Rewkowski

dblp:198/3879 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
9 papers
Audio and music processing · 34% Virtual and augmented reality · 32% Computer animation and physical simulation · 29%
Human-computer interaction and pervasive computing
4 papers
Immersive interaction · 35% Learning and educational technologies · 23% Haptics and multimodal interaction · 20%
Artificial intelligence
4 papers
Face, body and person analysis · 38% 3D vision · 37% Generative modeling · 25%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 50% Embedded and real-time systems · 50%

Topics — the 30 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Virtual and augmented reality
virtual agents
0.922021
Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents**This work has been supported in part by ARO Grants W911NF1910069 and W911NF1910315, and Intel. Code and additional materials available at: https: //gamma.umd.edu/t2g · VR 2021
Generating Emotive Gaits for Virtual Agents Using Affect-Based Autoregression · ISMAR 2020
Learning and educational technologies
distance learning
0.812024
An Overview of Enhancing Distance Learning Through Emerging Augmented and Virtual Reality Technologies · IEEE Trans. Vis. Comput. Graph. 2024
Audio and music processing › acoustic simulation
sound propagation
0.722019
P-Reverb: Perceptual Characterization of Early and Late Reflections for Auditory Displays · VR 2019
Diffraction Kernels for Interactive Sound Propagation in Dynamic Environments · IEEE Trans. Vis. Comput. Graph. 2018
Haptics and multimodal interaction › haptic feedback
haptic guidance
0.712023
A Framework for Active Haptic Guidance Using Robotic Haptic Proxies · ICRA 2023
Machine learning › Generative modeling › diffusion model › human motion generation
text-to-motion generation
0.512021
Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents**This work has been supported in part by ARO Grants W911NF1910069 and W911NF1910315, and Intel. Code and additional materials available at: https: //gamma.umd.edu/t2g · VR 2021
Computer animation and physical simulation › gesture generation
co-speech gesture generation
0.512021
Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning · ACM Multimedia 2021
Computer animation and physical simulation
gesture generation
0.512021
Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents**This work has been supported in part by ARO Grants W911NF1910069 and W911NF1910315, and Intel. Code and additional materials available at: https: //gamma.umd.edu/t2g · VR 2021
Virtual and augmented reality
immersive interaction
0.512021
Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents**This work has been supported in part by ARO Grants W911NF1910069 and W911NF1910315, and Intel. Code and additional materials available at: https: //gamma.umd.edu/t2g · VR 2021
Human-robot interaction
emotion expression
0.512021
Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning · ACM Multimedia 2021
Computer animation and physical simulation
character animation
0.412020
Generating Emotive Gaits for Virtual Agents Using Affect-Based Autoregression · ISMAR 2020
Computer animation and physical simulation › character animation
gait generation
0.412020
Generating Emotive Gaits for Virtual Agents Using Affect-Based Autoregression · ISMAR 2020
Audio and music processing
auditory display
0.412019
P-Reverb: Perceptual Characterization of Early and Late Reflections for Auditory Displays · VR 2019
Audio and music processing › sound synthesis › physical modeling synthesis
modal sound synthesis
0.412019
Audio-Material Reconstruction for Virtualized Reality Using a Probabilistic Damping Model · IEEE Trans. Vis. Comput. Graph. 2019
Audio and music processing
sound synthesis
0.412019
Audio-Material Reconstruction for Virtualized Reality Using a Probabilistic Damping Model · IEEE Trans. Vis. Comput. Graph. 2019
Immersive interaction
locomotion
0.412019
Evaluating the Effectiveness of Redirected Walking with Auditory Distractors for Navigation in Virtual Environments · VR 2019
Immersive interaction › virtual reality locomotion
redirected walking
0.412019
Evaluating the Effectiveness of Redirected Walking with Auditory Distractors for Navigation in Virtual Environments · VR 2019
Immersive interaction
virtual reality
0.412019
Evaluating the Effectiveness of Redirected Walking with Auditory Distractors for Navigation in Virtual Environments · VR 2019
Computer vision › 3D vision
3d reconstruction
0.312018
Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras · IEEE Trans. Vis. Comput. Graph. 2018
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.312018
Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras · IEEE Trans. Vis. Comput. Graph. 2018
Computer vision › Face, body and person analysis › human pose estimation › 3d pose estimation
egocentric pose estimation
0.312018
Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras · IEEE Trans. Vis. Comput. Graph. 2018
Computer vision › Face, body and person analysis
human pose estimation
0.312018
Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras · IEEE Trans. Vis. Comput. Graph. 2018
Audio and music processing
acoustic simulation
0.312018
Diffraction Kernels for Interactive Sound Propagation in Dynamic Environments · IEEE Trans. Vis. Comput. Graph. 2018
Rendering › physically based rendering › wave optics
diffraction
0.312018
Diffraction Kernels for Interactive Sound Propagation in Dynamic Environments · IEEE Trans. Vis. Comput. Graph. 2018
Embedded and real-time systems › wireless communication › wireless sensor networks
sensor placement
0.312017
Optimizing placement of commodity depth cameras for known 3D dynamic scene capture · VR 2017
Performance modeling and evaluation › simulation
simulation optimization
0.312017
Optimizing placement of commodity depth cameras for known 3D dynamic scene capture · VR 2017
Virtual and augmented reality
mixed reality
0.212023
A Framework for Active Haptic Guidance Using Robotic Haptic Proxies · ICRA 2023
Machine learning › Generative modeling
autoregressive model
0.112020
Generating Emotive Gaits for Virtual Agents Using Affect-Based Autoregression · ISMAR 2020
Audio and music processing › room acoustics
reverberation
0.112019
P-Reverb: Perceptual Characterization of Early and Late Reflections for Auditory Displays · VR 2019
Usability and user experience research › quality of experience
motion sickness and presence
0.112019
Evaluating the Effectiveness of Redirected Walking with Auditory Distractors for Navigation in Virtual Environments · VR 2019
Interaction techniques and input › spatial interaction › navigation
virtual reality navigation
0.112019
Evaluating the Effectiveness of Redirected Walking with Auditory Distractors for Navigation in Virtual Environments · VR 2019

Methods — techniques the papers use, named apart from their topics

user study · 2.1literature review · 1.5robotic haptic proxies · 1.3transformer · 1.0mel-frequency cepstral coefficients · 1.0graph convolution · 1.0generative adversarial network · 1.0deep learning · 1.0regularization · 0.9autoregression network · 0.9simulated annealing · 0.6greedy algorithm · 0.6fitness metric · 0.6distraction ratio measurement · 0.4parametric model-based motion estimation · 0.3convolutional neural network · 0.3
YearPublicationVenuePosition
2024 An Overview of Enhancing Distance Learning Through Emerging Augmented and Virtual Reality Technologies
abstract
Although distance learning presents a number of interesting educational advantages as compared to in-person instruction, it is not without its downsides. We first assess the educational challenges presented by distance learning as a whole and identify 4 main challenges that distance learning currently presents as compared to in-person instruction: the lack of social interaction, reduced student engagement and focus, reduced comprehension and information retention, and the lack of flexible and customizable instructor resources. After assessing each of these challenges in-depth, we examine how AR/VR technologies might serve to address each challenge along with their current shortcomings, and finally outline the further research that is required to fully understand the potential of AR/VR technologies as they apply to distance learning.
Elizabeth Childs, Ferzam Mohammad, Logan Stevens, Hugo Burbelo, Amanuel Awoke, Nicholas Rewkowski, Dinesh Manocha
IEEE Trans. Vis. Comput. Graph.6
2023 A Framework for Active Haptic Guidance Using Robotic Haptic Proxies
abstract
Haptic feedback is an important component of creating an immersive mixed reality experience. Traditionally, haptic forces are rendered in response to the user's interactions with the virtual environment. In this work, we explore the idea of rendering haptic forces in a proactive manner, with the explicit intention to influence the user's behavior through compelling haptic forces. To this end, we present a framework for active haptic guidance in mixed reality, using one or more robotic haptic proxies to influence user behavior and deliver a safer and more immersive virtual experience. We provide details on common challenges that need to be overcome when implementing active haptic guidance, and discuss example applications that show how active haptic guidance can be used to influence the user's behavior. Finally, we apply active haptic guidance to a virtual reality navigation problem, and conduct a user study that demonstrates how active haptic guidance creates a safer and more immersive experience for users.
Niall L. Williams, Nicholas Rewkowski, Ming C. Lin
ICRA2
2022 Audio-Visual Depth and Material Estimation for Robot Navigation
abstract
Reflective and textureless surfaces such as windows, mirrors, and walls can be a challenge for scene reconstruction, due to depth discontinuities and holes. We propose an audio-visual method that uses the reflections of sound to aid in depth estimation and material classification for 3D scene reconstruction in robot navigation and AR/VR applications. The mobile phone prototype emits pulsed audio, while recording video for audio-visual classification for 3D scene reconstruction. Reflected sound and images from the video are input into our audio (EchoCNN-A) and audio-visual (EchoCNN-AV) convolutional neural networks for surface and sound source detection, depth estimation, and material classification. The inferences from these classifications enhance 3D scene reconstructions containing open spaces and reflective surfaces by depth filtering, inpainting, and placement of unmixed sound sources in the scene. Our prototype, demos, and experimental results from real-world with challenging surfaces and sound, also validated with virtual scenes, indicate high success rates on classification of material, depth estimation, and closed/open surfaces, leading to considerable improvement in 3D scene reconstruction for robot navigation.
Justin Wilson, Nicholas Rewkowski, Ming C. Lin
IROS2
2021 Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning
abstract
We present a generative adversarial network to synthesize 3D pose sequences of co-speech upper-body gestures with appropriate affective expressions. Our network consists of two components: a generator to synthesize gestures from a joint embedding space of features encoded from the input speech and the seed poses, and a discriminator to distinguish between the synthesized pose sequences and real 3D pose sequences. We leverage the Mel-frequency cepstral coefficients and the text transcript computed from the input speech in separate encoders in our generator to learn the desired sentiments and the associated affective cues. We design an affective encoder using multi-scale spatial-temporal graph convolutions to transform 3D pose sequences into latent, pose-based affective features. We use our affective encoder in both our generator, where it learns affective features from the seed poses to guide the gesture synthesis, and our discriminator, where it enforces the synthesized gestures to contain the appropriate affective expressions. We perform extensive evaluations on two benchmark datasets for gesture synthesis from the speech, the TED Gesture Dataset and the GENEA Challenge 2020 Dataset. Compared to the best baselines, we improve the mean absolute joint error by 10-33%, the mean acceleration difference by 8-58%, and the Fréchet Gesture Distance by 21-34%. We also conduct a user study and observe that compared to the best current baselines, around 15.28% of participants indicated our synthesized gestures appear more plausible, and around 16.32% of participants felt the gestures had more appropriate affective expressions aligned with the speech.
Uttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, Dinesh Manocha
ACM Multimedia3
2021 Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents**This work has been supported in part by ARO Grants W911NF1910069 and W911NF1910315, and Intel. Code and additional materials available at: https: //gamma.umd.edu/t2g
abstract
We present Text2Gestures, a transformer-based learning method to interactively generate emotive full-body gestures for virtual agents aligned with natural language text inputs. Our method generates emotionally expressive gestures by utilizing the relevant biomechanical features for body expressions, also known as affective features. We also consider the intended task corresponding to the text and the target virtual agents' intended gender and handedness in our generation pipeline. We train and evaluate our network on the MPI Emotional Body Expressions Database and observe that our network produces state-of-the-art performance in generating gestures for virtual agents aligned with the text for narration or conversation. Our network can generate these gestures at interactive rates on a commodity GPU. We conduct a web-based user study and observe that around 91% of participants indicated our generated gestures to be at least plausible on a five-point Likert Scale. The emotions perceived by the participants from the gestures are also strongly positively correlated with the corresponding intended emotions, with a minimum Pearson coefficient of 0.77 in the valence dimension.
Uttaran Bhattacharya, Nicholas Rewkowski, Abhishek Banerjee 0006, Pooja Guhan, Aniket Bera, Dinesh Manocha
VR2
2020 Generating Emotive Gaits for Virtual Agents Using Affect-Based Autoregression
abstract
We present a novel autoregression network to generate virtual agents that convey various emotions through their walking styles or gaits. Given the 3D pose sequences of a gait, our network extracts pertinent movement features and affective features from the gait. We use these features to synthesize subsequent gaits such that the virtual agents can express and transition between emotions represented as combinations of happy, sad, angry, and neutral. We incorporate multiple regularizations in the training of our network to simultaneously enforce plausible movements and noticeable emotions on the virtual agents. We also integrate our approach with an AR environment using a Microsoft HoloLens and can generate emotive gaits at interactive rates to increase the social presence. We evaluate how human observers perceive both the naturalness and the emotions from the generated gaits of the virtual agents in a web-based study. Our results indicate around 89% of the users found the naturalness of the gaits satisfactory on a five-point Likert scale, and the emotions they perceived from the virtual agents are statistically similar to the intended emotions of the virtual agents. We also use our network to augment existing gait datasets with emotive gaits and will release this augmented dataset for future research in emotion prediction and emotive gait synthesis. Our project website is available at https://gamma.umd.edu/gen-emotive-gaits/.
Uttaran Bhattacharya, Nicholas Rewkowski, Pooja Guhan, Niall L. Williams, Trisha Mittal, Aniket Bera, Dinesh Manocha
ISMAR2
2019 Evaluating the Effectiveness of Redirected Walking with Auditory Distractors for Navigation in Virtual Environments
abstract
Many virtual locomotion interfaces allowing users to move in virtual reality have been built and evaluated, such as redirected walking (RDW), walking-in-place (WIP), and joystick input. RDW has been shown to be among the most natural and immersive as it supports real walking, and many newer methods further adapt RDW to allow for customization and greater immersion. Most of these methods have been demonstrated to work with vision, in this paper we evaluate the ability for a general distractor-based RDW framework to be used with only auditory display. We conducted two studies evaluating the differences between RDW with auditory distractors and other distractor modalities using distraction ratio, virtual and physical path information, immersion, simulator sickness, and other measurements. Our results indicate that auditory RDW has the potential to be used with complex navigational tasks, such as crossing streets and avoiding obstacles. It can be used without designing the system specifically for audio-only users. Additionally, sense of presence and simulator sickness remain reasonable across all user groups.
Nicholas Rewkowski, Atul Rungta, Mary C. Whitton, Ming C. Lin
VR1
2019 P-Reverb: Perceptual Characterization of Early and Late Reflections for Auditory Displays
abstract
We introduce a novel, perceptually derived metric (P - Reverb) that relates the just-noticeable difference (JND) of the early sound field (also called early reflections) to the late sound field (known as late reflections or reverberation). Early and late reflections are crucial components of the sound field and provide multiple perceptual cues for auditory displays. We conduct two extensive user evaluations that relate the JNDs of early reflections and late reverberation in terms of the mean-free path of the environment and present a novel P - Reverb metric. Our metric is used to estimate dynamic reverberation characteristics efficiently in terms of important parameters like reverberation time (RT60). We show the numerical accuracy of our P - Reverb metric in estimating RT60. Finally, we use our metric to design an interactive sound propagation algorithm and demonstrate its effectiveness on various benchmarks.
Atul Rungta, Nicholas Rewkowski, Roberta L. Klatzky, Dinesh Manocha
VR2
2019 Audio-Material Reconstruction for Virtualized Reality Using a Probabilistic Damping Model
abstract
Modal sound synthesis has been used to create realistic sounds from rigid-body objects, but requires accurate real-world material parameters. These material parameters can be estimated from recorded sounds of an impacted object, but external factors can interfere with accurate parameter estimation. We present a novel technique for estimating the damping parameters of materials from recorded impact sounds that probabilistically models these external factors. We represent the combined effects of material damping, support damping, and sampling inaccuracies with a probabilistic generative model, then use maximum likelihood estimation to fit a damping model to recorded data. This technique greatly reduces the human effort needed and does not require the precise object geometry or the exact hit location. We validate the effectiveness of this technique with a comprehensive analysis of a synthetic dataset and a perceptual study on object identification. We also present a study establishing human performance on the same parameter estimation task for comparison.
Auston Sterling, Nicholas Rewkowski, Roberta L. Klatzky, Ming C. Lin
IEEE Trans. Vis. Comput. Graph.2
2018 Effects of virtual acoustics on target-word identification performance in multi-talker environments
abstract
Many virtual reality applications let multiple users communicate in a multi-talker environment, recreating the classic cocktail-party effect. While there is a vast body of research focusing on the perception and intelligibility of human speech in real-world scenarios with cocktail party effects, there is little work in accurately modeling and evaluating the effect in virtual environments. Given the goal of evaluating the impact of virtual acoustic simulation on the cocktail party effect, we conducted experiments to establish the signal-to-noise ratio (SNR) thresholds for target-word identification performance. Our evaluation was performed for sentences from the coordinate response measure corpus in presence of multi-talker babble. The thresholds were established under varying sound propagation and spatialization conditions. We used a state-of-the-art geometric acoustic system integrated into the Unity game engine to simulate varying conditions of reverberance (direct sound, direct sound & early reflections, direct sound and early reflections and late reverberation) and spatialization (mono, stereo, and binaural). Our results show that spatialization has the biggest effect on the ability of listeners to discern the target words in multi-talker virtual environments. Reverberance, on the other hand, slightly affects the target word discerning ability negatively.
Atul Rungta, Nicholas Rewkowski, Carl Schissler, Philip W. Robinson, Ravish Mehra, Dinesh Manocha
SAP2
2018 Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras
abstract
We propose a new approach for 3D reconstruction of dynamic indoor and outdoor scenes in everyday environments, leveraging only cameras worn by a user. This approach allows 3D reconstruction of experiences at any location and virtual tours from anywhere. The key innovation of the proposed ego-centric reconstruction system is to capture the wearer's body pose and facial expression from near-body views, e.g. cameras on the user's glasses, and to capture the surrounding environment using outward-facing views. The main challenge of the ego-centric reconstruction, however, is the poor coverage of the near-body views - that is, the user's body and face are observed from vantage points that are convenient for wear but inconvenient for capture. To overcome these challenges, we propose a parametric-model-based approach to user motion estimation. This approach utilizes convolutional neural networks (CNNs) for near-view body pose estimation, and we introduce a CNN-based approach for facial expression estimation that combines audio and video. For each time-point during capture, the intermediate model-based reconstructions from these systems are used to re-target a high-fidelity pre-scanned model of the user. We demonstrate that the proposed self-sufficient, head-worn capture system is capable of reconstructing the wearer's movements and their surrounding environment in both indoor and outdoor situations without any additional views. As a proof of concept, we show how the resulting 3D-plus-time reconstruction can be immersively experienced within a virtual reality system (e.g., the HTC Vive). We expect that the size of the proposed egocentric capture-and-reconstruction system will eventually be reduced to fit within future AR glasses, and will be widely useful for immersive 3D telepresence, virtual tours, and general use-anywhere 3D content creation.
Young-Woon Cha, True Price, Xinran Lu, Nicholas Rewkowski, Rohan Chabra, Zihe Qin, Hyounghun Kim, Zhaoqi Su, Yebin Liu, Adrian Ilie, Andrei State, Zhenlin Xu, Jan-Michael Frahm, Henry Fuchs
IEEE Trans. Vis. Comput. Graph.5
2018 Diffraction Kernels for Interactive Sound Propagation in Dynamic Environments
abstract
We present a novel method to generate plausible diffraction effects for interactive sound propagation in dynamic scenes. Our approach precomputes a diffraction kernel for each dynamic object in the scene and combines them with interactive ray tracing algorithms at runtime. A diffraction kernel encapsulates the sound interaction behavior of individual objects in the free field and we present a new source placement algorithm to significantly accelerate the precomputation. Our overall propagation algorithm can handle highly-tessellated or smooth objects undergoing rigid motion. We have evaluated our algorithm's performance on different scenarios with multiple moving objects and demonstrate the benefits over prior interactive geometric sound propagation methods. We also performed a user study to evaluate the perceived smoothness of the diffracted field and found that the auditory perception using our approach is comparable to that of a wave-based sound propagation method.
Atul Rungta, Carl Schissler, Nicholas Rewkowski, Ravish Mehra, Dinesh Manocha
IEEE Trans. Vis. Comput. Graph.3
2017 Optimizing placement of commodity depth cameras for known 3D dynamic scene capture
abstract
Commodity depth cameras, such as the Microsoft Kinect®, have been widely used for the capture and reconstruction of the 3D structure of room-sized dynamic scenes. Camera placement and coverage during capture significantly impact the quality of the resulting reconstruction. In particular, dynamic occlusions and sensor interference have been shown to result in poor resolution and holes in the reconstruction results. This paper presents a novel algorithmic framework and a method for off-line optimization of depth cameras placements for a given 3D dynamic scene, simulated using virtual 3D models. We derive a fitness metric for a particular configuration of sensors by combining factors such as visibility and resolution of the entire dynamic scene with probabilities of interference between sensors. We employ this fitness metric both in a greedy algorithm that determines the number of depth cameras needed to cover the scene, and in a simulated annealing algorithm that optimizes the placements of those sensors. We compare our algorithm's optimized placements with manual sensor placements for a real dynamic scene. We present quantitative assessments using our fitness metric, as well as qualitative assessments to demonstrate that our algorithm not only enhances the resolution and total coverage of the reconstruction, but also fills in voids by avoiding occlusions and sensor interference when compared with the reconstruction of the same scene using mual sensor placement.
Rohan Chabra, Adrian Ilie, Nicholas Rewkowski, Young-Woon Cha, Henry Fuchs
VR3
2017 Glass half full: sound synthesis for fluid-structure coupling using added mass operator
Justin Wilson, Auston Sterling, Nicholas Rewkowski, Ming C. Lin
Vis. Comput.3