Gurunandan Krishnan

dblp:80/6021 · also Guru Krishnan · DBLP profile ↗
← Back
19ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0002-1533-2169ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Velocity Disambiguation for Video Frame Interpolation
abstract
Existing video frame interpolation (VFI) methods blindly predict where each object is at a specific timestep $t$t ("time indexing"), which struggles to predict precise object movements. Given two images of a baseball, there are infinitely many possible trajectories: accelerating or decelerating, straight or curved. This often results in blurry frames as the method averages out these possibilities. Instead of forcing the network to learn this complicated time-to-location mapping implicitly together with predicting the frames, we provide the network with an explicit hint on how far the object has traveled between start and end frames, a novel approach termed "distance indexing". This method offers a clearer learning goal for models, reducing the uncertainty tied to object speeds. We further observed that, even with this extra guidance, objects can still be blurry especially when they are equally far from both input frames (i.e., halfway in-between), due to the directional ambiguity in long-range motion. To solve this, we propose an iterative reference-based estimation strategy that breaks down a long-range prediction into several short-range steps. When integrating our plug-and-play strategies into state-of-the-art learning-based models, they exhibit markedly sharper outputs and superior perceptual quality in arbitrary time interpolations, using a uniform distance indexing map in the same format as time indexing without requiring extra computation. Furthermore, we demonstrate that if additional latency is acceptable, a continuous map estimator can be employed to compute a pixel-wise dense distance indexing using multiple nearby frames. Combined with efficient multi-frame refinement, this extension can further disambiguate complex motion, thus enhancing performance both qualitatively and quantitatively. Additionally, the ability to manually specify distance indexing allows for independent temporal manipulation of each object, providing a novel tool for video editing tasks such as re-timing.
Zhihang Zhong, Wei Wang 0333, Xiao Sun 0001, Yu Qiao 0001, Gurunandan Krishnan, Sizhuo Ma, Jian Wang 0100
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Privacy-Enabled Parallax Display
abstract
Privacy filters for displays are designed to obfuscate or hide visual content from unintended observers, while making the displayed information visible only to selected viewers. Existing privacy filters suffer from either wide viewing field (low selectivity) or limited user positioning. To solve this dilemma, we propose a display technology that allows a narrow but adaptive viewing field that can be directed to arbitrary user location. While conventional parallax barriers provide such capability of modulating the light field according to the user location, it suffers from repeated views. Our key observation is that this view repetition originates from the periodicity of barrier patterns, and we propose a privacy-enabled parallax display based on randomized barrier design. In addition to randomizing the locations of 1D slits, we also propose breaking down the slits into pinholes and randomizing their 2D locations, which results in privacy-preservation along the vertical direction as well. We build a hardware prototype using two off-the-shelf liquid-crystal displays. Experiments show that the proposed randomized parallax barrier can direct to the user a narrow viewing field of about ±6°, providing a significantly improved privacy protection as compared to traditional privacy screens.
Sizhuo Ma, Karl Bayer, Gurunandan Krishnan, Mohit Gupta 0001, Shree K. Nayar
VR3
2025 Copy or Not? Reference-Based Face Image Restoration with Fine Details
Min Jin Chong, Dejia Xu, Zhangyang Wang, David A. Forsyth, Gurunandan Krishnan
WACV6
2025 Personalized Restoration via Dual-Pivot Tuning
abstract
Generative diffusion models can serve as priors, ensuring that image restoration solutions adhere to natural image manifolds. For facial images, however, personalized priors are essential to accurately reconstruct individual-specific facial features. We propose Dual-Pivot Tuning - a simple yet effective two-stage approach to personalize blind restoration systems while preserving general prior integrity. Our key observation is that for efficient personalization, the diffusion model should be tuned around a fixed textual pivot in the first step, while in the second step a guiding network should be tuned in a generic (non-personalized) manner, using the personalized diffusion model as a fixed "pivot". This approach ensures that personalization does not interfere with the restoration process, producing results with a natural appearance that show high fidelity to both identity and degraded image attributes. We conducted extensive experiments with images of widely recognized individuals, evaluating our approach both qualitatively and quantitatively against relevant baselines. Notably, our personalized prior not only achieves superior identity fidelity, but also outperforms state-of-the-art generic priors in terms of overall image quality. Project webpage is https://personalized-restoration.github.io/ and code is available at https://github.com/personalized-restoration/personalized-restoration.
Pradyumna Chari, Sizhuo Ma, Daniil Ostashev, Achuta Kadambi, Gurunandan Krishnan, Jian Wang 0100, Kfir Aberman
IEEE Trans. Image Process.5
2025 Privacy-Preserving Visual Localization With Event Cameras
abstract
We consider the problem of client-server localization, where edge device users communicate visual data with the service provider for locating oneself against a pre-built 3D map. This localization paradigm is a crucial component for location-based services in AR/VR or mobile applications, as it is not trivial to store large-scale 3D maps and process fast localization on resource-limited edge devices. Nevertheless, conventional client-server localization systems possess numerous challenges in computational efficiency, robustness, and privacy-preservation during data transmission. Our work aims to jointly solve these challenges with a localization pipeline based on event cameras. By using event cameras, our system consumes low energy and maintains small memory bandwidth. Then during localization, we propose applying event-to-image conversion and leverage mature image-based localization, which achieves robustness even in low-light or fast-moving scenes. To further enhance privacy protection, we introduce privacy protection techniques at two levels. Network level protection aims to hide the entire user's view in private scenes using a novel split inference approach, while sensor level protection aims to hide sensitive user details such as faces with light-weight filtering. Both methods involve small client-side computation and localization performance loss, while significantly mitigating the feeling of insecurity as revealed in our user study. We thus project our method to serve as a building block for practical location-based services using event cameras.
Young Min Kim 0001, Ramzi Zahreddine, Weston A. Welge, Gurunandan Krishnan, Sizhuo Ma, Jian Wang 0100
IEEE Trans. Image Process.5
2024 DSL-FIQA: Assessing Facial Image Quality via Dual-Set Degradation Learning and Landmark-Guided Transformer
abstract
Generic Face Image Quality Assessment (GFIQA) evalu-ates the perceptual quality of facial images, which is crucial in improving image restoration algorithms and selecting high-quality face images for downstream tasks. We present a novel transformer-based method for GFIQA, which is aided by two unique mechanisms. First, a “Dual-Set Degradation Representation Learning” (DSL) mechanism uses facial images with both synthetic and real degradations to decouple degradation from content, ensuring gen-eralizability to real-world scenarios. This self-supervised method learns degradation features on a global scale, pro-viding a robust alternative to conventional methods that use local patch information in degradation learning. Second, our transformer leverages facial landmarks to emphasize visually salient parts of a face image in evaluating its per-ceptual quality. We also introduce a balanced and diverse Comprehensive Generic Face IQA (CGFIQA-40k) dataset of 40K images carefully designed to overcome the biases, in particular the imbalances in skin tone and gender represen-tation, in existing datasets. Extensive analysis and evaluation demonstrate the robustness of our method, marking a significant improvement over prior methods.
Gurunandan Krishnan, Sy-Yen Kuo, Sizhuo Ma, Jian Wang 0100
CVPR2
2024 Delving Deep into Engagement Prediction of Short Videos
Dasong Li, Baili Lu, Hongsheng Li 0001, Sizhuo Ma, Gurunandan Krishnan, Jian Wang 0100
ECCV (54)6
2024 Clearer Frames, Anytime: Resolving Velocity Ambiguity in Video Frame Interpolation
Zhihang Zhong, Gurunandan Krishnan, Xiao Sun 0001, Yu Qiao 0001, Sizhuo Ma, Jian Wang 0100
ECCV (33)2
2024 DisCO: Portrait Distortion Correction with Perspective-Aware 3D GANs
Zhixiang Wang 0001, Yu-Lun Liu 0001, Jia-Bin Huang 0001, Shin'ichi Satoh 0001, Sizhuo Ma, Gurunandan Krishnan, Jian Wang 0100
Int. J. Comput. Vis.6
2024 Perspective-Aligned AR Mirror with Under-Display Camera
abstract
Augmented reality (AR) mirrors are novel displays that have great potential for commercial applications such as virtual apparel try-on. Typically the camera is placed beside the display, leading to distorted perspectives during user interaction. In this paper, we present a novel approach to address this problem by placing the camera behind a transparent display, thereby providing users with a perspective-aligned experience. Simply placing the camera behind the display can compromise image quality due to optical effects. We meticulously analyze the image formation process, and present an image restoration algorithm that benefits from physics-based data synthesis and network design. Our method significantly improves image quality and outperforms existing methods especially on the underexplored wire and backscatter artifacts. We then carefully design a full AR mirror system including display and camera selection, real-time processing pipeline, and mechanical design. Our user study demonstrates that the system is exceptionally well-received by users, highlighting its advantages over existing camera configurations not only as an AR mirror, but also for video conferencing. Our work represents a step forward in the development of AR mirrors, with potential applications in retail, cosmetics, fashion, etc. The image restoration dataset and code are available at https://perspective-armirror.github.io/.
Jian Wang 0100, Sizhuo Ma, Karl Bayer, Yi Zhang 0108, Peihao Wang, Bing Zhou 0001, Shree K. Nayar, Gurunandan Krishnan
ACM Trans. Graph.8
2023 AO-Finger: Hands-free Fine-grained Finger Gesture Recognition via Acoustic-Optic Sensor Fusing
abstract
Finger gesture recognition is gaining great research interest for wearable device interactions such as smartwatches and AR/VR headsets. In this paper, we propose a hands-free fine-grained finger gesture recognition system AO-Finger based on acoustic-optic sensor fusing. Specifically, we design a wristband with a modified stethoscope microphone and two high-speed optic motion sensors to capture signals generated from finger movements. We propose a set of natural, inconspicuous and effortless micro finger gestures that can be reliably detected from the complementary signals from both sensors. We design a multi-modal CNN-Transformer model for fast gesture recognition (flick/pinch/tap), and a finger swipe contact detection model to enable fine-grained swipe gesture tracking. We built a prototype which achieves an overall accuracy of 94.83% in detecting fast gestures and enables fine-grained continuous swipe gestures tracking. AO-Finger is practical for use as a wearable device and ready to be integrated into existing wrist-worn devices such as smartwatches.
Chenhan Xu, Bing Zhou 0001, Gurunandan Krishnan, Shree K. Nayar
CHI3
2023 Energy-Efficient Adaptive 3D Sensing
abstract
Active depth sensing achieves robust depth estimation but is usually limited by the sensing range. Naively increasing the optical power can improve sensing range but induces eye-safety concerns for many applications, including autonomous robots and augmented reality. In this paper, we propose an adaptive active depth sensor that jointly optimizes range, power consumption, and eye-safety. The main observation is that we need not project light patterns to the entire scene but only to small regions of interest where depth is necessary for the application and passive stereo depth estimation fails. We theoretically compare this adaptive sensing scheme with other sensing strategies, such as full-frame projection, line scanning, and point scanning. We show that, to achieve the same maximum sensing distance, the proposed method consumes the least power while having the shortest (best) eye-safety distance. We implement this adaptive sensing scheme with two hardware prototypes, one with a phase-only spatial light modulator (SLM) and the other with a micro-electro-mechanical (MEMS) mirror and diffractive optical elements (DOE). Experimental results validate the advantage of our method and demonstrate its capability of acquiring higher quality geometry adaptively. Please see our project website for video results and code: https://btilmon.github.io/e3d.html.
Brevin Tilmon, Zhanghao Sun, Sanjeev J. Koppal, Georgios Evangelidis 0002, Ramzi Zahreddine, Gurunandan Krishnan, Sizhuo Ma, Jian Wang 0100
CVPR7
2023 Personalized Dereverberation of Speech
Ruilin Xu 0001, Gurunandan Krishnan, Changxi Zheng, Shree K. Nayar
INTERSPEECH2
2023 Be Real in Scale: Swing for True Scale in Dual Camera Mode
abstract
Many mobile AR apps that use the front-facing camera can benefit significantly from knowing the metric scale of the user’s face. However, the true scale of the face is hard to measure because monocular vision suffers from a fundamental ambiguity in scale. The methods based on prior knowledge about the scene either have a large error or are not easily accessible. In this paper, we propose a new method to measure the face scale by a simple user interaction: the user only needs to swing the phone to capture two selfies while using the recently popular Dual Camera mode. This mode allows simultaneous streaming of the front camera and the rear cameras and has become a key feature in many social apps. A computer vision method is applied to first estimate the absolute motion of the phone from the images captured by two rear cameras, and then calculate the point cloud of the face by triangulation. We develop a prototype mobile app to validate the proposed method. Our user study shows that the proposed method is favored compared to existing methods because of its high accuracy and ease of use. Our method can be built into Dual Camera mode and can enable a wide range of applications (e.g., virtual try-on for online shopping, true-scale 3D face modeling, gaze tracking, and face anti-spoofing) by introducing true scale to smartphone-based XR. The code is available at https://github.com/ruiyu0/Swing-for-True-Scale.
Rui Yu 0002, Jian Wang 0100, Sizhuo Ma, Sharon X. Huang, Gurunandan Krishnan
ISMAR5
2008 Cata-Fisheye Camera for Panoramic Imaging
abstract
We present a novel panoramic imaging system which uses a curved mirror as a simple optical attachment to a fish- eye lens. When compared to existing panoramic cameras, our "'cata-fisheye" camera has a simple, compact and inexpensive design, and yet yields high optical performance. It captures the desired panoramic field of view in two parts. The upper part is obtained directly by the fisheye lens and the lower part after reflection by the curved mirror. These two parts of the field of view have a small overlap that is used to stitch them into a single seamless panorama. The cata-fisheye concept allows us to design cameras with a wide range of fields of view by simply varying the parameters and position of the curved mirror. We provide an automatic method for the one-time calibration needed to stitch the two parts of the panoramic field of view. We have done a complete performance evaluation of our concept with respect to (i) the optical quality of the captured images, (ii) the working range of the camera over which the parallax is negligible, and (iii) the spatial resolution of the computed panorama. Finally, we have built a prototype cata-fisheye video camera with a spherical mirror that can capture high resolution panoramic images (3600times550pixels) with a 360deg (horizontal) x 55deg (vertical) field of view.
Gurunandan Krishnan, Shree K. Nayar
WACV1
2007 Visual Chatter in the Real World
Shree K. Nayar, Gurunandan Krishnan, Michael D. Grossberg, Ramesh Raskar
ISRR2
2007 Material Based Splashing of Water Drops
Kshitiz Garg, Gurunandan Krishnan, Shree K. Nayar
Rendering Techniques2
2006 Visual Chatter in the Real World
Shree K. Nayar, Gurunandan Krishnan
Rendering Techniques2
2006 Fast separation of direct and global components of a scene using high frequency illumination
abstract
We present fast methods for separating the direct and global illumination components of a scene measured by a camera and illuminated by a light source. In theory, the separation can be done with just two images taken with a high frequency binary illumination pattern and its complement. In practice, a larger number of images are used to overcome the optical and resolution limitations of the camera and the source. The approach does not require the material properties of objects and media in the scene to be known. However, we require that the illumination frequency is high enough to adequately sample the global components received by scene points. We present separation results for scenes that include complex interreflections, subsurface scattering and volumetric scattering. Several variants of the separation approach are also described. When a sinusoidal illumination pattern is used with different phase shifts, the separation can be done using just three images. When the computed images are of lower resolution than the source and the camera, smoothness constraints are used to perform the separation using a single image. Finally, in the case of a static scene that is lit by a simple point source, such as the sun, a moving occluder and a video camera can be used to do the separation. We also show several simple examples of how novel images of a scene can be computed from the separation results.
Shree K. Nayar, Gurunandan Krishnan, Michael D. Grossberg, Ramesh Raskar
ACM Trans. Graph.2