P. J. Narayanan

dblp:n/PJNarayanan · DBLP profile ↗
← Back
69ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-7164-4917ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 46 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 29 · 4 first-author · 8 since 2021Systems, architecture and hardware · 11 · 3 first-authorDatabases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 STRinGS: Selective Text Refinement in Gaussian Splatting
Abhinav Raundhal, Gaurav Behera, P. J. Narayanan, Ravi Kiran Sarvadevabhatla, Makarand Tapaswi
WACV3
2025 Prototype Guided Backdoor Defense via Activation Space Manipulation
Venkat Adithya Amula, Sunayana Samavedam, Saurabh Saini, Avani Gupta, P. J. Narayanan
ICCV5
2025 Linearly Transformed Spherical Distributions for Interactive Single Scattering with Area Lights
abstract
Abstract Single scattering in scenes with participating media is challenging, especially in the presence of area lights. Considerable variance still remains, in spite of good importance sampling strategies. Analytic methods that render unshadowed surface illumination have recently gained interest since they achieve biased but noise‐free plausible renderings while being computationally efficient. In this work, we extend the theory of Linearly Transformed Spherical Distributions (LTSDs) which is a well‐known analytic method for surface illumination, to work with phase functions. We show that this is non‐trivial, and arrive at a solution with in‐depth analysis. This enables us to analytically compute in‐scattered radiance, which we build on to semi‐analytically render unshadowed single scattering. We ground our derivations and formulations on the Volume Rendering Equation (VRE) which paves the way for realistic renderings despite the biased nature of our method. We also formulate ratio estimators for the VRE to work in conjunction with our formulation, enabling the rendering of shadows. We extensively validate our method, analyze its characteristics and demonstrate better performance compared to Monte Carlo single‐scattering.
Aakash KT, Ishaan Nikhil Shah, P. J. Narayanan
Comput. Graph. Forum3
2025 Fast self-supervised 3D mesh object retrieval for geometric similarity
Kajal Sanklecha, Prayushi Mathur, P. J. Narayanan
Comput. Vis. Image Underst.3
2024 GSN: Generalisable Segmentation in Neural Radiance Field
abstract
Traditional Radiance Field (RF) representations capture details of a specific scene and must be trained afresh on each scene. Semantic feature fields have been added to RFs to facilitate several segmentation tasks. Generalised RF representations learn the principles of view interpolation. A generalised RF can render new views of an unknown and untrained scene, given a few views. We present a way to distil feature fields into the generalised GNT representation. Our GSN representation generates new views of unseen scenes on the fly along with consistent, per-pixel semantic features. This enables multi-view segmentation of arbitrary new scenes. We show different semantic features being distilled into generalised RFs. Our multi-view segmentation results are on par with methods that use traditional RFs. GSN closes the gap between standard and generalisable RF methods significantly. Project Page: https://vinayak-vg.github.io/GSN/
Rahul Goel, Dhawal Sirikonda, P. J. Narayanan
AAAI4
2024 Specularity Factorization for Low-Light Enhancement
abstract
We present a new additive image factorization technique that treats images to be composed of multiple latent specular components which can be simply estimated recursively by modulating the sparsity during decomposition. Our model-driven RSFNet estimates these factors by unrolling the optimization into network layers requiring only a few scalars to be learned. The resultant factors are interpretable by design and can be fused for different image enhancement tasks via a network or combined directly by the user in a controllable fashion. Based on RSFNet, we detail a zero-reference Low Light Enhancement (LLE) application trained without paired or unpaired supervision. Our system improves the state-of-the-art performance on standard benchmarks and achieves better generalization on multiple other datasets. We also integrate our factors with other task specific fusion networks for applications like deraining, deblurring and dehazlng with negligible overhead thereby highlighting the multi-domain and multi-task generalizability of our proposed RSFNet. The code and data is released for reproducibility on the project homepage11https://sophont01.github.io/data/projects/RSFNet/.
Saurabh Saini, P. J. Narayanan
CVPR2
2024 Neural Histogram-Based Glint Rendering of Surfaces With Spatially Varying Roughness
abstract
Abstract The complex, glinty appearance of detailed normal‐mapped surfaces at different scales requires expensive per‐pixel Normal Distribution Function computations. Moreover, large light sources further compound this integration and increase the noise in the Monte Carlo renderer. Specialized rendering techniques that explicitly express the underlying normal distribution have been developed to improve performance for glinty surfaces controlled by a fixed material roughness. We present a new method that supports spatially varying roughness based on a neural histogram that computes per‐pixel NDFs with arbitrary positions and sizes. Our representation is both memory and compute efficient. Additionally, we fully integrate direct illumination for all light directions in constant time. Our approach decouples roughness and normal distribution, allowing the live editing of the spatially varying roughness of complex normal‐mapped objects. We demonstrate that our approach improves on previous work by achieving smaller footprints while offering GPU‐friendly computation and compact representation.
Ishaan Nikhil Shah, Luis E. Gamboa, Adrien Gruson, P. J. Narayanan
Comput. Graph. Forum4
2023 Interactive Segmentation of Radiance Fields
abstract
Radiance Fields (RF) are popular to represent casually-captured scenes for new view synthesis and several applications beyond it. Mixed reality on personal spaces needs understanding and manipulating scenes represented as RFs, with semantic segmentation of objects as an important step. Prior segmentation efforts show promise but don't scale to complex objects with diverse appearance. We present the ISRF method to interactively segment objects with fine structure and appearance. Nearest neighbor feature matching using distilled semantic features identifies high-confidence seed regions. Bilateral search in a joint spatio-semantic space grows the region to recover accurate segmentation. We show state-of-the-art results of segmenting objects from RFs and compositing them to another scene, changing appearance, etc., and an interactive segmentation tool that others can use.
Rahul Goel, Dhawal Sirikonda, Saurabh Saini, P. J. Narayanan
CVPR4
2023 Concept Distillation: Leveraging Human-Centered Explanations for Model Improvement
abstract
Humans use abstract *concepts* for understanding instead of hard features. Recent interpretability research has focused on human-centered concept explanations of neural networks. Concept Activation Vectors (CAVs) estimate a model's sensitivity and possible biases to a given concept. We extend CAVs from post-hoc analysis to ante-hoc training to reduce model bias through fine-tuning using an additional *Concept Loss*. Concepts are defined on the final layer of the network in the past. We generalize it to intermediate layers, including the last convolution layer. We also introduce *Concept Distillation*, a method to define rich and effective concepts using a pre-trained knowledgeable model as the teacher. Our method can sensitize or desensitize a model towards concepts. We show applications of concept-sensitive training to debias several classification problems. We also show a way to induce prior knowledge into a reconstruction problem. We show that concept-sensitive training can improve model interpretability, reduce biases, and induce prior knowledge.
Avani Gupta, Saurabh Saini, P. J. Narayanan
NeurIPS3
2023 Accelerating Hair Rendering by Learning High-Order Scattered Radiance
abstract
Abstract Efficiently and accurately rendering hair accounting for multiple scattering is a challenging open problem. Path tracing in hair takes long to converge while other techniques are either too approximate while still being computationally expensive or make assumptions about the scene. We present a technique to infer the higher order scattering in hair in constant time within the path tracing framework, while achieving better computational efficiency. Our method makes no assumptions about the scene and provides control over the renderer's bias & speedup. We achieve this by training a small multilayer perceptron (MLP) to learn the higher‐order radiance online, while rendering progresses. We describe how to robustly train this network and thoroughly analyze our resulting renderer's characteristics. We evaluate our method on various hairstyles and lighting conditions. We also compare our method against a recent learning based & a traditional real‐time hair rendering method and demonstrate better quantitative & qualitative results. Our method achieves a significant improvement in speed with respect to path tracing, achieving a run‐time reduction of 40%‐70% while only introducing a small amount of bias.
Aakash KT, Adrián Jarabo, Carlos Aliaga, Matt Jen-Yuan Chiang, Olivier Maury, Christophe Hery, P. J. Narayanan, Giljoo Nam
Comput. Graph. Forum7
2023 SHARP: Shape-Aware Reconstruction of People in Loose Clothing
Sai Sagar Jinka, Astitva Srivastava, Chandradeep Pokhariya, Avinash Sharma 0001, P. J. Narayanan
Int. J. Comput. Vis.5
2022 Casual Indoor HDR Radiance Capture from Omnidirectional Images
Pulkit Gera, Mohammad Reza Karimi Dastjerdi, Charles Renaud, P. J. Narayanan, Jean-François Lalonde
BMVC4
2020 PeeledHuman: Robust Shape Representation for Textured 3D Human Body Reconstruction
abstract
We introduce PeeledHuman - a novel shape representation of the human body that is robust to self-occlusions. PeeledHuman encodes the human body as a set of Peeled Depth and RGB maps in 2D, obtained by performing raytracing on the 3D body model and extending each ray beyond its first intersection. This formulation allows us to handle self-occlusions efficiently compared to other representations. Given a monocular RGB image, we learn these Peeled maps in an end-to-end generative adversarial fashion using our novel framework - PeelGAN. We train PeelGAN using a 3D Chamfer loss and other 2D losses to generate multiple depth values per-pixel and a corresponding RGB field per-vertex in a dual-branch setup. In our simple non-parametric solution, the generated Peeled Depth maps are back-projected to 3D space to obtain a complete textured 3D shape. The corresponding RGB maps provide vertex-level texture details. We compare our method with current parametric and non-parametric methods in 3D reconstruction and find that we achieve state-of-the-art-results. We demonstrate the effectiveness of our representation on publicly available BUFF and MonoPerfCap datasets as well as loose clothing data collected by our calibrated multi-Kinect setup.
Sai Sagar Jinka, Rohan Chacko, Avinash Sharma 0001, P. J. Narayanan
3DV4
2020 Unsupervised Image Style Embeddings for Retrieval and Recognition Tasks
abstract
We propose an unsupervised protocol for learning a neural embedding of visual style of images. Style similarity is an important measure for many applications such as style transfer, fashion search, art exploration, etc. However, computational modeling of style is a difficult task owing to its vague and subjective nature. Most methods for style based retrieval use supervised training with pre-defined categorization of images according to style. While this paradigm is suitable for applications where style categories are well-defined and curating large datasets according to such a categorization is feasible, in several other cases such a categorization is either ill-defined or does not exist. Our protocol for learning style based representations does not leverage categorical labels but a proxy measure for forming triplets of anchor, similar, and dissimilar images. Using these triplets, we learn a compact style embedding that is useful for style-based search and retrieval. The learned embeddings outperform other unsupervised representations for style-based image retrieval task on six datasets that capture different meanings of style. We also show that by fine-tuning the learned features with dataset-specific style labels, we obtain best results for image style recognition task on five of the six datasets.
Siddhartha Gairola, Rajvi Shah, P. J. Narayanan
WACV3
2019 Nose, Eyes and Ears: Head Pose Estimation by Locating Facial Keypoints
abstract
Monocular head pose estimation requires learning a model that computes the intrinsic Euler angles for pose (yaw, pitch, roll) from an input image of human face. Annotating ground truth head pose angles for images in the wild is difficult and requires ad-hoc fitting procedures (which provides only coarse and approximate annotations). This highlights the need for approaches which can train on data captured in controlled environment and generalize on the images in the wild (with varying appearance and illumination of the face). Most present day deep learning approaches which learn a regression function directly on the input images fail to do so. To this end, we propose to use a higher level representation to regress the head pose while using deep learning architectures. More specifically, we use the uncertainty maps in the form of 2D soft localization heatmap images over five facial key-points, namely left ear, right ear, left eye, right eye and nose, and pass them through an convolutional neural network to regress the head-pose. We show head pose estimation results on two challenging benchmarks BIWI and AFLW and our approach surpasses the state of the art on both the datasets.
Aryaman Gupta, Kalpit C. Thakkar, Vineet Gandhi, P. J. Narayanan
ICASSP4
2019 Defocus Magnification Using Conditional Adversarial Networks
abstract
Defocus magnification is the process of rendering a shallow depth-of-field in an image captured using a camera with a narrow aperture. Defocus magnification is a useful tool in photography for emphasis on the subject and for highlighting background bokeh. Estimating the per-pixel blur kernel or the depth-map of the scene followed by spatially-varying re-blurring is the standard approach to defocus magnification. We propose a single-step approach that directly converts a narrow-aperture image to a wide-aperture image. We use a conditional adversarial network trained on multi-aperture images created from light-fields. We use a novel loss term based on a composite focus measure to improve generalization and show high quality defocus magnification.
Parikshit Sakurikar, Ishit Mehta, P. J. Narayanan
WACV3
2018 Structured Adversarial Training for Unsupervised Monocular Depth Estimation
abstract
The problem of estimating scene-depth from a single image has seen great progress lately. Recent unsupervised methods are based on view-synthesis and learn depth by minimizing photometric reconstruction error. In this paper, we introduce Structured Adversarial Training (StrAT) to this problem. We generate multiple novel views using depth (or disparity), with the stereo-baseline changing in an increasing order. Adversarial training that goes from easy examples to harder ones produces richer losses and better models. The impact of StrAT is shown to exceed traditional data augmentation using random new views. The combination of an adversarial framework, multiview learning, and structured adversarial training produces state-of-the-art performance on unsupervised depth estimation for monocular images. The StrAT framework can benefit several problems that use adversarial training.
Ishit Mehta, Parikshit Sakurikar, P. J. Narayanan
3DV3
2018 Semantic Priors for Intrinsic Image Decomposition
Saurabh Saini, P. J. Narayanan
BMVC2
2018 Part-based Graph Convolutional Network for Action Recognition
Kalpit C. Thakkar, P. J. Narayanan
BMVC2
2018 RefocusGAN: Scene Refocusing Using a Single Image
Parikshit Sakurikar, Ishit Mehta, Vineeth N. Balasubramanian, P. J. Narayanan
ECCV (4)4
2018 View-Graph Selection Framework for SfM
Rajvi Shah, Visesh Chari, P. J. Narayanan
ECCV (5)3
2018 Find Me a Sky: A Data-Driven Method for Color-Consistent Sky Search and Replacement
Saumya Rawat, Siddhartha Gairola, Rajvi Shah, P. J. Narayanan
MMM (1)4
2018 Human Shape Capture and Tracking at Home
abstract
Human body tracking typically requires specialized capture set-ups. Although pose tracking is available in consumer devices like Microsoft Kinect, it is restricted to stick figures visualizing body part detection. In this paper, we propose a method for full 3D human body shape and motion capture of arbitrary movements from the depth channel of a single Kinect, when the subject wears casual clothes. We do not use the RGB channel or an initialization procedure that requires the subject to move around in front of the camera. This makes our method applicable for arbitrary clothing textures and lighting environments, with minimal subject intervention. Our method consists of 3D surface feature detection and articulated motion tracking, which is regularized by a statistical human body model [26]. We also propose the idea of a Consensus Mesh (CMesh) which is the 3D template of a person created from a single view point. We demonstrate tracking results on challenging poses and argue that using CMesh along with statistical body models can improve tracking accuracies. Quantitative evaluation of our dense body tracking shows that our method has very little drift which is improved by the usage of CMesh.
Saurabh Saini, Kiran Varanasi, P. J. Narayanan
WACV4
2017 Composite Focus Measure for High Quality Depth Maps
abstract
Depth from focus is a highly accessible method to estimate the 3D structure of everyday scenes. Today's DSLR and mobile cameras facilitate the easy capture of multiple focused images of a scene. Focus measures (FMs) that estimate the amount of focus at each pixel form the basis of depth-from-focus methods. Several FMs have been proposed in the past and new ones will emerge in the future, each with their own strengths. We estimate a weighted combination of standard FMs that outperforms others on a wide range of scene types. The resulting composite focus measure consists of FMs that are in consensus with one another but not in chorus. Our two-stage pipeline first estimates fine depth at each pixel using the composite focus measure. A cost-volume propagation step then assigns depths from confident pixels to others. We can generate high quality depth maps using just the top five FMs from our composite focus measure. This is a positive step towards depth estimation of everyday scenes with no special equipment.
Parikshit Sakurikar, P. J. Narayanan
ICCV2
2017 SynCam: Capturing sub-frame synchronous media using smartphones
abstract
Smartphones have become the de-facto capture devices for everyday photography. Unlike traditional digital cameras, smartphones are versatile devices with auxiliary sensors, processing power, and networking capabilities. In this work, we harness the communication capabilities of smartphones and present a synchronous/co-ordinated multi-camera capture system. Synchronous capture is important for many image/video fusion and 3D reconstruction applications. The proposed system provides an inexpensive and effective means to capture multi-camera media for such applications. Our coordinated capture system is based on a wireless protocol that uses NTP based synchronization and device specific lag compensation. It achieves sub-frame synchronization across all participating smartphones of even heterogeneous make and model. We propose a new method based on fiducial markers displayed on an LCD screen to temporally calibrate smart-phone cameras. We demonstrate the utility and versatility of this system to enhance traditional videography and to create novel visual representations such as panoramic videos, HDR videos, multi-view 3D reconstruction, multi-flash imaging, and multi-camera social media.
Ishit Mehta, Parikshit Sakurikar, Rajvi Shah, P. J. Narayanan
ICME4
2015 Geometry-Aware Feature Matching for Structure from Motion Applications
abstract
We present a two-stage, geometry-aware approach for matching SIFT-like features in a fast and reliable manner. Our approach first uses a small sample of features to estimate the epipolar geometry between the images and leverages it for guided matching of the remaining features. This simple and generalized two-stage matching approach produces denser feature correspondences while allowing us to formulate an accelerated search strategy to gain significant speedup over the traditional matching. The traditional matching punitively rejects many true feature matches due to a global ratio test. The adverse effect of this is particularly visible when matching image pairs with repetitive structures. The geometry-aware approach prevents such pre-emptive rejection using a selective ratio-test and works effectively even on scenes with repetitive structures. We also show that the proposed algorithm is easy to parallelize and implement it on the GPU. We experimentally validate our algorithm on publicly available datasets and compare the results with state-of-the-art methods.
Rajvi Shah, Vanshika Srivastava, P. J. Narayanan
WACV3
2014 Multistage SFM: Revisiting Incremental Structure from Motion
abstract
In this paper, we present a new multistage approach for SfM reconstruction of a single component. Our method begins with building a coarse 3D reconstruction using high-scale features of given images. This step uses only a fraction of features and is fast. We enrich the model in stages by localizing remaining images to it and matching and triangulating remaining features. Unlike traditional incremental SfM, localization and triangulation steps in our approach are made efficient and embarrassingly parallel using geometry of the coarse model. The coarse model allows us to use 3D-2D correspondences based direct localization techniques to register remaining images. We further utilize the geometry of the coarse model to reduce the pair-wise image matching effort as well as to perform fast guided feature matching for majority of features. Our method produces similar quality models as compared to incremental SfM methods while being notably fast and parallel. Our algorithm can reconstruct a 1000 images dataset in 15 hours using a single core, in about 2 hours using 8 cores and in a few minutes by utilizing full parallelism of about 200 cores.
Rajvi Shah, Aditya Deshpande, P. J. Narayanan
3DV3
2013 Can GPUs sort strings efficiently?
abstract
String sorting or variable-length key sorting has lagged in performance on the GPU even as the fixed-length key sorting has improved dramatically. Radix sorting is the fastest on the GPUs. In this paper, we present a fast and efficient string sort on the GPU that is built on the available radix sort. Our method sorts strings from left to right in steps, moving only indexes and small prefixes for efficiency. We reduce the number of sort steps by adaptively consuming maximum string bytes based on the number of segments in each step. Performance is improved by using Thrust primitives for most steps and by removing singleton segments from consideration. Over 70% of the string sort time is spent on Thrust primitives. This provides high performance along with high adaptability to future GPUs. We achieve speed of up to 10 over current GPU methods, especially on large datasets. We also scale to much larger input sizes. We present results on easy and difficult strings defined using their after-sort tie lengths.
Aditya Deshpande, P. J. Narayanan
HiPC2
2013 Interactive Video Manipulation Using Object Trajectories and Scene Backgrounds
abstract
Traditional video editing interfaces model and represent videos as a collection of frames against a timeline, which makes the object-centric manipulation of videos a laborious task. We enable a simple and meaningful interaction for object-centric navigation and manipulation of long shot videos by introducing operators on three high-level video semantics: background mosaics, object motions, and camera motions. We estimate the scene background and represent the object motion using 3-D space-time trajectories. We use the 3-D object trajectories as basic interaction elements, and define several object and camera operations as simple and intuitive curve manipulations. These allow users to perform various video object temporal manipulations by interactively manipulating the object trajectories. The camera operations model the camera as a movable and scalable aperture and allow the users to simulate pan, tilt, and zoom effects by creating new camera trajectories. With several example compositions, we demonstrate that our representation and operations allow users to simply and interactively perform numerous seemingly complex, high-level video manipulation tasks.
Rajvi Shah, P. J. Narayanan
IEEE Trans. Circuits Syst. Video Technol.2
2013 Designing Perspectively Correct Multiplanar Displays
abstract
Displays remain flat and passive amidst the many changes in their fundamental technologies. One natural step ahead is to create displays that merge seamlessly in shape and appearance with one’s natural surroundings. In this paper, we present a system to design, render to, and build view-dependent multiplanar displays of arbitrary piecewise-planar shapes, built using polygonal facets. Our system provides high quality, interactive rendering of 3D environments to a head-tracked viewer on arbitrary multiplanar displays. We develop a novel rendering scheme that produces exact image and depth map at each facet, producing artifact-free images on and across facet boundaries. The system scales to a large number of display facets by rendering all facets in a single pass of rasterization. This is achieved using a parallel, perframe, view-dependent binning and prewarping of scene triangles. The display is driven using one or more target quilt images into which facet pixels are packed. Our method places no constraints on the scene or the display and allows for fully dynamic scenes to be rendered interactively at high resolutions. The steps of our system are implemented efficiently on commodity GPUs. We present a few prototype displays to establish the scalability of our system on different display shapes, form factors, and complexity: from a cube made out of LCD panels to spherical/cylindrical projected setups to arbitrary complex shapes in simulation. Performance of our system is demonstrated for both rendering quality and speed, for increasing scene and display facet sizes. A subjective user study is also presented to evaluate the user experience using a walk-around display compared to a flat panel for a game-like setting.
Pawan Harish, P. J. Narayanan
IEEE Trans. Vis. Comput. Graph.2
2012 Visibility Probability Structure from SfM Datasets and Applications
Siddharth Choudhary, P. J. Narayanan
ECCV (5)2
2012 Mixed-Resolution Patch-Matching
Harshit Sureka, P. J. Narayanan
ECCV (6)2
2012 Fast graph cuts using shrink-expand reparameterization
abstract
Global optimization of MRF energy using graph cuts is widely used in computer vision. As the images are getting larger, faster graph cuts are needed without sacrificing optimality. Initializing or reparameterizing a graph using results of a similar one has provided efficiency in the past. In this paper, we present a method to speedup graph cuts using shrink-expand reparameterization. Our scheme merges the nodes of a given graph to shrink it. The resulting graph and its mincut are expanded and used to reparameterize the original graph for faster convergence. Graph shrinking can be done in different ways. We use a block-wise shrinking similar to multiresolution processing of images in our Multiresolution Cuts algorithm. We also develop a hybrid approach that can mix nodes from different levels without affecting optimality. Our algorithm is particularly suited for processing large images. The processing time on the full detail graph reduces nearly by a factor of 4. The overall application time including all book-keeping is faster by a factor of 2 on various types of images.
Parikshit Sakurikar, P. J. Narayanan
WACV2
2012 Raytracing Dynamic Scenes on the GPU Using Grids
abstract
Raytracing dynamic scenes at interactive rates have received a lot of attention recently. We present a few strategies for high performance raytracing on a commodity GPU. The construction of grids needs sorting, which is fast on today's GPUs. The grid is thus the acceleration structure of choice for dynamic scenes as per-frame rebuilding is required. We advocate the use of appropriate data structures for each stage of raytracing, resulting in multiple structure building per frame. A perspective grid built for the camera achieves perfect coherence for primary rays. A perspective grid built with respect to each light source provides the best performance for shadow rays. Spherical grids handle lights positioned inside the model space and handle spotlights. Uniform grids are best for reflection and refraction rays with little coherence. We propose an Enforced Coherence method to bring coherence to them by rearranging the ray to voxel mapping using sorting. This gives the best performance on GPUs with only user-managed caches. We also propose a simple, Independent Voxel Walk method, which performs best by taking advantage of the L1 and L2 caches on recent GPUs. We achieve over 10 fps of total rendering on the Conference model with one light source and one reflection bounce, while rebuilding the data structure for each stage. Ideas presented here are likely to give high performance on the future GPUs as well as other manycore architectures.
Sashidhar Guntury, P. J. Narayanan
IEEE Trans. Vis. Comput. Graph.2
2011 Hybrid implementation of error diffusion dithering
abstract
Many image filtering operations provide ample parallelism, but progressive non-linear processing of images is among the hardest to parallelize due to long, sequential, and non-linear data dependency. A typical example of such an operation is error diffusion dithering, exemplified by the Floyd-Steinberg algorithm. In this paper, we present its parallelization on multicore CPUs using a block-based approach and on the GPU using a pixel based approach. We also present a hybrid approach in which the CPU and the GPU operate in parallel during the computation. High Performance Computing has traditionally been associated with high end CPUs and GPUs. Our focus is on everyday computers such as laptops and desktops, where significant compute power is available on the GPU as on the CPU. Our implementation can dither an 8K × 8K image on an off-the-shelf laptop with an Nvidia 8600M GPU in about 400 milliseconds when the sequential implementation on its CPU took about 4 seconds.
Aditya Deshpande, Ishan Misra, P. J. Narayanan
HiPC3
2011 Scalable clustering using multiple GPUs
abstract
K-Means is a popular clustering algorithm with wide applications in Computer Vision, Data mining, Data Visualization, etc. Clustering is an important step for indexing and searching of documents, images, video, etc. Clustering large numbers of high-dimensional vectors is very computation intensive. In this paper, we present the design and implementation of the K-Means clustering algorithm on the modern GPU. All steps are performed entirely on the GPU efficiently in our approach. We also present a load balanced multi-node, multi-GPU implementation which can handle up to 6 million, 128-dimensional vectors. We use efficient memory layout for all steps to get high performance. The GPU accelerators are now present on high-end workstations and low-end laptops. Scalability in the number and dimensionality of the vectors, the number of clusters, as well as in the number of cores available for processing are important for usability to different users. Our implementation scales linearly or near-linearly with different problem parameters. We achieve up to 2 times increase in speed compared to the best GPU implementation for K-Means on a single GPU. We obtain a speed up of over 170 on a single Nvidia Fermi GPU compared to a standard sequential implementation. We are able to execute one iteration of K-Means in 136 seconds on off-the-shelf GPUs to cluster 6 million vectors of 128 dimensions into 4K clusters and in 2.5 seconds to cluster 125K vectors of 128 dimensions into 2K clusters.
K. Wasif Mohiuddin, P. J. Narayanan
HiPC2
2011 Trajectory based video object manipulation
abstract
We propose an object centric representation for easy and intuitive navigation and manipulation of videos. Object centric representation allows a user to directly access and process objects as basic video components. We demonstrate a trajectory based interface and example operations, which allow users to retime, reorder, remove or clone video objects in a 'click and drag' fashion. This interface is created by extracting object motion information from the video. We use object detection and tracking to obtain spatiotemporal video object tube. The corresponding object motion trajectories are represented in a 3D (x, y, t) grid. Users can navigate and manipulate video objects by scrubbing or manipulating corresponding trajectories. We show some example applications of proposed inter face like object synchronization, saliency magnification, visual effects and composite video creation.
Rajvi Shah, P. J. Narayanan
ICME2
2011 Person De-Identification in Videos
abstract
Advances in cameras and web technology have made it easy to capture and share large amounts of video data over to a large number of people. A large number of cameras oversee public and semi-public spaces today. These raise concerns on the unintentional and unwarranted invasion of the privacy of individuals caught in the videos. To address these concerns, automated methods to de-identify individuals in these videos are necessary. De-identification does not aim at destroying all information involving the individuals. Its ideal goals are to obscure the identity of the actor without obscuring the action. This paper outlines the scenarios in which de-identification is required and the issues brought out by those. We also present an approach to de-identify individuals from videos. Our approach involves tracking and segmenting individuals in a conservative voxel space involving x, y , and time. A de-identification transformation is applied per frame using these voxels to obscure the identity. Face, silhouette, gait, and other characteristics need to be obscured, ideally. We show results of our scheme on a number of videos and for several variations of the transformations. We present the results of applying algorithmic identification on the transformed videos. We also present the results of a user-study to evaluate how well humans can identify individuals from the transformed videos.
Prachi Agrawal, P. J. Narayanan
IEEE Trans. Circuits Syst. Video Technol.2
2011 Editorial introduction to the special issue
Sharat Chandran, P. J. Narayanan, A. N. Rajagopalan 0001
Vis. Comput.2
2010 Efficient Discrete Range Searching primitives on the GPU with applications
abstract
Graphics processing units provide a large computational power at a very low price which position them as an ubiquitous accelerator. Efficient primitives that can expand the r ange of operations performed on the GPU are thus important. Discrete Range Searching(DRS) is one such primitive with direct applications to string processing, document and text retrieval systems, and least common ancestor queries. In this work, we present a GPU specific implementation of DRS with an optimal space-time trade off. Toward this end, we also present GPU amenable succinct representations and discuss limitations on the GPU. Our method uses 7.5 bits of additional space per element. The speedup achieved by our method is in the range of 20-25 for preprocessing, and 25-35 for batch querying over a sequential implementation. Compared to an 8-threaded implementation, our methods obtain a speedup of 6-8. We study applications of the DRS on the GPU. Also, we suggest that most graph algorithms which focus on using least common ancestor, can easily be enabled on the GPU based on range minima primitive. Beyond this, we show applications of DRS in string querying and tree queries, and suggest how DRS can be helpful in implementing tree based graph algorithms on the GPU.
Jyothish Soman, Kiran Kumar Matam, Kishore Kothapalli, P. J. Narayanan
HiPC4
2010 Real-Time Ray Tracing of Implicit Surfaces on the GPU
abstract
Compact representation of geometry using a suitable procedural or mathematical model and a ray-tracing mode of rendering fit the programmable graphics processor units (GPUs) well. Several such representations including parametric and subdivision surfaces have been explored in recent research. The important and widely applicable category of the general implicit surface has received less attention. In this paper, we present a ray-tracing procedure to render general implicit surfaces efficiently on the GPU. Though only the fourth or lower order surfaces can be rendered using analytical roots, our adaptive marching points algorithm can ray trace arbitrary implicit surfaces without multiple roots, by sampling the ray at selected points till a root is found. Adapting the sampling step size based on a proximity measure and a horizon measure delivers high speed. The sign test can handle any surface without multiple roots. The Taylor test that uses ideas from interval analysis can ray trace many surfaces with complex roots. Overall, a simple algorithm that fits the SIMD architecture of the GPU results in high performance. We demonstrate the ray tracing of algebraic surfaces up to order 50 and nonalgebraic surfaces including a Blinn's blobby with 75 spheres at better than interactive frame rates.
Jag Mohan Singh, P. J. Narayanan
IEEE Trans. Vis. Comput. Graph.2
2009 Person De-identification in Videos
Prachi Agrawal, P. J. Narayanan
ACCV (3)2
2009 Solving Multilabel MRFs Using Incremental alpha-Expansion on the GPUs
Vibhav Vineet, P. J. Narayanan
ACCV (3)2
2009 A performance prediction model for the CUDA GPGPU platform
abstract
The significant growth in computational power of modern Graphics Processing Units (GPUs) coupled with the advent of general purpose programming environments like NVIDIA's CUDA, has seen GPUs emerging as a very popular parallel computing platform. Till recently, there has not been a performance model for GPGPUs. The absence of such a model makes it difficult to definitively assess the suitability of the GPU for solving a particular problem and is a significant impediment to the mainstream adoption of GPUs as a massively parallel (super)computing platform. In this paper we present a performance prediction model for the CUDA GPGPU platform. This model encompasses the various facets of the GPU architecture like scheduling, memory hierarchy, and pipelining among others. We also perform experiments that demonstrate the effects of various memory access strategies. The proposed model can be used to analyze pseudo code for a CUDA kernel to obtain a performance estimate, in a way that is similar to performing asymptotic analysis. We illustrate the usage of our model and its accuracy with three case studies: matrix multiplication, list ranking, and histogram generation.
Kishore Kothapalli, Rishabh Mukherjee, M. Suhail Rehman, Suryakant Patidar, P. J. Narayanan, K. Srinathan 0001
HiPC5
2009 Fast and scalable list ranking on the GPU
abstract
General purpose programming on the graphics processing units (GPGPU) has received a lot of attention in the parallel computing community as it promises to offer the highest performance per dollar. The GPUs have been used extensively on regular problems that can be easily parallelized. In this paper, we describe two implementations of List Ranking, a traditional irregular algorithm that is difficult to parallelize on such massively multi-threaded hardware. We first present an implementation of Wyllie's algorithm based on pointer jumping. This technique does not scale well to large lists due to the suboptimal work done. We then present a GPU-optimized, Recursive Helman-JaJa (RHJ) algorithm. Our RHJ implementation can rank a random list of 32 million elements in about a second and achieves a speedup of about 8-9 over a CPU implementation as well as a speedup of 3-4 over the best reported implementation on the Cell Broadband engine. We also discuss the practical issues relating to the implementation of irregular algorithms on massively multi-threaded architectures like that of the GPU. Regular or coalesced memory accesses pattern and balanced load are critical to achieve good performance on the GPU.
M. Suhail Rehman, Kishore Kothapalli, P. J. Narayanan
ICS3
2009 Singular value decomposition on GPU using CUDA
abstract
Linear algebra algorithms are fundamental to many computing applications. Modern GPUs are suited for many general purpose processing tasks and have emerged as inexpensive high performance co-processors due to their tremendous computing power. In this paper, we present the implementation of singular value decomposition (SVD) of a dense matrix on GPU using the CUDA programming model. SVD is implemented using the twin steps of bidiagonalization followed by diagonalization. It has not been implemented on the GPU before. Bidiagonalization is implemented using a series of householder transformations which map well to BLAS operations. Diagonalization is performed by applying the implicitly shifted QR algorithm. Our complete SVD implementation outperforms the Matlab and Intel regMath kernel library (MKL) LAPACK implementation significantly on the CPU. We show a speedup of upto 60 over the MATLAB implementation and upto 8 over the Intel MKL implementation on a Intel Dual Core 2.66 GHz PC on NVIDIA GTX 280 for large matrices. We also give results for very large matrices on NVIDIA Tesla S1070.
Sheetal Lahabar, P. J. Narayanan
IPDPS2
2008 A Parametric Proxy-Based Compression of Depth Movies
abstract
Depth movies provide 2-D representations of a time-varying 3D scene. Multistream depth movies arise in many structure-capturing setups and have been used for image based rendering. They are bulky and carry redundant information. We propose a proxy-based compression scheme for multistream depth movies of a scene involving dynamic human actors. The input to our system is depth movies from different viewpoints and the calibration parameters to relate depths to 3D points. We use an articulated human model as a proxy to represent the common structure of the scene. The proxy is parametrized by various bone angles.
Pooja Verlani, P. J. Narayanan
DCC2
2007 Accelerating Large Graph Algorithms on the GPU Using CUDA
Pawan Harish, P. J. Narayanan
HiPC2
2007 On Using Classical Poetry Structure for Indian Language Post-Processing
abstract
Post-processors are critical to the performance of language recognizers like OCRs, speech recognizers, etc. Dictionary-based post-processing commonly employ either an algorithmic approach or a statistical approach. Other linguistic features are not exploited for this purpose. The language analysis is also largely limited to the prose form. This paper proposes a framework to use the rich metric and formal structure of classical poetic forms in Indian languages for post-processing a recognizer like an OCR engine. We show that the structure present in the form of the vrtta and prasa can be efficiently used to disambiguate some cases that may be difficult for an OCR. The approach is efficient, and complementary to other post-processing approaches and can be used in conjunction with them.
Anoop M. Namboodiri, P. J. Narayanan, C. V. Jawahar
ICDAR2
2007 A Vision System for Monitoring Intermodal Freight Trains
abstract
We describe the design and implementation of a vision based Intermodal Train Monitoring System (ITMS) for extracting various features like length of gaps in an intermodal (IM) train which can later be used for higher level inferences. An intermodal train is a freight train consisting of two basic types of loads - containers and trailers. Our system first captures the video of an IM train, and applies image processing and machine learning techniques developed in this work to identify the various types of loads as containers and trailers. The whole process relies on a sequence of following tasks -robust background subtraction in each frame of the video, estimation of train velocity, creation of mosaic of the whole train from the video and classification of train loads into containers and trailers. Finally, the length of gaps between the loads of the IM train is estimated and is used to analyze the aerodynamic efficiency of the loading pattern of the train, which is a critical aspect of freight trains. This paper focusses on the machine vision aspect of the whole system
Avinash Kumar 0001, Narendra Ahuja, John M. Hart, Visesh Chari, P. J. Narayanan, C. V. Jawahar
WACV5
2007 Garuda: A Scalable Tiled Display Wall Using Commodity PCs
abstract
Cluster-based tiled display walls can provide cost-effective and scalable displays with high resolution and a large display area. The software to drive them needs to scale too if arbitrarily large displays are to be built. Chromium is a popular software API used to construct such displays. Chromium transparently renders any OpenGL application to a tiled display by partitioning and sending individual OpenGL primitives to each client per frame. Visualization applications often deal with massive geometric data with millions of primitives. Transmitting them every frame results in huge network requirements that adversely affect the scalability of the system. In this paper, we present Garuda, a client-server-based display wall framework that uses off-the-shelf hardware and a standard network. Garuda is scalable to large tile configurations and massive environments. It can transparently render any application built using the Open Scene Graph (OSG) API to a tiled display without any modification by the user. The Garuda server uses an object-based scene structure represented using a scene graph. The server determines the objects visible to each display tile using a novel adaptive algorithm that culls the scene graph to a hierarchy of frustums. Required parts of the scene graph are transmitted to the clients, which cache them to exploit the interframe redundancy. A multicast-based protocol is used to transmit the geometry to exploit the spatial redundancy present in tiled display systems. A geometry push philosophy from the server helps keep the clients in sync with one another. Neither the server nor a client needs to render the entire scene, making the system suitable for interactive rendering of massive models. Transparent rendering is achieved by intercepting the cull, draw, and swap functions of OSG and replacing them with our own. We demonstrate the performance and scalability of the Garuda system for different configurations of display wall. We also show that the server and network loads grow sublinearly with the increase in the number of tiles, which makes our scheme suitable to construct very large displays.
Nirnimesh, Pawan Harish, P. J. Narayanan
IEEE Trans. Vis. Comput. Graph.3
2005 Compression of multiple depth maps for IBR
Sashi Kumar Penta, P. J. Narayanan
Vis. Comput.2
2004 Constraints on Coplanar Moving Points
Sujit Kuthirummal, C. V. Jawahar, P. J. Narayanan
ECCV (4)3
2004 Building blocks for autonomous navigation using contour correspondences
abstract
We address a few problems in navigation of automated vehicles using images captured by a mounted camera. Specifically, we look at the recognition of sign boards, rectification of planar objects imaged by the camera and estimation of the position of a vehicle with respect to a fixed sign board. Our solutions are based on contour correspondence between a reference view and the current view. A mapping between corresponding points of a planar object in two different views is a 3/spl times/3 matrix called the homography. A novel two-step linear algorithm for homography calculation from contour correspondence is developed first. Our algorithm requires the identification of an image contour as the projections of a known planar world contour and the selection of a known starting point. The homography between the reference view and the target view is applied to several real-life navigation applications, results of which are presented in this paper.
M. Pawan Kumar, C. V. Jawahar, P. J. Narayanan
ICIP3
2004 Discrete contours in multiple views: approximation and recognition
M. Pawan Kumar, Saurabh Goyal, Sujit Kuthirummal, C. V. Jawahar, P. J. Narayanan
Image Vis. Comput.5
2004 Fourier domain representation of planar curves for recognition in multiple views
Sujit Kuthirummal, C. V. Jawahar, P. J. Narayanan
Pattern Recognit.3
2002 Video frame alignment in multiple views
abstract
Many events are captured using multiple cameras today. Frames of each video stream have to be synchronized and aligned to a common time axis before processing them. Synchronization of the video streams necessarily needs a hardware based solution that is applied while capturing. The alignment problem between the frames of multiple videos can be posed as a search using traditional measures for image similarity. Multiview relations and constraints developed in Computer Vision recently can provide more elegant solutions to this problem. In this paper, we provide two solutions for the video frame alignment problem using two view and three view constraints. We present solutions to this problem for the case when the videos are taken using affine cameras and for general projective cameras. Excellent experimental results are achieved by our algorithms.
Sujit Kuthirummal, C. V. Jawahar, P. J. Narayanan
ICIP (3)3
2002 Generalised correlation for multi-feature correspondence
C. V. Jawahar, P. J. Narayanan
Pattern Recognit.2
2002 An adaptive multifeature correspondence algorithm for stereo using dynamic programming
C. V. Jawahar, P. J. Narayanan
Pattern Recognit. Lett.2
1998 Constructing Virtual Worlds Using Dense Stereo
abstract
We present Virtualized Reality, a technique to create virtual worlds out of dynamic events using densely distributed stereo views. The intensity image and depth map for each camera view at each time instant are combined to form a Visible Surface Model. Immersive interaction with the virtualized event is possible using a dense collection of such models. Additionally, a Complete Surface Model of each instant can be built by merging the depth maps from different cameras into a common volumetric space. The corresponding model is compatible with traditional virtual models and can be interacted with immersively using standard tools. Because both VSMs and CSMs are fully three-dimensional, virtualized models can also be combined and modified to build larger, more complex environments, an important capability for many non-trivial applications. We present results from 3D Dome, our facility to create virtualized models.
P. J. Narayanan, Peter Rander, Takeo Kanade
ICCV1
1997 Virtualized reality: constructing time-varying virtual worlds from real world events
abstract
Virtualized reality is a modeling technique that constructs full 3D virtual representations of dynamic events from multiple video streams. Image-based stereo is used to compute a range image corresponding to each intensity image in each video stream. Each range and intensity image pair encodes the scene structure and appearance of the scene visible to the camera at that moment, and is therefore called a visible surface model (VSM). A single time instant of the dynamic event can be modeled as a collection of VSMs from different viewpoints, and the full event can be modeled as a sequence of static scenes-the 3D equivalent of video. Alternatively, the collection of VSMs at a single time can be fused into a global 3D surface model, thus creating a traditional virtual representation out of real world events. Global modeling has the added benefit of eliminating the need to hand-edit the range images to correct errors made in stereo, a drawback of previous techniques. Like image-based rendering models, these virtual representations can be used to synthesize nearly any view of the virtualized event. For this reason, the paper includes a detailed comparison of existing view synthesis techniques with the authors' own approach. In the virtualized representations, however, scene structure is explicitly represented and therefore easily manipulated, for example by adding virtual objects to (or removing virtualized objects from) the model without interfering with real event. Virtualized reality, then, is a platform not only for image-based rendering but also for 3D scene manipulation.
Peter Rander, P. J. Narayanan, Takeo Kanade
IEEE Visualization2
1994 Parallel search for the interpretation of aerial images
abstract
Abstract In this paper, we present a parallel search scheme for model‐based interpretation of aerial images, following a focus‐of‐attention paradigm. Interpretation is performed using the gray level image of an aerial scene and its segmentation into connected components of almost constant gray level. Candidate objects are generated from the window as connected combinations of its components. Each candidate is matched against the model by checking if the model constraints are satisfied by the parameters computed from the region. The problem of candidate generation and matching is posed as searching in the space of combinations of connected components in the image, with finding an (optimally) successful region as the goal. Our implementation exploits parallelism at multiple levels by parallelizing the management of the open list and other control tasks as well as the task of model matching. We discuss and present the implementation of the interpretation system on a Connection Machine CM‐2. The implementation reported a successful match in a few hundred milliseconds whenever they existed.
P. J. Narayanan, Larry Davis 0001
Concurr. Pract. Exp.1
1993 Processor Autonomy on SIMD Architectures
abstract
Flynn classified high speed (parallel) computers into four categories. Of these, the single instruction stream, multiple data stream (SIMD) processor array machines have become very popular in practical parallel processing. The commercially available processor array machines display important architectural variety, while belonging to SIMD category of machines. In this paper, we further categorize the SIMD class of machines on the basis of processor autonomy of the machines, which is the capability of the individual processing elements (PEs) to act autonomously in some significant way. For each autonomy class, we provide examples and illustrate some of its important algorithmic features. We also discuss how each type of autonomy can be simulated on machines without it. We study the addressing autonomous class of machines in greater detail by discussing three algorithms on machines with and without that type of autonomy. A discussion on how processor autonomy appears in algorithms in the literature and what impact they can have in the future machines also is provided.
P. J. Narayanan
International Conference on Supercomputing1
1992 Rank order filtering on SIMD machines
abstract
Rank order filters form an important class of low level image operations that have widespread applications in image smoothing, texture analysis, etc. In the paper, the authors study several ways of computing rank order filters on processor array architectures. They also present a replicated data algorithm for efficient processing of small images on relatively large processor arrays. Results of implementing the algorithms on a Connection Machine CM-2 and a Mas-Par MP-1 are presented.>
P. J. Narayanan, Larry Davis 0001
ICPR (4)1
1992 Replicated data algorithms in image processing
P. J. Narayanan, Larry Davis 0001
CVGIP Image Underst.1
1992 Replicated Image Algorithms and Their Analyses on SIMD Machines
abstract
Data parallel processing on processor array architectures has gained popularity in data intensive applications, such as image processing and scientific computing, as massively parallel processor array machines became feasible commercially. The data parallel paradigm of assigning one processing element to each data element results in an inefficient utilization of a large processor array when a relatively small data structure is processed on it. The large degree of parallelism of a massively parallel processor array machine does not result in a faster solution to a problem involving relatively small data structures than the modest degree of parallelism of a machine that is just as large as the data structure. We presented data replication technique to speed up the processing of small data structures on large processor arrays. In this paper, we present replicated data algorithms for digital image convolutions and median filtering, and compare their performance with conventional data parallel algorithms for the same on three popular array interconnection networks, namely, the 2-D mesh, the 3-D mesh, and the hypercube.
P. J. Narayanan, Larry Davis 0001
Int. J. Pattern Recognit. Artif. Intell.1
1992 Surface reconstruction (of rough terrain) in range image shadows
Behzad Kamgar-Parsi, P. J. Narayanan, Larry Davis 0001
Pattern Recognit. Lett.2
1991 Analysis of replicated data algorithms on processor array architectures
abstract
Processor array machines have gained popularity in practical parallel processing, particularly in data parallel application areas such as image processing.nected Connection Machine.
P. J. Narayanan
SC1
1990 Connection machine vision-Replicated data structures
abstract
The problem of efficiently processing small data structures on massively parallel single-instruction multiple-data machines using replication methods is discussed. The problem stems from considerations of both multiresolution vision systems and focus of attention vision systems. A general framework for developing replicated algorithms, based on the four steps of embedding, distribution, decomposition, and collection, is described. A simple example is provided based on computing the histogram of a gray-level image. Replicated chain processing is discussed, and an efficient algorithm for ranking the elements in a chain in log (n) time on a concurrent write parallel random access machine is presented.>
Larry Davis 0001, Ling Tony Chen, P. J. Narayanan
ICPR (2)3