EDBT 2026 Demo / reviewers in the wild / expert
Pavan Turaga
dblp:80/1682 · also Pavan K. Turaga
· DBLP profile ↗
78ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0002-5263-5943ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 40 · 5 first-author · 11 since 2021Computer networks · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Perceptual Observatory Characterizing Robustness and Grounding in MLLMsabstractRecent advances in multimodal large language models (MLLMs) have yielded increasingly powerful models, yet their perceptual capacities remain poorly characterized. In practice, most model families scale language component while reusing nearly identical vision encoders (e.g., Qwen2.5-VL 3B/7B/72B), which raises pivotal concerns about whether progress reflects genuine visual grounding or reliance on internet-scale textual world knowledge. Existing evaluation methods emphasize end-task accuracy, overlooking robustness, attribution fidelity, and reasoning under controlled perturbations. We present The Perceptual Observatory, a framework that characterizes MLLMs across verticals like: (i) simple vision tasks, such as face matching and text-in-vision comprehension capabilities; (ii) local-to-global understanding, encompassing image matching, grid pointing game, and attribute localization, which tests general visual grounding. Each vertical is instantiated with ground-truth datasets of faces and words, systematically perturbed through pixel-based augmentations and diffusion-based stylized illusions. The Perceptual Observatory moves beyond leaderboard accuracy to yield insights into how MLLMs preserve perceptual grounding and relational structure under perturbations, providing a principled foundation for analyzing strengths and weaknesses of current and future models. Tejas Anvekar, Fenil Denish Bardoliya, Pavan Turaga, Chitta Baral, Vivek Gupta 0001 |
WACV | 3 |
| 2026 | Topological feature guided knowledge distillation for improved wearable sensor data analysis
Jun Soo Kim, Jae Chan Jeong, Matthew P. Buman, Pavan Turaga, Eun Som Jeon |
Neurocomputing | 4 |
| 2026 | Improved Knowledge Distillation Based on Global Latent Workspace With Multimodal Knowledge Fusion for Understanding Topological Guidance on Wearable Sensor DataabstractWearable sensors have found numerous applications in health and wellness promotion and have achieved great success leveraging advancements in deep learning. However, the development of robust continues to be hindered by issues related to sensor noise, inconsistent sampling rates, and individual differences. Topological data analysis (TDA) has emerged as a viable solution to extract robust features from such time-series data by converting them into persistence images (PIs), which capture intrinsic characteristics and demonstrate resilience to noise and signal variations. However, the computational costs of TDA pose significant challenges for small devices with limited resources. To more efficiently incorporate topological features, we utilize knowledge distillation (KD), which is a promising way to generate a smaller model using larger models. Multiple teachers can be adopted to enrich features in KD. However, this approach has presented two key challenges: 1) differences in feature dimensions from multimodal data and 2) conflicting knowledge provided by the different teachers, both of which can degrade the student model's performance. To address these issues, we propose a novel KD framework called multimodal global latent workspace-based KD (mGLW-KD) that is motivated by global workspace theory (GTW) from cognitive neuroscience. GWT models how the brain integrates and distributes relevant information across different neural modules through a shared workspace, and it includes attentional control and working memory to prioritize and retain key information. Inspired by this theory, mGLW-KD incorporates a working memory module to unify diverse knowledge from multiple teacher models into a shared latent workspace, facilitating efficient knowledge transfer to the student model. By integrating topological insights with cognitive principles, mGLW-KD addresses the challenges posed by wearable sensor data and enables the student model to achieve superior performance using only time-series input during inference. Jinyung Hong, Eun Som Jeon, Matthew P. Buman, Pavan Turaga, Theodore P. Pavlic |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Deep Geometric Moments Promote Shape Consistency in Text-to-3D GenerationabstractTo address the data scarcity associated with 3D assets, 2D-lifting techniques such as Score Distillation Sampling (SDS) have become a widely adopted practice in text-to-3D generation pipelines. However, the diffusion models used in these techniques are prone to viewpoint bias and thus lead to geometric inconsistencies such as the Janus problem. To counter this, we introduce MT3D, a text-to-3D generative model that leverages a high-fidelity 3D object to overcome viewpoint bias and explicitly infuse geometric understanding into the generation pipeline. Firstly, we employ depth maps derived from a high-quality 3D model as control signals to guarantee that the generated 2D images preserve the funda-mental shape and structure, thereby reducing the inherent viewpoint bias. Next, we utilize deep geometric moments to ensure geometric consistency in the 3D representation explicitly. By incorporating geometric details from a 3D asset, MT3D enables the creation of diverse and geometri-cally consistent objects, thereby improving the quality and usability of our 3D representations. Project page and code: https://moment-3d.github.io/ Utkarsh Nath, Rajeev Goel, Eun Som Jeon, Changhoon Kim, Kyle Min 0001, Yezhou Yang, Yingzhen Yang, Pavan Turaga |
WACV | 8 |
| 2025 | Polynomial Implicit Neural Framework for Promoting Shape Awareness in Generative Models
Utkarsh Nath, Rajhans Singh, Ankita Shukla, Kuldeep Kulkarni, Pavan Turaga |
Int. J. Comput. Vis. | 5 |
| 2025 | Intra-class patch swap for self-distillation
Hongjun Choi, Eun Som Jeon, Ankita Shukla, Pavan Turaga |
Neurocomputing | 4 |
| 2025 | Ground Reaction Force Estimation via Time-Aware Knowledge DistillationabstractHuman gait analysis with wearable sensors has been widely used in various applications, such as daily life healthcare, rehabilitation, physical therapy, and clinical diagnostics and monitoring. In particular, ground reaction force (GRF) provides critical information about how the body interacts with the ground during locomotion. Although instrumented treadmills have been widely used as the gold standard for measuring GRF during walking, their lack of portability and high cost make them impractical for many applications. As an alternative, low-cost, portable, wearable insole sensors have been utilized to measure GRF; however, these sensors are susceptible to noise and disturbance and are less accurate than treadmill measurements. Deep learning has shown potential in addressing these issues, but such methods are computationally expensive and often require extensive computing resources, limiting their feasibility for real-time and portable systems. To address these challenges, we propose a Time-aware Knowledge Distillation framework for GRF estimation from insole sensor data. This framework leverages similarity and temporal features within a mini-batch during the knowledge distillation process, effectively capturing the complementary relationships between features and the sequential properties of the target and input data. The performance of the lightweight models distilled through this framework was evaluated by comparing GRF estimations from insole sensor data against measurements from an instrumented treadmill. Various teacher-student model architectures and learning strategies were evaluated across multiple performance metrics using data collected at different walking speeds. Empirical results demonstrated that Time-aware Knowledge Distillation outperforms current baselines in GRF estimation from wearable sensor data. Moreover, our method significantly reduces the number of training parameters needed for GRF estimation, offering a data- and resource-efficient solution for human gait analysis while achieving excellent accuracy and model reliability. Eun Som Jeon, Sinjini Mitra, Jisoo Lee, Omik M. Save, Ankita Shukla, Hyunglae Lee, Pavan Turaga |
IEEE Internet Things J. | 7 |
| 2024 | Learning Geometry of Pose Image Manifolds in Latent Spaces Using Geometry-Preserving GANs
Shenyuan Liang, Benjamin Beaudett, Pavan Turaga, Saket Anand, Anuj Srivastava |
ICPR (27) | 3 |
| 2024 | Topological persistence guided knowledge distillation for wearable sensor data
Eun Som Jeon, Hongjun Choi, Ankita Shukla, Yuan Wang 0057, Hyunglae Lee, Matthew P. Buman, Pavan Turaga |
Eng. Appl. Artif. Intell. | 7 |
| 2024 | RNAS-CL: Robust Neural Architecture Search by Cross-Layer Knowledge Distillation
Utkarsh Nath, Pavan Turaga, Yingzhen Yang |
Int. J. Comput. Vis. | 3 |
| 2024 | Uncertainty-Aware Topological Persistence Guided Knowledge Distillation on Wearable Sensor DataabstractIn applications involving analysis of wearable sensor data, machine learning techniques that use features from topological data analysis (TDA) have demonstrated remarkable performance. Persistence images (PIs) generated through TDA prove effective in capturing robust features, especially to signal perturbations, thus complementing classical time-series features. Despite its promising performance, utilizing TDA to create PI entails significant computational resources and time, posing challenges for applications on small devices. Knowledge distillation (KD) emerges as a solution to address these challenges, as it can produce a compact model. Using multiple teachers one trained with raw time-series and another with topological features, is a viable approach to distill a single compact student model. In such a case, the two teachers will have different statistical characteristics and need some form of feature harmonization. To tackle these issues, we propose uncertainty-aware topological persistence guided knowledge distillation. This approach involves separating common and distinct components between teachers and applying varying weights to control their effects. To enhance the knowledge provided to a student, uncertain features from teachers are rectified using uncertainty scores. We leverage feature similarities to offer more valuable information and employ relationships computed based on orthogonal properties to prevent excessive feature transformation. Ultimately, our method yields a robust single student that operates solely on time-series data at test-time. We validate the effectiveness of the proposed approach through empirical evaluations across various combinations of models and datasets, demonstrating its robustness and efficacy in different scenarios. The proposed method enhances the classification performance of a student model by approximately 4.3% compared to a model learned from scratch on GENEActiv. Eun Som Jeon, Matthew P. Buman, Pavan Turaga |
IEEE Internet Things J. | 3 |
| 2023 | Polynomial Implicit Neural Representations For Large Diverse DatasetsabstractImplicit neural representations (INR) have gained significant popularity for signal and image representation for many end-tasks, such as superresolution, 3D modeling, and more. Most INR architectures rely on sinusoidal positional encoding, which accounts for high-frequency information in data. However, the finite encoding size restricts the model's representational power. Higher representational power is needed to go from representing a single given image to representing large and diverse datasets. Our approach addresses this gap by representing an image with a polynomial function and eliminates the need for positional encodings. Therefore, to achieve a progressively higher degree of polynomial representation, we use element-wise multiplications between features and affine-transformed coordinate locations after every ReLU layer. The proposed method is evaluated qualitatively and quantitatively on large datasets like ImageNet. The proposed Poly-INR model performs comparably to state-of-the-art generative models without any convolution, normalization, or self-attention layers, and with far fewer trainable parameters. With much fewer training parameters and higher representative power, our approach paves the way for broader adoption of INR models for generative modeling tasks in complex domains. The code is available at https://github.com/Rajhans0/Poly_INR Rajhans Singh, Ankita Shukla, Pavan Turaga |
CVPR | 3 |
| 2023 | Robust Time Series Recovery and Classification Using Test-Time Noise Simulator NetworksabstractTime-series are commonly susceptible to various types of corruption due to sensor-level changes and defects which can result in missing samples, sensor and quantization noise, unknown calibration, unknown phase shifts etc. These corruptions cannot be easily corrected as the noise model may be unknown at the time of deployment. This also results in the inability to employ pre-trained classifiers, trained on (clean) source data. In this paper, we present a general framework and models for time-series that can make use of (unlabeled) test samples to estimate the noise model-entirely at test time. To this end, we use a coupled decoder model and an additional neural network which acts as a learned noise model simulator. We show that the framework is able to "clean" the data so as to match the source training data statistics and the cleaned data can be directly used with a pre-trained classifier for robust predictions. We perform empirical studies on diverse application domains with different types of sensors, clearly demonstrating the effectiveness and generality of this method. Eun Som Jeon, Suhas Lohit, Rushil Anirudh, Pavan Turaga |
ICASSP | 4 |
| 2023 | Single-Shot Domain Adaptation via Target-Aware Generative AugmentationsabstractThe problem of adapting models from a source domain using data from any target domain of interest has gained prominence, thanks to the brittle generalization in deep neural networks. While several test-time adaptation techniques have emerged, they typically rely on synthetic data augmentations in cases of limited target data availability. In this paper, we consider the challenging setting of single-shot adaptation and explore the design of augmentation strategies. We argue that augmentations utilized by existing methods are insufficient to handle large distribution shifts, and hence propose a new approach SiSTA (Single-Shot Target Augmentations), which first fine-tunes a generative model from the source domain using a single-shot target, and then employs novel sampling strategies for curating synthetic target data. Using experiments with a state-of-the-art domain adaptation method, we find that SiSTA produces improvements as high as 20% over existing baselines under challenging shifts in face attribute detection, and that it performs competitively to oracle models obtained by training on a larger target dataset. Our codes can be accessed at github.com/kowshikthopalli/SISTA. Rakshith Subramanyam, Kowshik Thopalli, Spring Berman, Pavan Turaga, Jayaraman J. Thiagarajan |
ICASSP | 4 |
| 2023 | Target-Aware Generative Augmentations for Single-Shot AdaptationabstractIn this paper, we address the problem of adapting models from a source domain to a target domain, a task that has become increasingly important due to the brittle generalization of deep neural networks. While several test-time adaptation techniques have emerged, they typically rely on synthetic toolbox data augmentations in cases of limited target data availability. We consider the challenging setting of single-shot adaptation and explore the design of augmentation strategies. We argue that augmentations utilized by existing methods are insufficient to handle large distribution shifts, and hence propose a new approach SiSTA, which first fine-tunes a generative model from the source domain using a single-shot target, and then employs novel sampling strategies for curating synthetic target data. Using experiments on a variety of benchmarks, distribution shifts and image corruptions, we find that SiSTA produces significantly improved generalization over existing baselines in face attribute detection and multi-class object recognition. Furthermore, SiSTA performs competitively to models obtained by training on larger target datasets. Our codes can be accessed at https://github.com/Rakshith-2905/SiSTA Kowshik Thopalli, Rakshith Subramanyam, Pavan Turaga, Jayaraman J. Thiagarajan |
ICML | 3 |
| 2023 | Understanding the Role of Mixup in Knowledge Distillation: An Empirical StudyabstractMixup is a popular data augmentation technique based on creating new samples by linear interpolation between two given data samples, to improve both the generalization and robustness of the trained model. Knowledge distillation (KD), on the other hand, is widely used for model compression and transfer learning, which involves using a larger network’s implicit knowledge to guide the learning of a smaller network. At first glance, these two techniques seem very different, however, we found that "smoothness" is the connecting link between the two and is also a crucial attribute in understanding KD’s interplay with mixup. Although many mixup variants and distillation methods have been proposed, much remains to be understood regarding the role of a mixup in knowledge distillation. In this paper, we present a detailed empirical study on various important dimensions of compatibility between mixup and knowledge distillation. We also scrutinize the behavior of the networks trained with a mixup in the light of knowledge distillation through extensive analysis, visualizations, and comprehensive experiments on image classification. Finally, based on our findings, we suggest improved strategies to guide the student network to enhance its effectiveness. Additionally, the findings of this study provide insightful suggestions to researchers and practitioners that commonly use techniques from KD. Our code is available at https://github.com/hchoi71/MIX-KD. Hongjun Choi, Eun Som Jeon, Ankita Shukla, Pavan Turaga |
WACV | 4 |
| 2023 | Leveraging angular distributions for improved knowledge distillation
Eun Som Jeon, Hongjun Choi, Ankita Shukla, Pavan Turaga |
Neurocomputing | 4 |
| 2022 | Domain Alignment Meets Fully Test-Time Adaptation
Kowshik Thopalli, Pavan Turaga, Jayaraman J. Thiagarajan |
ACML | 2 |
| 2022 | Role of Data Augmentation Strategies in Knowledge Distillation for Wearable Sensor DataabstractDeep neural networks are parametrized by several thousands or millions of parameters, and have shown tremendous success in many classification problems. However, the large number of parameters makes it difficult to integrate these models into edge devices such as smartphones and wearable devices. To address this problem, knowledge distillation (KD) has been widely employed, that uses a pre-trained high capacity network to train a much smaller network, suitable for edge devices. In this paper, for the first time, we study the applicability and challenges of using KD for time-series data for wearable devices. Successful application of KD requires specific choices of data augmentation methods during training. However, it is not yet known if there exists a coherent strategy for choosing an augmentation approach during KD. In this paper, we report the results of a detailed study that compares and contrasts various common choices and some hybrid data augmentation strategies in KD based human activity analysis. Research in this area is often limited as there are not many comprehensive databases available in the public domain from wearable devices. Our study considers databases from small scale publicly available to one derived from a large scale interventional study into human activity and sedentary behavior. We find that the choice of data augmentation techniques during KD have a variable level of impact on end performance, and find that the optimal network choice as well as data augmentation strategies are specific to a dataset at hand. However, we also conclude with a general set of recommendations that can provide a strong baseline performance across databases. Eun Som Jeon, Anirudh Som, Ankita Shukla, Kristina Hasanaj, Matthew P. Buman, Pavan Turaga |
IEEE Internet Things J. | 6 |
| 2021 | Generative Patch Priors for Practical Compressive Image RecoveryabstractIn this paper, we propose the generative patch prior (GPP) that defines a generative prior for compressive image recovery, based on patch-manifold models. Unlike learned, image-level priors that are restricted to the range space of a pre-trained generator, GPP can recover a wide variety of natural images using a pre-trained patch generator. Additionally, GPP retains the benefits of generative priors like high reconstruction quality at extremely low sensing rates, while also being much more generally applicable. We show that GPP outperforms several unsupervised and supervised techniques on three different sensing models - linear compressive sensing with known, and unknown calibration settings, and the non-linear phase retrieval problem. Finally, we propose an alternating optimization strategy using GPP for joint calibration-and-reconstruction which performs favorably against several baselines on a real world, uncalibrated compressive sensing dataset. The code and models for GPP are available on github.1. Rushil Anirudh, Suhas Lohit, Pavan Turaga |
WACV | 3 |
| 2021 | Recovering Trajectories of Unmarked Joints in 3D Human Actions Using Latent Space OptimizationabstractMotion capture (mocap) and time-of-flight based sensing of human actions are becoming increasingly popular modalities to perform robust activity analysis. Applications range from action recognition to quantifying movement quality for health applications. While marker-less motion capture has made great progress, in critical applications such as healthcare, marker-based systems, especially active markers, are still considered gold-standard. However, there are several practical challenges in both modalities such as visibility, tracking errors, and simply the need to keep marker setup convenient wherein movements are recorded with a reduced marker-set. This implies that certain joint locations will not even be marked-up, making downstream analysis of full body movement challenging. To address this gap, we first pose the problem of reconstructing the unmarked joint data as an ill-posed linear inverse problem. We recover missing joints for a given action by projecting it onto the manifold of human actions, this is achieved by optimizing the latent space representation of a deep autoencoder. Experiments on both mocap and Kinect datasets clearly demonstrate that the proposed method performs very well in recovering semantics of the actions and dynamics of missing joints. We will release all the code and models publicly. Suhas Lohit, Rushil Anirudh, Pavan Turaga |
WACV | 3 |
| 2020 | Rate-Invariant Autoencoding of Time-SeriesabstractFor time-series classification and retrieval applications, an important requirement is to develop representations/metrics that are robust to re-parametrization of the time-axis. Temporal re-parametrization as a model can account for variability in the underlying generative process, sampling rate variations, or plain temporal mis-alignment. In this paper, we extend prior work in disentangling latent spaces of autoencoding models, to design a novel architecture to learn rate-invariant latent codes in a completely unsupervised fashion. Unlike conventional neural network architectures, this method allows to explicitly disentangle temporal parameters in the form of order-preserving diffeomorphisms with respect to a learnable template. This makes the latent space more easily interpretable. We show the efficacy of our approach on a synthetic dataset and a real dataset for hand action-recognition. Kaushik Koneripalli, Suhas Lohit, Rushil Anirudh, Pavan Turaga |
ICASSP | 4 |
| 2019 | PrOSe: Product of Orthogonal Spheres Parameterization for Disentangled Representation Learning
Ankita Shukla, Sarthak Bhagat, Shagun Uppal, Saket Anand, Pavan Turaga |
BMVC | 5 |
| 2019 | Temporal Transformer Networks: Joint Learning of Invariant and Discriminative Time WarpingabstractMany time-series classification problems involve developing metrics that are invariant to temporal misalignment. In human activity analysis, temporal misalignment arises due to various reasons including differing initial phase, sensor sampling rates, and elastic time-warps due to subject-specific biomechanics. Past work in this area has only looked at reducing intra-class variability by elastic temporal alignment. In this paper, we propose a hybrid model-based and data-driven approach to learn warping functions that not just reduce intra-class variability, but also increase inter-class separation. We call this a temporal transformer network (TTN). TTN is an interpretable differentiable module, which can be easily integrated at the front end of a classification network. The module is capable of reducing intra-class variance by generating input-dependent warping functions which lead to rate-robust representations. At the same time, it increases inter-class variance by learning warping functions that are more discriminative. We show improvements over strong baselines in 3D action recognition on challenging datasets using the proposed framework. The improvements are especially pronounced when training sets are smaller. Suhas Lohit, Pavan Turaga |
CVPR | 3 |
| 2019 | An REU Experience in Machine Learning and Computational CamerasabstractIn this work in progress paper, we describe an REU summer experience on imaging sensors that involved a female junior level Electrical Engineering student, a graduate student advisor, and three faculty. A research plan was designed to embed the student in a sensor and machine learning research with specific emphasis on energy-efficient cameras. The motivation for submitting this paper is the unique planning and the quality of the overall student experience which resulted in continuous engagement of the REU student with the faculty after the REU summer program completed. The program resulted in a major presentation at an international event, an NSF I/UCRC poster presentation, a research conference submission which is remarkable for an undergraduate student, and finally a new research direction for the graduate mentor and faculty. This paper describes successful strategies for research engagement for undergraduates in state-of-the-art research fields which yield positive outcomes for all participants, and is grounded in contemporary educational methodology and theoretical frameworks. Divya Mohan, Sameeksha Katoch, Suren Jayasuriya, Pavan Turaga, Andreas Spanias |
FIE | 4 |
| 2019 | Multiple Subspace Alignment Improves Domain AdaptationabstractWe present a novel unsupervised domain adaptation (DA) method for cross-domain visual recognition. Though subspace methods have found success in DA, their performance is often limited due to the assumption of approximating an entire dataset using a single low-dimensional subspace. Instead, we develop a method to effectively represent the source and target datasets via a collection of low-dimensional subspaces, and subsequently align them by exploiting the natural geometry of the space of subspaces, on the Grassmann manifold. We demonstrate the effectiveness of this approach, using empirical studies on two widely used benchmarks,with performance on par or better than the performance of the state of the art domain adaptation methods. Kowshik Thopalli, Rushil Anirudh, Jayaraman J. Thiagarajan, Pavan Turaga |
ICASSP | 4 |
| 2019 | Spatially-Varying Sharpness Map Estimation Based on the Quotient of Spectral BandsabstractNatural images suffer from defocus blur due to the presence of objects at different depths from the camera. Automatic estimation of spatially-varying sharpness has several applications including depth estimation, image quality assessment, information retrieval, image restoration among others. In this paper, we propose a sharpness metric based on the quotient of high- to low-frequency bands of the log-spectrum of the image gradients. Using the proposed sharpness metric, we obtain a descriptive dense sharpness map. We also propose a simple yet effective method to segment out-of-focus regions using a global threshold which is defined using weak textured regions present in the input image. Results over two publicly available databases show that the proposed method provides competitive performance when compared with state-of-the-art methods. Juan Andrade, Pavan Turaga, Andreas Spanias |
ICIP | 2 |
| 2019 | Non-Parametric Priors For Generative Adversarial NetworksabstractThe advent of generative adversarial networks (GAN) has enabled new capabilities in synthesis, interpolation, and data augmentation heretofore considered very challenging. However, one of the common assumptions in most GAN architectures is the assumption of simple parametric latent-space distributions. While easy to implement, a simple latent-space distribution can be problematic for uses such as interpolation. This is due to distributional mismatches when samples are interpolated in the latent space. We present a straightforward formalization of this problem; using basic results from probability theory and off-the-shelf-optimization tools, we develop ways to arrive at appropriate non-parametric priors. The obtained prior exhibits unusual qualitative properties in terms of its shape, and quantitative benefits in terms of lower divergence with its mid-point distribution. We demonstrate that our designed prior helps improve image generation along any Euclidean straight line during interpolation, both qualitatively and quantitatively, without any additional training or architectural modifications. The proposed formulation is quite flexible, paving the way to impose newer constraints on the latent-space statistics. Rajhans Singh, Pavan Turaga, Suren Jayasuriya, Ravi Garg, Martin W. Braun |
ICML | 2 |
| 2018 | Perturbation Robust Representations of Topological Persistence Diagrams
Anirudh Som, Kowshik Thopalli, Karthikeyan Natesan Ramamurthy, Vinay Venkataraman, Ankita Shukla, Pavan Turaga |
ECCV (7) | 6 |
| 2018 | CS-VQA: Visual Question Answering with Compressively Sensed ImagesabstractVisual Question Answering (VQA) is a complex semantic task requiring both natural language processing and visual recognition. In this paper, we explore whether VQA is solvable when images are captured in a sub-Nyquist compressive paradigm. We develop a series of deep-network architectures that exploit available compressive data to increasing degrees of accuracy, and show that VQA is indeed solvable in the compressed domain. Our results show that there is nominal degradation in VQA performance when using compressive measurements, but that accuracy can be recovered when VQA pipelines are used in conjunction with state-of-the-art deep neural networks for CS reconstruction. The results presented yield important implications for resource-constrained VQA applications. Li-Chi Huang, Kuldeep Kulkarni, Anik Jha, Suhas Lohit, Suren Jayasuriya, Pavan Turaga |
ICIP | 6 |
| 2018 | Fast Non-Linear Methods for Dynamic Texture PredictionabstractThis paper aims to develop a fast dynamic-texture prediction method, using tools from non-linear dynamical modeling, and fast approaches for approximate regression. We consider dynamic textures to be described by patch-level non-linear processes, thus requiring tools such as delay-embedding to uncover a phase-space where dynamical evolution can be more easily modeled. After mapping the observed time-series from a dynamic texture video to its recovered phase-space, a time-efficient approximate prediction method is presented which utilizes locality-sensitive hashing approaches to predict possible phase-space vectors, given the current phase-space vector. Our experiments show the favorable performance of the proposed approach, both in terms of prediction fidelity, and computational time. The proposed algorithm is applied to shading prediction in utility scale solar arrays. Sameeksha Katoch, Pavan Turaga, Andreas Spanias, Cihan Tepedelenlioglu |
ICIP | 2 |
| 2018 | Tracking, Animating, and 3D Printing Elements of the Fine Arts Freehand Drawing ProcessabstractDynamic elements of traditional drawing processes such as the order of compilation, and speed, length, and pressure of strokes can be as important as the final art piece because they can reveal the technique, process, and emotions of the artist. In this paper, we present an interactive system that unobtrusively tracks the freehand drawing process (movement and pressure of artist»s pencil) on a regular easel. The system outputs captured information using 2D video renderings and 3D-printed sculptures. We also present a summery of findings from a user study with 6 experienced artists who created multiple pencil drawings using our system. The resulting digital and physical outputs from our system revealed vast differences in drawing speeds, styles, and techniques. At TEI art track, the attendees will likely engage in lively discussion around the analog, digital, and tangible aspects of our exhibit. We believe that such a discussion will be critical not only in shaping the future of our work, but also in understanding novel research directions at the intersection of art and computation. Piyum Fernando, Jennifer Weiler, Stacey Kuznetsov, Pavan Turaga |
TEI | 4 |
| 2018 | Introduction to the Special Issue on Representation, Analysis, and Recognition of 3D HumansabstractNo abstract available. Stefano Berretti, Mohamed Daoudi, Pavan Turaga, Anup Basu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2018 | Representation, Analysis, and Recognition of 3D Humans: A SurveyabstractComputer Vision and Multimedia solutions are now offering an increasing number of applications ready for use by end users in everyday life. Many of these applications are centered for detection, representation, and analysis of face and body. Methods based on 2D images and videos are the most widespread, but there is a recent trend that successfully extends the study to 3D human data as acquired by a new generation of 3D acquisition devices. Based on these premises, in this survey, we provide an overview on the newly designed techniques that exploit 3D human data and also prospect the most promising current and future research directions. In particular, we first propose a taxonomy of the representation methods, distinguishing between spatial and temporal modeling of the data. Then, we focus on the analysis and recognition of 3D humans from 3D static and dynamic data, considering many applications for body and face. Stefano Berretti, Mohamed Daoudi, Pavan Turaga, Anup Basu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2017 | Elastic Functional Coding of Riemannian TrajectoriesabstractVisual observations of dynamic phenomena, such as human actions, are often represented as sequences of smoothly-varying features. In cases where the feature spaces can be structured as Riemannian manifolds, the corresponding representations become trajectories on manifolds. Analysis of these trajectories is challenging due to non-linearity of underlying spaces and high-dimensionality of trajectories. In vision problems, given the nature of physical systems involved, these phenomena are better characterized on a low-dimensional manifold compared to the space of Riemannian trajectories. For instance, if one does not impose physical constraints of the human body, in data involving human action analysis, the resulting representation space will have highly redundant features. Learning an effective, low-dimensional embedding for action representations will have a huge impact in the areas of search and retrieval, visualization, learning, and recognition. Traditional manifold learning addresses this problem for static points in the euclidean space, but its extension to Riemannian trajectories is non-trivial and remains unexplored. The difficulty lies in inherent non-linearity of the domain and temporal variability of actions that can distort any traditional metric between trajectories. To overcome these issues, we use the framework based on transported square-root velocity fields (TSRVF); this framework has several desirable properties, including a rate-invariant metric and vector space representations. We propose to learn an embedding such that each action trajectory is mapped to a single point in a low-dimensional euclidean space, and the trajectories that differ only in temporal rates map to the same point. We utilize the TSRVF representation, and accompanying statistical summaries of Riemannian trajectories, to extend existing coding methods such as PCA, KSVD and Label Consistent KSVD to Riemannian trajectories or more generally to Riemannian functions. We show that such coding efficiently captures trajectories in applications such as action recognition, stroke rehabilitation, visual speech recognition, clustering and diverse sequence sampling. Using this framework, we obtain state-of-the-art recognition results, while reducing the dimensionality/ complexity by a factor of 100-250x. Since these mappings and codes are invertible, they can also be used to interactively-visualize Riemannian trajectories and synthesize actions. Rushil Anirudh, Pavan Turaga, Jingyong Su, Anuj Srivastava |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | ReconNet: Non-Iterative Reconstruction of Images from Compressively Sensed MeasurementsabstractThe goal of this paper is to present a non-iterative and more importantly an extremely fast algorithm to reconstruct images from compressively sensed (CS) random measurements. To this end, we propose a novel convolutional neural network (CNN) architecture which takes in CS measurements of an image as input and outputs an intermediate reconstruction. We call this network, ReconNet. The intermediate reconstruction is fed into an off-the-shelf denoiser to obtain the final reconstructed image. On a standard dataset of images we show significant improvements in reconstruction results (both in terms of PSNR and time complexity) over state-of-the-art iterative CS reconstruction algorithms at various measurement rates. Further, through qualitative experiments on real data collected using our block single pixel camera (SPC), we show that our network is highly robust to sensor noise and can recover visually better quality images than competitive algorithms at extremely low sensing rates of 0.1 and 0.04. To demonstrate that our algorithm can recover semantically informative images even at a low measurement rate of 0.01, we present a very robust proof of concept real-time visual tracking application. Kuldeep Kulkarni, Suhas Lohit, Pavan Turaga, Ronan Kerviche, Amit Ashok |
CVPR | 3 |
| 2016 | Consensus inference on mobile phone sensors for activity recognitionabstractThe pervasive use of wearable sensors in activity and health monitoring presents a huge potential for building novel data analysis and prediction frameworks. In particular, approaches that can harness data from a diverse set of low-cost sensors for recognition are needed. Many of the existing approaches rely heavily on elaborate feature engineering to build robust recognition systems, and their performance is often limited by the inaccuracies in the data. In this paper, we develop a novel two-stage recognition system that enables a systematic fusion of complementary information from multiple sensors in a linear graph embedding setting, while employing an ensemble classifier phase that leverages the discriminative power of different feature extraction strategies. Experimental results on a challenging dataset show that our framework greatly improves the recognition performance when compared to using any single sensor. Huan Song, Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Andreas Spanias, Pavan Turaga |
ICASSP | 5 |
| 2016 | Diversity promoting online sampling for streaming video summarizationabstractMany applications benefit from sampling algorithms where a small number of well chosen samples are used to generalize different properties of a large dataset. In this paper, we use diverse sampling for streaming video summarization. Several emerging applications support streaming video, but existing summarization algorithms need access to the entire video which requires a lot of memory and computational power. We propose a memory efficient and computationally fast, online algorithm that uses competitive learning for diverse sampling. Our algorithm is a generalization of online K-means such that the cost function reduces clustering error, while also ensuring a diverse set of samples. The diversity is measured as the volume of a convex hull around the samples. Finally, the performance of the proposed algorithm is measured against human users for 50 videos in the VSUMM dataset. The algorithm performs better than batch mode summarization, while requiring significantly lower memory and computational requirements. Rushil Anirudh, Ahnaf Masroor, Pavan Turaga |
ICIP | 3 |
| 2016 | Direct inference on compressive measurements using convolutional neural networksabstractCompressive imagers, e.g. the single-pixel camera (SPC), acquire measurements in the form of random projections of the scene instead of pixel intensities. Compressive Sensing (CS) theory allows accurate reconstruction of the image even from a small number of such projections. However, in practice, most reconstruction algorithms perform poorly at low measurement rates and are computationally very expensive. But perfect reconstruction is not the goal of high-level computer vision applications. Instead, we are interested in only determining certain properties of the image. Recent work has shown that effective inference is possible directly from the compressive measurements, without reconstruction, using correlational features. In this paper, we show that convolutional neural networks (CNNs) can be employed to extract discriminative non-linear features directly from CS measurements. Using these features, we demonstrate that effective high-level inference can be performed. Experimentally, using hand written digit recognition (MNIST dataset) and image recognition (ImageNet) as examples, we show that recognition is possible even at low measurement rates of about 0.1. Suhas Lohit, Kuldeep Kulkarni, Pavan Turaga |
ICIP | 3 |
| 2016 | Persistent homology of attractors for action recognitionabstractIn this paper, we propose a novel framework for dynamical analysis of human actions from 3D motion capture data using topological data analysis. We model human actions using the topological features of the attractor of the dynamical system. We reconstruct the phase-space of time series corresponding to actions using time-delay embedding, and compute the persistent homology of the phase-space reconstruction. In order to better represent the topological properties of the phase-space, we incorporate the temporal adjacency information when computing the homology groups. The persistence of these homology groups encoded using persistence diagrams are used as features for the actions. Our experiments with action recognition using these features demonstrate that the proposed approach outperforms other baseline methods. Vinay Venkataraman, Karthikeyan Natesan Ramamurthy, Pavan Turaga |
ICIP | 3 |
| 2016 | Geometry-Based Symbolic Approximation for Fast Sequence Matching on Manifolds
Rushil Anirudh, Pavan Turaga |
Int. J. Comput. Vis. | 2 |
| 2016 | Reconstruction-Free Action Inference from Compressive ImagersabstractPersistent surveillance from camera networks, such as at parking lots, UAVs, etc., often results in large amounts of video data, resulting in significant challenges for inference in terms of storage, communication and computation. Compressive cameras have emerged as a potential solution to deal with the data deluge issues in such applications. However, inference tasks such as action recognition require high quality features which implies reconstructing the original video data. Much work in compressive sensing (CS) theory is geared towards solving the reconstruction problem, where state-of-the-art methods are computationally intensive and provide low-quality results at high compression rates. Thus, reconstruction-free methods for inference are much desired. In this paper, we propose reconstruction-free methods for action recognition from compressive cameras at high compression ratios of 100 and above. Recognizing actions directly from CS measurements requires features which are mostly nonlinear and thus not easily applicable. This leads us to search for such properties that are preserved in compressive measurements. To this end, we propose the use of spatio-temporal smashed filters, which are compressive domain versions of pixel-domain matched filters. We conduct experiments on publicly available databases and show that one can obtain recognition rates that are comparable to the oracle method in uncompressed setup, even for high compression ratios. Kuldeep Kulkarni, Pavan Turaga |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Shape Distributions of Nonlinear Dynamical Systems for Video-Based InferenceabstractThis paper presents a shape-theoretic framework for dynamical analysis of nonlinear dynamical systems which appear frequently in several video-based inference tasks. Traditional approaches to dynamical modeling have included linear and nonlinear methods with their respective drawbacks. A novel approach we propose is the use of descriptors of the shape of the dynamical attractor as a feature representation of nature of dynamics. The proposed framework has two main advantages over traditional approaches: a) representation of the dynamical system is derived directly from the observational data, without any inherent assumptions, and b) the proposed features show stability under different time-series lengths where traditional dynamical invariants fail. We illustrate our idea using nonlinear dynamical models such as Lorenz and Rossler systems, where our feature representations (shape distribution) support our hypothesis that the local shape of the reconstructed phase space can be used as a discriminative feature. Our experimental analyses on these models also indicate that the proposed framework show stability for different time-series lengths, which is useful when the available number of samples are small/variable. The specific applications of interest in this paper are: 1) activity recognition using motion capture and RGBD sensors, 2) activity quality assessment for applications in stroke rehabilitation, and 3) dynamical scene classification. We provide experimental validation through action and gesture recognition experiments on motion capture and Kinect datasets. In all these scenarios, we show experimental evidence of the favorable properties of the proposed representation. Vinay Venkataraman, Pavan Turaga |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Component-Level Tuning of Kinematic Features From Composite Therapist Impressions of Movement QualityabstractIn this paper, we propose a general framework for tuning component-level kinematic features using therapists' overall impressions of movement quality, in the context of a home-based adaptive mixed reality rehabilitation (HAMRR) system. We propose a linear combination of nonlinear kinematic features to model wrist movement, and propose an approach to learn feature thresholds and weights using high-level labels of overall movement quality provided by a therapist. The kinematic features are chosen such that they correlate with the quality of wrist movements to clinical assessment scores. Further, the proposed features are designed to be reliably extracted from an inexpensive and portable motion capture system using a single reflective marker on the wrist. Using a dataset collected from ten stroke survivors, we demonstrate that the framework can be reliably used for movement quality assessment in HAMRR systems. The system is currently being deployed for large-scale evaluations, and will represent an increasingly important application area of motion capture and activity analysis. Vinay Venkataraman, Pavan Turaga, Michael Baran, Nicole Lehrer, Tingfang Du, Thanassis Rikakis, Steven L. Wolf |
IEEE J. Biomed. Health Informatics | 2 |
| 2015 | Dynamical Regularity for Action AnalysisabstractIn this paper, we propose a new approach for quantification of ‘dynamical regularity’ as applied to modeling human actions. We use approximate entropy-based feature representation to model the dynamics in human movement to achieve temporal segmentation in untrimmed motion capture data and fine-grained quality assessment of diving actions in videos. The principle herein is to quantify regularity (frequency of typical patterns) in the dynamical space computed from trajectories of action data. We extend conventional ideas for modeling dynamics in human movement by introducing multivariate and cross approximate entropy features. Our experimental evaluation on theoretical models and two publicly available databases show that the proposed features can achieve state-ofthe-art results on applications such as temporal segmentation and quality assessment of actions. Vinay Venkataraman, Ioannis Vlachos 0001, Pavan Turaga |
BMVC | 3 |
| 2015 | Elastic functional coding of human actions: From vector-fields to latent variablesabstractHuman activities observed from visual sensors often give rise to a sequence of smoothly varying features. In many cases, the space of features can be formally defined as a manifold, where the action becomes a trajectory on the manifold. Such trajectories are high dimensional in addition to being non-linear, which can severely limit computations on them. We also argue that by their nature, human actions themselves lie on a much lower dimensional manifold compared to the high dimensional feature space. Learning an accurate low dimensional embedding for actions could have a huge impact in the areas of efficient search and retrieval, visualization, learning, and recognition. Traditional manifold learning addresses this problem for static points in ℝn, but its extension to trajectories on Riemannian manifolds is non-trivial and has remained unexplored. The challenge arises due to the inherent non-linearity, and temporal variability that can significantly distort the distance metric between trajectories. To address these issues we use the transport square-root velocity function (TSRVF) space, a recently proposed representation that provides a metric which has favorable theoretical properties such as invariance to group action. We propose to learn the low dimensional embedding with a manifold functional variant of principal component analysis (mfPCA). We show that mf-PCA effectively models the manifold trajectories in several applications such as action recognition, clustering and diverse sequence sampling while reducing the dimensionality by a factor of ~ 250×. The mfPCA features can also be reconstructed back to the original manifold to allow for easy visualization of the latent variable space. Rushil Anirudh, Pavan Turaga, Jingyong Su, Anuj Srivastava |
CVPR | 2 |
| 2015 | Geometric Compression of Orientation Signals for Fast Gesture AnalysisabstractThis paper concerns itself with compression strategies for orientation signals, seen as signals evolving on the space of quaternion's. The compression techniques extend classical signal approximation strategies used in data mining, by explicitly taking into account the quotient-space properties of the quaternion space. The approximation techniques are applied to the case of human gesture recognition from cell phone-based orientation sensors. Results indicate that the proposed approach results in high recognition accuracies, with low storage requirements, with the geometric computations providing added robustness than classical vector-space computations. Aswin Sivakumar, Rushil Anirudh, Pavan Turaga |
DCC | 3 |
| 2015 | Discriminative feature learning from big data for visual recognition
Zhuolin Jiang, Zhe Lin 0001, Haibin Ling, Fatih Porikli, Ling Shao 0001, Pavan Turaga |
Pattern Recognit. | 6 |
| 2014 | Direct tracking from compressive imagers: A proof of conceptabstractThe compressive sensing paradigm holds promise for more cost-effective imaging outside of the visible range, particularly in infrared wavelengths. However, the process of reconstructing compressively sensed images remains computationally expensive. The proof-of-concept tracker described here uses a particle filter with a likelihood update based on a “smashed filter” which estimates correlation directly, avoiding the reconstruction step. This approach leads to increased noise in correlation estimates, but by implementing the track-before-detect concept in the particle filter, tracker convergence may still be achieved with reasonable sensing rates. The tracker has been successfully tested on sequences of moving cars in the PETS2000 dataset. Henry Braun, Pavan Turaga, Andreas Spanias |
ICASSP | 2 |
| 2014 | Interactively test driving an object detector: Estimating performance on unlabeled dataabstractIn this paper, we study the problem of `test-driving' a detector, i.e. allowing a human user to get a quick sense of how well the detector generalizes to their specific requirement. To this end, we present the first system that estimates detector performance interactively without extensive ground truthing using a human in the loop. We approach this as a problem of estimating proportions and show that it is possible to make accurate inferences on the proportion of classes or groups within a large data collection by observing only 5 - 10% of samples from the data. In estimating the false detections (for precision), the samples are chosen carefully such that the overall characteristics of the data collection are preserved. Next, inspired by its use in estimating disease propagation we apply pooled testing approaches to estimate missed detections (for recall) from the dataset. The estimates thus obtained are close to the ones obtained using ground truth, thus reducing the need for extensive labeling which is expensive and time consuming. Rushil Anirudh, Pavan Turaga |
WACV | 2 |
| 2014 | Differential geometric representations and algorithms for some pattern recognition and computer vision problems
Pavan Turaga, Anuj Srivastava, Rama Chellappa |
Pattern Recognit. Lett. | 2 |
| 2013 | A heterogeneous dictionary model for representation and recognition of human actionsabstractIn this paper, we consider low-dimensional and sparse representation models for human actions, that are consistent with how actions evolve in high-dimensional feature spaces. We first show that human actions can be well approximated by piecewise linear structures in the feature space. Based on this, we propose a new dictionary model that considers each atom in the dictionary to be an affine subspace defined by a point and a corresponding line. When compared to centered clustering approaches such as K-means, we show that the proposed dictionary is a better generative model for human actions. Furthermore, we demonstrate the utility of this model in efficient representation and recognition of human activities that are not available in the training set. Rushil Anirudh, Karthikeyan Natesan Ramamurthy, Jayaraman J. Thiagarajan, Pavan Turaga, Andreas Spanias |
ICASSP | 4 |
| 2013 | Optical flow for compressive sensing video reconstructionabstractAlthough considerable effort has been devoted to the problem of reconstructing compressively sensed video, no existing algorithm achieves results comparable to commonly available video compression methods such as H.264. One possible avenue for improving compressively sensed video reconstruction is the use of optical flow information. Current efforts reported in the literature have not fully utilized optical flow information, instead focusing on limited cases such as stationary backgrounds with sparse foreground motion. In this paper, a reconstruction method is presented which fully utilizes optical flow information to increase the quality of reconstruction. The special cases of known image motion and constant global image motion are presented, and the performance of the algorithm on existing datasets is evaluated. Henry Braun, Pavan Turaga, Cihan Tepedelenlioglu, Andreas Spanias |
ICASSP | 2 |
| 2013 | Compressive Acquisition of Linear Dynamical SystemsabstractCompressive sensing (CS) enables the acquisition and recovery of sparse signals and images at sampling rates significantly below the classical Nyquist rate. Despite significant progress in the theory and methods of CS, little headway has been made in compressive video acquisition and recovery. Video CS is complicated by the ephemeral nature of dynamic events, which makes direct extensions of standard CS imaging architectures and signal models difficult. In this paper, we develop a new framework for video CS for dynamic textured scenes that models the evolution of the scene as a linear dynamical system (LDS). This reduces the video recovery problem to first estimating the model parameters of the LDS from compressive measurements and then reconstructing the image frames. We exploit the low-dimensional dynamic parameters (the state sequence) and high-dimensional static parameters (the observation matrix) of the LDS to devise a novel compressive measurement strategy that measures only the time-varying parameters at each instant and accumulates measurements over time to estimate the time-invariant parameters. This enables us to lower the compressive measurement rate considerably. We validate our approach and demonstrate its effectiveness with a range of experiments involving video recovery and scene classification. Aswin C. Sankaranarayanan, Pavan Turaga, Rama Chellappa, Richard G. Baraniuk |
SIAM J. Imaging Sci. | 2 |
| 2012 | Domain Adaptive Dictionary Learning
Qiang Qiu 0002, Vishal M. Patel, Pavan Turaga, Rama Chellappa |
ECCV (4) | 3 |
| 2012 | Recurrence textures for human activity recognition from compressive camerasabstractRecent advances in camera architectures and associated mathematical representations now enable compressive acquisition of images and videos at low data-rates. In such a setting, we consider the problem of human activity recognition, which is an important inference problem in many security and surveillance applications. We propose a framework for understanding human activities as a non-linear dynamical system, and propose a robust, generalizable feature that can be extracted directly from the compressed measurements without reconstructing the original video frames. The proposed feature is termed recurrence texture and is motivated from recurrence analysis of non-linear dynamical systems. We show that it is possible to obtain discriminative features directly from the compressed stream and show its utility in recognition of activities at very low data rates. Kuldeep Kulkarni, Pavan Turaga |
ICIP | 2 |
| 2012 | On advances in differential-geometric approaches for 2D and 3D shape analyses and activity recognition
Anuj Srivastava, Pavan Turaga, Sebastian Kurtek |
Image Vis. Comput. | 2 |
| 2012 | A Blur-Robust Descriptor with Applications to Face RecognitionabstractUnderstanding the effect of blur is an important problem in unconstrained visual analysis. We address this problem in the context of image-based recognition by a fusion of image-formation models and differential geometric tools. First, we discuss the space spanned by blurred versions of an image and then, under certain assumptions, provide a differential geometric analysis of that space. More specifically, we create a subspace resulting from convolution of an image with a complete set of orthonormal basis functions of a prespecified maximum size (that can represent an arbitrary blur kernel within that size), and show that the corresponding subspaces created from a clean image and its blurred versions are equal under the ideal case of zero noise and some assumptions on the properties of blur kernels. We then study the practical utility of this subspace representation for the problem of direct recognition of blurred faces by viewing the subspaces as points on the Grassmann manifold and present methods to perform recognition for cases where the blur is both homogenous and spatially varying. We empirically analyze the effect of noise, as well as the presence of other facial variations between the gallery and probe images, and provide comparisons with existing approaches on standard data sets. Raghuraman Gopalan, Sima Taheri, Pavan Turaga, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Age Estimation and Face Verification Across Aging Using LandmarksabstractAge estimation and face verification across aging are important problems with a wide range of applications. It is well known that age and identity information are encoded in both texture and shape of the face. Building on recent advances in landmark extraction and statistical techniques for landmark-based shape analysis, we consider these problems using facial shapes. We show that by using well-defined shape spaces and their associated geometry, one can obtain significant performance improvements in both age estimation and face verification. Toward this end, we propose to model the facial shapes as points on a Grassmann manifold. Age estimation and face verification are then considered as regression and classification problems on this manifold. Algorithms for regression and classification are designed to take into account the geometry of the underlying space. The proposed method is flexible and can be used as a standalone age estimator or classifier, and we also present methods for fusion with texture-based algorithms. Tao Wu 0009, Pavan Turaga, Rama Chellappa |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2011 | Recent advances in age and height estimation from still images and videoabstractSoft-biometrics such as gender, age, race, etc have been found to be useful characterizations that enable fast pre-filtering and organization of data for biometric applications. In this paper, we focus on two useful soft-biometrics - age and height. We discuss their utility and the factors involved in their estimation from images and videos. In this context, we highlight the role that geometric constraints such as multiview-geometry, and shape-space geometry play. Then, we present methods based on these geometric constraints for age and height-estimation. These methods provide a principled means by fusing image-formation models, multi-view geometric constraints, and robust statistical methods for inference. Rama Chellappa, Pavan Turaga |
FG | 2 |
| 2011 | Towards view-invariant expression analysis using analytic shape manifoldsabstractFacial expression analysis is one of the important components for effective human-computer interaction. However, to develop robust and generalizable models for expression analysis one needs to break the dependence of the models on the choice of the coordinate frame of the camera i.e. expression models should generalize across facial poses. To perform this systematically, one needs to understand the space of observed images subject to projective transformations. However, since the projective shape-space is cumbersome to work with, we address this problem by deriving models for expressions on the affine shape-space as an approximation to the projective shape-space by using a Riemannian interpretation of deformations that facial expressions cause on different parts of the face. We use landmark configurations to represent facial deformations and exploit the fact that the affine shape-space can be studied using the Grassmann manifold. This representation enables us to perform various expression analysis and recognition algorithms without the need for the normalization as a preprocessing step. We extend some of the available approaches for expression analysis to the Grassmann manifold and experimentally show promising results, paving the way for a more general theory of view-invariant expression analysis. Sima Taheri, Pavan Turaga, Rama Chellappa |
FG | 2 |
| 2011 | Blurring-invariant Riemannian metrics for comparing signals and imagesabstractWe propose a novel Riemannian framework for comparing signals and images in a manner that is invariant to their levels of blur. This framework uses a log-Fourier representation of signals/images in which the set of all possible Gaussian blurs of a signal, i.e. its orbits under semigroup action of Gaussian blur functions, is a straight line. Using a set of Riemannian metrics under which the group actions are by isometries, the orbits are compared via distances between orbits. We demonstrate this framework using a number of experimental results involving 1D signals and 2D images. Zhengwu Zhang, Eric Klassen, Anuj Srivastava, Pavan Turaga, Rama Chellappa |
ICCV | 4 |
| 2011 | Manifold Precis: An Annealing Technique for Diverse Sampling of ManifoldsabstractIn this paper, we consider the 'Precis' problem of sampling K representative yet diverse data points from a large dataset. This problem arises frequently in applications such as video and document summarization, exploratory data analysis, and pre-filtering. We formulate a general theory which encompasses not just traditional techniques devised for vector spaces, but also non-Euclidean manifolds, thereby enabling these techniques to shapes, human activities, textures and many other image and video based datasets. We propose intrinsic manifold measures for measuring the quality of a selection of points with respect to their representative power, and their diversity. We then propose efficient algorithms to optimize the cost function using a novel annealing-based iterative alternation algorithm. The proposed formulation is applicable to manifolds of known geometry as well as to manifolds whose geometry needs to be estimated from samples. Experimental results show the strength and generality of the proposed approach. Nitesh Shroff, Pavan Turaga, Rama Chellappa |
NIPS | 2 |
| 2011 | Statistical Computations on Grassmann and Stiefel Manifolds for Image and Video-Based RecognitionabstractIn this paper, we examine image and video-based recognition applications where the underlying models have a special structure—the linear subspace structure. We discuss how commonly used parametric models for videos and image sets can be described using the unified framework of Grassmann and Stiefel manifolds. We first show that the parameters of linear dynamic models are finite-dimensional linear subspaces of appropriate dimensions. Unordered image sets as samples from a finite-dimensional linear subspace naturally fall under this framework. We show that an inference over subspaces can be naturally cast as an inference problem on the Grassmann manifold. To perform recognition using subspace-based models, we need tools from the Riemannian geometry of the Grassmann manifold. This involves a study of the geometric properties of the space, appropriate definitions of Riemannian metrics, and definition of geodesics. Further, we derive statistical modeling of inter and intraclass variations that respect the geometry of the space. We apply techniques such as intrinsic and extrinsic statistics to enable maximum-likelihood classification. We also provide algorithms for unsupervised clustering derived from the geometry of the manifold. Finally, we demonstrate the improved performance of these methods in a wide variety of vision applications such as activity recognition, video-based face recognition, object recognition from image sets, and activity-based video clustering. Pavan Turaga, Ashok Veeraraghavan, Anuj Srivastava, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Example-Driven Manifold Priors for Image DeconvolutionabstractImage restoration methods that exploit prior information about images to be estimated have been extensively studied, typically using the Bayesian framework. In this paper, we consider the role of prior knowledge of the object class in the form of a patch manifold to address the deconvolution problem. Specifically, we incorporate unlabeled image data of the object class, say natural images, in the form of a patch-manifold prior for the object class. The manifold prior is implicitly estimated from the given unlabeled data. We show how the patch-manifold prior effectively exploits the available sample class data for regularizing the deblurring problem. Furthermore, we derive a generalized cross-validation (GCV) function to automatically determine the regularization parameter at each iteration without explicitly knowing the noise variance. Extensive experiments show that this method performs better than many competitive image deconvolution methods. Jie Ni, Pavan Turaga, Vishal M. Patel, Rama Chellappa |
IEEE Trans. Image Process. | 2 |
| 2010 | Moving vistas: Exploiting motion for describing scenesabstractScene recognition in an unconstrained setting is an open and challenging problem with wide applications. In this paper, we study the role of scene dynamics for improved representation of scenes. We subsequently propose dynamic attributes which can be augmented with spatial attributes of a scene for semantically meaningful categorization of dynamic scenes. We further explore accurate and generalizable computational models for characterizing the dynamics of unconstrained scenes. The large intra-class variation due to unconstrained settings and the complex underlying physics present challenging problems in modeling scene dynamics. Motivated by these factors, we propose using the theory of chaotic systems to capture dynamics. Due to the lack of a suitable dataset, we compiled a dataset of `in-the-wild' dynamic scenes. Experimental results show that the proposed framework leads to the best classification rate among other well-known dynamic modeling techniques. We also show how these dynamic features provide a means to describe dynamic scenes with motion-attributes, which then leads to meaningful organization of the video data. Nitesh Shroff, Pavan Turaga, Rama Chellappa |
CVPR | 2 |
| 2010 | Shape-based similarity retrieval of Doppler images for clinical decision supportabstractFlow Doppler imaging has become an integral part of an echocardiographic exam. Automated interpretation of flow doppler imaging has so far been restricted to obtaining hemodynamic information from velocity-time profiles depicted in these images. In this paper we exploit the shape patterns in Doppler images to infer the similarity in valvular disease labels for purposes of automated clinical decision support. Specifically, we model the similarity in appearance of Doppler images from the same disease class as a constrained non-rigid translation transform of the velocity envelopes embedded in these images. The shape similarity between two Doppler images is then judged by recovering the alignment transform using a variant of dynamic shape warping. Results of similarity retrieval of doppler images for cardiac decision support on a large database of images are presented. Tanveer F. Syeda-Mahmood, Pavan Turaga, David Beymer, Fei Wang 0002, Arnon Amir, Hayit Greenspan, Kilian M. Pohl |
CVPR | 2 |
| 2010 | Articulation-Invariant Representation of Non-planar Shapes
Raghuraman Gopalan, Pavan Turaga, Rama Chellappa |
ECCV (3) | 2 |
| 2010 | Compressive Acquisition of Dynamic Scenes
Aswin C. Sankaranarayanan, Pavan Turaga, Richard G. Baraniuk, Rama Chellappa |
ECCV (1) | 2 |
| 2010 | The role of geometry in age estimationabstractUnderstanding and modeling of aging in human faces is an important problem in many real-world applications such as biometrics, authentication, and synthesis. In this paper, we consider the role of geometric attributes of faces, as described by a set of landmark points on the face, in age perception. Towards this end, we show that the space of landmarks can be interpreted as a Grassmann manifold. Then the problem of age estimation is posed as a problem of function estimation on the manifold. The warping of an average face to a given face is quantified as a velocity vector that transforms the average to a given face along a smooth geodesic in unit-time. This deformation is then shown to contain important information about the age of the face. We show in experiments that exploiting geometric cues in a principled manner provides comparable performance to several systems that utilize both geometric and textural cues. We show results on age estimation using the standard FG-Net dataset and a passport dataset which illustrate the effectiveness of the approach. Pavan Turaga, Soma Biswas, Rama Chellappa |
ICASSP | 1 |
| 2010 | Video Précis: Highlighting Diverse Aspects of VideosabstractSummarizing long unconstrained videos is gaining importance in surveillance, web-based video browsing, and video-archival applications. Summarizing a video requires one to identify key aspects that contain the essence of the video. In this paper, we propose an approach that optimizes two criteria that a video summary should embody. The first criterion, “coverage,” requires that the summary be able to represent the original video well. The second criterion, “diversity,” requires that the elements of the summary be as distinct from each other as possible. Given a user-specified summary length, we propose a cost function to measure the quality of a summary. The problem of generating a précis is then reduced to a combinatorial optimization problem of minimizing the proposed cost function. We propose an efficient method to solve the optimization problem. We demonstrate through experiments (on KTH data, unconstrained skating video, a surveillance video, and a YouTube home video) that optimizing the proposed criterion results in meaningful video summaries over a wide range of scenarios. Summaries thus generated are then evaluated using both quantitative measures and user studies. Nitesh Shroff, Pavan Turaga, Rama Chellappa |
IEEE Trans. Multim. | 2 |
| 2009 | Locally time-invariant models of human activities using trajectories on the grassmannianabstractHuman activity analysis is an important problem in computer vision with applications in surveillance and summarization and indexing of consumer content. Complex human activities are characterized by non-linear dynamics that make learning, inference and recognition hard. In this paper, we consider the problem of modeling and recognizing complex activities which exhibit time-varying dynamics. To this end, we describe activities as outputs of linear dynamic systems (LDS) whose parameters vary with time, or a time-varying linear dynamic system (TV-LDS). We discuss parameter estimation methods for this class of models by assuming that the parameters are locally time-invariant. Then, we represent the space of LDS models as a Grassmann manifold. Then, the TV-LDS model is defined as a trajectory on the Grassmann manifold. We show how trajectories on the Grassmannian can be characterized using appropriate distance metrics and statistical methods that reflect the underlying geometry of the manifold. This results in more expressive and powerful models for complex human activities. We demonstrate the strength of the framework for activity-based summarization of long videos and recognition of complex human actions on two datasets. Pavan Turaga, Rama Chellappa |
CVPR | 1 |
| 2009 | Unsupervised view and rate invariant clustering of video sequences
Pavan Turaga, Ashok Veeraraghavan, Rama Chellappa |
Comput. Vis. Image Underst. | 1 |
| 2008 | Statistical analysis on Stiefel and Grassmann manifolds with applications in computer visionabstractMany applications in computer vision and pattern recognition involve drawing inferences on certain manifold-valued parameters. In order to develop accurate inference algorithms on these manifolds we need to a) understand the geometric structure of these manifolds b) derive appropriate distance measures and c) develop probability distribution functions (pdf) and estimation techniques that are consistent with the geometric structure of these manifolds. In this paper, we consider two related manifolds - the Stiefel manifold and the Grassmann manifold, which arise naturally in several vision applications such as spatio-temporal modeling, affine invariant shape analysis, image matching and learning theory. We show how accurate statistical characterization that reflects the geometry of these manifolds allows us to design efficient algorithms that compare favorably to the state of the art in these very different applications. In particular, we describe appropriate distance measures and parametric and non-parametric density estimators on these manifolds. These methods are then used to learn class conditional densities for applications such as activity recognition, video based face recognition and shape classification. Pavan Turaga, Ashok Veeraraghavan, Rama Chellappa |
CVPR | 1 |
| 2008 | Learning action dictionaries from videoabstractSummarizing the contents of a video containing human activities is an important problem in computer vision and has important applications in automated surveillance systems. Summarizing a video requires one to identify and learn a 'vocabulary' of action-phrases corresponding to specific events and actions occurring in the video. We propose a generative model for dynamic scenes containing human activities as a composition of independent action-phrases - each of which is derived from an underlying vocabulary. Given a long video sequence, we propose a completely unsupervised approach to learn the vocabulary. Once the vocabulary is learnt, a video segment can be decomposed into a collection of phrases for summarization. We then describe methods to learn the correlations between activities and sequentiality of events. We also propose a novel method for building invariances to spatial transforms in the summarization scheme. Pavan Turaga, Rama Chellappa |
ICIP | 1 |
| 2008 | An ontology based approach for activity recognition from videoabstractRepresentation and recognition of human activities is an important problem for video surveillance and security applications. Considering the wide variety of settings in which surveillance systems are being deployed, it is necessary to create a common knowledge-base or ontology of human activities. Most current attempts at ontology design in computer vision for human activities have been empirical in nature. In this paper, we present a more systematic approach to address the problem of designing ontologies for visual activity recognition. We draw on general ontology design principles and adapt them to the specific domain of human activity ontologies. Then, we discuss qualitative evaluation principles and provide several examples from existing ontologies and how they can be improved upon. Finally, we demonstrate quantitatively in terms of recognition performance, the efficacy and validity of our approach for bank and airport tarmac surveillance domains. Umut Akdemir, Pavan Turaga, Rama Chellappa |
ACM Multimedia | 2 |
| 2008 | Machine Recognition of Human Activities: A SurveyabstractThe past decade has witnessed a rapid proliferation of video cameras in all walks of life and has resulted in a tremendous explosion of video content. Several applications such as content-based video annotation and retrieval, highlight extraction and video summarization require recognition of the activities occurring in the video. The analysis of human activities in videos is an area with increasingly important consequences from security and surveillance to entertainment and personal archiving. Several challenges at various levels of processing-robustness against errors in low-level processing, view and rate-invariant representations at midlevel processing and semantic representation of human activities at higher level processing-make this problem hard to solve. In this review paper, we present a comprehensive survey of efforts in the past couple of decades to address the problems of representation, recognition, and learning of human activities from video and related applications. We discuss the problem at two major levels of complexity: 1) "actions" and 2) "activities." "Actions" are characterized by simple motion patterns typically executed by a single human. "Activities" are more complex and involve coordinated actions among a small number of humans. We will discuss several approaches and classify them according to their ability to handle varying degrees of complexity as interpreted above. We begin with a discussion of approaches to model the simplest of action classes known as atomic or primitive actions that do not require sophisticated dynamical modeling. Then, methods to model actions with more complex dynamics are discussed. The discussion then leads naturally to methods for higher level representation of complex activities. Pavan Turaga, Rama Chellappa, V. S. Subrahmanian, Octavian Udrea |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | From Videos to Verbs: Mining Videos for Activities using a Cascade of Dynamical SystemsabstractClustering video sequences in order to infer and extract activities from a single video stream is an extremely important problem and has significant potential in video indexing, surveillance, activity discovery and event recognition. Clustering a video sequence into activities requires one to simultaneously recognize activity boundaries (activity consistent subsequences) and cluster these activity subsequences. In order to do this, we build a generative model for activities (in video) using a cascade of dynamical systems and show that this model is able to capture and represent a diverse class of activities. We then derive algorithms to learn the model parameters from a video stream and also show how a single video sequence may be clustered into different clusters where each cluster represents an activity. We also propose a novel technique to build affine, view, rate invariance of the activity into the distance metric for clustering. Experiments show that the clusters found by the algorithm correspond to semantically meaningful activities. Pavan Turaga, Ashok Veeraraghavan, Rama Chellappa |
CVPR | 1 |