Narendra Ahuja

dblp:30/3572 · DBLP profile ↗
← Back
329ranked-venue papers
25as first author
17since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 238 · 11 first-author · 11 since 2021Artificial intelligence and machine learning · 223 · 15 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-authorSystems, architecture and hardware · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 4
YearPublicationVenuePosition
2025 Potential Field Based Deep Metric Learning
abstract
Deep metric learning (DML) involves training a network to learn a semantically meaningful representation space. Many current approaches mine n-tuples of examples and model interactions within each tuplets. We present a novel, compositional DML model that instead of in tuples, represents the influence of each example (embedding) by a continuous potential field, and superposes the fields to obtain their combined global potential field. We use attractive/repulsive potential fields to represent interactions among embeddings from images of the same/different classes. Contrary to typical learning methods, where mutual influence of samples is proportional to their distance, we enforce reduction in such influence with distance, leading to a decaying field. We show that such decay helps improve performance on real world datasets with large intra-class variations and label noise. Like other proxy-based methods, we also use proxies to succinctly represent sub-populations of examples. We evaluate our method on three standard DML benchmarks-Cars-196, CUB-200-2011, and SOP datasets where it outperforms state-of-the-art baselines.
Shubhang Bhatnagar, Narendra Ahuja
CVPR2
2025 PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object Modeling
Hao Zhang 0122, Haolan Xu, Varun Jampani, Narendra Ahuja
ICCV5
2025 RGB-Only Supervised Camera Parameter Optimization in Dynamic Scenes
abstract
Although COLMAP has long remained the predominant method for camera parameter optimization in static scenes, it is constrained by its lengthy runtime and reliance on ground truth (GT) motion masks for application to dynamic scenes. Many efforts attempted to improve it by incorporating more priors as supervision such as GT focal length, motion masks, 3D point clouds, camera poses, and metric depth, which, however, are typically unavailable in casually captured RGB videos. In this paper, we propose a novel method for more accurate and efficient camera parameter optimization in dynamic scenes solely supervised by a single RGB video, dubbed $\textbf{\textit{ROS-Cam}}$. Our method consists of three key components: (1) Patch-wise Tracking Filters, to establish robust and maximally sparse hinge-like relations across the RGB video. (2) Outlier-aware Joint Optimization, for efficient camera parameter optimization by adaptive down-weighting of moving outliers, without reliance on motion priors. (3) A Two-stage Optimization Strategy, to enhance stability and optimization speed by a trade-off between the Softplus limits and convex minima in losses. We visually and numerically evaluate our camera estimates. To further validate accuracy, we feed the camera estimates into a 4D reconstruction method and assess the resulting 3D scenes, and rendered 2D RGB and depth maps. We perform experiments on 4 real-world datasets (NeRF-DS, DAVIS, iPhone, and TUM-dynamics) and 1 synthetic dataset (MPI-Sintel), demonstrating that our method estimates camera parameters more efficiently and accurately with a single RGB video as the only supervision.
Fang Li 0012, Hao Zhang 0122, Narendra Ahuja
NeurIPS3
2025 Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
abstract
We present Stable Part Diffusion 4D (SP4D), a framework for generating paired RGB and kinematic part videos from monocular inputs. Unlike conventional part segmentation methods that rely on appearance-based semantic cues, SP4D learns to produce kinematic parts --- structural components aligned with object articulation and consistent across views and time. SP4D adopts a dual-branch diffusion model that jointly synthesizes RGB frames and corresponding part segmentation maps. To simplify architecture and flexibly enable different part counts, we introduce a spatial color encoding scheme that maps part masks to continuous RGB-like images. This encoding allows the segmentation branch to share the latents VAE from the RGB branch, while enabling part segmentation to be recovered via straightforward post-processing. A Bidirectional Diffusion Fusion (BiDiFuse) module enhances cross-branch consistency, supported by a contrastive part consistency loss to promote spatial and temporal alignment of part predictions. We demonstrate that the generated 2D part maps can be lifted to 3D to derive skeletal structures and harmonic skinning weights with few manual adjustments. To train and evaluate SP4D, we construct KinematicParts20K, a curated dataset of over 20K rigged objects selected and processed from Objaverse XL, each paired with multi-view RGB and part video sequences. Experiments show that SP4D generalizes strongly to diverse scenarios, including real-world videos, novel generated objects, and rare articulated poses, producing kinematic-aware outputs suitable for downstream animation and motion-related tasks.
Hao Zhang 0122, Chun-Han Yao, Simon Donné, Narendra Ahuja, Varun Jampani
NeurIPS4
2025 PositiveCoOp: Rethinking Prompting Strategies for Multi-Label Recognition with Partial Annotations
abstract
Vision-language models (VLMs) like CLIP have been adapted for Multi-Label Recognition (MLR) with partial annotations by leveraging prompt-learning, where positive and negative prompts are learned for each class to associate their embeddings with class presence or absence in the shared vision-text feature space. While this approach improves MLR performance by relying on VLM priors, we hypothesize that learning negative prompts may be suboptimal, as the datasets used to train VLMs lack imagecaption pairs explicitly focusing on class absence. To analyze the impact of positive and negative prompt learning on MLR, we introduce PositiveCoOp and NegativeCoOp, where only one prompt is learned with VLM guidance while the other is replaced by an embedding vector learned directly in the shared feature space without relying on the text encoder. Through empirical analysis, we observe that negative prompts degrade MLR performance, and learning only positive prompts, combined with learned negative embeddings (PositiveCoOp), outperforms dual prompt learning approaches. Moreover, we quantify the performance benefits that prompt-learning offers over a simple vision-features-only baseline, observing that the baseline displays strong performance comparable to dual prompt learning approach (DualCoOp), when the proportion of missing labels is low, while requiring half the training compute and 16 times fewer parameters. Our code is available at https://github.com/Samyakkv'l/PositivcCoOp
Samyak Rawlekar, Shubhang Bhatnagar, Narendra Ahuja
WACV3
2024 CSL: Class-Agnostic Structure-Constrained Learning for Segmentation Including the Unseen
abstract
Addressing Out-Of-Distribution (OOD) Segmentation and Zero-Shot Semantic Segmentation (ZS3) is challenging, necessitating segmenting unseen classes. Existing strategies adapt the class-agnostic Mask2Former (CA-M2F) tailored to specific tasks. However, these methods cater to singular tasks, demand training from scratch, and we demonstrate certain deficiencies in CA-M2F, which affect performance. We propose the Class-Agnostic Structure-Constrained Learning (CSL), a plug-in framework that can integrate with existing methods, thereby embedding structural constraints and achieving performance gain, including the unseen, specifically OOD, ZS3, and domain adaptation (DA) tasks. There are two schemes for CSL to integrate with existing methods (1) by distilling knowledge from a base teacher network, enforcing constraints across training and inference phrases, or (2) by leveraging established models to obtain per-pixel distributions without retraining, appending constraints during the inference phase. Our soft assignment and mask split methodologies enhance OOD object segmentation. Empirical evaluations demonstrate CSL's prowess in boosting the performance of existing algorithms spanning OOD segmentation, ZS3, and DA segmentation, consistently transcending the state-of-art across all three tasks.
Hao Zhang 0122, Fang Li 0012, Lu Qi 0001, Ming-Hsuan Yang 0001, Narendra Ahuja
AAAI5
2024 Learning Implicit Representation for Reconstructing Articulated Objects
abstract
3D Reconstruction of moving articulated objects without additional information about object structure is a challenging problem. Current methods overcome such challenges by employing category-specific skeletal models. Consequently, they do not generalize well to articulated objects in the wild. We treat an articulated object as an unknown, semi-rigid skeletal structure surrounded by nonrigid material (e.g., skin). Our method simultaneously estimates the visible (explicit) representation (3D shapes, colors, camera parameters) and the underlying (implicit) skeletal representation, from motion cues in the object video without 3D supervision. Our implicit representation consists of four parts. (1) skeleton, which specifies which semi-rigid parts are connected. (2) Semi-rigid Part Assignment, which associates each surface vertex with a semi-rigid part. (3) Rigidity Coefficients, specifying the articulation of the local surface. (4) Time-Varying Transformations, which specify the skeletal motion and surface deformation parameters. We introduce an algorithm that uses these constraints as regularization terms and iteratively estimates both implicit and explicit representations. Our method is category-agnostic, thus eliminating the need for category-specific skeletons, we show that our method outperforms state-of-the-art across standard video datasets.
Hao Zhang 0122, Fang Li 0012, Samyak Rawlekar, Narendra Ahuja
ICLR4
2024 S3O: A Dual-Phase Approach for Reconstructing Dynamic Shape and Skeleton of Articulated Objects from Single Monocular Video
abstract
Reconstructing dynamic articulated objects from a singular monocular video is challenging, requiring joint estimation of shape, motion, and camera parameters from limited views. Current methods typically demand extensive computational resources and training time, and require additional human annotations such as predefined parametric models, camera poses, and key points, limiting their generalizability. We propose Synergistic Shape and Skeleton Optimization (S3O), a novel two-phase method that forgoes these prerequisites and efficiently learns parametric models including visible shapes and underlying skeletons. Conventional strategies typically learn all parameters simultaneously, leading to interdependencies where a single incorrect prediction can result in significant errors. In contrast, S3O adopts a phased approach: it first focuses on learning coarse parametric models, then progresses to motion learning and detail addition. This method substantially lowers computational complexity and enhances robustness in reconstruction from limited viewpoints, all without requiring additional annotations. To address the current inadequacies in 3D reconstruction from monocular video benchmarks, we collected the PlanetZoo dataset. Our experimental evaluations on standard benchmarks and the PlanetZoo dataset affirm that S3O provides more accurate 3D reconstruction, and plausible skeletons, and reduces the training time by approximately 60% compared to the state-of-the-art, thus advancing the state of the art in dynamic object reconstruction.
Hao Zhang 0122, Fang Li 0012, Samyak Rawlekar, Narendra Ahuja
ICML4
2024 Improving Multi-label Recognition using Class Co-Occurrence Probabilities
Samyak Rawlekar, Shubhang Bhatnagar, Vishnuvardhan Pogunulu Srinivasulu, Narendra Ahuja
ICPR (10)4
2024 Open-NeRF: Towards Open Vocabulary NeRF Decomposition
abstract
In this paper, we address the challenge of decomposing Neural Radiance Fields (NeRF) into objects from an open vocabulary, a critical task for object manipulation in 3D reconstruction and view synthesis. Current techniques for NeRF decomposition involve a trade-off between the flexibility of processing open-vocabulary queries and the accuracy of 3D segmentation. We present, Open-vocabulary Embedded Neural Radiance Fields (Open-NeRF), that leverage large-scale, off-the-shelf, segmentation models like the Segment Anything Model (SAM) and introduce an integrate-and-distill paradigm with hierarchical embeddings to achieve both the flexibility of open-vocabulary querying and 3D segmentation accuracy. Open-NeRF first utilizes large-scale foundation models to generate hierarchical 2D mask proposals from varying viewpoints. These proposals are then aligned via tracking approaches and integrated within the 3D space and subsequently distilled into the 3D field. This process ensures consistent recognition and granularity of objects from different viewpoints, even in challenging scenarios involving occlusion and indistinct features. Our experimental results show that the proposed Open-NeRF1outperforms state-of-the-art methods such as LERF [16] and FFD [18] in open-vocabulary scenarios. Open-NeRF offers a promising solution to NeRF decomposition, guided by open-vocabulary queries, enabling novel applications in robotics and vision-language interaction in open-world 3D scenes. Please find the code at https://github.com/haoz19/Open-NeRF.
Hao Zhang 0122, Fang Li 0012, Narendra Ahuja
WACV3
2023 Long-Distance Gesture Recognition Using Dynamic Neural Networks
abstract
Gestures form an important medium of communication between humans and machines. An overwhelming majority of existing gesture recognition methods are tailored to a scenario where humans and machines are located very close to each other. This short-distance assumption does not hold true for several types of interactions, for example gesture-based interactions with a floor cleaning robot or with a drone. Methods made for short-distance recognition are unable to perform well on long-distance recognition due to gestures occupying only a small portion of the input data. Their performance is especially worse in resource constrained settings where they are not able to effectively focus their limited compute on the gesturing subject. We propose a novel, accurate and efficient method for the recognition of gestures from longer distances. It uses a dynamic neural network to select features from gesture-containing spatial regions of the input sensor data for further processing. This helps the network focus on features important for gesture recognition while discarding background features early on, thus making it more compute efficient compared to other techniques. We demonstrate the performance of our method on the LD-ConGR long-distance dataset where it outperforms previous state-of-the-art methods on recognition accuracy and compute efficiency.
Shubhang Bhatnagar, Sharath Gopal, Narendra Ahuja, Liu Ren 0001
IROS3
2022 Detection of Covid-19 from Joint Time and Frequency Analysis of Speech, Breathing and Cough Audio
abstract
The distinct cough sounds produced by a variety of respiratory diseases suggest the potential for the development of a new class of audio bio-markers for the detection of COVID-19. Accurate audio biomarker-based COVID-19 tests would be inexpensive, readily scalable, and non-invasive. Audio biomarker screening could also be utilized in resource-limited settings prior to traditional diagnostic testing. Here we explore the possibility of leveraging three audio modalities: cough, breathing, and speech to determine COVID-19 status. We train a separate neural classification system on each modality, as well as a fused classification system on all three modalities together. Ablation studies are performed to understand the relationship between individual and collective performance of the modalities. Additionally, we analyze the extent to which temporal and spectral features contribute to COVID-19 status information contained in the audio signals.
John B. Harvill, Yash R. Wani, Moitreya Chatterjee, Mustafa Alam, David G. Beiser, David Chestek, Mark Hasegawa-Johnson, Narendra Ahuja
ICASSP8
2022 Learning Audio-Visual Dynamics Using Scene Graphs for Audio Source Separation
abstract
There exists an unequivocal distinction between the sound produced by a static source and that produced by a moving one, especially when the source moves towards or away from the microphone. In this paper, we propose to use this connection between audio and visual dynamics for solving two challenging tasks simultaneously, namely: (i) separating audio sources from a mixture using visual cues, and (ii) predicting the 3D visual motion of a sounding source using its separated audio. Towards this end, we present Audio Separator and Motion Predictor (ASMP) -- a deep learning framework that leverages the 3D structure of the scene and the motion of sound sources for better audio source separation. At the heart of ASMP is a 2.5D scene graph capturing various objects in the video and their pseudo-3D spatial proximities. This graph is constructed by registering together 2.5D monocular depth predictions from the 2D video frames and associating the 2.5D scene regions with the outputs of an object detector applied on those frames. The ASMP task is then mathematically modeled as the joint problem of: (i) recursively segmenting the 2.5D scene graph into several sub-graphs, each associated with a constituent sound in the input audio mixture (which is then separated) and (ii) predicting the 3D motions of the corresponding sound sources from the separated audio. To empirically evaluate ASMP, we present experiments on two challenging audio-visual datasets, viz. Audio Separation in the Wild (ASIW) and Audio Visual Event (AVE). Our results demonstrate that ASMP achieves a clear improvement in source separation quality, outperforming prior works on both datasets, while also estimating the direction of motion of the sound sources better than other methods.
Moitreya Chatterjee, Narendra Ahuja, Anoop Cherian
NeurIPS2
2021 A Hierarchical Variational Neural Uncertainty Model for Stochastic Video Prediction
abstract
Predicting the future frames of a video is a challenging task, in part due to the underlying stochastic real-world phenomena. Prior approaches to solve this task typically estimate a latent prior characterizing this stochasticity, however do not account for the predictive uncertainty of the (deep learning) model. Such approaches often derive the training signal from the mean-squared error (MSE) between the generated frame and the ground truth, which can lead to sub-optimal training, especially when the predictive uncertainty is high. Towards this end, we introduce Neural Uncertainty Quantifier (NUQ) - a stochastic quantification of the model’s predictive uncertainty, and use it to weigh the MSE loss. We propose a hierarchical, variational framework to derive NUQ in a principled manner using a deep, Bayesian graphical model. Our experiments on three benchmark stochastic video prediction datasets show that our proposed framework trains more effectively compared to the state-of-the-art models (especially when the training sets are small), while demonstrating better video generation quality and diversity against several evaluation metrics.
Moitreya Chatterjee, Narendra Ahuja, Anoop Cherian
ICCV2
2021 Visual Scene Graphs for Audio Source Separation
abstract
State-of-the-art approaches for visually-guided audio source separation typically assume sources that have characteristic sounds, such as musical instruments. These approaches often ignore the visual context of these sound sources or avoid modeling object interactions that may be useful to better characterize the sources, especially when the same object class may produce varied sounds from distinct interactions. To address this challenging problem, we propose Audio Visual Scene Graph Segmenter (AVSGS), a novel deep learning model that embeds the visual structure of the scene as a graph and segments this graph into subgraphs, each subgraph being associated with a unique sound obtained by co-segmenting the audio spectrogram. At its core, AVSGS uses a recursive neural network that emits mutually-orthogonal sub-graph embeddings of the visual graph using multi-head attention. These embeddings are used for conditioning an audio encoder-decoder towards source separation. Our pipeline is trained end-to-end via a self-supervised task consisting of separating audio sources using the visual graph from artificially mixed sounds.In this paper, we also introduce an “in the wild” video dataset for sound source separation that contains multiple non-musical sources, which we call Audio Separation in the Wild (ASIW). This dataset is adapted from the AudioCaps dataset, and provides a challenging, natural, and daily-life setting for source separation. Thorough experiments on the proposed ASIW and the standard MUSIC datasets demonstrate state-of-the-art sound separation performance of our method against recent prior approaches.
Moitreya Chatterjee, Jonathan Le Roux, Narendra Ahuja, Anoop Cherian
ICCV3
2021 Unsupervised 3D Pose Estimation for Hierarchical Dance Video Recognition *
abstract
Dance experts often view dance as a hierarchy of information, spanning low-level (raw images, image sequences), mid-levels (human poses and bodypart movements), and high-level (dance genre). We propose a Hierarchical Dance Video Recognition framework (HDVR). HDVR estimates 2D pose sequences, tracks dancers, and then simultaneously estimates corresponding 3D poses and 3D-to-2D imaging parameters, without requiring ground truth for 3D poses. Unlike most methods that work on a single person, our tracking works on multiple dancers, under occlusions. From the estimated 3D pose sequence, HDVR extracts body part movements, and therefrom dance genre. The resulting hierarchical dance representation is explainable to experts. To overcome noise and interframe correspondence ambiguities, we enforce spatial and temporal motion smoothness and photometric continuity over time. We use an LSTM network to extract 3D movement subsequences from which we recognize dance genre. For experiments, we have identified 154 movement types, of 16 body parts, and assembled a new University of Illinois Dance (UID) Dataset, containing 1143 video clips of 9 genres covering 30 hours, annotated with movement and genre labels. Our experimental results demonstrate that our algorithms outperform the state-of-the-art 3D pose estimation methods, which also enhances our dance recognition performance.
Xiaodan Hu, Narendra Ahuja
ICCV2
2021 Classification of COVID-19 from Cough Using Autoregressive Predictive Coding Pretraining and Spectral Data Augmentation
abstract
Serum and saliva-based testing methods have been crucial to slowing the COVID-19 pandemic, yet have been limited by slow throughput and cost.A system able to determine COVID-19 status from cough sounds alone would provide a low cost, rapid, and remote alternative to current testing methods.We explore the applicability of recent techniques such as pre-training and spectral augmentation in improving the performance of a neural cough classification system.We use Autoregressive Predictive Coding (APC) to pre-train a unidirectional LSTM on the COUGHVID dataset.We then generate our final model by finetuning added BLSTM layers on the DiCOVA challenge dataset.We perform various ablation studies to see how each component impacts performance and improves generalization with a small dataset.Our final system achieves an AUC of 85.35 and places third out of 29 entries in the DiCOVA challenge.
John B. Harvill, Yash R. Wani, Mark Hasegawa-Johnson, Narendra Ahuja, David G. Beiser, David Chestek
Interspeech4
2020 Low-level multiscale image segmentation and a benchmark for its evaluation
Emre Akbas, Narendra Ahuja
Comput. Vis. Image Underst.2
2020 Tracking Persons-of-Interest via Unsupervised Representation Adaptation
Jia-Bin Huang 0001, Jongwoo Lim, Yihong Gong, Jinjun Wang, Narendra Ahuja, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.6
2019 Fast and Accurate Image Super-Resolution with Deep Laplacian Pyramid Networks
abstract
Convolutional neural networks have recently demonstrated high-quality reconstruction for single image super-resolution. However, existing methods often require a large number of network parameters and entail heavy computational loads at runtime for generating high-accuracy super-resolution results. In this paper, we propose the deep Laplacian Pyramid Super-Resolution Network for fast and accurate image super-resolution. The proposed network progressively reconstructs the sub-band residuals of high-resolution images at multiple pyramid levels. In contrast to existing methods that involve the bicubic interpolation for pre-processing (which results in large feature maps), the proposed method directly extracts features from the low-resolution input space and thereby entails low computational loads. We train the proposed network with deep supervision using the robust Charbonnier loss functions and achieve high-quality image reconstruction. Furthermore, we utilize the recursive layers to share parameters across as well as within pyramid levels, and thus drastically reduce the number of parameters. Extensive quantitative and qualitative evaluations on benchmark datasets show that the proposed algorithm performs favorably against the state-of-the-art methods in terms of run-time and image quality.
Wei-Sheng Lai, Jia-Bin Huang 0001, Narendra Ahuja, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Joint Image Filtering with Deep Convolutional Networks
abstract
Joint image filters leverage the guidance image as a prior and transfer the structural details from the guidance image to the target image for suppressing noise or enhancing spatial resolution. Existing methods either rely on various explicit filter constructions or hand-designed objective functions, thereby making it difficult to understand, improve, and accelerate these filters in a coherent framework. In this paper, we propose a learning-based approach for constructing joint filters based on Convolutional Neural Networks. In contrast to existing methods that consider only the guidance image, the proposed algorithm can selectively transfer salient structures that are consistent with both guidance and target images. We show that the model trained on a certain type of data, e.g., RGB and depth images, generalizes well to other modalities, e.g., flash/non-Flash and RGB/NIR images. We validate the effectiveness of the proposed joint filter through extensive experimental evaluations with state-of-the-art methods.
Yijun Li 0001, Jia-Bin Huang 0001, Narendra Ahuja, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Clustering as physically inspired energy minimization
Huiguang Yang, Narendra Ahuja
Pattern Recognit.2
2018 DeepMVS: Learning Multi-View Stereopsis
abstract
We present DeepMVS, a deep convolutional neural network (ConvNet) for multi-view stereo reconstruction. Taking an arbitrary number of posed images as input, we first produce a set of plane-sweep volumes and use the proposed DeepMVS network to predict high-quality disparity maps. The key contributions that enable these results are (1) supervised pretraining on a photorealistic synthetic dataset, (2) an effective method for aggregating information across a set of unordered images, and (3) integrating multi-layer feature activations from the pre-trained VGG-19 network. We validate the efficacy of DeepMVS using the ETH3D Benchmark. Our results show that DeepMVS compares favorably against state-of-the-art conventional MVS algorithms and other ConvNet based methods, particularly for near-textureless regions and thin structures.
Kevin Matzen, Johannes Kopf 0001, Narendra Ahuja, Jia-Bin Huang 0001
CVPR4
2018 Coreset-Based Neural Network Compression
Abhimanyu Dubey, Moitreya Chatterjee, Narendra Ahuja
ECCV (7)3
2018 QL-Net: Quantized-by-LookUp CNN
abstract
Convolutional Neural Networks (CNNs) have achieved a state-of-the-art performance in the different computer vision tasks. However, CNN algorithms are computationally and power intensive, which makes them difficult to run on wearable and embedded systems. One way to address this constraint is to reduce the number of computational operations performed. Recently, several approaches addressed the problem of the computational complexity in the CNNs. Most of these methods, however, require a dedicated hardware. We propose a new method for the computation reduction in CNNs that substitutes Multiply and Accumulate (MAC) operations with a codebook lookup and can be executed on the generic hardware. The proposed method called QL-Net combines several concepts: (i) a codebook construction, (ii) a layer-wise retraining strategy, and (iii) a substitution of the MAC operations with the lookup of the convolution responses at inference time. The proposed QL-Net achieves a 98.6% accuracy on the MNIST dataset with a 5.8x reduction in runtime, when compared to MAC-based CNN model that achieved a 99.2% accuracy.
Kamila Abdiyeva, Kim-Hui Yap, Gang Wang 0012, Narendra Ahuja, Martin Lukac
ICARCV4
2018 Joint Estimation of Human Pose and Conversational Groups from Social Scenes
Jagannadan Varadarajan, Subramanian Ramanathan, Samuel Rota Bulò, Narendra Ahuja, Oswald Lanz, Elisa Ricci 0001
Int. J. Comput. Vis.4
2018 Superpixel Hierarchy
abstract
Superpixel segmentation has been one of the most important tasks in computer vision. In practice, an object can be represented by a number of segments at finer levels with consistent details or included in a surrounding region at coarser levels. Thus, a superpixel segmentation hierarchy is of great importance for applications that require different levels of image details. However, there is no method that can generate all scales of superpixels accurately in real time. In this paper, we propose the superhierarchy algorithm which is able to generate multi-scale superpixels as accurately as the state-of-the-art methods but with one to two orders of magnitude speed-up. The proposed algorithm can be directly integrated with recent efficient edge detectors to significantly outperform the state-of-the-art methods in terms of segmentation accuracy. Quantitative and qualitative evaluations on a number of applications demonstrate that the proposed algorithm is accurate and efficient in generating a hierarchy of superpixels.
Xing Wei 0001, Qingxiong Yang, Yihong Gong, Narendra Ahuja, Ming-Hsuan Yang 0001
IEEE Trans. Image Process.4
2017 On the Essence of Unsupervised Detection of Anomalous Motion in Surveillance Videos
Abdullah A. Abuolaim, Wee Kheng Leow, Jagannadan Varadarajan, Narendra Ahuja
CAIP (1)4
2017 Deep Laplacian Pyramid Networks for Fast and Accurate Super-Resolution
abstract
Convolutional neural networks have recently demonstrated high-quality reconstruction for single-image super-resolution. In this paper, we propose the Laplacian Pyramid Super-Resolution Network (LapSRN) to progressively reconstruct the sub-band residuals of high-resolution images. At each pyramid level, our model takes coarse-resolution feature maps as input, predicts the high-frequency residuals, and uses transposed convolutions for upsampling to the finer level. Our method does not require the bicubic interpolation as the pre-processing step and thus dramatically reduces the computational complexity. We train the proposed LapSRN with deep supervision using a robust Charbonnier loss function and achieve high-quality reconstruction. Furthermore, our network generates multi-scale predictions in one feed-forward pass through the progressive reconstruction, thereby facilitates resource-aware applications. Extensive quantitative and qualitative evaluations on benchmark datasets show that the proposed algorithm performs favorably against the state-of-the-art methods in terms of speed and accuracy.
Wei-Sheng Lai, Jia-Bin Huang 0001, Narendra Ahuja, Ming-Hsuan Yang 0001
CVPR3
2017 Robust Visual Tracking Using Oblique Random Forests
abstract
Random forest has emerged as a powerful classification technique with promising results in various vision tasks including image classification, pose estimation and object detection. However, current techniques have shown little improvements in visual tracking as they mostly rely on piece wise orthogonal hyperplanes to create decision nodes and lack a robust incremental learning mechanism that is much needed for online tracking. In this paper, we propose a discriminative tracker based on a novel incremental oblique random forest. Unlike conventional orthogonal decision trees that use a single feature and heuristic measures to obtain a split at each node, we propose to use a more powerful proximal SVM to obtain oblique hyperplanes to capture the geometric structure of the data better. The resulting decision surface is not restricted to be axis aligned and hence has the ability to represent and classify the input data better. Furthermore, in order to generalize to online tracking scenarios, we derive incremental update steps that enable the hyperplanes in each node to be updated recursively, efficiently and in a closed-form fashion. We demonstrate the effectiveness of our method using two large scale benchmark datasets (OTB-51 and OTB-100) and show that our method gives competitive results on several challenging cases by relying on simple HOG features as well as in combination with more sophisticated deep neural network based models. The implementations of the proposed random forest are available at https://github.com/ZhangLeUestc/ Incremental-Oblique-Random-Forest.
Le Zhang 0001, Jagannadan Varadarajan, Ponnuthurai N. Suganthan, Narendra Ahuja, Pierre Moulin
CVPR4
2017 Active Online Anomaly Detection Using Dirichlet Process Mixture Model and Gaussian Process Classification
abstract
We present a novel anomaly detection (AD) system for streaming videos. Different from prior methods that rely on unsupervised learning of clip representations, that are usually coarse in nature, and batch-mode learning, we propose the combination of two non-parametric models for our task: (i) Dirichlet process mixture models (DPMM) based modeling of object motion and directions in each cell, and (ii) Gaussian process based active learning paradigm involving labeling by a domain expert. Whereas conventional clip representation methods adopt quantizing only motion directions leading to a lossy, coarse representation that are inadequate, our clip representation approach results in fine grained clusters at each cell that model the scene activities (both direction and speed) more effectively. For active anomaly detection, we adapt a Gaussian Process framework to process incoming samples (video snippets) sequentially, seek labels for confusing or informative samples and and update the AD model online. Furthermore, the proposed video representation along with a novel query criterion to select informative samples for labeling that incorporates both exploration and exploitation criteria is proposed, and is found to outperform competing criteria on two challenging traffic scene datasets.
Jagannadan Varadarajan, Subramanian Ramanathan, Narendra Ahuja, Pierre Moulin, Jean-Marc Odobez
WACV3
2017 Clustering Through Hybrid Network Architecture With Support Vectors
abstract
In this paper, we propose a clustering algorithm based on a two-phased neural network architecture. We combine the strength of an autoencoderlike network for unsupervised representation learning with the discriminative power of a support vector machine (SVM) network for fine-tuning the initial clusters. The first network is referred as prototype encoding network, where the data reconstruction error is minimized in an unsupervised manner. The second phase, i.e., SVM network, endeavors to maximize the margin between cluster boundaries in a supervised way making use of the first output. Both the networks update the cluster centroids successively by establishing a topology preserving scheme like self-organizing map on the latent space of each network. Cluster fine-tuning is accomplished in a network structure by the alternate usage of the encoding part of both the networks. In the experiments, challenging data sets from two popular repositories with different patterns, dimensionality, and the number of clusters are used. The proposed hybrid architecture achieves comparatively better results both visually and analytically than the previous neural network-based approaches available in the literature.
Emrah Ergul, Nafiz Arica, Narendra Ahuja, Sarp Ertürk
IEEE Trans. Neural Networks Learn. Syst.3
2016 Detecting Migrating Birds at Night
abstract
Bird migration is a critical indicator of environmental health, biodiversity, and climate change. Existing techniques for monitoring bird migration are either expensive (e.g., satellite tracking), labor-intensive (e.g., moon watching), indirect and thus less accurate (e.g., weather radar), or intrusive (e.g., attaching geolocators on captured birds). In this paper, we present a vision-based system for detecting migrating birds in flight at night. Our system takes stereo videos of the night sky as inputs, detects multiple flying birds and estimates their orientations, speeds, and altitudes. The main challenge lies in detecting flying birds of unknown trajectories under high noise level due to the low-light environment. We address this problem by incorporating stereo constraints for rejecting physically implausible configurations and gathering evidence from two (or more) views. Specifically, we develop a robust stereo-based 3D line fitting algorithm for geometric verification and a deformable part response accumulation strategy for trajectory verification. We demonstrate the effectiveness of the proposed approach through quantitative evaluation of real videos of birds migrating at night collected with near-infrared cameras.
Jia-Bin Huang 0001, Rich Caruana, Andrew Farnsworth, Steve Kelling, Narendra Ahuja
CVPR5
2016 A Comparative Study for Single Image Blind Deblurring
abstract
Numerous single image blind deblurring algorithms have been proposed to restore latent sharp images under camera motion. However, these algorithms are mainly evaluated using either synthetic datasets or few selected real blurred images. It is thus unclear how these algorithms would perform on images acquired "in the wild" and how we could gauge the progress in the field. In this paper, we aim to bridge this gap. We present the first comprehensive perceptual study and analysis of single image blind deblurring using real-world blurred images. First, we collect a dataset of real blurred images and a dataset of synthetically blurred images. Using these datasets, we conduct a large-scale user study to quantify the performance of several representative state-of-the-art blind deblurring algorithms. Second, we systematically analyze subject preferences, including the level of agreement, significance tests of score differences, and rationales for preferring one method over another. Third, we study the correlation between human subjective scores and several full-reference and noreference image quality metrics. Our evaluation and analysis indicate the performance gap between synthetically blurred images and real blurred image and sheds light on future research in single image blind deblurring.
Wei-Sheng Lai, Jia-Bin Huang 0001, Narendra Ahuja, Ming-Hsuan Yang 0001
CVPR4
2016 Deep Joint Image Filtering
Yijun Li 0001, Jia-Bin Huang 0001, Narendra Ahuja, Ming-Hsuan Yang 0001
ECCV (4)3
2016 Unsupervised Visual Representation Learning by Graph-Based Consistent Constraints
Dong Li 0025, Wei-Chih Hung, Jia-Bin Huang 0001, Shengjin Wang, Narendra Ahuja, Ming-Hsuan Yang 0001
ECCV (4)5
2016 Tracking Persons-of-Interest via Adaptive Discriminative Features
Yihong Gong, Jia-Bin Huang 0001, Jongwoo Lim, Jinjun Wang, Narendra Ahuja, Ming-Hsuan Yang 0001
ECCV (5)6
2016 Robust Visual Tracking via Exclusive Context Modeling
abstract
In this paper, we formulate particle filter-based object tracking as an exclusive sparse learning problem that exploits contextual information. To achieve this goal, we propose the context-aware exclusive sparse tracker (CEST) to model particle appearances as linear combinations of dictionary templates that are updated dynamically. Learning the representation of each particle is formulated as an exclusive sparse representation problem, where the overall dictionary is composed of multiple group dictionaries that can contain contextual information. With context, CEST is less prone to tracker drift. Interestingly, we show that the popular L1 tracker is a special case of our CEST formulation. The proposed learning problem is efficiently solved using an accelerated proximal gradient method that yields a sequence of closed form updates. To make the tracker much faster, we reduce the number of learning problems to be solved by using the dual problem to quickly and systematically rank and prune particles in each frame. We test our CEST tracker on challenging benchmark sequences that involve heavy occlusion, drastic illumination changes, and large pose variations. Experimental results show that CEST consistently outperforms state-of-the-art trackers.
Tianzhu Zhang 0001, Bernard Ghanem, Si Liu 0001, Changsheng Xu, Narendra Ahuja
IEEE Trans. Cybern.5
2016 Temporally coherent completion of dynamic video
abstract
We present an automatic video completion algorithm that synthesizes missing regions in videos in a temporally coherent fashion. Our algorithm can handle dynamic scenes captured using a moving camera. State-of-the-art approaches have difficulties handling such videos because viewpoint changes cause image-space motion vectors in the missing and known regions to be inconsistent. We address this problem by jointly estimating optical flow and color in the missing regions. Using pixel-wise forward/backward flow fields enables us to synthesize temporally coherent colors. We formulate the problem as a non-parametric patch-based optimization. We demonstrate our technique on numerous challenging videos.
Jia-Bin Huang 0001, Sing Bing Kang, Narendra Ahuja, Johannes Kopf 0001
ACM Trans. Graph.3
2015 Single image super-resolution from transformed self-exemplars
abstract
Self-similarity based super-resolution (SR) algorithms are able to produce visually pleasing results without extensive training on external databases. Such algorithms exploit the statistical prior that patches in a natural image tend to recur within and across scales of the same image. However, the internal dictionary obtained from the given image may not always be sufficiently expressive to cover the textural appearance variations in the scene. In this paper, we extend self-similarity based SR to overcome this drawback. We expand the internal patch search space by allowing geometric variations. We do so by explicitly localizing planes in the scene and using the detected perspective geometry to guide the patch search process. We also incorporate additional affine transformations to accommodate local shape variations. We propose a compositional model to simultaneously handle both types of transformations. We extensively evaluate the performance in both urban and natural scenes. Even without using any external training databases, we achieve significantly superior results on urban scenes, while maintaining comparable performance on natural scenes as other state-of-the-art SR algorithms.
Jia-Bin Huang 0001, Abhishek Singh 0002, Narendra Ahuja
CVPR3
2015 Structural Sparse Tracking
abstract
Sparse representation has been applied to visual tracking by finding the best target candidate with minimal reconstruction error by use of target templates. However, most sparse representation based trackers only consider holistic or local representations and do not make full use of the intrinsic structure among and inside target candidates, thereby making the representation less effective when similar objects appear or under occlusion. In this paper, we propose a novel Structural Sparse Tracking (SST) algorithm, which not only exploits the intrinsic relationship among target candidates and their local patches to learn their sparse representations jointly, but also preserves the spatial layout structure among the local patches inside each target candidate. We show that our SST algorithm accommodates most existing sparse trackers with the respective merits. Both qualitative and quantitative evaluations on challenging benchmark image sequences demonstrate that the proposed SST algorithm performs favorably against several state-of-the-art methods.
Tianzhu Zhang 0001, Si Liu 0001, Changsheng Xu, Shuicheng Yan, Bernard Ghanem, Narendra Ahuja, Ming-Hsuan Yang 0001
CVPR6
2015 On the Equivalence of Moving Entrance Pupil and Radial Distortion for Camera Calibration
abstract
Radial distortion for ordinary (non-fisheye) camera lenses has traditionally been modeled as an infinite series function of radial location of an image pixel from the image center. While there has been enough empirical evidence to show that such a model is accurate and sufficient for radial distortion calibration, there has not been much analysis on the geometric/physical understanding of radial distortion from a camera calibration perspective. In this paper, we show using a thick-lens imaging model, that the variation of entrance pupil location as a function of incident image ray angle is directly responsible for radial distortion in captured images. Thus, unlike as proposed in the current state-of-the-art in camera calibration, radial distortion and entrance pupil movement are equivalent and need not be modeled together. By modeling only entrance pupil motion instead of radial distortion, we achieve two main benefits, first, we obtain comparable if not better pixel re-projection error than traditional methods, second, and more importantly, we directly back-project a radially distorted image pixel along the true image ray which formed it. Using a thick-lens setting, we show that such a back-projection is more accurate than the two-step method of undistorting an image pixel and then back-projecting it. We have applied this calibration method to the problem of generative depth-from-focus using focal stack to get accurate depth estimates.
Avinash Kumar 0001, Narendra Ahuja
ICCV2
2015 Uncovering Interactions and Interactors: Joint Estimation of Head, Body Orientation and F-Formations from Surveillance Videos
abstract
We present a novel approach for jointly estimating targets' head, body orientations and conversational groups called F-formations from a distant social scene (e.g., a cocktail party captured by surveillance cameras). Differing from related works that have (i) coupled head and body pose learning by exploiting the limited range of orientations that the two can jointly take, or (ii) determined F-formations based on the mutual head (but not body) orientations of interactors, we present a unified framework to jointly infer both (i) and (ii). Apart from exploiting spatial and orientation relationships, we also integrate cues pertaining to temporal consistency and occlusions, which are beneficial while handling low-resolution data under surveillance settings. Efficacy of the joint inference framework reflects via increased head, body pose and F-formation estimation accuracy over the state-of-the-art, as confirmed by extensive experiments on two social datasets.
Elisa Ricci 0001, Jagannadan Varadarajan, Subramanian Ramanathan, Samuel Rota Bulò, Narendra Ahuja, Oswald Lanz
ICCV5
2015 Action Recognition Using Discriminative Structured Trajectory Groups
abstract
In this paper, we develop a novel framework for action recognition in videos. The framework is based on automatically learning the discriminative trajectory groups that are relevant to an action. Different from previous approaches, our method does not require complex computation for graph matching or complex latent models to localize the parts. We model a video as a structured bag of trajectory groups with latent class variables. We model action recognition problem in a weakly supervised setting and learn discriminative trajectory groups by employing multiple instance learning (MIL) based Support Vector Machine (SVM) using pre-computed kernels. The kernels depend on the spatio-temporal relationship between the extracted trajectory groups and their associated features. We demonstrate both quantitatively and qualitatively that the classification performance of our proposed method is superior to baselines and several state-of-the-art approaches on three challenging standard benchmark datasets.
Indriyati Atmosukarto, Narendra Ahuja, Bernard Ghanem
WACV2
2015 Learning ramp transformation for single image super-resolution
Abhishek Singh 0002, Narendra Ahuja
Comput. Vis. Image Underst.2
2015 Constant Time Median and Bilateral Filtering
Qingxiong Yang, Narendra Ahuja, Kar-Han Tan
Int. J. Comput. Vis.2
2015 Robust Visual Tracking Via Consistent Low-Rank Sparse Learning
Tianzhu Zhang 0001, Si Liu 0001, Narendra Ahuja, Ming-Hsuan Yang 0001, Bernard Ghanem
Int. J. Comput. Vis.3
2015 Efficient and Robust Specular Highlight Removal
abstract
A robust and effective specular highlight removal method is proposed in this paper. It is based on a key observation--the maximum fraction of the diffuse colour component in diffuse local patches in colour images changes smoothly. The specular pixels can thus be treated as noise in this case. This property allows the specular highlights to be removed in an image denoising fashion: an edge-preserving low-pass filter (e.g., the bilateral filter) can be used to smooth the maximum fraction of the colour components of the original image to remove the noise contributed by the specular pixels. Recent developments in fast bilateral filtering techniques enable the proposed method to run over 200× faster than state-of-the-art techniques on a standard CPU and differentiates it from previous work.
Qingxiong Yang, Jinhui Tang 0001, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Super-Resolution Using Sub-Band Self-Similarity
Abhishek Singh 0002, Narendra Ahuja
ACCV (2)2
2014 Generalized Pupil-centric Imaging and Analytical Calibration for a Non-frontal Camera
abstract
We consider the problem of calibrating a small field of view central perspective non-frontal camera whose lens and sensor planes may not be parallel to each other. This can be due to manufacturing defects or intentional tilting. Thus, as such all cameras can be modeled as being non-frontal with varying degrees. There are two approaches to model non- frontal cameras. The first one based on rotation parameterization of sensor non-frontalness/tilt increases the number of calibration parameters, thus requiring heuristics to initialize a few calibration parameters for the final non-linear optimization step. Additionally, for this parameterization, while it has been shown that pupil-centric imaging model leads to more accurate rotation estimates than a thin-lens imaging model, it has only been developed for a single axis lens-sensor tilt. But, in real cameras we can have arbitrary tilt. The second approach based on decentering distortion modeling is approximate as it can only handle small tilts and cannot explicitly estimate the sensor tilt. In this paper, we focus on rotation based non-frontal camera calibration and address the aforementioned problems of over-parameterization and inadequacy of existing pupil-centric imaging model. We first derive a generalized pupil-centric imaging model for arbitrary axis lens-sensor tilt. We then derive an analytical solution, in this setting, for a subset of calibration parameters including sensor rotation angles as a function of center of radial distortion (CoD). A radial alignment based constraint is then proposed to computationally estimate CoD leveraging on the proposed analytical solution. Our analytical technique also estimates pupil-centric parameters of entrance pupil location and optical focal length, which have typically been done optically. Given these analytical and computational calibration parameter estimates, we initialize the non-linear calibration optimization for a set of synthetic and real data captured from a non-frontal camera and show reduced pixel re-projection and undistortion errors compared to state of the art techniques in rotation and decentering based approaches to non-frontal camera calibration.
Avinash Kumar 0001, Narendra Ahuja
CVPR2
2014 Robust Orthonormal Subspace Learning: Efficient Recovery of Corrupted Low-Rank Matrices
abstract
Low-rank matrix recovery from a corrupted observation has many applications in computer vision. Conventional methods address this problem by iterating between nuclear norm minimization and sparsity minimization. However, iterative nuclear norm minimization is computationally prohibitive for large-scale data (e.g., video) analysis. In this paper, we propose a Robust Orthogonal Subspace Learning (ROSL) method to achieve efficient low-rank recovery. Our intuition is a novel rank measure on the low-rank matrix that imposes the group sparsity of its coefficients under orthonormal subspace. We present an efficient sparse coding algorithm to minimize this rank measure and recover the low-rank matrix at quadratic complexity of the matrix size. We give theoretical proof to validate that this rank measure is lower bounded by nuclear norm and it has the same global minimum as the latter. To further accelerate ROSL to linear complexity, we also describe a faster version (ROSL+) empowered by random sampling. Our extensive experiments demonstrate that both ROSL and ROSL+ provide superior efficiency against the state-of-the-art methods at the same level of recovery accuracy.
Xianbiao Shu, Fatih Porikli, Narendra Ahuja
CVPR3
2014 Super-resolving Noisy Images
abstract
Our goal is to obtain a noise-free, high resolution (HR) image, from an observed, noisy, low resolution (LR) image. The conventional approach of preprocessing the image with a denoising algorithm, followed by applying a super-resolution (SR) algorithm, has an important limitation: Along with noise, some high frequency content of the image (particularly textural detail) is invariably lost during the denoising step. This 'denoising loss' restricts the performance of the subsequent SR step, wherein the challenge is to synthesize such textural details. In this paper, we show that high frequency content in the noisy image (which is ordinarily removed by denoising algorithms) can be effectively used to obtain the missing textural details in the HR domain. To do so, we first obtain HR versions of both the noisy and the denoised images, using a patch-similarity based SR algorithm. We then show that by taking a convex combination of orientation and frequency selective bands of the noisy and the denoised HR images, we can obtain a desired HR image where (i) some of the textural signal lost in the denoising step is effectively recovered in the HR domain, and (ii) additional textures can be easily synthesized by appropriately constraining the parameters of the convex combination. We show that this part-recovery and part-synthesis of textures through our algorithm yields HR images that are visually more pleasing than those obtained using the conventional processing pipeline. Furthermore, our results show a consistent improvement in numerical metrics, further corroborating the ability of our algorithm to recover lost signal.
Abhishek Singh 0002, Fatih Porikli, Narendra Ahuja
CVPR3
2014 Partial Occlusion Handling for Visual Tracking via Robust Part Matching
abstract
Part-based visual tracking is advantageous due to its ro-bustness against partial occlusion. However, how to effec-tively exploit the confidence scores of individual parts to construct a robust tracker is still a challenging problem. In this paper, we address this problem by simultaneously matching parts in each of multiple frames, which is realized by a locality-constrained low-rank sparse learning method that establishes multi-frame part correspondences through optimization of partial permutation matrices. The proposed part matching tracker (PMT) has a number of attractive properties. (1) It exploits the spatial-temporal locality-constrained property for robust part matching. (2) It match-es local parts from multiple frames jointly by considering their low-rank and sparse structure information, which can effectively handle part appearance variations due to occlu-sion or noise. (3) The proposed PMT model has the inbuilt mechanism of leveraging multi-mode target templates, so that the dilemma of template updating when encountering occlusion in tracking can be better handled. This contrasts with existing methods that only do part matching between a pair of frames. We evaluate PMT and compare with 10 pop-ular state-of-the-art methods on challenging benchmarks. Experimental results show that PMT consistently outperfor-m these existing trackers. 1.
Tianzhu Zhang 0001, Kui Jia, Changsheng Xu, Yi Ma 0001, Narendra Ahuja
CVPR5
2014 Towards accurate and robust cross-ratio based gaze trackers through learning from simulation
abstract
Cross-ratio (CR) based methods offer many attractive properties for remote gaze estimation using a single camera in an uncalibrated setup by exploiting invariance of a plane projectivity. Unfortunately, due to several simplification assumptions, the performance of CR-based eye gaze trackers decays significantly as the subject moves away from the calibration position. In this paper, we introduce an adaptive homography mapping for achieving gaze prediction with higher accuracy at the calibration position and more robustness under head movements. This is achieved with a learning-based method for compensating both spatially-varying gaze errors and head pose dependent errors simultaneously in a unified framework. The model of adaptive homography is trained offline using simulated data, saving a tremendous amount of time in data collection. We validate the effectiveness of the proposed approach using both simulated and real data from a physical setup. We show that our method compares favorably against other state-of-the-art CR based methods.
Jia-Bin Huang 0001, Qin Cai, Zicheng Liu 0001, Narendra Ahuja, Zhengyou Zhang
ETRA4
2014 Non-local compressive sampling recovery
abstract
Compressive sampling (CS) aims at acquiring a signal at a sampling rate below the Nyquist rate by exploiting prior knowledge that a signal is sparse or correlated in some domain. Despite the remarkable progress in the theory of CS, the sampling rate on a single image required by CS is still very high in practice. In this paper, a non-local compressive sampling (NLCS) recovery method is proposed to further reduce the sampling rate by exploiting non-local patch correlation and local piecewise smoothness present in natural images. Two non-local sparsity measures, i.e., non-local wavelet sparsity and non-local joint sparsity, are proposed to exploit the patch correlation in NLCS. An efficient iterative algorithm is developed to solve the NLCS recovery problem, which is shown to have stable convergence behavior in experiments. The experimental results show that our NLCS significantly improves the state-of-the-art of image compressive sampling.
Xianbiao Shu, Jianchao Yang, Narendra Ahuja
ICCP3
2014 Improving head and body pose estimation through semi-supervised manifold alignment
abstract
In this paper, we explore the use of a semi-supervised manifold alignment method for domain adaptation in the context of human body and head pose estimation in videos. We build upon an existing state-of-the-art system that leverages on external labelled datasets for the body and head features, and on the unlabelled test data with weak velocity labels to do a coupled estimation of the body and head pose. While this previous approach showed promising results, the learning of the underlying manifold structure of the features in the train and target data and the need to align them were not explored despite the fact that the pose features between two datasets may vary according to the scene, e.g. due to different camera point of view or perspective. In this paper, we propose to use a semi-supervised manifold alignment method to bring the train and target samples closer within the resulting embedded space. To this end, we consider an adaptation set from the target data and rely on (weak) labels, given for example by the velocity direction whenever they are reliable. These labels, along with the training labels are used to bias the manifold distance within each manifold and to establish correspondences for alignment.
Alexandre Heili, Jagannadan Varadarajan, Bernard Ghanem, Narendra Ahuja, Jean-Marc Odobez
ICIP4
2014 Generalized Radial Alignment Constraint for Camera Calibration
abstract
In camera calibration, the radial alignment constraint (RAC) has been proposed as a technique to obtain closed form solution to calibration parameters when the image distortion is purely radial about an axis normal to the sensor plane. But, in real images this normality assumption might be violated due to manufacturing limitations or intentional sensor tilt. A misaligned optic axis results in traditional formulation of RAC not holding for real images leading to calibration errors. In this paper, we propose a generalized radial alignment constraint (gRAC), which relaxes the optic axis-sensor normality constraint by explicitly modeling their configuration via rotation parameters which form a part of camera calibration parameter set. We propose a new analytical solution to solve the gRAC for a subset of calibration parameters. We discuss the resulting ambiguities in the analytical approach and propose methods to overcome them. The analytical solution is then used to compute the intersection of optic axis and the sensor about which overall distortion is indeed radial. Finally, the analytical estimates from gRAC are used to initialize the nonlinear refinement of calibration parameters. Using simulated and real data, we show the correctness of the proposed gRAC and the analytical solution in achieving accurate camera calibration.
Avinash Kumar 0001, Narendra Ahuja
ICPR2
2014 Non-frontal Camera Calibration Using Focal Stack Imagery
abstract
A non-frontal camera has its lens and sensor plane misaligned either due to manufacturing limitations or an intentional tilting as in tilt-shift cameras. Under ideal perspective imaging, a geometric calibration of tilt is impossible as tilt parameters are correlated with the principal point location parameter. In other words, there are infinite combinations of principal point and sensor tilt parameters such that the perspective imaging equations are satisfied equally well. Previously, the non-frontal calibration problem (including sensor tilt estimation) has been solved by introducing constraints to align the principal point with the center of radial distortion. In this paper, we propose an additional constraint which incorporates image blur/defocus present in non-frontal camera images into the calibration framework. Specifically, it has earlier been shown that a non-frontal camera rotating about its center of projection captures images with varying focus. This stack of images is referred to as a focal stack. Given a focal stack of a known checkerboard (CB) pattern captured from a non-frontal camera, we combine geometric re-projection error and image bur error computed from current estimate of sensor tilt as the calibration optimization criteria. We show that the combined technique outperforms geometry-only methods while also additionally yielding blur kernel estimates at CB corners.
Avinash Kumar 0001, Narendra Ahuja
ICPR2
2014 Sub-band Energy Constraints for Self-Similarity Based Super-resolution
abstract
In this paper, we propose a new self-similarity based single image super-resolution (SR) algorithm that is able to better synthesize fine textural details of the image. Conventional self-similarity based SR typically uses scaled down version(s) of the given image to first build a dictionary of low-resolution (LR) and high-resolution (HR) image patches, which is then used to predict the HR patches for each LR patch of the given image. However, metrics like pixel wise sum of squared differences (L2distance) make it difficult to find matches for high frequency textured patches in the dictionary. Textural details are thus often smoothed out in the final image. In this paper, we propose a method to compensate for this loss of textural detail. Our algorithm uses the responses of a bank of orientation selective band pass filters to represent texture instead of the spatial variation of intensity values directly. Specifically, we use the energies contained in different sub-bands of an image patch to separate different types of details of a texture, which we then impose as additional priors on the patches of the super-resolved image. Our experiments show that for each patch, the low energy sub-bands (which correspond to fine textural details) get severely attenuated during conventional L2distance based SR. We propose a method to learn this attenuation of sub-band energies in the patches, using scaled down version(s) of the given image itself (without requiring external training databases), and thus propose a way of compensating for the energy loss in these sub-bands. We demonstrate that as a consequence, our SR results appear richer in texture and closer to the ground truth as compared to several other state-of-the-art methods.
Abhishek Singh 0002, Narendra Ahuja
ICPR2
2014 Low-Level Hierarchical Multiscale Segmentation Statistics of Natural Images
abstract
This paper is aimed at obtaining the statistics as a probabilistic model pertaining to the geometric, topological and photometric structure of natural images. The image structure is represented by its segmentation graph derived from the low-level hierarchical multiscale image segmentation. We first estimate the statistics of a number of segmentation graph properties from a large number of images. Our estimates confirm some findings reported in the past work, as well as provide some new ones. We then obtain a Markov random field based model of the segmentation graph which subsumes the observed statistics. To demonstrate the value of the model and the statistics, we show how its use as a prior impacts three applications: image classification, semantic image segmentation and object detection.
Emre Akbas, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Automatic segmentation of granular objects in images: Combining local density clustering and gradient-barrier watershed
Huiguang Yang, Narendra Ahuja
Pattern Recognit.2
2014 Image completion using planar structure guidance
abstract
We propose a method for automatically guiding patch-based image completion using mid-level structural cues. Our method first estimates planar projection parameters, softly segments the known region into planes, and discovers translational regularity within these planes. This information is then converted into soft constraints for the low-level completion algorithm by defining prior probabilities for patch offsets and transformations. Our method handles multiple planes, and in the absence of any detected planes falls back to a baseline fronto-parallel image completion algorithm. We validate our technique through extensive comparisons with state-of-the-art algorithms on a variety of scenes.
Jia-Bin Huang 0001, Sing Bing Kang, Narendra Ahuja, Johannes Kopf 0001
ACM Trans. Graph.3
2013 A Topic Model Approach to Representing and Classifying Football Plays
abstract
We address the problem of modeling and classifying American Football offense \nteams’ plays in video, a challenging example of group activity analysis. Automatic play \nclassification will allow coaches to infer patterns and tendencies of opponents more ef- \nficiently, resulting in better strategy planning in a game. We define a football play as a \nunique combination of player trajectories. To this end, we develop a framework that uses \nplayer trajectories as inputs to MedLDA, a supervised topic model. The joint maximiza- \ntion of both likelihood and inter-class margins of MedLDA in learning the topics allows \nus to learn semantically meaningful play type templates, as well as, classify different \nplay types with 70% average accuracy. Furthermore, this method is extended to analyze \nindividual player roles in classifying each play type. We validate our method on a large \ndataset comprising 271 play clips from real-world football games, which will be made \npublicly available for future comparisons.
Jagannadan Varadarajan, Indriyati Atmosukarto, Shaunak Ahuja, Bernard Ghanem, Narendra Ahuja
BMVC5
2013 A generative focus measure with application to omnifocus imaging
abstract
Given a stack of registered images acquired using a range of focus settings (focal stack images), we propose a new focus measure to identify the most focused image. Although, most of the paper is concerned with the new focus measure, for evaluation purposes, we will present it in the context of an application to generating omnifocus images. An omnifocus image is the composite image in which each pixel is selected form the frame in the stack in which it appears to be in best focus. Conventional focus measures usually maximize some measure of image gradient in a window. They tend to fail when one of the edges of the window lies near the boundary of an intensity edge, or the pixel is near other complex edge patterns. This leads to the misidentification of the focused frame and formation of artifacts in omnifocused image. Our proposed measure does not attempt to identify the focused frame by calculating the degree of defocus, like the gradient based methods. Rather, it hypothesizes that a specific frame is in focus and then validates or rejects this hypothesis by recreating the defocused frames in the vicinity, and comparing them with the observed de-focused frames. This forward generative process leads to correct focus frame selection in regions where typical measures fail. This is because the conventional measures try to identify the focused frame from its distorted version which is the result of a complex convolution process. This involves a backward estimation for a many-to-one transformation. On the other hand, the generation of defocused frames from a hypothesized focused frame is more accurate since it involves applying an operator in the forward direction. We analytically show that under ideal imaging conditions, the proposed focus measure is unimodal in nature. This makes the search for the best focused image unambiguous. We evaluate our focus measure by generating omnifocus images from real focal stack images, and show that it performs better than all the existing focus measures.
Avinash Kumar 0001, Narendra Ahuja
ICCP2
2013 Transformation guided image completion
abstract
In this paper, we describe a new interactive image completion system that allows users to easily specify various forms of mid-level structures in the image. Our system supports the specification of four basic symmetric types: reflection, translation, rotation, and glide. The user inputs are automatically converted into guidance maps that encode possible candidate shifts and, indirectly, local transformations of rotation and scale. These guidance maps are used in conjunction with a color matching cost for image completion. We show that our system is capable of handling a variety of challenging examples.
Jia-Bin Huang 0001, Johannes Kopf 0001, Narendra Ahuja, Sing Bing Kang
ICCP3
2013 Low-Rank Sparse Coding for Image Classification
abstract
In this paper, we propose a low-rank sparse coding (LRSC) method that exploits local structure information among features in an image for the purpose of image-level classification. LRSC represents densely sampled SIFT descriptors, in a spatial neighborhood, collectively as low-rank, sparse linear combinations of code words. As such, it casts the feature coding problem as a low-rank matrix learning problem, which is different from previous methods that encode features independently. This LRSC has a number of attractive properties. (1) It encourages sparsity in feature codes, locality in codebook construction, and low-rankness for spatial consistency. (2) LRSC encodes local features jointly by considering their low-rank structure information, and is computationally attractive. We evaluate the LRSC by comparing its performance on a set of challenging benchmarks with that of 7 popular coding and other state-of-the-art methods. Our experiments show that by representing local features jointly, LRSC not only outperforms the state-of-the-art in classification accuracy but also improves the time complexity of methods that use a similar sparse linear representation model for feature coding.
Tianzhu Zhang 0001, Bernard Ghanem, Si Liu 0001, Changsheng Xu, Narendra Ahuja
ICCV5
2013 Motion-based background subtraction and panoramic mosaicing for freight train analysis
abstract
We propose a new motion-based background removal technique which along with panoramic mosaicing forms the core of a vision system we have developed for analyzing the loading efficiency of intermodal freight trains. This analysis is critical for estimating the aerodynamic drag caused by air gaps present between loads in freight trains. The novelty of our background removal technique lies in using conventional motion estimates to design a cost function which can handle challenging textureless background regions, e.g. clear blue sky. Supplemented with domain knowledge, we have built a system which has outperformed some recent background removal methods applied to our problem. We also build an orthographic mosaic of the freight train allowing identification of load types and gap lengths between them. The complete system has been installed near Sibley, Missouri, US and processes about 20-30 (5-10 GB/train video data depending on train length) trains per day with high accuracy.
Avinash Kumar 0001, John M. Hart, Narendra Ahuja
ICIP3
2013 Single image super-resolution using adaptive domain transformation
abstract
In this paper we propose a new image domain prior term for regularizing the super-resolution reconstruction algorithm. This term encourages preserving the local ramp structure around edges, in the reconstruction algorithm. Ramp at a pixel is defined as the steepest sequence of monotonically increasing (or decreasing) pixels among all feasible directions around the pixel. As described in previous work, ramp based modeling is a richer characterization of local image structure than conventional gradients. Our proposed ramp-preserving constraint image is obtained by first running an accurate segmentation algorithm (which is itself obtained by ramp based modeling) on the low resolution image. We then perform a domain transformation of the pixels belonging to the steepest ramps at the edge pixels, in order to preserve sharpness. The resulting non-uniformly spaced image is then upscaled to a uniform, high resolution grid, using an edge preserving non-uniform interpolation scheme. This image is then used both as the prior constraint as well as the initial guess for the iterative super-resolution reconstruction algorithm. Our results compare favorably to the classical back-projection algorithm as well as newer methods which use learning based gradient domain priors.
Abhishek Singh 0002, Narendra Ahuja
ICIP2
2013 On stochastic gradient descent and quadratic mutual information for image registration
abstract
Mutual information (MI) is quite popular as a cost function for intensity based registration of images due to its ability to handle highly non-linear relationships between intensities of the two images. More recently, quadratic mutual information (QMI) has been proposed as an alternative measure that computes Euclidean distance instead of KL divergence between the joint and the product of the marginal densities of pixel intensities. In this paper, we examine the conditions under which QMI is advantageous over the classical MI measure, for the image registration problem. We show that QMI is a better cost function to use for optimization methods such as stochastic gradient descent. We show that the QMI cost function remains much smoother than the classical MI measure on stochastic subsampling of the image data. As a consequence, QMI has a higher probability of convergence, even for larger degrees of initial misalignment of the images.
Abhishek Singh 0002, Narendra Ahuja
ICIP2
2013 Rotation-invariant texture recognition by rotation compensation and wavelet analysis
abstract
A new rotation-invariant wavelet-based texture recognition scheme is proposed. In the previous rotation-invariant approaches, the focus is on adapting the wavelet transform or filter to rotated texture. In our approach, instead, we estimate the rotation of the texture with respect to some reference orientation, and then rotate the texture image back to the reference orientation before applying the wavelet analysis to extract features. With such rotation compensation, even very simple features (such as 1-level DWT and the subband energy) can be effective in achieving high classification accuracy as we demonstrate through our experiments.
Huiguang Yang, Narendra Ahuja
ICIP2
2013 Modeling dynamic swarms
Bernard Ghanem, Narendra Ahuja
Comput. Vis. Image Underst.2
2013 Robust Visual Tracking via Structured Multi-Task Sparse Learning
Tianzhu Zhang 0001, Bernard Ghanem, Si Liu 0001, Narendra Ahuja
Int. J. Comput. Vis.4
2013 Fusion of Median and Bilateral Filtering for Range Image Upsampling
abstract
We present a new upsampling method to enhance the spatial resolution of depth images. Given a low-resolution depth image from an active depth sensor and a potentially high-resolution color image from a passive RGB camera, we formulate it as an adaptive cost aggregation problem and solve it using the bilateral filter. The formulation synergistically combines the median and bilateral filters thus it better preserves the depth edges and is more robust to noise. Numerical and visual evaluations on a total of 37 Middlebury data sets demonstrate the effectiveness of our method. A real-time high-resolution depth capturing system is also developed using commercial active depth sensor based on the proposed upsampling method.
Qingxiong Yang, Narendra Ahuja, Ruigang Yang, Kar-Han Tan, James Davis 0001, W. Bruce Culbertson, John G. Apostolopoulos, Gang Wang 0012
IEEE Trans. Image Process.2
2013 Automated Visual Inspection of Railroad Tracks
abstract
Thousands of miles of railroad track must be inspected twice weekly by a human inspector to maintain safety standards. A computer vision system, consisting of field-acquired video and subsequent analysis, could improve the efficiency of the current methods. Such a system is prototyped, and the following challenges are addressed: the detection, segmentation, and defect assessment of track components whose appearance vary across different tracks and the identification and inspection of special track areas such as track turnouts. An algorithm that utilizes the periodic manner in which track components repeat in an inspection video is developed. Spectral estimation and signal-processing methods are used to provide robust detection of the periodically occurring track components. Results are demonstrated on field-acquired images and video.
Esther Resendiz, John M. Hart, Narendra Ahuja
IEEE Trans. Intell. Transp. Syst.3
2012 Exploiting nonlocal spatiotemporal structure for video segmentation
abstract
Unsupervised video segmentation is a challenging problem because it involves a large amount of data, and image segments undergo noisy variations in color, texture and motion with time. However, there are significant redundancies that can help disambiguate the effects of noise. To exploit these redundancies and obtain the most spatio-temporally consistent video segmentation, we formulate the problem as a consistent labeling problem by exploiting higher order image structure. A label stands for a specific moving segment. Each segment (or region) is treated as a random variable which is to be assigned a label. Regions assigned the same label comprise a 3D space-time segment, or a region tube. The labels can also be automatically created or terminated at any frame in the video sequence, to allow objects entering or leaving the scene. To formulate this problem, we use the CRF (conditional random field) model. Unlike conventional CRF which has only unary and binary potentials, we also use higher order potentials to favor label consistency among disconnected spatial and temporal segments. Compared to region tracking based methods, the main advantages of the proposed algorithm are two fold: (1) the label consistency constraints are imposed on multiple regions but in a soft manner, and (2) the labeling decision is postponed until the confidence in the labeling is high. We compare our results with a recent state-of-the-art video segmentation algorithm and show that our results are quantitatively and qualitatively better.
Hsien-Ting (Tim) Cheng, Narendra Ahuja
CVPR2
2012 Robust visual tracking via multi-task sparse learning
abstract
In this paper, we formulate object tracking in a particle filter framework as a multi-task sparse learning problem, which we denote as Multi-Task Tracking (MTT). Since we model particles as linear combinations of dictionary templates that are updated dynamically, learning the representation of each particle is considered a single task in MTT. By employing popular sparsity-inducing ℓp, qmixed norms (p ∈ {2, ∞} and q = 1), we regularize the representation problem to enforce joint sparsity and learn the particle representations together. As compared to previous methods that handle particles independently, our results demonstrate that mining the interdependencies between particles improves tracking performance and overall computational complexity. Interestingly, we show that the popular L1tracker [15] is a special case of our MTT formulation (denoted as the L11tracker) when p = q = 1. The learning problem can be efficiently solved using an Accelerated Proximal Gradient (APG) method that yields a sequence of closed form updates. As such, MTT is computationally attractive. We test our proposed approach on challenging sequences involving heavy occlusion, drastic illumination changes, and large pose variations. Experimental results show that MTT methods consistently outperform state-of-the-art trackers.
Tianzhu Zhang 0001, Bernard Ghanem, Si Liu 0001, Narendra Ahuja
CVPR4
2012 Low-Rank Sparse Learning for Robust Visual Tracking
Tianzhu Zhang 0001, Bernard Ghanem, Si Liu 0001, Narendra Ahuja
ECCV (6)4
2012 Robust multi-object tracking via cross-domain contextual information for sports video analysis
abstract
Multiple player tracking is one of the main building blocks needed in a sports video analysis system. In an uncalibrated camera setting, robust mutli-object tracking can be very difficult due to a number of reasons including the presence of noise, occlusion, fast camera motion, low-resolution image capture, varying viewpoints and illumination changes. To address the problem of multi-object tracking in sports videos, we go beyond the video frame domain and make use of information in a homography transform domain that is denoted the homography field domain. We propose a novel particle filter based tracking algorithm that uses both object appearance information (e.g. color and shape) in the image domain and cross-domain contextual information in the field domain to improve object tracking. In the field domain, the effect of fast camera motion is significantly alleviated since the underlying homography transform from each frame to the field domain can be accurately estimated. We use contextual trajectory information (intra-trajectory and inter-trajectory context) to further improve the prediction of object states within an particle filter framework. Here, intra-trajectory contextual information is based on history tracking results in the field domain, while inter-trajectory contextual information is extracted from a compiled trajectory dataset based on tracks computed from videos depicting the same sport. Experimental results on real world sports data show that our system is able to effectively and robustly track a variable number of targets regardless of background clutter, camera motion and frequent mutual occlusion between targets.
Tianzhu Zhang 0001, Bernard Ghanem, Narendra Ahuja
ICASSP3
2012 Trajectory-based Fisher kernel representation for action recognition in videos
Indriyati Atmosukarto, Bernard Ghanem, Narendra Ahuja
ICPR3
2012 An edge-preserving filtering framework for visibility restoration
Linchao Bao, Yibing Song, Qingxiong Yang, Narendra Ahuja
ICPR4
2012 Saliency detection via divergence analysis: A unified perspective
Jia-Bin Huang 0001, Narendra Ahuja
ICPR2
2012 Learning human preferences to sharpen images
Myra Nam, Narendra Ahuja
ICPR2
2012 Exploiting ramp structures for improving optical flow estimation
Abhishek Singh 0002, Narendra Ahuja
ICPR2
2012 Aperture access and manipulation for computational imaging
Xianbiao Shu, Chunyu Gao, Narendra Ahuja
Comput. Vis. Image Underst.3
2012 Surface reflectance and normal estimation from photometric stereo
Qingxiong Yang, Narendra Ahuja
Comput. Vis. Image Underst.2
2012 Stereo Matching Using Epipolar Distance Transform
abstract
In this paper, we propose a simple but effective image transform, called the epipolar distance transform, for matching low-texture regions. It converts image intensity values to a relative location inside a planar segment along the epipolar line, such that pixels in the low-texture regions become distinguishable. We theoretically prove that the transform is affine invariant, thus the transformed images can be directly used for stereo matching. Any existing stereo algorithms can be directly used with the transformed images to improve reconstruction accuracy for low-texture regions. Results on real indoor and outdoor images demonstrate the effectiveness of the proposed transform for matching low-texture regions, keypoint detection, and description for low-texture scenes. Our experimental results on Middlebury images also demonstrate the robustness of our transform for highly textured scenes. The proposed transform has a great advantage, its low computational complexity. It was tested on a MacBook Air laptop computer with a 1.8 GHz Core i7 processor, with a speed of about 9 frames per second for a video graphics array-sized image.
Qingxiong Yang, Narendra Ahuja
IEEE Trans. Image Process.2
2012 Shadow Removal Using Bilateral Filtering
abstract
In this paper, we propose a simple but effective shadow removal method using a single input image. We first derive a 2-D intrinsic image from a single RGB camera image based solely on colors, particularly chromaticity. We next present a method to recover a 3-D intrinsic image based on bilateral filtering and the 2-D intrinsic image. The luminance contrast in regions with similar surface reflectance due to geometry and illumination variances is effectively reduced in the derived 3-D intrinsic image, while the contrast in regions with different surface reflectance is preserved. However, the intrinsic image contains incorrect luminance values. To obtain the correct luminance, we decompose the input RGB image and the intrinsic image. Each image is decomposed into a base layer and a detail layer. We obtain a shadow-free image by combining the base layer from the input RGB image and the detail layer from the intrinsic image such that the details of the intrinsic image are transferred to the input RGB image from which the correct luminance values can be obtained. Unlike previous methods, the presented technique is fully automatic and does not require shadow detection.
Qingxiong Yang, Kar-Han Tan, Narendra Ahuja
IEEE Trans. Image Process.3
2011 Imaging via three-dimensional compressive sampling (3DCS)
abstract
Compressive sampling (CS) aims at acquiring a signal at a sampling rate that is significantly below the Nyquist rate. Its main idea is that a signal can be decoded from incomplete linear measurements by seeking its sparsity in some domain. Despite the remarkable progress in the theory of CS, little headway has been made in the compressive imaging (CI) camera. In this paper, a three-dimensional compressive sampling (3DCS) approach is proposed to reduce the required sampling rate of the CI camera to a practical level. In 3DCS, a generic three-dimensional sparsity measure (3DSM) is presented, which decodes a video from incomplete samples by exploiting its 3D piecewise smoothness and temporal low-rank property. In addition, an efficient decoding algorithm is developed for this 3DSM with guaranteed convergence. The experimental results show that our 3DCS requires a much lower sampling rate than the existing CS methods without compromising recovery accuracy.
Xianbiao Shu, Narendra Ahuja
ICCV2
2011 A Uniform Framework for Estimating Illumination Chromaticity, Correspondence, and Specular Reflection
abstract
Based upon a new correspondence matching invariant called illumination chromaticity constancy, we present a new solution for illumination chromaticity estimation, correspondence searching, and specularity removal. Using as few as two images, the core of our method is the computation of a vote distribution for a number of illumination chromaticity hypotheses via correspondence matching. The hypothesis with the highest vote is accepted as correct. The estimated illumination chromaticity is then used together with the new matching invariant to match highlights, which inherently provides solutions for correspondence searching and specularity removal. Our method differs from the previous approaches: those treat these vision problems separately and generally require that specular highlights be detected in a preprocessing step. Also, our method uses more images than previous illumination chromaticity estimation methods, which increases its robustness because more inputs/constraints are used. Experimental results on both synthetic and real images demonstrate the effectiveness of the proposed method.
Qingxiong Yang, Narendra Ahuja, Ruigang Yang
IEEE Trans. Image Process.3
2010 Pedestrian Recognition with a Learned Metric
Mert Dikmen, Emre Akbas, Thomas S. Huang, Narendra Ahuja
ACCV (4)4
2010 A constant-space belief propagation algorithm for stereo matching
abstract
In this paper, we consider the problem of stereo matching using loopy belief propagation. Unlike previous methods which focus on the original spatial resolution, we hierarchically reduce the disparity search range. By fixing the number of disparity levels on the original resolution, our method solves the message updating problem in a time linear in the number of pixels contained in the image and requires only constant memory space. Specifically, for a 800 × 600 image with 300 disparities, our message updating method is about 30× faster (1.5 second) than standard method, and requires only about 0.6% memory (9 MB). Also, our algorithm lends itself to a parallel implementation. Our GPU implementation (NVIDIA Geforce 8800GTX) is about 10× faster than our CPU implementation. Given the trend toward higher-resolution images, stereo matching using belief propagation with large number of disparity levels as efficient as the small ones makes our method future-proof. In addition to the computational and memory advantages, our method is straightforward to implement.
Qingxiong Yang, Liang Wang 0002, Narendra Ahuja
CVPR3
2010 SVM for edge-preserving filtering
abstract
In this paper, we propose a new method to construct an edge-preserving filter which has very similar response to the bilateral filter. The bilateral filter is a normalized convolution in which the weighting for each pixel is determined by the spatial distance from the center pixel and its relative difference in intensity range. The spatial and range weighting functions are typically Gaussian in the literature. In this paper, we cast the filtering problem as a vector-mapping approximation and solve it using a support vector machine (SVM). Each pixel will be represented as a feature vector comprising of the exponentiation of the pixel intensity, the corresponding spatial filtered response, and their products. The mapping function is learned via ϵ-SVM regression using the feature vectors and the corresponding bilateral filtered values from the training image. The major computation involved is the computation of the spatial filtered responses of the exponentiation of the original image which is invariant to the filter size given that an IIR O(1) solution is available for the spatial filtering kernel. To our knowledge, this is the first learning-based O(1) bilateral filtering method. Unlike previous O(1) methods, our method is valid for both low and high range variance Gaussian and the computational complexity is independent of the range variance value. Our method is also the fastest O(1) bilateral filtering yet developed. Besides, our method allows varying range variance values, based on which we propose a new bilateral filtering method avoiding the over-smoothing or under-smoothing artifacts in traditional bilateral filter.
Qingxiong Yang, Narendra Ahuja
CVPR3
2010 Maximum Margin Distance Learning for Dynamic Texture Recognition
Bernard Ghanem, Narendra Ahuja
ECCV (2)2
2010 Supervised and Unsupervised Clustering with Probabilistic Shift
Sanketh Shetty, Narendra Ahuja
ECCV (5)2
2010 Hybrid Compressive Sampling via a New Total Variation TVL1
Xianbiao Shu, Narendra Ahuja
ECCV (6)2
2010 Real-Time Specular Highlight Removal Using Bilateral Filtering
Qingxiong Yang, Narendra Ahuja
ECCV (4)3
2010 Low-Level Image Segmentation Based Scene Classification
abstract
This paper is aimed at evaluating the semantic information content of multiscale, low-level image segmentation. As a method of doing this, we use selected features of segmentation for semantic classification of real images. To estimate the relative measure of the information content of our features, we compare the results of classifications we obtain using them with those obtained by others using the commonly used patch/grid based features. To classify an image using segmentation based features, we model the image in terms of a probability density function, a Gaussian mixture model (GMM) to be specific, of its region features. This GMM is fit to the image by adapting a universal GMM which is estimated so it fits all images. Adaptation is done using a maximum-aposteriori criterion. We use kernelized versions of Bhattacharyya distance to measure the similarity between two GMMs and support vector machines to perform classification. We outperform previously reported results on a publicly available scene classification dataset. These results suggest further experimentation in evaluating the promise of low level segmentation in image classification.
Emre Akbas, Narendra Ahuja
ICPR2
2010 Sparse Coding of Linear Dynamical Systems with an Application to Dynamic Texture Recognition
abstract
Given a sequence of observable features of a linear dynamical system (LDS), we propose the problem of finding a representation of the LDS which is sparse in terms of a given dictionary of LDSs. Since LDSs do not belong to Euclidean space, traditional sparse coding techniques do not apply. We propose a probabilistic framework and an efficient MAP algorithm to learn this sparse code. Since dynamic textures (DTs) can be modeled as LDSs, we validate our framework and algorithm by applying them to the problems of DT representation and DT recognition. In the case of occlusion, we show that this sparse coding scheme outperforms conventional DT recognition methods.
Bernard Ghanem, Narendra Ahuja
ICPR2
2010 A hemispherical imaging camera
Chunyu Gao, Hong Hua, Narendra Ahuja
Comput. Vis. Image Underst.3
2010 Dinkelbach NCUT: An Efficient Framework for Solving Normalized Cuts Problems with Priors and Convex Constraints
Bernard Ghanem, Narendra Ahuja
Int. J. Comput. Vis.2
2009 From Ramp Discontinuities to Segmentation Tree
Emre Akbas, Narendra Ahuja
ACCV (1)2
2009 Real-time O(1) bilateral filtering
abstract
We propose a new bilateral filtering algorithm with computational complexity invariant to filter kernel size, so-called O(1) or constant time in the literature. By showing that a bilateral filter can be decomposed into a number of constant time spatial filters, our method yields a new class of constant time bilateral filters that can have arbitrary spatial and arbitrary range kernels. In contrast, the current available constant time algorithm requires the use of specific spatial or specific range kernels. Also, our algorithm lends itself to a parallel implementation leading to the first real-time O(1) algorithm that we know of. Meanwhile, our algorithm yields higher quality results since we are effectively quantizing the range function instead of quantizing both the range function and the input image. Empirical experiments show that our algorithm not only gives higher PSNR, but is about 10× faster than the state-of-the-art. It also has a small memory footprint, needed only 2% of the memory required by the state-of-the-art for obtaining the same quality as exact using 8-bit images. We also show that our algorithm can be easily extended for O(1) median filtering. Our bilateral filtering algorithm was tested in a number of applications, including HD video conferencing, video abstraction, highlight removal, and multi-focus imaging.
Qingxiong Yang, Kar-Han Tan, Narendra Ahuja
CVPR3
2009 Non-uniform sampling: A novel approach
abstract
In this paper a novel approach to non-uniform sampling is proposed. Two engineering methods are discussed.
Garimella Rama Murthy, Narendra Ahuja
ICASSP2
2009 Texel-based texture segmentation
abstract
Given an arbitrary image, our goal is to segment all distinct texture subimages. This is done by discovering distinct, cohesive groups of spatially repeating patterns, called texels, in the image, where each group defines the corresponding texture. Texels occupy image regions, whose photometric, geometric, structural, and spatial-layout properties are samples from an unknown pdf. If the image contains texture, by definition, the image will also contain a large number of statistically similar texels. This, in turn, will give rise to modes in the pdf of region properties. Texture segmentation can thus be formulated as identifying modes of this pdf. To this end, first, we use a low-level, multiscale segmentation to extract image regions at all scales present. Then, we use the meanshift with a new, variable-bandwidth, hierarchical kernel to identify modes of the pdf defined over the extracted hierarchy of image regions. The hierarchical kernel is aimed at capturing texel substructure. Experiments demonstrate that accounting for the structural properties of texels is critical for texture segmentation, leading to competitive performance vs. the state of the art.
Sinisa Todorovic, Narendra Ahuja
ICCV2
2009 Robust segmentation of freight containers in train monitoring videos
abstract
This paper is about a vision-based system that automatically monitors intermodal freight trains for the quality of how the loads (containers) are placed along the train. An accurate and robust algorithm to segment the foreground of containers in videos of the moving train is indispensable for this purpose. Given a video of a moving train consisting of containers of different types, this paper presents a method exploiting the information in both frequency and spatial domains to segment these containers. This method can accurately segment all types of containers under a variety of background conditions, e.g illumination variations and moving clouds, in the train videos shot by a fixed camera. The accuracy and robustness of the proposed method are substantiated through a large number of experiments on real data of train videos.
Qing-Jie Kong, Avinash Kumar 0001, Narendra Ahuja, Yuncai Liu
WACV3
2009 Search strategies for shape regularized active contour
Tian-Li Yu 0002, Jiebo Luo 0001, Narendra Ahuja
Comput. Vis. Image Underst.3
2008 Connected Segmentation Tree - A joint representation of region layout and hierarchy
abstract
This paper proposes a new object representation, called connected segmentation tree (CST), which captures canonical characteristics of the object in terms of the photometric, geometric, and spatial adjacency and containment properties of its constituent image regions. CST is obtained by augmenting the objectpsilas segmentation tree (ST) with inter-region neighbor links, in addition to their recursive embedding structure already present in ST. This makes CST a hierarchy of region adjacency graphs. A regionpsilas neighbors are computed using an extension to regions of the Voronoi diagram for point patterns. Unsupervised learning of the CST model of a category is formulated as matching the CST graph representations of unlabeled training images, and fusing their maximally matching subgraphs. A new learning algorithm is proposed that optimizes the model structure by simultaneously searching for both the most salient nodes (regions) and the most salient edges (containment and neighbor relationships of regions) across the image graphs. Matching of the category model to the CST of a new image results in simultaneous detection, segmentation and recognition of all occurrences of the category, and a semantic explanation of these results.
Narendra Ahuja, Sinisa Todorovic
CVPR1
2008 Extracting a fluid dynamic texture and the background from video
abstract
Given the video of a still background occluded by a fluid dynamic texture (FDT), this paper addresses the problem of separating the video sequence into its two constituent layers. One layer corresponds to the video of the unoccluded background, and the other to that of the dynamic texture, as it would appear if viewed against a black background. The model of the dynamic texture is unknown except that it represents fluid flow. We present an approach that uses the image motion information to simultaneously obtain a model of the dynamic texture and separate it from the background which is required to be still. Previous methods have considered occluding layers whose dynamics follows simple motion models (e.g. periodic or 2D parametric motion). FDTs considered in this paper exhibit complex stochastic motion. We consider videos showing an FDT layer (e.g. pummeling smoke or heavy rain) in front of a static background layer (e.g. brick building). We propose a novel method for simultaneously separating these two layers and learning a model for the FDT. Due to the fluid nature of the DT, we are required to learn a model for both the spatial appearance and the temporal variations (due to changes in density) of the FDT, along with a valid estimate of the background. We model the frames of a sequence as being produced by a continuous HMM, characterized by transition probabilities based on the Navier-Stokes equations for fluid dynamics, and by generation probabilities based on the convex matting of the FDT with the background. We learn the FDT appearance, the FDT temporal variations, and the background by maximizing their joint probability using interactive conditional modes (ICM). Since the learned model is generative, it can be used to synthesize new videos with different backgrounds and density variations. Experiments on videos that we compiled demonstrate the performance of our method.
Bernard Ghanem, Narendra Ahuja
CVPR2
2008 Matching images under unstable segmentations
abstract
Region based features are getting popular due to their higher descriptive power relative to other features. However, real world images exhibit changes in image segments capturing the same scene part taken at different time, under different lighting conditions, from different viewpoints, etc. Segmentation algorithms reflect these changes, and thus segmentations exhibit poor repeatability. In this paper we address the problem of matching regions of similar objects under unstable segmentations. Merging and splitting of regions makes it difficult to find such correspondences using one-to-one matching algorithms. We present partial region matching as a solution to this problem. We assume that the high contrast, dominant contours of an object are fairly repeatable, and use them to compute partial matching cost (PMC) between regions. Region correspondences are obtained under region adjacency constraints encoded by region adjacency graph (RAG). We integrate PMC in a many-to-one label assignment framework for matching RAGs, and solve it using belief propagation. We show that our algorithm can match images of similar objects across unstable image segmentations. We also compare the performance of our algorithm with that of the standard one-to-one matching algorithm on three motion sequences. We conclude that our partial region matching approach is robust under segmentation irrepeatabilities.
Varsha Hedau, Himanshu Arora, Narendra Ahuja
CVPR3
2008 Learning subcategory relevances for category recognition
abstract
A real-world object category can be viewed as a characteristic configuration of its parts, that are themselves simpler, smaller (sub)categories. Recognition of a category can therefore be made easier by detecting its constituent subcategories and combing these detection results. Given a set of training images, each labeled by an object category contained in it, we present an approach to learning: (1) Taxonomy defined by recursive sharing of subcategories by multiple image categories; (2) Subcategory relevance as the degree of evidence a subcategory offers for the presence of its parent; (3) Likelihood that the image contains a subcategory; and (4) Prior that a subcategory occurs. The images are represented as points in a feature space spanned by confidences in the occurrences of the subcategories. The subcategory relevances are estimated as weights, necessary to rescale the corresponding axes of the feature space so that the images with the same label are closer to each other than to those with different labels. When a new image is encountered, the learned taxonomy, relevances, likelihoods, and priors are used by a linear classifier to categorize the image. On the challenging Caltech-256 dataset, the proposed approach significantly outperforms the best categorizations reported. This result is significant in that it not only demonstrates the advantages of exploiting subcategory taxonomy for recognition, but also suggests that a feature space spanned by part properties, instead of direct object properties, allows for linear separation of image classes.
Sinisa Todorovic, Narendra Ahuja
CVPR2
2008 Segmentation-based Perceptual Image Quality Assessment (SPIQA)
abstract
Computational representation of perceived image quality is a fundamental problem in computer vision and image processing, which has assumed increased importance with the growing role of images and video in human-computer interaction. It is well-known that the commonly used Peak Signal-to-noise ratio (PSNR), although analysis-friendly, falls far short of this need. We propose a perceptual image quality measure (IQM) in terms of an image's region structure. Given a reference image and its "distorted" version, we propose a "full-reference" IQM, called segmentation-based perceptual image quality as sessment (SPIQA), which quantifies this quality reduction, while minimizing the disparity between human judgment and automated prediction of image quality. One novel feature of SPIQA is that it enables the use of inter- and intra- region attributes in a way that closely resembles how the human visual system (HVS) perceives distortion. Experimental results over a number of images and distortion types demonstrate SPIQA's performance benefits.
Bernard Ghanem, Esther Resendiz, Narendra Ahuja
ICIP3
2008 Segmentation of periodically moving objects
abstract
We present a new approach for the identification and segmentation of objects undergoing periodic motion. Our method uses a combination of maximum likelihood estimation of the period, and segments moving objects using correlation of image segments over an estimated period of interest. Correlation provides the best locations of the moving objects in each frame. Segmentation tree provides the image segments at multiple resolutions. We ensure that children regions and their parent regions have the same period estimates. We show results of testing our method on real videos.
Ousman Azy, Narendra Ahuja
ICPR2
2008 A unified model for activity recognition from video sequences
abstract
We propose an activity recognition algorithm that utilizes a unified spatial-frequency model of motion to recognize large-scale differences in action using global statistics, and subsequently distinguishes between motions with similar global statistics by spatially localizing the moving objects. We model the Fourier transforms of translating rigid objects in a video, since the Fourier domain inherently groups regions of the video with similar motion in high energy concentrations within its domain to make global motion detectable. Frequency-domain statistics can be used to isolate the frames that both adhere to our model and contain similar global motion, thus we can separate activities into broader classes based on their global motion. A least-squares solution is then solved to isolate the spatially discriminative object configurations that produce similar global motion statistics. This model provides a unified framework to form concise globally-optimal spatial and motion descriptors necessary for discriminating activities. Experimental results are demonstrated on a human activity dataset.
Esther Resendiz, Narendra Ahuja
ICPR2
2008 A uniformity criterion and algorithm for data clustering
abstract
We propose a novel multivariate uniformity criterion for testing uniformity of point density in an arbitrary dimensional point pattern. An unsupervised, nonparametric data clustering algorithm, using this criterion, is also presented. The algorithm relies on a relatively general notion of cluster so that it is applicable to clusters of relatively unrestricted shapes, densities and sizes. We define a cluster as a set of contiguous interior points surrounded by border points. We use our uniformity test to differentiate between interior and border points. We group interior points to form cluster cores, and then identify cluster borders as formed by the border points neighboring the cluster cores. The algorithm is effective in resolving clusters of different shapes, sizes and densities. It is relatively insensitive to outliers. We present results for experiments performed on artificial and real data sets.
Sanketh Shetty, Narendra Ahuja
ICPR2
2008 Scale-invariant region-based hierarchical imagematching
abstract
This paper presents an approach to scale-invariant image matching. Given two images, the goal is to find correspondences between similar subimages, e.g., representing similar objects, even when the objects are captured under large variations in scale. As in previous work: similarity is defined in terms of geometric, photometric and structural properties of regions, and images are represented by segmentation trees that capture region properties and their recursive embedding. Matching two regions thus amounts to matching their corresponding subtrees. Scale invariance is aimed at overcoming two challenges in matching two images of similar objects. First, the absolute values of many object image properties may change with scale. Second, some of the finest details visible in the high-zoom image may not be visible in the coarser scale image. We normalize the region properties associated with one of the subtrees to the corresponding properties of the root of the other subtree. This makes the scales of objects represented by the two subtrees equal, and also decouples this scale from that of the entire scene. We also weight contributions of subregions to the total similarity of their parent regions by the relative area the subregions occupy within the parents. This reduces the penalty for not being able to match fine-resolution details present within only one of the two regions, since the penalty will be down-weighted by the relatively small area of these details. Our experiments demonstrate invariance of the proposed algorithm to large changes in scale.
Sinisa Todorovic, Narendra Ahuja
ICPR2
2008 Region-Based Hierarchical Image Matching
Sinisa Todorovic, Narendra Ahuja
Int. J. Comput. Vis.2
2008 A Tensor Approximation Approach to Dimensionality Reduction
Narendra Ahuja
Int. J. Comput. Vis.2
2008 Unsupervised Category Modeling, Recognition, and Segmentation in Images
abstract
Suppose a set of arbitrary (unlabeled) images contains frequent occurrences of 2D objects from an unknown category. This paper is aimed at simultaneously solving the following related problems: (1) unsupervised identification of photometric, geometric, and topological properties of multiscale regions comprising instances of the 2D category; (2) learning a region-based structural model of the category in terms of these properties; and (3) detection, recognition and segmentation of objects from the category in new images. To this end, each image is represented by a tree that captures a multiscale image segmentation. The trees are matched to extract the maximally matching subtrees across the set, which are taken as instances of the target category. The extracted subtrees are then fused into a tree-union that represents the canonical category model. Detection, recognition, and segmentation of objects from the learned category are achieved simultaneously by finding matches of the category model with the segmentation tree of a new image. Experimental validation on benchmark datasets demonstrates the robustness and high accuracy of the learned category models, when only a few training examples are used for learning without any human supervision.
Sinisa Todorovic, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Integration of Frequency and Space for Multiple Motion Estimation and Shape-Independent Object Segmentation
abstract
A video containing multiple objects undergoing independent translational and rotational motions is analyzed through a combination of spatial- and frequency-domain representations. The Fourier transform of the sequence is used to estimate the multiple translations and rotations in a computationally efficient manner, which is also robust to local inaccuracies and global illumination changes. A novel algorithm is presented for the simultaneous extraction of all objects undergoing translation and the background via a least squares technique that takes place entirely in the Fourier domain. Spatial information is combined with the frequency domain object extraction results, to further refine them. For the case of rotational or combined, rotational and translational motions, the moving objects are segmented using purely spatial information. We show that the combined analysis takes advantage of the strengths of both representations, by providing reliable and computationally efficient motion estimates and object segmentation. The proposed algorithm is shown to be robust to local noise and occlusion, because of its global nature. Experiments are performed on synthetic and real video sequences to demonstrate the capabilities of our approach.
Alexia Briassouli, Narendra Ahuja
IEEE Trans. Circuits Syst. Video Technol.2
2007 Modelling Objects using Distribution and Topology of Multiscale Region Pairs
abstract
We propose a method for simultaneous detection, localization and segmentation of objects of a known category. We show that this is possible by using segments as features. To this end, we propose an object model in which the image is represented as a tree, that captures containment relationships among the segments. Using segments as features has the advantage that object detection and segmentation is done simultaneously, forgoing the need for a separate sophisticated model for object segmentation. A generative model of an object category is estimated in a supervised mode, in terms of the characteristics of its constituent regions, their relative locations, and their mutual containment. The novel aspect of this work lies in simplifying the description of the hierarchy in terms of constraints that apply to only pairs of nodes, instead of all nodes in the tree. We show that this indeed improves the speed of learning algorithm. Inference is done using graph cuts. We report the performance of the model on standard datasets.
Himanshu Arora, Narendra Ahuja
CVPR2
2007 Unsupervised Segmentation of Objects using Efficient Learning
abstract
We describe an unsupervised method to segment objects detected in images using a novel variant of an interest point template, which is very efficient to train and evaluate. Once an object has been detected, our method segments an image using a conditional random field (CRF) model. This model integrates image gradients, the location and scale of the object, the presence of object parts, and the tendency of these parts to have characteristic patterns of edges nearby. We enhance our method using multiple unsegmented images of objects to learn the parameters of the CRF, in an iterative conditional maximization framework. We show quantitative results on images of real scenes that demonstrate the accuracy of segmentation.
Himanshu Arora, Nicolas Loeff, David A. Forsyth, Narendra Ahuja
CVPR4
2007 Active Aperture Control and Sensor Modulation for Flexible Imaging
abstract
In the paper, we describe an optical system which is capable of providing external access to both the sensor and the lens aperture (i.e., projection center) of a conventional camera. The proposed optical system is attached in front of the camera, and is the equivalent of adding externally accessible intermediate image plane and projection center. The system offers controls of the response of each pixel which could be used to realize many added imaging functions, such as high dynamic range imaging, image modulation, and optical computation. The ability to access the optical center could enable a wide variety of applications by simply allowing manipulation of the geometric properties of the optical center. For instance, panoramic imaging can be implemented by rotating a planar mirror about the camera axis; and small base-line stereo can be implemented by shifting the camera center. We have implemented a bench setup to demonstrate some of these functions. The experimental results are included.
Chunyu Gao, Narendra Ahuja, Hong Hua
CVPR2
2007 Extracting Texels in 2.1D Natural Textures
abstract
This paper proposes the problem of unsupervised extraction of texture elements, called texels, which repeatedly occur in the image of a frontally viewed, homogeneous, 2.1D, planar texture, and presents a solution. 2.1D texture here means that the physical texels are thin objects lying along a surface that may partially occlude one another. The image texture is represented by the segmentation tree whose structure captures the recursive embedding of regions obtained from a multiscale image segmentation. In the segmentation tree, the texels appear as subtrees with similar structure, with nodes having similar photometric and geometric properties. A new learning algorithm is proposed for fusing these similar subtrees into a tree-union, which registers all visible texel parts, and thus represents a statistical, generative model of the complete (unoccluded) texel. The learning algorithm involves concurrent estimation of texel tree structure, as well as the probability distributions of its node properties. Texel detection and segmentation are achieved simultaneously by matching the segmentation tree of a new image with the texel model. Experiments conducted on a newly compiled dataset containing 2.1D natural textures demonstrate the validity of our approach.
Narendra Ahuja, Sinisa Todorovic
ICCV1
2007 Learning the Taxonomy and Models of Categories Present in Arbitrary Images
abstract
This paper proposes, and presents a solution to, the problem of simultaneous learning of multiple visual categories present in an arbitrary image set and their inter-category relationships. These relationships, also called their taxonomy, allow categories to be defined recursively, as spatial configurations of (simpler) subcategories each of which may be shared by many categories. Each image is represented by a segmentation tree, whose structure captures recursive embedding of image regions in a multiscale segmentation, and whose nodes contain the associated region properties. The presence of any occurring categories is reflected in the occurrence of associated, similar subtrees within the image trees. Similar subtrees across the entire image set are clustered. Each cluster corresponds to a discovered category, represented by the cluster properties. A (subcategory) cluster of small matching subtrees may occur within multiple clusters (categories) of larger matching subtrees, in different spatial relationships with subtrees from other small clusters. Such recursive embedding, grouping and intersection of clusters is captured in a directed acyclic graph (DAG) which represents the discovered taxonomy. Detection, recognition and segmentation of any of the learned categories present in a new image are simultaneously conducted by matching the segmentation tree of the new image with the learned DAG. This matching also yields a semantic explanation of the recognized category, in terms of the presence of its subcategories. Experiments with a newly compiled dataset of four-legged animals demonstrate good cross-category resolvability.
Narendra Ahuja, Sinisa Todorovic
ICCV1
2007 Phase Based Modelling of Dynamic Textures
abstract
This paper presents a model of spatiotemporal variations in a dynamic texture (DT) sequence. Most recent work on DT modelling represents images in a DT sequence as the responses of a linear dynamical system (LDS) to noise. Despite its merits, this model has limitations because it attempts to model temporal variations in pixel intensities which do not take advantage of global motion coherence. We propose a model that relates texture dynamics to the variation of the Fourier phase, which captures the relationships among the motions of all pixels (i.e. global motion) within the texture, as well as the appearance of the texture. Unlike LDS, our model does not require segmentation or cropping during the training stage, which allows it to handle DT sequences containing a static background. We test the performance of this model on recognition and synthesis of DT's. Experiments with a dataset that we have compiled demonstrate that our phase based model outperforms LDS.
Bernard Ghanem, Narendra Ahuja
ICCV2
2007 Phase PCA for Dynamic Texture Video Compression
abstract
Temporal or dynamic textures (DT's) are video sequences that are spatially repetitive and temporally stationary. DT's are temporal analogs of the well known spatial still image texture. Examples of DT's include moving water, foliage, smoke, clouds, etc. We present a new DT model that can efficiently compress DT sequences. Our proposed method compactly represents the spatiotemporal properties of a DT by modelling its varying Fourier phase content, which can be shown to be the major determinant of both its dynamics and appearance. This is possible because this method combines both temporal and spatial properties in a compact spectral framework. Making use of the benefits inherent to working in the frequency domain, this model provides a significant improvement in DT compression, which can be used to improve the performance of MPEG-2 encoding. We will present experimental evidence that validates this method for a variety of complex sequences, while also comparing it to the most recent DT representational model that is based on modelling a DT as a linear dynamical system (LDS).
Bernard Ghanem, Narendra Ahuja
ICIP (3)2
2007 A Vision System for Monitoring Intermodal Freight Trains
abstract
We describe the design and implementation of a vision based Intermodal Train Monitoring System (ITMS) for extracting various features like length of gaps in an intermodal (IM) train which can later be used for higher level inferences. An intermodal train is a freight train consisting of two basic types of loads - containers and trailers. Our system first captures the video of an IM train, and applies image processing and machine learning techniques developed in this work to identify the various types of loads as containers and trailers. The whole process relies on a sequence of following tasks -robust background subtraction in each frame of the video, estimation of train velocity, creation of mosaic of the whole train from the video and classification of train loads into containers and trailers. Finally, the length of gaps between the loads of the IM train is estimated and is used to analyze the aerodynamic efficiency of the loading pattern of the train, which is a critical aspect of freight trains. This paper focusses on the machine vision aspect of the whole system
Avinash Kumar 0001, Narendra Ahuja, John M. Hart, Visesh Chari, P. J. Narayanan, C. V. Jawahar
WACV2
2007 Videoshop: A new framework for spatio-temporal video editing in gradient domain
Ning Xu 0005, Ramesh Raskar, Narendra Ahuja
Graph. Model.4
2007 Object segmentation using graph cuts based active contours
Ning Xu 0005, Narendra Ahuja, Ravi Bansal
Comput. Vis. Image Underst.2
2007 Shape and View Independent Reflectance Map from Multiple Views
Tian-Li Yu 0002, Ning Xu 0005, Narendra Ahuja
Int. J. Comput. Vis.3
2007 Extraction and Analysis of Multiple Periodic Motions in Video Sequences
abstract
The analysis of periodic or repetitive motions is useful in many applications, such as the recognition and classification of human and animal activities. Existing methods for the analysis of periodic motions first extract motion trajectories using spatial information and then determine if they are periodic. These approaches are mostly based on feature matching or spatial correlation, which are often infeasible, unreliable, or computationally demanding. In this paper, we present a new approach, based on the time-frequency analysis of the video sequence as a whole. Multiple periodic trajectories are extracted and their periods are estimated simultaneously. The objects that are moving in a periodic manner are extracted using the spatial domain information. Experiments with synthetic and real sequences display the capabilities of this approach.
Alexia Briassouli, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Design Analysis of a High-Resolution Panoramic Camera Using Conventional Imagers and a Mirror Pyramid
abstract
Wide field of view (FOV) and high-resolution image acquisition is highly desirable in many vision-based applications. Several systems have reported the use of reflections off mirror pyramids to capture high-resolution, single-viewpoint, and wide-FOV images. Using a dual mirror pyramid (DMP) panoramic camera as an example, in this paper, we examine how the pyramid geometry, and the selection and placement of imager clusters can be optimized to maximize the overall panoramic FOV, sensor utilization efficiency, and image uniformity. The analysis can be generalized and applied to other pyramid-based designs.
Hong Hua, Narendra Ahuja, Chunyu Gao
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Calibration of an HMPD-Based Augmented Reality System
abstract
In augmented reality (AR) applications, accurately registering a virtual object with its real counterpart is a challenging problem. The size, depth, and geometry, as well as the physical attributes of the virtual object, have to be rendered precisely relative to a physical reference. This paper presents a systematic calibration process to address the registration challenge in a custom-designed AR system, which is based upon recent head-mounted projective display (HMPD) technology. Following a concise review of the HMPD concept and our system configuration, we first present a computational model of the HMPD viewing system and requirements for the system calibration. Then, we describe, in detail, the calibration procedures to obtain estimates of unknown transformations, summarize the application of the estimates in a customized graphics rendering toolkit, and discuss the evaluation experiments and observations. Finally, the implementation of a testbed to demonstrate a successful registration is briefly described, and experimental results are presented
Hong Hua, Chunyu Gao, Narendra Ahuja
IEEE Trans. Syst. Man Cybern. Part A3
2006 Dynamic Textures Synthesis as Nonlinear Manifold Learning and Traversing
abstract
We formulate the problem of dynamic texture synthesis as a nonlinear manifold learning and traversing problem. We characterize dynamic textures as the temporal changes in spectral parameters of image sequences. For continuous changes of such parameters, it is commonly assumed that all these parameters lie on or close to a low-dimensional manifold embedded in the original configuration space. For complex dynamic data, the manifolds are usually nonlinear and we propose to use a mixture of linear subspaces to model a nonlinear manifold. These locally linear subspaces are further aligned within a global coordinate system. With the nonlinear manifold being globally parameterized, we overcome motion discontinuity problems encountered in switching linear models and dynamics. We present a nonparametric method to describe the complex dynamics of data sequences on the manifold. We also apply such approach to dynamic spatial parameters such as motion capture data. The experimental results suggest that our approach is able to synthesize smooth, complex dynamic textures and human motions, and has potential applications to other dynamic data synthesis problems. 1
Che-Bin Liu, Ruei-Sung Lin, Narendra Ahuja, Ming-Hsuan Yang 0001
BMVC3
2006 A Refractive Camera for Acquiring Stereo and Super-resolution Images
abstract
We propose a novel depth sensing system composed of a single camera, and a transparent plate which is placed in front of the camera and rotates about the optical axis of the camera. The camera takes a sequence of images as the plate rotates, which provide the equivalent of a large number of stereo pairs. Compared with conventional multi-camera stereo systems, the use of a single camera for capturing stereo pairs helps improve the accuracy of detecting correspondences. The availability of the large number of stereo pairs also reduces matching ambiguities, even for objects with low texture. By using both the images and the estimated depth map, we show that the proposed system is also capable of generating super-resolution images. Experimental results on reconstructing 3D structures and recovering highresolution images are presented.
Chunyu Gao, Narendra Ahuja
CVPR (2)2
2006 Extracting Subimages of an Unknown Category from a Set of Images
abstract
Suppose a set of images contains frequent occurrences of objects from an unknown category. This paper is aimed at simultaneously solving the following related problems: (1) unsupervised identification of photometric, geometric, and topological (mutual containment) properties of multiscale regions defining objects in the category; (2) learning a region-based structural model of the category in terms of these properties from a set of training images; and (3) segmentation and recognition of objects from the category in new images. To this end, each image is represented by a tree that captures a multiscale image segmentation. The trees are matched to find the maximally matching subtrees across the set, the existence of which is itself viewed as evidence that a category is indeed present. The matched subtrees are fused into a canonical tree, which represents the learned model of the category. Recognition of objects in a new image and image segmentation delineating all object parts are achieved simultaneously by finding matches of the model with subtrees of the new image. Experimental comparison with state-of-the-art methods shows that the proposed approach has similar recognition and superior localization performance while it uses fewer training examples.
Sinisa Todorovic, Narendra Ahuja
CVPR (1)2
2006 SDG Cut: 3D Reconstruction of Non-lambertian Objects Using Graph Cuts on Surface Distance Grid
abstract
We show that the approaches to 3D reconstruction that use volumetric graph cuts to minimize a cost function over the object surface have two types of biases, the minimal surface bias and the discretization bias. These biases make it difficult to recover surface extrusions and other details, especially when a non-lambertian photo-consistency measure is used. To reduce these biases, we propose a new iterative graph cuts based algorithm that operates on the Surface Distance Grid (SDG), which is a special discretization of the 3Dspace, constructed using a signed distance transform of the current surface estimate. It can be shown that SDG significantly reduces the minimal surface bias, and transforms the discretization bias into a controllable degree of surface smoothness. Experiments on 3D reconstruction of non-lambertian objects confirm the effectiveness of our algorithm over previous methods.
Tian-Li Yu 0002, Narendra Ahuja, Wei-Chao Chen
CVPR (2)2
2006 Estimation of Multiple Periodic Motions from Video
Alexia Briassouli, Narendra Ahuja
ECCV (1)2
2006 Learning Nonlinear Manifolds from Time Series
Ruei-Sung Lin, Che-Bin Liu, Ming-Hsuan Yang 0001, Narendra Ahuja, Stephen E. Levinson
ECCV (2)4
2006 Gradient Adaptive Image Restoration and Enhancement
abstract
Various methods have been proposed for image enhancement and restoration. The main difficulty is how to enhance the structures uniformly while suppressing the noise without artifacts. In this paper, we tackle this problem in the gradient domain instead of the traditional intensity domain. By enhancing the gradient field, we can enhance the structure uniformly without overshooting at the boundary. Because the gradient field is very sensitive to noise, we apply an orientation-isotropy adaptive filter to the gradient field, suppressing the gradients in the noise regions while enhancing along the object boundaries. Thus we obtain a modulated gradient field, which is usually not integrable. We reconstruct the enhanced image from the modulated gradient field with least square errors by solving a Poisson equation. This method can enhance the object contrast uniformly, suppress the noise with no artifacts, and avoid setting stopping time as in PDE methods. Experiments on noisy images show the efficacy of our method.
Yunqiang Chen, Tong Fang, Jason Tyan, Narendra Ahuja
ICIP5
2006 Sparse Lumigraph Relighting by Illumination and Reflectance Estimation from Multi-View Images
Tian-Li Yu 0002, Narendra Ahuja, Wei-Chao Chen
Rendering Techniques3
2005 Rank-R Approximation of Tensors: Using Image-as-Matrix Representation
abstract
We present a novel multilinear algebra based approach for reduced dimensionality representation of image ensembles. We treat an image as a matrix, instead of a vector as in traditional dimensionality reduction techniques like PCA, and higher-dimensional data as a tensor. This helps exploit spatio-temporal redundancies with less information loss than image-as-vector methods. The challenges lie in the computational and memory requirements for large ensembles. Currently, there exists a rank-R approximation algorithm which, although applicable to any number of dimensions, is efficient for only low-rank approximations. For larger dimensionality reductions, the memory and time costs of this algorithm become prohibitive. We propose a novel algorithm, for rank-R approximations of third-order tensors, which is efficient for arbitrary R but for the important special case of 2D image ensembles, e.g. video. Both of these algorithms reduce redundancies present in all dimensions. Rank-R tensor approximation yields the most compact data representation among all known image-as-matrix methods. We evaluated the performance of our algorithm vs. other approaches on a number of datasets with the following two main results. First, for a fixed compression ratio, the proposed algorithm yields the best representation of image ensembles visually as well as in the least squares sense. Second, proposed representation gives the best performance for object classification.
Narendra Ahuja
CVPR (2)2
2005 Videoshop: A New Framework for Spatio-Temporal Video Editing in Gradient Domain
abstract
Our goal is to develop tools that go beyond frame-constrained manipulation such as resizing, color correction, and simple transitions, and provide object-level operations within frames. Some of our targeted video editing tasks includes transferring a motion picture to a new still picture, importing a moving object into a new background, and compositing two video sequences. The challenges behind this kind of complex video editing tasks lie in two constraints: 1) Spatial consistency: imported objects should blend with the background seamlessly. Hence pixel replacement, which creates noticeable seams, is problematic. 2) Temporal coherency: successive frames should display smooth transitions. Hence frame-by-frame editing, which results in visual flicker, is inappropriate. Our work is aimed at providing an easy-to-use video editing tool that maximally satisfies the spatial and temporal constraints mentioned above and requires minimum user interaction. We propose a new framework for video editing in gradient domain. The spatio-temporal gradient fields of target videos are modified and/or mixed to generate a new gradient field which is usually not integrable. We propose a 3D video integration algorithm, which uses the variational method, to find the potential function whose gradient field is closest to the mixed gradient field in the sense of least squares. The video is reconstructed by solving a 3D Poisson equation. We derive an extension of current 2D gradient technique to 3D space, yielding in a novel video editing framework, which is very different from all current video editing software.
Ning Xu 0005, Ramesh Raskar, Narendra Ahuja
CVPR (2)4
2005 Shape Regularized Active Contour Using Iterative Global Search and Local Optimization
abstract
Recently, nonlinear shape models have been shown to improve the robustness and flexibility of segmentation. In this paper, we propose shape regularized active contour (ShRAC) that incorporates existing nonlinear shape models into the classical active contour approach. ShRAC uses a discrete representation of the contour to allow efficient combinatorial search. The search for optimal contour is performed by coarse-to-fine algorithm that iterates between combinatorial search and gradient-based local optimization. First, multi-solution dynamic programming (MSDP) is used to generate initial candidates by minimizing only the image energy. In the second step, a combination of image energy and shape energy determined by a given prior shape model is minimized for the initial candidates using a local optimization method and the best one is selected. To have diverse initial candidates, we employ a clustered solution pruning procedure in the MSDP search space. Finally, local shape regularization is used to feed shape constraints back into the new MSDP search space of the next iteration. Our search strategy combines the advantages of global combinatorial search and local optimization, and has shown excellent robustness to local minima caused by distracting suboptimal segmentations. Experimental results on segmentation of different anatomical structures using ShRAC are provided.
Tian-Li Yu 0002, Jiebo Luo 0001, Narendra Ahuja
CVPR (2)3
2005 Integrated Spatial and Frequency Domain 2D Motion Segmentation and Estimation
abstract
A video containing multiple objects in rotational and translational motion is analyzed through a combination of spatial and frequency domain representations. It is argued that the combined analysis can take advantage of the strengths of both representations. Initial estimates of constant, as well as time-varying, translation and rotation velocities are obtained from frequency analysis. Improved motion estimates and motion segmentation for the case of translation are achieved by integrating spatial and Fourier domain information. For combined rotational and translational motions, the frequency representation is used for motion estimation, but only spatial information can be used to separate and extract the independently moving objects. The proposed algorithms are tested on synthetic and real videos.
Alexia Briassouli, Narendra Ahuja
ICCV2
2005 Two-channel predictive multiple description coding
abstract
This paper presents a multiple description (MD) video codec based on the principles side-information coding. In particular, we highlight certain key components of the codec design that contribute significantly to the rate-distortion performance of the proposed codec. These include the use of randomized permutations of the quantization codebook in conjunction with binary LDPC codes for partitioning the available bit-rate among the coefficient bit-planes. Another key component of the proposed codec is the use of pdf estimation for improved decoder reconstruction. Lastly, we use a bank of sequential LDPC decoders to efficiently decode the transmitted coset information. Empirical evaluation demonstrates the superior performance of the proposed codec for the communication of encoded video over packet erasure channels.
Ashish Jagmohan, Anshul Sehgal, Narendra Ahuja
ICIP (2)3
2005 Modeling Dynamic Textures Using Subspace Mixtures
abstract
In this paper, we aim at modeling video sequences that exhibit temporal appearance variation. The dynamic texture model proposed in [6] is effective to model simple dynamic scenes. However, because of its over-simplified appearance model and under-constrained dynamics model, the visual quality of its synthesized video sequences is often not satisfactory. This leads to our new model. We parameterize the nonlinear image manifold using mixtures of probabilistic principal component analyzers. We then align coefficients from different mixture components in a global coordinate system, and model the image dynamics in the global coordinate using an autoregressive process. The experimental results show that our method is capable of capturing complex temporal appearance variation and offers improved synthesis results over previous works.
Che-Bin Liu, Ruei-Sung Lin, Narendra Ahuja
ICME3
2005 Out-of-core tensor approximation of multi-dimensional matrices of visual data
abstract
Tensor approximation is necessary to obtain compact multilinear models for multi-dimensional visual datasets. Traditionally, each multi-dimensional data item is represented as a vector. Such a scheme flattens the data and partially destroys the internal structures established throughout the multiple dimensions. In this paper, we retain the original dimensionality of the data items to more effectively exploit existing spatial redundancy and allow more efficient computation. Since the size of visual datasets can easily exceed the memory capacity of a single machine, we also present an out-of-core algorithm for higher-order tensor approximation. The basic idea is to partition a tensor into smaller blocks and perform tensor-related operations blockwise. We have successfully applied our techniques to three graphics-related data-driven models, including 6D bidirectional texture functions, 7D dynamic BTFs and 4D volume simulation sequences. Experimental results indicate that our techniques can not only process out-of-core data, but also achieve higher compression ratios and quality than previous methods.
Qing Wu 0006, Yizhou Yu, Narendra Ahuja
ACM Trans. Graph.5
2004 A Model for Dynamic Shape and Its Applications
Che-Bin Liu, Narendra Ahuja
CVPR (2)2
2004 Recovering Shape and Reflectance Model of Non-Lambertian Objects from Multiple Views
Tian-Li Yu 0002, Ning Xu 0005, Narendra Ahuja
CVPR (2)3
2004 A Robust Probabilistic Estimation Framework for Parametric Image Models
Maneesh Kumar Singh 0001, Himanshu Arora, Narendra Ahuja
ECCV (1)3
2004 Shape and View Independent Reflectance Map from Multiple Views
Tian-Li Yu 0002, Ning Xu 0005, Narendra Ahuja
ECCV (4)3
2004 Motion based retrieval of dynamic objects in videos
abstract
Most existing video retrieval systems use low-level visual features such as color histogram, shape, texture, or motion. In this paper, we explore the use of higher-level motion representation for video retrieval of dynamic objects. We use three motion representations, which together can retrieve a large variety of motion patterns. Our approach works on top of a tracking unit and assumes that each dynamic object has been tracked and circumscribed in a minimal bounding box in each video frame. We represent the motion attributes of each object in terms of changes in the image context of its circumscribing box. The changes are described via motion templates [4], self-similarity plots [3], and image dynamics [9]. Initially, defined criteria of the retrieval process are interactively refined using relevance feedback from the user. Experimental results demonstrate the use of the proposed motion models in retrieving objects undergoing complex motion.
Che-Bin Liu, Narendra Ahuja
ACM Multimedia2
2004 A potential-based generalized cylinder representation
Jen-Hui Chuang, Narendra Ahuja, Chien-Chou Lin, Chi-Hao Tsai, Cheng-Hui Chen
Comput. Graph.2
2004 Split Aperture Imaging for High Dynamic Range
Manoj Aggarwal, Narendra Ahuja
Int. J. Comput. Vis.2
2004 Multiview Panoramic Cameras Using Mirror Pyramids
abstract
A mirror pyramid consists of a set of planar mirror faces arranged around an axis of symmetry and inclined to form a pyramid. By strategically positioning a number of conventional cameras around a mirror pyramid, the viewpoints of the cameras' mirror images can be located at a single point within the pyramid and their optical axes pointed in different directions to effectively form a virtual camera with a panoramic field of view. Mirror pyramid-based panoramic cameras have a number of attractive properties, including single-viewpoint imaging, high resolution, and video rate capture. It is also possible to place multiple viewpoints within a single mirror pyramid, yielding compact designs for simultaneous multiview panoramic video rate imaging. Nalwa [4] first described some of the basic ideas behind mirror pyramid cameras. In this paper, we analyze the general class of multiview panoramic cameras, provide a method for designing these cameras, and present experimental results using a prototype we have developed to validate single-pyramid multiview designs. We first give a description of mirror pyramid cameras, including the imaging geometry, and investigate the relationship between the placement of viewpoints within the pyramid and the cameras' field of view (FOV), using simulations to illustrate the concepts. A method for maximizing sensor utilization in a mirror pyramid-based multiview panoramic camera is also presented. Images acquired using the experimental prototype for two viewpoints are shown.
Kar-Han Tan, Hong Hua, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Wyner-Ziv coding of video: an error-resilient compression framework
abstract
This paper addresses the problem of video coding in a joint source-channel setting. In particular, we propose a video encoding algorithm that prevents the indefinite propagation of errors in predictively encoded video-a problem that has received considerable attention over the last decade. This is accomplished by periodically transmitting a small amount of additional information, termed coset information, to the decoder, as opposed to the popular approach of periodic insertion of intra-coded frames. Perhaps surprisingly, the coset information is capable of correcting for errors, without the encoder having a precise knowledge of the lost packets that resulted in the errors. In the context of real-time transmission, the proposed approach entails a minimal loss in performance over conventional encoding in the absence of channel losses, while simultaneously allowing error recovery in the event of channel losses. We demonstrate the efficacy of the proposed approach through experimental evaluation. In particular, the performance of the proposed framework is 3-4 dB superior to the conventional approach of periodic insertion of intra-coded frames, and 1.5-2 dB away from an ideal system, with infinite decoding delay, operating at Shannon capacity.
Anshul Sehgal, Ashish Jagmohan, Narendra Ahuja
IEEE Trans. Multim.3
2003 Object Segmentation Using Graph Cuts Based Active Contours
abstract
In this paper we present a graph cuts based active contours (GCBAC) approach to object segmentation problems. Our method is a combination of active contours and the optimization tool of graph cuts and differs fundamentally from traditional active contours in that it uses graph cuts to iteratively deform the contour. Consequently, it has the following advantages. (1) It has the ability to jump over local minima and provide a more global result. (2) Graph cuts guarantee continuity and lead to smooth contours free of self-crossing and uneven spacing problems. Therefore, the internal force, which is commonly used in traditional energy functions to control the smoothness, is no longer needed, and hence the number of parameters is greatly reduced. (3). Our approach easily extends to the segmentation of three and higher dimensional objects. In addition, the algorithm is suitable for interactive correction and is shown to always converge. Experimental results and analyses are provided.
Ning Xu 0005, Ravi Bansal, Narendra Ahuja
CVPR (2)3
2003 Wyner-Ziv Encoded Predictive Multiple Descriptions
abstract
The predictive multiple description coding problem can be posed as a variant of the well-known Wyner-Ziv side-information problem. Predictive MD coding in this framework (termed the WYZE-PMD) eliminates the problem of predictive mismatch without requiring restrictive channel assumptions or high latency. The performance of two-channel one-step predictive MD coding was analyzed within the WYZE-PMD framework. Achievable rate-distortion (R-D) regions were obtained for the problem of MD coding in the presence of correlated decoder side-information. These were used to obtain the operational R-D performance for predictive MD coding under certain restrictions. Practical code constructions were proposed within the WYZE-PMD framework. Performance comparisons between the proposed codes and conventional approaches were presented for communication of a first-order Gauss-Markov source over two erasure channels with independent failure probabilities. Results indicated that the proposed approach significantly out-performs conventional approaches in terms of R-D performance.
Ashish Jagmohan, Narendra Ahuja
DCC2
2003 Robust Predictive Coding and the Wyner-Ziv Problem
abstract
The problem of robust communication of predictive encoded video in a joint source-channel setting is addressed. Specifically, the problem of predictive mismatch, where there is a drift between the state of the encoder and the decoder is addressed as a variant of the Wyner-Ziv problem. A video encoding algorithm based on the H.26L video codec, which presents the propagation of error in predictively encoded video in the event of predictive mismatch (or drift) between the encoder and the decoder, is proposed. One of the main advantages of the proposed approach is that there is minimal loss in performance over the standard H.26L encoder during error-free transmission, while simultaneously allowing error recovery in the event of errors. Using turbo codes as coset codes, the performance of the proposed codec is evaluated and the efficacy of the proposed framework is demonstrated. The performance of the proposed approach can only improve with the use of superior coset codes.
Anshul Sehgal, Narendra Ahuja
DCC2
2003 Regression based Bandwidth Selection for Segmentation using Parzen Windows
abstract
We consider the problem of segmentation of images that can be modelled as piecewise continuous signals having unknown, nonstationary statistics. We propose a solution to this problem which first uses a regression framework to estimate the image PDF, and then mean-shift to find the modes of this PDF. The segmentation follows from mode identification wherein pixel clusters or image segments are identified with unique modes of the multimodal PDF. Each pixel is mapped to a mode using a convergent, iterative process. The effectiveness of the approach depends upon the accuracy of the (implicit) estimate of the underlying multimodal density function and thus on the bandwidth parameters used for its estimate using Parzen windows. Automatic selection of bandwidth parameters is a desired feature of the algorithm. We show that the proposed regression-based model admits a realistic framework to automatically choose bandwidth parameters which minimizes a global error criterion. We validate the theory presented with results on real images.
Maneesh Kumar Singh 0001, Narendra Ahuja
ICCV2
2003 Facial Expression Decomposition
abstract
In this paper, we propose a novel approach for facial expression decomposition - higher-order singular value decomposition (HOSVD), a natural generalization of matrix SVD. We learn the expression subspace and person subspace from a corpus of images showing seven basic facial expressions, rather than resort to expert-coded facial expression parameters. We propose a simultaneous face and facial expression recognition algorithm, which can classify the given image into one of the seven basic facial expression categories, and then other facial expressions of the new person can be synthesized using the learned expression subspace model. The contributions of this work lie mainly in two aspects. First, we propose a new HOSVD based approach to model the mapping between persons and expressions, used for facial expression synthesis for a new person. Second, we realize simultaneous face and facial expression recognition as a result of facial expression decomposition. Experimental results are presented that illustrate the capability of the person subspace and expression subspace in both synthesis and recognition tasks. As a quantitative measure of the quality of synthesis, we propose using gradient minimum square error (GMSE) which measures the gradient difference between the original and synthesized images.
Narendra Ahuja
ICCV2
2003 A state-free causal video encoding paradigm
abstract
A commonly encountered problem in the communication of predictively encoded video is that of predictive mismatch or drift. The problem of predictive mismatch manifests itself in numerous communication scenarios, including on-demand streaming, real-time streaming and multicast streaming. This paper proposes a state-free video encoding architecture that alleviates this problem. The main benefit of state-free encoding is that there is no need for the encoder and the decoder to maintain the same state, or equivalently, predict using the same predictor. This facilitates robust communication of causally encoded media. The proposed approach is based on the Wyner-Ziv theorem in information theory. Consequently, it leverages the superior performance of coset codes for the Wyner-Ziv problem for predictive coding. A video codec, with state-free functionality, based on the H.26L encoding standard is proposed. The performance of the proposed codec is within 1-2.5 dB of the H.26L encoder.
Anshul Sehgal, Ashish Jagmohan, Narendra Ahuja
ICIP (1)3
2003 WYZE-PMD based multiple description video codec
abstract
The main hindrance to the development of efficient low-latency multiple description (MD) video coders are the problem of predictive mismatch. In this paper, we present a two-channel predictive MD video codec architecture based on the recently proposed WYZE-PMD framework. The proposed codec transmits coset information to curtail error-propagation caused by predictive mismatch, without requiring high latency or restrictive channel assumptions. MD scalar quantizers are used to generate multiple descriptions, low-density parity check (LDPC) codes are used to generate coset information, and the H.263 video coding standard is used for efficient motion compensation. The proposed codec is used to code descriptions of CIF video for communication over two erasure channels with independent failure probabilities. Results indicate that the proposed codec provides efficient, drift-free predictive MD coding.
Ashish Jagmohan, Anshul Sehgal, Narendra Ahuja
ICME3
2003 Easy Calibration of a Head-Mounted Projective Display for Augmented Reality Systems
abstract
Augmented reality (AR) superimposes computer-generated virtual images on the real world to allow users exploring both virtual and real worlds simultaneously. For a successful augmented reality application, an accurate registration of a virtual object with its physical counterpart has to be achieved, which requires precise knowledge of the projection information of the viewing device. The paper proposes a fast and easy off-line calibration strategy based on well-established camera calibration methods. Our method does not need exhausting effort on the collection of world-to-image correspondence data. All the correspondence data are sampled with an image based method and they are able to achieve sub-pixel accuracy. The method is applicable for all AR systems based on optical see-through head-mounted display (HMD), though we took a head-mounted projective display (HMPD) as the example. We first review the calibration requirements for an augmented reality system and the existing calibration methods. Then a new view projection model for optical see through HMD is addressed in detail, and proposed calibration method and experimental result are presented. Finally, the evaluation experiments and error analysis are also included. The evaluation results show that our calibration method is fairly accurate and consistent.
Chunyu Gao, Hong Hua, Narendra Ahuja
VR3
2003 A New Collaborative Infrastructure: SCAP
abstract
This paper presents a multi-user collaborative infrastructure, SCAPE (an acronym for Stereoscopic Collaboration in Augmented and Projective Environments), which is based on recent advancement in head-mounted projective display (HMPD) technology. The SCAPE mainly consists of a 3'/spl times/5' interactive workbench and a 12'/spl times/12'/spl times/9' room-sized walk-through display environment, multiple head-tracked HMPDs, multi-modality interface devices, and a generic application-programming interface (API) designed to coordinate the components. The infrastructure provides a shared space in which multiple users can simultaneously interact with a 3D synthetic environment from their individual viewpoints. We detail the SCAPE implementation and include an application example that demonstrates major interface and cooperation features.
Hong Hua, Leonard D. Brown, Chunyu Gao, Narendra Ahuja
VR4
2003 A scheme for spatial scalability using nonscalable encoders
abstract
We describe a scheme that achieves spatially scalable coding of video by employing nonscalable video encoders (e.g., MPEG-2 main profile), along with a downsampler and an upsampler. The scheme is illustrated for the case of coding video at two resolutions. The enhancement layer is coded in two steps by first exploiting the spatial redundancy and then exploiting the temporal redundancy. Hence, the scheme has a separable implementation. Results are presented for five different sequences, coded for three different combinations of base and enhancement layer bit rates. When MPEG-2 main profile is used for the nonscalable encoders, the results obtained are comparable to the performance of MPEG-2 spatial scalability profile.
Rakesh Dugad, Narendra Ahuja
IEEE Trans. Circuits Syst. Video Technol.2
2002 A Tale of Two Classifiers: SNoW vs. SVM in Visual Recognition
Ming-Hsuan Yang 0001, Dan Roth 0001, Narendra Ahuja
ECCV (4)3
2002 Isotropic error diffusion halftoning
abstract
The inherently causal nature of conventional single-pass error diffusion (ED) halftoning results in asymmetric diffusion of error. This results in the introduction of directional artifacts in the output halftone. In this paper we propose a novel two-pass algorithm which achieves symmetric error diffusion by using a zero-phase signal transfer function. We determine conditions under which isotropic diffusion of error and noise suppression are achieved. Experimental results demonstrate that the proposed algorithm breaks up worms and randomizes their direction, thus making the output halftone more visually appealing as compared to conventional error diffusion.
Ashish Jagmohan, Anshul Sehgal, Narendra Ahuja
ICASSP3
2002 Face recognition using feature extraction based on independent component analysis
abstract
We have explored a new method of feature extraction for face recognition. It is based on independent component analysis (ICA), but unlike original ICA, one of the unsupervised learning methods, it is developed to be well suited for classification problems by utilizing class information. By using ICA in solving supervised classification problems, we can obtain new features which are made as independent from each other as possible and which convey the class information faithfully. We have applied this method on Yale face databases and AT and T face databases and compared the performance with those of conventional methods such as principal component analysis (PCA), Fisher's linear discriminant (FLD), and so on. The experimental results show that for both databases the proposed method outperforms the others.
Nojun Kwak, Chong-Ho Choi, Narendra Ahuja
ICIP (2)3
2002 Predictive encoding using coset codes
abstract
Predictive encoding with respect to multiple possible predictors is a common scenario encountered in many digital set-top box applications, such as redundant storage of video/audio data, real-time robust communication with peripherals and Internet video/audio telephony. A key problem associated with this scenario is that of predictive mismatch or drift. In the present paper, we pose the problem of predictive encoding with multiple possible predictors as a variant of the well-known Wyner-Ziv side-information problem. We propose an approach based on the use of coset codes for predictive encoding, for mitigating the effect of drift without overly sacrificing compression efficiency. The proposed approach can be used to improve coding performance in a wide range of practical applications such as multiple description coding, scalable coding and redundant storage of video/audio streams. We illustrate the efficacy of the proposed approach through a simple example based on the application of low-delay Internet telephony. Our results indicate that the proposed approach significantly outperforms conventional predictive encoding for communication over lossy channels.
Anshul Sehgal, Ashish Jagmohan, Narendra Ahuja
ICIP (2)3
2002 Object contour tracking using graph cuts based active contours
abstract
In this paper, we present an object contour tracking approach using graph cuts based active contours (GCBAC). Our proposed algorithm does not need any a priori global shape model, which makes it useful for tracking objects with deformable shapes and appearances. GCBAC are not sensitive to initial conditions and always converge to the optimal contour within the dilated neighborhood of itself. Given an initial boundary near the object in the first frame, GCBAC can iteratively converge to an optimal object boundary. In each frame thereafter, the resulting contour in the previous frame is taken as initialization and the algorithm consists of two steps. In the first step, GCBAC are applied to the difference between this frame and its previous one. The resulting contour is taken as initialization of the second step, which applies GCBAC to current frame directly. To evaluate the tracking performance, we apply the algorithm to several real world video sequences. Experimental results are provided.
Ning Xu 0005, Narendra Ahuja
ICIP (3)2
2002 Iterative 3D surface modelling from a sparse set of matched feature points
abstract
We present an iterative algorithm to reconstruct a 3D object surface from a sparse set of matched feature points on the input stereo images of the object. The initial matches are sparse and do not have to be accurate. The reconstructed 3D surface is represented in terms of triangular polygons whose vertices are initially the 3D points corresponding to these matched feature points. In order to render photorealistic images of the surface, these feature points are iteratively updated. New feature points are added into the feature point set as well as the depth estimates of the feature points are refined. Experimental results showing the updated correspondences, reconstructed surfaces and virtual views rendered from new directions are presented.
Ning Xu 0005, Narendra Ahuja
ICME (1)2
2002 Calibration of a Head-Mounted Projective Display for Augmented Reality Systems
abstract
In augmented reality (AR) applications, registering a virtual object with its real counterpart accurately and comfortably is one of the basic and challenging issues in the sense that the size, depth, geometry, as well as physical attributes of the virtual objects have to be rendered precisely relative to a physical reference, which is well known as the calibration or registration problem. This paper presents a systematic calibration process to address static registration in a custom-designed augmented reality system, which is based upon the recent advancement of head-mounted projective display (HMPD) technology. Following a concise review of the HMPD concept and system configuration, we present in detail a computational model for system calibration, describe the calibration procedures to obtain estimations of the unknown transformations, and include the calibration results, evaluation experiments and results.
Hong Hua, Chunyu Gao, Narendra Ahuja
ISMAR3
2002 A Testbed for Precise Registration, Natural Occlusion and Interaction in an Augmented Environment Using a Head-Mounted Projective Display (HMPD)
abstract
A head-mounted projective display (HMPD) consists of a pair of miniature projection lenses, beam splitters and displays mounted on the helmet and retro-reflective sheeting materials placed strategically in the environment. This has recently been proposed as an alternative to existing 3D visualization devices. In this paper, we first briefly review HMPD technology, including its featured capabilities and the recent development in both display implementations and applications. Then the implementation of a testbed for playing a "Go" game with a remote opponent in a 3D augmented environment is described. The testbed not only demonstrates the capabilities of virtual-real augmentation and registration, the natural occlusion of virtual objects by real ones, interaction with augmented environments and networking collaboration, but also embodies part of our long-term objective to develop a collaborative framework in 3D augmented environments. Through the testbed, major calibration issues, such as accommodation/convergence considerations and determination of viewing transformations, are studied and discussed in detail. Both calibration methods and results are included, which are applicable to other applications. Finally, experimental results of the testbed implementation are presented.
Hong Hua, Chunyu Gao, Leonard D. Brown, Narendra Ahuja, Jannick P. Rolland
VR4
2002 Mean-Shift Segmentation with Wavelet-based Bandwidth Selection
abstract
Recently, various non-linear techniques for segmentation have been proposed based on non-parametric density estimation. These approaches model image data as clusters of pixels in the combined range-domain space, using kernel based techniques to represent the underlying, multi-modal Probability Density Function (PDF). In Mean-shift based segmentation, pixel clusters or image segments are identified with unique modes of the multi-modal PDF by mapping each pixel to a mode using a convergent, iterative process. The advantages of such approaches include flexible modeling of the image and noise processes and consequent robustness in segmentation. An important issue is the automatic selection of scale parameters a problem far from satisfactorily addressed. In this paper, we propose a regression-based model which admits a realistic framework to choose scale parameters. Results on real images are presented.
Maneesh Kumar Singh 0001, Narendra Ahuja
WACV2
2002 Appearance-based Eye Gaze Estimation
abstract
We present a method for estimating eye gaze direction, which represents a departure from conventional eye gaze estimation methods, the majority of which are based on tracking specific optical phenomena like corneal reflection and the Purkinje images. We employ an appearance manifold model, but instead of using a densely sampled spline to perform the nearest manifold point query, we retain the original set of sparse appearance samples and use linear interpolation among a small subset of samples to approximate the nearest manifold point. The advantage of this approach is that since we are only storing a sparse set of samples, each sample can be a high dimensional vector that retains more representational accuracy than short vectors produced with dimensionality reduction methods. The algorithm was tested with a set of eye images labelled with ground truth point-of-regard coordinates. We have found that the algorithm is capable of estimating eye gaze with a mean angular error of 0.38 degrees, which is comparable to that obtained by commercially available eye trackers.
Kar-Han Tan, David J. Kriegman, Narendra Ahuja
WACV3
2002 A Pupil-Centric Model of Image Formation
Manoj Aggarwal, Narendra Ahuja
Int. J. Comput. Vis.2
2002 Learning to Recognize Three-Dimensional Objects
abstract
A learning account for the problem of object recognition is developed within the probably approximately correct (PAC) model of learnability. The key assumption underlying this work is that objects can be recognized (or discriminated) using simple representations in terms of syntactically simple relations over the raw image. Although the potential number of these simple relations could be huge, only a few of them are actually present in each observed image, and a fairly small number of those observed are relevant to discriminating an object. We show that these properties can be exploited to yield an efficient learning approach in terms of sample and computational complexity within the PAC model. No assumptions are needed on the distribution of the observed objects, and the learning performance is quantified relative to its experience. Most important, the success of learning an object representation is naturally tied to the ability to represent it as a function of some intermediate representations extracted from the image. We evaluate this approach in a large-scale experimental study in which the SNoW learning architecture is used to learn representations for the 100 objects in the Columbia Object Image Library. Experimental results exhibit good generalization and robustness properties of the SNoW-based method relative to other approaches. SNoW's recognition rate degrades more gracefully when the training data contains fewer views, and it shows similar behavior in some preliminary experiments with partially occluded objects.
Dan Roth 0001, Ming-Hsuan Yang 0001, Narendra Ahuja
Neural Comput.3
2002 Extraction of 2D Motion Trajectories and Its Application to Hand Gesture Recognition
abstract
We present an algorithm for extracting and classifying two-dimensional motion in an image sequence based on motion trajectories. First, a multiscale segmentation is performed to generate homogeneous regions in each frame. Regions between consecutive frames are then matched to obtain two-view correspondences. Affine transformations are computed from each pair of corresponding regions to define pixel matches. Pixels matches over consecutive image pairs are concatenated to obtain pixel-level motion trajectories across the image sequence. Motion patterns are learned from the extracted trajectories using a time-delay neural network. We apply the proposed method to recognize 40 hand gestures of American Sign Language. Experimental results show that motion patterns of hand gestures can be extracted and recognized accurately using motion trajectories.
Ming-Hsuan Yang 0001, Narendra Ahuja, Mark Tabb
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Detecting Faces in Images: A Survey
abstract
Images containing faces are essential to intelligent vision-based human-computer interaction, and research efforts in face processing include face recognition, face tracking, pose estimation and expression recognition. However, many reported methods assume that the faces in an image or an image sequence have been identified and localized. To build fully automated systems that analyze the information contained in face images, robust and efficient face detection algorithms are required. Given a single image, the goal of face detection is to identify all image regions which contain a face, regardless of its 3D position, orientation and lighting conditions. Such a problem is challenging because faces are non-rigid and have a high degree of variability in size, shape, color and texture. Numerous techniques have been developed to detect faces in a single image, and the purpose of this paper is to categorize and evaluate these algorithms. We also discuss relevant issues such as data collection, evaluation metrics and benchmarking. After analyzing these algorithms and identifying their limitations, we conclude with several promising directions for future research.
Ming-Hsuan Yang 0001, David J. Kriegman, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.3
2002 Lossless image compression with multiscale segmentation
abstract
This paper is concerned with developing a lossless image compression method which employs an optimal amount of segmentation information to exploit spatial redundancies inherent in image data. Multiscale segmentation is obtained using a previously proposed transform which provides a tree-structured segmentation of the image into regions characterized by grayscale homogeneity. In the proposed algorithm we prune the tree to control the size and number of regions thus obtaining a rate-optimal balance between the overhead inherent in coding the segmented data and the coding gain that we derive from it. Another novelty of the proposed approach is that we use an image model comprising separate descriptions of pixels lying near the edges of a region and those lying in the interior. Results show that the proposed algorithm can provide performance comparable to the best available methods and 15-20% better compression when compared with the JPEG lossless compression standard for a wide range of images.
Krishna Ratakonda, Narendra Ahuja
IEEE Trans. Image Process.2
2001 A High-Resolution Panoramic Camer
abstract
Wide field of view (FOV) and high resolution are two desirable properties in many vision-based applications such as tele-conferencing, surveillance, and robot navigation. In some applications, such as 3D reconstruction and rendering, it is also desired that all viewing directions share a single viewpoint, the entire FOV be imaged simultaneously, in real-time, and the depth of field be large. In this paper, we review such a panoramic camera proposed by Nalwa in 1996 that uses reflections off planar mirrors to achieve the first four of the aforementioned capabilities He uses a single mirror pyramid (SMP) and a number of cameras that point to the individual pyramid faces. Together the cameras yield a visual field having a width of 360 degrees and a height same as that of the individual cameras. We propose a double mirror-pyramid (DALP) design that still achieves a 360-degree FOV horizontally but doubles the vertical FOV. It retains the other three capabilities namely high resolution, a single apparent viewpoint across the entire FOV, and real-time panoramic capture. We specify the visual field mapping from the scene to the sensor realized by, the proposed camera. Finally, an implementation of the proposed DMP design is described and examples of preliminary panoramic images obtained are included.
Hong Hua, Narendra Ahuja
CVPR (1)2
2001 A Representation of Image Structure and Its Application to Object Selection Using Freehand Sketches
abstract
We present an algorithm for computing a representation of image structure, or image segmentation, and use it for selecting objects in the image with freehand sketches drawn by the user over the image. The sketches are mapped onto image segments whose union forms the intended object. The mapping operation is performed with the aid of a simplicial decomposition of the image segmentation-a triangulation formed with vertices chosen to lie along the medial axes of the segments. Each edge of a triangle lies entirely inside the two segments that contains its vertices. This decomposition captures the adjacency information about the segments as well as the shape of the segment boundaries. Any object boundary is completely contained in a set of triangles. The triangles are also used to formulate the problem of estimating gradual photometric transition across an object boundary, called alpha channel estimation, as a set of local, intratriangle alpha channel estimation problems that can then be solved more accurately, independently, and in parallel. Experimental results are included to show how the algorithm allows selection of image objects with complex boundaries using roughly drawn simple sketches.
Kar-Han Tan, Narendra Ahuja
CVPR (2)2
2001 Spin Discriminant Analysis (SDA) - Using a One-Dimensional Classifier for High Dimensional Classification Problems
abstract
In this paper we discuss how to use a one-dimensional classifier for solving high dimensional classification problems. We propose Spin Discriminant Analysis (SDA), which enables us to construct a family of new classifiers. We prove that SDA is equivalent to ridged Linear Discriminant Analysis (LDA) when two classes are Gaussians with common covariance matrices. Moreover, we prove that classification based on Parzen's window, is a special case of SDA. In addition to theoretical investigations, we conduct extensive empirical studies, implementing SDA using Support Vector Machines (SVMs) as its one-dimensional classifiers. This SVM-based SDA implementation is named SpinSVM. Our experiments show that SpinSVM outperforms traditional high dimensional classifiers like SVMs, Classification Using Spline (CUS), classification-based Parzen's window, and LDA on most standard and synthetic datasets we tested.
Huaxin You, Hong Hua, Narendra Ahuja, Edward Y. Chang
CVPR (1)3
2001 High Dynamic Range Panoramic Imaging
Manoj Aggarwal, Narendra Ahuja
ICCV2
2001 A New Imaging Model
Manoj Aggarwal, Narendra Ahuja
ICCV2
2001 Split Aperture Imaging for High Dynamic Range
abstract
Most imaging sensors have limited dynamic range and hence are sensitive to only a part of the illumination range present in a natural scene. The dynamic range can be improved by acquiring multiple images of the same scene under different exposure settings and then combining them. In this paper, we describe a camera design for simultaneously acquiring multiple images of the same scene under different exposure settings. The cross-section of the incoming beam from a scene point is partitioned into as many parts as the desired degree of split. This is done by splitting the aperture into multiple parts and directing the light exiting from each in a different direction using an assembly of mirrors. A sensor is placed in the path of each beam and exposure of each sensor is controlled either by appropriately setting its exposure parameter, or by splitting the incoming beam unevenly. The resulting multiple exposure images are used to construct a high dynamic range image. We have implemented a video-rate camera based on this design and the results obtained are presented.
Manoj Aggarwal, Narendra Ahuja
ICCV2
2001 On Cosine-Fourth and Vignetting Effects in Real Lenses
Manoj Aggarwal, Hong Hua, Narendra Ahuja
ICCV3
2001 Selecting Objects With Freehand Sketches
Kar-Han Tan, Narendra Ahuja
ICCV2
2001 High capacity data embedding in the wavelet domain
abstract
A novel solution to the problem of data embedding in images is proposed in this paper The proposed algorithm allows high capacity data embedding and is robust to JPEG image compression data is embedded in the wavelet domain which provides better perceptual masking compared to the DCT domain set partitioning in hierarchical trees (SPIHT) is used to control the distortion (in the sense of PSNR) in the embedded host. Unlike other data embedding algorithms available in literature, the proposed algorithm provides control over the BER of the embedded data by appropriately choosing the JPEG quantization matrix. Preliminary results of an implementation of the algorithm are also presented.
Anshul Sehgal, Ashish Jagmohan, Narendra Ahuja
ICIP (3)3
2001 Face detection using large margin classifiers
abstract
Large margin classifiers have demonstrated their advantages in many visual learning tasks, and have attracted much attention in vision and image processing communities. We apply and compare two large margin classifiers, support vector machines and sparse network of winnows, to detect faces in still gray scale images Furthermore, we study the theoretical frameworks of these classifiers and analyze the empirical results. Experiments on a test set of 24,045 images exhibit good generalization and robustness, and conform to theoretical analysis.
Ming-Hsuan Yang 0001, Dan Roth 0001, Narendra Ahuja
ICIP (2)3
2001 Frame interpolation using transmitted block-based motion vectors
abstract
The proposed technique is designed for interpolating video frames often dropped during compression using standards, such as MPEG-4 and H.263, at very low bit rates. The coarse motion vector field is refined at the receiving side using mesh-based motion estimation instead of using computationally intensive dense motion estimation. We propose a framework for detecting and utilizing local motion boundaries in terms of an explicit model. Motion boundaries are modeled using edge detection and Hough transform. The motion of the occluding side is represented by affine mapping. The newly appearing region is also detected. Each pixel is interpolated differently according to adaptive interpolation based on the property of the pixel: moving, static and disoccluded.
Seung Chul Yoon, Narendra Ahuja
ICIP (3)2
2001 Face Detection Using Multimodal Density Models
Ming-Hsuan Yang 0001, David J. Kriegman, Narendra Ahuja
Comput. Vis. Image Underst.3
2001 A fast scheme for image size change in the compressed domain
abstract
Given a video frame in terms of its 8/spl times/8 block-DCT coefficients, we wish to obtain a downsized or upsized version of this frame also in terms of 8/spl times/8 block-DCT coefficients. The DCT being a linear unitary transform is distributive over matrix multiplication. This fact has been used for downsampling video frames in the DCT domain. However, this involves matrix multiplication with the DCT of the downsampling matrix. This multiplication can be costly enough to trade off any gains obtained by operating directly in the compressed domain. We propose an algorithm for downsampling and also upsampling in the compressed domain which is computationally much faster, produces visually sharper images, and gives significant improvements in PSNR (typically 4-dB better compared to bilinear interpolation). Specifically the downsampling method requires 1.25 multiplications and 1.25 additions per pixel of original image compared to 4.00 multiplications and 4.75 additions required by the method of Chang et al. (1995). Moreover, the downsampling and upsampling schemes combined together preserve all the low-frequency DCT coefficients of the original image. This implies tremendous savings for coding the difference between the original frame (unsampled image) and its prediction (the upsampled image). This is desirable for many applications based on scalable encoding of video. The method presented can also be used with transforms other than DCT, such as Hadamard or Fourier.
Rakesh Dugad, Narendra Ahuja
IEEE Trans. Circuits Syst. Video Technol.2
2000 Region Correspondence by Global Configuration Matching and Progressive Delaunay Triangulation
abstract
In this paper, we present a novel algorithm for establishing region correspondences across images by first matching global region configuration and then propagating the matches locally constrained by Delaunay triangulation. We exploit a global configuration constraint, which has not been explicitly used in existing matching algorithms. The proposed algorithm is comprised of two stages: first, stable regions are matched by enforcing the global configuration constraint. This yields a set of global matches corresponding to stable regions distributed over the images. In the second stage, these matches are used to guide the matching of the remaining unmatched regions in the intervening spaces. This is done by enforcing local positioning constraint, which starts with the Delaunay triangulation defined by the global matches and performs progressive Delaunay triangulation for local matching. Experiments on both stereo and motion images are presented to show the effectiveness of the proposed algorithm.
Narendra Ahuja
CVPR2
2000 Learning to Recognize Objects
abstract
A learning account for the problem of object recognition is developed within the PAC (Probably Approximately Correct) model of learnability. The proposed approach makes no assumptions on the distribution of the observed objects, but quantifies success relative to its past experience. Most importantly, the success of learning an object representation is naturally tied to the ability to represent at as a function of some intermediate representations extracted from the image. We evaluate this approach an a large scale experimental study in which the SNoW learning architecture is used to learn representations for the 100 objects in the Columbia Object Image Database (COIL-100). The SNoW-based method is shown to outperform other methods in terms of recognition rates; its performance degrades gracefully when the training data contains fewer views and in the presence of occlusion noise.
Dan Roth 0001, Ming-Hsuan Yang 0001, Narendra Ahuja
CVPR3
2000 A Geometric Approach to Train Support Vector Machines
abstract
Support Vector Machines (SVMs) have shown great potential in numerous visual learning and pattern recognition problems. The optimal decision surface of a SVM is constructed from its support vectors which are conventionally determined by solving a quadratic programming (QP) problem. However, solving a large optimization problem is challenging since it is computationally intensive and the memory requirement grows with square of the training vectors. In this paper, we propose a geometric method to extract a small superset of support vectors, which we call guard vectors, to construct the optimal decision surface. Specifically, the guard vectors are found by solving a set of linear programming problems. Experimental results on synthetic and real data sets show that the proposed method is more efficient than conventional methods using QPs and requires much less memory.
Ming-Hsuan Yang 0001, Narendra Ahuja
CVPR2
2000 Learning to Recognize 3D Objects with SNoW
Ming-Hsuan Yang 0001, Dan Roth 0001, Narendra Ahuja
ECCV (1)3
2000 Face Detection Using Mixtures of Linear Subspaces
abstract
We present two methods using mixtures of linear sub-spaces for face detection in gray level images. One method uses a mixture of factor analyzers to concurrently perform clustering and, within each cluster, perform local dimensionality reduction. The parameters of the mixture model are estimated using an EM algorithm. A face is detected if the probability of an input sample is above a predefined threshold. The other mixture of subspaces method uses Kohonen's self-organizing map for clustering and Fisher linear discriminant to find the optimal projection for pattern classification, and a Gaussian distribution to model the class-conditioned density function of the projected samples for each class. The parameters of the class-conditioned density functions are maximum likelihood estimates and the decision rule is also based on maximum likelihood. A wide range of face images including ones in different poses, with different expressions and under different lighting conditions are used as the training set to capture the variations of human faces. Our methods have been tested on three sets of 225 images which contain 871 faces. Experimental results on the first two datasets show that our methods perform as well as the best methods in the literature, yet have fewer false detects.
Ming-Hsuan Yang 0001, Narendra Ahuja, David J. Kriegman
FG2
2000 A Scheme for Joint Watermarking and Compression of Video
abstract
We present a scheme for jointly watermarking and compressing digital video. The amount of watermark added is adapted to the expected degradation of the watermark due to compression. This results in a more robust watermark. This is achieved without any appreciable decrease in the quality of the decoded video compared to the case when the watermark is not adaptive. Results are presented for the flower garden sequence.
Rakesh Dugad, Narendra Ahuja
ICIP2
2000 Recovering Frontal-Pose Image from a Single Profile Image
abstract
In appearance based face recognition, lip reading, etc., eigen face and eigen lip are used for recognition. The pose changes of the human head in a video sequence often cause errors in the eigen space comparison stage, because the frontal-pose assumption has been violated. We propose a new method to compensate the pose changes by exploiting the general symmetry of human face. From the imaging geometry we show that a frontal pose can be recovered from only one profile view. The resulting pose compensation method has the following advantages: (1) it only requires one profile image; (2) it does not need any 3D model; (3) it does not need accurate feature detection. Experimental results in the context of lip images are given to show the effectiveness of our method.
Narendra Ahuja, Chalapathy Neti, Andrew W. Senior
ICIP2
2000 Face Recognition Using Kernel Eigenfaces
abstract
Eigenface or principal component analysis (PCA) methods have demonstrated their success in face recognition, detection, and tracking. The representation in PCA is based on the second order statistics of the image set, and does not address higher order statistical dependencies such as the relationships among three or more pixels. Higher order statistics (HOS) have been used as a more informative low dimensional representation than PCA for face and vehicle detection. We investigate a generalization of PCA, kernel principal component analysis (kernel PCA), for learning low dimensional representations in the context of face recognition. In contrast to HOS, kernel PCA computes the higher order statistics without the combinatorial explosion of time and memory complexity. While PCA aims to find a second order correlation of patterns, kernel PCA provides a replacement which takes into account higher order correlations. We compare the recognition results using kernel methods with eigenface methods on two benchmarks. Empirical results show that kernel PCA outperforms the eigenface method in face recognition.
Ming-Hsuan Yang 0001, Narendra Ahuja, David J. Kriegman
ICIP2
2000 On Generating Seamless Mosaics with Large Depth of Field
abstract
Imaging cameras have only finite depth of field and only those objects within that depth range are simultaneously in focus. The depth of field of a camera can be improved by mosaicing a sequence of images taken under different focal settings. In conventional mosaicing schemes, a focus measure is computed for every scene point across the image sequence and the point is selected from that image where the focus measure is highest. We have, however, proved in this paper that the focus measure is not the highest in the best focussed frame for a certain class of scene points. The incorrect selection of image frames for these points, causes visual artifacts to appear in the resulting mosaic. We have also proposed a method to isolate such scene points, and an algorithm to compose large depth of field mosaics without the undesirable artifacts.
Manoj Aggarwal, Narendra Ahuja
ICPR2
2000 Camera Center Estimation
abstract
A fast camera calibration technique to estimate the center of perspective projection of an imaging camera has been described. The proposed technique requires a single image of two planar calibration charts arranged in a special manner. The special arrangement of the calibration charts simplifies the projection equations relating the 3-D scene coordinates to the 2-D image coordinates and they can be suitably combined to eliminate the unknown intrinsic parameters other than the desired center. We have analyzed the error in the center estimate due to various alignment errors in the experimental setup, and shown that the scheme is quite robust.
Manoj Aggarwal, Narendra Ahuja
ICPR2
2000 Estimating Sensor Orientation in Cameras
abstract
For most imaging cameras, it is desirable that the sensor plane be perpendicular to the optical axis. Such an orientation ensures that the imaging configuration is perspective and planar scene objects perpendicular to the optical axis can be focussed in their entirety. In this paper, we present an image processing method to estimate and subsequently correct the sensor tilt with precision. We propose to measure the tilt by measuring the variation of defocus in an image of a planar calibration chart placed perpendicular to the optical axis. We show that the proposed defocusing based method is inherently more accurate than geometry based techniques which estimate tilt by measuring the deviation in the geometry of a scene pattern. We analyze the sensitivity of the tilt estimates to errors in the experimental setup and show that the proposed technique is quite robust to errors even as large as 1 degree in the orientation of the calibration chart.
Manoj Aggarwal, Narendra Ahuja
ICPR2
1999 A Fast Scheme for Altering Resolution in the Compressed Domain
abstract
Given a video frame or image in terms of its 8/spl times/8 block-DCT coefficients we wish to obtain a downsized (lower resolution) or upsized (higher resolution) version of this frame also in terms of 8/spl times/8 block -DCT coefficients. We propose an algorithm for achieving this directly in the compressed domain which is computationally much faster, produces visually sharper images and gives significant improvements in PSNR (typically 4 dB better compared to other compressed domain methods based on bilinear interpolation). The downsampling and upsampling schemes combined together preserve all the low-frequency DCT coefficients of the original signal. This implies tremendous savings for coding the difference between the original (unsampled image) and its prediction (the upsampled image). This is desirable for many applications based on scalable encoding of video.
Rakesh Dugad, Narendra Ahuja
CVPR2
1999 Recognizing Hand Gesture Using Motion Trajectories
abstract
We present an algorithm for extracting and classifying two-dimensional motion in an image sequence based on motion trajectories. First, a multiscale segmentation is performed to generate homogeneous regions in each frame. Regions between consecutive frames are then matched to obtain 2-view correspondences. Affine transformations are computed from each pair of corresponding regions to define pixel matches. Pixels matches over consecutive images pairs are concatenated to obtain pixel-level motion trajectories across the image sequence. Motion patterns are learned from the extracted trajectories using a time-delay neural network. We apply the proposed method to recognize 40 hand gestures of American Sign Language. Experimental results show that motion patterns in hand gestures can be extracted and recognized with high recognition rate using motion trajectories.
Ming-Hsuan Yang 0001, Narendra Ahuja
CVPR2
1999 A Fast Scheme for Downsampling and Upsampling in the DCT Domain
abstract
Given a video frame or image in terms of its 8×8 block-DCT coefficients, we wish to obtain a downsized or upsized (by factor of two) version of this frame also in terms of 8×8 block-DCT coefficients. We propose an algorithm for achieving this directly in the DCT domain which is computationally much faster, produces visually sharper images and gives significant improvements in PSNR (typically 4 dB better), compared to other compressed domain methods based on bilinear interpolation. The downsampling and upsampling schemes combined together preserve all the low-frequency DCT coefficients of the original image. This implies tremendous savings for coding the difference between the original (unsampled image) and its prediction (the upsampled image). This is desirable for many applications based on scalable encoding of video.
Rakesh Dugad, Narendra Ahuja
ICIP (2)2
1999 Video Denoising by Combining Kalman and Wiener Estimates
abstract
The paper proposes a computationally fast scheme for denoising a video sequence. Temporal processing is done separately from spatial processing and the two are then combined to get the denoised frame. The temporal redundancy is exploited using a scalar state 1D Kalman filter. A novel way is proposed to estimate the variance of the state noise from the noisy frames. The spatial redundancy is exploited using an adaptive edge-preserving Wiener filter. These two estimates are then combined using simple averaging to get the final denoised frame. Simulation results for the foreman, trevor and susie sequences show an improvement of 6 to 8 dB in PSNR over the noisy frames at PSNR of 28 and 24 dB.
Rakesh Dugad, Narendra Ahuja
ICIP (4)2
1999 Segmentation Based Denoising Using Multiple Compaction Domains
abstract
In this paper, we propose a novel segmentation based denoising algorithm. Segmentation yields intrinsically homogeneous and extrinsically heterogeneous regions. A denoising algorithm that uses Multiple Compaction Domains (MCD) is then applied on each of the resulting segments. Such a scheme retains important perceptual information in the segment boundaries while the denoising algorithm operates only on homogeneous segments. Further, the MCD algorithm is demonstrably superior to the classical denoising algorithms using transform domain thresholding. Our algorithm yields better perceptual quality and superior PSNR as compared to MATLAB's adaptive Wiener filter.
Maneesh Kumar Singh 0001, Prakash Ishwar, Krishna Ratakonda, Narendra Ahuja
ICIP (1)4
1999 Face Detection Using a Mixture of Factor Analyzers
abstract
We present a probabilistic method to detect human faces using a mixture of factor analyzers. One characteristic of this mixture model is that it concurrently performs clustering and, within each cluster, local dimensionality reduction. A wide range of face images including ones in different poses, with different expressions and under different lighting conditions are used as the training set to capture the variations of human faces. In order to fit the mixture model to the sample face images, the parameters are estimated using an EM algorithm. Experimental results show that faces in different poses, with different facial expressions, and under different lighting conditions are accurately detected by our method.
Ming-Hsuan Yang 0001, Narendra Ahuja, David J. Kriegman
ICIP (3)2
1999 A data partition method for parallel self-organizing map
abstract
We propose a method to partition training vectors into clusters for a parallel implementation of self-organizing map (SOM) algorithm. The proposed algorithm assigns a cluster to a processor such that, in updating weights, the neighbourhoods of a winning node in a cluster do not overlap the neighboring nodes of some winning nodes in other clusters. It reduces the overheads caused by synchronization (i.e., maintaining coherency) of the weight matrices in the processors since the proposed algorithm allows multiple vectors to find their winning nodes and update weights in parallel. Our experimental results show that an average speedup of 3.15 for a parallel implementation of a four processor simulation.
Ming-Hsuan Yang 0001, Narendra Ahuja
IJCNN2
1999 A SNoW-Based Face Detector
Ming-Hsuan Yang 0001, Dan Roth 0001, Narendra Ahuja
NIPS3
1999 Structure and Motion Estimation from Dynamic Silhouettes under Perspective Projection
Tanuja Joshi, Narendra Ahuja, Jean Ponce
Int. J. Comput. Vis.2
1999 Low bit-rate video coding with implicit multiscale segmentation
abstract
Discusses a multiscale segmentation based video compression algorithm aimed at very low bit-rate applications such as video teleconferencing and video phones. We introduce novel techniques for multiscale segmentation based motion compensation and residual coding. Our region based forward motion compensation strategy (in terms of direction of motion vector, which is from the previous frame to the current frame) regulates the size and number of regions used, by pruning a multiscale segmentation of video frames. Since regions used for motion compensation are obtained by segmenting the previously decoded frame, the shape of the regions need not be transmitted to the decoder. Furthermore, our hierarchical motion compensation strategy refines an initial region level, coarse motion field to obtain a dense motion field which provides pixel level motion vectors. The refinement procedure does not require any additional information to be transmitted. This motion compensation technique effectively addresses the problem of dealing with "holes" and "overlapping regions" which are inherent to forward motion compensation. Residual coding is performed using a novel method which exploits the fact that the energy of the residual resulting from motion compensation is concentrated in a priori predictable positions. We show that this residual coding technique can also be extrapolated to improve the performance of coders using a block based motion compensation strategy. A fusion of these concepts leads to a gain of 2-3 dB in peak signal-to-noise ratio, apart from significant perceptual improvement, over a generic video coding algorithm using a block based motion compensation strategy (such as H.261 or H.263) for a variety of test sequences.
Seung Chul Yoon, Krishna Ratakonda, Narendra Ahuja
IEEE Trans. Circuits Syst. Video Technol.3
1999 A topological and temporal correlator network for spatiotemporal pattern learning, recognition, and recall
abstract
In this paper, we describe the design of an artificial neural network for spatiotemporal pattern recognition and recall. This network has a five-layered architecture and operates in two modes: pattern learning and recognition mode, and pattern recall mode. In pattern learning and recognition mode, the network extracts a set of topologically and temporally correlated features from each spatiotemporal input pattern based on a variation of Kohonen's self-organizing maps. These features are then used to classify the input into categories based on the fuzzy ART network. In the pattern recall mode, the network can reconstruct any of the learned categories when the appropriate category node is excited or probed. The network performance was evaluated via computer simulations of time-varying, two-dimensional and three-dimensional data. The results show that the network is capable of both recognition and recall of spatiotemporal data in an on-line and self-organized fashion. The network can also classify repeated events in the spatiotemporal input and is robust to noise in the input such as distortions in the spatial and temporal content.
Narayan Srinivasa, Narendra Ahuja
IEEE Trans. Neural Networks2
1998 Hierarchical Texture Segmentation
Peter Bajcsy, Narendra Ahuja
ACCV (2)2
1998 Learning Multiscale Image Models of 2D Object Classes
Benoit Perrin, Narendra Ahuja, Narayan Srinivasa
ACCV (2)2
1998 Restoring Image Quality Through Structure Preserving De-noising
Krishna Ratakonda, Narendra Ahuja
ACCV (2)2
1998 A Learning Approach to Fixating on 3D Targets with Active Cameras
Narayan Srinivasa, Narendra Ahuja
ACCV (1)2
1998 Dense Shape and Motion from Region Correspondences by Factorization
abstract
In this paper, we propose an algorithm for estimating dense shape and motion of dynamic piecewise planar scenes from region correspondences using factorization. Region correspondences are used since they are easier to establish and more reliable than either line or point correspondences. The image measurements required are the centroid and area for each region. Singular value decomposition is employed to find the basis of range space of the motion, shape, and surface normal matrices. By imposing model constraints, motion, shape, and surface normal can be recovered only from region correspondences.
Narendra Ahuja
CVPR2
1998 Extraction and Classification of Visual Motion Patterns for Hand Gesture Recognition
abstract
We present a new method for extracting and classifying motion patterns to recognize hand gestures. First, motion segmentation of the image sequence is generated based on a multiscale transform and attributed graph matching of regions across frames. This produces region correspondences and their affine transformations. Second, color information of motion regions is used to determine skin regions. Third, human head and palm regions are identified based on the shape and size of skin areas in motion. Finally, affine transformations defining a region's motion between successive frames are concatenated to construct the region's motion trajectory. Gestural motion trajectories are then classified by a time-delay neural network trained with backpropagation learning algorithm. Our experimental results show that hand gestures can be recognized well using motion patterns.
Ming-Hsuan Yang 0001, Narendra Ahuja
CVPR2
1998 Extracting Gestural Motion Trajectories
Ming-Hsuan Yang 0001, Narendra Ahuja
FG2
1998 Improving the throughput of flexible-precision DSPS via algorithm transformation
abstract
In this paper, we have presented a systematic technique to improve throughput of signal/image processing algorithms when implemented on flexible precision hardware. Many image/signal processing algorithms need 8-16 bit precision while the DSPs available are of much higher precision (32 bit). Significant performance gain can be obtained if multiple low precision computations can be performed in one cycle of a high precision DSP. We have proposed a framework based on algorithm transformation techniques of unfolding and retiming to systematically map low precision algorithms onto high precision DSPs. The improvement in throughput obtained by this framework is linearly related to the ratio of precision used by the processor and that required by the algorithm. The efficacy of this technique has been demonstrated on a IIR filter. We have also established some theoretical bounds on the maximum throughput that can be achieved using the proposed methodology.
Manoj Aggarwal, Naresh R. Shanbhag, Narendra Ahuja
ICASSP3
1998 Unsupervised multidimensional hierarchical clustering
abstract
A method for multidimensional hierarchical clustering that is invariant to monotonic transformations of the distance metric is presented. The method derives a tree of clusters organized according to the homogeneity of intracluster and interpoint distances. Higher levels correspond to coarser clusters. At any level the method can detect clusters of different densities, shapes and sizes. The number of clusters and the parameters for clustering are determined automatically and adaptively for a given data set which makes it unsupervised and non-parametric. The method is simple, noniterative and requires low computation. Results on various sample data sets are presented.
Rakesh Dugad, Narendra Ahuja
ICASSP2
1998 Image denoising using multiple compaction domains
abstract
We present a novel framework for denoising signals from their compact representation in multiple domains. Each domain captures, uniquely, certain signal characteristics better than others. We define confidence sets around data in each domain and find sparse estimates that lie in the intersection of these sets, using a POCS algorithm. Simulations demonstrate the superior nature of the reconstruction (both in terms of mean-square error and perceptual quality) in comparison to the adaptive Wiener filter.
Prakash Ishwar, Krishna Ratakonda, Pierre Moulin, Narendra Ahuja
ICASSP4
1998 A New Wavelet-based Scheme for Watermarking Images
abstract
A new method for digital image watermarking which does not require the original image for watermark detection is presented. Assuming that we are using a transform domain spread spectrum watermarking scheme, it is important to add the watermark in select coefficients with significant image energy in the transform domain in order to ensure non-erasability of the watermark. Previous methods, which did not use the original in the detection process, could not selectively add the watermark to the significant coefficients, since the locations of such selected coefficients can change due to image manipulations. Since watermark verification typically consists of a process of correlation which is extremely sensitive to the relative order in which the watermark coefficients are placed within the image, such changes in the location of the watermarked coefficients was unacceptable. We present a scheme which overcomes this problem of "order sensitivity". Advantages of the proposed method include (i) improved resistance to attacks on the watermark, (ii) implicit visual masking utilizing the time-frequency localization property of the wavelet transform and (iii) a robust definition for the threshold which validates the watermark. We present results comparing our method with previous techniques, which clearly validate our claims.
Rakesh Dugad, Krishna Ratakonda, Narendra Ahuja
ICIP (2)3
1998 POCS based Adaptive Image Magnification
Krishna Ratakonda, Narendra Ahuja
ICIP (3)2
1998 Digital Image Watermarking: Issues in Resolving Rightful Ownership
abstract
In the literature many strategies for attacking and subverting a watermark have been presented. Such attacks suggest that the ability to embed non-erasable watermarks does not necessarily imply that the watermarking scheme can be used to establish ownership. The main aim of this paper is to formulate necessary and sufficient requirements for a watermarking scheme to be able to resolve rightful ownership. It is shown that some of the popular schemes proposed in literature for watermarking images do not satisfy these requirements. Finally, we show that a modification of a watermarking scheme in literature (Piva et al., 1997) performs satisfactorily.
Krishna Ratakonda, Rakesh Dugad, Narendra Ahuja
ICIP (2)3
1998 Detecting Human Faces in Color Images
abstract
We propose a new method to detect human faces in color images. A human skin color model is built to capture the chromatic properties based on multivariate statistical analysis. Given a color image, multiscale segmentation is used to generate homogeneous regions at multiple different scales. From the coarsest to the finest scale, regions of skin color are merged until the shape is approximately elliptic. Postprocessing is performed to determine whether a merged region contains a human face and include the facial features of non-skin color such as eyes and mouth if necessary. Experimental results show that human faces in color images can be detected regardless of size, orientation and viewpoint.
Ming-Hsuan Yang 0001, Narendra Ahuja
ICIP (1)2
1998 Robust video shot change detection
abstract
We present a novel improvement to existing schemes for abrupt shot change detection. Existing schemes declare a shot change whenever the frame to frame histogram difference (FFD) value is above a particular threshold. In such an approach, a high value for the threshold results in a small number of false alarms and a large number of missed detections while a low value for the threshold decreases the number of missed detections at the expense of increasing the false alarms. We attribute this situation to the fact that the FFD cannot be reliably used as the sole indicator for the presence of a shot change. In the proposed method a two-step shot detection strategy is used which selectively uses a likelihood ratio (computed directly from the frames and not from the histograms) to confirm the presence of a shot change. Such a two-step checking increases the probability of detection without increasing the probability of false alarm. The improvement proposed is simple and computationally cheap. Tests with a wide variety of video sequences prove the efficacy of the proposed approach.
Rakesh Dugad, Krishna Ratakonda, Narendra Ahuja
MMSP3
1998 Location- and Density-Based Hierarchical Clustering Using Similarity Analysis
abstract
This paper presents a new approach to hierarchical clustering of point patterns. Two algorithms for hierarchical location- and density-based clustering are developed. Each method groups points such that maximum intracluster similarity and intercluster dissimilarity are achieved for point locations or point separations. Performance of the clustering methods is compared with four other methods. The approach is applied to a two-step texture analysis, where points represent centroid and average color of the regions in image segmentation.
Peter Bajcsy, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Segmentation and Factorization-Based Motion and Structure Estimation for Long Image Sequences
abstract
This paper presents a computer algorithm which, given a dense temporal sequence of intensity images of multiple moving objects, can separate the images into regions showing distinct objects and, for those objects which are rotating, calculate the three-dimensional structure and motion.
Christian Debrunner, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 An analytically tractable potential field model of free space and its application in obstacle avoidance
abstract
An analytically tractable potential field model of free space is presented. The model assumes that the border of every two dimensional (2D) region is uniformly charged. It is shown that the potential and the resulting repulsion (force and torque) between polygonal regions can he calculated in closed form. By using the Newtonian potential function, collision avoidance between object and obstacle thus modeled is guaranteed in a path planning problem. A local planner is developed for finding object paths going through narrow areas of free space where the obstacle avoidance is most important. Simulation results show that not only does individual object configuration of a path obtained with the proposed approach avoid obstacles effectively, the configurations also connect smoothly into a path.
Jen-Hui Chuang, Narendra Ahuja
IEEE Trans. Syst. Man Cybern. Part B2
1997 Discrete multi-dimensional linear transforms over arbitrarily shaped supports
abstract
In order to apply a multi-dimensional linear transform, over an arbitrarily shaped support, the usual practice is to fill out the support to a hypercube by zero padding. This does not however yield a satisfactory definition for transforms in two or more dimensions. The problem that we tackle is: how do we redefine the transform over an arbitrary shaped region suited to a given application? We present a novel iterative approach to define any multi-dimensional linear transform over an arbitrary shape given that we know its definition over a hypercube. The proposed solution is (1) extensible to all possible shapes of support (whether connected or unconnected) (2) adaptable to the needs of a particular application. We also present results for the Fourier transform, for a specific adaptation of the general definition of the transform which is suitable for compression or segmentation algorithms.
Krishna Ratakonda, Narendra Ahuja
ICASSP2
1997 Automated registration of multimodality images by maximization of a region similarity measure
abstract
This paper presents a robust algorithm for automated registration of images related by rigid-body transformations. This algorithm uses a new region-based similarity metric, which enables accurate registration of images of large contrast differences. Region segmentation required by the metric is accomplished using a multiscale segmentation algorithm, and minimization of this metric is done using the Powell direction set method. Experimental results are presented to demonstrate that the algorithm is effective for aligning images from single or multiple imaging modalities without the use of any fiducial markers.
Zhi-Pei Liang, Hao Pan 0002, Richard L. Magin, Narendra Ahuja, Thomas S. Huang
ICIP (3)4
1997 Coding the Displaced Frame Difference for Video Compression
abstract
Popular techniques employed to code the displaced frame difference (DFD) treat it no differently from an ordinary image for coding purposes. Since the DFD is generated by the process of motion compensation, such methods do not fully exploit the underlying redundancies. This paper proposes a DFD coding method which exploits such redundancies while incurring negligible information overhead. The key idea is to predict locations of high DFD concentration which occupy small portions of the image and use this predicted information (which is also available to the decoder without additional information transmission) to improve the quality of the decoded image. Two key features of the proposed approach are its compatibility with any transform based DFD coding scheme and negligible information overhead. Tests with a fully functional video coder show the efficacy of the proposed approach.
Krishna Ratakonda, S. Yoon, Narendra Ahuja
ICIP (1)3
1997 Region-Based Video Coding Using a Multiscale Image Segmentation
abstract
This paper proposes a novel region-based video coding technique using a multiscale image segmentation method thus obtaining better quality at the same bit rate. In most of the previous region-based video coding techniques, occlusion caused degradation in terms of both the PSNR and perceptual video quality. We propose a new motion estimation and compensation algorithm which solves occlusion related problems effectively. The proposed motion estimation and compensation is a two stage procedure. The first stage uses a coarse motion model while the second stage uses a dense motion model. The coarse motion model generates region level motion vectors which are then fine tuned by the dense motion model which produces pixel level motion vectors. A fusion of these concepts leads to a gain of 2/spl sim/3 dB in PSNR over the block-based algorithm for a variety of test sequences using a fully functional video coder.
Seung Chul Yoon, Krishna Ratakonda, Narendra Ahuja
ICIP (2)3
1997 Learning Recognition and Segmentation Using the Cresceptron
Juyang Weng, Narendra Ahuja, Thomas S. Huang
Int. J. Comput. Vis.2
1997 Shape Representation Using a Generalized Potential Field Model
abstract
This paper is concerned with efficient derivation of the medial axis transform of a 2D polygonal region. Instead of using the shortest distance to the region border, a potential field model is used for computational efficiency. The region border is assumed to be charged and the valleys of the resulting potential field are used to estimate the axes for the medial axis transform. The potential valleys are found by following the force field, thus, avoiding 2D search. The potential field is computed in closed form using equations of the border segments. The simple Newtonian potential is shown to be inadequate for this purpose. A higher order potential is defined which decays faster with distance than the inverse of distance. It is shown that as the potential order becomes arbitrarily large, the axes approach those computed using the shortest distance to the border. Algorithms are given for the computation of axes, which can run in linear parallel time for part of the axes having initial guesses. Experimental results are presented for a number of examples.
Narendra Ahuja, Jen-Hui Chuang
IEEE Trans. Pattern Anal. Mach. Intell.1
1997 Transitory Image Sequences, Asymptotic Properties, and Estimation of Motion and Structure
abstract
A transitory image sequence is one in which no scene element is visible through the entire sequence. This article deals with some major theoretical and algorithmic issues associated with the task of estimating structure and motion from transitory image sequences. It is shown that integration with a transitory sequence has properties that are very different from those with a nontransitory one. Two representations, world-centered (WC) and camera-centered (CC), behave very differently with a transitory sequence. The asymptotic error rates derived in this article indicate that one representation is significantly superior to the other, depending on whether one needs camera-centered or world-centered estimates. We introduce an efficient "cross-frame" estimation technique for the CC representation. For the WC representation, our analysis indicates that a good technique should be based on camera global pose instead of interframe motions. Rigorous experiments were conducted with real-image sequences taken by a fully calibrated camera system.
Juyang Weng, Yuntao Cui, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.3
1997 Multiscale image segmentation by integrated edge and region detection
abstract
This paper is concerned with the detection of low-level structure in images. It describes an algorithm for image segmentation at multiple scales. The detected regions are homogeneous and surrounded by closed edge contours. Previous approaches to multiscale segmentation represent an image at different scales using a scale-space. However, structure is only represented implicitly in this representation, structures at coarser scales are inherently smoothed, and the problem of structure extraction is unaddressed. This paper argues that the issues of scale selection and structure detection cannot be treated separately. A new concept of scale is presented that represents image structures at different scales, and not the image itself. This scale is integrated into a nonlinear transform which makes structure explicit in the transformed domain. Structures that are stable (locally invariant) to changes in scale are identified as being perceptually relevant. The transform can be viewed as collecting spatially distributed evidence for edges and regions, and making it available at contour locations, thereby facilitating integrated detection of edges and regions without restrictive models of geometry or homogeneity. In this sense, it performs Gestalt analysis. All scale parameters of the transform are automatically determined, and the structure of any arbitrary geometry can be identified without any smoothing, even at coarse scales.
Mark Tabb, Narendra Ahuja
IEEE Trans. Image Process.2
1996 Panoramic Image Acquisition
abstract
This paper is concerned with acquiring panoramic focused images using a small field of view video camera. When scene points are distributed over a range of distances from the sensor, obtaining a focused composite image involves focus computations and mechanically changing some sensor parameters (translation of sensor plane, panning of camera etc.) which can be time intensive. In this paper we present methods to optimize the image acquisition strategy in order to reduce redundancy. We show that panning a camera about a point f (focal length) in front of the camera eliminates redundancy. The non-frontal imaging camera (NICAM) with tilted sensor plane has been previously introduced as a sensor that can acquire focused panoramic images. In this paper we also describe strategies for optimal selection of panning angle increments and sensor plane tilt for NICAM. Experimental results are presented for panoramic image acquisition using a regular camera as well as using NICAM.
Arun Krishnan, Narendra Ahuja
CVPR2
1996 Shape from Appearance: A Statistical Approach to Surface Shape Estimation
Darrell R. Hougen, Narendra Ahuja
ECCV (1)2
1996 Segmentation based reversible image compression
abstract
Reversible compression of images has been the topic of considerable research as it finds applications in many fields in which the deviation of the reproduced image from the original image is intolerable, however small be the deviation. This paper is concerned with the problem of reducing spatial redundancies in gray scale images, thus providing effective lossless compression, using segmentation information. We will present new edge models that deal effectively with two issues that make such models normally unsuitable for compression applications: local applicability and large number of parameters needed for representation. Segmentation information is provided by a recent transform (1993), which we found to possess qualities making it especially suitable for compression. The final residual image is obtained using autocorrelation-based 2-D linear prediction. Different implementations providing lossless compression are presented along with results over a number of common test images. Results show that the proposed approach can be used to yield robust lossless compression, while providing consistently and significantly better results than the best possible JPEG lossless coder.
Krishna Ratakonda, Narendra Ahuja
ICIP (1)2
1996 Uniformity and homogeneity-based hierarchical clustering
abstract
This paper presents a clustering algorithm for dot patterns in n-dimensional space. The n-dimensional space often represents a multivariate (n/sub f/-dimensional) function in a n/sub s/-dimensional space (n/sub s/+n/sub f/=n). The proposed algorithm decomposes the clustering problem into the two lower dimensional problems. Clustering in n/sub f/-dimensional space is performed to detect the sets of dots in n-dimensional space having similar n/sub f/-variate function values (location based clustering using a homogeneity model). Clustering in n/sub s/ dimensional space is performed to detect the sets of dots in n-dimensional space having similar interneighbor distances (density based clustering with a uniformity model). Clusters in the n-dimensional space are obtained by combining the results in the two subspaces.
Peter Bajcsy, Narendra Ahuja
ICPR2
1996 Segmentation of volume images using a multiscale transform
abstract
This paper presents a new method for multiscale segmentation of volume images. The segmentation is achieved using a recent nonlinear transform which leads to well-characterized regions at different spatial and intensity scales. The detected three-dimensional regions are closed and are homogeneous relative to their surround. A pyramid is generated containing the region information extracted across a range of homogeneity scales. The pyramid represents the multiscale volumetric structure. Experimental results are given for magnetic resonance data as well as video sequences.
Tod Courtney, Narendra Ahuja
ICPR2
1996 Elliptical Gaussian filters
abstract
A multiscale region detector for low-level image analysis is described. The basis of the detector is a set of filters similar to the Laplacian of an elliptical Gaussian. The responses of these filters to ideal ellipses are derived, and equations for determining the parameters of detected ellipses from the filter responses are found. The use of scale-space techniques to eliminate false ellipse sites in real images is described.
Scott A. Jackson, Narendra Ahuja
ICPR2
1996 Active Surface Estimation: Integrating Coarse-to-Fine Image Acquisition and Estimation from Multiple Cues
Subhodev Das, Narendra Ahuja
Artif. Intell.2
1996 Range estimation from focus using a non-frontal imaging camera
Arun Krishnan, Narendra Ahuja
Int. J. Comput. Vis.2
1996 A Transform for Multiscale Image Segmentation by Integrated Edge and Region Detection
abstract
Describes a transform to extract image regions at all geometric and photometric scales. It is argued that linear approaches have the shortcoming that they require a priori models of region shape. The proposed transform avoids this by letting the structure emerge, bottom-up, from interactions among pixels. The transform involves global computations on pairs of pixels followed by vector integration of the results. An attraction force field is computed over the image in which pixels belonging to the same region are mutually attracted and the region is characterized by a convergent flow. It is shown that the transform possesses properties that allow multiscale segmentation, or extraction of original, unblurred structure at all different geometric and photometric scales present. This is in contrast with much previous work wherein multiscale structure is viewed as the smoothed structure in a multiscale signal decimation. Scale is an integral parameter of the force computation, and the number and values of scale parameters associated with the image can be estimated automatically. Regions are detected at all a priori unknown scales resulting in automatic construction of a segmentation tree, in which each pixel is annotated with descriptions of all the regions it belongs to. Transform properties are presented for piecewise-constant images but hold for more general ones. Thus the proposed method is intended as a solution to the problem of multiscale, integrated edge and region detection, or low-level image segmentation. Experimental results are given.
Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 Parallel distributed detection of feature trajectories in multiple discontinuous motion image sequences
abstract
Concerns the 3D interpretation of image sequences showing multiple objects in motion. Each object exhibits smooth motion except at certain time instants when a motion discontinuity may occur. The objects are assumed to contain point features which are detected as the images are acquired. Estimating feature trajectories in the first two frames amounts to feature matching. As more images are acquired, existing trajectories are extended. Both initial detection and extension of trajectories are done by enforcing pertinent constraints from among the following: similarity of the image plane arrangement of neighboring features, smoothness of the 3D motion and smoothness of the image plane motion. The constraints are incorporated into energy functions which are minimized using 2D Hopfield networks. Wrong matches that result from convergence to local minima are eliminated using a 1D Hopfield-like network. Experimental results on several image sequences are shown.
Srikanth Thirumalai, Narendra Ahuja
IEEE Trans. Neural Networks2
1995 Pixel Matching and Motion Segementation in Image Sequences
Narendra Ahuja, Ram Charan
ACCV1
1995 Structure and Motion Estimation from Dynamic Silhouettes under Perspective Projection
abstract
Addresses the problem of estimating the structure and motion of a smooth curved object from its silhouettes observed over time by a trinocular stereo rig under perspective projection. We first construct a model for the local structure along the silhouette for each frame in the temporal sequence. Successive local models are then integrated into a global surface description by estimating the motion between successive time instants. The algorithm tracks certain surface features (parabolic points) and image features (silhouette inflections and frontier points) which are used to bootstrap the motion estimation process. The entire silhouette along with the reconstructed local structure are then used to refine the initial motion estimate. We have implemented the proposed approach and report results on real images.>
Tanuja Joshi, Narendra Ahuja, Jean Ponce
ICCV2
1995 Introduction to the Special Volume on Computer Vision
Narendra Ahuja, Radu Horaud
Artif. Intell.1
1995 Integrated Matching and Segmentation of Multiple Features in Two Views
Sanghoon Sull, Narendra Ahuja
Comput. Vis. Image Underst.2
1995 Performance Analysis of Stereo, Vergence, and Focus as Depth Cues for Active Vision
abstract
This paper compares the performances of the binocular cues of stereo and vergence, and the monocular cue of focus for range estimation using an active vision system. The performance of each cue is characterized in terms of sensitivity to errors in the imaging parameters. The effects of random, quantization errors are expressed in terms of the standard deviation of the resulting depth error. The effect of systematic, calibration errors on estimation using each cue is also studied. Performance characterization of each cue is utilized to evaluate the relative performance of the cues. Also discussed, based on such characterization, are ways to select a cue taking into account the computational and reliability aspects of the corresponding estimation process.
Subhodev Das, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 Necessary and sufficient conditions for a unique solution of plane motion and structure
abstract
This paper presents necessary and sufficient conditions for uniquely determining the motion and structure of a planar surface from any number of monocular images. The authors first show an uncertain situation for the two-view plane motion problem in which an infinite number of solutions may result. Although this situation rarely occurs in real situations, it does raise visual ambiguities under certain circumstances, e.g., when some images are observed through a mirror. The understanding of this situation therefore helps analyze reflected images which are often seen in daily life. Then, the authors present necessary and sufficient conditions for a unique solution of plane motion and structure from any number of views. These algorithm-independent conditions enhance understanding about the problem of estimating plane motion and structure from image sequences and may be used to determine the uniqueness of the solution of the motion and structure of a planar surface.>
Xiaoping Hu 0004, Narendra Ahuja
IEEE Trans. Robotics Autom.2
1994 Adaptive polynomial modelling of the reflectance map for shape estimation from stereo and shading
abstract
This paper is concerned with estimation of the reflectance map and its use in surface shape estimation by a method that combines stereo and shading information. There are many advantages to the integrated approach. However, shape from shading algorithms are limited in their applicability by the assumption of idealized reflectance and lighting models involving many preset or hard to estimate parameters. In this paper, it is shown that the difficulties involved in estimation of the reflectance function and light source distribution can be eliminated through direct estimation of the reflectance map using adaptive polynomial models. The reflectance map estimate is used with a surface estimate provided by stereo in an integrated approach to surface shape estimation.>
Darrell R. Hougen, Narendra Ahuja
CVPR2
1994 Integration of transitory image sequences
abstract
A transitory image sequence is one in which no scene element is visible through the entire sequence. This article deals with some major theoretical and algorithmic issues associated with the task of estimating structure and motion from transitory image sequences. Two representations, world-centered (WC) and camera-centered (CC), behave very differently with a transitory sequence. The asymptotical error properties derived in this article indicate that one representation is significantly superior to the other, depending on whether one uses camera-centered or world-centered estimates. Rigorous experiments were conducted with real-image sequences taken by a fully calibrated camera system. The comparison demonstrated that a good accuracy can be obtained from transitory image sequences.>
Juyang Weng, Yuntao Cui, Narendra Ahuja, Ajit Singh
CVPR3
1994 Integrated 3D Analysis of Flight Image Sequences
Sanghoon Sull, Narendra Ahuja
ECCV (1)2
1994 Estimation and Segmentation of Displacement Field Using Multiple Features
abstract
We present an approach for estimating and segmenting the displacement field (DF) between two frames. Our method is based on local affine (first-order) approximation of the displacement field which is derived under the assumption of locally rigid motion. Each distinct motion is represented in the image plane by a distinct set of values of affine coefficients. All sets of values supported by the feature locations in two frames are identified by exhaustive coarse-to-fine search. The integrated use of multiple features (points, regions and lines) increases the probability of finding well-supported sets. The sets of coefficients thus obtained are used to describe the DF. Two experimental results with real images are presented to demonstrate the feasibility of our approach.>
Sanghoon Sull, Narendra Ahuja
ICIP (3)2
1994 Efficient collision detection among objects in arbitrary motion using multiple shape representations
abstract
We propose an efficient method for detecting potential collisions among multiple objects with arbitrary motion (translation and rotation) in 3D space. The method is useful for online monitoring and path planning in a 3D environment in which there are multiple independently-moving objects. The method consists of two main stages: 1) the coarse stage, an approximate test is performed to identify interfering objects in the entire workspace using octree representation of object shapes; and 2) the fine stage, polyhedral representation of object shapes is used to more accurately identify any object parts that might cause interference and collisions. For this purpose, specific pairs of faces belonging to any of the interfering objects found in the first stage are tested, thus performing detailed computation on a reduced amount of data. Experimental results, which demonstrate the efficiency of the proposed collision detection method, are given.
Yoshifumi Kitamura, Haruo Takemura, Narendra Ahuja, Fumio Kishino
ICPR (1)3
1994 A multiscale region-based approach to image matching
abstract
This paper presents a new technique for the estimation of 2D motion fields from image sequences. Homogeneous regions are identified and matched using a multiscale segmentation algorithm. Many matching difficulties result from structural changes in an image which occur only at certain scales and, hence, a multiscale approach results in a denser set of matched regions. Pixel correspondences are then obtained by estimating an affine transformation between matched regions. A motion field is then calculated from these correspondences. Areas where occlusion is present are identified and the effects of the occlusion on the affine parameter estimation process are compensated, resulting in a more accurate estimated motion field.
Mark Tabb, Narendra Ahuja
ICPR (1)2
1994 Mirror uncertainty and uniqueness conditions for determining shape and motion from orthographic projection
Xiaoping Hu 0004, Narendra Ahuja
Int. J. Comput. Vis.2
1994 Feature Extraction and Matching as Signal Detection
abstract
This paper discusses detection and matching of arbitrary image features or patterns. The common characteristics of feature extraction and matching are summarized which show that they can be considered as special cases of a more general problem—signal detection. However, the existing signal detection theories do not solve feature extraction and matching problems readily. Therefore, a general formulation of feature extraction and matching as a problem of signal detection is presented. This formulation unifies feature extraction and matching into a more general framework so that the two can be better integrated to form an automatic system for image matching or object recognition. Following this formulation, guidelines for designing algorithms for detection or matching of arbitrary image features or patterns which can be easily implemented or reconfigurated for many practical applications are derived. Sample algorithms resulting from this formulation and the associated experimental results with real image data are provided which demonstrate the performance and robustness of the methods.
Xiaoping Hu 0004, Narendra Ahuja
Int. J. Pattern Recognit. Artif. Intell.2
1994 Matching Point Features with Ordered Geometric, Rigidity, and Disparity Constraints
abstract
This correspondence presents a matching algorithm for obtaining feature point correspondences across images containing rigid objects undergoing different motions. First point features are detected using newly developed feature detectors. Then a variety of constraints are applied starting with simplest and following with more informed ones. First, an intensity-based matching algorithm is applied to the feature points to obtain unique point correspondences. This is followed by the application of a sequence of newly developed heuristic tests involving geometry, rigidity, and disparity. The geometric tests match two-dimensional geometrical relationships among the feature points, the rigidity test enforces the three dimensional rigidity of the object, and the disparity test ensures that no matched feature point in an image could be rematched with another feature, if reassigned another disparity value associated with another matched pair or an assumed match on the epipolar line. The computational complexity is proportional to the numbers of detected feature points in the two images. Experimental results with indoor and outdoor images are presented, which show that the algorithm yields only correct matches for scenes containing rigid objects.>
Xiaoping Hu 0004, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.2
1994 Integrated 3-D Analysis and Analysis-Guided Synthesis of Flight Image Sequences
abstract
This paper is concerned with three-dimensional (3D) analysis, and analysis-guided syntheses, of images showing 3-D motion of an observer relative to a scene. There are two objectives of the paper. First, it presents an approach to recovering 3D motion and structure parameters from multiple cues present in a monocular image sequence, such as point features, optical flow, regions, lines, texture gradient, and vanishing line. Second, it introduces the notion that the cues that contribute the most to 3-D interpretation are also the ones that would yield the most realistic synthesis, thus suggesting an approach to analysis guided 3-D representation. For concreteness, the paper focuses on flight image sequences of a planar, textured surface. The integration of information in these diverse cues is carried out using optimization. For reliable estimation, a sequential batch method is used to compute motion and structure. Synthesis is done by using (i) image attributes extracted from the image sequence, and (ii) simple, artificial image attributes which are not present in the original images. For display, real and/or artificial attributes are shown as a monocular or a binocular sequence. Performance evaluation is done through experiments with one synthetic sequence, and two real image sequences digitized from a commercially available video tape and a laserdisc. The attribute based representation of these sequences compressed their sizes by 502 and 367. The visualization sequence appears very similar to the original sequence in informal, monocular as well as stereo viewing on a workstation monitor.>
Sanghoon Sull, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.2
1993 Range Estimation From Focus Using a Non-frontal Imaging Camera
Arun Krishnan, Narendra Ahuja
AAAI2
1993 Necessary and Sufficient Conditions for a Unique Solution of Plane Motion and Structure
Xiaoping Hu 0004, Narendra Ahuja
CAIP2
1993 A transform for detection of multiscale image structure
abstract
A new transform is introduced to facilitate integrated edge and region detection in an image at all geometric and photometric scales at which structure is present. The proposed transform allows the structure to emerge from interactions among the pixels. A pixel's interaction with all other pixels is considered instead of testing specific neighborhoods for a specific structure. Computations are performed on pairs of pixels followed by vector integration of the results, rather than scalar, weighted averaging over pixel neighborhoods. For pixels on either side of a region boundary, the transformation yields high affinities, but there is little affinity between pixels across regions.>
Narendra Ahuja
CVPR1
1993 A comparative study of stereo, vergence, and focus as depth cues for active vision
abstract
The performances of the binocular cues of stereo, vergence, and the monocular cue of focus for range estimation using an active vision system are compared. The performance of each cue is characterized by its sensitivity to errors in the imaging parameters. The effect of random quantization errors is expressed in terms of the standard derivation of the resulting depth error. The effect of systematic calibration errors on estimation using each cue is studied. Performance characterization of each cue is shown to be useful for active control of the imaging parameters to improve the accuracy of the estimated range. Methods to integrate the use of the cues in order to overcome their individual limitations are discussed.>
Subhodev Das, Narendra Ahuja
CVPR2
1993 Estimation of the light source distribution and its use in integrated shape recovery from stereo and shading
abstract
The authors deal with estimation of the light source distribution and its use in surface shape estimation by the methods of stereo and shading. There are many advantages to integrating stereo and shading information. However, shape from shading algorithms are limited in their applicability by the assumption of overly simplistic models of the light source distribution. A more complete representation is described in which the lighting model makes use of multiple fixed point sources located at infinity. Methods of estimating the model parameters are developed, and a method of estimating the surface shape given the albedo and source distribution is presented. The shading algorithm is combined with a stereo algorithm in an integrated approach that is designed to handle albedo variations through the use of color images.>
Darrell R. Hougen, Narendra Ahuja
ICCV2
1993 Learning recognition and segmentation of 3-D objects from 2-D images
abstract
A framework called Cresceptron is introduced for automatic algorithm design through learning of concepts and rules, thus deviating from the traditional mode in which humans specify the rules constituting a vision algorithm. With the Cresceptron, humans as designers need only to provide a good structure for learning, but they are relieved of most design details. The Cresceptron has been tested on the task of visual recognition by recognizing 3-D general objects from 2-D photographic images of natural scenes and segmenting the recognized objects from the cluttered image background. The Cresceptron uses a hierarchical structure to grow networks automatically, adaptively, and incrementally through learning. The Cresceptron makes it possible to generalize training exemplars to other perceptually equivalent items. Experiments with a variety of real-world images are reported to demonstrate the feasibility of learning in the Cresceptron.>
Juyang Weng, Narendra Ahuja, Thomas S. Huang
ICCV2
1993 Motion and structure estimation using long sequence motion models
Xiaoping Hu 0004, Narendra Ahuja
Image Vis. Comput.2
1993 Active Stereo: Integrating Disparity, Vergence, Focus, Aperture and Calibration for Surface Estimation
abstract
An approach to integrating stereo disparity, camera vergence, and lens focus to exploit their complementary strengths and weaknesses through active control of camera focus and orientations is presented. In addition, the aperture and zoom settings of the cameras are controlled. The result is an active vision system that dynamically and cooperatively interleaves image acquisition with surface estimation. A dense composite map of a single contiguous surface is synthesized by automatically scanning the surface and combining estimates of adjacent, local surface patches. This problem is formulated as one of minimizing a pair of objective functions. The first such function is concerned with the selection of a target for fixation. The second objective function guides the surface estimation process in the vicinity of the fixation point. Calibration parameters of the cameras are treated as variables during optimization, thus making camera calibration an integral, flexible component of surface estimation. An implementation of this method is described, and a performance evaluation of the system is presented. An average absolute error of less than 0.15% in estimated depth was achieved for a large surface having a depth of approximately 2 m.>
Narendra Ahuja, A. Lynn Abbott
IEEE Trans. Pattern Anal. Mach. Intell.1
1993 Optimal Motion and Structure Estimation
abstract
The causes of existing linear algorithms exhibiting various high sensitivities to noise are analyzed. It is shown that even a small pixel-level perturbation may override the epipolar information that is essential for the linear algorithms to distinguish different motions. This analysis indicates the need for optimal estimation in the presence of noise. Methods are introduced for optimal motion and structure estimation under two situations of noise distribution: known and unknown. Computationally, the optimal estimation amounts to minimizing a nonlinear function. For the correct convergence of this nonlinear minimization, a two-step approach is used. The first step is using a linear algorithm to give a preliminary estimate for the parameters. The second step is minimizing the optimal objective function starting from that preliminary estimate as an initial guess. A remarkable accuracy improvement has been achieved by this two-step approach over using the linear algorithm alone.>
Juyang Weng, Narendra Ahuja, Thomas S. Huang
IEEE Trans. Pattern Anal. Mach. Intell.2
1993 NETRA: A Hierarchical and Partitionable Architecture for Computer Vision Systems
abstract
Computer vision is regarded as one of the most complex and computationally intensive problems. In general, a Computer Vision System (CVS) attempts to relate scene(s) in terms of model(s). A typical CVS employs algorithms from a very broad spectrum such as numerical, image processing, graph algorithms, symbolic processing, and artificial intelligence. The authors present a multiprocessor architecture, called "NETRA," for computer vision systems. NETRA is a highly flexible architecture. The topology of NETRA is recursively defined, and hence, is easily scalable from small to large systems. It is a hierarchical architecture with a tree-type control hierarchy. Its leaf nodes consists of a cluster of processors connected with a programmable crossbar with selective broadcast capability to provide the desired flexibility. The processors in clusters can operate in SIMD-, MIMD- or Systolic-like modes. Other features of the architecture include integration of limited data-driven computation within a primarily control flow mechanism, block-level control and data flow, decentralization of memory management functions, and hierarchical load balancing and scheduling capabilities. The paper also presents a qualitative evaluation and preliminary performance results of a cluster of NETRA.>
Alok N. Choudhary, Janak H. Patel, Narendra Ahuja
IEEE Trans. Parallel Distributed Syst.3
1992 Motion and Structure Factorization and Segmentation of Long Multiple Motion Image Sequences
Christian Debrunner, Narendra Ahuja
ECCV2
1992 Estimating motion of constant acceleration from image sequences
abstract
Presents a model-based algorithm for estimating motion from monocular image sequences. The authors first present a two-view motion algorithm and then extend it to multiple views. The two-view algorithm requires generally 6 pairs of point correspondences to give unique solution of the motion parameters. However, when the used points lie on a Maybank quadric, the algorithm requires 7 pairs of point correspondences to give double solutions. Object-centered motion representations and a motion model of constant acceleration are used to estimate motion parameters from long image sequences. The algorithm guarantees globally optimal solution. Since the algorithm does not involve structure parameters, it contains the least number of unknowns and is hence more efficient and robust than the existing ones. Experimental results with real image data are presented. The same method can be applied to solve for motions described by second or higher orders of polynomials.>
Xiaoping Hu 0004, Narendra Ahuja
ICPR (1)2
1992 Supervised classification of early perceptual structure in dot patterns
abstract
A supervised algorithm for computing perceptual groupings in dot patterns is presented. The algorithm uses shape features of the polygons in the Voronoi tessellation of the input pattern. The training patterns identified by humans are used to obtain an initial nocontextual classification which is then refined by a probabilistic relaxation labeling.>
Mihran Tüceryan, Anil K. Jain 0001, Narendra Ahuja
ICPR (2)3
1992 Why aspect graphs are not (yet) practical for computer vision
Olivier D. Faugeras, Joseph L. Mundy, Narendra Ahuja, Charles R. Dyer, Alex Pentland, Ramesh Jain 0001, Katsushi Ikeuchi, Kevin W. Bowyer
CVGIP Image Underst.3
1992 Matching Two Perspective Views
abstract
A computational approach to image matching is described. It uses multiple attributes associated with each image point to yield a generally overdetermined system of constraints, taking into account possible structural discontinuities and occlusions. In the algorithm implemented, intensity, edgeness, and cornerness attributes are used in conjunction with the constraints arising from intraregional smoothness, field continuity and discontinuity, and occlusions to compute dense displacement fields and occlusion maps along the pixel grids. The intensity, edgeness, and cornerness are invariant under rigid motion in the image plane. In order to cope with large disparities, a multiresolution multigrid structure is employed. Coarser level edgeness and cornerness measures are obtained by blurring the finer level measures. The algorithm has been tested on real-world scenes with depth discontinuities and occlusions. A special case of two-view matching is stereo matching, where the motion between two images is known. The algorithm can be easily specialized to perform stereo matching using the epipolar constraint.>
Juyang Weng, Narendra Ahuja, Thomas S. Huang
IEEE Trans. Pattern Anal. Mach. Intell.2
1992 Motion and Structure from Line Correspondences; Closed-Form Solution, Uniqueness, and Optimization
abstract
This work discusses estimating motion and structure parameters from line correspondences of a rigid scene. The authors present a closed-form solution to motion and structure parameters from line correspondences through three monocular perspective views. The algorithm makes use of redundancy in the data to improve the accuracy of the solutions. The uniqueness of the solution is established, and necessary and sufficient conditions for degenerate spatial line configurations are given. Optimization has been employed to further improve the accuracy of the estimates in the presence of noise. Simulations have shown that the errors of the optimized estimates are close to the theoretical lower error bound.>
Juyang Weng, Thomas S. Huang, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.3
1992 A potential field approach to path planning
abstract
A path-planning algorithm for the classical mover's problem in three dimensions using a potential field representation of obstacles is presented. A potential function similar to the electrostatic potential is assigned to each obstacle, and the topological structure of the free space is derived in the form of minimum potential valleys. Path planning is done at two levels. First, a global planner selects a robot's path from the minimum potential valleys and its orientations along the path that minimize a heuristic estimate of the path length and the chance of collision. Then, a local planner modifies the path and orientations to derive the final collision-free path and orientations. If the local planner fails, a new path and orientations are selected by the global planner and subsequently examined by the local planner. This process is continued until a solution is found or there are no paths left to be examined. The algorithm solves a much wider class of problems than other heuristic algorithms and at the same time runs much faster than exact algorithms (typically 5 to 30 min on a Sun 3/260).>
Yong Koo Hwang, Narendra Ahuja
IEEE Trans. Robotics Autom.2
1991 Estimation of motion and structure of planar surfaces from a sequence of monocular images
abstract
An algorithm is presented which estimates 10 parameters for motion and structure of a rigid planar patch given point correspondences in a monocular image sequence under perspective projection. The rotational velocity is assumed constant and the rotation center arbitrary. The algorithm mainly consists of two steps. First, the 3-D space of ( omega /sub x/, omega /sub y/, omega /sub z/) is searched exhaustively and for each ( omega /sub x/, omega /sub y/, omega /sub z/) all the other parameters with the value of an objective function are computed linearly. some of the ( omega /sub x/, omega /sub y/, omega /sub z/) and the corresponding structure values are used as the initial guesses in the second step. The objective function is iteratively minimized with respect to five variables for rotation and structure. The solution corresponding to the global minimum is used to obtain least squares estimates of the remaining unknowns, for translation and rotation center. It was experimentally found that the objective function converges well so that the 3-D space need not be densely searched. Results are presented for three image sequences, two simulated and one real.>
Sanghoon Sull, Narendra Ahuja
CVPR2
1991 Sufficient conditions for double or unique solution of motion and structure
abstract
Several sufficient conditions are presented for a double or unique solution of the problem of motion and structure from two monocular images. It is shown that: five correspondences of points that do not lie on two lines in the image plane suffice to determine a pure rotation uniquely; six correspondences of points that do not lie on two lines in the image plane and do not correspond to space points lying on a specific quadric surface suffice to determine a motion with nonzero translation uniquely; each Maybank quadric can sustain at most two physically acceptable motion solutions and surface interpretations, provided that a sufficient number of correspondences are present; in the plane motion case, six correspondences of points that do not lie on a quadric curve in the image plane will only admit the true motion and structure and their duals as solutions. Several properties of the essential matrix T*R and the plane motion matrix R+TN/sup T/, both of which are frequently used in the motion and structure estimation problem, are listed.>
Xiaoping Hu 0004, Narendra Ahuja
ICASSP2
1991 Path planning using the Newtonian potential
abstract
Newtonian potential function is used to represent polygonal objects and obstacles. The closed-form expression of this potential field and other gradient-related quantities are derived. Such results not only eliminate the problems associated with the discretization of the object and obstacles in evaluating the risk of collision, but also make the search for the optimal object configurations efficient. The object skeleton, a shape description of the moving object, is introduced to guide the moving object through narrow regions while the search is done at different stages. The free space can then be divided by the narrow regions where the path planning takes place-a very simple free space decomposition scheme. Successful global strategies are developed to connect the local plans into a safe and smooth global path.>
Jen-Hui Chuang, Narendra Ahuja
ICRA2
1991 Motion estimation under orthographic projection
abstract
Some new results for the problem of motion estimation under orthographic projection are presented. Some basic results obtained by previous researchers are refined and more detailed and precise results are provided. It is shown that, in the two-view problem, when the rotation is around the optical axis, the motion (but not the structure) is uniquely determined. In the three-view problem, only under certain conditions are the motion and structure uniquely determined. For any motion problem, if two-view matching cannot determine the motion, only under certain conditions can three-view or multiview matching help.>
Xiaoping Hu 0004, Narendra Ahuja
IEEE Trans. Robotics Autom.2
1990 Active surface reconstruction by integrating focus, vergence, stereo, and camera calibration
abstract
A method is described for estimating surfaces from stereo images. A single pair of stereo images can yield surface reconstruction for only a small volume of space. Since the scene may be large, images are obtained dynamically using different camera configurations through the control of focus and camera vergence. For smooth objects, this results in the reconstruction of surface patches that are contiguous and overlapping. These sequentially acquired surface maps are merged into a central. composite representation. During camera reconfiguration, unpredictable calibration errors may arise. Therefore, this method permits small changes in calibration parameter values to achieve agreement between adjacent reconstructed patches in the areas of overlap. To integrate the different sources of depth information, and to balance the effects they have on the resulting surface estimate, the surface reconstruction problem is formulated as one of optimization.>
A. Lynn Abbott, Narendra Ahuja
ICCV2
1990 Multiresolution image acquisition and surface reconstruction
abstract
Consideration is given to the problem of surface reconstruction from stereo images for large scenes having large depth ranges where it is necessary to aim cameras in different directions and to fixate at different objects. The authors concentrate on the selection of new fixation points from among the nonfixated, low resolution scene parts, and subsequent processing for surface reconstruction. The coarse stereo estimates in the vicinity of the new fixation point are refined as the images of the new fixation point gradually deblur during the process of refixation and are subsequently used to analyze the fixated parts of the scene.>
Subhodev Das, Narendra Ahuja
ICCV2
1990 A reconfigurable and hierarchical parallel processing architecture: performance results for stereo vision
abstract
A multiprocessor architecture called NETRA is discussed. It is highly reconfigurable and does not involve the use of complex interconnection schemes. The topology of this multiprocessor is recursively defined and is therefore easily scalable from small to large systems. It has a tree-type hierarchical architecture featuring leaf nodes that consist of a cluster of small but powerful processors connected via a programmable crossbar with selective broadcast capability. The architecture is simulated on a hypercube multiprocessor and the performance of one processor cluster is evaluated for stereo-vision tasks. The particular stereo algorithm selected for implementation requires computation of the two-dimensional fast Fourier transform (2-D FFT), template matching, histogram computation, and least-squares surface fitting. Static partitioning of data is used for the data-independent tasks such as 2-D FFT and dynamic scheduling, and load balancing is used for the data-dependent tasks of feature matching and disambiguation.>
Alok N. Choudhary, Subhodev Das, Narendra Ahuja, Janak H. Patel
ICPR (2)3
1990 A direct data approximation based motion estimation algorithm
abstract
A novel algorithm for estimating motion and structure from orthographic projections is presented. The estimate is based on the image locations of a set of points on a rigid object viewed in many frames. The object can have an arbitrary shape. Motion is assumed to consist of rotation about a fixed-direction axis and linear translation, each at a constant rate. The algorithm uses closed-form expressions to find estimates of the motion and the structure parameters and can accept any number of points over any number of frames as input. The performance of the algorithm on synthetic data generated to represent a range of motion and structure conditions and on data extracted from real images is evaluated. It is shown that the algorithm produces accurate results from synthetic data given the positions of 10 points over 10 frames, with a total rotation of 90 degrees and with uniformly distributed noise of 0.9%.>
Christian Debrunner, Narendra Ahuja
ICPR (1)2
1990 Estimating motion and structure from line matches: performance obtained and beyond
abstract
The performance issues of estimating motion and structure from line correspondences are studied. An approach to optimal estimation of motion and structure using line correspondences is presented. To minimize the expected errors in the estimated parameters, it is necessary to minimize the matrix-weighted discrepancy between the computed lines and the observed lines. In order to reliably reach the global minimum solution, a closed-form solution is computed and then used as the initial starting condition for an iterative optimal estimation algorithm. Simulation results show that, in the presence of noise, the accuracy of the optimal solution is not only considerably better than that of the closed-form solutions, but it has also reached a level that it is comparable with that of point-based optimal algorithms. Simulations also show that the error of the optimal solution is close to a theoretical lower error bound, the Cramer-Rao bound, which implies that there exists little room for accuracy improvement beyond the performance obtained.>
Juyang Weng, Thomas S. Huang, Narendra Ahuja
ICPR (1)3
1990 Octree Generation from Object Silhouettes in Perspective Views
Sanjay K. Srivastava, Narendra Ahuja
Comput. Vis. Graph. Image Process.2
1989 Path planning using a potential field representation
abstract
The findpath problem is the problem of moving an object to the desired position and orientation while avoiding obstacles. The authors present an approach to this problem using a potential-field representation of obstacles. A potential function similar to an electrostatic potential is assigned to each obstacle, and the topological structure of the free space is derived in the form of minimum potential valleys. A path specified by a subset of valley segments and associated object orientations, which minimizes a heuristic estimate of path length and the chance of collision, is selected as the initial guess of the solution. Then, the selected path as well as the orientation of the moving object along the path is modified to minimize a defined cost of the path. Findpath problems possessing three different levels of difficulty are identified. Path optimization is performed in up to three stages, according to the level of difficulty of the problem. These three stages are addressed by three separate algorithms which are automatically selected. The performance of the algorithms is illustrated.>
Yong Koo Hwang, Narendra Ahuja
CVPR2
1989 Optimal motion and structure estimation
abstract
The problem of estimating motion and structure of a rigid scene from two perspective monocular views is studied. The optimization approach presented is motivated by the following observations of linear algorithms: (1) for certain types of motion, even pixel-level perturbations (such as digitization noise) may override the information characterized by epipolar constraint; (2) existing linear algorithms do not use the constraints in the essential parameter matrix E in solving for this matrix. The authors present approaches to estimating errors in the optimal solutions, investigate the theoretical lower bounds on the errors in the solutions and compare them with actual errors, and analyze two types of algorithms of optimization: batch and sequential. The analysis and experiments show that, in general, a batch technique performs better than a sequential technique for any nonlinear problems. A recursive batch processing technique is proposed for nonlinear problems that require recursive estimation.>
Juyang Weng, Narendra Ahuja, Thomas S. Huang
CVPR2
1989 Extraction of early perceptual structure in dot patterns: Integrating region, boundary, and component gestalt
Narendra Ahuja, Mihran Tuceryan
Comput. Vis. Graph. Image Process.1
1989 A multiscale region detector
Dorothea Blostein, Narendra Ahuja
Comput. Vis. Graph. Image Process.2
1989 Generating Octrees from Object Silhouettes in Orthographic Views
abstract
An algorithm to construct the octree representation of a three-dimensional object from silhouette images of the object is described. The images must be obtained from thirteen viewing directions corresponding to the three face views, six edge views, and four corner views of an upright cube. These views where chosen because they provide a simple relationship between pixels in the image and the octant labels in the octree, thus replacing the computation of detecting intersections between the octree space and the objects by a table lookup operation. The average ratio of the object volume to the octree volume is found to be greater than 90%. The sequential use made of the chosen viewing directions results in a coarse-to-fine acquisition of occupancy information. The number and order of the viewpoints used provides a mechanism for trading accuracy of the representation against the computational effort needed to obtain the representation.>
Narendra Ahuja, Jack Veenstra
IEEE Trans. Pattern Anal. Mach. Intell.1
1989 Shape From Texture: Integrating Texture-Element Extraction and Surface Estimation
abstract
A method is presented for identifying texture elements while simultaneously recovering the orientation of textured surfaces. A multiscale region detector, based on measurements in a Del /sup 2/G (Laplacian-of-Gaussian) scale space, is used to construct a set of candidate texture elements. True elements are selected from the set of candidate elements by finding the planar surface that best predicts the observed areas of the latter. Results are shown for a variety of natural textures, including waves, flowers, rocks, clouds, and dirt clods.>
Dorothea Blostein, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.2
1989 Surfaces from Stereo: Integrating Feature Matching, Disparity Estimation, and Contour Detection
abstract
An approach is described that integrates the processes of feature matching, contour detection, and surface interpolation to determine the three-dimensional distance, or depth, of objects from a stereo pair of images. Integration is necessary to ensure that the detected surfaces are smooth. Surface interpolation takes into account detected occluding and ridge contours in the scene; interpolation is performed within regions enclosed by these contours. Planar and quadratic patches are used as local models of the surface. Occluded regions in the image are identified, and are not used for matching and interpolation. A coarse-to-fine algorithm is presented that generates a multiresolution hierarchy of surface maps, one at each level of resolution. Experimental results are given for a variety of stereo images.>
William A. Hoff, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.2
1989 Motion and Structure From Two Perspective Views: Algorithms, Error Analysis, and Error Estimation
abstract
Deals with estimating motion parameters and the structure of the scene from point (or feature) correspondences between two perspective views. An algorithm is presented that gives a closed-form solution for motion parameters and the structure of the scene. The algorithm utilizes redundancy in the data to obtain more reliable estimates in the presence of noise. An approach is introduced to estimating the errors in the motion parameters computed by the algorithm. Specifically, standard deviation of the error is estimated in terms of the variance of the errors in the image coordinates of the corresponding points. The estimated errors indicate the reliability of the solution as well as any degeneracy or near degeneracy that causes the failure of the motion estimation algorithm. The presented approach to error estimation applies to a wide variety of problems that involve least-squares optimization or pseudoinverse. Finally the relationships between errors and the parameters of motion and imaging system are analyzed. The results of the analysis show, among other things, that the errors are very sensitive to the translation direction and the range of field view. Simulations are conducted to demonstrate the performance of the algorithms and error estimation as well as the relationships between the errors and the parameters of motion and imaging systems. The algorithms are tested on images of real-world scenes with point of correspondences computed automatically.>
Juyang Weng, Thomas S. Huang, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.3
1988 Closed-form solution+maximum likelihood: a robust approach to motion and structure estimation
abstract
A robust approach is presented to estimation motion and structure from image sequences. The approach consists of two steps. The first step is estimating the motion parameters using a robust linear algorithm that gives a closed-form solution for motion parameters and scene structure. The second step is improving the results from the linear algorithm using maximum-likelihood estimation. An algorithm using point correspondences from monocular images is discussed in detail and experimented with. An algorithm using line correspondences is briefly discussed. The simulations show that maximum-likelihood estimation achieves remarkable improvement over the preliminary estimates given by the linear algorithm. The algorithm is also tested on images of real scenes from automatically computed displacement field. The proposed approach is independent of the exact tokens used to establish correspondences, e.g. displacement flow, optical flow, or discrete features. Two or more types of tokens may be used, for monocular or binocular images.>
Juyang Weng, Narendra Ahuja, Thomas S. Huang
CVPR2
1988 Estimating motion/structure from line correspondences: a robust linear algorithm and uniqueness theorems
abstract
A closed-form solution to motion and structure from line correspondences in monocular perspective image sequences is presented. The algorithm requires a minimum of 13 lines over three perspective views. Redundancy in the data provides overdetermination to combat noise. The estimates can be used as an initial guess for further optimization. A unique solution to motion and structure is guaranteed if and only if the line configuration is not degenerate and the translation between any two views does not vanish. Necessary and sufficient conditions for degenerate spatial line configurations have been derived. Simulations are performed which show the performance of the algorithm in the presence of noise.>
Juyang Weng, Yuncai Liu, Thomas S. Huang, Narendra Ahuja
CVPR4
1988 Surface Reconstruction By Dynamic Integration Of Focus, Camera Vergence, And Stereo
abstract
This paper concerns estimation of surface maps for real scenes having a wide field of view and a wide range of depths. Much research has emphasized stereo disparity as a source of depth information. To a lesser extent, camera focus and camera vergence have also been investigated for their utility in depth recovery. We argue that these sources of visual information have mutually complementary strengths and weaknesses, and to obtain surface maps for real scenes these processes must be integrated. Such integration requires active control of camera orientations and imaging parameters to dynamically and cooperatively interleave image acquisition with surface estimation. Accordingly, a global surface map of the visual field is synthesized by systematically scanning the scene, and combining estimates of adjacent, local surface patches, each acquired by an intermediate camera configuration and having a small depth range. We present an algorithm to perform this integration, and describe its implementation on a dynamic stereo-camera imaging system. Experimental results are presented to demonstrate the superior performance of the integrated system over that of each of its components.
A. Lynn Abbott, Narendra Ahuja
ICCV2
1988 Two-view Matching
abstract
Establishing correspondences between images of the same scene is one of the most challenging and critical stcps in motion and scene analysis. Part of the difficulty is due to a wide variety of three-dimension structural discontinuities and occlusions that occur in real world scenes. This paper describes a computational approach to image matching that uses multiple attributes associated with a pixel to yield a generally overdetermined system of constraints, taking into account possible structural discontinuities and occlusions. In the algorithm implemented, intensity, edgeness, and comemess attributes are used in conjunction with the constraints arising from intraregional smoothness, field continuity and discontinuity, and occlusions to compute dense displacement fields and occlusion maps at pixel grids. A multiresolution multigrid structure is employed to deal with large disparities. Coarser level attributes are obtained by blurring the finer level attributes. The algorithms are tested on real world scenes containing depth discontinuities and occlusions. A special case of two-view matching is stereo matching where the motion between two images is known. The general algorithm given here can be easily spccialized to perfonn stereo matching using epipolar line constraint.
Juyang Weng, Narendra Ahuja, Thomas S. Huang
ICCV2
1988 Motion and structure from point correspondences: a robust algorithm for planar case with error estimation
abstract
The problem of determining motion and structure for a planar surface and the error estimation are discussed. Since the motion of a planar patch is a degenerate case for linear algorithms (algorithms that consist of solving mainly linear equations and give a closed-form solution) for general surfaces, the motion of such a planar surface is considered separately. An algorithm is introduced that gives a closed-form solution to motion parameters using monocular perspective images of the points on a planar surface. The algorithm is simpler and more reliable, in the presence of noise, than existing ones. There are generally two solutions for two image frames. For three image frames the solution is generally unique. An approach is proposed to test whether the points are coplanar. The errors in the motion parameters and surface structure can be estimated for each pair of images. Specifically, the standard deviation of the errors is calculated in terms of the variance of the errors in the image coordinates. This approach to estimating errors is applicable to least-squares, pseudo-inverse and eigenvalue-eigenvector problems.>
Juyang Weng, Narendra Ahuja, Thomas S. Huang
ICPR2
1988 Path planning using a potential field representation
abstract
An approach to two-dimensional as well as three-dimensional findpath problems that divides each problem into two steps is presented. Rough paths are found based only on topological information. This is accomplished by assigning to each obstacle an artificial potential similar to electrostatic potential to prevent the moving object from colliding with the obstacles, and then locating minimum-potential valleys. The paths defined by the minimum-potential valleys are modified to obtain an optimal collision-free path and orientations of the moving object along the path. Three algorithms are given to accomplish this second step. These three algorithms based on potential fields are nearly complete in scope, and solve a large variety of problems.>
Yong Koo Hwang, Narendra Ahuja
ICRA2
1988 A simplified linear optic flow-motion algorithm
Xinhua Zhuang, Thomas S. Huang, Narendra Ahuja, Robert M. Haralick
Comput. Vis. Graph. Image Process.3
1988 Line drawings of octree-represented objects
abstract
The octree structure represents the space occupied by an object as a juxtaposition of cubes, where the sizes and position coordinates of the cubes are integer powers of 2 and are defined by a recursive decomposition of three-dimensional space. This makes the octree structure highly sensitive to object location and orientation, and the three-dimensional shape of the represented object obscure. It is helpful to be able to see the actual object represented by an octree, especially for visual performance evaluation of octree algorithms. Presented in this paper is a display algorithm that helps visualize the three-dimensional space represented by the octree. Given an octree, the algorithm produces a line drawing of the objects represented by the octree, using parallel projection, from any specified viewpoint with hidden lines removed. The order in which the algorithm traverses the octree has the property that if node x occludes node y , then node x is visited before node y . The algorithm produces a set of long, straight visible edge segments corresponding to the visible surface of the polyhedral object represented by the octree. Examples of some line drawing produced by the algorithm are given. The complexity of the algorithm is also discussed.
Jack Veenstra, Narendra Ahuja
ACM Trans. Graph.2
1987 3-D Motion Estimation, Understanding, and Prediction from Noisy Image Sequences
abstract
This paper presents an approach to understanding general 3-D motion of a rigid body from image sequences. Based on dynamics, a locally constant angular momentum (LCAM) model is introduced. The model is local in the sense that it is applied to a limited number of image frames at a time. Specifically, the model constrains the motion, over a local frame subsequence, to be a superposition of precession and translation. Thus, the instantaneous rotation axis of the object is allowed to change through the subsequence. The trajectory of the rotation center is approximated by a vector polynomial. The parameters of the model evolve in time so that they can adapt to long term changes in motion characteristics. The nature and parameters of short term motion can be estimated continuously with the goal of understanding motion through the image sequence. The estimation algorithm presented in this paper is linear, i.e., the algorithm consists of solving simultaneous linear equations. Based on the assumption that the motion is smooth, object positions and motion in the near future can be predicted, and short missing subsequences can be recovered. Noise smoothing is achieved by overdetermination and a leastsquares criterion. The framework is flexible in the sense that it allows both overdetermination in number of feature points and the number of image frames.
Juyang Weng, Thomas S. Huang, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.3
1986 Efficient planar embedding of trees for VLSI layouts
Narendra Ahuja
Comput. Vis. Graph. Image Process.1
1985 Octree generation from silhouette views of an object
abstract
Octrees are used in many 3-D representation problems because they provide a compact data structure, allow rapid access to information, and implement efficient data manipulation algorithms. The initial acquisition of the 3-D information, however, is a common problem. This paper describes an algorithm to construct the octree representation of a 3-D object from silhouette images of the object. The images must be obtained from nine viewing directions corresponding to the three "face-on" and six "edge-on" views of an upright cube. The execution time is found to be linear in the number of nodes in the octree.
Jack Veenstra, Narendra Ahuja
ICRA2
1985 Image representation using Voronoi tessellation
Narendra Ahuja, Byong An, Bruce J. Schachter
Comput. Vis. Graph. Image Process.1
1984 Octree representations of moving objects
Narendra Ahuja, Charles Nash
Comput. Vis. Graph. Image Process.1
1984 Multiprocessor Pyramid Architectures for Bottom-Up Image Analysis
abstract
This paper describes three hierarchical organizations of small processors for bottom-up image analysis:pyramids, interleaved pyramids, and pyramid trees. Progressively lower levels in the hierarchies process image windows of decreasing size. Bottom-up analysis is made feasible by transmitting up the levels quadrant borders and border-related information that captures quadrant interaction of interest for a given computation. The operation of the pyramid is illustrated by examples of standard algorithms for interior-based computations (e.g., area) and border-based computations of local properties (e.g., perimeter). A connected component counting algorithm is outlined that illustrates the role of border-related information in representing quadrant interaction. Interleaved pyramids are obtained by sharing processors among several pyramids. They increase processor utilization and throughput rate at the cost of increased hardware. Trees of shallow interleaved pyramids, calld pyramid trees, are introduced to reduce the hardware requirements of large interleaved pyramids at the expense of increased processing time, without sacrificing processor utilization. The three organizations are compared with respect to several performance measures.
Narendra Ahuja, Sowmitri Swamy
IEEE Trans. Pattern Anal. Mach. Intell.1
1983 On approaches to polygonal decomposition for hierarchical image representation
Narendra Ahuja
Comput. Vis. Graph. Image Process.1
1983 Octree representations of moving objects
Charles Nash, Narendra Ahuja
Comput. Vis. Graph. Image Process.2
1982 Dot Pattern Processing Using Voronoi Neighborhoods
abstract
A sound notion of the neighborhood of a point is essential for analyzing dot patterns. The past work in this direction has concentrated on identifying pairs of points that are neighbors. Examples of such methods include those based on a fixed radius, k-nearest neighbors, minimal spanning tree, relative neighborhood graph, and the Gabriel graph. This correspondence considers the use of the region enclosed by a point's Voronoi polygon as its neighborhood. It is argued that the Voronoi polygons possess intuitively appealing characteristics, as would be expected from the neighborhood of a point. Geometrical characteristics of the Voronoi neighborhood are used as features in dot pattern processing. Procedures for segmentation, matching, and perceptual border extraction using the Voronoi neighborhood are outlined. Extensions of the Voronoi definition to other domains are discussed.
Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.1
1981 Mosaic models for images - III. Spatial correlation in mosaics
Narendra Ahuja
Inf. Sci.1
1981 Mosaic models for images - I. Geometric properties of components in cell-structure mosaics
Narendra Ahuja
Inf. Sci.1
1981 Mosaic models for images--II. Geometric properties of components in coverage mosaics
Narendra Ahuja
Inf. Sci.1
1981 Mosaic Models for Textures
abstract
This paper deals with a class of image models based on random geometric processes. Theoretical and empirical results on properties of patterns generated using these models are summarized. These properties can be used as aids in fitting the models to images.
Narendra Ahuja, Azriel Rosenfeld
IEEE Trans. Pattern Anal. Mach. Intell.1
1980 Interference Detection and Collision Avoidance Among Three Dimensional Objects
Narendra Ahuja, Robert T. Chien, R. Yen, N. Bridwell
AAAI1
1980 Neighbor gray levels as features in pixel classification
Narendra Ahuja, Azriel Rosenfeld, Robert M. Haralick
Pattern Recognit.1
1980 Some Experiments with Mosaic Models for Images
abstract
Experimental results are presented on some properties of random mosaic models for textures. These observations are compared with the theoretically predicted values. The predictions are also compared with observations on a real Image.
Narendra Ahuja, Tsvi Dubitzki, Azriel Rosenfeld
IEEE Trans. Syst. Man Cybern.1
1978 Piecewise Approximation of Pictures Using Maximal Neighborhoods
abstract
Suppose that we are given a picture having approximately piecewise constant gray leveL Each point P has a largest neighborhood N(P) that is entirely contained in one of the constant regions, and the set of maximal N(P)'s (i.e., N(P)'s not contained in other N(P)'s) constitutes an economical description of the picture, generalizing the Blum "skeleton" or medial axis transformation. This description can be used to construct approximations to the picture (e.g., by discarding small N(P)'s). The picture can be smoothed, without excessive blurring, by averaging over each N(P). By taking differences between pairs of touching maximal N(P)'s, the edges between the regions can be detected; since this edge detection scheme is not based on symmetrical detection operators, it is not handicapped when two adjacent regions differ greatly in size.
Narendra Ahuja, Larry Davis 0001, David L. Milgram, Azriel Rosenfeld
IEEE Trans. Computers1