Raghuveer M. Rao

dblp:r/RaghuveerMRao · also M. R. Raghuveer, Mysore R. Raghuveer, Raghuveer Rao · DBLP profile ↗
← Back
69ranked-venue papers
7as first author
19since 2021 · last 2025
0000-0001-6481-4175ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 14 · 13 since 2021Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
abstract
Prevalent text-to-video retrieval systems mainly adopt embedding models for feature extraction and compute cosine similarities for ranking.However, this design presents two limitations.Low-quality text-video data pairs could compromise the retrieval, yet are hard to identify and examine.Cosine similarity alone provides no explanation for the ranking results, limiting the interpretability.We ask that can we interpret the ranking results, so as to assess the retrieval models and examine the text-video data?This work proposes X-CoT, an explainable retrieval framework upon LLM CoT reasoning in place of the embedding model-based similarity ranking.We first expand the existing benchmarks with additional video annotations to support semantic understanding and reduce data bias.We also devise a retrieval CoT consisting of pairwise comparison steps, yielding detailed reasoning and complete ranking.X-CoT empirically improves the retrieval performance and produces detailed rationales.It also facilitates the model behavior and data quality analysis.Code and data are available at: github.com
Prasanna Reddy Pulakurthi, Jiamian Wang, Majid Rabbani, Sohail A. Dianat, Raghuveer M. Rao, Zhiqiang Tao
EMNLP5
2025 MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
abstract
Runjia Zeng, Guangyan Sun, Qifan Wang, Tong Geng, Sohail Dianat, Xiaotian Han, Raghuveer Rao, Xueling Zhang, Cheng Han, Lifu Huang, Dongfang Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Runjia Zeng, Guangyan Sun, Qifan Wang 0001, Tong Geng, Sohail A. Dianat, Raghuveer M. Rao, Xueling Zhang, Cheng Han 0001, Lifu Huang, Dongfang Liu
EMNLP7
2025 Structured Policy Optimization: Enhance Large Vision-Language Model via Self-Referenced Dialogue
Can Qin, Yihao Feng, Zeyuan Chen 0001, Ran Xu 0001, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Zhiqiang Tao
ICCV8
2025 Shuffle PatchMix Augmentation with Confidence-Margin Weighted Pseudo-Labels for Enhanced Source-Free Domain Adaptation
abstract
This work investigates Source-Free Domain Adaptation (SFDA), where a model adapts to a target domain without access to source data. A new augmentation technique, Shuffle PatchMix (SPM), and a novel reweighting strategy are introduced to enhance performance. SPM shuffles and blends image patches to generate diverse and challenging augmentations, while the reweighting strategy prioritizes reliable pseudo-labels to mitigate label noise. These techniques are particularly effective on smaller datasets like PACS, where overfitting and pseudo-label noise pose greater risks. State-of-the-art results are achieved on three major benchmarks: PACS, VisDA-C, and DomainNet-126. Notably, on PACS, improvements of 7.3% (79.4% to 86.7%) and 7.2% are observed in single-target and multi-target settings, respectively, while gains of 2.8% and 0.7% are attained on DomainNet-126 and VisDA-C. This combination of advanced augmentation and robust pseudo-label reweighting establishes a new benchmark for SFDA. The code is available at: https://github.com/PrasannaPulakurthi/SPM.
Prasanna Reddy Pulakurthi, Majid Rabbani, Jamison Heard, Sohail A. Dianat, Celso de Melo, Raghuveer M. Rao
ICIP6
2025 Re-Imagining Multimodal Instruction Tuning: A Representation View
abstract
Multimodal instruction tuning has proven to be an effective strategy for achieving zero-shot generalization by fine-tuning pre-trained Large Multimodal Models (LMMs) with instruction-following data. However, as the scale of LMMs continues to grow, fully fine-tuning these models has become highly parameter-intensive. Although Parameter-Efficient Fine-Tuning (PEFT) methods have been introduced to reduce the number of tunable parameters, a significant performance gap remains compared to full fine-tuning. Furthermore, existing PEFT approaches are often highly parameterized, making them difficult to interpret and control. In light of this, we introduce Multimodal Representation Tuning (MRT), a novel approach that focuses on directly editing semantically rich multimodal representations to achieve strong performance and provide intuitive control over LMMs. Empirical results show that our method surpasses current state-of-the-art baselines with significant performance gains (e.g., 1580.40 MME score) while requiring substantially fewer tunable parameters (e.g., 0.03% parameters). Additionally, we conduct experiments on editing instrumental tokens within multimodal representations, demonstrating that direct manipulation of these representations enables simple yet effective control over network behavior.
Yiyang Liu 0003, James Liang, Ruixiang Tang, Yugyung Lee, Majid Rabbani, Sohail A. Dianat, Raghuveer M. Rao, Lifu Huang, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001
ICLR7
2025 Latent Chain-of-Thought for Visual Reasoning
abstract
Chain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such as SFT, PPO, and GRPO may not generalize well across unseen reasoning tasks and heavily rely on a biased reward model. To address this challenge, we reformulate reasoning in LVLMs as posterior inference and propose a scalable training algorithm based on amortized variational inference. By leveraging diversity-seeking reinforcement learning algorithms, we introduce a novel sparse reward function for token-level learning signals that encourage diverse, high-likelihood latent CoT, overcoming deterministic sampling limitations and avoiding reward hacking. Additionally, we implement a Bayesian inference-scaling strategy that replaces costly Best-of-N and Beam Search with a marginal likelihood to efficiently rank optimal rationales and answers. We empirically demonstrate that the proposed method enhances the state-of-the-art LVLMs on four reasoning benchmarks, in terms of effectiveness, generalization, and interpretability.
Hang Hua, Jiebo Luo 0001, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Zhiqiang Tao
NeurIPS7
2024 Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
abstract
The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, relying on consistent embedding representations to compute similarity. However, the text content in existing datasets is generally short and concise, making it hard to fully describe the redundant semantics of a video. Correspondingly, a single text embedding may be less expressive to capture the video embedding and empower the retrieval. In this study, we propose a new stochastic text modeling method T-MASS, i.e., text is modeled as a stochastic embedding, to enrich text embedding with a flexible and re-silient semantic range, yielding a text mass. To be specific, we introduce a similarity-aware radius module to adapt the scale of the text mass upon the given text-video pairs. Plus, we design and develop a support text regularization to further control the text mass during the training. The inference pipeline is also tailored to fully exploit the text mass for accurate retrieval. Empirical evidence suggests that T-MASS not only effectively attracts relevant text-video pairs while distancing irrelevant ones, but also enables the de-termination of precise text embeddings for relevant pairs. Our experimental results show a substantial improvement of T-MASS over baseline (3% ~ 6.3% by R@1). Also, T-MASS achieves state-of-the-art performance on five bench-mark datasets, including MSRVTT, LSMDC, DiDeMo, VA-TEX, and Charades. Code and models are available here.
Jiamian Wang, Pichao Wang, Dongfang Liu, Sohail A. Dianat, Raghuveer M. Rao, Majid Rabbani, Zhiqiang Tao
CVPR6
2024 AMD: Automatic Multi-step Distillation of Large-Scale Vision Models
Cheng Han 0001, Qifan Wang 0001, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Yi Fang 0008, Qiang Guan, Lifu Huang, Dongfang Liu
ECCV (65)5
2024 Enhancing GAN Performance Through Neural Architecture Search and Tensor Decomposition
abstract
Generative Adversarial Networks (GANs) have emerged as a powerful tool for generating high-fidelity content. This paper presents a new training procedure that leverages Neural Architecture Search (NAS) to discover the optimal architecture for image generation while employing the Maximum Mean Discrepancy (MMD) repulsive loss for adversarial training. Moreover, the generator network is compressed using tensor decomposition to reduce its computational footprint and inference time while preserving its generative performance. Experimental results show improvements of 34% and 28% in the FID score on the CIFAR-10 and STL-10 datasets, respectively, with corresponding footprint reductions of 14× and 31× compared to the best FID score method reported in the literature. The implementation code is available at: https://github.com/PrasannaPulakurthi/MMD-AdversarialNAS.
Prasanna Reddy Pulakurthi, Mahsa Mozaffari, Sohail A. Dianat, Majid Rabbani, Jamison Heard, Raghuveer M. Rao
ICASSP6
2024 Tokenmotion: Motion-Guided Vision Transformer for Video Camouflaged Object Detection VIA Learnable Token Selection
abstract
The area of Video Camouflaged Object Detection (VCOD) presents unique challenges in the field of computer vision due to texture similarities between target objects and their surroundings, as well as irregular motion patterns caused by both objects and camera movement. In this paper, we introduce TokenMotion (TMNet), which employs a transformer-based model to enhance VCOD by extracting motion-guided features using a learnable token selection. Evaluated on the challenging MoCA-Mask dataset, TMNet achieves state-of-the-art performance in VCOD. It outperforms the existing state-of-the-art method by a 12.8% improvement in weighted F-measure, an 8.4% enhancement in S-measure, and a 10.7% boost in mean IoU. The results demonstrate the benefits of utilizing motion-guided features via learnable token selection within a transformer-based framework to tackle the intricate task of VCOD. The code of our work will be available when the paper is accepted.
Zifan Yu, Erfan Bank Tavakoli, Meida Chen, Suya You, Raghuveer M. Rao, Sanjeev Agarwal, Fengbo Ren
ICASSP5
2024 Image Translation as Diffusion Visual Programmers
abstract
We introduce the novel Diffusion Visual Programmer (DVP), a neuro-symbolic image translation framework. Our proposed DVP seamlessly embeds a condition-flexible diffusion model within the GPT architecture, orchestrating a coherent sequence of visual programs ($i.e.$, computer vision models) for various pro-symbolic steps, which span RoI identification, style transfer, and position manipulation, facilitating transparent and controllable image translation processes. Extensive experiments demonstrate DVP’s remarkable performance, surpassing concurrent arts. This success can be attributed to several key features of DVP: First, DVP achieves condition-flexible translation via instance normalization, enabling the model to eliminate sensitivity caused by the manual guidance and optimally focus on textual descriptions for high-quality content generation. Second, the frame work enhances in-context reasoning by deciphering intricate high-dimensional concepts in feature spaces into more accessible low-dimensional symbols ($e.g.$, [Prompt], [RoI object]), allowing for localized, context-free editing while maintaining overall coherence. Last but not least, DVP improves systemic controllability and explainability by offering explicit symbolic representations at each programming stage, empowering users to intuitively interpret and modify results. Our research marks a substantial step towards harmonizing artificial image translation processes with cognitive intelligence, promising broader applications.
Cheng Han 0001, James Liang, Qifan Wang 0001, Majid Rabbani, Sohail A. Dianat, Raghuveer M. Rao, Ying Nian Wu, Dongfang Liu
ICLR6
2024 Prototypical Transformer As Unified Motion Learners
abstract
In this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoFormer seamlessly integrates prototype learning with Transformer by thoughtfully considering motion dynamics, introducing two innovative designs. First, Cross-Attention Prototyping discovers prototypes based on signature motion patterns, providing transparency in understanding motion scenes. Second, Latent Synchronization guides feature representation learning via prototypes, effectively mitigating the problem of motion uncertainty. Empirical results demonstrate that our approach achieves competitive performance on popular motion tasks such as optical flow and scene depth. Furthermore, it exhibits generality across various downstream tasks, including object tracking and video stabilization.
Cheng Han 0001, Yawen Lu, James Liang, Zhiwen Cao, Qifan Wang 0001, Qiang Guan, Sohail A. Dianat, Raghuveer M. Rao, Tong Geng, Zhiqiang Tao, Dongfang Liu
ICML9
2024 Diffusion-Inspired Truncated Sampler for Text-Video Retrieval
abstract
Prevalent text-to-video retrieval methods represent multimodal text-video data in a joint embedding space, aiming at bridging the relevant text-video pairs and pulling away irrelevant ones. One main challenge in state-of-the-art retrieval methods lies in the modality gap, which stems from the substantial disparities between text and video and can persist in the joint space. In this work, we leverage the potential of Diffusion models to address the text-video modality gap by progressively aligning text and video embeddings in a unified space. However, we identify two key limitations of existing Diffusion models in retrieval tasks: The L2 loss does not fit the ranking problem inherent in text-video retrieval, and the generation quality heavily depends on the varied initial point drawn from the isotropic Gaussian, causing inaccurate retrieval. To this end, we introduce a new Diffusion-Inspired Truncated Sampler (DITS) that jointly performs progressive alignment and modality gap modeling in the joint embedding space. The key innovation of DITS is to leverage the inherent proximity of text and video embeddings, defining a truncated diffusion flow from the fixed text embedding to the video embedding, enhancing controllability compared to adopting the isotropic Gaussian. Moreover, DITS adopts the contrastive loss to jointly consider the relevant and irrelevant pairs, not only facilitating alignment but also yielding a discriminatively structured embedding. Experiments on five benchmark datasets suggest the state-of-the-art performance of DITS. We empirically find that DITS can also improve the structure of the CLIP embedding space. Code is available at https://github.com/Jiamian- Wang/DITS-text-video-retrieval
Jiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Zhiqiang Tao
NeurIPS7
2024 DeepFTSG: Multi-stream Asymmetric USE-Net Trellis Encoders with Shared Decoder Feature Fusion Architecture for Video Motion Segmentation
abstract
Abstract Discriminating salient moving objects against complex, cluttered backgrounds, with occlusions and challenging environmental conditions like weather and illumination, is essential for stateful scene perception in autonomous systems. We propose a novel deep architecture, named DeepFTSG, for robust moving object detection that incorporates single and multi-stream multi-channel USE-Net trellis asymmetric encoders extending U-Net with squeeze and excitation (SE) blocks and a single shared decoder network for fusing multiple motion and appearance cues. DeepFTSG is a deep learning based approach that builds upon our previous hand-engineered flux tensor split Gaussian (FTSG) change detection video analysis algorithm which won the CDNet CVPR Change Detection Workshop challenge competition. DeepFTSG generalizes much better than top-performing motion detection deep networks, such as the scene-dependent ensemble-based FgSegNet_v2, while using an order of magnitude fewer weights. Short-term motion and longer-term change cues are estimated using general-purpose unsupervised methods—flux tensor and multi-modal background subtraction, respectively. DeepFTSG was evaluated using the CDnet-2014 change detection challenge dataset, the largest change detection video sequence benchmark with 12.3 billion labeled pixels, and had an overall F-measure of 97%. We also evaluated the cross-dataset generalization capability of DeepFTSG trained solely on CDnet-2014 short video segments and then evaluated on unseen SBI-2015, LASIESTA and LaSOT benchmark videos. On the unseen SBI-2015 dataset, DeepFTSG had an F-measure accuracy of 87%, more than 30% higher compared to the top-performing deep network FgSegNet_v2 and outperforms the recently proposed KimHa method by 17%. On the unseen LASIESTA, DeepFTSG had an F-measure of 88% and outperformed the best recent deep learning method BSUV-Net2.0 by 3%. On the unseen LaSOT with axis-aligned bounding box ground-truth, network segmentation masks were converted to bounding boxes for evaluation, DeepFTSG had an F-Measure of 55%, outperforming KimHa method by 14% and FgSegNet_v2 by almost 1.5%. When a customized single DeepFTSG model is trained in a scene-dependent manner for comparison with state-of-the-art approaches, then DeepFTSG performs significantly better, reaching an F-Measure of 97% on SBI-2015 (+ 10%) and 99% on LASIESTA (+ 11%). The source code, pre-trained weights, and video demo for DeepFTSG are available at https://github.com/CIVA-Lab/DeepFTSG .
Gani Rahmon, Kannappan Palaniappan, Imad Eddine Toubal, Filiz Bunyak, Raghuveer M. Rao, Guna Seetharaman
Int. J. Comput. Vis.5
2024 Geometrical Interpretation and Design of Multilayer Perceptrons
abstract
The multilayer perceptron (MLP) neural network is interpreted from the geometrical viewpoint in this work, that is, an MLP partition an input feature space into multiple nonoverlapping subspaces using a set of hyperplanes, where the great majority of samples in a subspace belongs to one object class. Based on this high-level idea, we propose a three-layer feedforward MLP (FF-MLP) architecture for its implementation. In the first layer, the input feature space is split into multiple subspaces by a set of partitioning hyperplanes and rectified linear unit (ReLU) activation, which is implemented by the classical two-class linear discriminant analysis (LDA). In the second layer, each neuron activates one of the subspaces formed by the partitioning hyperplanes with specially designed weights. In the third layer, all subspaces of the same class are connected to an output node that represents the object class. The proposed design determines all MLP parameters in a feedforward one-pass fashion analytically without backpropagation. Experiments are conducted to compare the performance of the traditional backpropagation-based MLP (BP-MLP) and the new FF-MLP. It is observed that the FF-MLP outperforms the BP-MLP in terms of design time, training time, and classification performance in several benchmarking datasets. Our source code is available at https://colab.research.google.com/drive/1Gz0L8AnT4ijrUchrhEXXsnaacrFdenn?usp = sharing.
Ruiyuan Lin, Zhiruo Zhou, Suya You, Raghuveer M. Rao, C.-C. Jay Kuo
IEEE Trans. Neural Networks Learn. Syst.4
2023 Enhanced Low-Resolution LiDAR-Camera Calibration via Depth Interpolation and Supervised Contrastive Learning
abstract
Motivated by the increasing application of low-resolution LiDAR, we target the problem of low-resolution LiDAR-camera calibration in this work. The main challenges are two-fold: sparsity and noise in point clouds. To address the problem, we propose to apply depth interpolation to increase the point density and supervised contrastive learning to learn noise-resistant features. The experiments on RELLIS-3D demonstrate that our approach achieves an average mean absolute rotation/translation errors of 0.15cm/0.33° on 32-channel LiDAR point cloud data, which significantly outperforms all reference methods.
Zifan Yu, Suya You, Raghuveer M. Rao, Sanjeev Agarwal, Fengbo Ren
ICASSP4
2023 TransUPR: A Transformer-based Plug-and-Play Uncertain Point Refiner for LiDAR Point Cloud Semantic Segmentation
abstract
Common image-based LiDAR point cloud semantic segmentation (LiDAR PCSS) approaches have bottlenecks resulting from the boundary-blurring problem of convolution neural networks (CNNs) and quantitation loss of spherical projection. In this work, we propose a transformer-based plug-and-play uncertain point refiner, i.e., TransUPR, to refine selected uncertain points in a learnable manner, which leads to an improved segmentation performance. Uncertain points are sampled from coarse semantic segmentation results of 2D image segmentation where uncertain points are located close to the object boundaries in the 2D range image representation and 3D spherical projection background points. Following that, the geometry and coarse semantic features of uncertain points are aggregated by neighbor points in 3D space without adding expensive computation and memory footprint. Finally, the transformer-based refiner, which contains four stacked self-attention layers, along with an MLP module, is utilized for uncertain point classification on the concatenated features of self-attention layers. As the proposed refiner is independent of 2D CNNs, our TransUPR can be easily integrated into any existing image-based LiDAR PCSS approaches, e.g., CENet. Our TransUPR with the CENet achieves state-of-the-art performance, i.e., 68.2% mean Intersection over Union (mIoU) on the Semantic KITTI benchmark, which provides a performance improvement of 0.6% on the mIoU compared to the original CENet.
Zifan Yu, Meida Chen, Suya You, Raghuveer M. Rao, Sanjeev Agarwal, Fengbo Ren
IROS5
2022 Semi-supervised Multi-source Domain Adaptation in Wearable Activity Recognition
abstract
The scarcity of labeled data has traditionally been the primary hindrance in building scalable supervised deep learning models that can retain adequate performance in the presence of various heterogeneities in sample distributions. Domain adaptation tries to address this issue by adapting features learned from a smaller set of labeled samples to that of the incoming unlabeled samples. The traditional domain adaptation approaches normally consider only a single source of labeled samples, but in real world use cases, labeled samples can originate from multiple-sources – providing motivation for multi-source domain adaptation (MSDA). Several MSDA approaches have been investigated for wearable sensor-based human activity recognition (HAR) in recent times, but their performance improvement compared to single source counterpart remained marginal. To remedy this performance gap that, we explore multiple avenues to align the conditional distributions in addition to the usual alignment of marginal ones. In our investigation, we extend an existing multi-source domain adaptation approach under semi-supervised settings. We assume the availability of partially labeled target domain data and further explore the pseudo labeling usage with a goal to achieve a performance similar to the former. In our experiments on three publicly available datasets, we find that a limited labeled target domain data and pseudo label data boost the performance over the unsupervised approach by 10-35% and 2-6%, respectively, in various domain adaptation scenarios.
Avijoy Chakma, Abu Zaher Md Faridee, Raghuveer M. Rao, Nirmalya Roy
DCOSS3
2021 On Relationship of Multilayer Perceptrons and Piecewise Polynomial Approximators
abstract
The relationship between a multilayer perceptron (MLP) regressor and a piecewise polynomial approximator is investigated in this work. We propose an MLP construction method, including the choice of activation, the specification of neuron numbers and filter weights. Through the construction, a one-to-one correspondence between an MLP and a piecewise polynomial is established. Especially, we point out that the form of nonlinear activation is related to the polynomial order. Since the approximation capability of piecewise polynomials is well understood, our study sheds new light on the universal approximation capability of an MLP.
Ruiyuan Lin, Suya You, Raghuveer M. Rao, C.-C. Jay Kuo
IEEE Signal Process. Lett.3
2020 Pixelhop++: A Small Successive-Subspace-Learning-Based (Ssl-Based) Model For Image Classification
abstract
The successive subspace learning (SSL) principle was developed and used to design an interpretable learning model, known as the PixelHop method, for image classification in our prior work. Here, we propose an improved PixelHop method and call it Pixel-Hop++. First, to make the PixelHop model size smaller, we decouple a joint spatial-spectral input tensor to multiple spatial tensors (one for each spectral component) under the spatial-spectral separability assumption and perform the Saab transform in a channel-wise manner, called the channel-wise (c/w) Saab transform. Second, by performing this operation from one hop to another successively, we construct a channel-decomposed feature tree whose leaf nodes contain features of one dimension (1D). Third, these 1D features are ranked according to their cross-entropy values, which allows us to select a subset of discriminant features for image classification. In Pixel-Hop++, one can control the learning model size of fine-granularity, offering a flexible tradeoff between the model size and the classification performance. We demonstrate the flexibility of Pixel-Hop++ on MNIST, Fashion MNIST, and CIFAR-10 three datasets.
Yueru Chen, Mozhdeh Rouhsedaghat, Suya You, Raghuveer M. Rao, C.-C. Jay Kuo
ICIP4
2020 Deep Learning Based Landmark Matching For Aerial Geolocalization
abstract
Visual odometry has gained increasing attention due to the proliferation of unmanned aerial vehicles, self-driving cars, and other autonomous robotics systems. Landmark detection and matching are critical for visual localization. While current methods rely upon point-based image features or descriptor mappings we consider landmarks at the object level. In this paper, we propose LMNet a deep learning based landmark matching pipeline for city-scale, aerial images of urban scenes. LMNet consists of a Siamese network, extended with a multi-patch based matching scheme, to handle offcenter landmarks, varying landmark scales, and occlusions of surrounding structures. While there exist a number of landmark recognition benchmark datasets for ground-based and nadir aerial or satellite imagery, there is a lack of datasets and results for oblique aerial imagery. We use a unique unsupervised multi-view landmark image generation pipeline for training and testing the proposed matching pipeline using over 0.5 million real landmark patches. Results for aerial landmark matching across four cities show promising results.
Koundinya Nouduri, Filiz Bunyak, Shizeng Yao, Hadi Aliakbarpour, Sanjeev Agarwal, Raghuveer M. Rao, Kannappan Palaniappan
ICIP6
2018 Why the Failure? How Adversarial Examples Can Provide Insights for Interpretable Machine Learning
abstract
Recent advances in Machine Learning (ML) have profoundly changed many detection, classification, recognition and inference tasks. Given the complexity of the battlespace, ML has the potential to revolutionise how Coalition Situation Understanding is synthesised and revised. However, many issues must be overcome before its widespread adoption. In this paper we consider two - interpretability and adversarial attacks. Interpretability is needed because military decision-makers must be able to justify their decisions. Adversarial attacks arise because many ML algorithms are very sensitive to certain kinds of input perturbations. In this paper, we argue that these two issues are conceptually linked, and insights in one can provide insights in the other. We illustrate these ideas with relevant examples from the literature and our own experiments.
Richard Tomsett, Amy Widdicombe, Tianwei Xing, Supriyo Chakraborty, Simon J. Julier, Prudhvi Gurram, Raghuveer M. Rao, Mani Srivastava 0001
FUSION7
2016 Moving object detection for vehicle tracking in Wide Area Motion Imagery using 4D filtering
abstract
Most Wide Area Motion Imagery (WAMI) based trackers use motion based cueing for detecting and tracking moving objects. The results are very high false alarm rates in urban environments with tall structures due to parallax effects. This paper proposes an accurate moving object detection method using a precise orthorectification approach for ground stabilization combined with accurate multiview depth maps to reduce the number of false positives induced by parallax effects by 90 percent. Proposed hybrid moving vehicle detection approach for large scale aerial urban imagery is based on fusion of motion detection mask obtained from median-based background subtraction and tall structures height mask provided by image depth map information. Using buildings mask enables us to improve the object level detection accuracy in terms of F-measure by 57 percent from 22.2% to 79.2%.
Kannappan Palaniappan, Mahdieh Poostchi, Hadi Aliakbarpour, Raphael Viguier, Joshua Fraser, Filiz Bunyak, Arslan Basharat, Steve Suddarth, Erik Blasch, Raghuveer M. Rao, Guna Seetharaman
ICPR10
2014 Multispectral Image Denoising With Optimized Vector Bilateral Filter
abstract
Vector bilateral filtering has been shown to provide good tradeoff between noise removal and edge degradation when applied to multispectral/hyperspectral image denoising. It has also been demonstrated to provide dynamic range enhancement of bands that have impaired signal to noise ratios (SNRs). Typical vector bilateral filtering described in the literature does not use parameters satisfying optimality criteria. We introduce an approach for selection of the parameters of a vector bilateral filter through an optimization procedure rather than by ad hoc means. The approach is based on posing the filtering problem as one of nonlinear estimation and minimization of the Stein's unbiased risk estimate of this nonlinear estimator. Along the way, we provide a plausibility argument through an analytical example as to why vector bilateral filtering outperforms bandwise 2D bilateral filtering in enhancing SNR. Experimental results show that the optimized vector bilateral filter provides improved denoising performance on multispectral images when compared with several other approaches.
Honghong Peng, Raghuveer M. Rao, Sohail A. Dianat
IEEE Trans. Image Process.2
2012 Optimized vector bilateral filter for multispectral image denoising
abstract
Vector bilateral filtering has been shown to provide several advantages in processing hyperspectral images such as good noise removal while minimizing edge degradation, and dynamic range enhancement of bands with impaired signal to noise ratios. This paper introduces an approach for selection of the parameters of a vector bilateral filter through an optimization procedure rather than by ad hoc means. The approach is based on posing the filtering problem as one of nonlinear estimation and minimizing the Stein's unbiased risk estimate (SURE) of this nonlinear estimator. Experimental results show that the optimized vector bilateral filter provides improved denoising performance on multispectral images when compared to several other approaches.
Honghong Peng, Raghuveer M. Rao, Sohail A. Dianat
ICIP2
2012 Nonnegative matrix factorization with deterministic annealing for unsupervised unmixing of hyperspectral imagery
abstract
Non-negative matrix factorization (NMF) technique and its extensions were developed to find part based, linear representations of non-negative multivariate data. They have been shown to provide more interpretable results with realistic non-negative constrain in unsupervised learning applications such as hyperspectral imagery unmixing, image feature extraction, and data mining. This paper extends the NMF method by incorporating deterministic annealing optimization procedure, which will help solve the non-convexity problem in NMF and provide a better choice of sparseness constrain. The approach is based on replacing the difficult non-convex optimization problem of NMF with an easier one by adding an auxiliary convex entropy constrain term and solving this first. Experiment results with hyperspectral unmixing application show that the proposed technique provides improved unmixing performance compared to other state-of-the-art methods.
Honghong Peng, Raghuveer M. Rao, Sohail A. Dianat
ICIP2
2011 Parallel flux tensor analysis for efficient moving object detection
Kannappan Palaniappan, Ilker Ersoy, Guna Seetharaman, Shelby R. Davis, Praveen Kumar 0005, Raghuveer M. Rao, Richard W. Linderman
FUSION6
2010 Efficient feature extraction and likelihood fusion for vehicle tracking in low frame rate airborne video
Kannappan Palaniappan, Filiz Bunyak, Praveen Kumar 0005, Ilker Ersoy, Stefan Jäger 0001, Koyeli Ganguli, Anoop Haridas, Joshua Fraser, Raghuveer M. Rao, Guna Seetharaman
FUSION9
2010 Bilateral kernel parameter optimization by risk minimization
abstract
A statistical approach is presented for determining the optimal parameters of a bilateral filter for suppressing noise in an image. A closed-form solution to obtaining Stein's unbiased estimate of the bilateral filter-based image estimation risk is derived. The values of the parameters are then determined by minimization of the estimated risk, which is equivalent to maximization of the estimated signal to noise ratio. Experimental results show that the method provides estimates of noise intensity that are close to the actual, and the performance of the bilateral filter, in terms of providing good tradeoff between noise suppression and edge preservation, is improved significantly by using the optimal parameters estimate.
Honghong Peng, Raghuveer M. Rao
ICIP2
2009 Segmentation of malaria parasites in peripheral blood smear images
abstract
Detection of malaria parasites in stained blood smears is critical for treatment of the disease. Automation of this process will help in reducing the time taken for diagnosis and the chance for human errors. However, the variability and artifacts in microscope images of blood samples pose significant challenges for accurate detection. A scheme based on HSV color space that segments Red Blood Cells and parasites by detecting dominant hue range and by calculating optimal saturation thresholds is presented in this paper. Methods that are less computation-intensive than existing approaches are proposed to remove artifacts. The scheme is evaluated using images taken from Leishman-stained blood smears. Sensitivity and specificity of the scheme are found to be 83% and 98% respectively.
Vishnu Vardhan Makkapati, Raghuveer M. Rao
ICASSP2
2009 Ordering random object poses
abstract
Complete or partial three-dimensional reconstruction of objects from multiple angle-views, or poses, is important in several applications such as photogrammetry, machine vision, and computer-aided design. Knowledge of the pose angles and their proper ordering are required for accurate reconstruction. When these multiple angle images are acquired in random order and the angle of view information is not available the poses have to be put into proper order. This work presents an approach based on principal component analysis (PCA) for automatic ordering of random object poses. A measure based on local curvature and correlation of the estimated pose trajectory in a multidimensional manifold is also developed to assess confidence in the ordering. In addition to providing a degree of confidence for pose ordering with single cameras, this measure enhances the pose estimation accuracy in double and multiple camera systems by providing a basis for camera selection for different poses. The paper presents theoretical development and experimental results.
James Massaro, Raghuveer M. Rao
ICASSP2
2009 Hyperspectral image enhancement with vector bilateral filtering
abstract
An approach is proposed to extend bilateral filtering to the vector case so as to simultaneously take spectral and spatial information into account by using spectral distances and multivariate Gaussian functions. To simplify the determination of the parameters of the corresponding covariance matrix, the data vectors are transformed to eigenspace through principal component analysis (PCA). By locally adapting to the spectral distribution in decorrelated PCA space, the proposed approach offers effective noise removal while keeping the spatial details in the band images. It also provides dynamic range enhancement of severely affected bands to make meaningful data extraction possible. Experimental results with the proposed approach using remote-sensed hyperspectral data demonstrate improved denoising and enhancement in comparison to existing methods.
Honghong Peng, Raghuveer M. Rao
ICIP2
2008 Constrained optimization algorithm for digital camera-based spectrometer
abstract
Spectrometers that provide spectral decomposition at different locations of a scene are required in many applications. A simple device based on a multi-channel digital camera system has recently been developed to provide pixel level spectrum estimation. The current spectrum estimation algorithm for the device is based on PCA analysis. This paper develops a new estimation approach in an optimization framework that is based on a model of the optical path. The new method provides better spectrum estimates. Theoretical development and experimental examples are provided.
Hongkong Peng, Raghuveer M. Rao
ICASSP2
2007 Image Enhancement of FOG-Impaired Scenes with Variable Visibility
abstract
Enhancement of images of foggy scenes is being investigated for application to navigational aids. Approaches based on atmospheric scatter models have previously been developed. In most cases, the assumption is made of uniform aerosol suspension leading to models of constant attenuation per unit distance. The paper addresses situations where such an assumption is not valid and proposes an approach for image enhancement when visibility varies with direction. The basis for the development of the approach is provided along with an example of application.
Honghong Peng, Raghuveer M. Rao
ICASSP (2)2
2006 A Video Processing Approach for Distance Estimation
abstract
The paper presents a passive ranging method for estimating distances to an object using a video sequence gathered from a moving platform. The motivation is provided by potential application to object distance estimation using video data from general aviation aircraft. The method exploits scale changes of an object in the video sequence, as inferred by processing wavelet transforms of video frames, to compute distance. The underlying principles are presented along with results of bench experiments.
Raghuveer M. Rao, Seungsin Lee
ICASSP (3)1
2006 Signal Processing for Hyperspectral Data
abstract
Hyperspectral data form a data-cube consisting of images of an object collected at several hundred, closely spaced wavelengths. They have been found to be of significant potential benefit in areas such as remote sensing of the Earth, medicine, and non-destructive evaluation. Effective extraction of information from the hyperspectral data cube presents several signal processing challenges, some of them unique to hyperspectral data. The problems involved range from registration and enhancement to development of statistical signal processing algorithms and models for object detection and classification. The focus of this paper is to provide an overview of select processing and modeling techniques for hyperspectral data
Pramod K. Varshney, Manoj K. Arora, Raghuveer M. Rao
ICASSP (5)3
2006 Object Labeling for 3-D Cross-Sectional Data using Trajectory Tracking
abstract
Volume visualization from object cross sections has created the need for object labeling in volume data sets. Labeling techniques for 3D objects vary from the memory-intensive connected-component labeling to the less intense overlap technique. In this paper, a novel labeling technique is presented that represents the objects with curves in 3D and then performs labeling by applying 3D curve tracing techniques and mapping the labels back to the 3D objects. The technique thus attempts to replicate aspects of the human visual approach to the task.
Brandon Mikulis, Raghuveer M. Rao
ICIP2
2006 Self-similar random field models in discrete space
abstract
Self-similar random fields are of interest in various areas of image processing since they fit certain types of natural patterns and textures. Current treatments of self-similarity in continuous two-dimensional (2-D) space use a definition that is a direct extension of the one-dimensional definition, which requires invariance of the statistics of a random process to time scaling. Current discrete-space 2-D approaches do not consider scaling, but, instead, are based on ad hoc formulations, such as digitizing continuous random fields. In this paper, we show that the current statistical self-similarity definition in continuous space is restrictive and provide an alternative, more general definition. We also provide a formalism for discrete-space statistical self-similarity that relies on a new scaling operator for discrete images. Within the new framework, it is possible to synthesize a wider class of discrete-space self-similar random fields and texture images.
Seungsin Lee, Raghuveer M. Rao
IEEE Trans. Image Process.2
2005 Factors influencing psycophysically valid taxonomies of image texture
abstract
Image texture is important in human and machine vision. A taxonomy of image texture that classifies textures the same way humans do psychophysically can be used in many fields. This paper deals with one attribute of texture, namely orderliness. To determine what underlying factors influence humans to perceive orderliness in textures, psychophysical direct magnitude estimation ratings of the orderliness of 27 Brodatz images were collected from 44 subjects. The images: a) were either tiled, locally oriented, or granular, b) had large, medium, or small scale elements, and c) contained either high, medium, or low regularity. Multidimensional scaling revealed three underlying factors that determined the perception of texture orderliness: uniformity of element shape and distribution, element size, and element dimensionality. Some of these results contradict predictions made by earlier computational models of texture. Such models should be revised to incorporate the results of our experiment.
Deepak Dewangan, Vincent J. Samar, Raghuveer M. Rao, Peter Paul
ICIP (3)3
2005 Algorithms for scene restoration and visibility estimation from aerosol scatter impaired images
abstract
Attenuation and backscatter from aerosol suspension in the atmosphere cause drop in visibility which manifests as haziness or fogginess in images captured of objects at a distance, especially from airborne cameras. Two iterative algorithms, based on a physical model relating image intensity to ground reflectance as a function of the scatter process, are developed to restore the images and estimate the scatter coefficient. The algorithms are guaranteed to converge to unique solutions. Examples are provided with synthetic and real data.
Raghuveer M. Rao, Seungsin Lee
ICIP (1)1
2004 Discrete space models for self-similar random images
abstract
Images exhibiting statistical self-similarity are of interest in various areas of image processing such as textures and scene synthesis. In continuous-space, statistical self-similarity is defined through statistics invariant to spatial scaling. However, because of lack of discrete-space scaling operation, statistical self-similarity in discrete-space has been characterized by approaches such as increments of fractional Brownian motion rather than scaling. We address these two issues regarding self-similar random fields through the paper. We show that the current self-similarity definition for continuous-space is somewhat restrictive, and we offer a new self-similarity definition in continuous-space more general than the current one. Furthermore, we provide a new formalism for statistical self-similarity in discrete-space by defining a scaling operation for discrete-space images. Consequently, a wider class of self-similar random images can be synthesized for different applications in discrete-space. The paper presents theoretical development and synthesis examples.
Seungsin Lee, Raghuveer M. Rao
ICASSP (3)2
2004 Scale-based formulations of statistical self-similarity in images
abstract
Statistically self-similar images that are segments of two dimensional self-similar random fields, have been useful in the analysis and synthesis of certain types of textures. Whereas a rigorous definition of self-similarity in continuous-space is based on spatial scaling, current treatments in digital image processing are based on ad-hoc approaches rather than on spatial scaling mainly because of the unavailability of continuous scaling in discrete-space. This paper presents a formulation based on such a continuous scaling operator leading to a more general and versatile characterization of statistical self-similarity in images.
Seungsin Lee, Raghuveer M. Rao
ICIP2
2003 Solution to the orthogonal M-channel bandlimited wavelet construction proposition
abstract
While bandlimited wavelets and associated IIR filters have shown serious potential in areas of pattern recognition and communications, the dyadic Meyer wavelet is the only known approach to construct bandlimited orthogonal decomposition. The sinc scaling function and wavelet are a special case of the Meyer. Previous works have proposed a M-Band extension of the Meyer wavelet without solving the problem. One key contribution of this paper is the derivation of the correct bandlimits for the scaling function and wavelets to guarantee an orthogonal basis. In addition, the actual construction of the wavelets based upon these bandlimits is developed. A composite wavelet is derived based on the M scale relationships from which we extract the wavelet functions. The proper solution to this task is proposed which generates associated filters with the knowledge of the scaling function and the constraints for M-band orthogonality.
Bryce Tennant, Raghuveer M. Rao
ICASSP (6)2
2003 Evolutionary algorithms for the edge biconnectivity augmentation problem
abstract
A graph is edge-biconnected if it requires the removal of at least two edges to disconnect it. Assume that we have weighted graph that is not biconnected, and an additional set of augmentation edges. The (NP-hard) edge biconnectivity augmentation problem is to select a minimal subset of the augmentation edges, whose inclusion will cause the graph to be biconnected. This paper explores the application of particle swarm optimization and genetic algorithms for this problem.
T. M. Rao, Raghuveer M. Rao, Ferat Sahin, Jason C. Tillett
SMC2
2003 Robust recruitment near the edge of chaos and an application to mine sweeping
abstract
The field of robotics is in rapid development. As robots become cheaper to build, new applications involving many robots systems can be envisioned. One reason for using many robots is to achieve robustness. Having many robots, however, does not ensure robustness. A control strategy and robot behaviors must be engineered to incorporate robustness into the system. Swarm intelligence based approaches are popular for developing optimal and robust control strategies for systems of robots. Here we analyze the behavior of a swarm of robots modeled after a swarm of ants, where tasks are spatially distributed in the environment and robots/ants are recruited through short-range recruitment. For ants that move probabilistically in response to the short range signal and who adjust their probabilities such that they are near a phase change boundary, or edge-of-chaos, in the mean field analysis of their motions, we find a significant improvement in the robustness of the system.
Jason C. Tillett, T. M. Rao, Raghuveer M. Rao, Ferat Sahin
SMC3
2003 Discrete-time self-similar systems and stable distributions: applications to VBR video modeling
abstract
This paper investigates the application of discrete-time statistically self-similar systems to modeling variable-bit rate (VBR) video traces. Potential application to classifying scenes in VBR video is explored. The work is motivated by the fact that while VBR video has been characterized as self-similar by various researchers, models based on self-similarity considerations have not been previously studied. This paper also shows that using heavy-tailed inputs these models can be used to match both the scene density time-series autocorrelation as well as its marginal distribution.
Rajesh Narasimha, Raghuveer M. Rao
IEEE Signal Process. Lett.2
2002 Construction of discrete-time linear scale-invariant systems using kernels
abstract
A linear scale invariant (LSI) system model in discrete-time was suggested by Zhao and Rao using a specialized dilation or scaling operation. This scaling operation relies on a frequency warping transformation between discrete-time and continuous-time frequencies. The DLSI systems have been shown to provide synthesis of self-similar random signals such as those seen in network traffic. In this paper, we provide a more general method for DLSI system modeling based on linear kernels. The alternative expression of DLSI system provided by linear kernels has several desirable properties lacking in the earlier formulation. In particular, the new formulation (a) allows commutativity between input and system characteristic sequence and (b) permits definition of a scale domain transform operator analogous to the Fourier transform for time-invariant systems.
Seungsin Lee, Raghuveer M. Rao
ICASSP2
2002 Discrete-time scale invariant systems: Relation to long-range dependence and farima models
abstract
A discrete-time version of linear scale-invariant systems was provided by Zhao and Rao. This formulation, called a DLSI system, was derived using a continuous dilation operator in discrete-time as a direct analog of the continuous-time linear scale-invariant system formulation of Wornell. While simulations had shown that DLSI systems were capable of generating self-similar data such as those found in network traffic, answers to questions regarding their relationship to other models had not been found. This paper investigates such relationships. It derives results establishing parameter ranges for DLSI systems that give rise to long-range dependent outputs for white noise inputs. Furthermore, it shows that the basic DLSI system is expressible as a special fractional ARIMA model. Finally, the DLSI model is applied to scene classification in variable bit rate (VBR) MPEG video to provide a preliminary indication of potential application.
Rajesh Narasimha, Seungsin Lee, Raghuveer M. Rao
ICASSP3
2002 Noise reduction and object enhancement in passive millimeter wave concealed weapon detection
abstract
Passive MM wave offers the advantage of penetration for concealed weapon detection. Its ability to penetrate through fog, smoke, clothing etc. makes it an attractive candidate to look for weapons concealed underneath a person's clothing. The sensor technology has advanced to a point where it is possible to generate real time video sequences. However, noise and blur are still severe problems. The work reported here investigates the problem of simultaneous noise suppression and object enhancement in passive millimeter wave video sequences. The basis for the approach is provided by undecimated wavelet transforms in the spatial dimension and motion compensated filtering in the temporal dimension. The paper presents the underlying principles of the approach as well as experimental results.
Seungsin Lee, Raghuveer M. Rao, Mohamed-Adel Slamani
ICIP (1)2
2001 Characterization of self-similarity properties of discrete-time linear scale-invariant systems
abstract
Discrete-time linear systems that possess scale-invariance properties even in the presence of continuous dilation were proposed by Zhao and Rao (1998, 1999). The principal purpose of this article is to describe results of subsequent investigation which have led to characterization of self-similarity properties of discrete-time signals synthesized by these systems. It is shown that white noise inputs to these linear scale invariant systems, which are unique in the DSP literature, produce self-similar outputs regardless of the marginal distribution of the noise. In most instances the output is fractional Gaussian. For heavy-tailed input distributions, the output is also heavy-tailed and self-similar. It is also shown that it is possible to synthesize statistically self-similar signals whose self-similarity parameters are consistent with those observed in network traffic.
Seungsin Lee, Raghuveer M. Rao, Rajesh Narasimha
ICASSP2
2000 Optimal filters for multidimensional sampling and reconstruction of image classes
abstract
This paper deals with multidimensional sampling and optimal reconstruction of the input random field or image class in the mean squared sense. It considers the interpolation and approximation sampling systems and presents closed form expressions for the filters in these systems as well as the mean squared errors between the input and the reconstructed fields. For the approximation sampling systems the optimal filters turn out to be spectral factors of an ideal brickwall filter whose support is determined by the sampling lattice. Finally, it presents examples of filters matched to an input image class derived from a multispectral LANDSAT image. In an image coding context, these filters can be considered as providing the best tradeoff between coding rate reduction by undersampling and distortion in reconstruction. The performance of these filters is tested in the presence of a quantizer for different bit rates and is compared with that of some standard filters to demonstrate this.
Ajit S. Bopardikar, Raghuveer M. Rao
ICASSP2
1999 A discrete-time wavelet transform based on a continuous dilation framework
abstract
We present a new form of wavelet transform. Unlike the continuous wavelet transform (CWT) or discrete wavelet transform (DWT), the mother wavelet is chosen to be a discrete-time signal and wavelet coefficients are computed by correlating a given discrete-time signal with continuous dilations of the mother wavelet. The results developed are based on the definition of a discrete-time scaling (dilation) operator through a mapping between the discrete and continuous frequencies. The forward and inverse wavelet transforms are formulated. The admissibility condition is derived, and examples of discrete-time wavelet construction are provided. The new form of wavelet transform is naturally suited for discrete-time signals and provides analysis and synthesis of such signals over a continuous range of scaling factors.
Raghuveer M. Rao
ICASSP2
1999 Sampling Systems Matched to the Input Image Class
abstract
The authors present closed form expressions for filters in interpolation and approximation sampling systems, given the sampling lattice and matched to the input image class in the mean squared sense. We also derive expressions for the mean squared error incurred in both the systems while reconstructing individual images of the input class from their samples. For the approximation sampling case, we show that the optimal filters are spectral factors of a brickwall-type filter whose support is determined by the sampling lattice. Finally, we present examples of filters in interpolation and approximation sampling cases that are matched to an image class derived from a multispectral image. The performance of these filters is tested in the presence of a quantizer for different bit rates and is compared with that of some standard filters.
Ajit S. Bopardikar, Raghuveer M. Rao
ICIP (2)2
1999 Image Processing Tools for the Enhancement of Concealed Weapon Detection
abstract
A number of technologies are being developed for Concealed Weapon Detection (CWD). Use of appropriate processing techniques will be very important to the success of such technologies. This article describes digital image processing procedures currently being investigated to enhance the detection of weapons concealed underneath clothing.
Mohamed-Adel Slamani, Pramod K. Varshney, Raghuveer M. Rao, Mark G. Alford, David Ferris
ICIP (3)3
1998 Continuous-dilation discrete-time self-similar signals and linear scale-invariant systems
abstract
In this paper we present a novel model for purely discrete-time self-similar processes and scale-invariant systems. The results developed are based on a new interpretation of the discrete-time scaling (equivalently dilation or contraction) operation which is defined through a mapping between discrete and continuous time. It is shown that it is possible to have continuous scaling factors through this operation even though the signal itself is discrete-time. We study both deterministic and stochastic discrete-time self-similar signals. We then derive the existence conditions of discrete-time deterministically self-similar signals with respect to some specific mappings. Finally, we discuss the construction of discrete-time linear scale-invariant system and present results related to white noise driven system models of stochastic self-similar signals. Unlike continuous-time self-similar signals, it is possible to construct a wide class of non-trivial discrete-time self-similar signals.
Raghuveer M. Rao
ICASSP2
1997 Perfect reconstruction circular convolution filter banks and their application to the implementation of bandlimited discrete wavelet transforms
abstract
This paper introduces a new filter bank structure called the perfect reconstruction circular convolution (PRCC) filter bank. These filter banks satisfy the perfect reconstruction properties, namely, the paraunitary properties in the discrete frequency domain. We further show how the PRCC analysis and synthesis filter banks are completely implemented in this domain and give a simple and a flexible method for the design of these filters. Finally, we use this filter bank structure for a frequency sampled implementation of the discrete wavelet transform based on orthogonal bandlimited scaling functions and wavelets.
Ajit S. Bopardikar, Raghuveer M. Rao, B. Suryanarayana Adiga
ICASSP2
1997 Wavelet transform based detection of photon-limited and low contrast objects
abstract
This paper presents methods for detection and localization of photon-limited objects in noise. As opposed to the correlation based or Fourier transform based techniques which exhibit sensitivity to object scaling, we propose a method based on the continuous wavelet transform with its ability to reject noise and to localize objects in space and time as well as in scale. An advantageous twist presented here is the use of the wavelet transform on the complex envelope of the signal of interest. This has the advantage of reducing "rippling" effects seen in the transform of the original waveform. An example of further post-processing on the wavelet-transformed data is provided.
James P. LeBlanc, Raghuveer M. Rao
ICASSP2
1996 Matched wavelets-their construction, and application to object detection
abstract
Wavelets matched to features in an image can be used to detect and/or classify objects. In this paper, image object features are generated by projecting the object at a finite number of angles. Wavelets matched to the projected object features in a training image are used to detect similar objects in test images. The detection is shown to be shift and scale insensitive and rotation insensitive to the number of discrete angles used for the projections.
Joseph O. Chapa, Raghuveer M. Rao
ICASSP2
1996 Design of biorthogonal MRA generating wavelets matched to specified signal
abstract
The problem of fitting a mother wavelet to a signal such that this wavelet along with a dual generate a biorthogonal wavelet decomposition is addressed. It is shown that when the problem is addressed in the frequency domain with bandlimited scaling functions and wavelets, a simple formulation is possible. Further, a closed form sub-optimal solution is possible when the problem is broken down into a series of optimization problems. In the process of the problem formulation, a generalization of the Meyer class of wavelets to the biorthogonal case is obtained.
Raghuveer M. Rao, Joseph O. Chapa
ICASSP1
1995 Digital image halftoning by noise thresholding
abstract
Approaches for digital halftoning of images using dithering threshold the input image with additive dithering noise. The paper presents a technique which thresholds the noise directly. The threshold is modified at each step such that the expected value of the output is equal to the input pixel's gray value. Further, error feedback is used to correct the threshold. Tests on images show the method's ability to retain features in the high frequency regions such as edges as well as low frequency features such as slow variations in intensity. One advantage the proposed method is its very general nature since it offers a wide choice of the three filters: noise shaping filter, feedforward filter and feedback filter.
Anamitru Makur, Raghuveer M. Rao
ICASSP2
1992 Cross-bispectrum computation for multichannel quadratic phase coupling estimation
abstract
A two-channel, quadratic, nonlinear process driven by sinusoidal signals is considered. Equations for estimating the parameters of the process are derived from colored Gaussian noise contaminated observations of the process. The derivation is based on the bispectrum and cross-power spectrum of the two channels. The resulting equations are nonlinear and iterative techniques are needed to solve for the parameters.>
Raghuveer M. Rao, Sohail A. Dianat
ICASSP1
1991 Restoration of speckle-degraded images using bispectra
abstract
Coherent speckle noise is modeled as a multiplicative noise process. Using a logarithmic transformation, this speckle noise is converted to a signal-independent additive process which is close to Gaussian when an integrating aperture is used. Bispectral reconstruction of speckle-degraded images is performed on such logarithmically transformed images when independent multiple snapshots are available.>
Raghuveer M. Rao, S. Wear
ICASSP1
1990 Fast algorithms for bispectral reconstruction of two-dimensional signals
abstract
Algorithms are developed to reconstruct the phase and magnitude of the Fourier transform of a two-dimensional, discrete, deterministic signal from samples of its bispectrum B( omega , lambda ). The samples are derived from the plane omega - lambda . The techniques involve finding the DFTs (discrete Fourier transforms) of the phase and magnitude of the bispectrum samples, and from them, recovering samples of the phase and magnitude of the Fourier transform of the signal at select frequencies.>
Sohail A. Dianat, Raghuveer M. Rao
ICASSP2
1989 Bispectrum phase transformation for non-minimum phase signal reconstruction
abstract
The authors present a Fourier-series-based approach for recovering the phase of a signal from samples of its bispectrum along the omega /sub 1/= omega /sub 2/ line in the bispectrum frequency plane. The Fourier series coefficients of the signal phase are recovered from the Fourier series of the bispectrum phase. The implementation of the method is simplified by using the FFT (fast Fourier transform) to compute the Fourier series coefficients. An important advantage of the approach is that while using a finite number of samples of the bispectrum phase, it provides estimates for the phase of the signal at all frequencies. Also there are computational advantages in that the FFT can be used and it may not be necessary to consider a large number of Fourier series coefficients.>
Sohail A. Dianat, Raghuveer M. Rao
ICASSP2
1989 Polyspectral factorization: necessary and sufficient condition for finite extent cumulant sequences
abstract
The authors provide necessary and sufficient conditions for bispectral and trispectral factorization for processes with finite-extent cumulant sequences. These conditions are derived entirely in terms of the cumulant sequences. They use the fact that for a finite-extent cumulant sequence factorability is equivalent to finding a finite-order moving-average process with an identical cumulant sequence. In principle the results can be extended to polyspectra of even higher orders. An interesting result of the investigation is that there exist processes generated by nonlinear mechanisms that are factorable.>
Sohail A. Dianat, Raghuveer M. Rao
ICASSP2
1987 Bispectrum estimation: A digital signal processing framework
abstract
It is the purpose of this tutorial paper to place bispectrum estimation in a digital signal processing framework in order to aid engineers in grasping the utility of the available bispectrum estimation techniques, to discuss application problems that can directly benefit from the use of the bispectrum, and to motivate research in this area. Three general reasons are behind the use of bispectrum in signal processing and are addressed in the paper: to extract information due to deviations from normality, to estimate the phase of parametric signals, and to detect and characterize the properties of nonlinear mechanisms that generate time series.
Chrysostomos L. Nikias, Raghuveer M. Rao
Proc. IEEE2
1985 Bispectrum estimation for short length data
abstract
Bispectrum estimation has been used to detect quadratic phase coupling among sinusoids in noise, in phase measurements for non-Gaussian processes, system identification etc. Most of the known estimation methods belong to the 'conventional' or 'Fourier-type' class. Recently, a parametric approach was proposed. All known methods perform poorly in terms of resolution when the given data records are short. The Constrained Third Order Mean (CTOM) method proposed in this paper estimates the parameters of an autoregressive (AR) model driven by non-Gaussian white Noise (NGWN) by setting the sample mean of the Third Order Recursion (TOR) error process to zero. The resulting bispectrum estimates perform very well in detecting quadratic phase coupling when the data records are short.
Raghuveer M. Rao, Chrysostomos L. Nikias
ICASSP1
1984 A parametric approach to bispectrum estimation
abstract
Higher order spectra provide information about processes not contained in the ordinary power spectrum such as the degree of nonlinearity and deviations from normality. The bispectrum which is a third order spectrum provides information about quadratic phase coupling among harmonic components. Bispectrum estimation has been applied in diverse fields principally to obtain such information. Existing methods for bispectrum estimation are patterned after the conventional methods for power spectrum estimation which are known to possess certain limitations. The paper proposes a parametric approach to bispectrum estimation based on AR modeling of time series. The definition and properties of a parametric bispectrum estimator in the general ARMA case are stated. The third moment recursion equations that follow from the model assumptions for the proposed AR bispectrum estimator are presented. The estimates are derived using biased estimates of the third moments in these equations. Results from preliminary experiments suggest that the resulting method does possess certain attractive features when compared with existing methods.
Raghuveer M. Rao, Chrysostomos L. Nikias
ICASSP1
1983 A new class of high-resolution and robust multi-dimensional spectral estimation algorithms
abstract
A new class of multi-dimensional (m-D,m=3,4) spectral estimation algorithms is introduced based on the minimum variance representations (MVR) of m-D (m=3,4) data fields. These representations are defined in the framework of linear prediction where it is shown that they may be classed into general categories depending upon the geometry of the prediction space. For example, in the 3-D case it is shown that there are four possible models: causal, semicausal I, semicausal II, and non-causal. The m-D (m=3,4) model formalisms and their linear-predictive and spectral interpretations are derived. The admissibility conditions of the spectral density function are also discussed. To obtain high-resolution spectral estimates from finite length m-D (m=3,4) data fields, the models are fitted to the data optimally in the sense of minimizing the covariance recursion errors within the prediction space considered. Computer-simulated short data fields consisting of two travelling waves embedded in noise are employed to demonstrate experimentally that the class of algorithms developed in this paper improves on the standard techniques for high-resolution and robustness in the presence of nonstationarities, such as envelope modulation.
Chrysostomos L. Nikias, Raghuveer M. Rao
ICASSP2