VLDB 2026 Research / reviewers in the wild / expert
Sohail A. Dianat
dblp:45/8109 · also Soheil A. Dianat
· DBLP profile ↗
30ranked-venue papers
7as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought ReasoningabstractPrevalent text-to-video retrieval systems mainly adopt embedding models for feature extraction and compute cosine similarities for ranking.However, this design presents two limitations.Low-quality text-video data pairs could compromise the retrieval, yet are hard to identify and examine.Cosine similarity alone provides no explanation for the ranking results, limiting the interpretability.We ask that can we interpret the ranking results, so as to assess the retrieval models and examine the text-video data?This work proposes X-CoT, an explainable retrieval framework upon LLM CoT reasoning in place of the embedding model-based similarity ranking.We first expand the existing benchmarks with additional video annotations to support semantic understanding and reduce data bias.We also devise a retrieval CoT consisting of pairwise comparison steps, yielding detailed reasoning and complete ranking.X-CoT empirically improves the retrieval performance and produces detailed rationales.It also facilitates the model behavior and data quality analysis.Code and data are available at: github.com Prasanna Reddy Pulakurthi, Jiamian Wang, Majid Rabbani, Sohail A. Dianat, Raghuveer M. Rao, Zhiqiang Tao |
EMNLP | 4 |
| 2025 | MEPT: Mixture of Expert Prompt Tuning as a Manifold MapperabstractRunjia Zeng, Guangyan Sun, Qifan Wang, Tong Geng, Sohail Dianat, Xiaotian Han, Raghuveer Rao, Xueling Zhang, Cheng Han, Lifu Huang, Dongfang Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Runjia Zeng, Guangyan Sun, Qifan Wang 0001, Tong Geng, Sohail A. Dianat, Raghuveer M. Rao, Xueling Zhang, Cheng Han 0001, Lifu Huang, Dongfang Liu |
EMNLP | 5 |
| 2025 | Structured Policy Optimization: Enhance Large Vision-Language Model via Self-Referenced Dialogue
Can Qin, Yihao Feng, Zeyuan Chen 0001, Ran Xu 0001, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Zhiqiang Tao |
ICCV | 6 |
| 2025 | Shuffle PatchMix Augmentation with Confidence-Margin Weighted Pseudo-Labels for Enhanced Source-Free Domain AdaptationabstractThis work investigates Source-Free Domain Adaptation (SFDA), where a model adapts to a target domain without access to source data. A new augmentation technique, Shuffle PatchMix (SPM), and a novel reweighting strategy are introduced to enhance performance. SPM shuffles and blends image patches to generate diverse and challenging augmentations, while the reweighting strategy prioritizes reliable pseudo-labels to mitigate label noise. These techniques are particularly effective on smaller datasets like PACS, where overfitting and pseudo-label noise pose greater risks. State-of-the-art results are achieved on three major benchmarks: PACS, VisDA-C, and DomainNet-126. Notably, on PACS, improvements of 7.3% (79.4% to 86.7%) and 7.2% are observed in single-target and multi-target settings, respectively, while gains of 2.8% and 0.7% are attained on DomainNet-126 and VisDA-C. This combination of advanced augmentation and robust pseudo-label reweighting establishes a new benchmark for SFDA. The code is available at: https://github.com/PrasannaPulakurthi/SPM. Prasanna Reddy Pulakurthi, Majid Rabbani, Jamison Heard, Sohail A. Dianat, Celso de Melo, Raghuveer M. Rao |
ICIP | 4 |
| 2025 | Re-Imagining Multimodal Instruction Tuning: A Representation ViewabstractMultimodal instruction tuning has proven to be an effective strategy for achieving zero-shot generalization by fine-tuning pre-trained Large Multimodal Models (LMMs) with instruction-following data. However, as the scale of LMMs continues to grow, fully fine-tuning these models has become highly parameter-intensive. Although Parameter-Efficient Fine-Tuning (PEFT) methods have been introduced to reduce the number of tunable parameters, a significant performance gap remains compared to full fine-tuning. Furthermore, existing PEFT approaches are often highly parameterized, making them difficult to interpret and control. In light of this, we introduce Multimodal Representation Tuning (MRT), a novel approach that focuses on directly editing semantically rich multimodal representations to achieve strong performance and provide intuitive control over LMMs. Empirical results show that our method surpasses current state-of-the-art baselines with significant performance gains (e.g., 1580.40 MME score) while requiring substantially fewer tunable parameters (e.g., 0.03% parameters). Additionally, we conduct experiments on editing instrumental tokens within multimodal representations, demonstrating that direct manipulation of these representations enables simple yet effective control over network behavior. Yiyang Liu 0003, James Liang, Ruixiang Tang, Yugyung Lee, Majid Rabbani, Sohail A. Dianat, Raghuveer M. Rao, Lifu Huang, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001 |
ICLR | 6 |
| 2025 | Latent Chain-of-Thought for Visual ReasoningabstractChain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such as SFT, PPO, and GRPO may not generalize well across unseen reasoning tasks and heavily rely on a biased reward model. To address this challenge, we reformulate reasoning in LVLMs as posterior inference and propose a scalable training algorithm based on amortized variational inference. By leveraging diversity-seeking reinforcement learning algorithms, we introduce a novel sparse reward function for token-level learning signals that encourage diverse, high-likelihood latent CoT, overcoming deterministic sampling limitations and avoiding reward hacking. Additionally, we implement a Bayesian inference-scaling strategy that replaces costly Best-of-N and Beam Search with a marginal likelihood to efficiently rank optimal rationales and answers. We empirically demonstrate that the proposed method enhances the state-of-the-art LVLMs on four reasoning benchmarks, in terms of effectiveness, generalization, and interpretability. Hang Hua, Jiebo Luo 0001, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Zhiqiang Tao |
NeurIPS | 5 |
| 2024 | Text Is MASS: Modeling as Stochastic Embedding for Text-Video RetrievalabstractThe increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, relying on consistent embedding representations to compute similarity. However, the text content in existing datasets is generally short and concise, making it hard to fully describe the redundant semantics of a video. Correspondingly, a single text embedding may be less expressive to capture the video embedding and empower the retrieval. In this study, we propose a new stochastic text modeling method T-MASS, i.e., text is modeled as a stochastic embedding, to enrich text embedding with a flexible and re-silient semantic range, yielding a text mass. To be specific, we introduce a similarity-aware radius module to adapt the scale of the text mass upon the given text-video pairs. Plus, we design and develop a support text regularization to further control the text mass during the training. The inference pipeline is also tailored to fully exploit the text mass for accurate retrieval. Empirical evidence suggests that T-MASS not only effectively attracts relevant text-video pairs while distancing irrelevant ones, but also enables the de-termination of precise text embeddings for relevant pairs. Our experimental results show a substantial improvement of T-MASS over baseline (3% ~ 6.3% by R@1). Also, T-MASS achieves state-of-the-art performance on five bench-mark datasets, including MSRVTT, LSMDC, DiDeMo, VA-TEX, and Charades. Code and models are available here. Jiamian Wang, Pichao Wang, Dongfang Liu, Sohail A. Dianat, Raghuveer M. Rao, Majid Rabbani, Zhiqiang Tao |
CVPR | 5 |
| 2024 | AMD: Automatic Multi-step Distillation of Large-Scale Vision Models
Cheng Han 0001, Qifan Wang 0001, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Yi Fang 0008, Qiang Guan, Lifu Huang, Dongfang Liu |
ECCV (65) | 3 |
| 2024 | Enhancing GAN Performance Through Neural Architecture Search and Tensor DecompositionabstractGenerative Adversarial Networks (GANs) have emerged as a powerful tool for generating high-fidelity content. This paper presents a new training procedure that leverages Neural Architecture Search (NAS) to discover the optimal architecture for image generation while employing the Maximum Mean Discrepancy (MMD) repulsive loss for adversarial training. Moreover, the generator network is compressed using tensor decomposition to reduce its computational footprint and inference time while preserving its generative performance. Experimental results show improvements of 34% and 28% in the FID score on the CIFAR-10 and STL-10 datasets, respectively, with corresponding footprint reductions of 14× and 31× compared to the best FID score method reported in the literature. The implementation code is available at: https://github.com/PrasannaPulakurthi/MMD-AdversarialNAS. Prasanna Reddy Pulakurthi, Mahsa Mozaffari, Sohail A. Dianat, Majid Rabbani, Jamison Heard, Raghuveer M. Rao |
ICASSP | 3 |
| 2024 | Image Translation as Diffusion Visual ProgrammersabstractWe introduce the novel Diffusion Visual Programmer (DVP), a neuro-symbolic image translation framework. Our proposed DVP seamlessly embeds a condition-flexible diffusion model within the GPT architecture, orchestrating a coherent sequence of visual programs ($i.e.$, computer vision models) for various pro-symbolic steps, which span RoI identification, style transfer, and position manipulation, facilitating transparent and controllable image translation processes. Extensive experiments demonstrate DVP’s remarkable performance, surpassing concurrent arts. This success can be attributed to several key features of DVP: First, DVP achieves condition-flexible translation via instance normalization, enabling the model to eliminate sensitivity caused by the manual guidance and optimally focus on textual descriptions for high-quality content generation. Second, the frame work enhances in-context reasoning by deciphering intricate high-dimensional concepts in feature spaces into more accessible low-dimensional symbols ($e.g.$, [Prompt], [RoI object]), allowing for localized, context-free editing while maintaining overall coherence. Last but not least, DVP improves systemic controllability and explainability by offering explicit symbolic representations at each programming stage, empowering users to intuitively interpret and modify results. Our research marks a substantial step towards harmonizing artificial image translation processes with cognitive intelligence, promising broader applications. Cheng Han 0001, James Liang, Qifan Wang 0001, Majid Rabbani, Sohail A. Dianat, Raghuveer M. Rao, Ying Nian Wu, Dongfang Liu |
ICLR | 5 |
| 2024 | Prototypical Transformer As Unified Motion LearnersabstractIn this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoFormer seamlessly integrates prototype learning with Transformer by thoughtfully considering motion dynamics, introducing two innovative designs. First, Cross-Attention Prototyping discovers prototypes based on signature motion patterns, providing transparency in understanding motion scenes. Second, Latent Synchronization guides feature representation learning via prototypes, effectively mitigating the problem of motion uncertainty. Empirical results demonstrate that our approach achieves competitive performance on popular motion tasks such as optical flow and scene depth. Furthermore, it exhibits generality across various downstream tasks, including object tracking and video stabilization. Cheng Han 0001, Yawen Lu, James Liang, Zhiwen Cao, Qifan Wang 0001, Qiang Guan, Sohail A. Dianat, Raghuveer M. Rao, Tong Geng, Zhiqiang Tao, Dongfang Liu |
ICML | 8 |
| 2024 | Diffusion-Inspired Truncated Sampler for Text-Video RetrievalabstractPrevalent text-to-video retrieval methods represent multimodal text-video data in a joint embedding space, aiming at bridging the relevant text-video pairs and pulling away irrelevant ones. One main challenge in state-of-the-art retrieval methods lies in the modality gap, which stems from the substantial disparities between text and video and can persist in the joint space. In this work, we leverage the potential of Diffusion models to address the text-video modality gap by progressively aligning text and video embeddings in a unified space. However, we identify two key limitations of existing Diffusion models in retrieval tasks: The L2 loss does not fit the ranking problem inherent in text-video retrieval, and the generation quality heavily depends on the varied initial point drawn from the isotropic Gaussian, causing inaccurate retrieval. To this end, we introduce a new Diffusion-Inspired Truncated Sampler (DITS) that jointly performs progressive alignment and modality gap modeling in the joint embedding space. The key innovation of DITS is to leverage the inherent proximity of text and video embeddings, defining a truncated diffusion flow from the fixed text embedding to the video embedding, enhancing controllability compared to adopting the isotropic Gaussian. Moreover, DITS adopts the contrastive loss to jointly consider the relevant and irrelevant pairs, not only facilitating alignment but also yielding a discriminatively structured embedding. Experiments on five benchmark datasets suggest the state-of-the-art performance of DITS. We empirically find that DITS can also improve the structure of the CLIP embedding space. Code is available at https://github.com/Jiamian- Wang/DITS-text-video-retrieval Jiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Zhiqiang Tao |
NeurIPS | 5 |
| 2014 | Multispectral Image Denoising With Optimized Vector Bilateral FilterabstractVector bilateral filtering has been shown to provide good tradeoff between noise removal and edge degradation when applied to multispectral/hyperspectral image denoising. It has also been demonstrated to provide dynamic range enhancement of bands that have impaired signal to noise ratios (SNRs). Typical vector bilateral filtering described in the literature does not use parameters satisfying optimality criteria. We introduce an approach for selection of the parameters of a vector bilateral filter through an optimization procedure rather than by ad hoc means. The approach is based on posing the filtering problem as one of nonlinear estimation and minimization of the Stein's unbiased risk estimate of this nonlinear estimator. Along the way, we provide a plausibility argument through an analytical example as to why vector bilateral filtering outperforms bandwise 2D bilateral filtering in enhancing SNR. Experimental results show that the optimized vector bilateral filter provides improved denoising performance on multispectral images when compared with several other approaches. Honghong Peng, Raghuveer M. Rao, Sohail A. Dianat |
IEEE Trans. Image Process. | 3 |
| 2012 | Optimized vector bilateral filter for multispectral image denoisingabstractVector bilateral filtering has been shown to provide several advantages in processing hyperspectral images such as good noise removal while minimizing edge degradation, and dynamic range enhancement of bands with impaired signal to noise ratios. This paper introduces an approach for selection of the parameters of a vector bilateral filter through an optimization procedure rather than by ad hoc means. The approach is based on posing the filtering problem as one of nonlinear estimation and minimizing the Stein's unbiased risk estimate (SURE) of this nonlinear estimator. Experimental results show that the optimized vector bilateral filter provides improved denoising performance on multispectral images when compared to several other approaches. Honghong Peng, Raghuveer M. Rao, Sohail A. Dianat |
ICIP | 3 |
| 2012 | Nonnegative matrix factorization with deterministic annealing for unsupervised unmixing of hyperspectral imageryabstractNon-negative matrix factorization (NMF) technique and its extensions were developed to find part based, linear representations of non-negative multivariate data. They have been shown to provide more interpretable results with realistic non-negative constrain in unsupervised learning applications such as hyperspectral imagery unmixing, image feature extraction, and data mining. This paper extends the NMF method by incorporating deterministic annealing optimization procedure, which will help solve the non-convexity problem in NMF and provide a better choice of sparseness constrain. The approach is based on replacing the difficult non-convex optimization problem of NMF with an easier one by adding an auxiliary convex entropy constrain term and solving this first. Experiment results with hyperspectral unmixing application show that the proposed technique provides improved unmixing performance compared to other state-of-the-art methods. Honghong Peng, Raghuveer M. Rao, Sohail A. Dianat |
ICIP | 3 |
| 2011 | Semi-automatic 3-D segmentation of Computed Tomographic imagery by iterative gradient-driven volume growingabstractWe propose a novel gradient driven methodology for three dimensional (3-D) segmentation of Computed Tomographic (CT) imagery. Our approach begins interactively where-in a user marks a set of voxels within the cross-section of a Sub-Volume Of Interest (SVOI), using a single slice of the CT volume. Subsequently, a 3-D gradient detection scheme is utilized to determine the radiodensity variations across the volume. The resultant gradient information is employed in an iterative volume growing procedure, which is initiated at voxels with small gradient magnitudes adjoining the user-selected voxels and culminates at voxels with large gradient magnitudes, to arrive at the final 3-D segmentation result of the SVOI. The aforementioned method was tested on multiple studies and the results show favorable performance against a state-of-the-art technique. Sreenath Rao Vantaram, Eli Saber, Sohail A. Dianat, Vishwas Abhyankar |
ICIP | 3 |
| 2010 | Point-and-click region based method for color editing and control for digital color printersabstractTo meet ever changing customer demand, the commercial printing industry requires the capability of producing colors accurately and consistently. In this presentation, we propose to use a region growing process with a spatial (i.e., 2-D) color gradient matrix in RGB or L*a*b* or CMYK space to pick a region/object of interest. A grid with Laplacian edge map is constructed for the whole image a priori using the spatial color gradient matrix. The Laplacian edge map (a template), if needed, is shown on the GUI to guide the user to point at his or her object of interest. After clicking the mouse, the algorithm will grow the region until it coincides with the Laplacian edge map. A color constrained cost function is introduced to prevent leakage through mild transitions between different objects. Once the object is selected, a customized rendering LUT is applied to improve the color quality of the object. We also present control-based techniques to create inverse transformation that produce more accurate CMYK values. Control based techniques minimize the rendering errors normally occur during inversion and gamut mapping process. Siyu Zhu 0005, Sohail A. Dianat, Lalit K. Mestha |
ICIP | 2 |
| 2009 | A multidimensional smoothing algorithm with applications to digital color printer calibrationabstractMultidimensional data smoothing has applications in signal and image processing, control, and other engineering disciplines. The two important data fitting criteria are the accuracy of the fit and the smoothness of the fit. In this paper, we discuss a tensor based algorithm for multidimensional smooth curve fitting where the cost function is a combination of two terms: one dealing with accuracy and the other one with the smoothness of the solution. We apply the proposed algorithm to digital color printer calibration in 1-d, 2-d and 3-d calibration techniques. Sohail A. Dianat, Bruce Brewington, Lalit K. Mestha |
ICIP | 1 |
| 2009 | An adaptive and progressive approach for efficient Gradient-based multiresolution color image segmentationabstractWe propose an image segmentation methodology which exploits gradient information in a multiresolution framework. The proposed algorithm commences with a wavelet decomposition procedure to obtain a pyramidal representation of the input image, accompanied by an adaptive threshold generation scheme required for segregating regions of varying gradient densities. At low (coarse) resolution levels, progressive region growth, texture characterization, and region merging modules are integrated together to provide interim segmentations. These interim results are transferred from one resolution level to another as a-priori information, until the final result at the highest (original) resolution is achieved. Performance evaluation on several hundred images demonstrates that our algorithm computationally outperforms various published techniques, with superior segmentation quality. Sreenath Rao Vantaram, Eli Saber, Sohail A. Dianat, Mark Q. Shaw, Ranjit Bhaskar |
ICIP | 3 |
| 2007 | SNR Estimation in Nakagami Fading Channels with Arbitrary ConstellationabstractIn this paper, a novel technique is proposed to estimate the average signal-to-noise ratio (SNR) for Nakagami-m fading channels with an arbitrary signal constellation. The Nakagami-m distribution is a good fit to empirical fading data obtained from radio communication channels. The proposed algorithm uses absolute second and fourth moments of the envelope of the received signal over a block of data as the sufficient statistics. The estimator is blind and no training sequence is used. It is also independent of signal constellation. Sohail A. Dianat |
ICASSP (2) | 1 |
| 2006 | Dynamic Optimization Algorithm For Generating Inverse Printer Map With Reduced MeasurementsabstractFor a color printer to attain good color rendering quality the image output terminal must be capable of producing the desired tone, i.e., the solidness, of each of primary color separations as requested. Calibration is a major task in providing high quality prints. A model (characterization) of the color printer is a first step for building profiles. Most of the methods used to solve the calibration problem utilize, in one way or another, a printer inverse map. Once the inverse map is constructed, the input image will be processed by the inverse map to obtain the correct CMY values that can produce the desired Lab values. A major drawback in this type of calibration is the need to measure N times N times N color patches each time the printer is calibrated. We propose an alternative approach, where few critical color patches are measured and the rest of the points in the forward printer map are obtained using interpolation Sohail A. Dianat, Lalit K. Mestha, Athimootitil Mathew |
ICASSP (3) | 1 |
| 2002 | Blind adaptive decision feedback multiuser detector for DS-CDMA with power estimationabstractIn this paper we propose a new blind adaptive decision-feedback multiuser detector (BADFP) for synchronous direct-sequence code-division multiple-access (DS-CDMA) systems. The BADFP algorithm is a stochastic adaptive gradient type algorithm that incorporates power estimation of the user of interest to enhance the performance of the detector. It is assumed that full power control is in place and all the users are equally powered. The proposed algorithm requires the knowledge of the signature sequence of the desired user. The simulation results show that the performance of the BADFP algorithm is comparable with the ideal blind linear minimum output energy (MOE) detector. Rajesh Narasimha, Sohail A. Dianat |
VTC Spring | 2 |
| 1992 | Cross-bispectrum computation for multichannel quadratic phase coupling estimationabstractA two-channel, quadratic, nonlinear process driven by sinusoidal signals is considered. Equations for estimating the parameters of the process are derived from colored Gaussian noise contaminated observations of the process. The derivation is based on the bispectrum and cross-power spectrum of the two channels. The resulting equations are nonlinear and iterative techniques are needed to solve for the parameters.> Raghuveer M. Rao, Sohail A. Dianat |
ICASSP | 2 |
| 1991 | A non-linear predictor for differential pulse-code encoder (DPCM) using artificial neural networksabstractA nonlinear predictor is designed for a DPCM encoder using artificial neural networks (ANN). The predictor is based on a multilayer perceptron with three input nodes, 30 hidden nodes and one output node. The back-propagation learning algorithm is used for the training of the network. Simulation results are presented to evaluate and compare the performance of the neural net based predictor (nonlinear) with that of an optimized linear predictor. Success in the use of the nonlinear predictor is demonstrated through the reduction in the entropy of the differential error signal as compared to that of a linear predictor. Also it is shown that the ANN predictor is much more robust for encoding noisy images compared to that of a linear predictor.> Sohail A. Dianat, Nasser M. Nasrabadi, S. Venkataraman |
ICASSP | 1 |
| 1990 | Fast algorithms for bispectral reconstruction of two-dimensional signalsabstractAlgorithms are developed to reconstruct the phase and magnitude of the Fourier transform of a two-dimensional, discrete, deterministic signal from samples of its bispectrum B( omega , lambda ). The samples are derived from the plane omega - lambda . The techniques involve finding the DFTs (discrete Fourier transforms) of the phase and magnitude of the bispectrum samples, and from them, recovering samples of the phase and magnitude of the Fourier transform of the signal at select frequencies.> Sohail A. Dianat, Raghuveer M. Rao |
ICASSP | 1 |
| 1989 | Bispectrum phase transformation for non-minimum phase signal reconstructionabstractThe authors present a Fourier-series-based approach for recovering the phase of a signal from samples of its bispectrum along the omega /sub 1/= omega /sub 2/ line in the bispectrum frequency plane. The Fourier series coefficients of the signal phase are recovered from the Fourier series of the bispectrum phase. The implementation of the method is simplified by using the FFT (fast Fourier transform) to compute the Fourier series coefficients. An important advantage of the approach is that while using a finite number of samples of the bispectrum phase, it provides estimates for the phase of the signal at all frequencies. Also there are computational advantages in that the FFT can be used and it may not be necessary to consider a large number of Fourier series coefficients.> Sohail A. Dianat, Raghuveer M. Rao |
ICASSP | 1 |
| 1989 | Polyspectral factorization: necessary and sufficient condition for finite extent cumulant sequencesabstractThe authors provide necessary and sufficient conditions for bispectral and trispectral factorization for processes with finite-extent cumulant sequences. These conditions are derived entirely in terms of the cumulant sequences. They use the fact that for a finite-extent cumulant sequence factorability is equivalent to finding a finite-order moving-average process with an identical cumulant sequence. In principle the results can be extended to polyspectra of even higher orders. An interesting result of the investigation is that there exist processes generated by nonlinear mechanisms that are factorable.> Sohail A. Dianat, Raghuveer M. Rao |
ICASSP | 1 |
| 1985 | Adaptive spectral estimation by the conjugate gradient methodabstractPisarenko's harmonic retrieval method provides unbiased spectral estimates of a signal consisting of the sum of n sinusoidal signals in additive white noise. Recently, a few adaptive versions based on Pisarenko's method have been developed. In this paper we propose an alternative technique for adaptive spectral estimation based on the method of conjugate gradient, which is used for iteratively finding the eigenvector corresponding to the minimum eigenvalue of the covariance matrix. The new method is an exact adaptive version of Pisarenko's method and converges in finite steps for any initial guesses. Simulations have been performed to compare the new method with the existing ones. It is seen that the new method improves the CPU time by a factor of 40. Also, the technique performs very well for both narrow band and wide band signals. Huanqun Chen, Tapan K. Sarkar, Sohail A. Dianat, John D. Brule |
ICASSP | 3 |
| 1985 | Deconvolution by the conjugate gradient methodabstractSince it is practically difficult to generate and propagate an impulse, often a system is excited by a narrow time domain pulse. The output is recorded and then a numerical deconvolution is often done to extract the impulse response of the object. Classically, the fast Fourier transform technique has been applied with much success to the above deconvolution problem. However, when the signal to noise ratio becomes small, sometimes one encounters instability with the FFT approach. In this paper, the method of conjugate gradient is applied to the deconvolution problem entirely in the time domain. The method converges for any initial guess in a finite number of steps. Also for the application of the conjugate gradient method the time samples need not be uniform like FFT. Computed impulse response utilizing this technique has been presented for measured incident and scattered fields from a sphere and a cylinder. Tapan K. Sarkar, Fung I. Tseng, Sohail A. Dianat, Bruce Z. Hollmann |
ICASSP | 3 |
| 1983 | A finite step adaptive implementation of the Pisarenko's harmonic retrieval method in colored noiseabstractAn adaptive spectral analysis technique is presented for estimating complex frequencies in colored noise. It is assumed that the noise covariance matrix of the colored noise is known. The method presented in this paper is similar to Thompson's technique of an on line estimation of the eigenvector of the covariance matrix corresponding to the minimum eigenvalue, without explicitly evaluating the covariance matrix. The method of conjugate gradient has been utilized to obtain the eigenvector corresponding to the minimum eigenvalue. The advantages of this technique over the method of steepest descent is that it is a finite step iterative method and secondly there is no arbitrary constants in the expression which dictates the overall rate of convergence. In the proposed method, the spread of the eigenvalues has no significant effect on the overall rate of convergence. The disadvantage of this technique is that one has to store a matrix of data instead of one row only, as is conventionally done. The proposed method however yields unbiased estimates for the frequencies in colored noise. Tapan K. Sarkar, Sohail A. Dianat, S. M. Rao |
ICASSP | 2 |