VLDB 2026 Research / reviewers in the wild / expert
Hassan Mansour
dblp:12/4863
· DBLP profile ↗
76ranked-venue papers
15as first author
23since 2021 · last 2026
0000-0002-1667-9885ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 65 · 14 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1Computer networks · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Recovering Pulse Waves From Video Using Deep Unrolling and Deep Equilibrium ModelsabstractCamera-based contactless monitoring of vital signs, also known as imaging photoplethysmography (iPPG), has seen applications in driver-monitoring, perfusion assessment, affective computing, and more. iPPG involves sensing the underlying cardiac pulse from video of the skin and estimating vital signs such as the pulse rate or a full pulse waveform. Some previous iPPG methods impose model-based sparse priors on the pulse signals and use iterative optimization for pulse wave recovery, while others use end-to-end black-box deep learning methods. In contrast, we introduce methods that combine signal processing and deep learning methods in an inverse problem framework. Our methods estimate the underlying pulse signal, pulse rate, and pulse rate variability from facial video by learning deep-network-based denoising operators that leverage deep algorithm unfolding and deep equilibrium models. Experiments show that our methods can denoise an acquired signal from the face and infer the correct underlying pulse rate and pulse rate variability, achieving pulse rate estimation performance consistent with the state-of-the-art on well-known benchmarks, all with less than one-fifth the number of learnable parameters as the closest competing method. Vineet R. Shenoy, Suhas Lohit, Hassan Mansour, Rama Chellappa, Tim K. Marks |
IEEE Trans. Image Process. | 3 |
| 2025 | Enabling DMG Wi-Fi Sensing in Data Transmission Intervals by Exploiting Beam Training CodebookabstractThis paper addresses the integration of millimeter-wave (mmWave) Wi-Fi communication and sensing during data transmission intervals (DTIs). We leverage prior knowledge from codebook beam training conducted during preceding beacon transmission intervals (BTIs) and association beamforming training (A-BFT) intervals to design a transceiver array response that meets both requirements on downlink communication SNR and targeted sensing area. By formulating it as a first-order array response optimization with constraints on power, codebook, communication SNR, and limited RF chains, this paper introduces a two-stage solution. First, we introduce a two-way communication-sensing matching pursuit to determine a set of codewords that prioritize the communication SNR constraint. Then, using the selected codewords, we employ an alternating minimization over an auxiliary phase term and beamforming weights to further minimize an array-response distance loss. Numerical results validate the effectiveness of the proposed DMG beamforming design over baseline methods. Kareem M. Attiah, Pu Wang 0004, Hassan Mansour, Toshiaki Koike-Akino, Petros Boufounos |
ICASSP | 3 |
| 2025 | Doppler Single-Photon LidarabstractSingle-photon lidar (SPL) can achieve high-accuracy, lowlight ranging; however, velocity estimation typically requires regression over multiple distance measurements. Here, we introduce Doppler SPL, which enables joint instantaneous velocity and range estimation. First, we derive a measurement model for SPL, showing that a target moving at a constant velocity introduces a Doppler shift into the sequence of photon detection times. We then introduce estimators for range and velocity based on Fourier analysis of the detection time sequence. Simulations show improved accuracy of our method over baseline approaches, and we further validate our approach on experimental SPL data for a moving target. Ruangrawee Kitichotkul, Joshua Rapp, Yanting Ma, Hassan Mansour |
ICASSP | 4 |
| 2025 | Indoor Airflow Imaging Using Physics-Informed Schlieren TomographyabstractRemote temperature sensing of volumetric flows has a variety of applications, such as promoting thermal comfort, heat dissipation, or data center cooling. The emergence of background-oriented schlieren (BOS) imaging in recent years has enabled transparent flow visualization at minor costs. In this paper, we develop a framework for non-invasive volumetric indoor airflow estimation from a single viewpoint using BOS measurements and physics-informed reconstruction. Our framework utilizes a light projector that projects a pattern onto a target back wall and a camera that observes small distortions in the light pattern due to the change in the refractive index of the air as a result of the temperature variation. While the single-view BOS tomography problem is severely ill-posed, we regularize the reconstruction using a physics-informed neural network (PINN) that ensures that the reconstructed airflow is consistent with the coupled Boussinesq approximation of the incompressible Navier– Stokes and the heat transfer equations. Arjun Teh, Wael H. Ali, Joshua Rapp, Hassan Mansour |
ICASSP | 4 |
| 2025 | Multi-Band Wi-Fi Neural Dynamic FusionabstractWi-Fi channel measurements across different bands, e.g., sub-7-GHz and 60-GHz bands, are asynchronous due to the uncoordinated nature of distinct standards protocols, e.g., 802.11ac/ax/be and 802.11ad/ay. Multi-band Wi-Fi fusion has been considered before on a frame-to-frame basis for simple classification tasks, which does not require fine-time-scale alignment. In contrast, this paper considers asynchronous sequence-to-sequence fusion between sub-7-GHz channel state information (CSI) and 60-GHz beam signal-to-noise-ratio (SNR)s for more challenging tasks, such as continuous coordinate estimation. To handle the timing disparity between asynchronous multi-band Wi-Fi channel measurements, this paper proposes a multi-band neural dynamic fusion (NDF) framework. This framework uses separate encoders to embed the multi-band Wi-Fi measurement sequences to separate initial latent conditions. Using a continuous-time ordinary differential equation (ODE) modeling, these initial latent conditions are propagated to the respective latent states of the multi-band channel measurements at the same time instances for a latent alignment and a post-ODE fusion, and at their original time instances for measurement reconstruction. We derive a customized loss function based on the variational evidence lower bound (ELBO) that balances between the multi-band measurement reconstruction and continuous coordinate estimation. We evaluate the NDF framework using an in-house multi-band Wi-Fi testbed and demonstrate substantial performance improvements over a comprehensive list of single-band and multi-band baseline methods. Sorachi Kato, Pu Wang 0004, Toshiaki Koike-Akino, Takuya Fujihashi, Hassan Mansour, Petros Boufounos |
IEEE Trans. Wirel. Commun. | 5 |
| 2024 | Tracking Beyond the Unambiguous Range with Modulo Single-Photon LidarabstractIn single photon lidar (SPL), the laser repetition rate sets the maximum distance that can be recovered unambiguously. Conventional SPL extends this maximum recordable depth by reducing the repetition rate; however, the slower acquisition speed limits the number of received photons, which may be insufficient to track fast-moving objects. Inspired by recent successes in modulo sensing, we leverage the smoothness of typical trajectories to achieve long-range tracking beyond the unambiguous range. Although SPL naturally acquires modulo time-of-flight measurements, it introduces several challenges—including random sampling times, multiple noise sources, and absolute distance uncertainty—that are not addressed by the current modulo sensing literature. Hence, we propose an interpolation and denoising method that operates directly over the modulo samples. We further disambiguate the absolute distance based on the changing reflectivity fall-off. Monte Carlo simulations considering realistic trajectories under practical conditions show that, when properly unwrapped, the normalized mean squared error of our depth estimate decreases by over 20 dB with respect to a lidar setup whose repetition period leads to no ambiguity. Samuel Fernández-Menduiña, Joshua Rapp, Hassan Mansour, M. Greiff, Kieran Parsons |
ICASSP | 3 |
| 2024 | Object Trajectory Estimation with Multi-Band Wi-Fi Neural Dynamic FusionabstractIn contrast to existing multi-band Wi-Fi fusion in a frame-to-frame basis for simple classification, this paper considers asynchronous sequence-to-sequence fusion between sub-7GHz channel state information (CSI) and 60GHz beam SNR for more challenging downstream tasks such as continuous regression. To handle the timing disparity between the two channel measurements, we extend our recently proposed dual-decoder neural dynamic (DDND) framework with latent ordinary differential equations (ODEs), align the distinct latent dynamic states at the same time instances, and introduce a post-ODE fusion framework. The resulting neural dynamic fusion (NDF) framework is trained in an end-to-end fashion with a modified variational autoencoder loss function. Evaluation over a newly collected in-house multi-band Wi-Fi dataset shows the advantage of the proposed NDF method over frame-based and DDND methods. Sorachi Kato, Pu Wang 0004, Toshiaki Koike-Akino, Takuya Fujihashi, Hassan Mansour, Petros Boufounos |
ICASSP | 5 |
| 2024 | Single-Pixel Imaging Of Dynamic Flows Using Neural Ode RegularizationabstractSingle-pixel imaging is an efficient image acquisition process where light from a target scene is passed through a spatial light modulator and then projected onto a single photodiode with a high temporal acquisition rate. The scene reconstruction is achieved using computational methods that leverage prior assumptions on the scene structure. In this paper, we propose to model the structure of a dynamic spatio-temporal scene using a reduced-order model that is learned from training data examples. Specifically, by combining single-pixel imaging methods with a reduced-order model prior implemented as a neural ordinary differential equation, image sequence reconstruction can be accomplished with significantly reduced data requirements while maintaining performance levels on par with leading methods. We demonstrate superior reconstruction at low sampling rates for simulated trajectories governed by Burgers’ equation and turbulent plumes emulating gas leaks. Aleksei Sholokhov, Joshua Rapp, Saleh Nabi, Steven L. Brunton, J. Nathan Kutz, Hassan Mansour |
ICASSP | 6 |
| 2023 | Deep Proximal Gradient Method for Learned Convex RegularizersabstractWe consider the problem of simultaneously learning a convex penalty function and its proximity operator for image reconstruction from incomplete measurements. Our goal is to apply Accelerated Proximal Gradient Method (APGM) using a learned proximity operator in place of the true proximity operator of the learned penalty function. Starting from a Gaussian image denoiser, we learn an associated penalty function and its proximity operator. The learned penalty function offers provable reconstruction guarantees, whereas access to its proximity operator presents the opportunity to achieve APGM convergence rates, which are faster than those of subgradient descent approaches. Aaron Berk, Yanting Ma, Petros Boufounos, Pu Wang 0004, Hassan Mansour |
ICASSP | 5 |
| 2023 | Phase Unwrapping in Correlated Noise for FMCW Lidar Depth EstimationabstractIn frequency-modulated continuous-wave (FMCW) lidar, the distance to an illuminated target is proportional to the beat frequency of the interference signal. Laser phase noise often limits the range accuracy of FMCW lidar, and existing frequency estimation methods make overly simplistic assumptions about the noise model. In this work, we propose an algorithm that performs frequency estimation via phase unwrapping by explicitly accounting for correlations in the phase noise. Given a candidate frequency, we approximately recover the maximum likelihood unwrapping sequence using the Viterbi algorithm and the phase noise statistics. The algorithm then alternates between unwrapping and frequency estimate refinement until convergence. Compared to state-of-the-art alternatives, our algorithm consistently achieves superior performance at long range or with large-linewidth lasers when the signal-to-noise ratio is sufficiently high. A. Ulvog, Joshua Rapp, Toshiaki Koike-Akino, Hassan Mansour, Petros Boufounos, Kieran Parsons |
ICASSP | 4 |
| 2023 | Deep Born Operator Learning for Reflection Tomographic ImagingabstractRecent developments in wave-based sensor technologies, such as ground penetrating radar (GPR), provide new opportunities for accurate imaging of underground scenes. Given measurements of the scattered electromagnetic wavefield, the goal is to estimate the spatial distribution of the permittivity of the underground scenes. However, such problems are highly ill-posed, difficult to formulate, and computationally expensive. In this paper, we propose a physics-inspired machine learning-based method to learn the wave-matter interaction under the GPR setting. The learned forward model is combined with a learned signal prior to recover the permittivity distribution of the unknown underground scenes. We test our approach on a dataset of 400 permittivity maps with a three-layer background, which is challenging to solve using existing methods. We demonstrate via numerical simulation that our method achieves a 50% improvement in mean squared error over benchmark machine learning-based solvers for reconstructing layered underground scenes. Yanting Ma, Petros Boufounos, Saleh Nabi, Hassan Mansour |
ICASSP | 5 |
| 2023 | Unrolled iPPG: Video Heart Rate Estimation via Unrolling Proximal Gradient DescentabstractImaging photoplethysmography (iPPG) is the process of estimating a person’s heart rate from video. In this work, we propose Unrolled iPPG, in which we integrate iterative optimization updates with deep learning-based signal priors to estimate the pulse waveform and heart rate from facial videos. We model the signal extracted from video as the sum of an underlying pulse signal and noise, but instead of explicitly imposing a handcrafted prior (e.g., sparsity in the frequency domain) on the signal, we learn priors on the signal and noise using neural networks. We solve for the underlying pulse signal by unrolling proximal gradient descent; the algorithm alternates between gradient descent steps and application of learned denoisers, which replace handcrafted priors and their proximal operators. Using this method, we achieve state-of-the-art heart rate estimation on the challenging MMSE-HR dataset. Vineet R. Shenoy, Tim K. Marks, Hassan Mansour, Suhas Lohit |
ICIP | 3 |
| 2022 | Learning Occlusion-Aware Dense Correspondences for Multi-Modal ImagesabstractWe introduce a scalable multi-modal approach to learn dense, i.e., pixel-level, correspondences and occlusion maps, between images in a video sequence. The problems of finding dense correspondences and occlusion maps are fundamental in computer vision. In this work we jointly train a deep network to tackle both, with a shared feature extraction stage. We use depth and color images with ground truth optical flow and occlusion maps to train the network end-to-end. From the multi-modal input, the network learns to estimate occlusion maps, optical flows, and a correspondence embedding providing a meaningful latent feature space. We evaluate the performance on a dataset of images derived from synthetic characters, and perform a thorough ablation study to demonstrate that the proposed components of our architecture combine to achieve the lowest correspondence error. The scalability of our proposed method comes from the ability to incorporate additional modalities, e.g., infrared images. Ryosuke Shimoya, Takashi Morimoto, Jeroen van Baar, Petros Boufounos, Yanting Ma, Hassan Mansour |
AVSS | 6 |
| 2022 | Distributed Radar Autofocus Imaging Using Deep PriorsabstractAntenna position ambiguity is a common problem that affects radar imaging systems that are mounted on mobile platforms. Existing approaches that aim to recover a sharp radar image despite this ambiguity aim to estimate the shift in the antenna position by modeling the radar scene as a sparse image with a small number of targets using explicit analytical models for the statistical distribution of the targets in a radar image. The radar imaging problem is then solved by alternating between estimating the radar image, followed by estimating the shift in the antenna positions, until convergence is reached. While such approaches have shown tremendous success, they still struggle to recover the true target positions and may arrive at incorrect local optima when the measurement noise level is high. In this work, we develop a data-driven learning-based strategy for modeling the image of the radar scene instead of relying on explicit analytical models. We adopt a residual Unet architecture of a neural network to act as a denoising operator which takes a backprojected radar image as input and outputs a true target image. While deep denoisers may generally result in unstable iterative algorithms, we introduce a simple filtering step that suppresses noise belonging to the null space of the radar operator from the iterates to stabilize the iterative procedure. We evaluate the effectiveness of our solution using simulated numerical experiments and demonstrate its superiority over the analytic signal prior. Hassan Mansour, Suhas Lohit, Petros Boufounos |
ICIP | 1 |
| 2022 | Maximum Likelihood Surface Profilometry Via Optical coherence TomographyabstractOptical coherence tomography (OCT) using Fourier domain processing can resolve micrometer-scale depth information. However, the conventional volumetric reconstruction approach is unnecessary for opaque samples with only one reflector per lateral position, and the required sample interpolation degrades performance. In this paper, we show that surface depth profilometery with a Fourier-domain OCT system simplifies to a sinusoidal parameter estimation problem. We derive approximate maximum likelihood estimators for the sample depth and reflectivity, which can easily be computed by backprojecting the data without interpolating. Iterative refinement further improves results at high signal-to-noise ratio (SNR). We demonstrate the performance of the technique compared to the conventional Fourier transform approach on both simulated and experimental data collected with a spectral-domain OCT system. Our results show that maximum likelihood profilometry is fast and more robust to noise than the Fourier approaches at moderate SNR. Joshua Rapp, Hassan Mansour, Petros Boufounos, Philip V. Orlik, Toshiaki Koike-Akino, Kieran Parsons |
ICIP | 2 |
| 2022 | Fast and High-Quality Blind Multi-Spectral Image PansharpeningabstractBlind pansharpening addresses the problem of generating a high spatial-resolution multi-spectral (HRMS) image given a low spatial-resolution multi-spectral (LRMS) image with the guidance of its associated spatially misaligned high spatial-resolution panchromatic (PAN) image without parametric side information. In this article, we propose a fast approach to blind pansharpening and achieve the state-of-the-art image reconstruction quality. Typical blind pansharpening algorithms are often computationally intensive since the blur kernel and the target HRMS image are often computed using iterative solvers and in an alternating fashion. To achieve fast blind pansharpening, we decouple the solution of the blur kernel and of the HRMS image. First, we estimate the blur kernel by computing the kernel coefficients with minimum total generalized variation that blur a downsampled version of the PAN image to approximate a linear combination of the LRMS image channels. Then, we estimate each channel of the HRMS image using local Laplacian prior (LLP) to regularize the relationship between each HRMS channel and the PAN image. Solving the HRMS image is accelerated by both parallelizing across the channels and by fast numerical algorithms for each channel. Due to the fast scheme and the powerful priors we used on the blur kernel coefficients (total generalized variation) and on the cross-channel relationship (LLP), numerical experiments demonstrate that our algorithm outperforms the state-of-the-art model-based counterparts in terms of both computational time and reconstruction quality of the HRMS images. Lantao Yu, Dehong Liu, Hassan Mansour, Petros Boufounos |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Near-Infrared Imaging Photoplethysmography During DrivingabstractImaging photoplethysmography (iPPG) could greatly improve driver safety systems by enabling capabilities ranging from identifying driver fatigue to unobtrusive early heart failure detection. Unfortunately, the driving context poses unique challenges to iPPG, including illumination and motion. First, drastic illumination variations present during driving can overwhelm the small intensity-based iPPG signals. Second, significant driver head motion during driving, as well as camera motion (e.g., vibration) make it challenging to recover iPPG signals. To address these two challenges, we present two innovations. First, we demonstrate that we can reduce most outside light variations using narrow-band near-infrared (NIR) video recordings and obtain reliable heart rate estimates. Second, we present a novel optimization algorithm, which we call AutoSparsePPG, that leverages the quasi-periodicity of iPPG signals and achieves better performance than the state-of-the-art methods. In addition, we release the first publicly available driving dataset that contains both NIR and RGB video recordings of a passenger’s face with simultaneous ground truth pulse oximeter recordings. Ewa Magdalena Nowara, Tim K. Marks, Hassan Mansour, Ashok Veeraraghavan |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Extended Object Tracking with Spatial Model Adaptation Using Automotive Radar
Pu Wang 0004, Karl Berntorp, Hassan Mansour, Petros Boufounos, Philip V. Orlik |
FUSION | 4 |
| 2021 | A Consensus Equilibrium Solution For Deep Image Prior Powered By RedabstractRecent advances in solving imaging inverse problems have witnessed the combination of deep learning models with classical image models for better signal representation. One such approach, DeepRED, combines the deep image prior (DIP) with the regularization by denoising (RED) framework to boost the performance of image deblurring and super resolution tasks. In this paper, we formulate DeepRED as a consensus equilibrium problem and set up a fixed-point algorithm for solving the equilibrium equations. We also derive sufficient conditions that the DIP generative prior should satisfy to ensure that the corresponding fixed-point operator is non-expansive. We then demonstrate that the fixed-point algorithm that solves the CE equations results in improved image reconstruction quality in a deblurring setting compared to state-of-the-art methods. Rakib Hyder, Hassan Mansour, Yanting Ma, Petros Boufounos, Pu Wang 0004 |
ICASSP | 2 |
| 2021 | Multiview Sensing with Unknown Permutations: an Optimal Transport ApproachabstractIn several applications, including imaging of deformable objects while in motion, simultaneous localization and mapping, and unlabeled sensing, we encounter the problem of recovering a signal that is measured subject to unknown permutations. In this paper we take a fresh look at this problem through the lens of optimal transport (OT). In particular, we recognize that in most practical applications the unknown permutations are not arbitrary but some are more likely to occur than others. We exploit this by introducing a regularization function that promotes the more likely permutations in the solution. We show that, even though the general problem is not convex, an appropriate relaxation of the resulting regularized problem allows us to exploit the well-developed machinery of OT and develop a tractable algorithm. Yanting Ma, Petros Boufounos, Hassan Mansour, Shuchin Aeron |
ICASSP | 3 |
| 2021 | Extended Object Tracking With Automotive Radar Using B-Spline Chained Ellipses ModelabstractThis paper introduces a B-spline chained ellipses model representation for extended object tracking (EOT) using high-resolution automotive radar measurements. With offline automotive radar training datasets, the proposed model parameters are learned using the expectation-maximization (EM) algorithm. Then the probabilistic multi-hypothesis tracking (PMHT) along with the unscented transform (UT) is proposed to deal with the nonlinear forward-warping coordinate transformation, the measurement-to-ellipsis association, and the state update step. Numerical validation is provided to verify the effectiveness of the proposed EOT framework with automotive radar measurements. Pu Wang 0004, Karl Berntorp, Hassan Mansour, Petros Boufounos, Philip V. Orlik |
ICASSP | 4 |
| 2021 | Turnip: Time-Series U-Net With Recurrence For NIR Imaging PPGabstractImaging photoplethysmography (iPPG) is the process of estimating the waveform of a person’s pulse by processing a video of their face to detect minute color or intensity changes in the skin. Typically, iPPG methods use three-channel RGB video to address challenges due to motion. In situations such as driving, however, illumination in the visible spectrum is often quickly varying (e.g., daytime driving through shadows of trees and buildings) or insufficient (e.g., night driving). In such cases, a practical alternative is to use active illumination and bandpass-filtering from a monochromatic near-infrared (NIR) light source and camera. Contrary to learning-based iPPG solutions designed for multi-channel RGB, previous work in single-channel NIR iPPG has been based on hand-crafted models (with only a few manually tuned parameters), exploiting the sparsity of the PPG signal in the frequency domain. In contrast, we propose a modular framework for iPPG estimation of the heartbeat signal, in which the first module extracts a time-series signal from monochromatic NIR face video. The second module consists of a novel time-series U-net architecture in which a GRU (gated recurrent unit) network has been added to the passthrough layers. We test our approach on the challenging MR-NIRP Car Dataset, which consists of monochromatic NIR videos taken in both stationary and driving conditions. Our model’s iPPG estimation performance on NIR video outperforms both the state-of-the-art model-based method and a recent end-to-end deep learning method that we adapted to monochromatic video. Armand Comas, Tim K. Marks, Hassan Mansour, Suhas Lohit, Yechi Ma, Xiaoming Liu 0002 |
ICIP | 3 |
| 2021 | Application-Agnostic Spatio-Temporal Hand Graph Representations For Stable Activity UnderstandingabstractUsing hand skeleton data to understand complex hand actions, such as assembly tasks or kitchen activities, is an important yet challenging task. This paper introduces an unsupervised hand graph-based spatio-temporal feature extraction method. To evaluate the efficacy of the proposed representation, we consider action segmentation and recognition tasks. The segmentation problem involves an assembling task in an industrial setting, while the recognition problem deals with kitchen and office activities. For both tasks, we propose novel notions of stability, loss function stability (LFS) and estimation stability with cross-validation (ESCV), that are used to quantify the robustness of achieved solutions. Our proposed feature extraction leads to classification performance comparable to state of the art methods, while achieving significantly better accuracy and stability in a cross-person setting. The proposed method also outperforms the existing methods in the segmentation task in terms of accuracy and shows robustness to any change in the input hyper-parameters. Pratyusha Das, Antonio Ortega, Siheng Chen, Hassan Mansour, Anthony Vetro |
ICIP | 4 |
| 2020 | Learning Plug-And-Play Proximal Quasi-Newton DenoisersabstractPlug-and-play (PnP) denoising for solving inverse problems has received significant attention recently thanks to its state of the art signal reconstruction performance. However, the performance improvement hinges on carefully choosing the noise level of the Gaus-sian denoiser and the descent step size in every iteration. We propose a strategy for training a Gaussian denoiser inspired by an unfolded proximal quasi-Newton algorithm, where the noise level of the input signal to the denoiser is estimated in each iteration and at every entry in the signal. Our scheme deploys a small convolutional neural network (mini-CNN) to estimate an element-wise noise level, mimicking a diagonal approximation of the Hessian matrix in quasi-Newton methods. Empirical simulation results on image deblurring demonstrate that our proposed approach achieves approximately 1dB improvement over state of the art methods, such as, BM3D-PnP and proximal gradient descent-PnP that are supplied with the true noise level, as well as over an end-to-end retrained FFDNet architecture that was trained to estimate the noise level and recover the deblurred images. Abdullah H. Al-Shabili, Hassan Mansour, Petros Boufounos |
ICASSP | 2 |
| 2020 | Inverse Multiple Scattering with Phaseless MeasurementsabstractWe study the problem of reconstructing an object from phaseless measurements in the context of inverse multiple scattering. Our formulation explicitly decouples the variables that represent the unknown object image and the unknown phase, respectively, in the forward model. This enables us to simultaneously optimize over both unknowns with appropriate regularization for each. The resulting optimization problem is nonconvex due to the nonlinear propagation model for multiple scattering and the nonconvex regularization of the phase variables. Nevertheless, we demonstrate experimentally that we can solve the optimization problem using a variation of the fast iterative shrinkage-thresholding algorithm (FISTA)-a convex algorithm, popular for its speed and simplicity-that converges well in our experiments. Numerical results with both simulated and experimentally measured data show that the proposed method outperforms the state-of-the-art phaseless inverse scattering method. Muhammad Asad Lodhi, Yanting Ma, Hassan Mansour, Petros Boufounos, Dehong Liu |
ICASSP | 3 |
| 2020 | Slow-Time MIMO-FMCW Automotive Radar Detection with Imperfect Waveform SeparationabstractThis paper considers object detection in the case of imperfect waveform separation, in the context of automotive radars with a slow-time MIMO-FMCW signaling scheme. We develop an explicit signal model that accounts for waveform separation residuals and propose a Kronecker subspace-based object detector in the framework of generalized likelihood ratio test (GLRT). Our exact theoretical analysis under both hypotheses shows that the proposed detector holds the desired property of constant false alarm rate (CFAR). Numerical simulations validate our proposed object detection scheme. Pu Wang 0004, Petros Boufounos, Hassan Mansour, Philip V. Orlik |
ICASSP | 3 |
| 2020 | Extended Object Tracking Using Hierarchical Truncation Measurement Model with Automotive RadarabstractMotivated by real-world automotive radar measurements that are distributed around object (e.g., vehicles) edges with a certain volume, a novel hierarchical truncated Gaussian measurement model is proposed to resemble the underlying spatial distribution of radar measurements. With the proposed measurement model, a modified random matrix-based extended object tracking algorithm is developed to estimate both kinematic and extent states. In particular, a new state update step and an online bound estimation step are proposed with the introduction of pseudo measurements. The effectiveness of the proposed algorithm is verified in simulations. Yuxuan Xia, Pu Wang 0004, Karl Berntorp, Toshiaki Koike-Akino, Hassan Mansour, Milutin Pajovic, Petros Boufounos, Philip V. Orlik |
ICASSP | 5 |
| 2020 | Robust Parameter Estimation of Contaminated Damped ExponentialsabstractParameter estimation of damped exponential signals has wide applications including fault detection and system parameter identification, etc. However, existing methods for estimating parameters of damped exponentials are either sensitive to noise or restricted to dealing with a certain type of noise such as Gaussian noise. In this paper we aim to estimate parameters of damped exponentials from contaminated signal, i.e., a mixture of damped exponentials, random Gaussian noise, and spike interference. We propose two robust approaches, a convex one solved by the alternating direction method of multipliers (ADMM) and a non-convex one solved by coordinate descent, to recovering a low-rank Hankel matrix of damped exponentials from noisy measurements for further parameter estimation using the matrix pencil technique. Numerical experiments show that our proposed methods outperform classical ones in detecting small damped fault signatures from noisy measurements. While the convex approach is amenable to theoretical analysis and global convergence guarantees, the non-convex one exhibits more robustness and computational efficiency. Youye Xie, Dehong Liu, Hassan Mansour, Petros Boufounos |
ICASSP | 3 |
| 2020 | Blind Multi-Spectral Image Pan-SharpeningabstractWe address the problem of sharpening low spatial-resolution multi-spectral (MS) images with their associated misaligned high spatial-resolution panchromatic (PAN) image, based on priors on the spatial blur kernel and on the cross-channel relationship. In particular, we formulate the blind pan-sharpening problem within a multi-convex optimization framework using total generalized variation for the blur kernel and local Laplacian prior for the cross-channel relationship. The problem is solved by the alternating direction method of multipliers (ADMM), which alternately updates the blur kernel and sharpens intermediate MS images. Numerical experiments demonstrate that our approach is more robust to large misalignment errors and yields better super resolved MS images compared to state-of-the-art optimization-based and deep-learning-based algorithms. Lantao Yu, Dehong Liu, Hassan Mansour, Petros Boufounos, Yanting Ma |
ICASSP | 3 |
| 2020 | Robust 3D Tomographic Imaging of the Ionospheric Electron DensityabstractIn this paper, we develop a robust three dimensional tomographic imaging framework to estimate the ionospheric electron density using ground-based total electron content (TEC) measurements from GPS receivers. In order to increase the sampling rate of the domain, we incorporate into the tomographic measurements the TEC readings observed from low-angle satellites that fall outside of the target ionospheric domain. We discount the proportion of the TEC measurements that originate outside of the target domain using the simulation-based NeQuick2 model as reference. We also employ a diffusion kernel regularization function to robustify the reconstruction against errors in the NeQuick2 model. Finally, we demonstrate through simulations that our framework delivers superior reconstruction of the ionospheric electron density compared to existing schemes. We also demonstrate the applicability of our approach on real TEC measurements. Xiaojian Xu 0002, Oussama Dhifallah, Hassan Mansour, Petros Boufounos, Philip V. Orlik |
IGARSS | 3 |
| 2019 | Hand Graph Representations for Unsupervised Segmentation of Complex ActivitiesabstractAnalysis of hand skeleton data can be used to understand patterns in manipulation and assembly tasks. This paper introduces a graph-based representation of hand skeleton data and proposes a method to perform unsupervised temporal segmentation of a sequence of sub-tasks in order to evaluate the efficiency of an assembly task. We explore the properties of different choices of hand graphs and their spectral decomposition. A comparative performance of these graphs is presented in the context of complex activity segmentation. We show that the spectral graph features extracted from 2D hand motion data outperform the direct use of motion vectors as features. We also make the collected hand position data available to the research community to facilitate further development in this direction. Pratyusha Das, Jiun-Yu Kao, Antonio Ortega, Tomoya Sawada, Hassan Mansour, Anthony Vetro, Akira Minezawa |
ICASSP | 5 |
| 2019 | Reflection Tomographic Imaging of Highly Scattering Objects Using Incremental Frequency InversionabstractReflection tomography is an inverse scattering technique that estimates the spatial distribution of an object's permittivity by illuminating it with a probing pulse and measuring the scattered wavefields by receivers located on the same side as the transmitter. Unlike conventional transmission tomography, the reflection regime is severely ill-posed since the measured wavefields contain far less spatial frequency information about the object. In this paper, we propose an incremental frequency inversion framework that requires no initial target model, and that leverages spatial regularization to reconstruct the permittivity distribution of highly scattering objects. Our framework solves a wave-equation constrained, total-variation (TV) regularized nonlinear least squares problem that solves a sequence of subproblems that incrementally enhance the resolution of the estimated object model. With each subproblem, higher frequency wavefield components are incorporated in the inversion to improve the recovered model resolution. We validate the performance of our approach using synthetically generated data for retrieving high-contrast material such as water in an underground radar imaging setup. Ajinkya Kadu, Hassan Mansour, Petros Boufounos, Dehong Liu |
ICASSP | 2 |
| 2019 | Coherent Radar Imaging Using Unsynchronized Distributed AntennasabstractIn this paper we develop an optimization-based solution to the problem of distributed radar imaging using antennas with asynchronous clocks. In particular, we consider a distributed radar imaging MIMO system observing a sparse scene under an unknown, but bounded, delay between the transmitter and receiver clocks. Most existing approaches pose the problem as the recovery of a phase shift, leading to non-convex formulations. Instead, inspired by recent work in blind deconvolution, we exploit the realization that synchronization errors in the received data can be modeled as a convolution with an unknown 1-sparse delay signal to be estimated in addition to the image. Thus, we formulate a convex optimization problem that simultaneously recovers all the pair-wise drifts between transmit/receive pairs, as well as the sparse scene being imaged. We verify the validity and performance of our proposed model and recovery method through numerical simulations on synthetic data. Muhammad Asad Lodhi, Hassan Mansour, Petros Boufounos |
ICASSP | 2 |
| 2019 | Unrolled Projected Gradient Descent for Multi-spectral Image FusionabstractIn this paper, we consider the problem of fusing low spatial resolution multi-spectral (MS) aerial images with their associated high spatial resolution panchromatic image. To solve this problem, various methods have been proposed, using either model-based or model-agnostic algorithms such as deep learning techniques. In this paper, we aim to utilize more interpretable architectures to solve the MS fusion problem by integrating existing ideas from image processing with deep learning. In particular, we develop a signal processing-inspired learning solution, where we unroll the iterations of the projected gradient descent (PGD) algorithm, and each iteration contains a projection operation carried out by a deep convolutional neural network. We observe that our proposed method provides a new perspective on existing deep-learning solutions, and under certain circumstance it reduces to current black-box deep learning methods. Our extensive experimental results show significant improvements of the proposed approach over several baselines. Suhas Lohit, Dehong Liu, Hassan Mansour, Petros Boufounos |
ICASSP | 3 |
| 2019 | Graph Based Skeleton Modeling for Human Activity AnalysisabstractUnderstanding human activity based on sensor information is required in many applications and has been an active research area. With the advancement of depth sensors and tracking algorithms, systems for human motion activity analysis can be built by combining off-the-shelf motion tracking systems with application-dependent learning tools to extract higher semantic level information. Many of these motion tracking systems provide raw motion data registered to the skeletal joints in the human body. In this paper, we propose novel representations for human motion data using the skeleton-based graph structure along with techniques in graph signal processing. Methods for graph construction and their corresponding basis functions are discussed. The proposed representations can achieve comparable classification performance in action recognition tasks while additionally being more robust to noise and missing data. Jiun-Yu Kao, Antonio Ortega, Dong Tian, Hassan Mansour, Anthony Vetro |
ICIP | 4 |
| 2019 | Robust Mutual Information-Based Multi-Image RegistrationabstractImage registration is of crucial importance in image fusion such as pan-sharpening. Mutual information (MI)-based methods have been widely used and demonstrated effectiveness in registering multi-spectral or multi-modal images. However, MI-based methods may fail to converge in searching registration parameters, resulting mis-registration. In this paper, we propose an outlier robust method to improve the robustness of MI-based registration for multiple rigid transformed images. In particular, we first generate registration parameter matrices using a MI-based approach, then we decompose each parameter matrix into a low-rank matrix of inlier registration parameters and a sparse matrix corresponding to outlier parameter errors. Results of registering multi-spectral images with random rigid transformations show significant improvement and robustness of our method. Dehong Liu, Hassan Mansour, Petros Boufounos |
IGARSS | 2 |
| 2018 | Online Detection of Action Start in Untrimmed, Streaming Videos
Zheng Shou 0001, Junting Pan, Kazuyuki Miyazawa, Hassan Mansour, Anthony Vetro, Xavier Giró-i-Nieto, Shih-Fu Chang |
ECCV (3) | 5 |
| 2018 | Accelerated Image Reconstruction for Nonlinear Diffractive ImagingabstractThe problem of reconstructing an object from the measurements of the light it scatters is common in numerous imaging applications. While the most popular formulations of the problem are based on linearizing the object-light relationship, there is an increased interest in considering nonlinear formulations that can account for multiple light scattering. In this paper, we propose an image reconstruction method, called CISOR, for nonlinear diffractive imaging, based on our new variant of fast iterative shrinkage/thresholding algorithm (FISTA) and total variation (TV) regularization. We prove that CISOR reliably converges for our nonconvex optimization problem, and systematically compare our method with other state-of-the-art methods on simulated as well as experimentally measured data. Yanting Ma, Hassan Mansour, Dehong Liu, Petros Boufounos, Ulugbek Kamilov |
ICASSP | 2 |
| 2018 | Radar Autofocus Using Sparse Blind DeconvolutionabstractThe radar autofocus problem arises in situations where radar measurements are acquired of a scene using antennas that suffer from position ambiguity. Current techniques model the antenna ambiguity as a global phase error affecting the received radar measurement at every antenna. However, the phase error signal model is only valid in the far field regime where the position error can be approximated by a one dimensional shift in the down-range direction. We propose in this paper an alternate formulation where the antenna position error is modeled using a two-dimensional shift operator in the image-domain. The radar autofocus problem then becomes a multichannel two-dimensional blind deconvolution problem where the static radar image is convolved with a two dimensional shift kernel for each antenna measurement. We develop an alternating minimization framework that leverages the sparsity and piece-wise smoothness of the radar scene, as well as the one-sparse property of the two dimensional shift kernels. Hassan Mansour, Dehong Liu, Petros Boufounos, Ulugbek Kamilov |
ICASSP | 1 |
| 2018 | Deepcasd: An End-to-End Approach for Multi-Spectral Image Super-ResolutionabstractMulti-spectral (MS) image super-resolution aims to reconstruct super-resolved multi-channel images from their low-resolution images by regularizing the image to be reconstructed. Recently data-driven regularization techniques based on sparse modeling and deep learning have achieved substantial improvements in single image reconstruction problems. Inspired by these data-driven methods, we develop a novel coupled analysis and synthesis dictionary (CASD) model for MS image super-resolution, by exploiting a regularizer that operates within, as well as across, multiple spectral channels using convolutional dictionaries. To learn the CASD model parameters, we propose a deep dictionary learning framework, named DeepCASD, by unfolding and training an end-to-end CASD based reconstruction network over an image data set. Experimental results show that the DeepCASD framework exhibits improved performance on multi-spectral image super-resolution compared to state-of-the-art learning based super-resolution algorithms. Bihan Wen, Ulugbek Kamilov, Dehong Liu, Hassan Mansour, Petros Boufounos |
ICASSP | 4 |
| 2018 | Robust Sensor Localization Based on Euclidean Distance MatrixabstractIn remote sensing systems, exact knowledge of the sensor locations is critical for generating focused images. In order to accurately locate misplaced or perturbed sensors from their received signal data, we proposed a robust sensor localization method based on low-rank Euclidean distance matrix (EDM) reconstruction. To this end, an EDM of sensors and objects under detection is defined and partially initialized by computing distances between the inaccurate sensor locations and distances from the sensors to the objects using signal coherence analysis. We then decompose the noisy EDM with missing entries into a low-rank EDM corresponding to true sensor locations and a sparse matrix of distance errors by solving a constrained optimization problem using the alternating direction method of multipliers (ADMM). We verify our method with simulations on a uniform linear array with unknown perturbations up to several wavelengths. Dehong Liu, Hassan Mansour, Petros Boufounos, Ulugbek Kamilov |
IGARSS | 2 |
| 2017 | Disc-GLasso: Discriminative graph learning with sparsity regularizationabstractLearning graph topology from data is challenging. Previous work leads to learning graphs on which the graph signals used for training are smooth. In this paper, we propose an optimization framework for learning multiple graphs, each associated to a class of signals, such that representation of signals within a class and discrimination of signals in different classes are both taken into consideration. A Fisher-LDA-like term is included in the optimization objective function in addition to the conventional Gaussian ML objective. A block coordinate descent algorithm is then developed to estimate optimal graphs for different categories of signals, which are then used to efficiently classify the different signals. Experiments on synthetic data demonstrate that our proposed method can achieve better discrimination between the learned graphs, leading to improvements in subsequent classification tasks. Jiun-Yu Kao, Dong Tian, Hassan Mansour, Antonio Ortega, Anthony Vetro |
ICASSP | 3 |
| 2017 | Jazz: A companion to music for frequency estimation with missing dataabstractFrequency estimation is a classical problem in signal processing, with applications ranging from sensor array processing to wireless communications and structural health monitoring. Modern algorithms based on atomic norm minimization can cope with missing data but incur a high computational cost. To recover missing data from an ensemble of frequency-sparse signals, we propose a computationally efficient low-rank tensor completion algorithm that exploits the fact that each signal in the ensemble can be associated with a Toeplitz matrix. We name our algorithm JAZZ in the spirit of the classical MUSIC algorithm for frequency estimation and in tribute to the random, improvisational nature of jazz music. Qiuwei Li, Shuang Li 0003, Hassan Mansour, Michael B. Wakin, Dehui Yang, Zhihui Zhu |
ICASSP | 3 |
| 2017 | Compressive imaging with iterative forward modelsabstractWe propose a new compressive imaging method for reconstructing 2D or 3D objects from their scattered wave-field measurements. Our method relies on a novel, nonlinear measurement model that can account for the multiple scattering phenomenon, which makes the method preferable in applications where linear measurement models are inaccurate. We construct the measurement model by expanding the scattered wave-field with an accelerated-gradient method, which is guaranteed to converge and is suitable for large-scale problems. We provide explicit formulas for computing the gradient of our measurement model with respect to the unknown image, which enables image formation with a sparsity-driven numerical optimization algorithm. We validate the method both analytically and with numerical simulations. Hsiou-Yuan Liu, Ulugbek Kamilov, Dehong Liu, Hassan Mansour, Petros Boufounos |
ICASSP | 4 |
| 2017 | Fusion of multi-angular aerial images based on epipolar geometry and matrix completionabstractWe consider the problem of fusing multiple cloud-contaminated aerial images of a 3D scene to generate a cloud-free image, where the images are captured from multiple unknown view angles. In order to fuse these images, we propose an end-to-end framework incorporating epipolar geometry and low-rank matrix completion. In particular, we first warp the multi-angular images to single-angle ones based on the estimated fundamental matrices that relate the multi-angular images according to their projective relations to the 3D scene. Then we formulate the fusion process of the warpped images as a low-rank matrix completion problem where each column of the matrix corresponds to a vectorized image with missing entries corresponding to cloud or occluded areas. Results using DigitalGlobe high spatial resolution images demonstrate that our algorithm outperforms existing approaches. Yanting Ma, Dehong Liu, Hassan Mansour, Ulugbek Kamilov, Yuichi Taguchi, Petros Boufounos, Anthony Vetro |
ICIP | 3 |
| 2017 | A Plug-and-Play Priors Approach for Solving Nonlinear Imaging Inverse ProblemsabstractIn the past two decades, nonlinear image reconstruction methods have led to substantial improvements in the capabilities of numerous imaging systems. Such methods are traditionally formulated as optimization problems that are solved iteratively by simultaneously enforcing data consistency and incorporating prior models. Recently, the Plug-and-Play Priors (PPP) framework suggested that by using more sophisticated denoisers, not necessarily corresponding to an optimization objective, it is possible to improve the quality of reconstructed images. In this letter, we show that the PPP approach is applicable beyond linear inverse problems. In particular, we develop the fast iterative shrinkage/thresholding algorithm variant of PPP for model-based nonlinear inverse scattering. The key advantage of the proposed formulation over the original ADMM-based one is that it does not need to perform an inversion on the forward model. We show that the proposed method produces high quality images using both simulated and experimentally measured data. Ulugbek Kamilov, Hassan Mansour, Brendt Wohlberg |
IEEE Signal Process. Lett. | 2 |
| 2016 | Geometric-guided label propagation for moving object detectionabstractMoving object segmentation in video has uses in many applications and is a particularly challenging task when the video is acquired by a moving camera. Typical approaches that rely on principal component analysis (PCA) tend to extract scattered sparse components of the moving objects and generally fail in extracting dense object segmentations. In this paper, a novel label propagation framework based on motion vanishing point (MVP) analysis is proposed to address the challenges. A weighted graph is constructed with image pixels as nodes and the MVP-guided approach is used to define the graph weights. Label propagation is then performed by incorporating the graph Laplacian. In addition, a PCA result is used to initialize the foreground/background labels. Experiments on the Hopkins data set of outdoor sequences captured by a hand-held moving camera demonstrate that the proposed label propagation method outperforms state-of-the-art PCA and spectral clustering methods for a dense segmentation task. Moreover, the framework is capable of correcting mislabeled foreground pixels and thus does not require accurate initial label assignment. Jiun-Yu Kao, Dong Tian, Hassan Mansour, Anthony Vetro, Antonio Ortega |
ICASSP | 3 |
| 2016 | Multipath removal by online blind deconvolution in through-the-wall-imagingabstractIn this paper, we propose an online radar imaging scheme that recovers a sparse scene and removes the multipath ringing induced by the front wall in a Through-the-Wall-Imaging (TWI) system without prior knowledge of the wall parameters. Our approach uses online measurements obtained from individual transmitter-receiver pairs to incrementally build the primary response of targets behind the front wall and find a corresponding delay convolution operator that generates the multi-path reflections available in the received signal. In order to perform online sparse imaging while removing wall clutter reflections, we developed a deconvolution extension of the Sparse Randomized Kaczmarz (SRK) algorithm that finds sparse solutions to under- and over-determined linear systems of equations. Our scheme allows for imaging with nonuniformly spaced antennas by building an explicit delay-and-sum imaging operator for each new measurement. Moreover, the active memory requirements remain small even for large scale MIMO systems since the imaging operators are only constructed for individual transmitter-receiver pairs. We test our approach on a simple FDTD simulated room with internal targets and demonstrate that our method successfully eliminates multipath reflections while correctly locating the targets. Hassan Mansour, Ulugbek Kamilov |
ICASSP | 1 |
| 2016 | Moving object segmentation using depth and optical flow in car driving sequencesabstractSegmentation of moving objects in a scene is difficult for non-stationary cameras, and especially challenging in the presence of fast and unstable egomotion, e.g., as encountered with car-mounted cameras or wearable devices. Based on an analysis of motion vanishing points of the scene and estimated depth, a geometric model that relates extracted 2D motion to a 3D motion field relative to the camera is derived. Observing that the 3D motion field is piece-wise smooth, a constrained optimization problem that considers group sparsity is formulated to recover the 3D motion field from the 2D motion. The recovered 3D motion field is then clustered to provide the segmentation of moving objects. Experiments are performed using the KITTI Vision Benchmark Suite and demonstrate that the proposed framework provides a dense segmentation of moving objects that is robust to the challenging conditions inherent with car driving sequences. Jiun-Yu Kao, Dong Tian, Hassan Mansour, Anthony Vetro, Antonio Ortega |
ICIP | 3 |
| 2016 | Robust low rank dynamic mode decomposition for compressed domain crowd and traffic flow analysisabstractIn this paper, we develop a dynamic mode decomposition algorithm that is robust to both inlier and outlier noise in the data. One application of our algorithm is the identification of multiple crowd or traffic flows from compressed video streams. Our method uses motion vectors that are readily available in the compressed bitstream, and do not require computationally expensive optical flow. These motion vectors are known to be very noisy, however, our algorithm is able to extract the underlying dynamical systems that define the flows. We formulate a rank regularized dynamic mode decomposition problem with total least squares constraints to estimate the Koopman modes of the motion dynamics. The estimated Koopman modes are then used to analyze the stability of the system and extract steady state and transient flows. We demonstrate the improved performance of our approach compared to state of the art schemes and illustrate it applicability in identifying transient and steady-state flows in real video sequences. Caglayan Dicle, Hassan Mansour, Dong Tian, Mouhacine Benosman, Anthony Vetro |
ICME | 2 |
| 2016 | A Recursive Born Approach to Nonlinear Inverse ScatteringabstractThe iterative Born approximation (IBA) is a well-known method for describing waves scattered by semitransparent objects. In this letter, we present a novel nonlinear inverse scattering method that combines IBA with an edge-preserving total variation regularizer. The proposed method is obtained by relating iterations of IBA to layers of an artificial multilayer neural network and developing a corresponding error backpropagation algorithm for efficiently estimating the permittivity of the object. Simulations illustrate that, by accounting for multiple scattering, the method successfully recovers the permittivity distribution where the traditional linear inverse scattering fails. Ulugbek Kamilov, Dehong Liu, Hassan Mansour, Petros Boufounos |
IEEE Signal Process. Lett. | 3 |
| 2016 | Learning Optimal Nonlinearities for Iterative Thresholding AlgorithmsabstractIterative shrinkage/thresholding algorithm (ISTA) is a well-studied method for finding sparse solutions to ill-posed inverse problems. In this letter, we present a data-driven scheme for learning optimal thresholding functions for ISTA. The proposed scheme is obtained by relating iterations of ISTA to layers of a simple feedforward neural network and developing a corresponding error backpropagation algorithm for fine-tuning the thresholding functions. Simulations on sparse statistical signals illustrate potential gains in estimation quality due to the proposed data adaptive ISTA. Ulugbek Kamilov, Hassan Mansour |
IEEE Signal Process. Lett. | 2 |
| 2015 | Kernel Machine Classification Using Universal EmbeddingsabstractSummary form only given. Visual inference over a transmission channel is increasingly becoming an important problem in a variety of applications. In such applications, low latency and bit-rate consumption are often critical performance metrics, making data compression necessary. In this paper, we examine feature compression for support vector machine (SVM)-based inference using quantized randomized embeddings. We demonstrate that embedding the features is equivalent to using the SVM kernel trick with a mapping to a lower dimensional space. Furthermore, we show that universal embeddings - a recently proposed quantized embedding design - approximate a radial basis function (RBF) kernel, commonly used for kernel-based inference. Our experimental results demonstrate that quantized embeddings achieve 50% rate reduction, while maintaining the same inference performance. Moreover, universal embeddings achieve a further reduction in bit-rate over conventional quantized embedding methods, validating the theoretical predictions. Petros Boufounos, Hassan Mansour |
DCC | 2 |
| 2015 | A robust online subspace estimation and tracking algorithmabstractIn this paper, we present a robust online subspace estimation and tracking algorithm (ROSETA) that is capable of identifying and tracking a time-varying low dimensional subspace from incomplete measurements and in the presence of sparse outliers. Our algorithm minimizes a robust ℓ1norm cost function between the observed measurements and their projection onto the estimated subspace. The projection coefficients and sparse outliers are computed using ADMM solver and the subspace estimate is updated using a proximal point iteration with adaptive parameter selection. We demonstrate using simulated experiments and a video background subtraction example that ROSETA succeeds in identifying and tracking low dimensional subspaces using fewer iterations than other state of art algorithms. Hassan Mansour |
ICASSP | 1 |
| 2015 | Weighted one-norm minimization with inaccurate support estimates: Sharp analysis via the null-space propertyabstractWe study the problem of recovering sparse vectors given possibly erroneous support estimates. First, we provide necessary and sufficient conditions for weighted ℓ1minimization to successfully recovery all sparse signals whose support estimate is sufficiently accurate. We relate these conditions to the analogous ones for ℓ1minimization, showing that they are equivalent when the support estimate is 50% accurate but that the weighted ℓ1conditions are easier to satisfy when the support is more than 50% accurate. Second, to quantify this improvement, we provide bounds on the number of Gaussian measurements that ensure, with high probability, that weighted ℓ1minimization succeeds. The resulting number of measurements can be significantly less than what is needed to ensure recovery via ℓ1minimization. Finally, we illustrate our results via numerical experiments. Hassan Mansour, Rayan Saab |
ICASSP | 1 |
| 2015 | Depth-weighted group-wise principal component analysis for video foreground/background separationabstractWe propose a depth-weighted group-wise PCA (DG-PCA) approach to separate moving foreground pixels from the background of a video acquired by a moving camera. Our approach utilizes a corresponding depth signal in addition to the video signal. The problem is formulated as a weighted l2,1-norm PCA problem with depth-based group sparsity being introduced. In particularly, dynamic groups are first generated solely based on depth, and then an iterative solution using depth to define the weights in l2,1-norm is developed. In addition, we propose a depth-enhanced homography model for global motion compensation before the DG-PCA method is executed. We demonstrate through experiments on an RGB-D dataset the superiority of the proposed DG-PCA approach over conventional robust PCA methods. Dong Tian, Hassan Mansour, Anthony Vetro |
ICIP | 2 |
| 2015 | Graph spectral motion segmentation based on motion vanishing point analysisabstractMotion segmentation relies on identifying coherent relationships between image pixels that are associated with motion vectors. However, perspective differences can often deteriorate the performance of conventional techniques. In this paper, we develop a motion segmentation scheme that utilizes the motion map of a single frame to identify motion representations based on motion vanishing points. Segmentation is achieved using graph spectral clustering where a novel graph is constructed using the motion representation distances in the motion vanishing point image associated with the image pixels. Experimental results show that the proposed graph spectral motion segmentation algorithm outperforms state-of-the-art methods for dense segmentation on image sequences with strong perspective effects using motion vectors between only two images. Dong Tian, Jiun-Yu Kao, Hassan Mansour, Anthony Vetro |
MMSP | 3 |
| 2014 | Video background subtraction using semi-supervised robust matrix completionabstractWe propose a factorized robust matrix completion (FRMC) algorithm with global motion compensation to solve the video background subtraction problem. The algorithm decomposes a sequence of video frames into the sum of a low rank background component and a sparse motion component. The algorithm alternates between the solution of each component following a Pareto curve trajectory for each subproblem. For videos with moving background, we utilize the motion vectors extracted from the coded video bitstream to compensate for the change in the camera perspective. Performance evaluations show that our approach is faster than state-of-the-art solvers and results in highly accurate motion segmentation. Hassan Mansour, Anthony Vetro |
ICASSP | 1 |
| 2014 | Video querying via compact descriptors of visually salient objectsabstractWe consider the problem of extracting descriptors that represent visually salient portions of a video sequence. Most state-of-the-art schemes generate video descriptors by extracting features, e.g., SIFT or SURF or other keypoint-based features, from individual video frames. This approach is wasteful in scenarios that impose constraints on storage, communication overhead and on the allowable computational complexity for video querying. More importantly, the descriptors obtained by this approach generally do not provide semantic clues about the video content. In this paper, we investigate new feature-agnostic approaches for efficient retrieval of similar video content. We evaluate the efficiency and accuracy of retrieval when k-means clustering is applied to image features extracted from video frames. We also propose a new approach in which the extraction of compact video descriptors is cast as a Non-negative Matrix Factorization (NMF) problem. Initial experiments on video-based matching suggest that compact descriptors obtained via low-rank matrix factorization improve discriminability and robustness to parameter selection compared to k-means clustering. Hassan Mansour, Shantanu Rane, Petros Boufounos, Anthony Vetro |
ICIP | 1 |
| 2014 | Depth-assisted stereo video enhancement using graph-based approachesabstractIn stereo video applications, the quality of the two views may vary based on different camera capturing conditions and setup, compression/transmission, and sensor noise. Although some studies show that the perceived video quality may not be significantly affected by the lower quality view, maintaining a similar video quality is still desired in order to prevent eye strain during extended viewing sessions. In this paper, we study a graph-based approach to enhance the lower quality views by referring to the high quality view in addition to an accompanying depth map. We construct a graphical signal model with joint bilateral edge weights and show that graph-based joint bilateral filtering can better suppress several types of noises, e.g., Gaussian, motion as well as quantization noise. Dong Tian, Hassan Mansour, Anthony Vetro, Yongzhe Wang, Antonio Ortega |
ICIP | 2 |
| 2013 | Visually Favorable Tone-Mapping With High Compression Performance in Bit-Depth Scalable Video CodingabstractIn bit-depth scalable video coding, the tone-mapping scheme used to convert high-bit-depth to eight-bit videos is an essential yet very often ignored component. In this paper, we demonstrate that an appropriate choice of a tone-mapping operator can improve the coding efficiency of bit-depth scalable encoders. We present a new tone-mapping scheme that delivers superior compression efficiency while adhering to a predefined base layer perceptual quality. We develop numerical models that estimate the base layer bit-rate (Rb), the enhancement layer bitrate (Re), and the mismatch (QL) between the resulting low dynamic range (LDR) base-layer signal and the predefined base layer representation. Our proposed tone curve is given by the solution of an optimization problem which minimizes a weighted sum of Rb, Re, and QL. The problem formulation also considers the temporal effect of tone-mapping by adding a constraint to the optimization problem that suppresses flickering artifacts. We also propose a technique with which to tone-map a high-bit-depth video directly in a compression-friendly color space (e.g., one luma and two chroma channels) without converting to the RGB domain. Experimental results show that we can save up to 40% of the total bit-rate (or 3.5 dB PSNR improvement for the same bitrate), and, in general, about 20% bit-rate savings can be achieved. Zicong Mai, Hassan Mansour, Panos Nasiopoulos, Rabab K. Ward |
IEEE Trans. Multim. | 2 |
| 2012 | Support driven reweighted ℓ1 minimizationabstractIn this paper, we propose a support driven reweighted ℓ1minimization algorithm (SDRL1) that solves a sequence of weighted ℓ1problems and relies on the support estimate accuracy. Our SDRL1 algorithm is related to the IRL1 algorithm proposed by Candès, Wakin, and Boyd. We demonstrate that it is sufficient to find support estimates with good accuracy and apply constant weights instead of using the inverse coefficient magnitudes to achieve gains similar to those of IRL1. We then prove that given a support estimate with sufficient accuracy, if the signal decays according to a specific rate, the solution to the weighted ℓ1minimization problem results in a support estimate with higher accuracy than the initial estimate. We also show that under certain conditions, it is possible to achieve higher estimate accuracy when the intersection of support estimates is considered. We demonstrate the performance of SDRL1 through numerical simulations and compare it with that of IRL1 and standard ℓ1minimization. Hassan Mansour, Özgür Yilmaz |
ICASSP | 1 |
| 2012 | Adaptive compressed sensing for video acquisitionabstractIn this paper, we propose an adaptive compressed sensing scheme that utilizes a support estimate to focus the measurements on the large valued coefficients of a compressible signal. We embed a “sparse-filtering” stage into the measurement matrix by weighting down the contribution of signal coefficients that are outside the support estimate. We present an application which can benefit from the proposed sampling scheme, namely, video compressive acquisition. We demonstrate that our proposed adaptive CS scheme results in a significant improvement in reconstruction quality compared with standard CS as well as adaptive recovery using weighted ℓ1minimization. Hassan Mansour, Özgür Yilmaz |
ICASSP | 1 |
| 2012 | Recovering Compressively Sampled Signals Using Partial Support InformationabstractWe study recovery conditions of weightedl1minimization for signal reconstruction from compressed sensing measurements when partial support information is available. We show that if at least 50% of the (partial) support information is accurate, then weightedl1minimization is stable and robust under weaker sufficient conditions than the analogous conditions for standardl1minimization. Moreover, weightedl1minimization provides better upper bounds on the reconstruction error in terms of the measurement noise and the compressibility of the signal to be recovered. We illustrate our results with extensive numerical experiments on synthetic data and real audio and video signals. Michael P. Friedlander, Hassan Mansour, Rayan Saab, Özgür Yilmaz |
IEEE Trans. Inf. Theory | 2 |
| 2011 | Optimizing a Tone Curve for Backward-Compatible High Dynamic Range Image and Video CompressionabstractFor backward compatible high dynamic range (HDR) video compression, the HDR sequence is reconstructed by inverse tone-mapping a compressed low dynamic range (LDR) version of the original HDR content. In this paper, we show that the appropriate choice of a tone-mapping operator (TMO) can significantly improve the reconstructed HDR quality. We develop a statistical model that approximates the distortion resulting from the combined processes of tone-mapping and compression. Using this model, we formulate a numerical optimization problem to find the tone-curve that minimizes the expected mean square error (MSE) in the reconstructed HDR sequence. We also develop a simplified model that reduces the computational complexity of the optimization problem to a closed-form solution. Performance evaluations show that the proposed methods provide superior performance in terms of HDR MSE and SSIM compared to existing tone-mapping schemes. It is also shown that the LDR image quality resulting from the proposed methods matches that produced by perceptually-based TMOs. Zicong Mai, Hassan Mansour, Rafal Mantiuk, Panos Nasiopoulos, Rabab K. Ward, Wolfgang Heidrich |
IEEE Trans. Image Process. | 2 |
| 2011 | Rate and Distortion Modeling of CGS Coded Scalable Video ContentabstractIn this paper, we derive single layer and scalable video rate and distortion models for video bitstreams encoded using the coarse grain quality scalability (CGS) feature of the scalable extension of H.264/AVC. In these models, we assume the source is Laplacian distributed and compensate for errors in the distribution assumption by linearly scaling the Laplacian parameter . Moreover, we present simplified approximations of the derived models that allow for a run-time calculation of sequence dependent model constants. Our models use the mean absolute difference (MAD) of the prediction residual signal and the encoder quantization parameter (QP) as input parameters. Consequently, we are able to estimate the residual MAD, bitrate, and distortion of a future video frame at any QP value and for both base-layer and CGS layer packets. We also present simulation results that demonstrate the accuracy of the proposed models. Hassan Mansour, Panos Nasiopoulos, Vikram Krishnamurthy |
IEEE Trans. Multim. | 1 |
| 2010 | Color image desaturation using sparse reconstructionabstractIn this paper, we propose an algorithm to estimate the true values of saturated pixels in color images. Pixel saturation occurs when at least one color channel is clipped at some value below the full dynamic range of the scene, resulting in a loss in image fidelity. The proposed algorithm is based on the assumptions that images are nearly sparse in an appropriate transform domain, and that saturated pixels can be inferred from the structure of non-saturated neighboring pixels. Consequently, we use a hierarchical windowing algorithm which selects image regions containing relatively few saturated pixels for processing. Starting with small sized regions, and progressively increasing the size, we solve a sparsity promoting constrained ℓ1minimization problem for each selected region to recover the saturated pixels. Moreover, we provide simulation results to show the effectiveness of our algorithm. Hassan Mansour, Rayan Saab, Panos Nasiopoulos, Rabab K. Ward |
ICASSP | 1 |
| 2010 | Visually-favorable tone-mapping with high compression performanceabstractWe develop a tone-mapping operator (TMO) that considers the perceptual quality of the tone-mapped image together with the compression efficiency. The proposed TMO is formulated as an optimization problem that incorporates statistical models of i) the quality of the tone-mapped image given a desired TMO, ii) the base layer bit-rate and iii) the enhancement layer bit-rate. The results show that our method achieves high coding gain while maintaining good quality tone-mapped images. Zicong Mai, Hassan Mansour, Panos Nasiopoulos, Rabab K. Ward |
ICIP | 2 |
| 2010 | HDR image construction from multi-exposed stereo LDR imagesabstractIn this paper, we present an algorithm that generates high dynamic range (HDR) images from multi-exposed low dynamic range (LDR) stereo images. The vast majority of cameras in the market only capture a limited dynamic range of a scene. Our algorithm first computes the disparity map between the stereo images. The disparity map is used to compute the camera response function which in turn results in the scene radiance maps. A refinement step for the disparity map is then applied to eliminate edge artifacts in the final HDR image. Existing methods generate HDR images of good quality for still or slow motion scenes, but give defects when the motion is fast. Our algorithm can deal with images taken during fast motion scenes and tolerate saturation and radiometric changes better than other stereo matching algorithms. Hassan Mansour, Rabab K. Ward |
ICIP | 2 |
| 2010 | On-the-fly tone mapping for backward-compatible high dynamic range image/video compressionabstractIn this paper, we propose a real-time tone-mapping scheme for backward compatible high dynamic range (HDR) video compression. The appropriate choice of a tone-mapping operator (TMO) can significantly improve the HDR quality reconstructed from a low dynamic range (LDR) version. We develop a statistical model that approximates the mean square error (MSE) distortion resulting from the combined processes of tone-mapping and compression. Using this model, we formulate a numerical optimization problem to find the tone-curve that minimizes the expected MSE in the reconstructed HDR sequence. We then simplify the developed model in order to reduce the computational complexity of the optimization problem to a closed-form solution. Performance evaluations show that the proposed methods provide superior performance in terms of HDR MSE and SSIM compared to existing tone-mapping schemes. It is also shown that the LDR image quality resulting from the proposed methods matches that produced by perceptually-based TMOs. Zicong Mai, Hassan Mansour, Rafal Mantiuk, Panos Nasiopoulos, Rabab K. Ward, Wolfgang Heidrich |
ISCAS | 2 |
| 2008 | Joint media-channel aware unequal error protection for wireless scalable video streamingabstractIn this paper, we propose a joint source-channel unequal error protection scheme for scalable video streaming over capacity constrained high speed packet access (HSPA) networks. Conventional link adaptation schemes in HSPA networks use the modulation and coding scheme (MCS) that achieves a preset channel frame error rate. Our scheme utilizes video priority information along with channel quality information to set the channel coding rate that maximizes the cumulative coding rate of channel coding and application layer unequal error protection. Performance evaluations show that that under the same constraints our scheme results in an average performance improvement of 0.5dB in video PSNR for different channel conditions and different video sequences. Hassan Mansour, Panos Nasiopoulos, Vikram Krishnamurthy |
ICASSP | 1 |
| 2008 | Rate and distortion modeling of medium grain scalable video codingabstractScalability in video coding is becoming the primary choice for providing quality of service (QoS) guarantees in wireless video communication. In this paper, we develop real-time rate and distortion prediction models for medium grained scalable (MGS) coded video streams. These models allow mobile video encoders to predict the packet size and corresponding distortion of a video frame using only the mean absolute difference (MAD) of the motion prediction and the quantization parameter (QP). The prediction of rate and distortion measures can be used in devices with cross layer optimization capabilities to choose the combination of base and enhancement layer packets that deliver the best picture quality given channel quality information. Performance evaluations demonstrate that our models accurately predict the size and distortion of base and enhancement layer MGS packets. Hassan Mansour, Vikram Krishnamurthy, Panos Nasiopoulos |
ICIP | 1 |
| 2008 | An optimized link adaptation scheme for efficient delivery of scalable H.264 Video over IEEE 802.11nabstractIn this paper, we propose a cross-layer optimization scheme for delivery of scalable video over variable bit-rate wireless networks, in particular 802.11 based wireless local area networks (WLAN). For scalable video streaming applications, the conventional solution to reduced throughput due to channel distortions is to reduce the video bitrate by dropping the higher enhancement layers of the scalable video. We show that video quality can be improved, without adding to traffic load, when the WLAN link adaptation scheme uses a temporal fairness criterion along with scalable video distortion estimates to adjust its physical (PHY) layer modulation and coding parameters used for delivering each video layer. We formulate the problem as an optimization problem for assigning different PHY modes to different layers of scalable video under temporal fairness constrains; the solution to this problem provides a set of PHY configuration parameters that achieve the highest possible video quality while meeting the admission control constraints. Performance evaluations demonstrate the effectiveness of our method and the accuracy of the models. Yaser P. Fallah, Hassan Mansour, Panos Nasiopoulos, Hussein M. Alnuweiri |
ISCAS | 2 |
| 2008 | Real-time joint rate and protection allocation for multi-user scalable video streamingabstractIn this paper, we present a real-time joint bit-rate and error protection allocation scheme for multiple scalable video streams sharing a single downlink channel. High speed downlink packet access (HSDPA) systems allow for multiple live video streams to share a common downlink channel among multiple mobile users. However, the unreliable nature of the wireless link results in packet losses and fluctuations in the available channel capacity. This calls for flexible error protection and rate control strategies implemented at the video encoders that can respond to the variation in channel conditions. In this paper, we formulate a global optimization problem, which is solved at every frame transmission instant and minimizes the expected sum of the video frame distortions of all users by adjusting the encoding quality at the base- and enhancement-layers as well as the application layer error protection overhead used to combat packet losses. We consider frame-level unequal erasure protection (UXP) as the application layer forward error correction scheme. Performance evaluations show that compared with existing schemes our proposed scheme delivers far superior decoded video quality averaging 1.2 dB in PSNR. Hassan Mansour, Panos Nasiopoulos, Vikram Krishnamurthy |
PIMRC | 1 |
| 2008 | A Link Adaptation Scheme for Efficient Transmission of H.264 Scalable Video Over Multirate WLANsabstractIn this paper, we propose a cross-layer optimization scheme for delivery of scalable video over multirate wireless networks, in particular the popular 802.11 based wireless local area network (WLAN). The 802.11 based networks use a link adaptation mechanism in the physical layer (PHY) to maintain the reliability of transmission under varying channel conditions. When channel condition worsens, the reliability is maintained by employing more robust modulation and coding schemes, at the cost of reduced PHY bit rate. The reduced bit rate will result in lower available throughput for applications. For scalable video streaming applications, the conventional solution to this problem is to reduce the video bit rate by dropping the higher enhancement layers of the scalable video. We show in this article that the video quality can be improved, if the link adaptation scheme uses more intelligent reliability criteria and adjusts the PHY parameters used for delivering each video layer, according to the relative importance of that layer. Our scheme achieves better video quality without increasing the traffic load of the WLAN. For this purpose we present temporal fairness constraints and formulate an optimization problem for assigning different PHY modes to different layers of scalable video; the solution to this problem provides a set of PHY configuration parameters that achieve the highest possible video quality while meeting the admission control constraints in the network. Performance evaluations demonstrate that our method outperforms the existing mechanisms. Yaser P. Fallah, Hassan Mansour, Panos Nasiopoulos, Hussein M. Alnuweiri |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Channel Aware Multiuser Scalable Video Streaming Over Lossy Under-Provisioned Channels: Modeling and AnalysisabstractIn this paper, we analyze the performance of media-aware multiuser video streaming strategies in capacity limited wireless channels suffering from latency problems and packet losses. Wireless video streaming applications are characterized by their bandwidth-intensity, delay-sensitivity, and loss-tolerance. Our main contributions include (i) a rate-minimized unequal erasure protection (UXP) scheme, (ii) an analytical expression for packet delay and play-out deadline of UXP protected scalable video, (iii) a loss-distortion model for hierarchical predictive video coders with picture copy concealment, (iv) an analysis of the performance and complexity of delay-aware, capacity-aware, and optimized UXP streaming scenarios, and (v) we show that the use of unequal error protection causes a rate-constrained optimization problem to be nonconvex. Performance evaluations using a 3GPP network simulator show that, for different channel capacities and packet loss rates, delay-aware nonstationary rate-allocation streaming policies deliver significant gains which range between 1.65 dB to 2 dB in average Y-PSNR of the received video streams over delay-unaware strategies. These gains come at a cost of increasedofflinecomputation which is performed prior to the start of the streaming session or in batches during transmission and therefore, do not affect the run-time performance of the streaming system. Hassan Mansour, Vikram Krishnamurthy, Panos Nasiopoulos |
IEEE Trans. Multim. | 1 |