Trac D. Tran

dblp:02/565 · also Trac Duy Tran · DBLP profile ↗
← Back
138ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0002-0421-8416ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 109 · 9 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 2 since 2021Systems, architecture and hardware · 7 · 1 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Theory of computation · 3Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 ReLACS: Responsive Learned Adaptive Compressive Subsampling for Efficient Readout of Large-Area Tactile Skins
Dylan Poppert, Ariel Slepyan, Nitish V. Thakor, Trac D. Tran
DCC4
2026 Frequency-Aware Domain Generalization
abstract
Deep Neural Networks (DNNs) exhibit surprising zero-shot generalization and emergent phenomena across various tasks. However, the underlying mechanisms behind these behaviors remain unclear. By analyzing the perception of image frequencies by DNNs, we establish the association between the generalization behavior and frequency-aware regions. DNNs with stronger generalization exhibit wider frequency-aware regions. Therefore, we improve the generalization performance by broadening the frequency awareness. Specifically, we enable DNNs to learn the relations between high-frequency components and semantic labels through frequency decomposition and mixup. Based on hierarchical feature alignment, we allow larger submodels to guide the frequency awareness of smaller submodels. Beyond training, we ensemble submodels to extract features from different frequency bands to enrich DNNs' frequency awareness during inference. We validate the effectiveness of our proposed method in image classification and object detection tasks in single and multi-source domain generalization scenarios. We also demonstrate the plug-and-play scalability of our method across existing approaches and different DNNs.
Xiang Xiang 0001, Trac D. Tran
IEEE Trans. Image Process.4
2025 SINR: Sparsity Driven Compressed Implicit Neural Representations
abstract
Implicit Neural Representations (INRs) are increasingly recognized as a versatile data modality for representing discretized signals, offering benefits such as infinite query resolution and reduced storage requirements. Existing signal compression approaches for INRs typically employ one of two strategies: 1. direct quantization with entropy coding of the trained INR; 2. deriving a latent code on top of the INR through a learnable transformation. Thus, their performance is heavily dependent on the quantization and entropy coding schemes employed. In this paper, we introduce SINR, an innovative compression algorithm that leverages the patterns in the vector spaces formed by weights of INRs. We compress these vector spaces using a high-dimensional sparse code within a dictionary. Further analysis reveals that the atoms of the dictionary used to generate the sparse code do not need to be learned or transmitted to successfully recover the INR weights. We demonstrate that the proposed approach can be integrated with any existing INR-based signal compression technique. Our results indicate that SINR achieves substantial reductions in storage requirements for INRs across various configurations, outperforming conventional INR-based compression baselines. Furthermore, SINR maintains high-quality decoding across diverse data modalities, including images, occupancy fields, and Neural Radiance Fields.
Dhananjaya Jayasundara, Sudarshan Rajagopalan, Yasiru Ranasinghe, Trac D. Tran, Vishal M. Patel
CVPR4
2025 Compressive Subsampling for Scalable Tactile Skin
abstract
Real-time robotic control relies on high-speed tactile arrays, but increasing the number of sensing pixels to cover large areas often leads to greater scanning delays, with readout speeds for large arrays rarely exceeding 100 Hz. To overcome this restriction, we developed compressive tactile subsampling methods that take advantage of spatial patterns in tactile data. By sampling fewer pixels in each frame and reconstructing the tactile signal using a learned tactile dictionary, these methods enable quicker readout. Using a$32\times 32$tactile sensor array, we evaluated classification accuracy and reconstruction error for tactile interactions with 30 daily and 3D printed objects using a robotic arm. Compared to traditional raster scanning, our method produced 18 times faster frame rates while maintaining minimal reconstruction and classification error. By implementing this scalable technique into software, low-cost tactile arrays may be transformed and robots can attain high-resolution, high-speed touch sensing across their bodies. More details in our preprint [1].
Ariel Slepyan, Trac D. Tran, Nitish V. Thakor
DCC3
2025 Live Demonstration: Compressive Subsampling for High-Speed Large-Area Tactile Sensing
abstract
This demonstration showcases how compressive subsampling enhances the temporal resolution of large-area tactile sensor arrays, achieving high spatiotemporal fidelity in systems previously limited by slow response times of raster scanning. By applying compressive subsampling, the tactile system can track dynamic interactions, such as a tennis ball ricochet, soft object deformation, and the impact of a fired foam projectile. Participants can interact with several tactile sensor arrays in compressive subsampling and traditional raster-scan modes, providing a direct comparison. Interactive experiments include: (1) indenting soft, rapidly deforming objects, (2) bouncing a tennis ball to observe impact angles, and (3) targeting a sensor ‘dartboard’ with a NERF gun. These hands-on experiences allow visitors to appreciate the ease of implementing compressive subsampling with standard tactile arrays and reveal its potential for high-temporal-resolution applications in robotics and tactile sensing.ISCAS Track: Sensory Circuits and Systems
Ariel Slepyan, Trac D. Tran, Nitish V. Thakor
ISCAS3
2022 Deep Filter Bank Regression for Super-Resolution of Anisotropic MR Brain Images
Samuel Remedios, Shuo Han 0001, Yuan Xue 0002, Aaron Carass, Trac D. Tran, Dzung L. Pham, Jerry L. Prince
MICCAI (6)5
2021 EffiScene: Efficient Per-Pixel Rigidity Inference for Unsupervised Joint Learning of Optical Flow, Depth, Camera Pose and Motion Segmentation
abstract
This paper addresses the challenging unsupervised scene flow estimation problem by jointly learning four low-level vision sub-tasks: optical flow F, stereo-depth D, camera pose P and motion segmentation S. Our key insight is that the rigidity of the scene shares the same inherent geometrical structure with object movements and scene depth. Hence, rigidity from S can be inferred by jointly coupling F, D and P to achieve more robust estimation. To this end, we propose a novel scene flow framework named EffiScene with efficient joint rigidity learning, going beyond the existing pipeline with independent auxiliary structures. In EffiScene, we first estimate optical flow and depth at the coarse level and then compute camera pose by Perspective-n-Points method. To jointly learn local rigidity, we design a novel Rigidity From Motion (RfM) layer with three principal components: (i) correlation extraction; (ii) boundary learning; and (iii) outlier exclusion. Final outputs are fused based on the rigid map MRfrom RfM at finer levels. To efficiently train EffiScene, two new losses ℒbndand ℒuncare designed to prevent trivial solutions and to regularize the flow boundary discontinuity. Extensive experiments on scene flow benchmark KITTI show that our method is effective and significantly improves the state-of-the-art approaches for all sub-tasks, i.e. optical flow (5.19→4.20), depth estimation (3.78→3.46), visual odometry (0.012→0.011) and motion segmentation (0.57→ 0.62).
Trac D. Tran, Guangming Shi
CVPR2
2021 Bayesian Massive MIMO Channel Estimation with Parameter Estimation Using Low-Resolution ADCs
abstract
In order to reduce hardware complexity and power consumption, massive multiple-input multiple-output (MIMO) systems employ low-resolution analog-to-digital converters (ADCs) to acquire quantized measurements y. This poses new challenges to the channel estimation problem, and the sparse prior on the channel coefficient vector x in the angle domain is often used to compensate for the information lost during quantization. By interpreting the sparse prior from a Bayesian perspective, we can assume x follows certain sparse prior distribution and recover it using approximate message passing (AMP). However, the distribution parameters are unknown in practice and need to be estimated. Due to the increased computational complexity in the quantization noise model, previous works either use an approximated noise model or manually tune the noise distribution parameters. In this paper, we treat both signals and parameters as random variables and recover them jointly within the AMP framework. The proposed approach leads to a much simpler parameter estimation method, allowing us to work with the quantization noise model directly. Experimental results show that the proposed approach achieves state-of-the-art performance under various noise levels and does not require parameter tuning, making it a practical and maintenance-free approach for channel estimation.
Deqiang Qiu, Trac D. Tran
ICASSP3
2021 A Scale Invariant Measure of Flatness for Deep Network Minima
abstract
It has been empirically observed that the flatness of minima obtained from training deep networks seems to correlate with better generalization. However, for deep networks with positively homogeneous activations, most measures of flatness are not invariant to rescaling of the network parameters. This means that the measure of flatness can be made as small or as large as possible through rescaling, rendering the quantitative measures meaningless. In this paper we show that for deep networks with positively homogenous activations, these rescalings constitute equivalence relations, and that these equivalence relations induce a quotient manifold structure in the parameter space. Using an appropriate Riemannian metric, we propose a Hessian-based measure for flatness that is invariant to rescaling and perform simulations to empirically verify our claim. Finally we perform experiments to verify that our flatness measure correlates with generalization by using minibatch stochastic gradient descent with different batch sizes to find deep network minima with different generalization properties.
Akshay Rangamani, Nam H. Nguyen, Dzung T. Phan, Sang (Peter) Chin, Trac D. Tran
ICASSP6
2021 Optical Flow Estimation Via Motion Feature Recovery
abstract
Optical flow estimation with occlusion or large displacement is a problematic challenge due to the lost of corresponding pixels between consecutive frames. In this paper, we discover that the lost information is related to a large quantity of motion features (more than 40%) computed from the popular discriminative cost-volume feature would completely vanish due to invalid sampling, leading to the low efficiency of optical flow learning. We call this phenomenon the Vanishing Cost Volume Problem. Inspired by the fact that local motion tends to be highly consistent within a short temporal window, we propose a novel iterative Motion Feature Recovery (MFR) method to address the vanishing cost volume via modeling motion consistency across multiple frames. In each MFR iteration, invalid entries from original motion features are first determined based on the current flow. Then, an efficient network is designed to adaptively learn the motion correlation to recover invalid features for lost-information restoration. The final optical flow is then decoded from the recovered motion features. Experimental results on Sintel and KITTI show that our method achieves state-of-the-art performances. In fact, MFR currently ranks second on Sintel public website.
Guangming Shi, Trac D. Tran
ICIP3
2021 Joint Down-Range and Cross-Range RFI Suppression in Ultra-Wideband SAR
abstract
Radio frequency interference (RFI) is a critical problem for ultra-wideband synthetic aperture radars (UWB SAR) because the VHF/UHF band used by them is shared by other systems as well. A number of solutions have been proposed over the years. Recently, sparsity and low-rank estimation-based solutions were shown to perform better than traditional methods such as adaptive notch filtering. These algorithms model the SAR signal to be sparse and the RFI to be either sparse or low-ranked in nature and solve an optimization problem to estimate the SAR signal and RFI simultaneously. Algorithms in this class share the common characteristic that the SAR signal sparsity is captured by modeling each data vector as a linear combination of shifted SAR pulses. This data model addresses the structure of SAR signals in the down-range direction, but the inter-aperture cross-range structure, i.e., the fact that SAR signals add coherently across the cross-range, has been completely ignored. In this work, we incorporate this “global” 2-D structure of the SAR data into the RFI mitigation problem. Two algorithms are proposed: 1) (2-D) sparse SAR and sparse RFI estimation and 2) (2-D) sparse SAR and low-rank RFI estimation. The experimental results demonstrate that the 2-D model does a much better job of capturing the sparsity of SAR and the 2-D algorithms consistently perform better than the “local” 1-D algorithms. The level of improvement rises significantly in challenging cases-when the noise level and/or the number of RFI bands increases. Experiments are conducted extensively on simulated data sets as well as real SAR and RFI data sets collected by the U.S. Army Research Laboratory (ARL) to validate the proposed framework.
Sonia Joy, Lam H. Nguyen, Trac D. Tran
IEEE Trans. Geosci. Remote. Sens.3
2020 Supervised Deep Sparse Coding Networks for Image Classification
abstract
In this paper, we propose a novel deep sparse coding network (SCN) capable of efficiently adapting its own regularization parameters for a given application. The network is trained end-to-end with a supervised task-driven learning algorithm via error backpropagation. During training, the network learns both the dictionaries and the regularization parameters of each sparse coding layer so that the reconstructive dictionaries are smoothly transformed into increasingly discriminative representations. In addition, the adaptive regularization also offers the network more flexibility to adjust sparsity levels. Furthermore, we have devised a sparse coding layer utilizing a 'skinny' dictionary. Integral to computational efficiency, these skinny dictionaries compress the high dimensional sparse codes into lower dimensional structures. The adaptivity and discriminability of our fifteen-layer sparse coding network are demonstrated on five benchmark datasets, namely Cifar-10, Cifar-100, STL-10, SVHN and MNIST, most of which are considered difficult for sparse coding models. Experimental results show that our architecture overwhelmingly outperforms traditional one-layer sparse coding architectures while using much fewer parameters. Moreover, our multilayer architecture exploits the benefits of depth with sparse coding's characteristic ability to operate on smaller datasets. In such data-constrained scenarios, our technique demonstrates highly competitive performance compared to the deep neural networks.
Xiaoxia Sun, Nasser M. Nasrabadi, Trac D. Tran
IEEE Trans. Image Process.3
2019 JOBS: Joint-Sparse Optimization from Bootstrap Samples
abstract
Classical sparse regression based on ℓ1minimization solves the least squares problem with all available measurements via sparsity-promoting regularization. In challenging practical applications with high levels of noise and missing or adversarial samples, solving the problem using all measurements simultaneously may fail. In this paper, we propose a robust global sparse recovery strategy, named JOBS, which uses bootstrap samples of measurements to improve sparse regression in difficult cases. K measurement vectors are generated from the original pool of m measurements via bootstrapping, with each bootstrap sample containing L elements, and then a joint-sparse constraint is enforced to ensure support consistency among multiple predictors. The final estimate is obtained by averaging over K estimators. The performance limits associated with finite bootstrap sampling ratio L/m and number of estimates K is analyzed theoretically. Simulation results validate the theoretical analysis of proper choice of (L,K) and show that the proposed method yields state-of-the-art recovery performance, outperforming ℓ1minimization and other existing bootstrap-based techniques, especially when the number of measurements are limited. With a proper choice of bootstrap sampling ratio (0.3-0.5) and a reasonably large number of estimates K (≥ 30), the SNR improvement over the baseline ℓ1-minimization algorithm can reach up to 336%.
Luoluo Liu, Sang (Peter) Chin, Trac D. Tran
ISIT3
2018 A Greedy Pursuit Algorithm for Separating Signals from Nonlinear Compressive Observations
abstract
In this paper we study the unmixing problem which aims to separate a set of structured signals from their superposition. In this paper, we consider the scenario in which the mixture is observed via nonlinear compressive measurements. We present a fast, robust, greedy algorithm called Unmixing Matching Pursuit (UnmixMP) to solve this problem. We prove rigorously that the algorithm can recover the constituents from their noisy nonlinear compressive measurements with arbitrarily small error. We compare our algorithm to the Demixing with Hard Thresholding (DHT) algorithm [1], in a number of experiments on synthetic and real data.
Sang (Peter) Chin, Trac D. Tran, Dung N. Tran, Akshay Rangamani
ICASSP2
2018 A Deep Learning Based Alternative to Beamforming Ultrasound Images
abstract
Deep learning methods are capable of performing sophisticated tasks when applied to a myriad of artificial intelligent (AI) research fields. In this paper, we introduce a novel approach to replace the inherently flawed beamforming step during ultrasound image formation by applying deep learning directly to RF channel data. Specifically, we pose the ultrasound beamforming process as a segmentation problem and apply a fully convolutional neural network architecture to segment anechoic cysts from surrounding tissue. We train our network on a dataset created using the Field II ultrasound simulation software to simulate plane wave imaging with a single insonification angle. We demonstrate the success of our architecture in extracting tissue information directly from the raw channel data, which completely bypasses the beamforming step that would otherwise require multiple insonification angles for plane wave imaging. Our simulated results produce mean Dice coefficient of 0.98 ± 0.02, when measuring the overlap between ground truth cyst locations and cyst locations determined by the network. The proposed approach is promising for developing dedicated deep-learning networks to improve the real-time ultrasound image formation process.
Arun Asokan Nair, Trac D. Tran, Austin Reiter, Muyinatu A. Lediju Bell
ICASSP2
2018 Recovery of UWB Radar Signals in Spectrally Restricted Environments
abstract
This paper presents a novel technique to recover the missing spectral information due to RF spectral restriction in ultra-wideband (UWB) synthetic aperture radar (SAR) imaging. We address a critical problem in UWB radar imaging: radar transmission is prohibited in reserved frequency bands specified by local frequency management agencies. We model the problem in a compressed sensing setup where a time-sparse received radar signal at each aperture is collected in the incoherent frequency domain using a stepped-frequency implementation. Using electromagnetic (EM) finite-difference, time-domain (FDTD) data of various targets and clutter objects computed at all viewing aspect angles, we show that the proposed technique can successfully recover the relevant information on targets of interest even from a large percentage of missing frequency bands.
Lam H. Nguyen, Trac D. Tran
ICIP2
2018 Supervised Deep Sparse Coding Networks
abstract
In this paper, we present the deep sparse coding network (DSCN) - a novel deep learning framework that encodes intermediate representations with nonnegative sparse coding. DSCN is constructed from a cascade of bottleneck modules, each of which consists of two sparse coding layers with relatively wide and slim dictionaries that are specialized to produce high dimensional discriminative features and low dimensional clustered representations, respectively. During training, all dictionaries at all depth levels along with all regularization parameters are optimized jointly with an end-to-end supervised learning algorithm based on multilevel optimization. The effectiveness of the proposed DSCN with seven bottleneck modules11Consisting 14 sparse coding layers. is verified on several popular benchmark datasets Remarkably, with few parameters to learn, our SCN achieves 5.81 % and 19.93% classification error rate on CIFAR-10 and CIFAR-100, respectively.
Xiaoxia Sun, Nasser M. Nasrabadi, Trac D. Tran
ICIP3
2018 S3D: Stacking Segmental P3D for Action Quality Assessment
abstract
Action quality assessment is crucial in areas of sports, surgery and assembly line where action skills can be evaluated. In this paper, we propose the Segment-based P3D-fused network S3D built-upon ED-TCN and push the performance on the UNLV-Dive dataset by a significant margin. We verify that segment-aware training performs better than full-video training which turns out to focus on the water spray. We show that temporal segmentation can be embedded with few efforts.
Xiang Xiang 0001, Austin Reiter, Gregory D. Hager, Trac D. Tran
ICIP5
2018 Using Deep Learning to Extract Scenery Information in Real Time Spatiotemporal Compressed Sensing
abstract
One of the problems of real time compressed sensing system is the computational cost of the reconstruction algorithms. It is especially problematic for close loop sensory applications where the sensory parameters needs to be constantly adjust to adapt to a dynamic scene. Through a preliminary experiment with MNIST dataset, we showed that we can extract some scene information (object recognition, scene movement direction and speed) based on the compressed samples using a deep convolutional neural network. It achieves 100% percent accuracy in distinguishing moving velocity, 96.22% in recognizing the digit and 90.04% in detecting moving direction after the code images are re-centered. Even though the classification accuracy drops slightly compared to using original videos, the computational speed is two time faster than classification on videos directly. This method also eliminates the need for sparse reconstruction prior to classification.
Xiao Wang 0028, Jie Zhang 0063, Trac D. Tran, Sang (Peter) Chin, Ralph Etienne-Cummings
ISCAS4
2018 Sparse Coding and Autoencoders
abstract
In this work we study the landscape of squared loss of an Autoencoder when the data generative model is that of “Sparse Coding”/“Dictionary Learning”. The neural net considered is an$\mathbb{R}^{n}\rightarrow \mathbb{R}^{n}$mapping and has a single ReLU activation layer of size$h > n$. The net has access to vectors$y\in \mathbb{R}^{n}$obtained as$y=A^{\ast}x^{\ast}$where$x^{\ast}\in \mathbb{R}^{h}$are sparse high dimensional vectors and$A^{\ast}\in \mathbb{R}^{n\times h}$is an overcomplete incoherent matrix. Under very mild distributional assumptions on$x^{\ast}$, we prove that the norm of the expected gradient of the squared loss function is asymptotically (in sparse code dimension) negligible for all points in a small neighborhood of$A^{\ast}$. This is supported with experimental evidence using synthetic data. We conduct experiments to suggest that$A^{\ast}$sits at the bottom of a well in the landscape and we also give experiments showing that gradient descent on this loss function gets columnwise very close to the original dictionary even with far enough initialization. Along the way we prove that a layer of ReLU gates can be set up to automatically recover the support of the sparse codes. Since this property holds independent of the loss function we believe that it could be of independent interest. A full version of this paper is accessible at: https://arxiv.org/abs/1708.03735
Akshay Rangamani, Anirbit Mukherjee, Amitabh Basu, Ashish Arora, Tejaswini Ganapathi, Sang (Peter) Chin, Trac D. Tran
ISIT7
2018 Linear Disentangled Representation Learning for Facial Actions
abstract
Limited annotated data available for the recognition of facial expression and particularly action units makes it hard to train a deep network which can learn disentangled invariant features. However, a supervised linear model is undemanding in terms of training data. In this paper, we propose an elegant linear model to untangle facial actions from expressive face videos which contain a mixture of linearly-representable attributes. Previous attempts require an explicit decoupling of identity and expression which is practically inexact. Instead, we exploit the low-rank property across frames to implicitly subtract the intrinsic neutral face, which are modeled jointly with sparse representation only on the residual expression components. On CK+, our one-shot C-HiSLR on raw-face pixel-intensities performs far more competitive than conventional shape+SVM models with landmark detection and two-stepped SRC of the same type yet applied on manually prepared expression components. It is also comparable with the piecewise linear model DCS and temporal models, such as CRF and Bayes nets. We apply it to action unit (AU) recognition on MPI-VDB achieving a decent performance. As expression is a mixture of AUs, the result gives hopes of approximating an expression using a piecewise linear model.
Xiang Xiang 0001, Trac D. Tran
IEEE Trans. Circuits Syst. Video Technol.2
2017 Sparse signal recovery using generalized approximate message passing with built-in parameter estimation
abstract
The generalized approximate message passing (GAMP) algorithm under the Bayesian setting shows significant advantages in recovering under-sampled sparse signals from corrupted observations. Compared to conventional convex optimization methods, it has a much lower complexity and is computationally tractable. Under the GAMP framework, the sparse signal and the observation are viewed to be generated according to some pre-specified probability distributions in the input and output channels. However, the parameters of the distributions are usually unknown in practice and need to be decided. In this paper, we propose an extended GAMP algorithm with built-in parameter estimation (PE-GAMP). Specifically, PE-GAMP treats the parameters as unknown random variables with simple priors and jointly estimates them with the sparse signals along the recovery process. Sparse signal recovery experiments confirm PE-GAMP's convergence behavior and show that its performance matches the oracle GAMP algorithm that has the knowledge of the true parameter values.
Trac D. Tran
ICASSP2
2017 A comprehensive performance comparison of RFI mitigation techniques for UWB radar signals
abstract
This paper presents a comprehensive benchmark comparison in objective as well as subjective performances of radio-frequency interference (RFI) suppression/extraction techniques for ultra-wideband (UWB) signals in synthetic aperture radar (SAR) imaging applications. In this study, we employ two sets of UWB SAR signals: one simulated from a step-frequency radar setup, whereas the other is collected on the testing field in a real-world setup from the U.S. Army Research Laboratory. Similarly, our RFI experiments involve two RFI data sets: one is simulated from a collection of randomly generated frequency bands and the other is the RFI data collected in a real-world environment with the radar receiving antenna pointing toward Washington DC. These SAR and RFI data sets represent four diverse experimental setups where we can carefully benchmark the denoising performance of several popular RFI-mitigation techniques in the current literature based on notch-filtering, principal component analysis (PCA), model-based sparse recovery, and simultaneous low-rank and sparse recovery or robust PCA (RPCA). We validate that RPCA and model-based sparse recovery consistently yields the best overall RFI separation performance on a wide range of settings in all data sets.
Lam H. Nguyen, Trac D. Tran
ICASSP2
2017 A provable nonconvex model for factoring nonnegative matrices
abstract
We study the Nonnegative Matrix Factorization problem which approximates a nonnegative matrix by a low-rank factorization. This problem is particularly important in Machine Learning, and finds itself in a large number of applications. Unfortunately, the original formulation is ill-posed and NP-hard. In this paper, we propose a row sparse model based on Row Entropy Minimization to solve the NMF problem under separable assumption which states that each data point is a convex combination of a few distinct data columns. We utilize the concentration of the entropy function and the ℓ∞norm to concentrate the energy on the least number of latent variables. We prove that under the separability assumption, our proposed model robustly recovers data columns that generate the dataset, even when the data is corrupted by noise. We empirically justify the robustness of the proposed model and show that it is significantly more robust than the state-of-the-art separable NMF algorithms.
Dung N. Tran, Sang (Peter) Chin, Trac D. Tran
ICASSP3
2017 Regularizing face verification nets for pain intensity regression
abstract
Limited labeled data are available for the research of estimating facial expression intensities. For instance, the ability to train deep networks for automated pain assessment is limited by small datasets with labels of patient-reported pain intensities. Fortunately, fine-tuning from a data-extensive pre-trained domain, such as face verification, can alleviate this problem. In this paper, we propose a network that fine-tunes a state-of-the-art face verification network using a regularized regression loss and additional data with expression labels. In this way, the expression intensity regression task can benefit from the rich feature representations trained on a huge amount of data for face verification. The proposed regularized deep regressor is applied to estimate the pain expression intensity and verified on the widely-used UNBC-McMaster Shoulder-Pain dataset, achieving the state-of-the-art performance. A weighted evaluation metric is also proposed to address the imbalance issue of different pain intensities.
Feng Wang 0015, Xiang Xiang 0001, Trac D. Tran, Austin Reiter, Gregory D. Hager, Harry Quon, Jian Cheng 0003, Alan L. Yuille
ICIP4
2017 Supervised hashing with jointly learning embedding and quantization
abstract
Compared with unsupervised hashing, supervised hashing commonly illustrates better accuracy in many real applications by leveraging semantic (label) information. However, it is tough to solve the supervised hashing problem directly because it is essentially a discrete optimization problem. Some other works try to solve the discrete optimization problem directly using binary quadratic programming, but they are typically too complicated and time-consuming while some supervised hashing methods have to solve a relaxed continuous optimization problem by dropping the discrete constraints. However, these methods typically suffer from poor performance due to the errors caused by the relaxation manner. In this paper based on the general two-step framework: learning binary embedded codes and learning hash functions, we propose a new method to solve the problem introduced by relaxing the cost function. Inspired by the property of rotation invariance of learning embedding features, our method tries to jointly learn similarity-preserving representation and rotation transformation for better quantization alternatively. In experiments, our method shows significant improvement. Compared with the methods based on discrete optimization our methods obtains the competitive performance and even achieves the state-of-the-art performance in some image retrieval applications.
Feng Wang 0015, Xiang Xiang 0001, Trac D. Tran
ICIP4
2017 Live demonstration: A compact all-CMOS spatiotemporal compressed sensing video camera
abstract
A compact all-CMOS spatiotemporal compressed sensing (CS) video camera is demonstrated. This CS-based framework [1], implemented on integrated circuits, is able to achieve 20-fold reduction in the readout speed and consumes only 14μW to provide 100 fps videos. Taking advantage of dictionary learning and sparse recovery, this prototype image sensor (127×90 pixels) can reconstruct 100 fps videos from the coded images sampled at 5 fps.
Jie Zhang 0063, Chetan Singh Thakur, John M. Rattray, Sang (Peter) Chin, Trac D. Tran, Ralph Etienne-Cummings
ISCAS6
2017 Deep Image-to-Image Recurrent Network with Shape Basis Learning for Automatic Vertebra Labeling in Large-Scale 3D CT Volumes
Dong Yang 0005, Daguang Xu, Shaohua Kevin Zhou, Zhoubing Xu, Mingqing Chen, Jin Hyeong Park, Sasa Grbic, Trac D. Tran, Sang (Peter) Chin, Dimitris N. Metaxas, Dorin Comaniciu
MICCAI (3)9
2016 Partial face recognition: A sparse representation-based approach
abstract
Partial face recognition is a problem that often arises in practical settings and applications. We propose a sparse representation-based algorithm for this problem. Our method firstly trains a dictionary and the classifier parameters in a supervised dictionary learning framework and then aligns the partially observed test image and seeks for the sparse representation with respect to the training data alternatively to obtain its label. We also analyze the performance limit of sparse representation-based classification algorithms on partial observations. Finally, face recognition experiments on the popular AR data-set are conducted to validate the effectiveness of the proposed method.
Luoluo Liu, Trac D. Tran, Sang (Peter) Chin
ICASSP2
2016 Sparse coding with fast image alignment via large displacement optical flow
abstract
Sparse representation-based classifiers have shown outstanding accuracy and robustness in image classification tasks even with the presence of intense noise and occlusion. However, it has been discovered that the performance degrades significantly either when test image is not aligned with the dictionary atoms or the dictionary atoms themselves are not aligned with each other, in which cases the sparse linear representation assumption fails. In this paper, having both training and test images misaligned, we introduce a novel sparse coding framework that is able to efficiently adapt the dictionary atoms to the test image via large displacement optical flow. In the proposed algorithm, every dictionary atom is automatically aligned with the input image and the sparse code is then recovered using the adapted dictionary atoms. A corresponding supervised dictionary learning algorithm is also developed for the proposed framework. Experimental results on digit datasets recognition verify the efficacy and robustness of the proposed algorithm.
Xiaoxia Sun, Nasser M. Nasrabadi, Trac D. Tran
ICASSP3
2016 Low-rank matrices recovery via entropy function
abstract
The low-rank matrix recovery problem consists of reconstructing an unknown low-rank matrix from a few linear measurements, possibly corrupted by noise. One of the most popular method in low-rank matrix recovery is based on nuclear-norm minimization, which seeks to simultaneously estimate the most significant singular values of the target low-rank matrix by adding a penalizing term on its nuclear norm. In this paper, we introduce a new method that requires substantially fewer measurements needed for exact matrix recovery compared to nuclear norm minimization. The proposed optimization program utilizes a sparsity promoting regularization in the form of the entropy function of the singular values. Numerical experiments on synthetic and real data demonstrates that the proposed method outperforms stage-of-the-art nuclear norm minimization algorithms.
Dung N. Tran, Sang (Peter) Chin, Trac D. Tran
ICASSP4
2016 Dimensionality reduction for image classification via mutual information maximization
abstract
Low-level feature encoding combined with Spatial Pyramid Matching (SPM) is widely adopted in the image classification system nowadays to extract features, which are usually high-dimensional. This not only makes the classification problem computationally prohibitive, but also raises other issues, such as the “curse of dimensionality”. In this paper we present supervised dimensionality reduction (DR) approaches that try to maximize the mutual information (MI) between the class labels and the low-dimensional features through an orthonormal linear projection. In addition to reduced computational load, using the extracted low-dimensional features during classification also leads to higher accuracies, as evident from a series of image classification experiments in this paper.
Trac D. Tran
ICIP2
2016 Sparse signal recovery based on nonconvex entropy minimization
abstract
We propose a new sparsity-promoting objective function to be used in sparse signal recovery. Specifically, the objective is an entropy function l1 defined on the sparse signal x. Compared to the conventional l1, it is a nonconvex function and the optimization problem can be solved based on the fast iterative shrinkage thresholding algorithm (FISTA). Experiments on 1-dimensional sparse signal recovery and 2-dimensional real image recovery show that minimizing lp favors sparse solutions, and that it could recover sparse signals better than the convex l1 norm minimization and the nonconvex lp-norm minimization.
Dung N. Tran, Trac D. Tran
ICIP3
2015 Targeted Dot Product Representation for Friend Recommendation in Online Social Networks
abstract
In this paper, we develop Targeted Dot Product Representation (TarDPR), a DPR-based feature selection and combination framework for friend recommendation in online social networks (OSNs). Our approach modifies conventional DPR techniques and makes itself applicable to OSNs by focusing on computing a consistent representation while minimizing unnecessary suggestions made outside these interested regions. A notable property of TarDPR is its ability to effectively incorporate different types of social features and produce new meaningful features that help competitive approaches to significantly improve their recommendation quality. We derive an iterative algorithm for TarDPR that is supported by mathematical analysis, and is efficient on large social traces. To certify the usability of our approach, we conduct empirical experiments on real social traces including Facebook and Foursquare social networks. The competitive experimental results show that TarDPR achieves up to 15% improvement in comparison with other competitive methods. These results consequently confirm the efficacy of our suggested framework.
Minh Dao, Akshay Rangamani, Sang (Peter) Chin, Nam P. Nguyen, Trac D. Tran
ASONAM5
2015 Multi-sensor classification via sparsity-based representation with low-rank interference
abstract
In this paper, we propose a general collaborative sparse representation framework for multi-sensor classification which exploits correlation as well as complementary information among homogeneous and heterogeneous sensors while simultaneously extracting the low-rank interference term. Specifically, we observe that incorporating the noise or interfered signal as a low-rank component is essential in a multi-sensor problem when multiple co-located sources/sensors simultaneously record the same physical event. We further extend our frameworks to kernelized models which rely on sparsely representing a test sample in terms of all the training samples in a feature space induced by a kernel function. A fast and efficient algorithm based on alternative direction method is proposed where its convergence to optimal solution is guaranteed. Extensive experiments are conducted on a real data set for a multi-sensor classification problem focusing on discriminating between human and animal footsteps. Results are compared with the conventional classifiers and existing sparsity-based representation methods to verify the effectiveness of our proposed models.
Minh Dao, Nasser M. Nasrabadi, Trac D. Tran
ICASSP3
2015 Nonnegative matrix factorization with gradient vertex pursuit
abstract
Nonnegative Matrix Factorization (NMF), defined as factorizing a nonnegative matrix into two nonnegative factor matrices, is a particularly important problem in machine learning. Unfortunately, it is also ill-posed and NP-hard. We propose a fast, robust, and provably correct algorithm, namely Gradient Vertex Pursuit (GVP), for solving a well-defined instance of the problem which results in a unique solution: there exists a polytope, whose vertices consist of a few columns of the original matrix, covering the entire set of remaining columns. Our algorithm is greedy: it detects, at each iteration, a correct vertex until the entire polytope is identified. We evaluate the proposed algorithm on both synthetic and real hyperspectral data, and show its superior performance compared with other state-of-the-art greedy pursuit algorithms.
Dung N. Tran, Sang (Peter) Chin, Trac D. Tran
ICASSP4
2015 Hierarchical Sparse and Collaborative Low-Rank representation for emotion recognition
abstract
In this paper, we design a Collaborative-Hierarchical Sparse and Low-Rank (C-HiSLR) model that is natural for recognizing human emotion in visual data. Previous attempts require explicit expression components, which are often unavailable and difficult to recover. Instead, our model exploits the low-rank property to subtract neutral faces from expressive facial frames as well as performs sparse representation on the expression components with group sparsity enforced. For the CK+ dataset, C-HiSLR on raw expressive faces performs as competitive as the Sparse Representation based Classification (SRC) applied on manually prepared emotions. Our C-HiSLR performs even better than SRC in terms of true positive rate.
Xiang Xiang 0001, Minh Dao, Gregory D. Hager, Trac D. Tran
ICASSP4
2015 Local sensing with global recovery
abstract
In this paper, we study Locally Compressed Sensing for images, where sampling process is allowed to be performed on arbitrary local regions of the images. We propose a fast and efficient reconstruction algorithm which utilizes local structures of images. Several numerical experiments on real images demonstrates that our algorithm yields better reconstruction quality than existing techniques at much lower computational complexity and memory requirement.
Dung N. Tran, Duyet N. Tran, Sang (Peter) Chin, Trac D. Tran
ICIP4
2015 Estimation and extraction of radio-frequency interference from ultra-wideband radar signals
abstract
This paper presents a simple adaptive framework for robust separation and extraction of multiple sources of radio-frequency interference (RFI) from raw ultra-wideband (UWB) radar signals in challenging bandwidth management environments. RFI sources pose critical challenges for UWB systems since (i) RFI often occupies a wide range of the radar's operating frequency spectrum; (ii) RFI might have significant power; and (iii) RFI signals are difficult to predict and model due to their non-stationary nature as well as the complexity of various communication devices. Our proposed framework involves a standard RFI detection step that operates directly on previously-collected contaminated radar signals to identify RFI-dominant frequency sub-bands. This vital information is then applied to construct an RFI dictionary with various sinusoidal patterns spanning these RFI bands. We then employ a sparsity-driven optimization to estimate and then extract RFI from the received radar signals. Our method can be considered as a de-noising preprocessing stage for raw radar signals prior to image formation and other follow-up tasks. Recovery results from extensive simulated as well as real-world UWB synthetic aperture radar (SAR) data sets illustrate the robustness and effectiveness of our framework.
Lam H. Nguyen, Trac D. Tran
IGARSS2
2015 An unsupervised dictionary learning algorithm for neural recordings
abstract
To meet the growing demand of wireless and power efficient neural recordings systems, we demonstrate an unsupervised dictionary learning algorithm in Compressed Sensing (CS) framework which can be implemented in VLSI systems. Without prior label information of neural spikes, we extend our previous work to unsupervised learning and construct a dictionary with discriminative structures for spike sorting. To further improve the reconstruction and classification performance, we proposed a joint prediction to determine the class of neural spikes in dictionary learning. When the neural spikes is compressed 50 times, our approach can achieve an average gain of 2 dB and 15 percentage units over state-of-the-art of CS approaches in terms of the reconstruction quality and classification accuracy respectively.
Jie Zhang 0063, Yuanming Suo, Dung N. Tran, Ralph Etienne-Cummings, Sang (Peter) Chin, Trac D. Tran
ISCAS7
2015 Iterative Convex Refinement for Sparse Recovery
abstract
In this letter, we address sparse signal recovery in a Bayesian framework where sparsity is enforced on reconstruction coefficients via probabilistic priors. In particular, we focus on the setup of Yenwho employ a variant of spike and slab prior to encourage sparsity. The optimization problem resulting from this model has broad applicability in recovery and regression problems and is known to be a hard non-convex problem whose existing solutions involve simplifying assumptions and/or relaxations. We propose an approach called Iterative Convex Refinement (ICR) that aims to solve the aforementioned optimization problem directly allowing for greater generality in the sparse structure. Essentially, ICR solves a sequence of convex optimization problems such that sequence of solutions converges to a sub-optimal solution of the original hard optimization problem. We propose two versions of our algorithm: a.) an unconstrained version, and b.) with a non-negativity constraint on sparse coefficients, which may be required in some real-world problems. Experimental validation is performed on both synthetic data and for a real-world image recovery problem, which illustrates merits of ICR over state of the art alternatives.
Hojjat Seyed Mousavi, Vishal Monga, Trac D. Tran
IEEE Signal Process. Lett.3
2015 Task-Driven Dictionary Learning for Hyperspectral Image Classification With Structured Sparsity Constraints
abstract
Sparse representation models a signal as a linear combination of a small number of dictionary atoms. As a generative model, it requires the dictionary to be highly redundant in order to ensure both a stable high sparsity level and a low reconstruction error for the signal. However, in practice, this requirement is usually impaired by the lack of labeled training samples. Fortunately, previous research has shown that the requirement for a redundant dictionary can be less rigorous if simultaneous sparse approximation is employed, which can be carried out by enforcing various structured sparsity constraints on the sparse codes of the neighboring pixels. In addition, numerous works have shown that applying a variety of dictionary learning methods for the sparse representation model can also improve the classification performance. In this paper, we highlight the task-driven dictionary learning (TDDL) algorithm, which is a general framework for the supervised dictionary learning method. We propose to enforce structured sparsity priors on the TDDL method in order to improve the performance of the hyperspectral classification. Our approach is able to benefit from both the advantages of the simultaneous sparse representation and those of the supervised dictionary learning. We enforce two different structured sparsity priors, the joint and Laplacian sparsities, on the TDDL method and provide the details of the corresponding optimization algorithms. Experiments on numerous popular hyperspectral images demonstrate that the classification performance of our approach is superior to that of the sparse representation classifier with structured priors or the TDDL method.
Xiaoxia Sun, Nasser M. Nasrabadi, Trac D. Tran
IEEE Trans. Geosci. Remote. Sens.3
2015 Structured Sparse Priors for Image Classification
abstract
Model-based compressive sensing (CS) exploits the structure inherent in sparse signals for the design of better signal recovery algorithms. This information about structure is often captured in the form of a prior on the sparse coefficients, with the Laplacian being the most common such choice (leading to l1 -norm minimization). Recent work has exploited the discriminative capability of sparse representations for image classification by employing class-specific dictionaries in the CS framework. Our contribution is a logical extension of these ideas into structured sparsity for classification. We introduce the notion of discriminative class-specific priors in conjunction with class specific dictionaries, specifically the spike-and-slab prior widely applied in Bayesian sparse regression. Significantly, the proposed framework takes the burden off the demand for abundant training image samples necessary for the success of sparsity-based classification schemes. We demonstrate this practical benefit of our approach in important applications, such as face recognition and object categorization.
Umamahesh Srinivas, Yuanming Suo, Minh Dao, Vishal Monga, Trac D. Tran
IEEE Trans. Image Process.5
2014 Subspace vertex pursuit for separable non-negative matrix factorization in hyperspectral unmixing
abstract
Recently, the separability assumption turns the nonnegative matrix factorization (NMF) into a tractable problem. The assumption coincides with the pixel purity assumption and provides new insights for the hyperspectral unmixing problem. In this paper, we present a quasi-greedy algorithm for solving the problem by employing a back-tracking strategy. Unlike the current greedy methods, the proposed method can refresh the endmember index set in every iteration. Therefore, our method has two important characteristics: (i) low computational complexity comparable to state-of-the-art greedy methods but (ii) empirically enhanced robustness against noise. Finally, computer simulations on synthetic hyperspectral data demonstrate the effectiveness of the proposed method.
Qing Qu 0001, Xiaoxia Sun, Nasser M. Nasrabadi, Trac D. Tran
ICASSP4
2014 Robust joint down- and cross-range sparse recovery of raw ultra-wideband SAR data
abstract
Ultra-wideband (UWB) Synthetic Aperture Radars (SAR) often operate in environments where other systems are present or some of the frequency bands may be reserved for other purposes. In this paper we propose a novel method based on sparse representation to recover SAR signals from radio frequency interference (RFI) and missing spectral information. The new method makes use of correlation between SAR data records in the cross-range direction in addition to the correlation in down-range direction as previously proposed by us[1, 2]. We demonstrate that the proposed method yields robust recovery results in the presence of high levels of RFI (up to 20dB) and up to 75% of frequency bands missing.
Sonia Joy, Lam H. Nguyen, Trac D. Tran
ICIP3
2014 Multi-task image classification via collaborative, hierarchical spike-and-slab priors
abstract
Promising results have been achieved in image classification problems by exploiting the discriminative power of sparse representations for classification (SRC). Recently, it has been shown that the use of class-specific spike-and-slab priors in conjunction with the class-specific dictionaries from SRC is particularly effective in low training scenarios. As a logical extension, we build on this framework for multitask scenarios, wherein multiple representations of the same physical phenomena are available. We experimentally demonstrate the benefits of mining joint information from different camera views for multi-view face recognition.
Hojjat Seyed Mousavi, Umamahesh Srinivas, Vishal Monga, Yuanming Suo, Minh Dao, Trac D. Tran
ICIP6
2014 Radio-frequency interference separation and suppression from ultrawideband radar data via low-rank modeling
abstract
Radio-frequency interference (RFI) is the most common, and also the most challenging type of interference or noise source that has a direct impact on the performance of ultrawideband radar systems in various practical application settings. Existing techniques for RFI suppression either employ filtering (notching) which introduces other harmful side-effects such as side-lobe distortion and target-amplitude reduction or RFI modeling/estimation/tracking which requires complicated narrow-band modulation models or even direct RFI sniffing. In this paper, we propose a robust and adaptive technique for the separation and then suppression of RFI signals from ultra-wideband (UWB) radar data via modeling RFI as low-rank components in a joint optimization framework. More specifically, we advocate a joint sparse-and-low-rank recovery approach that simultaneously solves for (i) UWB radar signals as sparse representations with respect to a dictionary containing transmitted waveforms; and (ii) RFI signals as a low-rank structure. The proposed technique is completely adaptive with highly time-varying environments, and does not require any prior knowledge of the RFI sources (other than the low-rank assumption). Both simulated data and real-world data measured by the U.S. Army Research Laboratory (ARL) Ultra-Wideband (UWB) synthetic aperture radar (SAR) confirm that the proposed RFI separation/suppression technique successfully recovers UWB radar signals embedded in large-amplitude RFI signals.
Lam H. Nguyen, Minh Dao, Trac D. Tran
ICIP3
2014 Task-driven dictionary learning for hyperspectral image classification with structured sparsity priors
abstract
In hyperspectral pixel classification, previous research have shown that the sparse representation classifier can achieve a better performance when exploiting the neighboring test pixels through enforcing different structured sparsity priors. In this paper, we propose a supervised sparse-representation-based dictionary learning method with joint or Laplacian s-parsity priors. The proposed method has numerous advantages over the existing dictionary learning techniques. It uses a structured sparsity and provides a more robust and stable sparse coefficients. Besides, it is capable of reducing the classification error by jointly optimizing the dictionary and the classifier's parameters during the dictionary training stage.
Xiaoxia Sun, Nasser M. Nasrabadi, Trac D. Tran
ICIP3
2014 Group structured dirty dictionary learning for classification
abstract
Dictionary learning techniques have gained tremendous success in many classification problems. Inspired by the dirty model for multi-task regression problems, we proposed a novel method called group-structured dirty dictionary learning (GDDL) that incorporates the group structure (for each task) with the dirty model (across tasks) in the dictionary training process. Its benefits are two-fold: 1) the group structure enforces implicitly the label consistency needed between dictionary atoms and training data for classification; and 2) for each class, the dirty model separates the sparse coefficients into ones with shared support and unique support, with the first set being more discriminative. We use proximal operators and block coordinate decent to solve the optimization problem. GDDL has been shown to give state-of-art result on both synthetic simulation and two face recognition datasets.
Yuanming Suo, Minh Dao, Trac D. Tran, Hojjat Seyed Mousavi, Umamahesh Srinivas, Vishal Monga
ICIP3
2014 Structured Priors for Sparse-Representation-Based Hyperspectral Image Classification
abstract
Pixelwise classification, where each pixel is assigned to a predefined class, is one of the most important procedures in hyperspectral image (HSI) analysis. By representing a test pixel as a linear combination of a small subset of labeled pixels, a sparse representation classifier (SRC) gives rather plausible results compared with that of traditional classifiers such as the support vector machine. Recently, by incorporating additional structured sparsity priors, the second-generation SRCs have appeared in the literature and are reported to further improve the performance of HSI. These priors are based on exploiting the spatial dependences between the neighboring pixels, the inherent structure of the dictionary, or both. In this letter, we review and compare several structured priors for sparse-representation-based HSI classification. We also propose a new structured prior called the low-rank (LR) group prior, which can be considered as a modification of the LR prior. Furthermore, we will investigate how different structured priors improve the result for the HSI classification.
Xiaoxia Sun, Qing Qu 0001, Nasser M. Nasrabadi, Trac D. Tran
IEEE Geosci. Remote. Sens. Lett.4
2014 Abundance Estimation for Bilinear Mixture Models via Joint Sparse and Low-Rank Representation
abstract
Sparsity-based unmixing algorithms, exploiting the sparseness property of the abundances, have recently been proposed with promising performances. However, these algorithms are developed for the linear mixture model (LMM), which cannot effectively handle the nonlinear effects. In this paper, we extend the current sparse regression methods for the LMM to bilinear mixture models (BMMs), where the BMMs introduce additional bilinear terms in the LMM in order to model second-order photon scattering effects. To solve the abundance estimation problem for the BMMs, we propose to perform a sparsity-based abundance estimation by using two dictionaries: a linear dictionary containing all the pure endmembers and a bilinear dictionary consisting of all the possible second-order endmember interaction components. Then, the abundance values can be estimated from the sparse codes associated with the linear dictionary. Moreover, to exploit the spatial data structure where the adjacent pixels are usually homogeneous and are often mixtures of the same materials, we first employ the joint-sparsity (row-sparsity) model to enforce structured sparsity on the abundance coefficients. However, the joint-sparsity model is often a strict assumption, which might cause some aliasing artifacts for the pixels that lie on the boundaries of different materials. To deal with this problem, the low-rank-representation model, which seeks the lowest rank representation of the data, is further introduced to better capture the spatial data structure. Our simulation results demonstrate that the proposed algorithms provide much enhanced performance compared with state-of-the-art algorithms.
Qing Qu 0001, Nasser M. Nasrabadi, Trac D. Tran
IEEE Trans. Geosci. Remote. Sens.3
2013 Video frame interpolation via weighted robust principal component analysis
abstract
In this paper, we propose a new video frame interpolation technique by a locally-adaptive robust principal component analysis (RPCA) with weight priors. The proposed algorithm relies on two main steps: 1. the pre-processing step initializes the new frame by a simplified motion-compensated frame interpolation and assigns each pixel a confident weight based on both the difference of motion estimation and local consistency; and 2. the refinement step updates the frame by a proposed weighted robust principal component analysis (WRPCA) algorithm. Experiments demonstrate that the proposed method outperforms the state-of-the-art algorithms, both in visual quality and PSNR performance.
Minh Dao, Yuanming Suo, Sang (Peter) Chin, Trac D. Tran
ICASSP4
2013 Hyperspectral abundance estimation for the generalized bilinear model with joint sparsity constraint
abstract
In this paper, we present a novel abundance estimation method for the generalized bilinear model (GBM) via sparse representation for hyperspectral imagery. Because the GBM generalizes the linear mixture model (LMM) by introducing an additional bilinear term, our sparsity-based abundance estimation is performed by utilizing two dictionaries-a linear dictionary containing all the pure endmembers and a bilinear dictionary consisting of all the possible bilinear interaction components. Because the components within the bilinear term are also linearly combined, by employing a composite dictionary made up by the concatenation of the linear and bilinear dictionaries we can reformulate the bilinear problem in a linear sparse regression framework. In this way, the abundance values are estimated from the sparse codes only associated with the linear dictionary. To further improve the estimation performance, we incorporate the joint-sparsity model to exploit the spatial information in the data. The experiments demonstrate the effectiveness of the proposed algorithms on both synthetic and real data.
Qing Qu 0001, Nasser M. Nasrabadi, Trac D. Tran
ICASSP3
2013 Hierarchical sparse modeling using Spike and Slab priors
abstract
Sparse modeling has demonstrated its superior performances in many applications. Compared to optimization based approaches, Bayesian sparse modeling generally provides a more sparse result with a knowledge of confidence. Using the Spike and Slab priors, we propose the hierarchical sparse models for the scenario of single task and multitask - Hi-BCS and CHi-BCS. We draw the connections of these two methods to their optimization based counterparts and use expectation propagation for inference. The experiment results using synthetic and real data demonstrate that the performance of Hi-BCS and Chi-BCS are comparable or better than their optimization based counterparts.
Yuanming Suo, Minh Dao, Trac D. Tran, Umamahesh Srinivas, Vishal Monga
ICASSP3
2013 Temporal rate up-conversion of synthetic aperture radar via low-rank matrix recovery
abstract
The radar data to form synthetic aperture radar (SAR) imagery is normally transmitted and received by moving platforms like aircraft or vehicles. In many situations, the platforms move at high speed; which reduces the number of sampling records collected to the synthetic aperture, hence degrades the quality of the reconstructed SAR images. Therefore, it is necessary to develop an algorithm that is capable of increasing the temporal frequency rate of the received data. In this paper, we propose a novel technique to generate intermediate records from the existing ones by a locally-adaptive low-rank matrix recovery framework. The system first fills in the blank records using a bi-directional motion estimation scheme. The initialized aperture records are then refined by a robust low-rank matrix completion algorithm using the reference from neighborhood clean records. Experiments demonstrate that the proposed method provides comparative results when up-converting the aperture rate by a factor of two or four, both in mean square error of the raw SAR signal and PSNR performance of the recovered SAR images.
Minh Dao, Lam H. Nguyen, Trac D. Tran
ICIP3
2013 Structured sparse priors for image classification
abstract
Model-based compressive sensing (CS) exploits the structure inherent in sparse signals for the design of better signal recovery algorithms. This information about structure is often captured in the form of a prior on the sparse coefficients, the Laplacian being the most common such choice (leading to l1-norm minimization). The recent seminal contribution by Wright et al. exploits the discriminative capability of sparse representations for image classification, specifically face recognition. Their approach employs the analytical framework of CS with class-specific dictionaries. Our contribution is a logical extension of these ideas into structured sparsity for classification. We use class-specific dictionaries in conjunction with discriminative class-specific priors, specifically the spike-and-slab prior widely applied in Bayesian regression. Significantly, the proposed framework takes the burden off the demand for abundant training image samples necessary for the success of sparsity-based classification schemes.
Umamahesh Srinivas, Yuanming Suo, Minh Dao, Vishal Monga, Trac D. Tran
ICIP5
2013 Reconstruction of neural action potentials using signal dependent sparse representations
abstract
We demonstrate a method to build signal dependent sparse representation dictionary for neural action potentials using K-SVD algorithm and Discrete Wavelets Transform. We also show a method to utilize this dictionary to recover the neural signal in the Compressive Sensing (CS) framework. Comparing against the non-signal dependent CS recovery algorithms, this new recovery algorithm can achieve same reconstruction quality with 2.5 times less compressed sensing measurements. For the same compression ratio, the purposed approach can increase recovery signal's signal to noise and distortion ratio (SNDR) by around 6 dB compare to non-signal dependent recovery method. We also evaluated the recovered signal using spike sorting techniques. The results have shown that the spikes clusters still maintain clear separation even when the compression ratio is at 15-20% of the Nyquist rate. This work also implies that any hardware implementation of compressed sensing could be scaled down in term of power and chip area by the same order if this signal dependent framework is used to recover the signal.
Jie Zhang 0063, Yuanming Suo, Srinjoy Mitra, Sang (Peter) Chin, Trac D. Tran, Refet Firat Yazicioglu, Ralph Etienne-Cummings
ISCAS5
2013 Exploiting Sparsity in Hyperspectral Image Classification via Graphical Models
abstract
A significant recent advance in hyperspectral image (HSI) classification relies on the observation that the spectral signature of a pixel can be represented by a sparse linear combination of training spectra from an overcomplete dictionary. A spatiospectral notion of sparsity is further captured by developing a joint sparsity model, wherein spectral signatures of pixels in a local spatial neighborhood (of the pixel of interest) are constrained to be represented by a common collection of training spectra, albeit with different weights. A challenging open problem is to effectively capture the class conditional correlations between these multiple sparse representations corresponding to different pixels in the spatial neighborhood. We propose a probabilistic graphical model framework to explicitly mine the conditional dependences between these distinct sparse features. Our graphical models are synthesized using simple tree structures which can be discriminatively learnt (even with limited training samples) for classification. Experiments on benchmark HSI data sets reveal significant improvements over existing approaches in classification rates as well as robustness to choice of training.
Umamahesh Srinivas, Yi Chen 0014, Vishal Monga, Nasser M. Nasrabadi, Trac D. Tran
IEEE Geosci. Remote. Sens. Lett.5
2013 Hyperspectral Image Classification via Kernel Sparse Representation
abstract
In this paper, a novel nonlinear technique for hyperspectral image (HSI) classification is proposed. Our approach relies on sparsely representing a test sample in terms of all of the training samples in a feature space induced by a kernel function. For each test pixel in the feature space, a sparse representation vector is obtained by decomposing the test pixel over a training dictionary, also in the same feature space, by using a kernel-based greedy pursuit algorithm. The recovered sparse representation vector is then used directly to determine the class label of the test pixel. Projecting the samples into a high-dimensional feature space and kernelizing the sparse representation improve the data separability between different classes, providing a higher classification accuracy compared to the more conventional linear sparsity-based classification algorithms. Moreover, the spatial coherency across neighboring pixels is also incorporated through a kernelized joint sparsity model, where all of the pixels within a small neighborhood are jointly represented in the feature space by selecting a few common training samples. Kernel greedy optimization algorithms are suggested in this paper to solve the kernel versions of the single-pixel and multi-pixel joint sparsity-based recovery problems. Experimental results on several HSIs show that the proposed technique outperforms the linear sparsity-based classification technique, as well as the classical support vector machines and sparse kernel logistic regression classifiers.
Yi Chen 0014, Nasser M. Nasrabadi, Trac D. Tran
IEEE Trans. Geosci. Remote. Sens.3
2013 Exact Recoverability From Dense Corrupted Observations via ℓ1-Minimization
abstract
This paper confirms a surprising phenomenon first observed by Wright under a different setting: givenmhighly corrupted measurementsy=AΩ·x*+e*, whereAΩ·is a submatrix whose rows are selected uniformly at random from rows of an orthogonal matrixAande*is an unknown sparse error vector whose nonzero entries may be unbounded, we show that with high probability, ℓ1-minimization can recover the sparse signal of interestx*exactly from onlym=Cμ2k(logn)2, wherekis the number of nonzero components ofx*and μ =nmaxij Aij2, even if a significant fraction of the measurements are corrupted. We further guarantee that stable recovery is possible when measurements are polluted by both gross sparse and small dense errors:y=AΩ·x*+e*+ ν, where ν is the small dense noise with bounded energy. Numerous simulation results under various settings are also presented to verify the validity of the theory as well as to illustrate the promising potential of the proposed framework.
Nam H. Nguyen, Trac D. Tran
IEEE Trans. Inf. Theory2
2013 Robust Lasso With Missing and Grossly Corrupted Observations
abstract
This paper studies the problem of accurately recovering ak-sparse vector β*∈ \BBRpfrom highly corrupted linear measurementsy=Xβ*+e*+w, wheree*∈ \BBRnis a sparse error vector whose nonzero entries may be unbounded andwis a stochastic noise term. We propose a so-called extended Lasso optimization which takes into consideration sparse prior information of both β*ande*. Our first result shows that the extended Lasso can faithfully recover both the regression as well as the corruption vector. Our analysis relies on the notion of extended restricted eigenvalue for the design matrixX. Our second set of results applies to a general class of Gaussian design matrixXwith i.i.d. rowsN(0,Σ), for which we can establish a surprising result: the extended Lasso can recover exact signed supports of both β*ande*from only Ω(klogplogn) observations, even when a linear fraction of observations is grossly corrupted. Our analysis also shows that this amount of observations required to achieve exact signed support is indeed optimal.
Nam H. Nguyen, Trac D. Tran
IEEE Trans. Inf. Theory2
2012 Kernel sparse representation for hyperspectral target detection
abstract
In this paper, we present a nonlinear kernel-based target detection algorithm for hyperspectral images. The proposed approach relies on the sparse representation of an unknown sample with respect to both background and target training samples in a high-dimensional feature space induced by a kernel function. The sparse representation vector can be recovered via a kernelized greedy algorithm, where the kernel trick is used to avoid explicit evaluations of the data in the feature space. The spatial smoothness in hyperspectral images is also taken into account through a kernelized joint sparsity model. The detection decision is then made by comparing the reconstruction accuracy in terms of the background and target sub-dictionaries. The detection algorithm in a high-dimensional feature space implicitly exploits the higher-order structure (correlations) within the data which cannot be captured by a linear model. Therefore, projecting the pixels into a kernel feature space and kernelizing the linear sparse representation model improves the separability between the background and target classes, leading to a more accurate detection performance.
Yi Chen 0014, Nasser M. Nasrabadi, Trac D. Tran
IGARSS3
2012 Robust and adaptive extraction of RFI signals from ultra-wideband radar data
abstract
In this paper, we propose a novel, robust, and adaptive technique for the extraction of radio frequency interference (RFI) signals from ultra-wideband (UWB) radar data via sparse recovery. Unlike notch-filtering techniques that have been widely employed in the past, our proposed technique directly estimates and suppresses RFI signals from the UWB radar signal directly in time domain. Therefore, it does not suffer from several detrimental side effects such as high-sidelobe distortion and target-amplitude reduction as often observed in notch-filtering approaches. In addition, the technique is completely adaptive with highly time-varying environments and does not assume any knowledge (from frequency band to modulation scheme) of the RFI sources. The proposed technique is based on a sparse-recovery approach that simultaneously solves for (i) the UWB radar signal embedded in RFI noise with large amplitudes and (ii) RFI signals. Using both simulated and real-world data measured by the U.S. Army Research Laboratory (ARL) UWB synthetic aperture radar (SAR), we show that our proposed RFI extraction technique successfully recovers the UWB radar signal embedded in large-amplitude RFI signals. An average of 12 dB of RFI suppression is consistently realized in the real radar data experiments.
Lam H. Nguyen, Trac D. Tran
IGARSS2
2012 Discriminative graphical models for sparsity-based hyperspectral target detection
abstract
The inherent discriminative capability of sparse representations has been exploited recently for hyperspectral target detection. This approach relies on the observation that the spectral signature of a pixel can be represented as a linear combination of a few training spectra drawn from both target and background classes. The sparse representation corresponding to a given test spectrum captures class-specific discriminative information crucial for detection tasks. Spatio-spectral information has also been introduced into this framework via a joint sparsity model that simultaneously solves for the sparse features for a group of spatially local pixels, since such pixels are highly likely to have similar spectral characteristics. In this paper, we propose a probabilistic graphical model framework that can explicitly learn the class conditional correlations between these distinct sparse representations corresponding to different pixels in a spatial neighborhood. Simulation results show that the proposed algorithm outperforms classical hyperspectral target detection algorithms as well as support vector machines.
Umamahesh Srinivas, Yi Chen 0014, Vishal Monga, Nasser M. Nasrabadi, Trac D. Tran
IGARSS5
2011 Robust multi-sensor classification via joint sparse representation
Nam H. Nguyen, Nasser M. Nasrabadi, Trac D. Tran
FUSION3
2011 Video error concealment using sparse recovery and local dictionaries
abstract
Video error concealment is a post-processing technique that conceals the errors in a decoded video sequence based on data available only at the decoder. Most of the current techniques adopt the approach that recovers the Motion Vector (MV) of a lost image block, uses that MV to look for data to fill in the blank then performs some refinements. We propose a method that does not rely on MV recovery, but essentially bases on sparse representation of image patches on local temporal dictionaries. Experiment results show a large improvement over Boundary Matching Algorithm (BMA), the standard method used in reference software for H.264 video codec.
Dzung T. Nguyen, Minh Dao, Trac D. Tran
ICASSP3
2011 Hyperspectral image classification via kernel sparse representation
abstract
In this paper, a new technique for hyperspectral image classification is proposed. Our approach relies on the sparse representation of a test sample with respect to all training samples in a feature space induced by a kernel function. Projecting the samples into the feature space and kernelizing the sparse representation improves the separability of the data and thus yields higher classification accuracy compared to the more conventional linear sparsity-based classification algorithm. Moreover, the spatial coherence across neighboring pixels is also incorporated through a kernelized joint sparsity model, where all of the pixels within a small neighborhood are sparsely represented in the feature space by selecting a few common training samples. Two greedy algorithms are also provided in this paper to solve the kernel versions of the pixel-wise and jointly sparse recovery problems. Experimental results show that the proposed technique outperforms the linear sparsity-based classification technique and the classical Support Vector Machine classifiers.
Yi Chen 0014, Nasser M. Nasrabadi, Trac D. Tran
ICIP3
2011 Error concealment via 3-mode tensor approximation
abstract
This paper presents a novel video error concealment method, which is essentially a combination of non-local grouping of image patches and low-rank tensor approximation. The proposed method, though does not require the knowledge of Motion Vectors (MVs) as in traditional video concealment techniques such as Boundary Matching Algorithm (BMA) and its derivations, gives striking results in restoration of sequences, especially in the challenging cases when the error rate is high and/or key frames are not available. Our proposed framework can also be customized to deal with single images as well (for example, in image inpainting tasks).
Dzung T. Nguyen, Minh Dao, Trac D. Tran
ICIP3
2011 Robust recovery of synthetic aperture radar data from uniformly under-sampled measurements
abstract
In this paper, we propose a novel robust sparse-recovery technique that allows sub-Nyquist uniform under-sampling of wide-bandwidth radar data in real time (single observation). Although much of the information is lost in the received signal due to the low sampling rate, we hypothesize that each wide- bandwidth radar data record can be modeled as a superposition of many backscattered signals from reflective point targets in the scene. In other words, our proposed technique is based on direct sparse recovery via orthogonal matching pursuit using a special dictionary containing many time-delayed versions of the transmitted probing signal. Using data from the U.S. Army Research Laboratory (ARL) Ultra-Wideband (UWB) synthetic aperture radar (SAR), we show that the proposed sparse-recovery model- based (SMB) technique successfully models and synthesizes the returned radar data from real-world scenes using only an analytical waveform that models the transmitted signal and a handful of reflectivity coefficients. More importantly, the reconstructed SAR imagery using the SBM technique with data sampled at only 20% of the original sampling rate has a comparable signal-to-noise ratio (SNR) to the original SAR imagery. For comparison purpose, the paper also presents SAR images recovered from conventional interpolation techniques and the standard random projection based compressed sensing technique, both of which resulted in very poor SAR image quality at the same sub-Nyquist sampling rate (20%).
Lam H. Nguyen, Trac D. Tran
IGARSS2
2011 Robust Lasso with missing and grossly corrupted observations
abstract
This paper studies the problem of accurately recovering a sparse vector $\beta^{\star}$ from highly corrupted linear measurements $y = X \beta^{\star} + e^{\star} + w$ where $e^{\star}$ is a sparse error vector whose nonzero entries may be unbounded and $w$ is a bounded noise. We propose a so-called extended Lasso optimization which takes into consideration sparse prior information of both $\beta^{\star}$ and $e^{\star}$. Our first result shows that the extended Lasso can faithfully recover both the regression and the corruption vectors. Our analysis is relied on a notion of extended restricted eigenvalue for the design matrix $X$. Our second set of results applies to a general class of Gaussian design matrix $X$ with i.i.d rows $\oper N(0, \Sigma)$, for which we provide a surprising phenomenon: the extended Lasso can recover exact signed supports of both $\beta^{\star}$ and $e^{\star}$ from only $\Omega(k \log p \log n)$ observations, even the fraction of corruption is arbitrarily close to one. Our analysis also shows that this amount of observations required to achieve exact signed support is optimal.
Nam H. Nguyen, Nasser M. Nasrabadi, Trac D. Tran
NIPS3
2011 Simultaneous Joint Sparsity Model for Target Detection in Hyperspectral Imagery
abstract
This letter proposes a simultaneous joint sparsity model for target detection in hyperspectral imagery (HSI). The key innovative idea here is that hyperspectral pixels within a small neighborhood in the test image can be simultaneously represented by a linear combination of a few common training samples but weighted with a different set of coefficients for each pixel. The joint sparsity model automatically incorporates the interpixel correlation within the HSI by assuming that neighboring pixels usually consist of similar materials. The sparse representations of the neighboring pixels are obtained by simultaneously decomposing the pixels over a given dictionary consisting of training samples of both the target and background classes. The recovered sparse coefficient vectors are then directly used for determining the label of the test pixels. Simulation results show that the proposed algorithm outperforms the classical hyperspectral target detection algorithms, such as the popular spectral matched filters, matched subspace detectors, and adaptive subspace detectors, as well as binary classifiers such as support vector machines.
Yi Chen 0014, Nasser M. Nasrabadi, Trac D. Tran
IEEE Geosci. Remote. Sens. Lett.3
2011 Stepwise Optimal Subspace Pursuit for Improving Sparse Recovery
abstract
We propose a new iterative algorithm to reconstruct an unknown sparse signal x from a set of projected measurements y = Φx . Unlike existing methods, which rely crucially on the near orthogonality of the sampling matrix Φ , our approach makes stepwise optimal updates even when the columns of Φ are not orthogonal. We invoke a block-wise matrix inversion formula to obtain a closed-form expression for the increase (reduction) in the L2-norm of the residue obtained by removing (adding) a single element from (to) the presumed support of x . We then use this expression to design a computationally tractable algorithm to search for the nonzero components of x . We show that compared to currently popular sparsity seeking matching pursuit algorithms, each step of the proposed algorithm is locally optimal with respect to the actual objective function. We demonstrate experimentally that the algorithm significantly outperforms conventional techniques in recovering sparse signals whose nonzero values have exponentially decaying magnitudes or are distributed N(0,1) .
Balakrishnan Varadarajan, Sanjeev Khudanpur, Trac D. Tran
IEEE Signal Process. Lett.3
2011 Hyperspectral Image Classification Using Dictionary-Based Sparse Representation
abstract
A new sparsity-based algorithm for the classification of hyperspectral imagery is proposed in this paper. The proposed algorithm relies on the observation that a hyperspectral pixel can be sparsely represented by a linear combination of a few training samples from a structured dictionary. The sparse representation of an unknown pixel is expressed as a sparse vector whose nonzero entries correspond to the weights of the selected training samples. The sparse vector is recovered by solving a sparsity-constrained optimization problem, and it can directly determine the class label of the test sample. Two different approaches are proposed to incorporate the contextual information into the sparse recovery optimization problem in order to improve the classification performance. In the first approach, an explicit smoothing constraint is imposed on the problem formulation by forcing the vector Laplacian of the reconstructed image to become zero. In this approach, the reconstructed pixel of interest has similar spectral characteristics to its four nearest neighbors. The second approach is via a joint sparsity model where hyperspectral pixels in a small neighborhood around the test pixel are simultaneously represented by linear combinations of a few common training samples, which are weighted with a different set of coefficients for each pixel. The proposed sparsity-based algorithm is applied to several real hyperspectral images for classification. Experimental results show that our algorithm outperforms the classical supervised classifier support vector machines in most cases.
Yi Chen 0014, Nasser M. Nasrabadi, Trac D. Tran
IEEE Trans. Geosci. Remote. Sens.3
2010 A sparsity-driven joint image registration and change detection technique for SAR imagery
abstract
This paper presents a novel Sparsity-driven joint Image REgistration and Change Detection (SIRE-CD) technique for SAR imagery. The proposed algorithm simultaneously performs two main tasks: (i) locally register the test and reference images; and (ii) perform the change detection between the two. The key innovative concept here is the sparsity-driven transformation of the signatures from the reference image to match to those of the test image at the local image patch level. In other words, we are constructing a large dictionary from the reference data and use that to find the sparsest representation that best approximates the new incoming test data. The accuracy level of the approximation determines the detected changes between the reference and the test image. We demonstrate the performance of this technique using both simulated data and real SAR imagery from the Army Research Laboratory ultra-wideband (UWB) SAR forward-looking radar.
Lam H. Nguyen, Trac D. Tran
ICASSP2
2010 Sparse coding for speech recognition
abstract
This paper proposes a novel feature extraction technique for speech recognition based on the principles of sparse coding. The idea is to express a spectro-temporal pattern of speech as a linear combination of an overcomplete set of basis functions such that the weights of the linear combination are sparse. These weights (features) are subsequently used for acoustic modeling. We learn a set of overcomplete basis functions (dictionary) from the training set by adopting a previously proposed algorithm which iteratively minimizes the reconstruction error and maximizes the sparsity of weights. Furthermore, features are derived using the learned basis functions by applying the well established principles of compressive sensing. Phoneme recognition experiments show that the proposed features outperform the conventional features in both clean and noisy conditions.
Garimella S. V. S. Sivaram, Sridhar Krishna Nemala, Mounya Elhilali, Trac D. Tran, Hynek Hermansky
ICASSP4
2010 Robust face recognition using locally adaptive sparse representation
abstract
This paper presents a block-based face-recognition algorithm based on a sparse linear-regression subspace model via locally adaptive dictionary constructed from past observable data (training samples). The local features of the algorithm provide an immediate benefit - the increase in robustness level to various registration errors. Our proposed approach is inspired by the way human beings often compare faces when presented with a tough decision: we analyze a series of local discriminative features (do the eyes match? how about the nose? what about the chin?...) and then make the final classification decision based on the fusion of local recognition results. In other words, our algorithm attempts to represent a block in an incoming test image as a linear combination of only a few atoms in a dictionary consisting of neighboring blocks in the same region across all training samples. The results of a series of these sparse local representations are used directly for recognition via either maximum likelihood fusion or a simple democratic majority voting scheme. Simulation results on standard face databases demonstrate the effectiveness of the proposed algorithm in the presence of multiple mis-registration errors such as translation, rotation, and scaling.
Yi Chen 0014, Thong T. Do, Trac D. Tran
ICIP3
2010 Fast dimension reduction through random permutation
abstract
This paper studies permutation-based dimension reduction, which can be implemented by first scrambling the input data, then applying the FFT, DCT or Walsh-Hadamard transform and finally using either uniformly random sampling or sparse random projection. By exploiting concentration inequalities of random permutation, we show that this subclass of operators can offer (near) optimal theoretical guarantee. Besides, as random permutation of N elements can be implemented in O(N) time, the proposed algorithm has very low complexity. Some numerical examples are presented to demonstrate the validity of our theoretical development and their promising applications in image processing.
Lu Gan 0002, Thong T. Do, Trac D. Tran
ICIP3
2010 Sparsity-based classification of hyperspectral imagery
abstract
In this paper, a new sparsity-based classification algorithm for hyperspectral imagery is proposed. This algorithm is based on the concept that a pixel in hyperspectral imagery lies in a low-dimensional subspace and thus can be represented by a sparse linear combination of the training samples. The sparse representation (a sparse vector representing the selected training samples) of a test sample can be recovered by solving a constrained optimization problem. Once the sparse vector is obtained, the class of the test sample can be directly determined by the behavior of the vector on reconstruction. In addition to the constraints on sparsity and reconstruction accuracy, we also exploit the fact that hyperspectral images are usually smooth within a neighborhood. In our proposed algorithm, a smoothness constraint is imposed by forcing the Laplacian of the reconstructed image to be minimum in the optimization process. The proposed sparsity-based algorithm is applied to several hyperspectral imagery to classify the pixels into target and background classes. Simulation results show that our algorithm outperforms the classical hyperspectral target detection algorithms, such as the popular spectral matched filters, matched subspace detectors, and adaptive subspace detectors.
Yi Chen 0014, Nasser M. Nasrabadi, Trac D. Tran
IGARSS3
2009 A fast and efficient heuristic nuclear-norm algorithm for affine rank minimization
abstract
The problem of affine rank minimization seeks to find the minimum rank matrix that satisfies a set of linear equality constraints. Generally, since affine rank minimization is NP-hard, a popular heuristic method is to minimize the nuclear norm that is a sum of singular values of the matrix variable. A recent intriguing paper shows that if the linear transform that defines the set of equality constraints is nearly isometrically distributed and the number of constraints is at least O(r(m + n) logmn), where r and m times n are the rank and size of the minimum rank matrix, minimizing the nuclear norm yields exactly the minimum rank matrix solution. Unfortunately, it takes a large amount of computational complexity and memory buffering to solve the nuclear norm minimization problem with known nearly isometric transforms. This paper presents a fast and efficient algorithm for nuclear norm minimization that employs structurally random matrices for its linear transform and a projected subgradient method that exploits the unique features of structurally random matrices to substantially speed up the optimization process. Theoretically, we show that nuclear norm minimization using structurally random linear constraints guarantees the minimum rank matrix solution if the number of linear constraints is at least O(r(m+n) log3mn). Extensive simulations verify that structurally random transforms still retain optimal performance while their implementation complexity is just a fraction of that of completely random transforms, making them promising candidates for large scale applications.
Thong T. Do, Yi Chen 0014, Nam H. Nguyen, Lu Gan 0002, Trac D. Tran
ICASSP5
2009 Fast and efficient dimensionality reduction using Structurally Random Matrices
abstract
Structurally Random Matrices (SRM) are first proposed in [1] as fast and highly efficient measurement operators for large scale compressed sensing applications. Motivated by the bridge between compressed sensing and the Johnson-Lindenstrauss lemma [2] , this paper introduces a related application of SRMs regarding to realizing a fast and highly efficient embedding. In particular, it shows that a SRM is also a promising dimensionality reduction transform that preserves all pairwise distances of high dimensional vectors within an arbitrarily small factor ∈, provided that the projection dimension is on the order of O(∈−2log3N), where N denotes the number of d-dimensional vectors. In other words, SRM can be viewed as the sub-optimal Johnson-Lindenstrauss embedding that, however, owns very low computational complexity O(d log d) and highly efficient implementation that uses only O(d) random bits, making it a promising candidate for practical, large scale applications where efficiency and speed of computation are highly critical.
Thong T. Do, Lu Gan 0002, Yi Chen 0014, Nam P. Nguyen, Trac D. Tran
ICASSP5
2009 A fast and efficient algorithm for low rank matrix recovery from incomplete observations
abstract
Minimizing the rank of a matrix X over certain constraints arises in diverse areas such as machine learning, control system and is known to be computationally NP-hard. In this paper, a new simple and efficient algorithm for solving this rank minimization problem with linear constraints is proposed. By using gradient projection method to optimize S while consecutively updating matrices U and V (where X = USVT) in combination with the use of an approximation function for l0-norm of singular values, our algorithm is shown to run significantly faster with much lower computational complexity than general-purpose interior-point solvers, for instance, the SeDuMi package. In addition, the proposed algorithm can recover the matrix exactly with much fewer measurements and is also appropriate for large-scale applications.
Nam P. Nguyen, Thong T. Do, Yi Chen 0014, Trac D. Tran
ICASSP4
2009 Rational canonical form of polyphase matrices with applications to designing paraunitary filter banks
abstract
In this paper we consider the rational canonical form of arbitrary polyphase matrices and use it to derive a simple implementation of paraunitary filter banks (PUFBs) based on a cascade of elementary building blocks. Furthermore, this decomposition is shown to be easily extendable to include a large class of perfect reconstruction filter banks (PRFBs) and can be especially useful for deriving the initial condition of PUFB design algorithms.
Peter G. Vouras, Trac D. Tran, Michael Ching
ICASSP2
2009 Distributed compressed video sensing
abstract
This paper proposes a novel framework called Distributed Compressed Video Sensing (DISCOS) - a solution for Distributed Video Coding (DVC) based on the recently emerging Compressed Sensing theory. The DISCOS framework compressively samples each video frame independently at the encoder. However, it recovers video frames jointly at the decoder by exploiting an interframe sparsity model and by performing sparse recovery with side information. In particular, along with global frame-based measurements, the DISCOS encoder also acquires local block-based measurements for block prediction at the decoder. Our interframe sparsity model mimics state-of-the-art video codecs: the sparsest representation of a block is a linear combination of a few temporal neighboring blocks that are in previously reconstructed frames or in nearby key frames. This model enables a block to be optimally predicted from its local measurements by l1-minimization. The DISCOS decoder also employs a sparse recovery with side information to jointly reconstruct a frame from its global measurements and its local block-based prediction. Simulation results show that the proposed framework outperforms the baseline compressed sensing-based scheme of intraframe-coding and intraframe-decoding by 8 – 10dB. Finally, unlike conventional DVC schemes, our DISCOS framework can perform most encoding operations in the analog domain with very low-complexity, making it be a promising candidate for real-time, practical applications where the analog to digital conversion is expensive, e.g., in Terahertz imaging.
Thong T. Do, Yi Chen 0014, Dzung T. Nguyen, Nam P. Nguyen, Lu Gan 0002, Trac D. Tran
ICIP6
2009 Implementation and application of local computation of wavelet coefficients in the dual-tree complex wavelets
abstract
The dual-tree complex wavelet transform (DT CWT) was introduced to overcome the disadvantages of the traditional fully decimated discrete wavelet transform (DWT), namely the shift-variance and the poor directional selectivity properties. Because of its improvements in these aspects, the dual-tree has been widely used in many image processing applications such as denoising, motion estimation, image classification and even compression despite its redundant representation. In our previous work, we were able to accurately and locally estimate the wavelet coefficients of one tree in the DT CWT, given a subset of the other tree coefficients. Our method is based on exploiting the orthogonality properties of one of the nicest dual-tree designs - the Q-shift complex wavelets. In this paper, we demonstrate the implementation of multiple level of decomposition as well as the two dimensional realization with application to region of interest (ROI) imaging applications such as denoising.
Iman A. El-Shehaby, Trac D. Tran
ICIP2
2009 Robust video transmission using Layered Compressed Sensing
abstract
We propose a novel Layered Compressed Sensing (CS) approach for robust transmission of video signals over packet loss channels. In our proposed method, the encoder consists of a base layer and an enhancement layer. The base layer is a conventionally encoded bitstream and transmitted without any error protection. The additional enhancement layer is a stream of compressed measurements taken across slices of video signals for error-resilience. The decoder regards the corrupted base layer as the side information (SI) and employs a sparse recovery with SI to recover approximation of lost packets. By exploiting the SI at the decoder, the enhancement layer is required to transmit a minimal amount of compressed measurements for error protection that is only proportional to the amount of lost packets. Simulation results show that both compression efficiency and error-resilience capacity of the proposed scheme are competitive with those of other state-of-the-art robust transmission methods, in which Wyner-Ziv (WZ) coders often generate an enhancement layer. Thanks to the soft-decoding feature of sparse recovery algorithms, our CS-based scheme can avoid the cliff effect that often occurs with otherWyner-Ziv based schemes when the error rate is over the error correction capacity of the channel code. In addition, our result suggests that compressed sensing is actually closer to source coding with decoder side information than to conventional source coding.
Thong T. Do, Yi Chen 0014, Dzung T. Nguyen, Nam P. Nguyen, Lu Gan 0002, Trac D. Tran
MMSP6
2009 A fast and efficient algorithm for low-rank approximation of a matrix
abstract
The low-rank matrix approximation problem involves finding of a rank k version of a m x n matrix A, labeled Ak, such that Ak is as "close" as possible to the best SVD approximation version of A at the same rank level. Previous approaches approximate matrix A by non-uniformly adaptive sampling some columns (or rows) of A, hoping that this subset of columns contain enough information about A. The sub-matrix is then used for the approximation process. However, these approaches are often computationally intensive due to the complexity in the adaptive sampling. In this paper, we propose a fast and efficient algorithm which at first pre-processes matrix A in order to spread out information (energy) of every columns (or rows) of A, then randomly selects some of its columns (or rows). Finally, a rank-k approximation is generated from the row space of these selected sets. The preprocessing step is performed by uniformly randomizing signs of entries of A and transforming all columns of A by an orthonormal matrix F with existing fast implementation (e.g. Hadamard, FFT, DCT...). Our main contribution is summarized as follows. 1) We show that by uniformly selecting at random d rows of the preprocessed matrix with d = ( 1/η k max {log k, log 1/β} ), we guarantee the relative Frobenius norm error approximation: (1 + η) norm{A - Ak}F with probability at least 1 - 5β. 2) With d above, we establish a spectral norm error approximation: (2 + √2m/d) norm{A - Ak}2 with probability at least 1 - 2β. 3) The algorithm requires 2 passes over the data and runs in time (mn log d + (m+n) d2) which, as far as the best of our knowledge, is the fastest algorithm when the matrix A is dense. 4) As a bonus, applying this framework to the well-known least square approximation problem min norm{A x - b} where A ∈ Rm x r, we show that by randomly choosing d = (1/η γ r log m), the approximation solution is proportional to the optimal one with a factor of η and with extremely high probability, (1 - 6 m-γ), say.
Nam H. Nguyen, Thong T. Do, Trac D. Tran
STOC3
2009 Multiple Description Coding With Prediction Compensation
abstract
A new multiple description coding paradigm is proposed by combining the time-domain lapped transform, block level source splitting, linear prediction, and prediction residual encoding. The method provides effective redundancy control and fully utilizes the source correlation. The joint optimization of all system components and the asymptotic performance analysis are presented. Image coding results demonstrate the superior performance of the proposed method, especially at low redundancies.
Guoqian Sun, Upul Samarawickrama, Jie Liang 0001, Chao Tian 0002, Chengjie Tu, Trac D. Tran
IEEE Trans. Image Process.6
2008 Fast compressive sampling with structurally random matrices
abstract
This paper presents a novel framework of fast and efficient compressive sampling based on the new concept of structurally random matrices. The proposed framework provides four important features. (i) It is universal with a variety of sparse signals. (ii) The number of measurements required for exact reconstruction is nearly optimal. (iii) It has very low complexity and fast computation based on block processing and linear filtering. (iv) It is developed on the provable mathematical model from which we are able to quantify trade-offs among streaming capability, computation/memory requirement and quality of reconstruction. All currently existing methods only have at most three out of these four highly desired features. Simulation results with several interesting structurally random matrices under various practical settings are also presented to verify the validity of the theory as well as to illustrate the promising potential of the proposed framework.
Thong T. Do, Trac D. Tran, Lu Gan 0002
ICASSP2
2008 Lifting-based Laplacian Pyramid reconstruction schemes
abstract
Laplacian Pyramid (LP) provides a redundant signal representation and can be characterized as an oversampled filter bank (FB). In this paper, a generic lifting-based parameterization reconstruction algorithm is proposed to characterize all LP synthesis banks that can satisfy the perfect reconstruction property. Two typical lifting-based LP reconstruction schemes are then derived from this general representation. The first scheme presents the dual frame LP reconstruction and its closed-form solutions for any LP filters. The second LP reconstruction scheme leads to an efficient FB, which demonstrates improvements over the usual LP reconstruction in the presence of noise.
Lijie Liu, Lu Gan 0002, Trac D. Tran
ICIP3
2008 Local computation and estimation of wavelet coefficients in the dual-tree complex wavelet transform
abstract
The dual-tree complex wavelet transform (DT CWT) was introduced to overcome the disadvantages of the traditional fully decimated discrete wavelet transform (DWT), namely the shift-variance and the poor directional selectivity properties. Because of its improvements in this regards, the dual-tree has been successfully demonstrated in many image processing applications such as denoising, motion estimation, image classification, and even compression despite its redundancy nature. Our goal is to be able to predict the wavelet coefficients of one tree knowing those of the other in the dual-tree complex wavelet transform. In other words, given a subset of real coefficients in the DTCWT, how can we compute or estimate accurately the imaginary coefficients in the same local neighborhood and vice versa? The proposed method is based on exploiting the orthogonality properties of one of the nicest dual-tree designs - the Q-shift complex wavelets.
Iman A. El-Shehaby, Trac D. Tran
ISCAS2
2007 Optical Flow Approximation of Sub-Pixel Accurate Block Matching for Video Coding
abstract
Video compression algorithms almost universally rely on block matching algorithms (BMA) to exploit the temporal redundancy in image sequences. However block matching is extremely computationally intensive, especially if sub-pixel accuracy is desired. We propose a fundamentally different approach to obtaining motion vectors using the principles from gradient based optical flow but in a form fully compatible with any macro-block based video compression scheme. Experimental results show that in many cases, the gradient approximation to block matching results in videos within 0.3 dB PSNR to BMA at sub-pixel resolution while being potentially much more computationally efficient and easier to implement at the hardware level.
Yu M. Chi, Trac D. Tran, Ralph Etienne-Cummings
ICASSP (1)2
2007 An 8×8 IEEE-Compliant Lifting-Based Multiplierless IDCT Structure and Algorithm
abstract
In this paper we propose a lifting-based 8times8 IDCT structure and its EEEE-1180 compliant approximation solution. Derived from an efficient Loeffler's 11-multiply IDCT structure, the proposed scheme comprises of butterflies and dyadic-rational lifting steps that can be implemented using only shift and add operations. Our approach also allows the computational scalability with different accuracy-versus-complexity trade-offs. Furthermore, the lifting construction allows a simple construction of the corresponding multiplierless forward DCT, providing bit-exact reconstruction if pairing with our proposed IDCT Our high-accuracy solution provides a very close approximation of the floating-point IDCT. The experiments in MPEG-2 and MPEG-4 video coders under the worst-case assumptions show almost drifting-free reconstructions.
Lijie Liu, Trac D. Tran
ICASSP (1)2
2007 Paraunitary Filter Bank Design using Derivative Constraints
abstract
In this paper two new algorithms are presented for designing finite impulse response (FIR) paraunitary (PU) filter banks. Each algorithm minimizes the mean square error between the desired response and the FIR PU approximation subject to constraints on either the derivative of the complex frequency response or the power response of individual channel filters. The derivative constraints are useful for shaping the response of a channel filter at particular frequencies of interest. An example illustrating the utility of derivative constraints is presented whereby a FIR PU approximation is derived for an ideal principal component filter bank (PCFB).
Peter G. Vouras, Trac D. Tran
ICASSP (3)2
2007 Multiple Description Image Codingwith Prediction Compensation
abstract
A new multiple description image coding paradigm is presented in this paper by combining the lapped transform, block level source splitting, inter-description prediction, and coding of the prediction residual. Jointly optimal designs of all system components are discussed. Compared with the best multiple description image coding algorithm in the literature, the new method can achieve significant improvement when one description is lost, given the same bit rate and the same central distortion.
Guoqian Sun, Upul Samarawickrama, Jie Liang 0001, Chengjie Tu, Trac D. Tran
ICIP (6)5
2007 A Complexity Scalable Universal DCT Domain Image Resizing Algorithm
abstract
This paper presents a computationally flexible method for producing a mapping from one discrete cosine transform (DCT) domain to another that results in a decoded image that has been arbitrarily resized in the spatial dimension. A notable feature of the proposed mapping is its computational scalability; final image quality can be traded off for lower implementation complexity and a wide range of complexity versus final quality operation points can be realized for each scale factor. Current existing methods often suffer from a lack of flexibility (i.e., work for only one or at most a few resizing factors, have only one or two levels of complexity) or require more operations to achieve similar levels of final image quality. When constructing the mapping, a multiplierless DCT approximation can also be employed, yielding fast implementation with excellent results. The use of the DCT approximation confers several benefits upon the proposed mapping including multiplierless implementation or at most integer operations rather than floating point operations
Carlos Salazar-Lazaro, Trac D. Tran
IEEE Trans. Circuits Syst. Video Technol.2
2007 Undersampled Boundary Pre-/Postfilters for Low Bit-Rate DCT-Based Block Coders
abstract
It has been well established that critically sampled boundary pre-/postfiltering operators can improve the coding efficiency and mitigate blocking artifacts in traditional discrete cosine transform-based block coders at low bit rates. In these systems, both the prefilter and the postfilter are square matrices. This paper proposes to use undersampled boundary pre- and postfiltering modules, where the pre-/postfilters are rectangular matrices. Specifically, the prefilter is a "fat" matrix, while the postfilter is a "tall" one. In this way, the size of the prefiltered image is smaller than that of the original input image, which leads to improved compression performance and reduced computational complexities at low bit rates. The design and VLSI-friendly implementation of the undersampled pre-/postfilters are derived. Their relations to lapped transforms and filter banks are also presented. Two design examples are also included to demonstrate the validity of the theory. Furthermore, image coding results indicate that the proposed undersampled pre-/postfiltering systems yield excellent and stable performance in low bit-rate image coding.
Lu Gan 0002, Chengjie Tu, Jie Liang 0001, Trac D. Tran, Kai-Kuang Ma
IEEE Trans. Image Process.4
2007 Wiener Filter-Based Error Resilient Time-Domain Lapped Transform
abstract
In this paper, the design of the error resilient time-domain lapped transform is formulated as a linear minimal mean-squared error problem. The optimal Wiener solution and several simplifications with different tradeoffs between complexity and performance are developed. We also prove the persymmetric structure of these Wiener filters. The existing mean reconstruction method is proven to be a special case of the proposed framework. Our method also includes as a special case the linear interpolation method used in DCT-based systems when there is no pre/postfiltering and when the quantization noise is ignored. The design criteria in our previous results are scrutinized and improved solutions are obtained. Various design examples and multiple description image coding experiments are reported to demonstrate the performance of the proposed method.
Jie Liang 0001, Chengjie Tu, Lu Gan 0002, Trac D. Tran, Kai-Kuang Ma
IEEE Trans. Image Process.4
2006 Two-Dimensional Wiener Filters for Error Resilient Time Domain Lapped Transform
abstract
This paper presents the design of two-dimensional Wiener filters for error resilient time domain lapped transform. Two solutions are discussed, and a multi-pass approach is also proposed to make the algorithm adaptive to input statistics. Design examples and image coding experiments show that the adaptive 2-D Wiener filters provide significant improvement over the existing 1-D Wiener filtering method.
Jie Liang 0001, Xin Li 0005, Guoqian Sun, Trac D. Tran
ICASSP (3)4
2006 Multiplierless Design of Biorthogonal Dual-Tree Complex Wavelet Transform using Lifting Scheme
abstract
In this work, we present the design, implementation and application of two families of biorthogonal dual-tree complex wavelet transform (CWT) filters using lifting scheme. The first design is achieved using exhaustive search with coding gain, DC leakage and directional selectivity as the fundamental criteria. The second set of filters are derived from the biorthogonal design procedure that was recently suggested by Selesnick. Furthermore, this paper also introduces a new theorem that suggests that lifting implementation of filters that are time-reversals of each other is closely related. Various applications are presented to validate the proposed design scheme, including performance in denoising as well as in the JPEG-2000 image coding standard.
Adeel Abbas, Trac D. Tran
ICIP2
2006 A More Efficient and Video Friendly Spatial Resizing Algorithm
abstract
Arbitrary spatial resizing of compressed video is a useful transcoding operation that can be used to match the video stream to a low bandwidth channel or to end user devices with a different resolution from that of the original source material. By performing this transcoding in the compressed domain the computational cost of the operation can be minimized. Our proposed algorithm resizes a video stream by modifying functional blocks in an existing transcoding architecture to permit spatial resizing. In order for the algorithm to function, a new motion compensation operation suitable for resized video was derived. The algorithm retains the advantages of permitting tradeoffs between computational cost and final quality while re sizing by any rational fraction. Additionally, the algorithm operates on an arbitrary supporting area and complements motion compensation.
Carlos Salazar-Lazaro, Trac D. Tran
ICIP2
2006 JPEG-compliant image coding with adaptive pre-/post-filtering
abstract
In this paper we propose an image coding scheme with adaptive pre-/post-filtering which produces a fully compliant JPEG bitstream. The basic idea is to introduce pre-filtering to improve the coding performance and post-filtering to reduce JPEG blocking artifacts. The adaptivity of the pre-/post-filters is achieved by varying their filter supports based on two criteria: rate-distortion optimization (RD-opt) and over-/under-flow. Experiments show that despite keeping intact JPEG baseline coding, our proposed coding scheme with these two criteria can improve not only the objective quality (0.3-1.5 dB PSNR gain), but also yield superior visual quality by preserving edge details and mitigating blocking artifacts. Our proposed algorithm is competitive with state-of-the-art deblocking algorithms.
Lijie Liu, Trac D. Tran
ISCAS3
2006 Error resilient pre/post-filtering for DCT-based block coding systems
abstract
Block coding based on the discrete cosine transform (DCT) is very popular in image and video compression. Pre/post-filtering can be attached to a DCT-based block coding system to improve coding efficiency as well as to mitigate blocking artifacts. Previously designed pre/post-filters are optimized to maximize coding efficiency solely. For image and video communication over unreliable channels, those pre/post-filters are sensitive to transmission errors. This paper addresses the problem of designing pre/post-filters which are more error resilient. Reconstruction performance is measured by how low the average reconstruction error is, and how uniformly the reconstruction error is distributed. A family of pre/post-filters is designed to provide desired tradeoffs between coding efficiency and robustness to transmission errors. Experiments show that these filtering operators can achieve superior reconstruction performance without sacrificing much coding performance.
Chengjie Tu, Trac D. Tran, Jie Liang 0001
IEEE Trans. Image Process.2
2005 Adaptive Block-Based Image Coding with Pre-/Post-Filtering
abstract
This paper presents an adaptive block-based image coding method, which combines the advantages of variable block size transform and adaptive pre-/post-filtering scheme. Our approach partitions an image into blocks with different sizes, which are best suitable for the characteristics of the underlying data in the rate-distortion (RD) sense. The adaptive block decomposition mitigates the ringing artifacts by adopting a small block size transform in nonstationary regions, and improves the coding efficiency by using a large block size transform in homogenous regions. Moreover, pre-/post-filtering is adaptively applied along the block boundaries to improve coding efficiency and minimize blocking artifacts. Simulation results show that the proposed coder can achieve competitive objective performance as well as yield superior reconstruction visual quality, compared with the RD-optimized JPEG2000 and H.264/AVC I-frame coder.
Lijie Liu, Trac D. Tran
DCC3
2005 Fast approximations of the orthogonal dual-tree wavelet bases
abstract
Recently, there has been a significant interest in the design of iterated filter banks in which the resulting wavelet bases form an approximate Hilbert transform pair. In this work, we propose three approximations of such dual-tree wavelet bases that satisfy Hilbert transform conditions. Our designs are derived from Selesnick's and Kingsbury's orthogonal wavelet filter solutions, and meet other desirable properties such as high coding gain, reduced computational complexity and sufficient regularity. The quantization is performed in the lattice domain using sum-of-power-of-two (SOPOT) coefficients. Several performance comparisons are presented. Furthermore, this paper introduces a proposition that lattice coefficients of filters that are time-reversals of each other are closely related.
Adeel Abbas, Trac D. Tran
ICASSP (4)2
2005 Wiener Filtering for Generalized Error Resilient Time Domain Lapped Transform
abstract
In this paper, we revisit the design of the time-domain lapped transform for error resilient image transmission. A general structure is first proposed whose solution is given by a Wiener filter. Two simplified schemes with different tradeoffs between complexity and performance are then developed, for which Wiener filter solutions also exist. We show that the existing method is a special case of the general scheme. Design examples and image coding experiments verify that the performance of our new approach is significantly better than existing techniques.
Jie Liang 0001, Chengjie Tu, Trac D. Tran, Lu Gan 0002
ICASSP (2)3
2005 Optimal block boundary pre/postfiltering for wavelet-based image and video compression
abstract
This paper presents a pre/postfiltering framework to reduce the reconstruction errors near block boundaries in wavelet-based image and video compression. Two algorithms are developed to obtain the optimal filter, based on boundary filter bank and polyphase structure, respectively. A low-complexity structure is employed to approximate the optimal solution. Performances of the proposed method in the removal of JPEG 2000 tiling artifact and the jittering artifact of three-dimensional wavelet video coding are reported. Comparisons with other methods demonstrate the advantages of our pre/postfiltering framework.
Jie Liang 0001, Chengjie Tu, Trac D. Tran
IEEE Trans. Image Process.3
2004 Optimal block boundary pre/post-filtering for wavelet-based image and video compression
abstract
This paper presents a pre/post-filtering method to reduce the reconstruction errors near block boundaries in wavelet-based image and video compression. It can be used effectively to mitigate the tiling artifact in JPEG2000 and the jittering artifact in 3D wavelet-based video compression. In this method, a short prefilter is applied across the boundaries of image tiles or video frame groups before wavelet compression, and a post-filter is applied at the same place after wavelet reconstruction. The optimal pre/postfilter is obtained by formulating and solving the corresponding rate-distortion optimization problem. A low-complexity structure is then proposed to approximate the optimal solution. The performance of the proposed method is demonstrated by both image and video coding examples.
Jie Liang 0001, Chengjie Tu, Trac D. Tran
ICIP3
2004 On resizing images in the dct domain
Carlos Salazar-Lazaro, Trac D. Tran
ICIP2
2004 Over-sampled and under-sampled Pre/post-filters for block DCT coders
abstract
Pre-post-filtering operators have been shown to improve the coding efficiency as well as to mitigate blocking artifacts in traditional DCT-based block coders. Pre-post-filters which preserve the system sampling rate have been extensively investigated. This paper explores the two noncritically-sampled signal decomposition cases - under-sampling and over-sampling - via the pre-post-processing perspective. We discuss various design issues, efficient structures, and present two application examples: under-sampled pre-post-filters for very low bit-rate image coding and over-sampled pre-post-filters for error resilient image transmission. Preliminary experimental results illustrate that noncritically-sampled pre-post-filtering does provide much improved coding performances than its critically-sampled counterpart for the aforementioned specific applications.
Chengjie Tu, Trac D. Tran, Jie Liang 0001
ICIP2
2003 On efficient implementation of oversampled linear phase perfect reconstruction filter banks
abstract
In this paper, we first present an alternative way of generating oversampled linear phase perfect reconstruction filter banks (OSLPPRFB). We show that this method provides the minimal factorization of a subset of existing OSLPPRFB. The combination of the new structure and the conventional one leads to efficient implementations of a general class of OSLPPRFB. Possible application of the new scheme is discussed.
Jie Liang 0001, Lu Gan 0002, Chengjie Tu, Trac D. Tran, Kai-Kuang Ma
ICASSP (6)4
2003 Error resilient pre-/post-filtering for DCT-based block coding systems
abstract
Pre-/post-filtering can be attached to a DCT-based block coding system to improve coding efficiency as well as to mitigate blocking artifacts. Previously designed pre-/post-filters are optimized to maximize coding efficiency solely. For image and video communication over unreliable channels, those pre-/post-filters are sensitive to transmission errors. This paper addresses the problem of designing pre-/post-filters which are more error resilient. A family of pre-/post-filters are designed to provide desired trade-offs between coding efficiency and robustness to transmission errors. These filters achieve superior reconstruction performance without sacrificing much coding performance.
Chengjie Tu, Trac D. Tran, Jie Liang 0001
ICASSP (3)2
2003 On efficient implementation of oversampled linear phase perfect reconstruction filter banks
abstract
In this paper, we first present an alternative way of generating over-sampled linear phase perfect reconstruction filter banks (OSLP-PRFB). We show that this method provides the minimal factorization of a subset of existing OSLPPRFB. The combination of the new structure and the conventional one leads to efficient implementations of a general class of OSLPPRFB. Possible application of the new scheme is discussed.
Jie Liang 0001, Lu Gan 0002, Chengjie Tu, Trac D. Tran, Kai-Kuang Ma
ICME4
2003 Error resilient pre-/post-filtering for DCT-based block coding systems
abstract
Pre-/post filtering can be attached to a DCT-based block coding system to improve the encoding efficiency as well as to mitigate blocking artifacts. Previously designed pre-/post-filters are optimized to maximize coding efficiency solely. For image and video communication over unreliable channels, those pre-/post-filters are sensitive to transmission errors. This paper addresses the problem of designing pre-/post-filters, which are more error resilient. A family of pre-/post-filters is designed to provide desired trade-offs between coding efficiency and robustness to transmission errors. These filters achieve superior reconstruction performance without sacrificing much coding performance.
Chengjie Tu, Trac D. Tran, Jie Liang 0001
ICME2
2003 Adaptive runlength coding
abstract
Runlength coding is the standard coding technique for block transform-based image/video compression. A block of quantized transform coefficients is first represented as a sequence of RUN/LEVEL pairs that are then entropy coded-RUN being the number of consecutive zeros and LEVEL being the value of the following nonzero coefficient. We point out the inefficiency of conventional runlength coding and introduce a novel adaptive runlength (ARL) coding scheme that encodes RUN and LEVEL separately using adaptive binary arithmetic coding and simple context modeling. We aim to maximize compression efficiency by adaptively exploiting the characteristics of block transform coefficients and the dependency between RUN and LEVEL. Coding results show that with the same level of complexity, the proposed ARL coding algorithm outperforms the conventional runlength coding scheme by a large margin in the rate-distortion sense.
Chengjie Tu, Jie Liang 0001, Trac D. Tran
IEEE Signal Process. Lett.3
2002 Multiplierless approximation of transforms using lifting scheme and coordinate descent with adder constraint
abstract
This paper describes an algorithm for systematically finding a multiplierless approximation of transforms where VLSI-friendly binary coefficients of the form k/2nare employed in the approximation. Assuming the cost of binary shifters is negligible in hardware, the total number of binary adders required to approximate the transform is used as the complexity constraint. The proposed algorithm is systematic and fast. It eliminates the need for trial-and-error binary approximations of the coefficients. Specifically, two types of multiplierless approximations of the discrete cosine transform (DCT) are presented to illustrate the algorithm.
Ying-Jui Chen, Soontorn Oraintara, Trac D. Tran, Kevin Amaratunga, Truong Q. Nguyen
ICASSP3
2002 Adaptive pre- and post-filtering for block based systems
abstract
This paper introduces an adaptive time-varying signal decomposition framework with perfect reconstruction via a combination of adaptive time-domain pre/post-processing and size-adaptive block DCTs. We explore different methods to produce time-varying basis functions and study various properties of the transition filter banks involved. Several criteria that could be used to select a certain set of basis functions are investigated. Promising coding gains over those of non-adaptive decompositions in an image coding setting are also presented.
Sachin Gangaputra, Trac D. Tran
ICASSP2
2002 DCT-based general structure for linear phase paraunitary filter banks
abstract
The factorization of linear-phase paraunitary filter banks (LPPUFB) has been well studied. In this paper, we show that it can be further simplified by fixing the last stage without losing its completeness. The structure can be viewed as the dual of the GenLOT when the last stage is chosen to be the DCT. However, the new structure is more flexible since it can generate basis functions of arbitrary length. The implementation for finite-length signals is discussed. A DCT-oriented initialization method for filter bank optimization is developed to improve its convergence. The proposed method leads to an effective way of handling the sign parameters when modeling orthogonal matrices via Givens rotations. As a result, better optimization results can be obtained.
Jie Liang 0001, Trac D. Tran
ICASSP2
2002 Further results on DCT-based linear phase paraunitary filter banks
abstract
A DCT-based simplified general structure for a linear phase paraunitary filter bank (LPPUFB) was developed by the pre- and post-processing of the DCT in the time domain. The new structure can be viewed as the dual of the generalized LOT (GenLOT). We generalize the result to odd-channel LPPUFB and LPPUFB with pair-wise mirror image property. Design examples and their application in image compression are presented.
Jie Liang 0001, Trac D. Tran
ICIP (2)2
2002 On context-based entropy coding of block transform coefficients
abstract
It has been well established that state-of-the-art wavelet image coders outperform block transform image coders in the rate-distortion (R-D) sense by a wide margin. An often asked question is: how much of the coding improvement is due to the transform and how much is due to the encoding strategy? A notable observation is that each block transform coefficient is highly correlated with its neighbors within the same block as well as its neighbors within the same subband. Current block transform coders suffer from poor context modeling and fail to take full advantage of intra- and inter-block correlation in both space and frequency sense. This paper presents a simple, fast and efficient adaptive block transform image coding algorithm based on high-order space-frequency, context modeling. Despite the simplicity constraints, coding results show that the proposed codec achieves competitive R-D performances comparing to the best wavelet codecs in the current literature.
Chengjie Tu, Trac D. Tran
ICIP (2)2
2002 Adaptive runlength coding
abstract
Runlength coding is the standard coding technique for block transform based image/video compression. A block of quantized transform coefficients is first represented as a sequence of RUN (number of consecutive zeros) / LEVEL (the value of the following nonzero coefficient) pairs which are then entropy coded. We point out in this paper the inefficiency of conventional runlength coding and introduce a novel adaptive runlength coding scheme that encodes RUN and LEVEL symbols separately using context based adaptive binary arithmetic coding. We aim to maximize compression efficiency by adaptively exploiting the characteristics of block transform coefficients and the dependency between RUN and LEVEL. Coding results show that, with the same level of complexity, the proposed adaptive runlength coding algorithm outperforms the conventional runlength coding scheme by a wide margin in the rate-distortion (R-D) sense.
Chengjie Tu, Trac D. Tran, Jie Bang
ICIP (2)2
2002 Multiplierless approximation of transforms with adder constraint
abstract
This letter describes an algorithm for systematically finding a multiplierless approximation of transforms by replacing floating-point multipliers with VLSI-friendly binary coefficients of the form k/2/sup n/. Assuming the cost of hardware binary shifters is negligible, the total number of binary adders employed to approximate the transform can be regarded as an index of complexity. Because the new algorithm is more systematic and faster than trial-and-error binary approximations with adder constraint, it is a much more efficient design tool. Furthermore, the algorithm is not limited to a specific transform; various approximations of the discrete cosine transform are presented as examples of its versatility.
Ying-Jui Chen, Soontorn Oraintara, Trac D. Tran, Kevin Amaratunga, Truong Q. Nguyen
IEEE Signal Process. Lett.3
2002 On the completeness of the lattice factorization for linear-phase perfect reconstruction filter banks
abstract
In this letter, we re-examine the completeness of the lattice factorization for M-channel linear-phase perfect reconstruction filter bank (LPPRFB) with filters of the same length L=KM as discussed by Tran et al. (see IEEE Trans. Signal Processing, vol.48, p.133-47, Jan. 2000). We point out that the assertion of completeness is incorrect. Examples are presented to show that the proposed lattice structure of Tran et al. is not complete when K>2. In addition, we verify that the lattice structure is complete only when K/spl les/2.
Lu Gan 0002, Kai-Kuang Ma, Truong Q. Nguyen, Trac D. Tran, Ricardo L. de Queiroz
IEEE Signal Process. Lett.4
2002 Context-based entropy coding of block transform coefficients for image compression
abstract
It has been well established that state-of-the-art wavelet image coders outperform block transform image coders in the rate-distortion (R-D) sense by a wide margin. Wavelet-based JPEG2000 is emerging as the new high-performance international standard for still image compression. An often asked question is: how much of the coding improvement is due to the transform and how much is due to the encoding strategy? Current block transform coders such as JPEG suffer from poor context modeling and fail to take full advantage of correlation in both space and frequency sense. This paper presents a simple, fast, and efficient adaptive block transform image coding algorithm based on a combination of prefiltering, postfiltering, and high-order space-frequency context modeling of block transform coefficients. Despite the simplicity constraints, coding results show that the proposed coder achieves competitive R-D performance compared to the best wavelet coders in the literature.
Chengjie Tu, Trac D. Tran
IEEE Trans. Image Process.2
2000 Seismic Data Compression Using GENLOT: Towards "Optimality"?
abstract
Summary form only given. Seismic data compression is desirable in geophysics for both storage and transmission stages. Wavelet coding methods have generated interesting developments, including a real-time field test trial in the North Sea in 1995. Previous work showed that GenLOT with basic optimization also outperforms state-of-the-art biorthogonal wavelet coders for seismic data. In this paper, we focus on the problem of filter bank optimization using various properties of seismic data. It is often desirable to evaluate the compression performance of a transform on a set of data using a priori objective measures, to reduce extensive testings by selecting only good a priori transforms, and to tailor transforms to the statistical properties of the data set. In the scope of this work, we use symmetric AR models up to order 4 to obtain an average model of the horizontal and vertical signals of a seismic stack section. Rosten et al. (1999), have already shown that order 1 or 2 models give good results in filter bank optimization for non-unitary filter banks, using coding gain optimization. Several other criteria may be used for transform optimization. Following the theory in Tran and Nguyen (1999), we use a weighted combination of C/sub o/=k/sub C/C/sub C/+k/sub S/C/sub S/+k/sub d/C/sub D/ of coding gain, stopband attenuation and DC leakage minimization functions.
Laurent Duval, Van Bui-Tran, Truong Q. Nguyen, Trac D. Tran
Data Compression Conference4
2000 GenLOT optimization techniques for seismic data compression
abstract
GenLOT coding has been shown an effective technique for seismic data compression, especially when compared to block-based algorithms (such as JPEG), or to wavelets. The transforms remove statistical redundancy and permit efficient compression, when used with advanced encoding techniques, such as the embedded zerotree coding framework. We derive a model for seismic data based on auto-regressive processes. This model is used to design GenLOT filter banks optimized for seismic data, using objective optimization criteria.
Laurent Duval, Van Bui-Tran, Truong Q. Nguyen, Trac D. Tran
ICASSP4
2000 Optimizing Block-Threshold Segmentation for MRC Compression
abstract
Compound document images contain graphic or textual content along with pictures. They are a very common form of documents, found in magazines, brochures, Web-sites, etc. We focus our attention on the mixed raster content (MRC) multi-layer approach for compound image compression. We study block thresholding as a means to segment an image for MRC. An attempt is made to optimize the block-threshold in a rate-distortion sense. Rate-distortion curves are presented to demonstrate the performance of the proposed algorithm.
Ricardo L. de Queiroz, Zhigang Fan 0001, Trac D. Tran
ICIP3
2000 The binDCT: fast multiplierless approximation of the DCT
abstract
This paper presents a family of fast biorthogonal block transforms called binDCT that can be implemented using only shift and add operations. The transform is based on a VLSI-friendly lattice structure that robustly enforces both linear phase and perfect reconstruction properties. The lattice coefficients are parameterized as a series of dyadic lifting steps providing fast, efficient, in place computation of the transform coefficients as well as the ability to map integers to integers. The new 8/spl times/8 transforms all approximate the popular 8/spl times/8 DCT closely, attaining a coding gain range of 8.77-8.82 dB, despite requiring as low as 14 shifts and 31 additions per eight input samples. Application of the binDCT in both lossy and lossless image coding yields very competitive results compared to the performance of the original floating-point DCT.
Trac D. Tran
IEEE Signal Process. Lett.1
2000 The LiftLT: fast-lapped transforms via lifting steps
abstract
This paper introduces a class of multiband linear phase-lapped biorthogonal transforms with fast, VLSI-friendly implementations via lifting steps called the LiftLT. The transform is based on a lattice structure that robustly enforces both linear phase and perfect reconstruction properties. The lattice coefficients are parameterized as a series of lifting steps, providing fast, efficient, in-place computation of the transform coefficients. The new transform is designed for applications in image and video coding. Compared to the popular 8/spl times/8 DCT, the 8/spl times/16 LiftLT only requires one more multiplication, 22 more additions, and six more shifting operations. However, image coding examples show that the LiftLT is far superior to the DCT in both objective and subjective coding performance. Thanks to properly designed overlapping basis functions, the LiftLT can completely eliminate annoying blocking artifacts. In fact, the novel LT's coding performance consistently surpasses that of the much more complex 9/7-tap biorthogonal wavelet with floating-point coefficients. More importantly, the transform's block-based nature facilitates one-pass sequential block coding, region-of-interest coding/decoding, and parallel processing.
Trac D. Tran
IEEE Signal Process. Lett.1
2000 Optimizing block-thresholding segmentation for multilayer compression of compound images
abstract
Compound document images contain graphic or textual content along with pictures. They are a very common form of documents, found in magazines, brochures, Web sites, etc. We focus our attention on the mixed raster content (MRC) multilayer approach for compound image compression. We study block thresholding as a means to segment an image for MRC. An attempt is made to optimize the block threshold in a rate-distortion sense. Also, a fast algorithm is presented to approximate the optimized method. Extensive results are presented including rate-distortion curves, segmentation masks and reconstructed images, showing the performance of the proposed algorithm.
Ricardo L. de Queiroz, Zhigang Fan 0001, Trac D. Tran
IEEE Trans. Image Process.3
1999 Local Zerotree Coding
abstract
In this paper, we introduce a novel transform-based image compression framework called local zerotree (LZT) coding. The main idea is to partition the transform coefficients into small groups, each of which is encoded independently using popular zerotree algorithms such as EZW or SPIHT. The advantage of the new coding algorithm is fourfold: (i) because of the reduction of memory buffering, LZT can reduce the complexity of the codec implementation and increase the speed of the zerotree algorithm significantly, especially in hardware; (ii) LZT is capable of processing large images under limited memory constraint; (iii) LZT supports parallel processing mode as long as the transform in use has that capability; and (iv) LZT facilitates the coding/decoding of regions of interest. Moreover, we shall demonstrate that the penalty in coding performance is minute comparing to its global zerotree predecessors: there are only a few extra bytes of side information. Essentially, the only property that LZT sacrifices is the embeddedness, which might not be crucial in many applications.
Trac D. Tran
ICIP (2)1
1999 The LIFTLT: Fast Lapped Transform Via Lifting Steps
abstract
This paper introduces a class of multi-band linear phase lapped biorthogonal transforms with fast, VLSI-friendly implementations via lifting steps called the LiftLT. The transform is based on a lattice structure which robustly enforces both linear phase and perfect reconstruction properties. The lattice coefficients are parameterized as a series of lifting steps, providing fast, efficient in-place computation of the transform coefficients. Our main motivation of the new transform is its application in image and video coding. Comparing to the popular 8/spl times/8 DCT, the 8/spl times/16 LiftLT only requires 1 more multiplication, 22 more additions, and 6 more shifting operations. However, image coding examples show that the LiftLT is far superior than the DCT in both objective and subjective coding performance. Thanks to properly designed overlapping basis functions, the LiftLT can completely eliminate annoying blocking artifacts. In fact, the novel LT's coding performance consistently surpasses that of the much more complex 9/7-tap biorthogonal wavelet with floating-point coefficients. More importantly, the transform's block-based nature facilitates one-pass sequential block coding, region-of-interest coding/decoding as well as parallel processing.
Trac D. Tran
ICIP (2)1
1999 A Fast Multiplierless Block Transform for Image and Video Compression
abstract
In this paper, we present a family of fast biorthogonal block transforms called binDCT that can be implemented using only shift and add operations. All transforms are based on a VLSI-friendly lattice structure which robustly enforces both linear phase and perfect reconstruction properties. The lattice coefficients are parameterized as a series of dyadic lifting steps, providing fast, efficient in-place computation of the transform coefficients as well as the ability to map integers to integers. The new 8/spl times/8 transforms approximate closely the popular 8/spl times/8 DCT, attaining coding gains from 8.77 to 8.82 dB, despite requiring a modest amount of computations: as low as 14 shifts and 31 additions per 8 input samples. Application of the novel transforms in both lossy and lossless image coding yields very competitive results compared to the performance of the original heating-point DCT.
Trac D. Tran
ICIP (3)1
1999 A progressive transmission image coder using linear phase uniform filterbanks as block transforms
abstract
This paper presents a novel image coding scheme using M-channel linear phase perfect reconstruction filterbanks (LPPRFBs) in the embedded zerotree wavelet (EZW) framework introduced by Shapiro (1993). The innovation here is to replace the EZWs dyadic wavelet transform by M-channel uniform-band maximally decimated LPPRFBs, which offer finer frequency spectrum partitioning and higher energy compaction. The transform stage can now be implemented as a block transform which supports parallel processing and facilitates region-of-interest coding/decoding. For hardware implementation, the transform boasts efficient lattice structures, which employ a minimal number of delay elements and are robust under the quantization of lattice coefficients. The resulting compression algorithm also retains all the attractive properties of the EZW coder and its variations such as progressive image transmission, embedded quantization, exact bit rate control, and idempotency. Despite its simplicity, our new coder outperforms some of the best image coders published previously in the literature, for almost all test images (especially natural, hard-to-code ones) at almost all bit rates.
Trac D. Tran, Truong Q. Nguyen
IEEE Trans. Image Process.1
1998 The generalized lapped biorthogonal transform
abstract
A lattice structure based on the singular value decomposition (SVD) is introduced. The lattice can be proven to use a minimal number of delay elements and to completely span a large class of M-channel linear phase perfect reconstruction filter banks (LPPRFB): all analysis and synthesis filters have the same FIR length of L=KM, sharing the same center of symmetry. The lattice also structurally enforces both linear phase and perfect reconstruction properties, is capable of providing fast and efficient implementation, and avoids the costly matrix inversion problem in the optimization process. From a block transform perspective, the new lattice represents a family of generalized lapped biorthogonal transforms (GLBT) with arbitrary integer overlapping factor K. The relaxation of the orthogonal constraint allows the GLBT to have significantly different analysis and synthesis basis functions which can then be tailored appropriately to fit a particular application. Several design examples are presented along with a high-performance GLBT-based progressive image coder to demonstrate the superiority of the new lapped transforms.
Trac D. Tran, Ricardo L. de Queiroz, Truong Q. Nguyen
ICASSP1
1998 A GenLOT-based Progressive Image Coder for Low Resolution Images
abstract
The popular EZW (embedded zerotree wavelet) and its improved version SPIHT (set partitioning in hierarchical trees) are high-performance progressive transmission image coders based on the wavelet transform which gives excellent compression results for images with significant low-frequency content. For images with high texture contents, the GenLOT-based coder outperforms SPIHT in PSNR measure by a wide margin. On the other hand, low bit rate video finds applications in videophone and surveillance systems, where smaller size image in QCIF format is often transmitted. We show that using the conventional zero-tree algorithm for the QCIF image is suboptimal and we propose several progressive algorithms with modified zero-tree structures. The extensive coding results using both the DCT and GenLOT transform confirms that our proposed modified zero-tree algorithm outperforms the conventional zerotree algorithm for QCIF-sized images.
Mika Helsingius, Trac D. Tran, Truong Q. Nguyen
ICIP (2)2
1998 Generalized Lapped Biorthogonal Transforms with Integer Coefficients
Masaaki Ikehara, Trac D. Tran, Truong Q. Nguyen
ICIP (3)2
1998 The Variable-Length Generalized Lapped Biorthogonal Transform
Trac D. Tran, Ricardo L. de Queiroz, Truong Q. Nguyen
ICIP (3)1
1996 A locally adaptive perceptual masking threshold model for image coding
abstract
This paper involves designing, implementing, and testing of a locally adaptive perceptual masking threshold model for image compression. This model computes, based on the contents of the original images, the maximum amount of noise energy that can be injected at each transform coefficient that results in perceptually distortion-free still images or sequences of images. The adaptive perceptual masking threshold model can be used as a pre-processor to a JPEG compression standard image coder. DCT coefficients less than their corresponding perceptual thresholds can be set to zero before the normal JPEG quantization and Huffman coding steps. The result is an image-dependent gain in the bit rate needed for transparent coding. In an informal subjective test involving 318 still images in the AT&T Bell Laboratory image database, this model provided a gain in bit-rate saving on the order of 10 to 30%.
Trac D. Tran, Robert J. Safranek
ICASSP1