EDBT 2026 Demo / reviewers in the wild / expert
Truong Q. Nguyen
dblp:23/5856 · also Truong Nguyen 0002
· DBLP profile ↗
316ranked-venue papers
9as first author
30since 2021 · last 2026
0000-0002-5022-063XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 293 · 5 first-author · 26 since 2021Systems, architecture and hardware · 13 · 4 first-authorArtificial intelligence and machine learning · 12 · 8 since 2021Databases, data management, data science and information retrieval · 4Computer networks · 3Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DynaGSLAM: Real-Time Gaussian-Splatting SLAM for Online Rendering, Tracking, Motion Predictions of Moving Objects in Dynamic ScenesabstractSimultaneous Localization and Mapping (SLAM) is one of the most important environment-perception and navigation algorithms for computer vision, robotics, and autonomous cars/drones. Hence, high quality and fast mapping becomes a fundamental problem. With the advent of 3D Gaussian Splatting (3DGS) as an explicit representation with excellent rendering quality and speed, state-of-the-art (SOTA) works introduce GS to SLAM. Compared to classical pointcloud-SLAM, GS-SLAM generates photometric information by learning from input camera views and synthesizing unseen views with high-quality textures. However, these GS-SLAM fail when moving objects occupy the scene that violates the static assumption of bundle adjustment. The failed updates of moving GS affects the static GS and contaminates the full map over the video sequence. Although some efforts have been made by concurrent works to consider moving objects for GS-SLAM, they simply detect and remove the moving regions from GS rendering ("anti" dynamic GS-SLAM), where only the static background could benefit from GS. To this end, we propose the first real-time GS-SLAM, "DynaGSLAM", that achieves high-quality online GS rendering, tracking, motion predictions of moving objects in dynamic scenes while jointly estimating accurate ego motion. Our DynaGSLAM outperforms SOTA static & "Anti" dynamic GS-SLAM on three dynamic real datasets, while keeping speed and memory efficiency in practice. https://blarklee.github.io/dynagslam/ Runfa Blark Li, Mahdi Shaghaghi, Keito Suzuki, Xinshuang Liu, Varun Moparthi, Bang Du, Walker Curtis, Martin Renschler, Ki Myung Brian Lee, Nikolay Atanasov 0001, Truong Q. Nguyen |
WACV | 11 |
| 2026 | KeyNode-Driven Geometry Coding for Real-World Scanned Human Dynamic Mesh CompressionabstractThe compression of real-world scanned 3D human dynamic meshes is an emerging research area, driven by applications such as telepresence, virtual reality, and 3D digital streaming. Unlike synthesized dynamic meshes with fixed topology, scanned dynamic meshes often not only have varying topology across frames but also scan defects such as holes and outliers, increasing the complexity of prediction and compression. Additionally, human meshes often combine rigid and non-rigid motions, making accurate prediction and encoding significantly more difficult compared to objects that exhibit purely rigid motion. To address these challenges, we propose a compression method designed for real-world scanned human dynamic meshes, leveraging embedded key nodes. The temporal motion of each vertex is formulated as a distance-weighted combination of transformations from neighboring key nodes, requiring the transmission of solely the key nodes' transformations. To enhance the quality of the KeyNode-driven prediction, we introduce an octree-based residual coding scheme and a Dual-direction prediction mode, which uses I-frames from both directions. Extensive experiments demonstrate that our method achieves significant improvements over the state-of-the-art, with an average bitrate savings of 58.43% across the evaluated sequences, particularly excelling at low bitrates. Huong Hoang, Truong Q. Nguyen, Pamela C. Cosman |
IEEE Trans. Multim. | 2 |
| 2025 | Open-Vocabulary Semantic Part Segmentation of 3D Humanabstract3D part segmentation is still an open problem in the field of 3D vision and AR/VR. Due to limited 3D labeled data, traditional supervised segmentation methods fall short in generalizing to unseen shapes and categories. Recently, the advancement in vision-language models' zero-shot abilities has brought a surge in open-world 3D segmentation methods. While these methods show promising results for$3 D$scenes or objects, they do not generalize well to 3D humans. In this paper, we present the first open-vocabulary segmentation method capable of handling 3D human. Our framework can segment the human category into desired fine-grained parts based on the textual prompt. We design a simple segmentation pipeline, leveraging SAM to generate multi-view proposals in 2D and proposing a novel Human-CLIP model to create unified embeddings for visual and textual inputs. Compared with existing pre-trained CLIP models, the HumanCLIP model yields more accurate embeddings for human-centric contents. We also design a simple-yet-effective MaskFusion module, which classifies and fuses multi-view features into 3D semantic masks without complex voting and grouping mechanisms. The design of decoupling mask proposals and text input also significantly boosts the efficiency of per-prompt inference. Experimental results on various 3D human datasets show that our method outperforms current state-of-the-art open-vocabulary 3D segmentation methods by a large margin. In addition, we showthat our method can be directly applied to various 3D representations including meshes, point clouds, and 3D Gaussian Splatting. Keito Suzuki, Bang Du, Girish Krishnan, Kunyao Chen, Runfa Blark Li, Truong Q. Nguyen |
3DV | 6 |
| 2025 | Class-Conditioned Image Synthesis with Diffusion for Imbalanced Diabetic Retinopathy Grading
Anna Heinke, Ines D. Nagel, Dirk-Uwe Bartsch, William R. Freeman, Truong Q. Nguyen, Cheolhong An |
MICCAI (4) | 6 |
| 2025 | Stop-band Energy Constraint for Orthogonal Tunable Wavelet Units in Convolutional Neural Networks for Computer Vision problemsabstractThis work introduces a stop-band energy constraint for filters in orthogonal tunable wavelet units with a lattice structure, aimed at improving image classification and anomaly detection in CNNs, especially on texture-rich datasets. Integrated into ResNet-18, the method enhances convolution, pooling, and downsampling operations, yielding accuracy gains of 2.48% on CIFAR-10 and 13.56% on the Describable Textures dataset. Similar improvements are observed in ResNet-34. On the MVTec hazelnut anomaly detection task, the proposed method achieves competitive results in both segmentation and detection, outperforming existing approaches. An Dinh Le, Sungbal Seo, You-Suk Bae, Truong Q. Nguyen |
VCIP | 5 |
| 2025 | DWTGS: Rethinking Frequency Regularization for Sparse-view 3D Gaussian SplattingabstractSparse-view 3D Gaussian Splatting (3DGS) presents significant challenges in reconstructing high-quality novel views, as it often overfits to the widely-varying high-frequency (HF) details of the sparse training views. While frequency regularization can be a promising approach, its typical reliance on Fourier transforms causes difficult parameter tuning and biases towards detrimental HF learning. We propose DWTGS, a framework that rethinks frequency regularization by leveraging wavelet-space losses that provide additional spatial supervision. Specifically, we supervise only the low-frequency (LF) LL subbands at multiple DWT levels, while enforcing sparsity on the HF HH subband in a self-supervised manner. Experiments across benchmarks show that DWTGS consistently outperforms Fourier-based counterparts, as this LF-centric strategy improves generalization and reduces HF hallucinations. Runfa Li, An Le, Truong Q. Nguyen |
VCIP | 4 |
| 2025 | Universal Vessel Segmentation for Multi-Modality Retinal ImagesabstractWe identify two major limitations in the existing studies on retinal vessel segmentation: 1) Most existing works are restricted to one modality, i.e., the Color Fundus (CF). However, multi-modality retinal images are used every day in the study of the retina and diagnosis of retinal diseases, and the study of vessel segmentation on other modalities is scarce; 2) Even though a few works extended their experiments to new modalities such as the Multi-Color Scanning Laser Ophthalmoscopy (MC), these works still require fine-tuning a separate model for the new modality. The fine-tuning will require extra training data, which is difficult to acquire. In this work, we present a novel universal vessel segmentation model (URVSM) for multi-modality retinal images. In addition to performing the study on a much wider range of image modalities, we also propose a universal model to segment the vessels in all these commonly used modalities. While being much more versatile compared with existing methods, our universal model also demonstrates comparable performance to the state-of-the-art fine-tuned methods. To the best of our knowledge, this is the first work that achieves modality-agnostic retinal vessel segmentation and the first to study retinal vessel segmentation in several novel modalities (Code, model and 3 new retinal vessel segmentation datasets are available at https://github.com/JRC-VPLab/URVSM). Anna Heinke, Akshay Agnihotri, Dirk-Uwe Bartsch, William R. Freeman, Truong Q. Nguyen, Cheolhong An |
IEEE Trans. Image Process. | 6 |
| 2024 | Select-Sliced Wasserstein Distance for Point Cloud LearningabstractQuantifying the discrepancy between point sets is a critical component for point cloud learning tasks. The mainstream point cloud learning tasks utilize Chamfer distance and Earth-Mover’s distance. Chamfer distance is computationally efficient but may not fully capture differences between sets of points. Earth-Mover’s distance, while precise, is computationally expensive and can be impractical to use with high-definition data. Several variants of Sliced Wasserstein distances (SW) are introduced to reduce the computation cost, but bring new problems to the situation: The vanilla SW treats sampled slices equally, resulting in redundant projections; Distributional Sliced Wasserstein distance requires gradient-based optimization, offsetting its benefits. To overcome this limitation and leverage the advantages of Sliced Wasserstein distance over EMD, we propose a novel metric, Select-Sliced Wasserstein distance. This new distance analyzes drawn samples of slices and quantifies their informativeness for each point in a single shot, which eliminates unnecessary projections as well as costly optimizations, but perpetuates the performance. Extensive experiments on various point cloud learning tasks to demonstrate the efficiency and effectiveness of the proposed distance metric. Our code is available at https://github.com/VideoProcessingLab/SSW_Distance Bang Du, Kunyao Chen, Truong Q. Nguyen |
3DV | 4 |
| 2024 | AUEditNet: Dual-Branch Facial Action Unit Intensity Manipulation with Implicit DisentanglementabstractFacial action unit (AU) intensity plays a pivotal role in quantifying fine-grained expression behaviors, which is an effective condition for facial expression manipulation. How-ever, publicly available datasets containing intensity annotations for multiple AUs remain severely limited, often featuring a restricted number of subjects. This limitation places challenges to the AU intensity manipulation in images due to disentanglement issues, leading researchers to resort to other large datasets with pretrained AU intensity estimators for pseudo labels. In addressing this constraint and fully leveraging manual annotations of AU intensities for precise manipulation, we introduce AUEditNet. Our proposed model achieves impressive intensity manipulation across 12 AUs, trained effectively with only 18 subjects. Utilizing a dual-branch architecture, our approach achieves comprehensive disentanglement of facial attributes and identity without necessitating additional loss functions or implementing with large batch sizes. This approach offers a potential solution to achieve desired facial attribute editing despite the dataset's limited subject count. Our experiments demonstrate AUEdit-Net's superior accuracy in editing AU intensities, affirming its capability in disentangling facial attributes and identity within a limited subject pool. AUEditNet allows conditioning by either intensity values or target images, eliminating the need for constructing AU combinations for specific facial expression synthesis. Moreover, AU intensity estimation, as a downstream task, validates the consistency between real and edited images, confirming the effectiveness of our proposed AU intensity manipulation method. Shiwei Jin, Zhen Wang 0009, Lei Wang 0018, Peng Liu 0039, Ning Bi, Truong Q. Nguyen |
CVPR | 6 |
| 2023 | ReDirTrans: Latent-to-Latent Translation for Gaze and Head RedirectionabstractLearning-based gaze estimation methods require large amounts of training data with accurate gaze annotations. Facing such demanding requirements of gaze data collection and annotation, several image synthesis methods were proposed, which successfully redirected gaze directions pre-cisely given the assigned conditions. However, these methods focused on changing gaze directions of the images that only include eyes or restricted ranges of faces with low res-olution (less than$128\times 128$) to largely reduce interference from other attributes such as hairs, which limits application scenarios. To cope with this limitation, we proposed a portable network, called ReDirTrans, achieving latent-to-latent translation for redirecting gaze directions and head orientations in an interpretable manner. ReDirTrans projects input latent vectors into aimed-attribute embed-dings only and redirects these embeddings with assigned pitch and yaw values. Then both the initial and edited embeddings are projected back (deprojected) to the initial latent space as residuals to modify the input latent vec-tors by subtraction and addition, representing old status re-moval and new status addition. The projection of aimed at-tributes only and subtraction-addition operations for status replacement essentially mitigate impacts on other attributes and the distribution of latent vectors. Thus, by combining ReDirTrans with a pretrained fixed e4e-StyleGAN pair, we created ReDirTrans-GAN, which enables accurately redi-recting gaze in full-face images with$1024\times 1024$resolution while preserving other attributes such as identity, expres-sion, and hairstyle. Furthermore, we presented improvements for the downstream learning-based gaze estimation task, using redirected samples as dataset augmentation. Shiwei Jin, Zhen Wang 0009, Lei Wang 0018, Ning Bi, Truong Q. Nguyen |
CVPR | 5 |
| 2023 | Encoder-Decoder Graph Convolutional Network for Automatic Timed-Up-and-Go and Sit-to-Stand SegmentationabstractVision-based action segmentation is an important tool in human movement analysis. In this work, we present a novel Encoder-Decoder Graph Convolutional Network (ED-GCN) to perform auto-segmentation on two widely accepted clinical tests for human mobility and balance assessment: the "Timed-Up-and-Go" (TUG) test and the "Sit-to-Stand" (STS) test. For STS, we perform a fine-grained segmentation that further segments the stand up and sit down actions into more sub-phases. To the best of our knowledge, this is the first work that analyzes such subtle segmentation with biomedical significance. We also propose two novel metrics for action segmentation, which overcome some key drawbacks in the popular F1 and Edit scores. Experiment shows that our network has superior performance over state-of-the-art action segmentation networks in TUG and STS segmentation. Truong Q. Nguyen |
ICASSP | 3 |
| 2023 | Accurate Registration between Ultra-Wide-Field and Narrow Angle Retina Images with 3D Eyeball Shape OptimizationabstractThe Ultra-Wide-Field (UWF) retina images have attracted wide attentions in recent years in the study of retina. However, accurate registration between the UWF images and the other types of retina images could be challenging due to the distortion in the peripheral areas of an UWF image, which a 2D warping can not handle. In this paper, we propose a novel 3D distortion correction method which sets up a 3D projection model and optimizes a dense 3D retina mesh to correct the distortion in the UWF image. The corrected UWF image can then be accurately aligned to the target image using 2D alignment methods. The experimental results show that our proposed method outperforms the state-of-the-art method by 30%. Junkang Zhang, Fritz Gerald P. Kalaw, Melina Cavichini, Dirk-Uwe Bartsch, William R. Freeman, Truong Q. Nguyen, Cheolhong An |
ICIP | 7 |
| 2023 | A Novel Learnable Orthogonal Wavelet Unit Neural Network with Perfection Reconstruction Constraint Relaxation for Image ClassificationabstractCNNs utilize lowpass frequency features, evidenced by max pooling operations that preserve dominant features. Addressing this, we used a one-layer FCN using both coarse and detailed components via DWT. We then developed UwU, a learnable wavelet-based unit that integrates the PR constraint relaxation, allowing feature map component fine-tuning. Distinctively, UwU’s coefficients are trainable, unlike prior studies. This is the first work that utilizes PR constraint relaxation to enhance CNNs. Our innovative techniques serve to enhance stride-convolution, pooling, and downsampling units in CNNs. We tested these improved units using the ResNet family architectures against traditional frequency and wavelet-based units. Performance metrics from CIFAR10, ImageNet1K, and the DTD show promising results. Particularly, while CIFAR10 results are on par with other methods, there’s a marked performance improvement on ImageNet1K and DTD. Further, integrating UwU into the ResNet18-based encoder of the CFLOW-AD system yields competitive anomaly detection results, particularly for hazelnut images in the MVTecAD dataset. An Dinh Le, Shiwei Jin, You Suk Bae, Truong Q. Nguyen |
VCIP | 4 |
| 2023 | AMCC-Net: An asymmetric multi-cross convolution for skin lesion segmentation on dermoscopic images
Chaitra Dayananda, Nagaraj Yamanakkanavar, Truong Q. Nguyen, Bumshik Lee |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | View-Invariant Center-of-Pressure Metrics Estimation With Monocular RGB CameraabstractCenter of pressure (CoP) metrics, including CoP path length and sway area, have been used as gold standard measurements of postural and balance control in biomechanical studies. A recent study of computer-vision-based CoP metrics estimation from 3D body landmark sequences offers a more portable and comprehensive solution than conventional force plate methods to obtain these important metrics for real-time evaluation of balance control. However, obtaining accurate 3D body landmarks requires a calibrated motion capture system or on-body markers, which involves lengthy data collection and processing time and limits their implementation in home and clinical environments. Existing methods that instead use 2D body landmarks fail to adapt to different camera positions. To overcome these challenges, we propose a view-invariant deep learning framework for video-level CoP metrics estimation, including CoP path length and sway area, using pose dimension lifting and graph convolutional network (GCN). This work is the first step toward obtaining gold-standard CoP metrics with an accessible, monocular RGB camera. We propose to use a dimension lifting convolutional neural network (CNN) to obtain view-invariant 3D body landmark features from 2D body landmarks. We also propose a two-stream regression model using GCN and discrete cosine transform (DCT) for a robust CoP metrics estimation. To facilitate the line of research, we release a novel multi-view body landmark dataset containing 2D body landmarks of a wide variety of action patterns from four different camera views with synchronized CoP labels and corresponding 3D body landmarks, which enables cross-view evaluation with different camera angles. We subsequently validate the proposed method through a cross-dataset training by training the dimension lifting model on an existing balance dataset and evaluating the CoP metrics estimation on the multi-view body landmark dataset. The experiments validate that our framework achieves state-of-the-art accuracy for both CoP path length and CoP sway area using a monocular RGB camera input for unseen views. Sarah Graham, Colin A. Depp, Truong Q. Nguyen |
IEEE Trans. Multim. | 4 |
| 2023 | Efficient Registration for Human Surfaces via Isometric Regularization on Embedded Deformationabstract3D registration is a fundamental step to obtain the correspondences between surfaces. Traditional mesh alignment methods tackle this problem through non-rigid deformation, mostly accomplished by applying ICP-based (Iterative Closest Point) optimization. The embedded deformation method is proposed for the purpose of acceleration, which enables various real-time applications. However, it regularizes on an underlying simplified structure, which could be problematic for intricate cases when the simplified graph doesn't fully represent the surface attributes. Moreover, without elaborate parameter-tuning, deformation usually performs suboptimally, leading to slow convergence or a local minimum if all regions on the surface are assumed to share the same rigidity during the optimization. In this article, we propose a novel solution that decouples regularization from the underlying deformation model by explicitly managing the rigidity of vertex clusters. We further design an efficient two-step solution that alternates between isometric deformation and embedded deformation with cluster-based regularization. Our method can easily support region-adaptive regularization with cluster refinement and execute efficiently. Extensive experiments demonstrate the effectiveness of our approach for mesh alignment tasks even under large-scale deformation and imperfect data. Our method outperforms state-of-the-art methods both numerically and visually. Kunyao Chen, Bang Du, Baichuan Wu, Truong Q. Nguyen |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Data Augmentation Methods For Object Detection and Segmentation In Ultrasound Scans: An Empirical Comparative StudyabstractIn ultrasound imaging, sonographers are tasked with analyzing scans for diagnostic purposes; a challenging task, especially for novice sonographers. Deep Learning methods have shown great potential in their ability to infer semantics and key information from scans to assist with these tasks. However, deep learning methods require large training sets to accomplish tasks such as segmentation and object detection. Generating these large datasets is a significant challenge in the medical domain due to the high cost of acquisition and annotation. Therefore, data augmentation is used to increase the size of training datasets to create the needed variability for deep learning models to generalize. These augmentation methods try to mimic differences among scans that result from noise, tissue movement, acquisition settings, and others. In this paper, we analyze the effectiveness of general augmentation methods that perform color, rigid, and non-rigid geometric transformation, to empirically analyze and compare their ability to improve the performance of three segmentation architectures on three different ultrasound datasets. We observe that non-rigid geometric transformations produce the best performance improvement. Sachintha R. Brandigampala, Abdullah F. Al-Battal, Truong Q. Nguyen |
CBMS | 3 |
| 2022 | MonoPLFlowNet: Permutohedral Lattice FlowNet for Real-Scale 3D Scene Flow Estimation with Monocular Images
Runfa Li, Truong Q. Nguyen |
ECCV (27) | 2 |
| 2022 | Object Detection and Tracking in Ultrasound Scans Using an Optical Flow and Semantic Segmentation Framework Based on Convolutional Neural NetworksabstractBased on non-ionizing radiation, ultrasound scanning is safe to image a specific region of the body repeatedly to identify and localize target anatomical structures during therapeutic and diagnostic procedures. However, it is labor intensive, and requires sonographers to have extensive experience to be able to identify and track these anatomical structures of interest, making the identification and tracking process highly prone to errors. In this paper, we propose a framework to autonomously detect, localize and track anatomical structures in ultrasound scans during scanning and therapeutic sessions in real-time. The proposed framework uses a segmentation-based convolutional neural network (CNN) to detect and localize the target anatomical structure within a scan. Concurrently, it uses an optical flow CNN to track the movement of this structure across frames to accurately guide therapeutic procedures. We tested the framework on detecting and tracking the Vagus nerve in ultrasound scans. It achieved state-of-the art localization and tracking accuracy with an average error of less than 1.25 mm for localization and 0.75 mm for tracking while maintaining an inference time of less than 35 ms. Abdullah F. Al-Battal, Imanuel R. Lerman, Truong Q. Nguyen |
ICASSP | 3 |
| 2022 | Joint Motion Correction and 3D Segmentation with Graph-Assisted Neural Networks for Retinal OCTabstractOptical Coherence Tomography (OCT) is a widely used non-invasive high resolution 3D imaging technique for biological tissues and plays an important role in ophthalmology. OCT retinal layer segmentation is a fundamental image processing step for OCT-Angiography projection, and disease analysis. A major problem in retinal imaging is the motion artifacts introduced by involuntary eye movements. In this paper, we propose neural networks that jointly correct eye motion and retinal layer segmentation utilizing 3D OCT information, so that the segmentation among neighboring B-scans would be consistent. The experimental results show both visual and quantitative improvements by combining motion correction and 3D OCT layer segmentation comparing to conventional and deep-learning based 2D OCT layer segmentation. Carlo Miguel B. Galang, William R. Freeman, Truong Q. Nguyen, Cheolhong An |
ICIP | 4 |
| 2022 | Self-Supervised Rigid Registration for Multimodal Retinal ImagesabstractThe ability to accurately overlay one modality retinal image to another is critical in ophthalmology. Our previous framework achieved the state-of-the-art results for multimodal retinal image registration. However, it requires human-annotated labels due to the supervised approach of the previous work. In this paper, we propose a self-supervised multimodal retina registration method to alleviate the burdens of time and expense to prepare for training data, that is, aiming to automatically register multimodal retinal images without any human annotations. Specially, we focus on registering color fundus images with infrared reflectance and fluorescein angiography images, and compare registration results with several conventional and supervised and unsupervised deep learning methods. From the experimental results, the proposed self-supervised framework achieves a comparable accuracy comparing to the state-of-the-art supervised learning method in terms of registration accuracy and Dice coefficient. Cheolhong An, Junkang Zhang, Truong Q. Nguyen |
IEEE Trans. Image Process. | 4 |
| 2022 | Two-Step Registration on Multi-Modal Retinal Images via Deep Neural NetworksabstractMulti-modal retinal image registration plays an important role in the ophthalmological diagnosis process. The conventional methods lack robustness in aligning multi-modal images of various imaging qualities. Deep-learning methods have not been widely developed for this task, especially for the coarse-to-fine registration pipeline. To handle this task, we propose a two-step method based on deep convolutional networks, including a coarse alignment step and a fine alignment step. In the coarse alignment step, a global registration matrix is estimated by three sequentially connected networks for vessel segmentation, feature detection and description, and outlier rejection, respectively. In the fine alignment step, a deformable registration network is set up to find pixel-wise correspondence between a target image and a coarsely aligned image from the previous step to further improve the alignment accuracy. Particularly, an unsupervised learning framework is proposed to handle the difficulties of inconsistent modalities and lack of labeled training data for the fine alignment step. The proposed framework first changes multi-modal images into a same modality through modality transformers, and then adopts photometric consistency loss and smoothness loss to train the deformable registration network. The experimental results show that the proposed method achieves state-of-the-art results in Dice metrics and is more robust in challenging cases. Junkang Zhang, Ji Dai, Melina Cavichini, Dirk-Uwe Bartsch, William R. Freeman, Truong Q. Nguyen, Cheolhong An |
IEEE Trans. Image Process. | 7 |
| 2022 | Multi-Task Center-of-Pressure Metrics Estimation With Graph Convolutional NetworkabstractCenter of pressure (CoP) metrics, including its path length, sway area, and position, are important measurements of postural and balance control in biomechanical studies. A computer-vision-based CoP metrics estimation system offers a portable solution to obtain these gold-standard metrics with 3D multi-joint coordination underlying body movements for real-time evaluation of balance control. In this paper, we propose an end-to-end framework for video-level estimation of CoP path length and sway area, as well as the frame-level estimation of CoP position, utilizing the spatial-temporal features and adaptive graph structure learned by graph convolution network. This work is the first step toward demonstrating that these gold-standard metrics can be obtained with a more comprehensive tool than current force plate technologies. We propose two single-task models for video-level and frame-level estimation, respectively, and a multi-task learning approach that jointly learns the two-temporal-level features. To facilitate this line of research, we release a novel computer-vision-based 3D body landmark dataset containing a wide variety of action patterns with synchronized CoP labels using pose estimation. We also adapt our framework on an existing kinematic dataset collected by wearable markers. The experiments on both datasets validate that our framework achieves state-of-the-art accuracies for all metric estimations, while the proposed multi-task approach yields the most accurate and robust performance on video-level estimation.1 Sarah Graham, Colin A. Depp, Truong Q. Nguyen |
IEEE Trans. Multim. | 4 |
| 2021 | In Defense of Scene Graphs for Image CaptioningabstractThe mainstream image captioning models rely on Convolutional Neural Network (CNN) image features to generate captions via recurrent models. Recently, image scene graphs have been used to augment captioning models so as to leverage their structural semantics, such as object entities, relationships and attributes. Several studies have noted that the naive use of scene graphs from a black-box scene graph generator harms image captioning performance and that scene graph-based captioning models have to incur the overhead of explicit use of image features to generate decent captions. Addressing these challenges, we propose SG2Caps, a framework that utilizes only the scene graph labels for competitive image captioning performance. The basic idea is to close the semantic gap between the two scene graphs - one derived from the input image and the other from its caption. In order to achieve this, we leverage the spatial location of objects and the Human-Object-Interaction (HOI) labels as an additional HOI graph. SG2Caps outperforms existing scene graph-only captioning models by a large margin, indicating scene graphs as a promising representation for image captioning. Direct utilization of scene graph labels avoids expensive graph convolutions over high-dimensional CNN features resulting in 49% fewer trainable parameters. Our code is available at: https://github.com/Kien085/SG2Caps Kien Nguyen 0006, Subarna Tripathi, Bang Du, Tanaya Guha, Truong Q. Nguyen |
ICCV | 5 |
| 2021 | Mesh Completion with Virtual ScansabstractMeshes generated by range scanners are often incomplete and contain complex holes due to limited input coverage and occlusion. In this paper, we present an effective method to fill the gap regions on meshes by leveraging the templates inferred from the learning-based method. We first segment both source and template models into corresponding parts. Each part will be aligned with non-rigid deformation. We then modify the gap regions by “virtual” depth maps rendered using the aligned parts from newly selected viewpoints. Comparing with the template-based mesh completion approaches, our algorithm can generate natural appearances without any user interaction. Comparing with the state-of-the-art volumetric fusion methods, our approach supports selective blending, which only modifies the regions of interest and prevents bad template inference from impacting the source. Kunyao Chen, Baichuan Wu, Bang Du, Truong Q. Nguyen |
ICIP | 5 |
| 2021 | SM3D: Simultaneous Monocular Mapping and 3D DetectionabstractMapping and 3D detection are two major issues in vision-based robotics, and self-driving. While previous works only focus on each task separately, we present an innovative and efficient multi-task deep learning framework (SM3D) for Simultaneous Mapping and 3D Detection by bridging the gap with robust depth estimation and “Pseudo-Lidar” point cloud for the first time. The Mapping module takes consecutive monocular frames to generate depth and pose estimation. In 3D Detection module, the depth estimation is projected into 3D space to generate “Pseudo-Lidar” point cloud, where Lidar-based 3D detector can be leveraged on point cloud for vehicular 3D detection and localization. By end-to-end training of both modules, the proposed mapping and 3D detection method outperforms the state-of-the-art baseline by 10.0% and 13.2% in accuracy, respectively. While achieving better accuracy, our monocular multi-task SM3D is more than 2 times faster than the state of the art pure stereo 3D detector, and 18.3% faster than using two modules separately. Runfa Li, Truong Q. Nguyen |
ICIP | 2 |
| 2021 | Learning to Correct Axial Motion in Oct for 3D Retinal ImagingabstractOptical Coherence Tomography (OCT) is a powerful technique for non-invasive 3D imaging of biological tissues at high resolution that has revolutionized retinal imaging. A major challenge in OCT imaging is the motion artifacts introduced by involuntary eye movements. In this paper, we propose a convolutional neural network that learns to correct axial motion in OCT based on a single volumetric scan. The proposed method is able to correct large motion, while preserving the overall curvature of the retina. The experimental results show significant improvements in visual quality as well as overall error compared to the conventional methods in both normal and disease cases. Alexandra Warter, Melina Cavichini, William R. Freeman, Dirk-Uwe Bartsch, Truong Q. Nguyen, Cheolhong An |
ICIP | 6 |
| 2021 | Compressed image restoration via deep deblocker driven unified framework
Chao Ren 0002, Qizhi Teng, Xiaohai He, Linbo Qing, Truong Q. Nguyen |
Knowl. Based Syst. | 5 |
| 2021 | Learning Image Profile Enhancement and Denoising Statistics Priors for Single-Image Super-ResolutionabstractSingle-image super-resolution (SR) has been widely used in computer vision applications. The reconstruction-based SR methods are mainly based on certain prior terms to regularize the SR problem. However, it is very challenging to further improve the SR performance by the conventional design of explicit prior terms. Because of the powerful learning ability, deep convolutional neural networks (CNNs) have been widely used in single-image SR task. However, it is difficult to achieve further improvement by only designing the network architecture. In addition, most existing deep CNN-based SR methods learn a nonlinear mapping function to directly map low-resolution (LR) images to desirable high-resolution (HR) images, ignoring the observation models of input images. Inspired by the split Bregman iteration (SBI) algorithm, which is a powerful technique for solving the constrained optimization problems, the original SR problem is divided into two subproblems: 1) inversion subproblem and 2) denoising subproblem. Since the inversion subproblem can be regarded as an inversion step to reconstruct an intermediate HR image with sharper edges and finer structures, we propose to use deep CNN to capture low-level explicit image profile enhancement prior (PEP). Since the denoising subproblem aims to remove the noise in the intermediate image, we adopt a simple and effective denoising network to learn implicit image denoising statistics prior (DSP). Furthermore, the penalty parameter in SBI is adaptively tuned during the iterations for better performance. Finally, we also prove the convergence of our method. Thus, the deep CNNs are exploited to capture both implicit and explicit image statistics priors. Due to SBI, the SR observation model is also leveraged. Consequently, it bridges between two popular SR approaches: 1) learning-based method and 2) reconstruction-based method. Experimental results show that the proposed method achieves the state-of-the-art SR results. Chao Ren 0002, Xiaohai He, Yi-Fei Pu, Truong Q. Nguyen |
IEEE Trans. Cybern. | 4 |
| 2021 | Robust Content-Adaptive Global Registration for Multimodal Retinal Images Using Weakly Supervised Deep-Learning FrameworkabstractMultimodal retinal imaging plays an important role in ophthalmology. We propose a content-adaptive multimodal retinal image registration method in this paper that focuses on the globally coarse alignment and includes three weakly supervised neural networks for vessel segmentation, feature detection and description, and outlier rejection. We apply the proposed framework to register color fundus images with infrared reflectance and fluorescein angiography images, and compare it with several conventional and deep learning methods. Our proposed framework demonstrates a significant improvement in robustness and accuracy reflected by a higher success rate and Dice coefficient compared with other methods. Junkang Zhang, Melina Cavichini, Dirk-Uwe Bartsch, William R. Freeman, Truong Q. Nguyen, Cheolhong An |
IEEE Trans. Image Process. | 6 |
| 2020 | Screen-space Regularization on Differentiable Rasterization
Kunyao Chen, Cheolhong An, Truong Q. Nguyen |
3DV | 3 |
| 2020 | Multi-Task Center-Of-Pressure Metrics Estimation from Skeleton Using Graph Convolutional NetworkabstractCenter of pressure (COP) is an important measurement of postural and gait control in human biomechanical studies. A vision-based estimation of COP metrics offers a way to obtain these gold-standard metrics for the detection of balance and gait problems. In this paper, we propose an end-to-end framework to estimate the COP path length and the COP positions from the 3D skeleton, utilizing the spatial-temporal features learned by graph convolutional networks. We propose two single-task models for each metric and a multi-task approach jointly learning two metrics. To facilitate this line of research, we also release a novel 3D skeleton dataset containing a wide variety of action patterns with synchronized COP labels. The experiments on the dataset validate that our framework achieves state-of-the-art accuracies for both COP path length and COP position estimations, while the multitask approach could yield more accurate and robust performance on COP path length estimation compared to the single-task model. Sarah Graham, Shiwei Jin, Colin A. Depp, Truong Q. Nguyen |
ICASSP | 5 |
| 2020 | A Segmentation Based Robust Deep Learning Framework for Multimodal Retinal Image RegistrationabstractMultimodal image registration plays an important role in diagnosing and treating ophthalmologic diseases. In this paper, a deep learning framework for multimodal retinal image registration is proposed. The framework consists of a segmentation network, feature detection and description network, and an outlier rejection network, which focuses only on the globally coarse alignment step using the perspective transformation. We apply the proposed framework to register color fundus images with infrared reflectance images and compare it with the state-of-the-art conventional and learning-based approaches. The proposed framework demonstrates a significant improvement in robustness and accuracy reflected by a higher success rate and Dice coefficient compared to other coarse alignment methods. Junkang Zhang, Cheolhong An, Melina Cavichini, Mahima Jhingan, Manuel J. Amador-Patarroyo, Christopher P. Long, Dirk-Uwe Bartsch, William R. Freeman, Truong Q. Nguyen |
ICASSP | 10 |
| 2020 | Adaptive image coding efficiency enhancement using deep convolutional neural networks
Honggang Chen, Xiaohai He, Cheolhong An, Truong Q. Nguyen |
Inf. Sci. | 4 |
| 2020 | IENet: Internal and External Patch Matching ConvNet for Web Image Guided DenoisingabstractFrom the non-local self-similarity (NSS)-based image denoising to the convolutional-network (ConvNet)-based image denoising, the denoising performance has been greatly improved. However, it is still not clear how to utilize similar web images to guide image denoising using ConvNet. This paper proposes a novel ConvNet for image denoising to explore both internal (NSS) and external correlations when external similar images are available. Since external similar images may be taken with different viewpoints, focal lengths, and may contain different objects, it is difficult to directly explore external correlations at image level using ConvNet. Therefore, we propose an internal and external patch matching ConvNet (IENet), whose inputs are similar patch cubes extracted from the noisy input and its external similar images. We design three different network structures, namely early-fusion, middle-fusion, and late-fusion of the internal and external cubes to fully combine the strengths of internal and external correlations. The experimental results demonstrate that the proposed method achieves the best denoising results compared with the seven state-of-the-art denoising methods. In specific, the proposed method outperforms the state-of-the-art web image guided denoising method by more than 1 dB on average, which further demonstrates the superiority of the proposed IENet-based filtering over the hand-crafted filtering methods. Huanjing Yue, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Truong Q. Nguyen, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Boosting Feature Matching Accuracy With Pairwise Affine EstimationabstractLocal image feature matching lies in the heart of many computer vision applications. Achieving high matching accuracy is challenging when significant geometric difference exists between the source and target images. The traditional matching pipeline addresses the geometric difference by introducing the concept of support region. Around each feature point, the support region defines a neighboring area characterized by estimated attributes like scale, orientation, affine shape, etc. To correctly assign support region is not an easy job, especially when each feature is processed individually. In this paper, we propose to estimate the relative affine transformation for every pair of to-be-compared features. This "tailored" measurement of geometric difference is more precise and helps improve the matching accuracy. Our pipeline can be incorporated into most existing 2D local image feature detectors and descriptors. We comprehensively evaluate its performance with various experiments on a diversified selection of benchmark datasets. The results show that the majority of tested detectors/descriptors gain additional matching accuracy with proposed pipeline. Ji Dai, Shiwei Jin, Junkang Zhang, Truong Q. Nguyen |
IEEE Trans. Image Process. | 4 |
| 2019 | Explicit Learning of Feature Orientation EstimationabstractWhile many learning-driven algorithms for local feature detection and description have submerged during recent years. One key component in the pipeline, namely orientation estimation, still remains underdeveloped. Among all sorts of difficulties, the impracticality and tedium of finding a "ground truth" feature orientation as a learning target is one big challenge. In this paper, we bypass this "thinking trap" and propose an unsupervised scheme that explicitly trains a simple convolutional neural network to predict orientations for feature points. Together with a carefully designed loss term, the network manages to provide accurate orientation estimations. We further evaluate the capability of this estimator in two experiments: orientation estimation and feature matching. Results showed the proposed method outperforms other compared methods on multiple benchmark datasets. The pretrained model is publicly available. Ji Dai, Junkang Zhang, Truong Q. Nguyen |
ICIP | 3 |
| 2019 | Joint Vessel Segmentation and Deformable Registration on Multi-Modal Retinal Images Based on Style TransferabstractIn multi-modal retinal image registration task, there are two major challenges, i.e., poor performance in finding correspondence due to inconsistent features, and lack of labeled data for training learning-based models. In this paper, we propose a joint vessel segmentation and deformable registration model based on CNN for this task, built under the framework of weakly supervised style transfer learning and perceptual loss. In vessel segmentation, a style loss guides the model to generate segmentation maps that look authentic, and helps transform images of different modalities into consistent representations. In deformable registration, a content loss helps find dense correspondence for multi-modal images based on their consistent representations, and improves the segmentation results simultaneously. Experiment results show that our model has better performance than other deformable registration methods in both quantitative and visual evaluations, and the segmentation results also help the rigid transformation1. Junkang Zhang, Cheolhong An, Ji Dai, Manuel Amador, Dirk-Uwe Bartsch, Shyamanga Borooah, William R. Freeman, Truong Q. Nguyen |
ICIP | 8 |
| 2019 | Deep Wide-Activated Residual Network Based Joint Blocking and Color Bleeding Artifacts Reduction for 4: 2: 0 JPEG-Compressed ImagesabstractBlocking and color bleeding are two well-known artifacts for 4:2:0 JPEG-compressed images. Blocking mainly results from the block-level quantization of the luma component, while color bleeding is mainly caused by the subsampling and quantization of chroma components. Restoring luma can reduce blocking distortion, but with little influence on color bleeding. On the contrary, color bleeding can be removed via chroma components restoration. This letter proposes a deep wide-activated residual network for reducing blocking and color bleeding artifacts simultaneously, in which the luma and chroma components are jointly restored. Chroma components usually suffer from more severe distortion than the luma component due to subsampling and coarse quantization. Thus, we use the luma component to guide the restoration of chroma components. Moreover, we reduce blocking and color bleeding artifacts in low-resolution space via pixel shuffle-based decimation and assembling, which allows to obtain high restoration speed. Experimental results show that the proposed approach achieves state-of-the-art performance on joint blocking and color bleeding artifacts reduction. Honggang Chen, Xiaohai He, Cheolhong An, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 4 |
| 2019 | Random Forest With Learned Representations for Semantic SegmentationabstractWe present a random forest framework that learns the weights, shapes, and sparsities of feature representations for real-time semantic segmentation. Typical filters (kernels) have predetermined shapes and sparsities and learn only weights. A few feature extraction methods fix weights and learn only shapes and sparsities. These predetermined constraints restrict learning and extracting optimal features. To overcome this limitation, we propose an unconstrained representation that is able to extract optimal features by learning weights, shapes, and sparsities. We, then, present the random forest framework that learns the flexible filters using an iterative optimization algorithm and segments input images using the learned representations. We demonstrate the effectiveness of the proposed method using a hand segmentation dataset for hand-object interaction and using two semantic segmentation datasets. The results show that the proposed method achieves real-time semantic segmentation using limited computational and memory resources. Byeongkeun Kang, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2019 | Accelerating GMM-Based Patch Priors for Image Restoration: Three Ingredients for a 100× Speed-UpabstractImage restoration methods aim to recover the underlying clean image from corrupted observations. The expected patch log-likelihood (EPLL) algorithm is a powerful image restoration method that uses a Gaussian mixture model (GMM) prior on the patches of natural images. Although it is very effective for restoring images, its high runtime complexity makes the EPLL ill-suited for most practical applications. In this paper, we propose three approximations to the original EPLL algorithm. The resulting algorithm, which we call the fast-EPLL (FEPLL), attains a dramatic speed-up of two orders of magnitude over EPLL while incurring a negligible drop in the restored image quality (less than 0.5 dB). We demonstrate the efficacy and versatility of our algorithm on a number of inverse problems, such as denoising, deblurring, super-resolution, inpainting, and devignetting. To the best of our knowledge, the FEPLL is the first algorithm that can competitively restore a pixel image in under 0.5 s for all the degradations mentioned earlier without specialized code optimizations, such as CPU parallelization or GPU implementation. Shibin Parameswaran, Charles-Alban Deledalle, Loïc Denis, Truong Q. Nguyen |
IEEE Trans. Image Process. | 4 |
| 2019 | Enhanced Non-Local Total Variation Model and Multi-Directional Feature Prediction Prior for Single Image Super ResolutionabstractIt is widely acknowledged that single image super-resolution (SISR) methods play a critical role in recovering the missing high-frequencies in an input low-resolution image. As SISR is severely ill-conditioned, image priors are necessary to regularize the solution spaces and generate the corresponding high-resolution image. In this paper, we propose an effective SISR framework based on the enhanced non-local similarity modeling and learning-based multi-directional feature prediction (ENLTV-MDFP). Since both the modeled and learned priors are exploited, the proposed ENLTV-MDFP method benefits from the complementary properties of the reconstruction-based and learning-based SISR approaches. Specifically, for the non-local similarity-based modeled prior [enhanced non-local total variation, (ENLTV)], it is characterized via the decaying kernel and stable group similarity reliability schemes. For the learned prior [multi-directional feature prediction prior, (MDFP)], it is learned via the deep convolutional neural network. The modeled prior performs well in enhancing edges and suppressing visual artifacts, while the learned prior is effective in hallucinating details from external images. Combining these two complementary priors in the MAP framework, a combined SR cost function is proposed. Finally, the combined SR problem is solved via the split Bregman iteration algorithm. Based on the extensive experiments, the proposed ENLTV-MDFP method outperforms many state-of-the-art algorithms visually and quantitatively. Chao Ren 0002, Xiaohai He, Yi-Fei Pu, Truong Q. Nguyen |
IEEE Trans. Image Process. | 4 |
| 2019 | High ISO JPEG Image Denoising by Deep Fusion of Collaborative and Convolutional FilteringabstractCapturing images at high ISO modes will introduce much realistic noise, which is difficult to be removed by traditional denoising methods. In this paper, we propose a novel denoising method for high ISO JPEG images via deep fusion of collaborative and convolutional filtering. Collaborative filtering explores the non-local similarity of natural images, while convolutional filtering takes advantage of the large capacity of convolutional neural networks (CNNs) to infer noise from noisy images. We observe that the noise variance map of a high ISO JPEG image is spatial-dependent and has a Bayer-like pattern. Therefore, we introduce the Bayer pattern prior in our noise estimation and collaborative filtering stages. Since collaborative filtering is good at recovering repeatable structures and convolutional filtering is good at recovering irregular patterns and removing noise in flat regions, we propose to fuse the strengths of the two methods via deep CNN. The experimental results demonstrate that our method outperforms the state-of-the-art realistic noise removal methods for a wide variety of testing images in both subjective and objective measurements. In addition, we construct a dataset with noisy and clean image pairs for high ISO JPEG images to facilitate research on this topic. Huanjing Yue, Jing-Yu Yang 0002, Truong Q. Nguyen, Feng Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Adjusted Non-Local Regression and Directional Smoothness for Image RestorationabstractImage restoration (IR) problems are very important in many low-level vision tasks. Due to their ill-posed natures, image priors are widely used to regularize the solution spaces. Recently, patch-based non-local self-similarity has shown great potential in IR problems, leading to many effective non-local priors. Their performance largely depends on whether the non-local self-similarity of the underlying image can be fully exploited. However, most of these priors, including non-local regression (NLR), only utilize the center pixel of each patch to model the non-local feature, which is suboptimal. We propose an effective overlap-based non-local regression (ONLR) to fully exploit the non-local similar patches: first, the concept of overlap-based similar pixels group (OSPG) is introduced; second, for each pixel within an OSPG, the non-local weight is obtained via a novel similarity measurement method; third, based on the consistency assumption, the non-local fitting deviations (NLFDs) by using OSPGs are uniformly constrained. Because of the uniform constraints, the restoration may be poor in regions where OSPGs are not reliable. Consequently, a weighting scheme is proposed to measure the OSPG reliability, leading to a novel adjusted non-local regression (ANLR). In addition, the integral image technique (IIT) is adopted to speed up the similar patches search process. To further boost the ANLR, a local directional smoothness (DS) prior is proposed as a good complement of the non-local feature. Finally, a fast split Bregman iteration algorithm is designed to solve the ANLR-DS minimization problem. Extensive experiments on two typical IR problems, that is, image deblurring and super resolution, demonstrate the superiority of the proposed method compared to many state-of-the-art IR methods. Chao Ren 0002, Xiaohai He, Truong Q. Nguyen |
IEEE Trans. Multim. | 3 |
| 2018 | Pyramid Structured Optical Flow Learning with Motion CuesabstractAfter the introduction of FlowNet and the large scale synthetic dataset Flying Chairs, we witnessed a rapid growth of deep learning based optical flow estimation algorithms. However, most of these algorithms rely on a very deep network to learn both large and small motions, making them less efficient. They also process each frame individually for the video dataset like MPI Sintel without using temporally correlated information across frames. This paper presents a pyramid structured network that estimates the optical flow from coarse to fine. We use a much shallower subnetwork at each pyramid level to predict an incremental flow, which contains relatively small motions, based on higher level's prediction. For video dataset, the network utilizes motion cues from previous frames' estimations for assistance. Evaluations show that the proposed network outperforms FlowNet on multiple benchmarks and has a slight edge on other similar pyramid structure networks. The shallow network design shrinks the parameter size by 88% comparing to FlowNet, allowing it to reach almost 100 frames per second prediction speed. Ji Dai, Truong Q. Nguyen |
ICIP | 3 |
| 2018 | Accurate and Efficient Video De-Fencing Using Convolutional Neural Networks and Temporal InformationabstractDe-fencing is to eliminate the captured fence on an image or a video, providing a clear view of the scene. It has been applied for many purposes including assisting photographers and improving the performance of computer vision algorithms such as object detection and recognition. However, the state-of-the-art de-fencing methods have limited performance caused by the difficulty of fence segmentation and also suffer from the motion of the camera or objects. To overcome these problems, we propose a novel method consisting of segmentation using convolutional neural networks and a fast/robust recovery algorithm. The segmentation algorithm using convolutional neural network achieves significant improvement in the accuracy of fence segmentation. The recovery algorithm using optical flow produces plausible de-fenced images and videos. The proposed method is experimented on both our diverse and complex dataset and publicly available datasets. The experimental results demonstrate that the proposed method achieves the state-of-the-art performance for both segmentation and content recovery. Byeongkeun Kang, Ji Dai, Truong Q. Nguyen |
ICME | 5 |
| 2018 | Image Denoising with Generalized Gaussian Mixture Model Patch PriorsabstractPatch priors have become an important component of image restoration. A powerful approach in this category of restoration algorithms is the popular expected patch log-likelihood (EPLL) algorithm. EPLL uses a Gaussian mixture model (GMM) prior learned on clean image patches as a way to regularize degraded patches. In this paper, we show that a generalized Gaussian mixture model (GGMM) captures the underlying distribution of patches better than a GMM. Even though GGMM is a powerful prior to combine with EPLL, the non-Gaussianity of its components presents major challenges to be applied to a computationally intensive process of image restoration. Specifically, each patch has to undergo a patch classification step and a shrinkage step. These two steps can be efficiently solved with a GMM prior but are computationally impractical when using a GGMM prior. In this paper, we provide approximations and computational recipes for fast evaluation of these two steps, so that EPLL can embed a GGMM prior on an image with more than tens of thousands of patches. Our main contribution is to analyze the accuracy of our approximations based on thorough theoretical analysis. Our evaluations indicate that the GGMM prior is consistently a better fit for modeling image patch distribution and performs better on average in image denoising task. Charles-Alban Deledalle, Shibin Parameswaran, Truong Q. Nguyen |
SIAM J. Imaging Sci. | 3 |
| 2018 | A unified framework for sparse non-negative least squares using multiplicative updates and the non-negative matrix factorization problem
Igor Fedorov, Alican Nalci, Ritwik Giri, Bhaskar D. Rao, Truong Q. Nguyen, Harinath Garudadri |
Signal Process. | 5 |
| 2018 | Patch Matching for Image Denoising Using Neighborhood-Based Collaborative FilteringabstractWe consider patch matching as a recommendation system problem and introduce a new patch-matching approach using nearest neighbor-based collaborative filtering (NN-CF). Our approach involves recommending similar patches to a query patch with the help of other similar patches in a noisy image or an external database. Using user-oriented and item-oriented formulations of NN-CF, we present two variations of CF-based patch-matching criterion. To demonstrate the superior matches found with our method, we apply the new patch-matching scheme to patch-based image denoising and evaluate its effect on the denoising performance. We test the methods on two data sets with varying background and image complexities and under different levels of noise. The proposed method not only improves robustness to patch matching but also provides a new formulation to seamlessly combine internal and external denoising. Shibin Parameswaran, Enming Luo, Truong Q. Nguyen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Camera-Aware Multi-Resolution Analysis for Raw Image Sensor Data Compressionabstractnalysis, or CAMRA. Specifically, by CAMRA we refer to modifications that we make to wavelet transform of CFA sampled images in order to achieve a very high degree of decorrelation at the finest scale wavelet coefficients; and a series of color processing steps applied to the coarse scale wavelet coefficients, aimed at limiting the propagation of lossy compression errors through the subsequent camera processing pipeline. We validated our theoretical analysis and the performance of the proposed compression schemes using the images of natural scenes captured in a raw format. The experimental results verify that our proposed methods improve coding efficiency relative to the standard and the state-of-the-art compression schemes for CFA sampled images. Yeejin Lee, Keigo Hirakawa, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2018 | Depth-Adaptive Deep Neural Network for Semantic SegmentationabstractIn this paper, we present the depth-adaptive deep neural network using a depth map for semantic segmentation. Typical deep neural networks receive inputs at the predetermined locations regardless of the distance from the camera. This fixed receptive field presents a challenge to generalize the features of objects at various distances in neural networks. Specifically, the predetermined receptive fields are too small at a short distance, and vice versa. To overcome this challenge, we develop a neural network that is able to adapt the receptive field not only for each layer but also for each neuron at the spatial location. To adjust the receptive field, we propose the depth-adaptive multiscale (DaM) convolution layer consisting of the adaptive perception neuron and the in-layer multiscale neuron. The adaptive perception neuron is to adjust the receptive field at each spatial location using the corresponding depth information. The in-layer multiscale neuron is to apply the different size of the receptive field at each feature space to learn features at multiple scales. The proposed DaM convolution is applied to two fully convolutional neural networks. We demonstrate the effectiveness of the proposed neural networks on the publicly available RGB-D dataset for semantic segmentation and the novel hand segmentation dataset for hand-object interaction. The experimental results show that the proposed method outperforms the state-of-the-art methods without any additional layers or preprocessing/postprocessing. Byeongkeun Kang, Yeejin Lee, Truong Q. Nguyen |
IEEE Trans. Multim. | 3 |
| 2017 | Multimodal sparse Bayesian dictionary learning applied to multimodal data classificationabstractIn this paper, we present a novel multimodal sparse dictionary learning algorithm based on a hierarchical sparse Bayesian framework. The framework allows for enforcing joint sparsity across dictionaries without restricting the actual entries to be equal. We show that the proposed method is able to learn dictionaries of higher quality than existing approaches. We validate our claims with extensive experiments on synthetic data as well as real-world data. Igor Fedorov, Bhaskar D. Rao, Truong Q. Nguyen |
ICASSP | 3 |
| 2017 | View synthesis with hierarchical clustering based occlusion fillingabstractThis paper presents a depth image based rendering algorithm for view synthesis task. We address the challenging occlusion filling problem with a hierarchical clustering approach. Depth distribution of neighboring pixels around each occlusion is explored and from which we determine the number of surrounding depth planes with agglomerative clustering. Pixels in the most distant plane are picked as candidates to restore that occlusion. The proposed algorithm is evaluated on Middlebury stereo dataset and Microsoft Research 3D video dataset. Results show that our method ranks among the best performers. Ji Dai, Truong Q. Nguyen |
ICIP | 2 |
| 2017 | Lossless compression of CFA sampled image using decorrelated Mallat wavelet packet decompositionabstractThis paper presents a rigorous analysis of wavelet transform on color filter array (CFA) sampled images. The presented analysis suggests that the wavelet coefficients of HL and LH subbands are highly correlated. Hence, we propose a novel lossless compression scheme for CFA sampled images using the decorrelated Mallat wavelet packet decomposition. We validated our theoretical analysis and the performance of the proposed compression scheme using images of natural scenes captured in a raw format. The experimental results verify that our proposed method improves coding efficiency relative to the standard and the state-of-the-art lossless compression schemes CFA sampled images. Yeejin Lee, Keigo Hirakawa, Truong Q. Nguyen |
ICIP | 3 |
| 2017 | Targeted video denoising for decompressed videosabstractThe paradigm of using clean patches from a targeted external database to design optimal denoising filters, called Targeted Image Denoising (TID), has been shown to outperform state-of-the-art denoising algorithms such as BM3D. In this paper, we introduce Targeted Video Denoising algorithm that extends the TID algorithm to denoise decompressed video without adding complexity. Our algorithm leverages the motion vectors generated during compression to establish temporal coherency between patches in consecutive frames. We test our algorithm on three decompressed video sequences with different foregrounds, backgrounds and movement patterns, and different noise level settings. Experimental results show that our approach is effective and performs better than original TID and the state-of-the-art video denoising algorithm. Shibin Parameswaran, Enming Luo, Truong Q. Nguyen |
ICIP | 3 |
| 2017 | Image noise estimation and removal considering the bayer pattern of noise varianceabstractTraditional image denoising methods are designed for Gaussian or Poisson noise, which are not suitable for realistic noise introduced in the complicated imaging pipeline. We observe that, due to the demosaicing process in imaging, the noise variance maps of captured JPEG images are characterized by Bayer patterns. In this paper, we propose a novel noise estimation and removal method based on the Bayer pattern of noise variance maps. There are two key contributions in the proposed method. First, to the best of our knowledge, we are the first to consider the Bayer patterns of noise variance maps in noise estimation and denoising. Second, we extend the state-of-the-art denoising method CBM3D to deal with realistic noise by integrating the estimated noise variance map and Bayer-pattern down-sampling into the denoising process. Experimental results show that the proposed method achieves the best noise estimation performance compared with two state-of-the-art methods. In addition, the denoising performance of CBM3D for realistic noise is significantly improved using the proposed approach and outperforms state-of-the-art blind denoising methods. Huanjing Yue, Jing-Yu Yang 0002, Truong Q. Nguyen, Chunping Hou |
ICIP | 4 |
| 2017 | Moving Object Detection With a Freely Moving Camera via Background Motion SubtractionabstractDetection of moving objects in a video captured by a freely moving camera is a challenging problem in computer vision. Most existing methods often assume that the background (BG) can be approximated by dominant single plane/multiple planes or impose significant geometric constraints on BG, or utilize a complex BG/foreground probabilistic model. Instead, we propose a computationally efficient algorithm that is able to detect moving objects accurately and robustly in a general 3D scene. This problem is formulated as a coarse-to-fine thresholding scheme on the particle trajectories in the video sequence. First, a coarse foreground (CFG) region is extracted by performing reduced singular value decomposition on multiple matrices that are built from bundles of particle trajectories. Next, the BG motion of pixels in the CFG region is reconstructed by a fast inpainting method. After subtracting the BG motion, the fine foreground is segmented out by an adaptive thresholding method that is capable of solving multiple-moving-objects scenarios. Finally, the detected foreground is further refined by the mean-shift segmentation method. Extensive simulations and a comparison with the state-of-the-art methods verify the effectiveness of the proposed method. Yuanyuan Wu 0001, Xiaohai He, Truong Q. Nguyen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Joint Defogging and DemosaickingabstractImage defogging is a technique used extensively for enhancing visual quality of images in bad weather conditions. Even though defogging algorithms have been well studied, defogging performance is degraded by demosaicking artifacts and sensor noise amplification in distant scenes. In order to improve the visual quality of restored images, we propose a novel approach to perform defogging and demosaicking simultaneously. We conclude that better defogging performance with fewer artifacts can be achieved when a defogging algorithm is combined with a demosaicking algorithm simultaneously. We also demonstrate that the proposed joint algorithm has the benefit of suppressing noise amplification in distant scenes. In addition, we validate our theoretical analysis and observations for both synthesized data sets with ground truth fog-free images and natural scene data sets captured in a raw format. Yeejin Lee, Keigo Hirakawa, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2017 | Single Image Super-Resolution via Adaptive High-Dimensional Non-Local Total Variation and Adaptive Geometric FeatureabstractSingle image super-resolution (SR) is very important in many computer vision systems. However, as a highly ill-posed problem, its performance mainly relies on the prior knowledge. Among these priors, the non-local total variation (NLTV) prior is very popular and has been thoroughly studied in recent years. Nevertheless, technical challenges remain. Because NLTV only exploits a fixed non-shifted target patch in the patch search process, a lack of similar patches is inevitable in some cases. Thus, the non-local similarity cannot be fully characterized, and the effectiveness of NLTV cannot be ensured. Based on the motivation that more accurate non-local similar patches can be found by using shifted target patches, a novel multishifted similar-patch search (MSPS) strategy is proposed. With this strategy, NLTV is extended as a newly proposed super-high-dimensional NLTV (SHNLTV) prior to fully exploit the underlying non-local similarity. However, as SHNLTV is very high-dimensional, applying it directly to SR is very difficult. To solve this problem, a novel statistics-based dimension reduction strategy is proposed and then applied to SHNLTV. Thus, SHNLTV becomes a more computationally effective prior that we call adaptive high-dimensional non-local total variation (AHNLTV). In AHNLTV, a novel joint weight strategy that fully exploits the potential of the MSPS-based non-local similarity is proposed. To further boost the performance of AHNLTV, the adaptive geometric duality (AGD) prior is also incorporated. Finally, an efficient split Bregman iteration-based algorithm is developed to solve the AHNLTV-AGD-driven minimization problem. Extensive experiments validate the proposed method achieves better results than many state-of-the-art SR methods in terms of both objective and subjective qualities. Chao Ren 0002, Xiaohai He, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2016 | Context Matters: Refining Object Detection in Video with Recurrent Neural Networks
Subarna Tripathi, Zachary C. Lipton, Serge J. Belongie, Truong Q. Nguyen |
BMVC | 4 |
| 2016 | Robust Bayesian method for simultaneous block sparse signal recovery with applications to face recognitionabstractIn this paper, we present a novel Bayesian approach to recover simultaneously block sparse signals in the presence of outliers. The key advantage of our proposed method is the ability to handle non-stationary outliers, i.e. outliers which have time varying support. We validate our approach with empirical results showing the superiority of the proposed method over competing approaches in synthetic data experiments as well as the multiple measurement face recognition problem. Igor Fedorov, Ritwik Giri, Bhaskar D. Rao, Truong Q. Nguyen |
ICIP | 4 |
| 2016 | Detecting temporally consistent objects in videos through object class label propagationabstractObject proposals for detecting moving or static video objects need to address issues such as speed, memory complexity and temporal consistency. We propose an efficient Video Object Proposal (VOP) generation method and show its efficacy in learning a better video object detector A deep-learning based video object detector learned using the proposed VOP achieves state-of-the-art detection performance on the Youtube-Objects dataset. We further propose a clustering of VOPs which can efficiently be used for detecting objects in video in a streaming fashion. As opposed to applying per-frame convolutional neural network (CNN) based object detection, our proposed method called Objects in Video Enabler thRough LAbel Propagation (OVERLAP) needs to classify only a small fraction of all candidate proposals in every video frame through streaming clustering of object proposals and class-label propagation. Source code for VOP clustering is available at https://github. com/subtri/streaming_VOP_clustering. Subarna Tripathi, Serge J. Belongie, Youngbae Hwang, Truong Q. Nguyen |
WACV | 4 |
| 2016 | Realistic surface geometry reconstruction using a hand-held RGB-D camera
Kyoung-Rok Lee, Truong Q. Nguyen |
Mach. Vis. Appl. | 2 |
| 2016 | Long-Range Motion Trajectories Extraction of Articulated Human Using Mesh EvolutionabstractThis letter presents a novel approach to extract reliable dense and long-range motion trajectories of articulated human in a video sequence. Compared with existing approaches that emphasize temporal consistency of each tracked point, we also consider the spatial structure of tracked points on the articulated human. We treat points as a set of vertices, and build a triangle mesh to join them in image space. The problem of extracting long-range motion trajectories is changed to the issue of consistency of mesh evolution over time. First, self-occlusion is detected by a novel mesh-based method and an adaptive motion estimation method is proposed to initialize mesh between successive frames. Furthermore, we propose an iterative algorithm to efficiently adjust vertices of mesh for a physically plausible deformation, which can meet the local rigidity of mesh and silhouette constraints. Finally, we compare the proposed method with the state-of-the-art methods on a set of challenging sequences. Evaluations demonstrate that our method achieves favorable performance in terms of both accuracy and integrity of extracted trajectories. Yuanyuan Wu 0001, Xiaohai He, Byeongkeun Kang, Haiying Song, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 5 |
| 2016 | A Framework for Depth Video Reconstruction From a Subset of Samples and Its ApplicationsabstractThe growth of depth sensing systems in this decade has facilitated a variety of applications in computer vision. Depending on the systematic configurations, both direct and indirect sensing techniques encounter image processing issues, such as hole filling and depth map super resolution. In this paper, a framework for depth video reconstruction from a subset of samples is proposed. By redefining classical dense depth estimation into two individual problems, sensing and synthesis, we propose a motion compensation-assisted sampling (MCAS) scheme and a spatio-temporal depth reconstruction (STDR) algorithm for reconstructing depth video sequences from a subset of samples. Using the 3-dimensional extensible dictionary, discrete wavelet transform (DWT), and applying alternating direction method of multiplier technique, the proposed STDR algorithm possesses scalability for temporal volume and efficiency for processing large scale depth data. Exploiting the temporal information and corresponding RGB images, the proposed MCAS achieves an efficient one-stage sampling scheme. Experimental results show that the proposed depth reconstruction framework outperforms the existing methods and is competitive compared with our previous work, which requires a pilot signal in the two-stage sampling scheme. Finally, to estimate missing reliable depth samples from varying input sources, we present an inference approach using geometrical and color similarities. Applications for depth video super resolution from uniform-grid subsampled data and dense disparity video estimation from a subset of reliable samples are presented. Lee-Kang Liu, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2016 | Adaptive Image Denoising by Mixture AdaptationabstractWe propose an adaptive learning procedure to learn patch-based image priors for image denoising. The new algorithm, called the expectation-maximization (EM) adaptation, takes a generic prior learned from a generic external database and adapts it to the noisy image to generate a specific prior. Different from existing methods that combine internal and external statistics in ad hoc ways, the proposed algorithm is rigorously derived from a Bayesian hyper-prior perspective. There are two contributions of this paper. First, we provide full derivation of the EM adaptation algorithm and demonstrate methods to improve the computational complexity. Second, in the absence of the latent clean image, we show how EM adaptation can be modified based on pre-filtering. The experimental results show that the proposed adaptation algorithm yields consistently better denoising results than the one without adaptation and is superior to several state-of-the-art algorithms. Enming Luo, Stanley H. Chan, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2016 | Single Image Super-Resolution Using Local Geometric Duality and Non-Local SimilarityabstractSuper-resolution (SR) from a single image plays an important role in many computer vision applications. It aims to estimate a high-resolution (HR) image from an input low- resolution (LR) image. To ensure a reliable and robust estimation of the HR image, we propose a novel single image SR method that exploits both the local geometric duality (GD) and the non-local similarity of images. The main principle is to formulate these two typically existing features of images as effective priors to constrain the super-resolved results. In consideration of this principle, the robust soft-decision interpolation method is generalized as an outstanding adaptive GD (AGD)-based local prior. To adaptively design weights for the AGD prior, a local non-smoothness detection method and a directional standard-deviation-based weights selection method are proposed. After that, the AGD prior is combined with a variational-framework-based non-local prior. Furthermore, the proposed algorithm is speeded up by a fast GD matrices construction method, which primarily relies on the selective pixel processing. The extensive experimental results verify the effectiveness of the proposed method compared with several state-of-the-art SR algorithms. Chao Ren 0002, Xiaohai He, Qizhi Teng, Yuanyuan Wu 0001, Truong Q. Nguyen |
IEEE Trans. Image Process. | 5 |
| 2016 | Analysis of Crosstalk in 3D Circularly Polarized LCDs Depending on the Vertical Viewing LocationabstractCrosstalk in circularly polarized (CP) liquid crystal display (LCD) with polarized glasses (passive 3D glasses) is mainly caused by two factors: 1) the polarizing system including wave retarders and 2) the vertical misalignment (VM) of light between the LC module and the patterned retarder. We show that the latter, which is highly dependent on the vertical viewing location, is a much more significant factor of crosstalk in CP LCD than the former. There are three contributions in this paper. Initially, a display model for CP LCD, which accurately characterizes VM, is proposed. A novel display calibration method for the VM characterization that only requires pictures of the screen taken at four viewing locations. In addition, we prove that the VM-based crosstalk cannot be efficiently reduced by either preprocessing the input images or optimizing the polarizing system. Furthermore, we derive the analytic solution for the viewing zone, where the entire screen does not have the VM-based crosstalk. Menglin Zeng, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2015 | Cubic Convolution Scaler Optimized for Local Property of Image DataabstractA scaler is one of the most important modules in various video applications, such as ultra-high definition TV and scalable video systems. A variety of scaling techniques have been used to increase the video quality when the resolution of the source image has to be up- and down-scaled. Some conventional schemes exploit the property of local block data. Others consider the edge information of the data to be scaled. In this paper, we formulate a scaling problem to minimize the information loss resulting from the resizing process. The loss is considered in both the spatial and the frequency domains, and then it is minimized to optimize the kernel of the scaler. The simulation results show that the proposed algorithm reduces the information loss more than conventional schemes. When compared with the conventional algorithms, the proposed method outperforms those with similar complexity. Jin-Kee Chae, Jae-Yung Lee, Min-Ho Lee, Jong-Ki Han, Truong Q. Nguyen, Woon-Young Yeo |
IEEE Trans. Image Process. | 5 |
| 2015 | Depth Reconstruction From Sparse Samples: Representation, Algorithm, and SamplingabstractThe rapid development of 3D technology and computer vision applications has motivated a thrust of methodologies for depth acquisition and estimation. However, existing hardware and software acquisition methods have limited performance due to poor depth precision, low resolution, and high computational cost. In this paper, we present a computationally efficient method to estimate dense depth maps from sparse measurements. There are three main contributions. First, we provide empirical evidence that depth maps can be encoded much more sparsely than natural images using common dictionaries, such as wavelets and contourlets. We also show that a combined wavelet-contourlet dictionary achieves better performance than using either dictionary alone. Second, we propose an alternating direction method of multipliers (ADMM) for depth map reconstruction. A multiscale warm start procedure is proposed to speed up the convergence. Third, we propose a two-stage randomized sampling scheme to optimally choose the sampling locations, thus maximizing the reconstruction performance for a given sampling budget. Experimental results show that the proposed method produces high-quality dense depth estimates, and is robust to noisy measurements. Applications to real data in stereo matching are demonstrated. Lee-Kang Liu, Stanley H. Chan, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2015 | Adaptive Image Denoising by Targeted DatabasesabstractWe propose a data-dependent denoising procedure to restore noisy images. Different from existing denoising algorithms which search for patches from either the noisy image or a generic database, the new algorithm finds patches from a database that contains relevant patches. We formulate the denoising problem as an optimal filter design problem and make two contributions. First, we determine the basis function of the denoising filter by solving a group sparsity minimization problem. The optimization formulation generalizes existing denoising algorithms and offers systematic analysis of the performance. Improvement methods are proposed to enhance the patch search process. Second, we determine the spectral coefficients of the denoising filter by considering a localized Bayesian prior. The localized prior leverages the similarity of the targeted database, alleviates the intensive Bayesian computation, and links the new method to the classical linear minimum mean squared error estimation. We demonstrate applications of the proposed method in a variety of scenarios, including text images, multiview images, and face images. Experimental results show the superiority of the new algorithm over existing methods. Enming Luo, Stanley H. Chan, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2015 | Vector Sparse Representation of Color Image Using Quaternion Matrix AnalysisabstractTraditional sparse image models treat color image pixel as a scalar, which represents color channels separately or concatenate color channels as a monochrome image. In this paper, we propose a vector sparse representation model for color images using quaternion matrix analysis. As a new tool for color image representation, its potential applications in several image-processing tasks are presented, including color image reconstruction, denoising, inpainting, and super-resolution. The proposed model represents the color image as a quaternion matrix, where a quaternion-based dictionary learning algorithm is presented using the K-quaternion singular value decomposition (QSVD) (generalized K-means clustering for QSVD) method. It conducts the sparse basis selection in quaternion space, which uniformly transforms the channel images to an orthogonal color space. In this new color space, it is significant that the inherent color structures can be completely preserved during vector reconstruction. Moreover, the proposed sparse model is more efficient comparing with the current sparse models for image restoration tasks due to lower redundancy between the atoms of different color channels. The experimental results demonstrate that the proposed sparse image model avoids the hue bias issue successfully and shows its potential as a general and powerful tool in color image analysis and processing domain. Yi Xu 0001, Licheng Yu, Hongteng Xu, Truong Q. Nguyen |
IEEE Trans. Image Process. | 5 |
| 2015 | Single Image Superresolution Based on Gradient Profile SharpnessabstractSingle image superresolution is a classic and active image processing problem, which aims to generate a high-resolution (HR) image from a low-resolution input image. Due to the severely under-determined nature of this problem, an effective image prior is necessary to make the problem solvable, and to improve the quality of generated images. In this paper, a novel image superresolution algorithm is proposed based on gradient profile sharpness (GPS). GPS is an edge sharpness metric, which is extracted from two gradient description models, i.e., a triangle model and a Gaussian mixture model for the description of different kinds of gradient profiles. Then, the transformation relationship of GPSs in different image resolutions is studied statistically, and the parameter of the relationship is estimated automatically. Based on the estimated GPS transformation relationship, two gradient profile transformation models are proposed for two profile description models, which can keep profile shape and profile gradient magnitude sum consistent during profile transformation. Finally, the target gradient field of HR image is generated from the transformed gradient profiles, which is added as the image prior in HR image reconstruction model. Extensive experiments are conducted to evaluate the proposed algorithm in subjective visual effect, objective quality, and computation time. The experimental results demonstrate that the proposed approach can generate superior HR images with better visual quality, lower reconstruction error, and acceptable computation efficiency as compared with state-of-the-art works. Yi Xu 0001, Xiaokang Yang 0001, Truong Q. Nguyen |
IEEE Trans. Image Process. | 4 |
| 2015 | Modeling, Prediction, and Reduction of 3D Crosstalk in Circular Polarized Stereoscopic LCDsabstractCrosstalk, which is the incomplete separation between the left and right views in 3D displays, induces ghosting and causes difficulty of the eyes to fuse the stereo image for depth perception. Circularly polarized (CP) liquid crystal display (LCD) is one of the main-stream consumer 3D displays with the prospering of 3D movies and gamings. The polarizing system including the patterned retarder is one of the major causes of crosstalk in CP LCD. The contributions of this paper are the modeling of the polarizing system of CP LCD, and a crosstalk reduction method that efficiently cancels crosstalk and preserves image contrast. For the modeling, the practical orientation of the polarized glasses (PG) is considered. In addition, this paper calculates the rotation of the light-propagation coordinate for the Stokes vector as light propagates from LCD to PG, and this calculation is missing in the previous works when applying Mueller calculus. The proposed crosstalk reduction method is formulated as a linear programming problem, which can be easily solved. In addition, we propose excluding the highly textured areas in the input images to further preserve image contrast in crosstalk reduction. Menglin Zeng, Alan E. Robinson, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2015 | Multi-Resolution Disparity Processing and Fusion for Large High-Resolution Stereo ImageabstractLarge panoramic views with high resolution have the advantage of a wide field of view over regular stereo views. However, the large size and high resolution impose difficulties on the stereo matching problem such as complexity and structure ambiguity, respectively. In this paper, effective multi-resolution disparity processing to resolve the difficulties is presented. We propose to adaptively determine the disparity search range based on the combined local structure from image intensity and initial disparity. The adaptive disparity range is able to propagate the smoothness property at low resolution to high resolution while preserving fine details. It reduces structure ambiguity as well as computational complexity. To reduce the disparity quantization error at the coarse level, we propose a reliable multiple fitting algorithm that is noticeably effective on the round surface. The spatial-multi-resolution total variation method is investigated to minimize inconsistency in space-scale dimension . The experimental results on the Middlebury datasets and real-world high-resolution images demonstrate that the proposed multi-resolution scheme produces high-quality and high-resolution disparity maps by fusing individual multi-scale disparity maps, while reducing complexity. Zucheul Lee, Truong Q. Nguyen |
IEEE Trans. Multim. | 2 |
| 2014 | Visual odometry for RGB-D cameras for dynamic scenesabstractIn this paper, we propose an accurate estimation of the camera motion in a dynamic environment from RGB-D videos. To better exclude the moving object portion of the scene from the stationary background, we use image segmentation. Next, dense pixel matching between the current and reference color images is performed to construct the 3D point cloud for dense motion estimation. At the end, we perform motion optimization, i.e., to find the combination of motion parameters that minimizes the remainder difference between the reference and the current image. We validate our proposed method across two benchmark sequences and show that our approach is more accurate than the existing solutions. We show that our method reduces the RMSE by 6.55% and 7.16% for stationary and dynamic scenes, respectively. Haleh Azartash, Kyoung-Rok Lee, Truong Q. Nguyen |
ICASSP | 3 |
| 2014 | Hierarchical depth processing with adaptive search range and fusionabstractIn this paper, we present an effective hierarchical depth processing and fusion for large stereo images. We propose the adaptive disparity search range based on the combined local structure from image and initial disparity. The adaptive search range can propagate the smoothness property at the coarse level to the fine level while preserving details and suppressing undesirable errors. The spatial-multiscale total variation method is investigated to enforce the spatial and scaling consistency of multi-scale depth estimates. The experimental results demonstrate that the proposed hierarchical scheme produces high quality and high resolution depth maps by fusing individual multi-scale depth maps, while reducing complexity. Zucheul Lee, Truong Q. Nguyen |
ICASSP | 2 |
| 2014 | Image denoising by targeted external databasesabstractClassical image denoising algorithms based on single noisy images and generic image databases will soon reach their performance limits. In this paper, we propose to denoise images using targeted external image databases. Formulating denoising as an optimal filter design problem, we utilize the targeted databases to (1) determine the basis functions of the optimal filter by means of group sparsity; (2) determine the spectral coefficients of the optimal filter by means of localized priors. For a variety of scenarios such as text images, multiview images, and face images, we demonstrate superior denoising results over existing algorithms. Enming Luo, Stanley H. Chan, Truong Q. Nguyen |
ICASSP | 3 |
| 2014 | Stereo image defoggingabstractThis paper presents a new approach to estimate fog-free images from stereo foggy images. We investigate a new way to estimate transmission by computing the scattering coefficient and depth information of a scene. However, most existing visibility restoration algorithms estimate transmission independently on scattering coefficient and object distance. In the proposed method, the natural color of a foggy image is recovered using depth information from a stereo image pair even though prior knowledge or multiple images taken at different times are not required. Furthermore, we explore a new way to measure the scattering coefficient by using a stereo image pair from an image processing perspective. Experimental results verify that the proposed method outperforms the conventional defogging methods. Yeejin Lee, Kristofor B. Gibson, Zucheul Lee, Truong Q. Nguyen |
ICIP | 4 |
| 2014 | An effective example-based learning method for denoising of medical images corrupted by heavy Gaussian noise and poisson noiseabstractDenoising is an essential application to improve image quality, especially in medical imaging. This paper introduces an example and patch-based learning method for reducing Gaussian noise and Poisson noise which often appear in medical imaging modalities using ionizing radiation. In the proposed method, denoising is performed by learning the regression model based on a set of the nearest neighbors of a given noisy patch, with the help of a given set of standard images. The method is evaluated and compared to several state-of-the-art denoising methods. The obtained results confirm its efficiency, especially for heavy noise. Dinh Hoan Trinh, Marie Luong, Françoise Dibos, Jean-Marie Rocchisani, Canh Duong Pham, Nguyen Linh-Trung, Truong Q. Nguyen |
ICIP | 7 |
| 2014 | Crosstalk modeling in circularly polarized stereoscopic LCDSabstractCrosstalk is the most critical artifact in stereoscopic 3D displays. Circularly polarized (CP) 3D LCDs with passive glasses has the issue of crosstalk especially when the viewer moves vertically away from the screen center. In this work, a new display model for CP LCD is proposed which considers both the intrinsic optical system including the polarizer and the wave retarder, and the vertical misalignment (VM) of light between the liquid crystal (LC) and the patterned retarder (PR) as the two major factors causing crosstalk in CP LCDs. The optical modeling for the display uses Mueller calculus and the characterization of VM is realized by the proposed user-calibration method. The display model is proved to be accurate by both measuring and subjective test in predicting the crosstalk resulted from different viewing scenarios. Additionally, we proposed considering image texture for the estimation of crosstalk perception. The proposed estimation is shown to be consistent with the viewer's observation. Menglin Zeng, Haleh Azartash, Truong Q. Nguyen |
ICIP | 3 |
| 2014 | Robust tracking and mapping with a handheld RGB-D cameraabstractIn this paper, we propose a robust method for camera tracking and surface mapping using a handheld RGB-D camera which is effective in challenging situations such as fast camera motion or geometrically featureless scenes. The main contributions are threefold. First, we introduce a robust orientation estimation based on quaternion method for initial sparse estimation. By using visual feature points detection and matching, no prior or small movement assumption is required to estimate a rigid transformation between frames. Second, a weighted ICP (Iterative Closest Point) method for better rate of convergence in optimization and accuracy in resulting trajectory is proposed. While the conventional ICP fails when there is no 3D features in the scene, our approach achieves robustness by emphasizing the influence of points that contain more geometric information of the scene. Finally, we show quantitative results on an RGB-D benchmark dataset. The experiments on an RGB-D trajectory benchmark dataset demonstrate that our method is able to track camera pose accurately. Kyoung-Rok Lee, Truong Q. Nguyen |
WACV | 2 |
| 2014 | Improving streaming video segmentation with early and mid-level visual processingabstractDespite recent advances in video segmentation, many opportunities remain to improve it using a variety of low and mid-level visual cues. We propose improvements to the leading streaming graph-based hierarchical video segmentation (streamGBH) method based on early and mid level visual processing. The extensive experimental analysis of our approach validates the improvement of hierarchical supervoxel representation by incorporating motion and color with effective filtering. We also pose and illuminate some open questions towards intermediate level video analysis as further extension to streamGBH. We exploit the supervoxels as an initialization towards estimation of dominant affine motion regions, followed by merging of such motion regions in order to hierarchically segment a video in a novel motion-segmentation framework which aims at subsequent applications such as foreground recognition. Subarna Tripathi, Youngbae Hwang, Serge J. Belongie, Truong Q. Nguyen |
WACV | 4 |
| 2014 | HEASK: Robust homography estimation based on appearance similarity and keypoint correspondences
Yi Xu 0001, Xiaokang Yang 0001, Truong Q. Nguyen |
Pattern Recognit. | 4 |
| 2014 | Separation of Weak Reflection from a Single Superimposed ImageabstractIt is an inherently ill-posed problem to separate a single superimposed image into a reflection image and a transmission image. In this letter, a novel algorithm is proposed based on the prior knowledge that edges of weak reflection are always smoother than most edges of observed objects. To filter out the edges of weak reflection, an MRF-EM (Markov Random Field and Expectation Maximization) framework is proposed. In the MRF model, a data energy function is established based on the edge smoothness metric GPS (Gradient Profile Sharpness), and a spatial smoothness energy function is formulated using a weighted Potts model. Moreover, the parameters in the data energy function are updated using the EM algorithm. Experimental results demonstrate that the proposed algorithm can produce superior separation results with less residuals and color distortions compared to state-of-the-art methods. Yi Xu 0001, Xiaokang Yang 0001, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 4 |
| 2014 | Depth-Assisted Frame Rate Up-Conversion for Stereoscopic VideoabstractIn this letter, we propose a depth-assisted frame rate up-conversion (DA-FRUC) scheme which finds applications in 3D video processing. By considering the depth cue in video plus depth representation, we categorize the blocks of the interpolated frame as depth-continuous and depth-discontinuous groups. The motion vector (MV) outliers of the depth continuous blocks are then detected and corrected by layer-constrained MV refinement method. Moreover, a depth-based adaptive interpolation and block segmentation method is proposed to deal with the disocclusion and occlusion at the boundary of the foreground area. Experimental results show that the proposed method achieves better subjective and objective performance for the interpolated color frames. Additionally, the interpolated depth frames obtained by the proposed method is more accurate, which benefit the video quality in view synthesis. Jiande Sun 0001, Yeejin Lee, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 5 |
| 2014 | Multiview Video Plus Depth Coding With Depth-Based Prediction ModeabstractIn this paper, we propose a novel multiview video plus depth (MVD) codec that uses a depth-based prediction mode (DBPM) both for texture videos and the depth maps. We show that the proposed codec can provide up to 9.2%, 9.9%, and 6.7% bitrate savings over H.264/MVC (MVC) for coding MVD data, depth maps, and multiview videos, respectively. In addition, we present subjective test results that compare the perceptual video quality of stereo videos coded with the DBPM-enabled codec and MVC. We also show that a typical Lagrangian rate-distortion optimization is effective to successfully choose between the prediction modes available to the proposed encoder. We provide a complexity analysis for the proposed codec as well and show that the encoder complexity is comparable to MVC, while the decoder complexity is about twice of MVC. Finally, we analyze the effects of different encoder control parameters on the virtual view synthesis quality. Results show that video QP has the most influence on the synthesis quality, and there is a tradeoff between the depth map QP and resolution. Lower depth map resolutions yield more bitrate savings, whereas for similar bitrates higher resolution and higher QP depth maps achieve better view synthesis performance. Can Bal, Truong Q. Nguyen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Comparing Perceived Quality and Fatigue for Two Methods of Mixed Resolution Stereoscopic CodingabstractMixed resolution stereoscopic coding exploits the perceptual phenomenon of binocular suppression by encoding each view of a stereo pair at a different resolution. Many investigations have shown that a large reduction in bandwidth can be gained for a relatively small compromise in image quality by this asymmetric representation. However, the viability of such an encoding method depends on the subjective response to viewing such videos at length. While methods for mixed resolution coding have been developed on the presumption of a certain visual fatigue response, none have actually examined it. In this paper, we address this shortcoming in three experiments comparing two methods of binocular suppression processing. The first two experiments reveal subjects' preferences in terms of overall quality between the two methods for short exposures, and the third experiment examines the fatigue resulting from 10-min exposures to mixed resolution encoded videos. Ankit K. Jain, Alan E. Robinson, Truong Q. Nguyen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | A Frame-Level Rate Control Scheme Based on Texture and Nontexture Rate Models for High Efficiency Video CodingabstractIn this paper, a frame-level rate control scheme is proposed based on texture and nontexture rate models for High Efficiency Video Coding (HEVC). Due to more complicated coding structures and the adoption of new coding tools, the statistical characteristics of transform residues are significantly different depending on the depth levels of coding units (CUs) from which the residues are obtained. A new texture rate model is constructed for the transform residues, which are categorized into three types of CUs: low-, medium- and high-textured CUs. One single Laplacian probability PDF model is used for each residue category to derive a rate-quantization model. Based on the Laplacian PDF, a simplified rate model for texture bits is derived using entropy. In addition, an analytic rate model for nontexture bits is proposed, which also takes into account the different characteristics of nontexture bits occurring in various depths of CUs in HEVC. The nontexture bitrates are modeled based on the linear relation between the total nontexture data and the dominant nontexture data in each CU category. Based on the proposed rate models for the texture and nontexture bits, accurate rate control can be achieved owing to more precise rate estimation. The experimental results show that the proposed rate control scheme achieves the average PSNR with 0.44 dB higher and the average PSNR standard deviation of 0.32 point lower with the buffer status levels maintained very close to target buffer levels, compared to the conventional methods. Finally, the proposed rate control scheme remarkably outperforms the conventional schemes especially for the sequences of complex texture and large motion. Bumshik Lee, Munchurl Kim, Truong Q. Nguyen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | An Analysis and Method for Contrast Enhancement Turbulence MitigationabstractA common problem for imaging in the atmosphere is fog and atmospheric turbulence. Over the years, many researchers have provided insight into the physics of either the fog or turbulence but not both. Most recently, researchers have proposed methods to remove fog in images fast enough for real-time processing. Additionally, methods have been proposed by other researchers that address the atmospheric turbulence problem. In this paper, we provide an analysis that incorporates both physics models: 1) fog and 2) turbulence. We observe how contrast enhancements (fog removal) can affect image alignment and image averaging. We present in this paper, a new joint contrast enhancement and turbulence mitigation (CETM) method that utilizes estimations from the contrast enhancement algorithm to improve the turbulence removal algorithm. We provide a new turbulent mitigation object metric that measures temporal consistency. Finally, we design the CETM to be efficient such that it can operate in fractions of a second for near real-time applications. Kristofor B. Gibson, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2014 | Discriminability Limits in Spatio-Temporal Stereo Block MatchingabstractDisparity estimation is a fundamental task in stereo imaging and is a well-studied problem. Recently, methods have been adapted to the video domain where motion is used as a matching criterion to help disambiguate spatially similar candidates. In this paper, we analyze the validity of the underlying assumptions of spatio-temporal disparity estimation, and determine the extent to which motion aids the matching process. By analyzing the error signal for spatio-temporal block matching under the sum of squared differences criterion and treating motion as a stochastic process, we determine the probability of a false match as a function of image features, motion distribution, image noise, and number of frames in the spatio-temporal patch. This performance quantification provides insight into when spatio-temporal matching is most beneficial in terms of the scene and motion, and can be used as a guide to select parameters for stereo matching algorithms. We validate our results through simulation and experiments on stereo video. Ankit K. Jain, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2014 | Novel Example-Based Method for Super-Resolution and Denoising of Medical ImagesabstractIn this paper, we propose a novel example-based method for denoising and super-resolution of medical images. The objective is to estimate a high-resolution image from a single noisy low-resolution image, with the help of a given database of high and low-resolution image patch pairs. Denoising and super-resolution in this paper is performed on each image patch. For each given input low-resolution patch, its high-resolution version is estimated based on finding a nonnegative sparse linear representation of the input patch over the low-resolution patches from the database, where the coefficients of the representation strongly depend on the similarity between the input patch and the sample patches in the database. The problem of finding the nonnegative sparse linear representation is modeled as a nonnegative quadratic programming problem. The proposed method is especially useful for the case of noise-corrupted and low-resolution image. Experimental results show that the proposed method outperforms other state-of-the-art super-resolution methods while effectively removing noise. Dinh Hoan Trinh, Marie Luong, Françoise Dibos, Jean-Marie Rocchisani, Canh Duong Pham, Truong Q. Nguyen |
IEEE Trans. Image Process. | 6 |
| 2014 | Multi-Array Camera Disparity EnhancementabstractMulti-array camera systems have greater potential for 3-D depth-based application development compared with stereo camera systems. However, there are very few research results on multi-array-based disparity enhancement, extending standard stereo matchings to multi-array systems. In this paper, we propose to alternately use local and global fusion of multi-array disparities to maximize the disparity enhancement in array camera systems. We propose a new cascade regularization based approach, which can restore diagonal structures better than conventional techniques. The detailed analysis and experimental results verify that the cascade approach better regularizes the diagonal variations and in turn yields better image enhancement. We adapt total variation for regularization to the multi-array camera systems in order to globally combine multiple disparity estimates. A local multiple cross-filling algorithm is proposed to achieve cross consistency between array disparity estimates by effectively filling the mismatches. Experimental results show that the proposed multi-array disparity enhancement algorithm can improve the accuracy of the initial array disparity estimates up to 65% while alleviating memory limitation. Zucheul Lee, Truong Q. Nguyen |
IEEE Trans. Multim. | 2 |
| 2013 | Subframe video synchronization by matching trajectoriesabstractWe propose a novel approach to align unsynchronized video sequences of the same dynamic scene that can be subframe accurate and is applicable for different frame rate problem. The proposed approach relies on matching motion trajectories and it is assumed that the object moves on a planar surface. By exploring the invariants of planar trajectories under projective transformation, the cross ratio as an invariant feature is computed for each point along the trajectories and the similarity between invariant features of different trajectories is measured with a distance that takes into account the statistical properties of the cross ratio. Then the smooth high frame rate trajectory is synthesized for searching subframe temporal displacement under our alignment framework. The experimental results with synthetic and real-world sequences show that our approach achieves fairly accuracy and efficiency in subframe temporal alignment of the multiple unsynchronized video sequences. Yuanyuan Wu 0001, Xiaohai He, Truong Q. Nguyen |
ICASSP | 3 |
| 2013 | Crosstalk modeling, analysis, simulation and cancellation in passive-type stereoscopic LCD displaysabstractPassive-type stereoscopic LCD display (or polarized 3D TV) generates 3D effect using quarter-wave patterned retarders and polarized glasses. However, passive-type stereo-scopic displays are susceptible to an undesirable effect called “crosstalk”, which degrades the 3D effect and causes increased eye strain and nausea. In this paper, crosstalk in passive 3D LCD is modeled, analyzed and simulated by implementing extended Jones matrix method. We present how the display intensity changes with viewing angles and simulate stereo images with crosstalk. Furthermore, a method of crosstalk cancellation based on linear programming in YCbCr domain is proposed. Results of our proposed method show that zero crosstalk can be achieved in passive 3D displays with less contrast loss compared with other methods. Menglin Zeng, Truong Q. Nguyen |
ICASSP | 2 |
| 2013 | Depth-based prediction mode for 3D video codingabstractWith emerging multiview displays there is an interest in a new 3D video (3DV) representation and compression standard, which contains depth information, so that virtual views can be synthesized at the decoder from a sparse set of anchor views. With such a representation the depth map can also be utilized to increase video coding efficiency. In this work, we propose a novel depth-based prediction mode that increases inter-view prediction efficiency. The proposed mode is designed such that the syntax is simple and efficient to signal, and it is compatible with existing H.264/MVC modes. This design allows the encoder to choose between the proposed and existing prediction modes by a simple rate-distortion optimization decision. We show that depending on the depth map quality, the proposed mode can provide up to 5.5% additional bitrate savings over coding the texture video and the depth map separately. We also show that when the depth map is accurate, the proposed codec can achieve gains even over coding the texture video only, using H.264/MVC. Can Bal, Truong Q. Nguyen |
ICIP | 2 |
| 2013 | Example based depth from fogabstractThe presence of fog in an image reduces contrast which can be considered a nuisance in imaging applications, however, we consider this useful information for image enhancement and scene understanding. In this paper, we present a new method for estimating depth from fog in a single image and single image fog removal. We use an example based approach that is trained from data with known fog and depth. A data driven method and physics based model are used to develop the example based learning framework for single image fog removal. In addition, we account for various colors of fog by using a linear transformation of the RGB colorspace. This approach has the flexibility to learn from various scenes and relaxes the common constraint of fixed camera position. We present depth estimations and fog removal from a single image with good results. Kristofor B. Gibson, Serge J. Belongie, Truong Q. Nguyen |
ICIP | 3 |
| 2013 | Fast single image fog removal using the adaptive Wiener filterabstractWe present in this paper a fast single image defogging method that uses a novel approach to refining the estimate of amount of fog in an image with the Locally Adaptive Wiener Filter. We provide a solution for estimating noise parameters for the filter when the observation and noise are correlated by decorrelating with a naively estimated defogged image. We demonstrate our method is 50 to 100 times faster than existing fast single image defogging methods and that our proposed method subjectively performs as well as the Spectral Matting smoothed Dark Channel Prior method. Kristofor B. Gibson, Truong Q. Nguyen |
ICIP | 2 |
| 2013 | Video super-resolution for mixed resolution stereoabstractMixed resolution stereoscopic coding is a compression method that reduces the bandwidth of stereo video by transmitting one of the two views at a lower resolution. While the visual system weights the sharper view more heavily resulting in a nearly fully sharp 3D percept, certain applications require, or would benefit from, having both views at full resolution. In this paper, we present a super-resolution method that recovers the lost frequency information in mixed resolution stereo video by exploiting stereo correspondence while emphasizing spatial, temporal, and spectral consistency. Our method is distinguished by operating on video modeled as a 3D Markov network, whereas previous approaches focused solely on images. Ankit K. Jain, Truong Q. Nguyen |
ICIP | 2 |
| 2013 | Adaptive non-local means for multiview image denoising: Searching for the right patches via a statistical approachabstractWe present an adaptive non-local means (NLM) denoising method for a sequence of images captured by a multiview imaging system, where direct extensions of existing single image NLM methods are incapable of producing good results. Our proposed method consists of three major components: (1) a robust joint-view distance metric to measure the similarity of patches; (2) an adaptive procedure derived from statistical properties of the estimates to determine the optimal number of patches to be used; (3) a new NLM algorithm to denoise using only a set of similar patches. Experimental results show that the proposed method is robust to disparity estimation error, out-performs existing algorithms in multiview settings, and performs competitively in video settings. Enming Luo, Stanley H. Chan, Shengjun Pan, Truong Q. Nguyen |
ICIP | 4 |
| 2013 | A No-Reference Perceptual Based Contrast Enhancement Metric for Ocean Scenes in FogabstractIn this paper, we develop a perceptually based contrast enhancement metric as a means to solve the problem of autonomously enhancing images degraded by fog that are perceptually pleasing to humans. A learning based approach is considered to develop the contrast enhancement using human observations and low-level contrast enhancement metrics based on the human vision system. In addition, we provide new low-level metrics based on the physics of the scene to improve the performance of existing contrast enhancement metrics. This paper shows that a contrast enhancement metric can be designed to mimic human preference. Kristofor B. Gibson, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2013 | Local Disparity Estimation With Three-Moded Cross Census and Advanced Support WeightabstractThe classical local disparity methods use simple and efficient structure to reduce the computation complexity. To increase the accuracy of the disparity map, new local methods utilize additional processing steps such as iteration, segmentation, calibration and propagation, similar to global methods. In this paper, we present an efficient one-pass local method with no iteration. The proposed method is also extended to video disparity estimation by using motion information as well as imposing spatial temporal consistency. In local method, the accuracy of stereo matching depends on precise similarity measure and proper support window. For the accuracy of similarity measure, we propose a novel three-moded cross census transform with a noise buffer, which increases the robustness to image noise in flat areas. The proposed similarity measure can be used in the same form in both stereo images and videos. We further improve the reliability of the aggregation by adopting the advanced support weight and incorporating motion flow to achieve better depth map near moving edges in video scene. The experimental results show that the proposed method is the best performing local method on the Middlebury stereo benchmark test and outperforms the other state-of-the-art methods on video disparity evaluation. Zucheul Lee, Jason Juang, Truong Q. Nguyen |
IEEE Trans. Multim. | 3 |
| 2012 | A perceptual based contrast enhancement metric using AdaBoostabstractThis paper is an attempt to explore a human element not easily solved in the image processing communities. The problem statement is vague but important to address. What is a good image? More specifically, if a low contrast image is presented, at what level of enhancement is good enough for a human observer? This of course depends on diverse elements, e.g., personal preference, emotional state, physical impairments, purpose (object recognition), etc. In this paper we present a new contrast enhancement metric (CEM) that is trained using several simple contrast measures and mean opinion scores obtained from human observations. Our goal is to train the algorithm to mimic a human when selecting an image with the best contrast between two images. For example, the algorithm will accept two images of the same scene with differing (unknown) contrast and will choose which of the two images is ‘better’ according to what a human believes is ‘better’. Kristofor B. Gibson, Truong Q. Nguyen |
ISCAS | 2 |
| 2012 | An Investigation of Dehazing Effects on Image and Video CodingabstractThis paper makes an investigation of the dehazing effects on image and video coding for surveillance systems. The goal is to achieve good dehazed images and videos at the receiver while sustaining low bitrates (using compression) in the transmission pipeline. At first, this paper proposes a novel method for single-image dehazing, which is used for the investigation. It operates at a faster speed than current methods and can avoid halo effects by using the median operation. We then consider the dehazing effects in compression by investigating the coding artifacts and motion estimation in cases of applying any dehazing method before or after compression. We conclude that better dehazing performance with fewer artifacts and better coding efficiency is achieved when the dehazing is applied before compression. Simulations for Joint Photographers Expert Group images in addition to subjective and objective tests with H.264 compressed sequences validate our conclusion. Kristofor B. Gibson, Dung Trung Vo, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2012 | An Online Learning Approach to Occlusion Boundary DetectionabstractWe propose a novel online learning-based framework for occlusion boundary detection in video sequences. This approach does not require any prior training and instead "learns" occlusion boundaries by updating a set of weights for the online learning Hedge algorithm at each frame instance. Whereas previous training-based methods perform well only on data similar to the trained examples, the proposed method is well suited for any video sequence. We demonstrate the performance of the proposed detector both for the CMU data set, which includes hand-labeled occlusion boundaries, and for a novel video sequence. In addition to occlusion boundary detection, the proposed algorithm is capable of classifying occlusion boundaries by angle and by whether the occluding object is covering or uncovering the background. Natan Jacobson, Yoav Freund, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2012 | Scale-Aware Saliency for Application to Frame Rate UpconversionabstractOur understanding of human visual perception has been paramount in the development of tools for digital video processing. For this reason, saliency detection, i.e., the determination of visual importance in a scene, has come to the forefront in recent literature. In the proposed work, a new method for scale-aware saliency detection is introduced. Scale determination is afforded through a scale-space model utilizing color and texture cues. Scale information is fed back to a discriminant saliency engine by automatically tuning center-surround parameters through a soft weighting. Excellent results are demonstrated for the proposed method through its performance against a database of measured human fixations. Further evidence of the proposed algorithm's performance is demonstrated through an application to frame rate upconversion. The ability of the algorithm to detect salient objects at multiple scales allows for class-leading performance both objectively, in terms of peak signal-to-noise ratio/structural similarity index, and subjectively. Finally, the need for operator tuning of saliency parameters is dramatically reduced by the inclusion of scale information. The proposed method is well suited for any application requiring automatic saliency determination for images or video. Natan Jacobson, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2012 | Generalized Block-Lifting Factorization of $M$-Channel Biorthogonal Filter Banks for Lossy-to-Lossless Image CodingabstractGeneralized block-lifting factorization of M-channel (M > 2) biorthogonal filter banks (BOFBs) for lossy-to-lossless image coding is presented in this paper. Since the proposed block-lifting structure is more general than the conventional lifting factorizations and does NOT require many restrictions such as paraunitary, number of channels, and McMillan degree in each building block unlike the conventional lifting factorizations, its coding gain is higher than that of the previous methods. Several proposed BOFBs are designed and applied to image coding. Comparing the results with conventional lossy-to-lossless image coding structures, including the 5/3- and 9/7-tap discrete wavelet transforms in JPEG 2000 and a 4 × 8 hierarchical lapped biorthogonal transform in JPEG XR, the proposed BOFBs achieve better result in both objective measure and perceptual visual quality for the images with a lot of high-frequency components. Taizo Suzuki, Masaaki Ikehara, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2011 | Stochastic Model on the Post-Fabrication Error for a Bragg Reflectors Based Photonic Allpass FilterabstractThis paper presents an analysis of the post fabrication performance for a Bragg-mirrors-based photonic allpass filter. The development subsequently leads to a first definition of the stochastic measures in a realistic system implementation. A Digital Signal Processing (DSP) approach to examining photonic integrated circuitry is explored in depth. The physical origins of error that manifests in an unbalanced phase error is mathematically described in detail. We describe the details of error analysis and demonstrate a stochastic measure for the expected behavior of first-order allpass filter as well as high complexity systems. Our findings show that the frequency performance under error can be characterized by a convolution of the ideal response with a probability function based kernel. Andrew Grieco, Boris Slutsky, Bhaskar D. Rao, Yeshaiahu Fainman, Truong Q. Nguyen |
GLOBECOM | 6 |
| 2011 | Detection and removal of binocular luster in compressed 3D imagesabstractBinocular luster is an extremely salient effect seen in 3D when an object in each stereo image exhibits a different contrast polarity relative to the background. The object appears to shimmer, a phenomenon seen in nature and on 3D displays, which the Human Visual System rapidly detects. This binocular luster is also induced by compression of stereo imagery where corresponding blocks are quantized to different values. In this paper, we discuss the psychovisual background of binocular luster, introduce the "shine" artifact induced by compression, and present an algorithm for detection and removal of shine from JPEG compressed stereo images without introducing additional bitrate. Can Bal, Ankit K. Jain, Truong Q. Nguyen |
ICASSP | 3 |
| 2011 | An augmented Lagrangian method for video restorationabstractThis paper presents a fast algorithm for restoring video sequences. The proposed algorithm, as opposed to existing methods, does not consider video restoration as a sequence of image restoration problems. Rather, it treats a video sequence as a space-time volume and poses a space-time total variation regularization to enhance the smoothness of the solution. The optimization problem is solved by transforming the original unconstrained minimization problem to an equivalent constrained minimization problem. An augmented Lagrangian method is used to handle the constraints, and an alternating direction method (ADM) is used to iteratively find solutions of the subproblems. The proposed algorithm has a wide range of applications, including video deblurring and denoising, disparity map refinement, and reducing hot-air turbulence effects. Stanley H. Chan, Ramsin Khoshabeh, Kristofor B. Gibson, Philip E. Gill, Truong Q. Nguyen |
ICASSP | 5 |
| 2011 | On the effectiveness of the Dark Channel Prior for single image dehazing by approximating with minimum volume ellipsoidsabstractThere is an increasing number of methods for removing haze and fog from a single image. One of such methods is Dark Channel Prior (DCP). The goal of this paper is to develop a mathematical explanation on why DCP works well by using principal component analysis, and minimum volume ellipsoid approximations. Kristofor B. Gibson, Truong Q. Nguyen |
ICASSP | 2 |
| 2011 | Occlusion boundary detection using an online learning frameworkabstractIn this work, a novel occlusion detection algorithm using online learning is proposed for video applications. Each frame of a video is considered as a time-step for which pixels are classified as being either occluded or non-occluded. The Hedge algorithm is employed to determine weights for a set of experts, each of which is tuned to detect a specific type of occlusion boundary. In contrast to previous training-based methods, the proposed algorithm does not require any training, and has a runtime linear with respect to the number of experts considered. Detection performance is excellent on novel video sequences for which training data does not exist. In addition, the proposed algorithm is easily extended to provide classification results supplementary to detection. We demonstrate results on a series of challenging video sequences including a dataset of hand-labelled occlusion boundaries. Natan Jacobson, Yoav Freund, Truong Q. Nguyen |
ICASSP | 3 |
| 2011 | Video processing with scale-aware saliency: Application to Frame Rate Up-ConversionabstractA new method for scale-aware saliency detection is introduced in this work. Scale determination is realized through a fast scale-space algorithm using color and texture. Scale information is fed back to a Discriminant Saliency engine by automatically tuning center-surround parameters. Excellent results are demonstrated for predicted fixations using a public database of measured human fixations. Further evidence of the proposed algorithm's performance is exhibited through an application to Frame Rate Up-Conversion (FRUC). The ability of the algorithm to detect salient objects at multiple scales allows for class-leading performance both objectively in terms of PSNR/SSIM as well as subjectively. Finally, the need for operator tuning of saliency parameters is dramatically reduced by the inclusion of scale information. The proposed method is well-suited for any application requiring automatic saliency determination for images or video. Natan Jacobson, Truong Q. Nguyen |
ICASSP | 2 |
| 2011 | Efficient stereo-to-multiview synthesisabstractFrom a rectified stereo image pair, the task of view synthesis is to generate images from any viewpoint along the baseline. The main difficulty of the problem is how to fill occluded regions. In this pa per, we present a new method for view synthesis that is both fast and accurate. Occlusions are filled using color and disparity information to produce consistent pixel estimates. Results are comparable to current state-of-the-art methods in terms of objective measures while computation time is drastically reduced. This work has applications in free-viewpoint television, angular scalability for 3D video coding/decoding, and stereo-to-multiview conversion. Ankit K. Jain, Lam C. Tran, Ramsin Khoshabeh, Truong Q. Nguyen |
ICASSP | 4 |
| 2011 | Spatio-temporal consistency in video disparity estimationabstractWe present a novel stereo video disparity estimation method. The proposed method is a two-stage algorithm. During the first stage, initial disparity maps are computed in a frame by-frame basis. In the second stage, the initial estimates are treated as a space-time volume. By setting up an l1-normed minimization problem with a novel three-dimensional total variation regularization, spatial smoothness and temporal consistency are handled simultaneously. Due to our unique formulation, any existing image disparity estimation technique may utilize our method as a post-processing step to refine noisy estimates or to be extended to videos. The proposed method shows superior speed, accuracy, and consistency compared to state-of-the-art algorithms. Ramsin Khoshabeh, Stanley H. Chan, Truong Q. Nguyen |
ICASSP | 3 |
| 2011 | Dual domain method for single image dehazing and enhancingabstractIn this paper, we propose a novel method for improving the visibility of an image (with fog or haze), as well as the image's details. The proposed method adjusts the global contrast in the spatial domain for dehazing fog or haze spread over images and then enhances the local contrast in the transform domain for reviving the details of images. Compared to the previous methods performing in the spatial domain only, the proposed method improves the visibility and quality of images significantly. Dongin Shin, Kristofor B. Gibson, Wonha Kim, Truong Q. Nguyen |
ICASSP | 4 |
| 2011 | Spatially consistent view synthesis with coordinate alignmentabstractIn this paper, we propose a novel method that uses coordinate alignment and background pixel extraction to synthesize highly accurate and spatially consistent intermediate views from a pair of stereo images and disparity maps. In contrast to the traditional depth image-based rendering (DIBR) method, where useful background pixels are discarded in the warping process, the proposed method extracts these background pixels and uses them as candidates for an exemplar-based image in-painting technique (EBIIT) to synthesize realistic content in disocclusion regions. Our second contribution is a coordinate alignment algorithm that aligns disocclusion regions in each view together and simultaneously synthesizes disocclusion regions to enhance spatial consistency across all virtual views. The proposed method compares favorably in quantitative measures to those obtained by existing techniques and has superior potential for stereo-to-multiview conversion. Lam C. Tran, Ramsin Khoshabeh, Ankit K. Jain, Christopher Joseph Pal, Truong Q. Nguyen |
ICASSP | 5 |
| 2011 | Localized filtering for artifact removal in compressed imagesabstractThe paper proposes a novel method for coding artifact reduction in compressed images. For removing blocking artifacts, a localized DCT-based filter with condition on the similarity between surrounding blocks is considered. To reduce ringing, a localized fuzzy filter is utilized to avoid the blurry effect of linear filter and painting-like effect of conventional fuzzy filter. To enhance chroma components and reduce the color bleeding, the localized filter for luma component are implemented for the chroma components. Simulations on a wide range of compressed images are performed to verify the effectiveness of the algorithm. Dung Trung Vo, Truong Q. Nguyen |
ICASSP | 2 |
| 2011 | Design and analysis of a narrowband filter for optical platformabstractThis paper presents an approach to designing narrowband digital filters that are realizable using optical allpass building blocks. We describe a top-down design method by explicitly examining the derivation of an Infinite Impulse Response (IIR) architecture. Our result demonstrates a design that can achieve a 0.0025π passband edge while providing 60dB stopband attenuation. The design is aimed to reduce filter pole magnitudes, providing tolerance for waveguide losses and fabrication errors. The narrowband filter is based on the foundation of latticed allpass sections, which makes it naturally realizable using basic photonic components. Furthermore, analysis is performed on delay length variations that can result from the fabrication process. Andrew Grieco, Boris Slutsky, Bhaskar D. Rao, Yeshaiahu Fainman, Truong Q. Nguyen |
ICASSP | 6 |
| 2011 | Motion-decision based spatiotemporal saliency for video sequencesabstractAn adaptive spatiotemporal saliency algorithm for video attention detection using motion vector decision is proposed, motivated by the importance of motion information in video sequences for human visual system. This novel system can detect the saliency regions quickly by using only part of the classic saliency features in each iteration. Motion vectors calculated by block matching and optical flow are used to determine the decision condition. When significant motion contrast occurs (decision condition is satisfied), the saliency area is detected by motion and intensity features. Otherwise, when motion contrast is low, color and orientation features are added to form a more detailed saliency map. Experimental results show that the proposed algorithm can detect salient objects and actions in video sequences robustly and efficiently. Natan Jacobson, Hong Pan 0001, Truong Q. Nguyen |
ICASSP | 4 |
| 2011 | Combining generic and class-specific codebooks for object categorization and detectionabstractCombining advantages of shape and appearance features, we propose a novel model that integrates these two complementary features into a common framework for object categorization and detection. In particular, generic shape features are applied as a pre-filter that produces initial detection hypotheses following a weak spatial model, then the learnt class-specific discriminative appearance-based SVM classifier using local kernels verifies these hypotheses with a stronger spatial model and filter out false positives. We also enhance the discriminability of appearance codebooks for the target object class by selecting several most discriminative part codebooks that are built upon a pool of heterogeneous local descriptors, using a classification likelihood criterion. Experimental results show that both improvements significantly reduce the number of false positives and cross-class confusions and perform better than methods using only one cue. Hong Pan 0001, Liang-Zheng Xia, Truong Q. Nguyen |
ICASSP | 4 |
| 2011 | Single image spatially variant out-of-focus blur removalabstractThis paper addresses an out-of-focus blur problem in which the fore ground object is in focus whereas the background scene is out of focus. To recover the details of the background scene, a spatially variant blind deconvolution problem must be solved. However, spatially variant deconvolution is computationally intensive because Fourier based methods cannot be used to handle spatially variant convolution operators. The proposed method exploits the invariant structure of the problem by first predicting the background. Then a blind deconvolution algorithm is applied to estimate the blur kernel and a coarse estimate of the image is found as a side product. Finally, the back ground is recovered using total variation minimization, and fused with the foreground to produce the final deblurred image. Stanley H. Chan, Truong Q. Nguyen |
ICIP | 2 |
| 2011 | Hazy image modeling using color ellipsoidsabstractSeveral methods exist that try to remove haze from a single image but they lack in the support for why a particular method was chosen. The idea of using color ellipsoids is presented in this paper for the purpose of developing a new framework for analyzing single image dehazing methods. Synthetic and real world images are used to support important properties of the color ellipsoids that give a queue to the amount of haze in an image. As an example of its usefulness, a single dehazing method is analyzed within the color ellipsoid framework. Kristofor B. Gibson, Truong Q. Nguyen |
ICIP | 2 |
| 2011 | Fast and high quality learning-based super-resolution utilizing TV regularization methodabstractSuper-resolution image reconstruction is an important technology in many image processing areas such as image sensing, medical imaging, satellite imaging, and television signal conversion. It is also a key word of a recent consumer HDTV set that utilizes the CELL processor. Among various super-resolution methods, the learning-based method is one of the most promising solutions. The problem of the learning-based method is its enormous computational time for image searching from the large database of training images. We have proposed a new Total Variation (TV) regularization super-resolution method that utilizes a learning-based super-resolution method. We have obtained excellent results in image quality improvement. However, our proposed method needs long computational time because of the learning-based method. In this paper, we examine two methods that reduce the computational time of the learning-based method. The resulting algorithms reduce complexity significantly while maintaining comparable image quality. This enables the adoption of learning-based super-resolution to the motion pictures such as HDTV and internet movies. Tomio Goto, Shotaro Suzuki, Satoshi Hirano, Masaru Sakurai, Truong Q. Nguyen |
ICIP | 5 |
| 2011 | An optimal design of FIR filters with discrete coefficients and image sampling applicationabstractThe paper proposes a new approach for the design of linear phase finite impulse response (FIR) filters with discrete co-efficient values. This problem is a very hard combinatoric discrete optimization, which results in the prohibitive computational complexity for solution. In this paper, we first explicitly express the discrete coefficients of filters as indefinite quadratic but continuous constraints. We then develop an efficient iterative algorithm to tackle the nonconvex optimization problem to locate optimal discrete filter coefficients. By numerical simulation results, we show that our proposed method significantly outperform the methods using quantized coefficients of filters. We also provide an image sampling application to illustrate the performance of our designed filters. Ha Hoang Kha, Hoang Duong Tuan, Truong Q. Nguyen |
ICIP | 3 |
| 2011 | Block-lifting factorization of M-channel biorthogonal filter banks with an arbitrary McMillan degreeabstractA block-lifting factorization of M-channel biorthogonal filter banks (BOFBs) with degree-N building blocks and even/odd M (M ≥ 2) for lossless-to-lossy image coding is introduced in this paper. In the previous work, block-lifting factorization of M-channel BOFBs has been proposed. Since the block-lifting structure does not require the restriction of determinant of each building block, it achieves better coding performance than the conventional methods. However, the factorization is not completed because McMillan degree in each building block is fixed M/2 (M is even). This paper proposes a block-lifting factorization without restrictions for a fixed degree and even block size. Our proposal is validated by several filter designs and their application to lossless-to-lossy image coding. Taizo Suzuki, Masaaki Ikehara, Truong Q. Nguyen |
ICIP | 3 |
| 2011 | Advances in multirate filter bank structures and multiscale representations
Thierry Blu, Laurent Duval, Truong Q. Nguyen, Jean-Christophe Pesquet |
Signal Process. | 3 |
| 2011 | No-reference image quality assessment using structural activity
Jing Zhang 0008, Thinh M. Le, Sim Heng Ong, Truong Q. Nguyen |
Signal Process. | 4 |
| 2011 | Adaptive Lagrange Multiplier Selection Using Classification-Maximization and Its Application to Chroma QP Offset DecisionabstractIn this paper, we propose bit allocation between luma samples and chroma samples using chroma quantization parameter (QP) offsets for Cb and Cr. For this work, we propose an efficient adaptive Lagrange multiplier selection method using classification-maximization, and then apply the proposed adaptive Lagrange multiplier selection to decide chroma QP offsets for Cb and Cr. To our knowledge, this is the first proposal to adaptively decide chroma QP offsets. Because the default mapping function between a chroma QP and a luma QP in H.264 is unbalanced at especially low QPs, the proposed chroma QP offset decision achieves improvement up to 0.8 dB from the experimental results. Cheolhong An, Truong Q. Nguyen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | An Augmented Lagrangian Method for Total Variation Video RestorationabstractThis paper presents a fast algorithm for restoring video sequences. The proposed algorithm, as opposed to existing methods, does not consider video restoration as a sequence of image restoration problems. Rather, it treats a video sequence as a space-time volume and poses a space-time total variation regularization to enhance the smoothness of the solution. The optimization problem is solved by transforming the original unconstrained minimization problem to an equivalent constrained minimization problem. An augmented Lagrangian method is used to handle the constraints, and an alternating direction method is used to iteratively find solutions to the subproblems. The proposed algorithm has a wide range of applications, including video deblurring and denoising, video disparity refinement, and hot-air turbulence effect reduction. Stanley H. Chan, Ramsin Khoshabeh, Kristofor B. Gibson, Philip E. Gill, Truong Q. Nguyen |
IEEE Trans. Image Process. | 5 |
| 2011 | LCD Motion Blur: Modeling, Analysis, and AlgorithmabstractLiquid crystal display (LCD) devices are well known for their slow responses due to the physical limitations of liquid crystals. Therefore, fast moving objects in a scene are often perceived as blurred. This effect is known as the LCD motion blur. In order to reduce LCD motion blur, an accurate LCD model and an efficient deblurring algorithm are needed. However, existing LCD motion blur models are insufficient to reflect the limitation of human-eye-tracking system. Also, the spatiotemporal equivalence in LCD motion blur models has not been proven directly in the discrete 2-D spatial domain, although it is widely used. There are three main contributions of this paper: modeling, analysis, and algorithm. First, a comprehensive LCD motion blur model is presented, in which human-eye-tracking limits are taken into consideration. Second, a complete analysis of spatiotemporal equivalence is provided and verified using real video sequences. Third, an LCD motion blur reduction algorithm is proposed. The proposed algorithm solves an l(1)-norm regularized least-squares minimization problem using a subgradient projection method. Numerical results show that the proposed algorithm gives higher peak SNR, lower temporal error, and lower spatial error than motion-compensated inverse filtering and Lucy-Richardson deconvolution algorithm, which are two state-of-the-art LCD deblurring algorithms. Stanley H. Chan, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2011 | Optimal Design of FIR Triplet Halfband Filter Bank and Application in Image CodingabstractThis correspondence proposes an efficient semidefinite programming (SDP) method for the design of a class of linear phase finite impulse response triplet halfband filter banks whose filters have optimal frequency selectivity for a prescribed regularity order. The design problem is formulated as the minimization of the least square error subject to peak error constraints and regularity constraints. By using the linear matrix inequality characterization of the trigonometric semi-infinite constraints, it can then be exactly cast as a SDP problem with a small number of variables and, hence, can be solved efficiently. Several design examples of the triplet halfband filter bank are provided for illustration and comparison with previous works. Finally, the image coding performance of the filter bank is presented. Ha Hoang Kha, Hoang Duong Tuan, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2011 | Real-Time Affine Global Motion Estimation Using Phase Correlation and its Application for Digital Image StabilizationabstractWe propose a fast and robust 2D-affine global motion estimation algorithm based on phase-correlation in the Fourier-Mellin domain and robust least square model fitting of sparse motion vector field and its application for digital image stabilization. Rotation-scale-translation (RST) approximation of affine parameters is obtained at the coarsest level of the image pyramid, thus ensuring convergence for a much larger range of motions. Despite working at the coarsest resolution level, using subpixel-accurate phase correlation provides sufficiently accurate coarse estimates for the subsequent refinement stage of the algorithm. The refinement stage consists of RANSAC based robust least-square model fitting for sparse motion vector field, estimated using block-based subpixel-accurate phase correlation at randomly selected high activity regions in finest level of image pyramid. Resulting algorithm is very robust to outliers such as foreground objects and flat regions. We investigate the robustness of the proposed method for digital image stabilization application. Experimental results show that the proposed algorithm is capable of estimating larger range of motions as compared to another phase correlation method and optical flow algorithm. Sanjeev Kumar 0003, Haleh Azartash, Mainak Biswas, Truong Q. Nguyen |
IEEE Trans. Image Process. | 4 |
| 2011 | Rate Control Scheme for Consistent Video Quality in Scalable Video CodecabstractMultimedia data delivered to mobile devices over wireless channels or the Internet are complicated by bandwidth fluctuation and the variety of mobile devices. Scalable video coding has been developed as an extension of H.264/AVC to solve this problem. Since scalable video codec provides various scalabilities to adapt the bitstream for the channel conditions and terminal types, scalable codec is one of the useful codecs for wired or wireless multimedia communication systems, such as IPTV and streaming services. In such scalable multimedia communication systems, video quality fluctuation degrades the visual perception significantly. It is important to efficiently use the target bits in order to maintain a consistent video quality or achieve a small distortion variation throughout the whole video sequence. The scheme proposed in this paper provides a useful function to control video quality in applications supporting scalability, whereas conventional schemes have been proposed to control video quality in the H.264 and MPEG-4 systems. The proposed algorithm decides the quantization parameter of the enhancement layer to maintain a consistent video quality throughout the entire sequence. The video quality of the enhancement layer is controlled based on a closed-form formula which utilizes the residual data and quantization error of the base layer. The simulation results show that the proposed algorithm controls the frame quality of the enhancement layer in a simple operation, where the parameter decision algorithm is applied to each frame. Chan-Won Seo, Jong-Ki Han, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2011 | Enhanced Adaptive Loop Filter for Motion Compensated FrameabstractWe propose an adaptive loop filter to remove the redundancy between current and motion compensated frames so that the residual signal is minimized, thus coding efficiency increases. The loop filter coefficients and offset are optimized for each frame or a set of blocks to minimize the total energy of the residual signal resulting from motion estimation and compensation. The optimized loop filter with offset is applied for the set of blocks where the filtering process gives coding gain based upon rate-distortion cost. The proposed loop filter is used for the motion compensated frame whereas the conventional adaptive interpolation filter (AIF) is applied to the reference frames to interpolate the subpixel values. Another conventional scheme adaptive loop filter (ALF), is used after deblocking filtering to enhance quality of reconstructed frames, not to minimize energy of residual signal. The proposed loop filter can be used in combination with the AIF and ALF. Experimental results show that proposed algorithm provides the averaged bit reduction of 8% compared to conventional H.264/AVC scheme. When the proposed scheme is combined with AIF and ALF, the coding gain increases even further. Young-Joe Yoo, Chan-Won Seo, Jong-Ki Han, Truong Q. Nguyen |
IEEE Trans. Image Process. | 4 |
| 2011 | Spatioangular Prefiltering for Multiview 3D DisplaysabstractIn this paper, we analyze the reproduction of light fields on multiview 3D displays. A three-way interaction between the input light field signal (which is often aliased), the joint spatioangular sampling grids of multiview 3D displays, and the interview light leakage in modern multiview 3D displays is characterized in the joint spatioangular frequency domain. Reconstruction of light fields by all physical 3D displays is prone to light leakage, which means that the reconstruction low-pass filter implemented by the display is too broad in the angular domain. As a result, 3D displays excessively attenuate angular frequencies. Our analysis shows that this reduces sharpness of the images shown in the 3D displays. In this paper, stereoscopic image recovery is recast as a problem of joint spatioangular signal reconstruction. The combination of the 3D display point spread function and human visual system provides the narrow-band low-pass filter which removes spectral replicas in the reconstructed light field on the multiview display. The nonideality of this filter is corrected with the proposed prefiltering. The proposed light field reconstruction method performs light field antialiasing as well as angular sharpening to compensate for the nonideal response of the 3D display. The union of cosets approach which has been used earlier by others is employed here to model the nonrectangular spatioangular sampling grids on a multiview display in a generic fashion. We confirm the effectiveness of our approach in simulation and in physical hardware, and demonstrate improvement over existing techniques. Vikas Ramachandra, Keigo Hirakawa, Matthias Zwicker, Truong Q. Nguyen |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2010 | Subpixel motion estimation without interpolationabstractWe propose a fast subpixel motion estimation method for motion deblurring, where conventional motion estimation algorithms used in video codings are too complex. The new algorithm is a combination of block matching and optical flow. It does not require any interpolation and it does not provide motion compensated frames. Thus it is much faster than conventional methods. Statistical results show that the new algorithm performs quickly and accurately. It also demonstrates compatible performance with the benchmarking full search algorithm, yet uses significantly less amount of time. Stanley H. Chan, Dung Trung Vo, Truong Q. Nguyen |
ICASSP | 3 |
| 2010 | Rate distortion optimization for bidirectional scalable motion modelabstractThe fully scalable motion model (SMM) is proposed for scalable video codec by taking advantage of motion information scalability. In previous work, SMM has been improved to support hierarchical B frame and bidirectional or multidirectional motion estimation. It yields better results comparing to method using unidirectional motion estimation. However, the process of bidirectional or multidirectional rate distortion optimization motion estimation is very complex and is the most time-consuming task in the encoding process. We present several algorithms to reduce its complexity and show corresponding simulation to compare their performance. Meng-Ping Kao, Truong Q. Nguyen |
ICASSP | 4 |
| 2010 | Demosaicking imageswith motion blurabstractIn standard digital color imaging, each pixel position acquires data for only one color plane and the remaining two color planes must be inferred through a process known as demosaicking. Furthermore, the image is susceptible to blurring artifacts due to a moving camera or fast moving subject. In this work we develop a robust framework to demosaick the color filter array (CFA) image while reducing the blur corrupting the image. We begin by defining a color motion blur model that describes the motion blur artifacts affecting color images. We then integrate the motion blur model in the demosaicking algorithm to obtain a computationally efficient framework for deblurring while demosaicking. Shay Har-Noy, Stanley H. Chan, Truong Q. Nguyen |
ICASSP | 3 |
| 2010 | High frame rate Motion Compensated Frame Interpolation in High-Definition video processingabstractNumerous MCFI methods have been proposed to increase the frame rate in the past ten years. However, these methods usually focus on how to double the frame rate and involve complex computation, complicated time-consuming iterations, and are difficult to implement for real-time High-Definition (HD) videos. In this paper, a fast one-pass processing method is proposed for high Motion Compensated Frame Interpolation in HD videos. This approach is based on our basic MCFI scheme, slightly increases complexity, and achieves 4x Frame Rate Up Conversion. Relative Motion Estimation (RME) is also proposed to enhance the accuracy of motion search. Yen-Lin Lee, Truong Q. Nguyen |
ICASSP | 2 |
| 2010 | Empirical Type-i filter design for image interpolationabstractEmpirical filter designs generalize relationships inferred from training data to effect realistic solutions that conform well to the human visual system. Complex algorithms involving multiple linear regressions produce optimal results, but a single zero-phase filter yields comparable image quality at a fraction of the computational load. We propose an algorithm that builds a single symmetrical linear filter based purely on collected training data. Such a filtering technique balances the tradeoff between performance and complexity. Previous implementations of zero-phase interpolation filters as well as other learning-based interpolating algorithms are analyzed and examined. The proposed algorithm utilizes a Type-I symmetrical filter, an improvement and alternative over previous work on Type-II empirically-based interpolating filters. Given image training patches, the work discusses the enforcement of our filter properties while simultaneously drawing information from the training set. Additionally, we describe the implementation of the designed filter, its application, and related considerations. Finally, advantages of the proposed algorithm are analyzed. Karl S. Ni, Truong Q. Nguyen |
ICASSP | 2 |
| 2010 | Spatio-angular sharpening for multiview 3D displaysabstractIn this paper, we analyze the reproduction of light fields on multiview 3D displays. A two-way interaction between the input light field signal (which is often aliased) and the interview light leakage in modern multiview 3D displays is characterized in the joint spatio-angular frequency domain. Reconstruction of light fields by all physical 3D displays is prone to light leakage. This means that the reconstruction low pass filter implemented by the display is too broad in the angular domain, which causes loss of image sharpness. The combination of the 3D display point spread function and human visual system provides the narrow band low pass filter which removes spectral replicas in the reconstructed light field on the multiview display. The non-ideality of this filter is corrected with the proposed prefiltering technique. Vikas Ramachandra, Keigo Hirakawa, Matthias Zwicker, Truong Q. Nguyen |
ICASSP | 4 |
| 2010 | 2-D two-fold symmetric circular shaped filter design with homomorphic processing applicationabstractA design method of a linear-phased, two-dimensional (2-D), two-fold symmetric circular shaped filter is presented in this paper. Although the proposed method designs a non-separable filter, its implementation has linear complexity. The shape of the passband and the stopband is expressed in terms of level sets of second order trigonometric polynomials. This enables the transformation of the filter specifications to a Semi-Definite Program (SDP) of moderate dimension. The proposed filter outperforms currently available filter design methods. We present a performance comparison, as well as a homomorphic processing image enhancement example to illustrate the effectiveness of this method. Akila J. Seneviratne, Ha Hoang Kha, Hoang Duong Tuan, Truong Q. Nguyen |
ICASSP | 4 |
| 2010 | Total subset variation priorabstractWe propose total subset variation (TSV), a convexity preserving generalization of the total variation (TV) prior, for higher order clique MRF. A proposed differentiable approximation of the TSV prior makes it amenable for use in large images (e.g. 1080p). A convex relaxation of sub-exponential distribution is proposed as a criterion to determine the parameters of the optimization problem resulting from the TSV prior. For the super-resolution application, experiments show reconstruction error improvement with respect to the TV and other methods. Sanjeev Kumar 0003, Truong Q. Nguyen |
ICIP | 2 |
| 2010 | Robust object detection scheme using feature selectionabstractFeature selection is an important issue for object detection. In this paper, we propose an effective wrapper-based feature selection scheme using Binary Particle Swarm Optimization (BPSO) and Support Vector Machine (SVM) for object detection. In our algorithm, Scale-Invariant Feature Transform (SIFT) descriptors in a patch around the keypoints are extracted as the initial feature representations. The initial feature set is fed into the feature selection module in which the BPSO searches the feature space, and a SVM classifier serves as an evaluator for the performance of the feature subset selected by the BPSO. We tested the proposed detection scheme on the UIUC car dataset and our results show that feature selection scheme not only improves the detection accuracy but also enhances the detection efficiency. Hong Pan 0001, Liang-Zheng Xia, Truong Q. Nguyen |
ICIP | 3 |
| 2010 | View synthesis based on Conditional Random Fields and graph cutsabstractWe propose a novel method to synthesize intermediate views from two stereo images and disparity maps that is robust to errors in disparity maps. The proposed method computes a placement matrix from each disparity map that can be used to correct errors when warping pixels from reference view to virtual view. The second contribution is a new hole filling method that uses depth, edge, and segmentation information to aid the process of filling disoccluded pixels. The proposed method selects pixels from segmented regions that are connected to the disoccluded region as candidates to fill the disoccluded pixels. We also provide an explicit probabilistic model to select the best candidate for each disoccluded pixel efficiently with Conditional Random Fields (CRFs) and graph-cuts. Lam C. Tran, Christopher Joseph Pal, Truong Q. Nguyen |
ICIP | 3 |
| 2010 | Using Adaboost on contourlet based image deblurring for Fluid Lens Camera SystemsabstractThe Fluidic Lens Camera System provides an exciting opportunity for the Image Processing Community. Designed for a surgical environment, this camera has higher magnification and has better portability than traditional laparoscopic cameras. From an image processing prospective, the fluid causes non-uniform blur of different color planes. While the green image is sharp, the red and blue images are blurred. Previous methods have been developed to separate out the edge and shading components of the green image and to use the edge information in green to replace the blurred blue edges. This algorithm succeed in most areas, however in some areas, color bleeding artifacts occurred. We restate this problem as a classification problem. Using the contourlet and wavelet coefficients as features, the proposed algorithm determines in what areas color bleeding will occur and does not apply the sharpening algorithm in these areas. By applying the previous contourlet method in areas where it succeeds, we can produce an overall sharper image with reduced color bleeding artifacts. The ability to correctly classify when the previous algorithm will succeed is crucial to the success of the algorithm. The principal application is medical imaging, however, the fields of satellite pan-sharpening and image denoising can benefit from the results found in this paper. Jack Tzeng, Yoav Freund, Truong Q. Nguyen |
ICIP | 3 |
| 2010 | LCD motion blur modeling and simulationabstractLiquid crystal display (LCD) devices are well known to have slow response due to the physical limitations of the liquid crystals. Therefore, fast moving objects in a scene are often seen blurred on an LCD. In order to reduce motion blur, an accurate LCD model and an efficient deblurring algorithm are needed. However, existing LCD models are inadequate to reflect the human eye tracking limitation. Also, the spatial-temporal equivalence in LCD models is widely used but not proven directly in the 2D discrete spatial domain. In this paper, we study the human eye tracking limit by reviewing a number of papers in the cognitive science literature. We provide both theoretical and experimental arguments to support our findings. Also, we prove the spatial-temporal equivalence rigourously and verify the results using real video sequences. Stanley H. Chan, Truong Q. Nguyen |
ICME | 2 |
| 2010 | Motion vector refinement for FRUC using saliency and segmentationabstractMotion-Compensated Frame Interpolation (MCFI) is a technique used extensively for increasing the temporal frequency of a video sequence. In order to obtain a high quality interpolation, the motion field between frames must be well-estimated. However, many current techniques for determining the motion are prone to errors in occlusion regions, as well as regions with repetitive structure. An algorithm is proposed for improving both the objective and subjective quality of MCFI by refining the motion vector field. A Discriminant Saliency classifier is employed to determine regions of the motion field which are most important to a human observer. These regions are refined using a multi-stage motion vector refinement which promotes candidates based on their likelihood given a local neighborhood. For regions which fall below the saliency threshold, frame segmentation is used to locate regions of homogeneous color and texture via Normalized Cuts. Motion vectors are promoted such that each homogeneous region has a consistent motion. Experimental results demonstrate an improvement over previous methods in both objective and subjective picture quality. Natan Jacobson, Yen-Lin Lee, Vijay Mahadevan, Nuno Vasconcelos, Truong Q. Nguyen |
ICME | 5 |
| 2010 | Filter banks for improved LCD motion
Shay Har-Noy, E. Martinez, Truong Q. Nguyen |
Signal Process. Image Commun. | 3 |
| 2010 | Comparison of Two Frame Rate Conversion Schemes for Reducing LCD Motion BlursabstractLiquid crystal display (LCD) is known to have motion blur due to the slow response and sample-hold characteristics of the liquid crystal (LC). To alleviate the LCD motion blur, improving the LC response is the most fundamental solution. However, if the response time is shortened, then more frames are needed and hence frame rate up conversion (FRUC) should be used. In this paper, we study two FRUC methods. We compare the output signal qualities by studying the temporal and spatial profile of the two methods. We use the solution of Erickson-Leslie equation to derive the step response, in contrast to the existing literature where the resistor-capacitor (RC) approximation and uniform averaging function are used. The step response we derived is able to model not only the general trend of the rising and falling edges, but also the effects of different gray level transitions. Based on the step response, we analyze the two methods by comparing the observed signal in both the spatial and temporal domain. Stanley H. Chan, Thomas X. Wu, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 3 |
| 2010 | Efficient Bit Allocation and Rate Control Algorithms for Hierarchical Video CodingabstractHierarchical structure is a useful tool for providing the necessary scalability in adapting to the variety of channel environments. For schemes involving hierarchical picture structures, bit allocation, and rate control algorithms are vital components for improving video codec performance. Since conventional bit allocation schemes do not efficiently consider the hierarchical structure characteristics, it is difficult to optimize the video quality at an arbitrary bitrate. Similarly, conventional quantization parameter decision methods are not appropriate for controlling the bitrate generated by a codec using a hierarchical encoding structure. In this paper, we propose an effective bit allocation scheme that assigns the target number of bits to pictures or macroblocks (MBs) and improves the overall quality of images encoded by a hierarchical-based encoder. A rate control scheme is also proposed to ensure that the generated bitrate is equal to the assigned target bitrate. From the simulation results, the proposed schemes outperformed conventional methods from a rate-distortion perspective, by efficiently controlling the bitrate of the MB unit. The algorithms regulated the generated bits to achieve the target bits by using the proposed linear R-Q model. Chan-Won Seo, Jong-Ki Han, Truong Q. Nguyen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Bidirectional Scalable Motion for Scalable Video CodingabstractMotion information scalability is an important requirement for a fully scalable video codec, especially in low bit rate or small resolution decoding scenarios, for which the fully scalable motion model (SMM) has been proposed. SMM can collaborate flawlessly with other scalabilities, such as spatial, temporal and quality, in a scalable video codec. It performs better than the nonscalable motion model. To further improve the SMM, this paper extends the algorithm to support the hierarchical B frame structure and bidirectional or multidirectional motion estimation. Furthermore, the corresponding rate distortion optimized estimation for improved efficiency in several scenarios is discussed. Several simulation results based on the updated framework are presented to verify the advantage of this extension. Meng-Ping Kao, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2010 | A Novel Approach to FRUC Using Discriminant Saliency and Frame SegmentationabstractMotion-compensated frame interpolation (MCFI) is a technique used extensively for increasing the temporal frequency of a video sequence. In order to obtain a high quality interpolation, the motion field between frames must be well-estimated. However, many current techniques for determining the motion are prone to errors in occlusion regions, as well as regions with repetitive structure. We propose an algorithm for improving both the objective and subjective quality of MCFI by refining the motion vector field. We first utilize a discriminant saliency classifier to determine which regions of the motion field are most important to a human observer. These regions are refined using a multistage motion vector refinement (MVR), which promotes motion vector candidates based upon their likelihood given a local neighborhood. For regions which fall below the saliency-threshold, a frame segmentation is used to locate regions of homogeneous color and texture via normalized cuts. Motion vectors are promoted such that each homogeneous region has a consistent motion. Experimental results demonstrate an improvement over previous frame rate up-conversion (FRUC) methods in both objective and subjective picture quality. Natan Jacobson, Yen-Lin Lee, Vijay Mahadevan, Nuno Vasconcelos, Truong Q. Nguyen |
IEEE Trans. Image Process. | 5 |
| 2010 | Rate-Distortion Optimized Bitstream Extractor for Motion Scalability in Wavelet-Based Scalable Video CodingabstractMotion scalability is designed to improve the coding efficiency of a scalable video coding framework, especially in the medium to low range of decoding bit rates and spatial resolutions. In order to fully benefit from the superiority of motion scalability, a rate-distortion optimized bitstream extractor, which determines the optimal motion quality layer for any specific decoding scenario, is required. In this paper, the determination process first starts off with a brute force searching algorithm. Although guaranteed by the optimal performance within the search domain, it suffers from high computational complexities. Two properties, i.e., the monotonically nondecreasing property and the unimodal property, are then derived to accurately describe the rate-distortion behavior of motion scalability. Based on these two properties, modified searching algorithms are proposed to reduce the complexity (up to five times faster) and to achieve the global optimality, even for those decoding scenarios outside the search domain. Meng-Ping Kao, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2010 | Adaptive Directional Wavelet Transform Based on Directional PrefilteringabstractThis paper proposes an efficient approach for adaptive directional wavelet transform (WT) based on directional prefiltering. Although the adaptive directional WT is able to transform an image along diagonal orientations as well as traditional horizontal and vertical directions, it sacrifices computation speed for good image coding performance. We present two efficient methods to find the best transform directions by prefiltering using 2-D filter bank or 1-D directional WT along two fixed directions. The proposed direction calculation methods achieve comparable image coding performance comparing to the conventional one with less complexity. Furthermore, transform direction data of the proposed method can be used for content-based image retrieval to increase retrieval ratio. Yuichi Tanaka 0001, Madoka Hasegawa, Shigeo Kato, Masaaki Ikehara, Truong Q. Nguyen |
IEEE Trans. Image Process. | 5 |
| 2010 | Contourlet Domain Multiband Deblurring Based on Color Correlation for Fluid Lens CamerasabstractDue to the novel fluid optics, unique image processing challenges are presented by the fluidic lens camera system. Developed for surgical applications, unique properties, such as no moving parts while zooming and better miniaturization than traditional glass optics, are advantages of the fluid lens. Despite these abilities, sharp color planes and blurred color planes are created by the nonuniform reaction of the liquid lens to different color wavelengths. Severe axial color aberrations are caused by this reaction. In order to deblur color images without estimating a point spread function, a contourlet filter bank system is proposed. Information from sharp color planes is used by this multiband deblurring method to improve blurred color planes. Compared to traditional Lucy-Richardson and Wiener deconvolution algorithms, significantly improved sharpness and reduced ghosting artifacts are produced by a previous wavelet-based method. Directional filtering is used by the proposed contourlet-based system to adjust to the contours of the image. An image is produced by the proposed method which has a similar level of sharpness to the previous wavelet-based method and has fewer ghosting artifacts. Conditions for when this algorithm will reduce the mean squared error are analyzed. While improving the blue color plane by using information from the green color plane is the primary focus of this paper, these methods could be adjusted to improve the red color plane. Many multiband systems such as global mapping, infrared imaging, and computer assisted surgery are natural extensions of this work. This information sharing algorithm is beneficial to any image set with high edge correlation. Improved results in the areas of deblurring, noise reduction, and resolution enhancement can be produced by the proposed algorithm. Jack Tzeng, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2010 | Selective Data Pruning-Based Compression Using High-Order Edge-Directed InterpolationabstractThis paper proposes a selective data pruning-based compression scheme to improve the rate-distortion relation of compressed images and video sequences. The original frames are pruned to a smaller size before compression. After decoding, they are interpolated back to their original size by an edge-directed interpolation method. The data pruning phase is optimized to obtain the minimal distortion in the interpolation phase. Furthermore, a novel high-order interpolation is proposed to adapt the interpolation to several edge directions in the current frame. This high-order filtering uses more surrounding pixels in the frame than the fourth-order edge-directed method and it is more robust. The algorithm is also considered for multiframe-based interpolation by using spatio-temporally surrounding pixels coming from the previous frame. Simulation results are shown for both image interpolation and coding applications to validate the effectiveness of the proposed methods. Dung Trung Vo, Joel Sole, Peng Yin 0002, Cristina Gomila, Truong Q. Nguyen |
IEEE Trans. Image Process. | 5 |
| 2009 | Entropy Coding via Parametric Source Model with Applications in Fast and Efficient Compression of Image and Video DataabstractIn this paper a framework is proposed for efficient entropy coding of data which can be represented by a parametric distribution model. Based on the proposed framework, an entropy coder achieves coding efficiency by estimating the parameters of the statistical model (for the coded data), either via maximum a posteriori (MAP) or Maximum Likelihood (ML) parameter estimation techniques. The problem of optimal entropy coding for transmission of a block of data x1,,x2,...xN, can be formulated by assuming that the data comes from a source with a parametric probability mass function (pmf) P(X1,X2,...XN;thetas) with parameter thetas (in general thetas is a vector). The parametric model assumption makes it possible to assign a probability to the event of observing x1,,x2,...xN, and use this probability for entropy coding of this data, only by conveying the parameter thetas.The impressive results from the simple parametric model, based on a geometric distribution of coded data for compression of natural images, are encouraging to further investigate the effect of more complicated data models such as Poisson distribution and mixture models. Koohyar Minoo, Truong Q. Nguyen |
DCC | 2 |
| 2009 | Fast LCD motion deblurring by decimation and optimizationabstractThe LCD deblurring problem is considered as a simple bounded quadratic programming problem and is solved using conjugate gradient with early stopping criteria to avoid excessive search. A decimation and interlace interpolation method is introduced to reduce the computing time. Solutions are competitive to those the generated by conventional Lucy Richardson algorithm, but using much shorter amount of time. The method can be extended to higher decimation factors. Visual subjective tests are conducted to justify our proposed method. Stanley H. Chan, Truong Q. Nguyen |
ICASSP | 2 |
| 2009 | Analog flat filter designabstractThis paper proposes a systematic approach for the design of a general class of analog infinite-impulse-response (IIR) filters, which includes all well-known classical analog filters as a special case. All specifications including the conventional ones and also filter flatness degrees are explicitly incorporated into design process. Several numerical examples are presented to demonstrate the efficiency and flexibility of the proposed method. Hung Gia Hoang, Hoang Duong Tuan, Truong Q. Nguyen |
ICASSP | 3 |
| 2009 | Rate-distortion optimized bitstream extractor for motion scalability in scalable video codingabstractMotion scalability is designed to improve the coding efficiency of a scalable video coding framework, especially in the medium to low range of decoding bit rates or spatial resolutions. In order to fully benefit from the superiority of motion scalability, a rate-distortion optimized bitstream extractor, which determines the optimal motion quality layer for each decoding scenario, is required. In this paper, the determination process first starts off with a brute force searching algorithm. Although guaranteed by the optimal performance within the search domain, it has high computational complexity. Two properties, i.e. the monotonically non-decreasing property and the unimodal property, are then derived to accurately describe the rate-distortion behavior of motion scalability. Based on these two properties, modified searching algorithms are proposed to reduce the complexity by a factor up to 5. Meng-Ping Kao, Truong Q. Nguyen |
ICASSP | 2 |
| 2009 | Polyphase interpretation of empirical image interpolationabstractWe observe several characteristics of empirical image interpolating algorithms and contribute four novel concepts and claims. First, we interpret well-known classification-based filtering algorithms in terms of their polyphase components. We examine the underlying principles behind the various fixed-scale linear interpolating kernels. Second, we conceptually extend the properties of the multiple filters to two dimensions to analyze frequency domain characteristics common to all empirically-designed interpolating filters. Third, we propose a general linear filter for image interpolation, which uses a universal magnitude response and zero-phase. Finally, the proposed filter is further generalized to support arbitrary scaling factors. We claim that at any scaling factor, the proposed algorithm yields low-complexity at a minimal loss of high image-quality with the ability to interpolate diverse image content. Karl S. Ni, Truong Q. Nguyen |
ICASSP | 2 |
| 2009 | Combined image plus depth seam carving for multiview 3D imagesabstractMultiview 3D displays have to multiplex a set of views on a single LCD panel. Due to this, each view has to be downsampled by a considerable amount leading to loss of details. In this paper, we extend the seam carving technique for adaptive resizing of images. It is proposed that the depth information be used along with the image pixel intensity values for resizing. This results in better resized multiview images. It is clear from the results presented that the object structure is maintained when the proposed method is used as compared to vanilla seam carving. Vikas Ramachandra, Matthias Zwicker, Truong Q. Nguyen |
ICASSP | 3 |
| 2009 | Using filter banks to enhance images for fluid lens cameras based on color correlationabstractThe novel field of fluid lens cameras introduces unique image processing challenges. Intended for surgical applications, these fluid optics systems have improved miniaturization over glass lenses and do not have moving parts while zooming. However, the liquid medium creates non-uniform color blur, which causes certain color planes to appear sharper than others. We propose an adapted perfect reconstruction filter bank that uses high frequency sub-bands of sharp color planes to improve blurred color planes. The approach is refined by adjusting the decomposition level based on limited channel information. This paper primarily considers the use of a sharp green color plane to improve a blurred blue color plane. More generally, these methods could improve the red color plane as well as any system with high edge correlation between two images. Jack Tzeng, Truong Q. Nguyen |
ICASSP | 2 |
| 2009 | Data pruning-based compression using high order edge-directed interpolationabstractThis paper proposes a data pruning-based compression scheme to improve the rate-distortion relation of compressed images and video sequences. The original frames are pruned to a smaller size before compression. After decoding, they are interpolated to their original size by an edge-directed interpolation. The data pruning is optimized to obtain the minimal distortion in the interpolation phase. Furthermore, a novel high order interpolation is proposed to adapt the interpolation to many edge directions. This high order filtering uses extra surrounding pixels and achieves more robust edge-directed image interpolation. Simulation results are shown for both image interpolation and coding applications. Dung Trung Vo, Joel Sole, Peng Yin 0002, Cristina Gomila, Truong Q. Nguyen |
ICASSP | 5 |
| 2009 | Bidirectional scalable motion for scalable video codingabstractMotion information scalability is an important requirement for a fully scalable video codec, especially in low bit rate or small resolution decoding scenarios, for which the fully scalable motion model (SMM) has been proposed. SMM can collaborate flawlessly with other scalabilities, such as spatial, temporal and quality, in a scalable video codec and it performs better than nonscalable motion model. To further improve the SMM, this paper further extends the algorithm to support hierarchical B frame structure and bidirectional or multidirectional motion estimation and, moreover, provides the corresponding rate distortion optimized estimation. Several simulation results are presented to verify the improvements of this extension. Meng-Ping Kao, Truong Q. Nguyen |
ICIP | 3 |
| 2009 | Efficient design of 2-D nonseparable filters of low complexityabstractBy using Semi-Definite Programming (SDP) as a tool, a new deign for Two-Dimensional (2-D) Diamond-Shaped (DS) filters is developed. Surprisingly, the diamond shape of the filter is exactly expressed by using simple 2-D trigonometric polynomial curves of second order. In contrast to the high-order polynomial transformation based methods, the order of the designed filters is kept moderate with no performance sacrifices. Unlike conventional non-separable 2-D filters, the designed filters allow fast digital implementation despite being nonseparable and hence they are of low complexity. Numerical Simulations with application to quincunx image sampling are also performed to illustrate the viability of our method. Kaveh Fanian, Hoang Duong Tuan, Truong Q. Nguyen |
ICIP | 3 |
| 2009 | LCD motion blur reduction using fir filter banksabstractDue to the sample-and-hold nature of liquid crystal display (LCD) image formation, LCDs suffer from motion picture blur. This is especially evident during scenes containing fast motion due to the inherent sample-and-hold nature of LCD image formation. Using models for the human visual system (HVS) we take a signal processing approach to solving this problem by pre-processing the data before it is sent to the display. Whereas previous pre-processing approaches either apply a simple high pass filter or an iterative deconvolution algorithm, this work uses a small collection of efficient linear FIR filters to reduce the amount of perceived motion blur. Specifically, we develop a two-channel non-perfect reconstruction filter bank to reduce the motion dependent low pass effects of the HVS. Perceptual tests indicate that our algorithm reduces the amount of perceived motion blur on LCDs at a lower complexity than the existing deconvolution approach. Shay Har-Noy, Truong Q. Nguyen |
ICIP | 2 |
| 2009 | Motion vector processing based on trajectory curve analysis for motion compensated frame interpolationabstractIn this paper, we propose a novel motion vector processing approach based on the temporal analysis of motion vector reliability for motion-compensated frame interpolation. We first address the problems of conventional methods that find motion by minimizing the difference between the forward and backward predictions. The ghost artifacts can be reduced due to this procedure, but the resulting motion vectors may not represent the actual motion. Therefore, we propose analyzing the temporal variations of the bidirectional prediction difference along the motion trajectory to detect temporally unreliable motion vectors and further improve the accuracy of the bidirectional motion vector processing. Experimental results show that the proposed method can effectively detect unreliable motion vectors and outperforms other methods in terms of visual quality, and structure similarity. Ai-Mei Huang, Truong Q. Nguyen |
ICIP | 2 |
| 2009 | Motion vector processing using the color informationabstractIn this paper, we investigate color distribution around object edges and further exploit this information for motion vector reliability analysis. Since the motion vectors are often estimated only based on luminance component, the resulting motion vectors may become unreliable for the areas with similar intensity. By making use of the color information for the received residual energy calculation, these unreliable motion vectors can effectively be identified. Moreover, the motion boundaries can be easily detected due to stronger residual energy distribution on the object boundaries. This characteristic can also be used to assist the motion vector processing in motion-compensated frame interpolation so that the interpolated object structures can be well maintained. Experimental results show that using same motion vector processing approach but with additional consideration on color information outperforms other methods. Ai-Mei Huang, Truong Q. Nguyen |
ICIP | 2 |
| 2009 | Fast one-pass motion compensated frame interpolation in high-definition video processingabstractIn this paper, a fast one-pass processing method with an efficient architecture is proposed for motion compensated frame interpolation (MCFI) in high-definition (HD) videos. Unlike previous works involving high complexity, complicated time-consuming iterations, and less practicability, the proposed method adopts one-pass and low-complexity concept. The proposed method operates a modified fast full-search algorithm, multi-level successive eliminate algorithm (MSEA), in a raster scan order as in the usual block-based processing order in popular codecs, such as H.264/AVC and VC-1. The proposed method analyzes and classifies temporal information to compensate insufficient spatial information based on preprocessed neighboring blocks. According to the analyzed temporal information, our method explores true motion candidates and refines the accuracy of true motions for sub-blocks. When searching motion candidates, the proposed method introduces an adaptive overlapped block matching algorithm called a multi-directional enlarged matching algorithm (MDEMA), and considers different overlapped types based on directions of current sought motion vector in order to enhance the searching accuracy and visual quality. Experimental results show that the proposed algorithm provides better video quality than conventional methods and shows satisfying performance. Yen-Lin Lee, Truong Q. Nguyen |
ICIP | 2 |
| 2009 | PPIQ: A probabilistic framework for Image Quality AssessmentabstractIn this paper a framework for image quality assessment (IQA) is introduced based on the properties of receptive fields (RFs) which are the primary mechanism for detection of visual patterns in the human visual system (HVS). The proposed framework offers a probabilistic approach to the perceptual IQA, based on the probability of detecting discrepancies (distortion) between the corresponding features of a test and a reference image. The proposed probabilistic perceptual image quality (PPIQ) framework facilitates defining specific perceptual metrics for specific applications. To give an example on how the PPIQ framework can be utilized to define an image quality metric (IQM), a sample IQM is introduced based on the properties of simple RFs of the early vision in the HVS. The sample IQM, based on the PPIQ framework, exhibits comparable accuracy to that of the legacy methods in terms of predicting the outcome of subjective image quality experiments. Koohyar Minoo, Truong Q. Nguyen |
ICIP | 2 |
| 2009 | Adaptive directionalwavelet transform using pre-directional filteringabstractThis paper proposes a computationally efficient approach of adaptive directional wavelet transform (AD WT). The AD WT is based on lifting implementation of WT, and it is able to transform an image along diagonal orientations as well as traditional horizontal and vertical directions. The AD WT sacrifices computational speed for its good image coding performance. We present an alternative method to find the best transform directions by pre-directional filterings for the AD WT. The proposed direction calculation method shows very comparable image coding performance to the conventional one, whereas its computational cost is relatively very low. Yuichi Tanaka 0001, Madoka Hasegawa, Shigeo Kato, Masaaki Ikehara, Truong Q. Nguyen |
ICIP | 5 |
| 2009 | Optimal motion compensated spatio-temporal filter for quality enhancement of H.264/AVC compressed video sequencesabstractThe paper proposes a novel algorithm to enhance the quality of H.264/AVC compressed video sequences by using an in-loop spatio-temporal motion compensated ??lter (MCSTF). Extra information from surrounding coded frames are used together with the information of the current coded frame to reduce the coding artifacts. With the availability of the original frame, the MCSTF coef??cients are optimized at the encoder and then are implemented at the decoder. Furthermore, an overlapped motion compensated scheme is proposed to reduce the blocking artifacts from surrounding motion compensated frames. Simulation results are judged by PSNR and ??ickering metric. Dung Trung Vo, Truong Q. Nguyen |
ICIP | 2 |
| 2009 | Method and Architecture Design for Motion Compensated Frame Interpolation in High-definition Video ProcessingabstractIn this paper, a novel, fast, and low complexity method is proposed for motion compensated frame interpolation (MCFI). Unlike previous works involving high complexity and complicated time-consuming iterations, the proposed method adopts one-pass and low-complexity approach without any repeated iteration and is capable of dealing with high definition (HD) video processing. The proposed method employs a unique true motion engine that explores at most nine motion candidates with different motion directions and then determines one true motion by referring to neighboring spatial information as well as temporal information. Besides, the proposed method introduces an adaptive overlapped block matching algorithm to enhance the searching accuracy and visual quality. An architecture design is also proposed, which employs a modified multi-level successive eliminate algorithm (MSEA) and has capability to reduce the heavy computation of full search. Experimental results show that the proposed algorithm provides better video quality than conventional methods and shows satisfying performance for HD1080p 30 fps at 180 MHz or HD720p 30 fps at 83 MHz. Yen-Lin Lee, Truong Q. Nguyen |
ISCAS | 2 |
| 2009 | An Adaptive Extension of Combined 2D and 1D-directional Filter BanksabstractIn this paper, we propose an extension of combined 2D and 1D-directional filter banks (TODFBs) with adaptive approach in their 1D-directional stages. TODFBs show better performance in image coding and denoising compared to the traditional wavelet transform (WT), however, they still have possibilities to improve their performance in image coding by adopting the adaptive directional WTs for their 1D-directional stages. An efficient method to determine the transform directions is also proposed. In image coding results, our proposed filter banks gain PSNR and visual quality improvements compared with the WTs and the non-adaptive TODFBs. Yuichi Tanaka 0001, Masaaki Ikehara, Truong Q. Nguyen |
ISCAS | 3 |
| 2009 | Higher-order feasible building blocks for lattice structure of oversampled linear-phase perfect reconstruction filter banks
Yuichi Tanaka 0001, Masaaki Ikehara, Truong Q. Nguyen |
Signal Process. | 3 |
| 2009 | Correlation-Based Motion Vector Processing With Adaptive Interpolation Scheme for Motion-Compensated Frame InterpolationabstractIn this paper, we address the problems of unreliable motion vectors that cause visual artifacts but cannot be detected by high residual energy or bidirectional prediction difference in motion-compensated frame interpolation. A correlation-based motion vector processing method is proposed to detect and correct those unreliable motion vectors by explicitly considering motion vector correlation in the motion vector reliability classification, motion vector correction, and frame interpolation stages. Since our method gradually corrects unreliable motion vectors based on their reliability, we can effectively discover the areas where no motion is reliable to be used, such as occlusions and deformed structures. We also propose an adaptive frame interpolation scheme for the occlusion areas based on the analysis of their surrounding motion distribution. As a result, the interpolated frames using the proposed scheme have clearer structure edges and ghost artifacts are also greatly reduced. Experimental results show that our interpolated results have better visual quality than other methods. In addition, the proposed scheme is robust even for those video sequences that contain multiple and fast motions. Ai-Mei Huang, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2009 | An Adaptable k -Nearest Neighbors Algorithm for MMSE Image InterpolationabstractWe propose an image interpolation algorithm that is nonparametric and learning-based, primarily using an adaptive k-nearest neighbor algorithm with global considerations through Markov random fields. The empirical nature of the proposed algorithm ensures image results that are data-driven and, hence, reflect "real-world" images well, given enough training data. The proposed algorithm operates on a local window using a dynamic k -nearest neighbor algorithm, where k differs from pixel to pixel: small for test points with highly relevant neighbors and large otherwise. Based on the neighbors that the adaptable k provides and their corresponding relevance measures, a weighted minimum mean squared error solution determines implicitly defined filters specific to low-resolution image content without yielding to the limitations of insufficient training. Additionally, global optimization via single pass Markov approximations, similar to cited nearest neighbor algorithms, provides additional weighting for filter generation. The approach is justified in using a sufficient quantity of training per test point and takes advantage of image properties. For in-depth analysis, we compare to existing methods and draw parallels between intuitive concepts including classification and ideas introduced by other nearest neighbor algorithms by explaining manifolds in low and high dimensions. Karl S. Ni, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2009 | Multiresolution Image Representation Using Combined 2-D and 1-D Directional Filter BanksabstractIn this paper, effective multiresolution image representations using a combination of 2-D filter bank (FB) and directional wavelet transform (WT) are presented. The proposed methods yield simple implementation and low computation costs compared to previous 1-D and 2-D FB combinations or adaptive directional WT methods. Furthermore, they are nonredundant transforms and realize quad-tree like multiresolution representations. In applications on nonlinear approximation, image coding, and denoising, the proposed filter banks show visual quality improvements and have higher PSNR than the conventional separable WT or the contourlet. Yuichi Tanaka 0001, Masaaki Ikehara, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2009 | Image Enhancement for Fluid Lens Camera Based on Color CorrelationabstractThe novel field of fluid lens cameras introduces unique image processing challenges. Intended for surgical applications, these fluid optics systems have a number of advantages over traditional glass lens systems. These advantages include improved miniaturization and no moving parts while zooming. However, the liquid medium creates two forms of image degradation: image distortion, which warps the image such that straight lines appear curved, and nonuniform color blur, which degrades the image such that certain color planes appear sharper than others. We propose the use of image processing techniques to reduce these degradations. To deal with image warping, we employ a conventional method that models the warping process as a degree-six polynomial in order to invert the effect. For image blur, we propose an adapted perfect reconstruction filter bank that uses high frequency sub-bands of sharp color planes to improve blurred color planes. The algorithm adjusts the number of levels in the decomposition and alters a prefilter based on crude knowledge of the blurring channel characteristics. While this paper primarily considers the use of a sharp green color plane to improve a blurred blue color plane, these methods can be applied to improve the red color plane as well, or more generally adapted to any system with high edge correlation between two images. Jack Tzeng, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2009 | Adaptive Fuzzy Filtering for Artifact Reduction in Compressed Images and VideosabstractA fuzzy filter adaptive to both sample's activity and the relative position between samples is proposed to reduce the artifacts in compressed multidimensional signals. For JPEG images, the fuzzy spatial filter is based on the directional characteristics of ringing artifacts along the strong edges. For compressed video sequences, the motion compensated spatiotemporal filter (MCSTF) is applied to intraframe and interframe pixels to deal with both spatial and temporal artifacts. A new metric which considers the tracking characteristic of human eyes is proposed to evaluate the flickering artifacts. Simulations on compressed images and videos show improvement in artifact reduction of the proposed adaptive fuzzy filter over other conventional spatial or temporal filtering approaches. Dung Trung Vo, Truong Q. Nguyen, Sehoon Yea, Anthony Vetro |
IEEE Trans. Image Process. | 2 |
| 2008 | Maximum Likelihood Rate Estimation: With Applications in Image and Video CompressionabstractTo overcome the computational complexity in making optimized rate-distortion coding decisions (due to the calculation of actual rate) we propose a rate estimation scheme based on the maximum likelihood parameter estimation (MLPE) method. Koohyar Minoo, Truong Q. Nguyen |
DCC | 2 |
| 2008 | Iterative R-D optimization of H.264abstractIn this paper, we apply the primal-dual decomposition and subgradient projection methods to solve the rate-distortion optimization problem with the constant bit rate constraint. The primal decomposition method enables spatial or temporal prediction dependency within a group of picture (GOP) to be processed in the master primal problem. As a result, we can apply the dual decomposition to minimize independently the Lagrangian cost of all the MBs using the reference software model of H.264. Furthermore, the optimal Lagrange multiplier lambda* is iteratively derived from the solution of the dual problem. As an example, we derive the optimal bit allocation condition with the consideration of temporal prediction dependency among the pictures. Experimental results show that the proposed method achieves better performance than the reference software model of H.264 with rate control for given bit constraint. Cheolhong An, Truong Q. Nguyen |
ICASSP | 2 |
| 2008 | Adaptable K-nearest neighbor for image interpolationabstractA variant of the k-nearest neighbor algorithm is proposed for image interpolation. Instead of using a static volume or static k, the proposed algorithm determines a dynamic k that is small for inputs whose neighbors are very similar and large for inputs whose neighbors are dissimilar. Then, based on the neighbors that the adaptable k provides and their corresponding similarity measures, a weighted MMSE solution defines filters specific to intrinsic content of a low-resolution input image patch without yielding to the limitations of a non-uniformly distributed training set. Finally, global optimization through a single pass Markovian-like network further imposes on filter weights. The approach is justified by a sufficient quantity of relevant training pairs per test input and compared to current state of the art nearest neighbor interpolation techniques. Kenta S. Ni, Truong Q. Nguyen |
ICASSP | 2 |
| 2008 | Statistical learning based intra prediction in H.264abstractIn this paper, we improve the performance of intra prediction and simplify mode decision procedure at the same time. For these works, we apply a statistical learning method such as Support Vector Machines for Regression (SVR) to improve the performance of current H.264 intra prediction via batch learning. In addition, we only use single Macro Block type and one intra prediction mode with high prediction performance to simplify mode decision procedure. In our knowledge, this work is the first approach to apply a statistical learning method for prediction of video sequences. Therefore, we introduce theoretical backgrounds of SVR, and show the possibility of this challenge for video compression. From the experimental results, statistical learning based intra prediction improves significantly the average Peak Signal-to-Noise Ratio of intra prediction than the performance of current H.264. Cheolhong An, Truong Q. Nguyen |
ICIP | 2 |
| 2008 | Correlation-based motion vector processing for motion compensated frame interpolationabstractIn this paper, a novel correlation-based motion vector processing method is proposed for motion compensated frame interpolation. We first address the problem of unreliable motion vectors due to low correlation. Unlike other motion vector processing methods using vector median filter, we proposed using bidirectional prediction difference to select the best motion vectors with constraint on increasing motion vector correlation. Furthermore, to reduce the blockiness artifacts while maintaining edges, we use an adaptive vector averaging filter to obtain a finer motion field by taking motion vector correlation into account. Experimental results show that the proposed scheme outperforms other methods in terms of visual quality and PSNR performance. Ai-Mei Huang, Truong Q. Nguyen |
ICIP | 2 |
| 2008 | A perceptual metric for blind measurement of blocking artifacts with applications in transform-block-based image and video codingabstractIn this paper, we analyze the formation of blocking artifacts as a result of quantization of the discrete cosine transform (DCT) coefficients. These artifacts are known to be the dominant distortion of lossy block-based hybrid image and video coding schemes when operating at very low data rates. Our analysis is carried out based on the theory of edge detection and by means of a mathematical framework to model the properties of receptive fields, which are known to be the main mechanism for sensing intensity contrasts and edges. Based on our analysis we propose a metric to quantify the blocking artifact distortion. The proposed distortion metric considers the susceptibility of a region in a reconstructed (decoded) image to the formation of blocking artifacts via a probabilistic model, which quantifies the possibility of a detected edge to be a distortion. Koohyar Minoo, Truong Q. Nguyen |
ICIP | 2 |
| 2008 | Image interpolation using classification and stitchingabstractImage interpolation is a well-studied signal processing application that continues to receive substantial attention from the research community. It has been recognized that taking into account the presence of edges in an image can significantly improve the resulting interpolated image. Many techniques modify the interpolation method in the presence of edges to avoid common artifacts such as blurring, blocking and ringing. We propose to improve upon this idea by fusing the best features of several interpolation methods. This can be done using a novel region classification algorithm to determine which method is best suited to a particular region. From this information, we can use image mosaic techniques to fuse the methods into a result that contains both sharp edges and detailed textures. Nickolaus Mueller, Truong Q. Nguyen |
ICIP | 2 |
| 2008 | A model for image patch-based algorithmsabstractAn empirical study of the domain of patch-based learning algorithms for image and video processing is conducted. As patch-based algorithms are commonly used, knowledge of the properties of fixed size image patches would prove particularly useful and interesting. We are concerned with investigating the overall distribution of vectorized patches of general images. A multivariate distribution model is proposed and analyzed using various techniques, which include univariate histograms and modified k-nearest neighbors. The model is verified and an application using the distribution model is introduced and compared. Karl S. Ni, Truong Q. Nguyen |
ICIP | 2 |
| 2008 | A block-based super-resolution for video sequencesabstractAn algorithm for video resolution enhancement is presented. The approach borrows from previous methods for still-image super- resolution, introducing modifications better suited for the characteristics specific to video domain problems. Each high-resolution (HR) frame is determined through a series of MMSE spatial interpolations based on the local features (statistics) of the frame. Cross-frame registration is estimated externally and the reconstruction algorithm does not limit the form of the motion model, unlike previous data-fusion/deconvolution approaches which have required motion models that do not alter the point-spread function (i.e., motion/blur commutability). This feature is made possible using a reverse motion model mapping the locations of desired HR pixels onto their corresponding locations in the observation frames. An ticipating the existence of registration error found in typical video sequences, the algorithm also provides an internal validation of the observation pixels, helping to reduce significant mis-registration artifacts. An arbitrary enhancement factor can be used, allowing an output at any desired resolution. Interpolation and deblurring are incorporated as a single MMSE filtering operation, providing a non-iterative one-step reconstruction process. Experimental results demonstrating the capabilities of the algorithm are made available. Ryan Strong Prendergast, Truong Q. Nguyen |
ICIP | 2 |
| 2008 | Display dependent coding for 3D video on automultiscopic displaysabstractA method to adaptively code multiview videos has been proposed which uses the depth characteristics of automultiscopic multiview displays. It is found that for the 3D scene seen on multiview displays, regions appearing at large depths are rendered blurry. The proposed method identifies such regions and uses fewer bits to code them. Also, greater number of bits are used for regions which appear sharp on the 3D displays. The overall quality is better than regular AVC/H.264. For compatibility with the framework of scalable multiview coding, we introduce depth scalability. This ensures that the (hierarchical layerwise) encoded video bitstream can be optimally displayed on various displays with different depth properties. Vikas Ramachandra, Matthias Zwicker, Truong Q. Nguyen |
ICIP | 3 |
| 2008 | A new combination of 1D and 2D filter banks for effective multiresolution image representationabstractIn this paper, an effective multiresolution image representation using the combination of 2D quincunx filter bank (FB) and directional wavelet transform (WT) is presented. The proposed method yields simple implementation and low calculation costs compared to the other 1D and 2D FB combinations or adaptive directional WTs. Furthermore, it is a nonredundant transform and realizes quad-tree like multiresolution representation. In applications on nonlinear approximation and image coding, the proposed filter bank shows visual quality improvements and has higher PSNR. Yuichi Tanaka 0001, Masaaki Ikehara, Truong Q. Nguyen |
ICIP | 3 |
| 2008 | Directional motion-compensated spatio-temporal fuzzy filtering for quality enhancement of compressed video sequencesabstractWe propose a directional fuzzy filter to enhance the quality of compressed video sequences. A motion compensated spatio-temporal filter (MCSTF) is applied to intra-frame and inter-frame pixels to reduce both spatial and temporal artifacts. The MCSTF is adaptive to the pixel location, the intensity and position of surrounding pixels. To evaluate the flickering artifacts, a new metric is proposed to take into account the tracking characteristic of human eyes. Simulations on H.264 compressed videos show improvement in artifact reduction of the proposed directional fuzzy filter over other conventional spatial or temporal filtering approaches. Dung Trung Vo, Truong Q. Nguyen |
ICIP | 2 |
| 2008 | Edge-based directional fuzzy filter for compression artifact reduction in JPEG imagesabstractWe propose a novel method to reduce both blocking and ringing artifacts in compressed images and videos. Based on the directional characteristics of ringing artifacts along edges, we use a directional fuzzy filter which is adaptive to the direction of the ringing artifact. The filter exploits the spatial order, the rank order and the spread information of the signal together with the position of the pixels to enhance the quality of the compressed image. Simulations results on compressed images and videos having simple and complex edges show the improvement of the proposed directional fuzzy filter over the conventional fuzzy filtering and other approaches. Dung Trung Vo, Truong Q. Nguyen, Sehoon Yea, Anthony Vetro |
ICIP | 2 |
| 2008 | Oversampled linear-phase perfect reconstruction filter banks with higher-order feasible building blocks: Structure and parameterizationabstractThis paper proposes new building blocks for the lattice structure of oversampled linear-phase perfect reconstruction filter banks (OLPPRFBs). The structure is an extended version of higher-order feasible building blocks for critically-sampled LPPRFBs. It uses fewer number of building blocks and design parameters than those of traditional OLPPRFBs, whereas the frequency characteristic of the new OLPPRFB is comparable to that of traditional one. Yuichi Tanaka 0001, Masaaki Ikehara, Truong Q. Nguyen |
ISCAS | 3 |
| 2008 | Motion vector processing using bidirectional frame difference in motion compensated frame interpolationabstractIn this paper, we address the potential issues in bidirectional motion compensated frame interpolation when the received motion vector field is directly used. Based on this motion vector analysis, we therefore propose using bidirectional motion vector processing method to eliminate the ghost artifacts around the moving objects. The artifacts caused by large motion vector magnitude can be effectively removed by correcting motion vectors bidirectionally. Moreover, since the proposed motion vector selection process chooses the best motion vector from adjacent motion vectors, the computational complexity can be greatly reduced comparing to motion estimation. We also demonstrate how to obtain clearer object edges by allowing displacement adjustment during the bidirectional motion vector processing. Experimental results show that the proposed algorithm outperforms other methods in terms of visual quality. Ai-Mei Huang, Truong Q. Nguyen |
WOWMOM | 2 |
| 2008 | Analysis and Efficient Architecture Design for VC-1 Overlap Smoothing and In-Loop Deblocking FilterabstractIn contrast to the macroblock-based in-loop deblocking filters, the filters of VC-1 perform all horizontal edges (for in-loop filtering) or vertical edges (for overlap smoothing) first and then the vertical edges (for in-loop filtering) or horizontal edges (for overlap smoothing) within a frame, field, or slice. These two filters of VC-1 perform filtering operations on many edges among reconstructed blocks in different processing orders. The entire procedure is very time-consuming and involves high memory access loading for the whole system. This paper analyzes the behavior of VC-1 filters and presents several efficient methods and an integrated architecture design, which involves an overlapped 12$\times$12 block that combines overlap smoothing with in-loop filtering for performance and cost by sharing circuits and system resources. In order to go a step further to efficiently utilize system resources, this paper also presents two other efficient methods, multiple processing order and modified chrominance processing order, which greatly reduce external memory cycles and on-chip memory size for filtered and temporal reconstructed pixels. The specification of the proposed architecture implemented with TSMC 90-nm multithreshold voltage technique has capability to process HDTV1080p 30-fps video and HDTV 2048$\times$1536 24-fps video at 200 MHz. The same concept is applicable to other video processing algorithms, especially in deblocking filter for video post-processing in a frame-based order. Yen-Lin Lee, Truong Q. Nguyen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Quality Enhancement for Motion JPEG Using Temporal RedundanciesabstractThe paper proposes a pixel-based post-processing algorithm to enhance the quality of motion JPEG (MJPEG) by exploiting the temporal redundancies of the decoded frames. The technique permits reconstruction of the high frequency coefficients lost during quantization, thereby reducing ringing artifacts. Based on the linearization of the quantization function, the error between the estimated and original coefficients is analyzed for both cases of ideal and real video sequences. Blocking artifact reduction is verified by a reduction in the variance of this coefficient error. The condition of valid motion vectors to get quality improvement is considered based on these errors. The algorithm is also extended to find the optimal filter for a general estimation scheme based on an arbitrary number of frames. Results in visual and peak signal-to-noise ratio improvement using both integer and subpixel motion vectors are verified by simulations on video sequences. Dung Trung Vo, Truong Q. Nguyen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Iterative Rate-Distortion Optimization of H.264 With Constant Bit Rate ConstraintabstractIn this paper, we apply the primal-dual decomposition and subgradient projection methods to solve the rate-distortion optimization problem with the constant bit rate constraint. The primal decomposition method enables spatial or temporal prediction dependency within a group of picture (GOP) to be processed in the master primal problem. As a result, we can apply the dual decomposition to minimize independently the Lagrangian cost of all the MBs using the reference software model of H.264. Furthermore, the optimal Lagrange multiplier lambda* is iteratively derived from the solution of the dual problem. As an example, we derive the optimal bit allocation condition with the consideration of temporal prediction dependency among the pictures. Experimental results show that the proposed method achieves better performance than the reference software model of H.264 with rate control. Cheolhong An, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2008 | Resource Allocation for Error Resilient Video Coding Over AWGN Using Optimization ApproachabstractThe number of slices for error resilient video coding is jointly optimized with 802.11a-like media access control and the physical layers with automatic repeat request and rate compatible punctured convolutional code over additive white gaussian noise channel as well as channel times allocation for time division multiple access. For error resilient video coding, the relation between the number of slices and coding efficiency is analyzed and formulated as a mathematical model. It is applied for the joint optimization problem, and the problem is solved by a convex optimization method such as the primal-dual decomposition method. We compare the performance of a video communication system which uses the optimal number of slices with one that codes a picture as one slice. From numerical examples, end-to-end distortion of utility functions can be significantly reduced with the optimal slices of a picture especially at low signal-to-noise ratio. Cheolhong An, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2008 | LCD Motion Blur Reduction: A Signal Processing ApproachabstractLiquid crystal displays (LCDs) have shown great promise in the consumer market for their use as both computer and television displays. Despite their many advantages, the inherent sample-and-hold nature of LCD image formation results in a phenomenon known as motion blur. In this work, we develop a method for motion blur reduction using the Richardson-Lucy deconvolution algorithm in concert with motion vector information from the scene. We further refine our approach by introducing a perceptual significance metric that allows us to weight the amount of processing performed on different regions in the image. In addition, we analyze the role of motion vector errors in the quality of our resulting image. Perceptual tests indicate that our algorithm reduces the amount of perceivable motion blur in LCDs. Shay Har-Noy, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2008 | A Multistage Motion Vector Processing Method for Motion-Compensated Frame InterpolationabstractIn this paper, a novel, low-complexity motion vector processing algorithm at the decoder is proposed for motion-compensated frame interpolation or frame rate up-conversion. We address the problems of having broken edges and deformed structures in an interpolated frame by hierarchically refining motion vectors on different block sizes. Our method explicitly considers the reliability of each received motion vector and has the capability of preserving the structure information. This is achieved by analyzing the distribution of residual energies and effectively merging blocks that have unreliable motion vectors. The motion vector reliability information is also used as a prior knowledge in motion vector refinement using a constrained vector median filter to avoid choosing identical unreliable one. We also propose using chrominance information in our method. Experimental results show that the proposed scheme has better visual quality and is also robust, even in video sequences with complex scenes and fast motion. Ai-Mei Huang, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2008 | A Fully Scalable Motion Model for Scalable Video CodingabstractMotion information scalability is an important requirement for a fully scalable video codec, especially for decoding scenarios of low bit rate or small image size. So far, several scalable coding techniques on motion information have been proposed, including progressive motion vector precision coding and motion vector field layered coding. However, it is still vague on the required functionalities of motion scalability and how it collaborates flawlessly with other scalabilities, such as spatial, temporal, and quality, in a scalable video codec. In this paper, we first define the functionalities required for motion scalability. Based on these requirements, a fully scalable motion model is proposed along with tailored encoding techniques to minimize the coding overhead of scalability. Moreover, the associated rate distortion optimized motion estimation algorithm will be provided to achieve better efficiency throughout various decoding scenarios. Simulation results will be presented to verify the superiorities of proposed scalable motion model over nonscalable ones. Meng-Ping Kao, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2008 | Markov Random Field Model-Based Edge-Directed Image InterpolationabstractThis paper presents an edge-directed image interpolation algorithm. In the proposed algorithm, the edge directions are implicitly estimated with a statistical-based approach. In opposite to explicit edge directions, the local edge directions are indicated by length-16 weighting vectors. Implicitly, the weighting vectors are used to formulate geometric regularity (GR) constraint (smoothness along edges and sharpness across edges) and the GR constraint is imposed on the interpolated image through the Markov random field (MRF) model. Furthermore, under the maximum a posteriori-MRF framework, the desired interpolated image corresponds to the minimal energy state of a 2-D random field given the low-resolution image. Simulated annealing methods are used to search for the minimal energy state from the state space. To lower the computational complexity of MRF, a single-pass implementation is designed, which performs nearly as well as the iterative optimization. Simulation results show that the proposed MRF model-based edge-directed interpolation method produces edges with strong geometric regularity. Compared to traditional methods and other edge-directed interpolation methods, the proposed method improves the subjective quality of the interpolated edges while maintaining a high PSNR level. Min Li 0014, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2008 | Resource Allocation for TDMA Video Commtmmunication Over AWGN Using Cross-Layer Optimization ApproachabstractCross-layer optimization approach is applied to allocate channel times of time-division multiple access (TDMA) to utility functions which send different video streams with different rate-distortion characteristics. Given a channel time, each utility function solves an optimization problem to obtain the optimal source code rate, channel code rate and media access control (MAC) frame size. In this paper, we derive mathematical models to represent end-to-end distortion of video streams and 802.11a-like MAC and PHY with automatic repeat reQuest (ARQ) and rate compatible punctured convolutional code (RCPC) over additive white Gaussian noise (AWGN) channel. These models are formulated as a convex optimization problem, and it is solved by the primal-dual decomposition method. In addition, coexistence among the proposed utility functions and conventional utility functions is discussed. Finally, numerical examples show that significant reduction of overall distortion of utility functions can be achieved. Cheolhong An, Truong Q. Nguyen |
IEEE Trans. Multim. | 2 |
| 2007 | Cross-Layer Optimization for Video Communication Over AWGN ChannelabstractIn this paper, cross-layer optimization approach is applied from video coding to the physical layer through media access control (MAC) layer. We mainly focus on resource allocation of source code rate and channel code rate with optimal MAC frame length. Therefore, elaborate mathematical models are derived to represent end-to-end distortion of video streams and 802.1 la-like MAC and PHY with automatic repeat request (ARQ) and rate compatible punctured convolutional code (RCPC) over additive white Gaussian noise (AWGN) channel. These models are formulated as a convex optimization problem, and it is solved by the dual decomposition method. From a numerical example, significant reduction of overall distortion can be achieved. Cheolhong An, Truong Q. Nguyen |
GLOBECOM | 2 |
| 2007 | Integer FFT with Optimized Coefficient SetsabstractIn this paper, the principle of finding the optimized coefficient set of integer fast Fourier transform (IntFFT) is introduced. IntFFT has been regarded as an approximation of original FFT since it utilizes lifting scheme (LS) and decomposes the complex multiplication of twiddle factor into three lifting steps. Based on the observation of the quantization loss model of lifting operations, we can select an optimized coefficient set and achieve better signal-to-quantization-noise ratio (SQNR). A mixed-radix 128-point FFT is used to compare the SQNR performance between IntFFT and other FFT implementations. A fixed-point simulation environment with the presence of additive white Gaussian noise (AWGN) channel is also constructed for comparison purposes. Wei-Hsin Chang, Truong Q. Nguyen |
ICASSP (2) | 2 |
| 2007 | Frequency Selective KYP Lemma and its Applications to IIR Filter Bank DesignabstractFor a transfer function/filter F(ejω) of order n, Kalman-Yakubovich-Popov (KYP) lemma characterizes the intractable semi-infinite programming (SIP) condition F(e-jω)1 Θ [F(ejω)1]T≥ 0 ∀ ω in frequency domain by a tractable semi-definite programming (SDP) in state-space domain. Some recent results generalize this lemma to SDP for SIP of frequency selectivity (FS-SIP). All these SDP characterizations are given at the expense of the introduced Lyapunov matrix variable of dimension n × n, making them impractical for high order problem. Moreover, the existing SDP characterizations for FS-SIP do not allow to formulate synthesis/design problems as SDPs. In this paper, we propose a completely new SDP characterization of general FS-SIP, which is of moderate size and is free from Lyapunov variables. Extensive examples are provided to validate the effectiveness of our result. Hung Gia Hoang, Hoang Duong Tuan, Truong Q. Nguyen |
ICASSP (3) | 3 |
| 2007 | Design of Half-Band Diamond and Fan Filters by SDPabstractA new design method for linear phase half-band diamond (DS) and fan-shaped (FS) 2-D filters is proposed. A general formulation for frequency mask constraints in different shapes using 2-D trigonometric curves is developed. This facilitates semi-definite programming (SDP) of moderate dimension for the design problem. Several examples are included to illustrate advantages of our method. T. Q. Hung, Hoang Duong Tuan, Truong Q. Nguyen |
ICASSP (3) | 3 |
| 2007 | An Efficient SDP Based Design for Prototype Filters of M-Channel Cosine-Modulated Filter BanksabstractThe paper presents an efficient semidefinite programming (SDP) based design for prototype filters of cosine-modulated filter banks (CMFBs). We consider a class of near-perfect reconstruction CMFBs with the linear phase prototype filter, which structurally eliminates the amplitude overall distortion. The prototype filter design problem is then formulated into a convex semi-infinite programming problem. Furthermore, to handle the semi-infinite constraints, we use the linear matrix inequality (LMI) characterization of positive trigonometric polynomials to cast the semi-infinite programming problem into SDP one. Finally, convex duality is applied to transform the SDP into another SDP with the minimal number of additional variables, which is efficiently solved. An additional advantage of the proposed method is that we can precisely control the filter specifications. Ha Hoang Kha, Hoang Duong Tuan, Truong Q. Nguyen |
ICASSP (3) | 3 |
| 2007 | Kernel Resolution Synthesis for SuperresolutionabstractThis work considers a combination classification-regression based framework with the proposal of using learned kernels in modified support vector regression to provide superresolution. The usage of both generative and discriminative learning techniques is examined first by assuming a distribution for image content for classification and then providing regression via semi-definite programming (SDP) and quadratically constrained quadratic programming (QCQP) problems. The advantage of the proposed method over other learning-based superresolution algorithms include reduced problem complexity, specificity with regard to image content, added degrees of freedom from the nonlinear approach, and excellent generalization that a combined methodology has over its individual counterparts. Karl S. Ni, Truong Q. Nguyen |
ICASSP (1) | 2 |
| 2007 | Design of Oversampled Filter Bank Based Widely Linear OQPSK EqualizerabstractIn this paper, the design and analysis of multirate filter bank (FB) based OQPSK widely linear (WL) oversampled equalizers for frequency selectively channels (FSCs) are studied. We focus on OQPSK signal transmissions in FSCs with various amount of fractionally spaced equalization using WL processing under the framework of FB. We explicitly represent a OQPSK communications system as a FB which allows the study of conditions for perfect reconstruction using a finite length polynomial equalizer. Furthermore, WL FB based MMSE oversampled equalizer designs operating at various sampling rates are presented. Computer simulations are simulated to study the performance of various types of equalizers in FSCs. Ka Shun Carson Pun, Truong Q. Nguyen |
ICASSP (3) | 2 |
| 2007 | Entropy of General Gaussian Distributions and MIMO Channel Capacity Maximizing Precoder and DecoderabstractExploiting channel state information at the transmitter and receiver to design an optimal linear precoder and decoder for a multiple-input multiple-output (MIMO) communication system is an active research area. The design is often based on the information rate criterion, that is to design the precoder and decoder such that the system capacity is maximized, subject to the average transmit power constraint. Although such an optimization problem has been considered intensively and there have been numerous proposals so far, they are not rigorously correct. We propose a mathematically rigorous framework for solving this optimization problem. Our proposed solution is applicable to both MIMO flat fading and frequency selective fading channels. Simulations verify the theoretical analysis. Hoang Duong Tuan, Duong-Hung Pham, Ba-Ngu Vo, Truong Q. Nguyen |
ICASSP (3) | 4 |
| 2007 | Analysis of Utility Functions for VideoabstractIn this paper, we formulate the utility functions of distortion and peak signal-to-noise ratio (PSNR) which are generally used for the performance evaluation of video coding applications as the utility functions. The convexity property of these utility functions is analyzed in the original domain and transformed domain of an optimization variable. From this analysis, we derive joint optimization scheme with congestion control through the utility matching between TCP layer and video coding layer. Experimental results show that the overall PSNR increases and the variation of quality among the utility functions is reduced. Cheolhong An, Truong Q. Nguyen |
ICIP (5) | 2 |
| 2007 | A Novel Multi-Stage Motion Vector Processing Method for Motion Compensated Frame InterpolationabstractIn this paper, a novel multi-stage motion vector processing algorithm at the decoder is proposed for motion compensated frame interpolation. We address the problems of discontinuous edges and deformed structures in an interpolated frame by explicitly considering reliability of each received motion vector. By hierarchically refining motion vectors with different block sizes, the proposed method is capable of preserving structure information. Experimental results show that the proposed scheme outperforms other methods in terms of visual quality and PSNR, and it is also robust when video sequences have complex scenes and fast motion. Ai-Mei Huang, Truong Q. Nguyen |
ICIP (5) | 2 |
| 2007 | A Fully Scalable Motion Model for Scalable Video CodingabstractMotion information scalability is an important requirement for a fully scalable video codec, especially in low bit rate or small resolution decoding scenarios. So far, several layered coding and motion vector precision scalability approaches have been proposed. However, it is still vague on the required functionalities of a fully scalable motion model and how it interacts with other scalabilities, such as spatial, temporal and quality, in a scalable video codec. In this paper, we first define the functionalities required by a fully scalable motion model. Based on these requirements, a fully scalable motion model will be proposed and, moreover, the associated rate distortion optimized estimation techniques will also be provided as a companion. Several simulation results will be presented to summarize the advantages of proposed motion model. Meng-Ping Kao, Truong Q. Nguyen |
ICIP (2) | 2 |
| 2007 | Analysis and Integrated Architecture Design for Overlap Smooth and in-Loop Deblocking Filter in VC-1abstractUnlike familiar macroblock-based in-loop deblocking filter in H.264, the filters of VC-1 perform all horizontal edges (for in-loop deblocking filtering) or vertical edges (for overlap smoothing) first and then the other directional filtering edges. The entire procedure is very time-consuming and with high memory access loading for the whole system. This paper presents a novel method and the efficient integrated architecture design, which involves an 12 times 12 overlapped block that combines overlap smoothing with loop filtering for performance and cost by sharing circuits and resources. This architecture has capability to process HDTV1080p 30 fps video and HDTV 2048 times 1536 24 fps video at 180 MHz. The same concept is applicable to other video processing algorithms, especially in deblocking filter for video post-processing in a frame-based order. Yen-Lin Lee, Truong Q. Nguyen |
ICIP (5) | 2 |
| 2007 | Markov Random Field Model-Based Edge-Directed Image InterpolationabstractThis paper presents an edge-directed image interpolation algorithm. In the proposed algorithm, the edge directions are implicitly estimated with a statistical-based approach. Consequently, the local edge directions are represented by length-16 vectors, which are denoted as weight vectors. The weight vectors are used to formulate geometric regularity constraint, which is imposed on the interpolated image through the Markov Random Field (MRF) model. Furthermore, the interpolation problem is formulated as a Maximum A Posterior (MAP)-MRF problem and, under the MAP-MRF framework, the desired interpolated image corresponds to the minimal energy state of a two-dimensional random held. Simulated Annealing method is used to search for the minimal energy state from a reasonable large state space. Simulation and comparison results show that the proposed MRF model-based edge-directed interpolation method produces edges with strong geometric regularity. Min Li 0014, Truong Q. Nguyen |
ICIP (2) | 2 |
| 2007 | Color Image Superresolution Based on a Stochastic Combinational Classification-Regression AlgorithmabstractThe proposed algorithm in this work provides superresolution for color images by using a learning based technique that utilizes both generative and discriminant approaches. The combination of the two approaches is designed with a stochastic classification-regression framework where a color image patch is first classified by its content, and then, based on the class of the patch, a learned regression provides the optimal solution. For good generalization, the classification portion of the algorithm determines the probability that the image patch is in a given class by modeling all possible image content (learned through a training set) as a Gaussian mixture, with each Gaussian of the mixture portraying a single class. The regression portion of the algorithm has been chosen to be a modified Support Vector Regression, where the kernel has been learned by solving a semi definite programming (SDP) and quadratically constrained quadratic programming (QCQP) problem. The SVR is further modified by scaling the training points in the SDP and QCQP problems by their relevance and importance to the examined regression. The result is a weighted average of different regressions depending on how much a single regression is likely to contribute, where advantages include reduced problem complexity, specificity with regard to image content, added degrees of freedom from a nonlinear approaches, and excellent generalization that a combined methodology has over its individual counterparts. Karl S. Ni, Truong Q. Nguyen |
ICIP (2) | 2 |
| 2007 | Unequal Length First-Order Linear-Phase Filter Banks for Efficient Image CodingabstractIn this paper, we present the structure and design method for a first-order linear-phase filter bank (FOLPFB) which has unequal filter lengths in its synthesis bank (UFLPFB). A FOLPFB is a generalized version of biorthogonal LPFBs regarding their synthesis filter lengths. Ringing artifact is the main disadvantage of image coding based on FOLPFBs. UFLPFBs can reduce the ringing artifacts as well as approximate smooth regions well. Yuichi Tanaka 0001, Masaaki Ikehara, Truong Q. Nguyen |
ICIP (4) | 3 |
| 2007 | Quality Enhancement for Motion JPEG using Temporal RedundanciesabstractWe propose a post processing algorithm to enhance the quality of Motion JPEG (MJPEG) by exploiting temporal redundancies. The error between the estimated and original blocks is analyzed using the translational relation in the discrete cosine transform (DCT) domain. The proposed algorithm permits reconstructing the high frequency coefficients lost during quantization and therefore reduces the ringing artifact. Blocking artifact reduction is verified by the decrease in the variance of this coefficient error. Results in quality and PSNR improvement for both cases of integer and sub-pixel motion vectors are verified by simulations on video sequences. Dung Trung Vo, Truong Q. Nguyen |
ICIP (4) | 2 |
| 2007 | A New Objective Quality Metric for Frame Interpolation using in Video CompressionabstractThis paper discusses the disadvantages of several existing objective quality evaluation methods for frame interpolation techniques. Samples show that these disadvantages lead the objective quality measurement inconsistent with what humans perceive. Based on these observations, a new metric designed to evaluate the performance of frame interpolation techniques is proposed. This metric combines the severity of interpolation artifacts and several human visual factors into a single quality score. Final implementation shows that the proposed metric out-performs other commonly used metrics. Kai-Chieh Yang, Ai-Mei Huang, Truong Q. Nguyen, Clark C. Guest, Pankaj K. Das |
ICIP (2) | 3 |
| 2007 | Design of Diamond and Circular Filters by Semi-definite ProgrammingabstractA new design for linear phase diamond-shaped (DS) and circular-shaped (CS) 2D filters is developed. First, the frequency masks are efficiently constrained by 2D second-order trigonometric polynomials. Then semi-definite programming (SDP) of reasonably low dimension is employed to effectively express the filter specifications. Several numerical examples are provided to demonstrate the superior performance of our design in comparison with all other existing designs. T. Q. Hung, Hoang Duong Tuan, Truong Q. Nguyen |
ISCAS | 3 |
| 2007 | Design of Cosine-Modulated Pseudo-QMF Banks Using Semidefinite Programming RelaxationabstractThe paper proposes a new approach for the design of M-channel pseudo-quadrature mirror filter (QMF) banks. First, the convex hull of 2Mth band linear phase filters admitting linear phase spectral factors is analytically described by semidefinite programming (SDP). Then, the prototype filter design is cast into an SDP problem, which is efficiently solved. Design examples are presented to illustrate the effectiveness of the proposed method and to evaluate the design performance in comparison with the existing designs. Ha Hoang Kha, Hoang Duong Tuan, Truong Q. Nguyen |
ISCAS | 3 |
| 2007 | Adaptive In-Loop Prediction Refinement for Video CodingabstractModern video compression codecs achieve high compression efficiency by exploiting the temporal and spatial redundancies present in video sequences. Although state of the nit block based intra and inter prediction have been refined to better adapt to the content in the scene, they still exhibit limitations in fully extracting information from the decoded data. We propose the addition of an in-loop prediction refinement stage that is initialized with the conventional intra or inter predicted result and is capable of further reducing the spatial redundancies present in the sequence. By extracting additional information from the decoded neighborhood, the refinement stage is able to bring the prediction closer to the original data, thus improving compression efficiency. In this work, we study one possible method of refinement for the particular case of intra prediction. This uses sparse decomposition of blocks that overlap decoded neighboring blocks and the current predicted block in order to make prediction more coherent with already encoded data. Shay Har-Noy, Òscar Divorra Escoda, Peng Yin 0002, Cristina Gomila, Truong Q. Nguyen |
MMSP | 5 |
| 2007 | A De-Interlacing Algorithm Using Markov Random Field ModelabstractIn this paper, a motion-compensated de-interlacing algorithm using the Markov random field (MRF) model is proposed. The de-interlacing problem is formulated as a maximum a posteriori (MAP) MRF problem. The MAP solution is the one that minimizes an energy function, which imposes discontinuity-adaptive smoothness (DAS) spatial constraint on the de-interlaced frame. The edge direction information, which is used to formulate the DAS constraint, is implicitly indicated by weight vectors (weights for 16 digitized directions). Generally, large weights are assigned to along-edge directions and relatively small weights are assigned to across-edge directions. As a local statistical-based method, the proposed weighting method should be more robust than traditional edge-directed interpolation methods in deciding local edge directions. The proposed algorithm is implemented by an iterative optimization process, which guarantees convergence. However, a global optimal solution is not guaranteed due to computational complexity concern. Simulation results compare the proposed algorithm to other motion compensated de-interlacing algorithms. Significant improvements of de-interlaced edges are observed. Min Li 0014, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2007 | Image Superresolution Using Support Vector RegressionabstractA thorough investigation of the application of support vector regression (SVR) to the superresolution problem is conducted through various frameworks. Prior to the study, the SVR problem is enhanced by finding the optimal kernel. This is done by formulating the kernel learning problem in SVR form as a convex optimization problem, specifically a semi-definite programming (SDP) problem. An additional constraint is added to reduce the SDP to a quadratically constrained quadratic programming (QCQP) problem. After this optimization, investigation of the relevancy of SVR to superresolution proceeds with the possibility of using a single and general support vector regression for all image content, and the results are impressive for small training sets. This idea is improved upon by observing structural properties in the discrete cosine transform (DCT) domain to aid in learning the regression. Further improvement involves a combination of classification and SVR-based techniques, extending works in resolution synthesis. This method, termed kernel resolution synthesis, uses specific regressors for isolated image content to describe the domain through a partitioned look of the vector space, thereby yielding good results. Karl S. Ni, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2006 | Architecture and Performance Analysis of Lossless FFT in OFDM SystemsabstractThe design of fast Fourier transform (FFT) architecture is one of the bottlenecks in the implementation of OFDM systems. With the recent progress on the development of lossless transform, its possible applications in communication systems have received more and more attention. In this paper, a lossless integer FFT (IntFFT) architecture based on radix-22FFT algorithm is analyzed and implemented. By exploring the symmetric property, the overall memory usage is reduced by 27.4% for a 64-point FFT design. The variance of quantization loss for both IntFFT and conventional fixed-point FFT (FxpFFT) is derived. The signal to quantization loss ratio (SQNR) and the bit error rate (BER) performance in system level is also simulated to test the accuracy of IntFFT. Based on the simulation results, IntFFT can yield comparative performance while using less memory usage than FxpFFT designs Wei-Hsin Chang, Truong Q. Nguyen |
ICASSP (3) | 2 |
| 2006 | SDP for 2-d Filter Design: General Formulation and Dimension Reduction TechniquesabstractIn this paper, a new technique for designing linear phase 2-D filter based on semi-definite programming (SDP) is proposed. This approach allows the design of 2-D filters with accurate cut-off frequency, subject to hard bounds on the frequency response to be achieved on a standard computer. Using the notion of 2-D trigonometric curves, we generalize the 2-D trigonometric Markov-Lukacs theorem to identify the pass-band and the stop-band in the region of support. The 2-D filter specifications are expressed as linear matrix inequalities. We also exploit convex duality to derive SDP formulations of reduced dimensions. Numerical examples illustrating the advantages of our method are also presented T. Q. Hung, Hoang Duong Tuan, Ba-Ngu Vo, Truong Q. Nguyen |
ICASSP (2) | 4 |
| 2006 | Symmetric Orthogonal Complex-Valued Filter Bank Design by Semidefinite ProgrammingabstractA new design method for complex-valued two-channel FIR filter banks with both orthogonality and symmetry properties is developed. Based on a novel linear matrix inequality (LMI) characterization of trigonometric curves, the optimal design of the perfect reconstruction filter bank is reformulated as a semi-definite programme. The dimension of the resulting semi-definite programme is further reduced by exploiting the strong convex duality. Consequently, the globally optimal solution can be effectively found for any practical filter length and desired regularity order Ha Hoang Kha, Hoang Duong Tuan, Ba-Ngu Vo, Truong Q. Nguyen |
ICASSP (3) | 4 |
| 2006 | Single Image Superresolution Based on Support Vector RegressionabstractSupport vector machine (SVM) regression is considered for a statistical method of single frame superresolution in both the spatial and discrete cosine transform (DCT) domains. As opposed to current classification techniques, regression allows considerably more freedom in the determination of missing high-resolution information. In addition, since SVM regression approaches the superresolution problem as an estimation problem with a criterion of image correctness rather than visual acceptableness, its optimization results have better mean-squared error. With the addition of structure in the DCT coefficients, DCT domain image superresolution is further improved Karl S. Ni, Sanjeev Kumar 0003, Nuno Vasconcelos, Truong Q. Nguyen |
ICASSP (2) | 4 |
| 2006 | Jointly Optimal Precoding/Postcoding for Colored MIMO SystemsabstractThe problem of designing a jointly optimal linear precoder and decoder for a multiple-input multiple-output (MIMO) channel has received much interest recently. However, most existing works only deal with white input signal. When the input signal is colored, pre-whitening and its inverse operation are often applied prior to precoding and after decoding, respectively. Consequently, the precoder and decoder are no longer optimal with respect to the original colored signal. In this paper, we propose a closed-form solution for optimal linear precoder and decoder for colored input signal. Our approach is based on minimizing the symbol mean squared error under an average output power constraint, and is applicable to both MIMO flat fading and frequency selective fading channels. Simulations show the advantage of our solution over prewhitening-based method. Duong-Hung Pham, Hoang Duong Tuan, Ba-Ngu Vo, Truong Q. Nguyen |
ICASSP (4) | 4 |
| 2006 | A Non-Isotropic Parametric Model for Image SpectraabstractImage restoration and enhancement problems are often considered using a solution based on a model for the image's power spectral density (PSD). This paper considers a new PSD model for images through modification of a previously considered isotropic model. The new model and an algorithm for estimating its parameters are presented. Several simulations considering PSD estimation for images corrupted by additive noise, distortion, and resolution reduction are presented to demonstrate the improvements made with this new modelling approach Ryan Strong Prendergast, Truong Q. Nguyen |
ICASSP (2) | 2 |
| 2006 | A Deconvolution Method for LCD Motion Blur ReductionabstractLiquid crystal displays (LCD) have shown great promise in the consumer market for their use as both computer and television displays. Despite their many advantages, the inherent nature of LCD image formation results in a phenomenon known as motion blur. LCDs emit light and aim to hold it constant during the entire frame time. In this paper we develop a method for motion blur reduction using the Richardson-Lucy deconvolution algorithm in concert with motion vector information from the scene. In addition, we derive a general lower bound on the performance of the deconvolution algorithm. Perceptual tests indicate that our algorithm is indeed effective at reducing the amount of motion blur on LCDs. Shay Har-Noy, Truong Q. Nguyen |
ICIP | 2 |
| 2006 | Motion Vector Processing Based on Residual Energy Information for Motion Compensated Frame InterpolationabstractIn this paper, we address the problem of motion compensated frame interpolation (MCFI) at the decoder by analyzing received motion vectors (MVs) and residual error. In the proposed method, the skipped frames are generated at the decoder by using received information from the bitstream. We combine a hierarchical motion vector correction algorithm and a residual-energy constrained median filter to obtain true motion. Experimental results show that the proposed motion vector processing algorithm improves the visual quality of interpolated frames, especially in the motion transition area and the motion boundary. Ai-Mei Huang, Truong Q. Nguyen |
ICIP | 2 |
| 2006 | Compression Artifact Reduction using Support Vector RegressionabstractIn this paper, we propose a compression artifact reduction algorithm based on ν support vector regression. It belongs to the broad family of regularized reconstruction methods but regularization model is learned from a set of training samples of original images and corresponding noise corrupted version. As opposed to artifact reduction methods specific to each type of compression artifact (e.g. blocking, ringing etc), we treat such different artifacts as symptoms of the same problem, quantization of DCT coefficients. In the testing step, algorithm tries to undo the effect of quantization using information (relationship between original and artifact-corrupted image) learned during the training step. Experimental results exhibit significant reduction in all types of compression artifacts. Sanjeev Kumar 0003, Truong Q. Nguyen, Mainak Biswas |
ICIP | 2 |
| 2006 | Discontinuity-Adaptive De-Interlacing Scheme Using Markov Random Field ModelabstractIn this paper, a de-interlacing algorithm to find the optimal deinterlaced results given accuracy-limited motion information is proposed. The de-interlacing process is formulated as a maximum a posteriori (MAP)-Markov random field (MRF) problem. The MAP solution is the one that minimizes an energy function. The energy function imposes discontinuity adaptive smoothness constraint upon the deinterlaced frame. Simulation results show that the MAP-MRF formulation is efficient and the high frequency noise is removed in a few iterations. Min Li 0014, Truong Q. Nguyen |
ICIP | 2 |
| 2006 | Reverse, Sub-Pixel Block Matching: Applications within H.264 and Analysis of LimitationsabstractIn this paper the concept of reverse block matching is introduced based on the relativity of motion, especially, the translational motion. As one of the main applications for this technique, sub-pixel motion estimation is revisited to achieve faster, yet efficient video coding when it is to be deployed on a system with limited memory and data caching requirements. Analysis of direct (conventional) and reverse (proposed) sub-pixel motion estimation at the encoder is provided to identify the sources of mismatch between these two techniques. Finally, experimental results are provided for sub-pixel motion estimation, using the reverse block matching scheme within the H.264 encoder. These results show a significant increase in speed and a reduction of required memory at a reasonable cost, in terms of PSNR degradation, when compared to the fastest direct sub-pixel motion estimation methods currently implemented in the reference H.264 codec by HHI. Koohyar Minoo, Truong Q. Nguyen |
ICIP | 2 |
| 2006 | A Novel Motion Compensated Frame Interpolation Based on Block-Merging and Residual EnergyabstractIn this paper, a novel motion compensated frame interpolation (MCFI) algorithm by merging blocks that have unreliable motion vectors (MVs) based on their residual errors is proposed. Unlike the conventional methods that find true motion using smaller blocks and vector median filter, we proposed to find one single motion vector to represent a group of adjacent macroblocks (MBs) where the conventional MCFI methods are likely to fail, likely to fail. The proposed method is able to preserve the structure of different objects and their edge information, without requiring complicated edge detection and object-based motion estimation. Experimental results show that the proposed scheme improves both visual quality and PSNR, especially in the areas with different motions and the motion boundary Ai-Mei Huang, Truong Q. Nguyen |
MMSP | 2 |
| 2006 | Learning the Kernel Matrix for SuperresolutionabstractThis paper proposes the application of learned kernels in support vector regression to superresolution in the discrete cosine transform (DCT) domain. Though previous works involve kernel learning, their problem formulation is examined to reformulate the semi-definite programming problem of finding the optimal kernel matrix. For the particular application to superresolution, downsampling properties derived in the DCT domain are exploited to add structure to the learning algorithm. The advantage of the proposed method over other learning-based superresolution algorithms include specificity with regard to image content, structured consideration of energy compaction, and the added degrees of freedom that regression has over classification-based algorithms Karl S. Ni, Sanjeev Kumar 0003, Truong Q. Nguyen |
MMSP | 3 |
| 2006 | Performance analysis of motion-compensated de-interlacing systemsabstractA lot of research has been conducted on motion-compensated (MC) de-interlacing, but there are very few publications that discuss the performances of de-interlacing quantatively. The various methods are compared through their performance on known video sequences. Linear system analysis of interlaced video and de-interlacer are proposed in. It is well established that the performance of the MC methods outperform the fixed or motion-adaptive methods when the motion vectors used are reliable and true to the scene content. Being an open-loop process the performance of the MC de-interlacers degrade drastically when there are motion vector errors. In this paper, a linear system analysis of MC video upconversion systems is presented and the effects of motion vector accuracy on system performance are analyzed. We investigate the various factors that contribute to the motion vector inaccuracy, such as incorrect motion modelling, acceleration between the frames, and insufficient interpolation kernel. Mainak Biswas, Sanjeev Kumar 0003, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2006 | Optimal temporal interpolation filter for motion-compensated frame rate up conversionabstractFrame rate up conversion (FRUC) methods that employ motion have been proven to provide better image quality compared to nonmotion-based methods. While motion-based methods improve the quality of interpolation, artifacts are introduced in the presence of incorrect motion vectors. In this paper, we study the design problem of optimal temporal interpolation filter for motion-compensated FRUC (MC-FRUC). The optimal filter is obtained by minimizing the prediction error variance between the original frame and the interpolated frame. In FRUC applications, the original frame that is skipped is not available at the decoder, so models for the power spectral density of the original signal and prediction error are used to formulate the problem. The closed-form solution for the filter is obtained by Lagrange multipliers and statistical motion vector error modeling. The effect of motion vector errors on resulting optimal filters and prediction error is analyzed. The performance of the optimal filter is compared to nonadaptive temporal averaging filters by using two different motion vector reliability measures. The results confirm that to improve the quality of temporal interpolation in MC, the interpolation filter should be designed based on the reliability of motion vectors and the statistics of the MC prediction error. Gökçe Dane, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2005 | Optimal filter bank reconstruction of periodically undersampled signalsabstractA multirate filter bank model is considered for reconstruction of periodically sampled signals. In contrast to many previous methods which considered perfect reconstruction of deterministic signals, this approach uses a known discrete-time cyclostationary signal model to find a minimum mean-squared error reconstruction solution. A primary advantage of this approach is that it does not require a minimum sampling density, allowing optimal solutions to be determined in cases of undersampling. This allows for consideration of a wide range of generalized sampling problems under a single framework. An example, with simulation results, is presented. Ryan Strong Prendergast, Truong Q. Nguyen |
ICASSP (4) | 2 |
| 2005 | LMI characterization for the convex hull of trigonometric curves and applicationsabstractIn this paper, we develop a new linear matrix inequality (LMI) technique, which is practical for solutions of the general trigonometric semi-infinite linear constraint (TSIC) of competitive orders. Based on the new full LMI characterization for the convex hull of a trigonometric curve, it is shown that the semi-infinite optimization problem involving TSIC can be solved by an LMI optimization problem with additional variables of dimension just n, the order of the trigonometric curve. Our solution method is very robust which allows us to address almost all practical filter design problems. Unlike most previous works involving several complex mathematical tools, our derivation arguments are based on simple results of the convex analysis and some formal elementary transforms. Furthermore, many filter/filterbank design problems can be reformulated as the optimization of linear/convex quadratic objectives over the TSIC. Based on this reformulation, these problems can be equivalently reduced to LMI optimization problems with the minimal size. Our examples of designing up to 1200-tap filters verifies the viability of our formulation. Hoang Duong Tuan, Tran Thai Son, Ba-Ngu Vo, Truong Q. Nguyen |
ICASSP (4) | 4 |
| 2005 | Spatio-temporal texture synthesis and image inpainting for video applicationsabstractIn this paper we investigate the application of texture synthesis and image inpainting techniques for video applications. Working in the non-parametric framework, we use 3D patches for matching and copying. This ensures temporal continuity to some extent which is not possible to obtain by working with individual frames. Since, in present application, patches might contain arbitrary shaped and multiple disconnected holes, fast Fourier transform (FFT) and summed area table based sum of squared difference (SSD) calculation (S.L. Kilthau, et al, 2002) cannot be used. We propose a modification of above scheme which allows its use in present application. This results in significant gain of efficiency since search space is typically huge for video applications. Sanjeev Kumar 0003, Mainak Biswas, Serge J. Belongie, Truong Q. Nguyen |
ICIP (2) | 4 |
| 2005 | Optimal wavelet filter design in scalable video codingabstractThe design of a class of wavelet filters and its application in scalable video coding (SVC) is discussed in detail in this paper. The design method uses maximal flat wavelet filters as prototype filters and incorporates all other wavelet filter design requirements. The designed wavelet filters are optimal in a sense that best tradeoff between high stopband attenuation of analysis lowpass filter H/sub 0/(z) and flat passband response of synthesis lowpass filter F/sub 0/(z) is achieved. The simulation that compares the performances of the designed filters and Daubechies (9,7) filters in SVC are illustrated. Min Li 0014, Truong Q. Nguyen |
ICIP (1) | 2 |
| 2005 | Improving frequency domain super-resolution via undersampling modelabstractThe super-resolution problem is considered using a mean-squared error minimizing solution of a generalized undersampling model in this non-iterative frequency domain approach. While previous frequency domain approaches have been based on a bandlimited image model, this approach uses a non-bandlimited stationary spectral model. This allows improved reconstruction of certain image features. The model and algorithm are presented along with an example Ryan Strong Prendergast, Truong Q. Nguyen |
ICIP (1) | 2 |
| 2004 | Linear system analysis of motion compensated de-interlacingabstractThe majority of the work on deinterlacing techniques is ad hoc in nature and lacks theoretical justification. The designs are heuristic and the results are in most cases experimental. It is very difficult to quantify the overall performance of a deinterlacing system based on its performance over a few video sequences. In this paper, we present a linear system analysis for a motion-compensated video upconversion system, analyze the effect of motion vector accuracy in system performance and investigate switching the algorithm between fixed and motion-compensated deinterlacing. Mainak Biswas, Truong Q. Nguyen |
ICASSP (3) | 2 |
| 2004 | Motion vector processing for frame rate up conversionabstractIn this paper, the effect of motion vector accuracy on the efficiency of motion compensated frame rate up conversion (MC-FRUC) is studied. The motion vector processing problem is formulated and analyzed via motion vector modelling. A processing scheme is proposed to improve the interpolated frame quality at the decoder. The practical application of integrating the motion vector processing algorithm in a standard H.263 decoder for MC-FRUC is also discussed. Experimental results show that a 0.4-0.6 dB gain in the H.263 codec, and a 0.5-2 dB gain in off-line frame rate conversion can be obtained by the proposed algorithm. Gökçe Dane, Truong Q. Nguyen |
ICASSP (3) | 2 |
| 2004 | Hidden Markov tree image denoising with redundant lapped transformsabstractHidden Markov tree (HMT) wavelet models have demonstrated superior performance in image filtering, by their ability to capture features across scales. Recently, we proposed to extend the HMT framework to the lapped transform domain, where lapped transforms (LT) are M-channel linear phase filterbanks. When the number of channels is a power of 2, the block partition provided by LT is remapped to an octave-like representation, where an HMT is able to model the statistical dependencies between intra- and interband coefficients. Due to better energy compaction and reduced aliasing properties, LT outperforms discrete wavelet transforms at moderate noise levels, both subjectively and objectively. However, critically-decimated LT suffers from a lack of shift-invariance, resulting in a degraded performance. We study the improvement of HMT modeling in the LT domain (HMT-LT), combined with a redundant decomposition, in order to increase its performance for image denoising. Laurent Duval, Truong Q. Nguyen |
ICASSP (3) | 2 |
| 2004 | Global motion estimation in frequency and spatial domainabstractWe propose a fast and robust global motion estimation (GME) algorithm, based on a 2-stage coarse-to-fine refinement strategy, which is capable of measuring large motions. A 6-parameter affine motion model has been used. Coarse estimation is carried out in the frequency domain using polar, log-polar or log-log sampling of the Fourier magnitude spectrum of the sub-sampled image. The sampling scheme is adaptively selected, based on past motion patterns. The refinement stage consists of a RANSAC based model fitting to motion vectors of randomly selected high-activity blocks, and hence is robust to outliers. The motion vector of blocks is measured using phase correlation, which offers two advantages in this context: sub-pixel accuracy without significant computational overhead; and if a particular block consists of background as well as foreground pixels, both motions are simultaneously measured. Due to its hardware-friendly nature, the proposed algorithm holds potential for real-time GME even for television images. Sanjeev Kumar 0003, Mainak Biswas, Truong Q. Nguyen |
ICASSP (3) | 3 |
| 2004 | Appearance model based face-to-face transformabstractIn this paper, a novel approach to face-to-face transform is presented. The face-to-face transform is a technique, which transforms one person’s facial actions to the others. In gen-eral, the 3D models of faces are used for such transforma-tion. Therefore the facial action parameters should be es-timated from the 2D input images, which is not an easy task. On the contraly, our proposed approach is based on the 2D appearance model instead of the 3D model so that the model is acquired by learning directly from training im-ages. To achieve this, we investigate making use of the Hid-den Markov Model (HMM) framework, which models the correspondence between an input face and the other’s one as well as the appearances of both faces. The experimental results show the effectiveness of the proposed method. 1. Takayuki Nagai, Truong Q. Nguyen |
ICASSP (5) | 2 |
| 2004 | The effect of global motion parameter accuracies on the efficiency of video codingabstractIn this paper, we present theoretical analysis on how the global motion parameter accuracies affect the efficiency of motion compensated video coding. The inaccurate global motion compensation is modelled by introducing probabilistic rotation, scale and translation parameter errors. Approximate expressions that relate the power spectrum of the prediction error to motion parameter errors is derived. By doubling the accuracy of the motion parameters, up to 6 dB theoretical gains can be obtained in prediction error variance. Gökçe Dane, Truong Q. Nguyen |
ICIP | 2 |
| 2004 | Motion wavelet difference reduction (MWDR) video codecabstractIn this paper, a new fast video codec using Wavelet Difference Reduction (WDR) algorithm is presented. This proposed video codec is inspired by the concept of motion JPEG; namely, we adapted the efficient WDR still image compression algorithm into a video compression algorithm without the motion estimation step. This approach significantly reduces the processing time for motion vector search and motion compensation procedures. In addition, we employ block-based WDR algorithm which significantly reduces the memory requirement while comparing to other wavelet based compression algorithm. Yee L. Law, Truong Q. Nguyen |
ICIP | 2 |
| 2004 | DCT-based phase correlation motion estimationabstractA DCT-based phase correlation motion estimation algorithm is proposed in this paper. By combining four real transforms, a new complex linear phase transform is obtained and is used in phase correlation motion estimation. The application in video compression is discussed in detail. The simulation results show that the proposed algorithm is robust and efficient. Min Li 0014, Mainak Biswas, Sanjeev Kumar 0003, Truong Q. Nguyen |
ICIP | 4 |
| 2004 | A novel boundary handling scheme for arbitrary object shape in video and image compressionabstractThis paper presents a novel scheme for handling image boundary for a class of biorthogonal multidimensional (MD) perfect reconstruction (PR) filter banks (FB) using lifting scheme which finds important applications in image and video compression applications. An intuitive example is given to demonstrate our proposed for obtaining a nonexpansive transformation using nonseparable FB on arbitrary image shape. Ka Shun Carson Pun, Truong Q. Nguyen |
ICIP | 2 |
| 2004 | A novel and efficient design of multidimensional PR two-channel filter banks with hourglass-shaped passband supportabstractIn this letter, we propose a new design technique for two-dimensional two-channel maximally decimated perfect reconstruction (PR) filter banks with hourglass-shaped spectral support. The proposed approach is different from the conventional design techniques in the sense that it does not require any modulator. Furthermore, the proposed design retains the separable filtering property, while ensuring that the overall system is structurally PR and maximally decimated. Moreover, we extend such structure to higher dimensional case and discuss several interesting properties. A design example is given to demonstrate the effectiveness of the proposed approach. Ka Shun Carson Pun, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 2 |
| 2003 | Deblocking of video sequences with lapped embedded IDCTabstractAt mid to low bit-rates, the standard DCT based video codecs produce, in the reconstructed sequences, annoying artifacts known as ringing and blocking effects. We are concerned in blocking artifact removal with lapped embedded IDCT. In intra frames, the deblocking always produces an increase in the peak signal-to-noise ratio; as a consequence of the processing on the last reference intra frame, inter frames, whose prediction error is decoded with IDCT, are free of blocking artifacts. Using the proposed algorithm on the inter frame decreases the resulting PSNR while improving the visual quality of the decoded images. We present a theoretical bound to the PSNR reduction. Lorenzo Cappellari, Truong Q. Nguyen |
ICASSP (3) | 2 |
| 2003 | Low-order IIR filter bank designabstractThe advantage of IIR filters over FIR ones is that the former require a much lower order to obtain the desired response specifications. However, the existing deterministic techniques for IIR filter bank design based on heuristic usually lead to too high order IIR filters and thus cannot be practically used. In this paper, we propose new method to solve the low-order IIR filter bank design, which is based on linear matrix inequalities (LMI) optimization. Our focus is the QMF bank design, although other IIR filter related problems can be treated and solved in similar way. Hoang Duong Tuan, Tran Thai Son, Truong Q. Nguyen |
ICASSP (6) | 3 |
| 2003 | Denoising in the lapped transform domainabstractIn this paper, we extend the results of the denoising in the wavelet domain to the lapped transform (LT) domain. The LT coefficients are rearranged into the octave based representation, and their statistics are modeled by the same methods as the wavelet coefficients. Since the part of the LT that follows the discrete cosine transform (DCT) is orthogonal, the problem of removing the quantization noise introduced in the DCT based codecs can be solved in the LT domain. The maximum a posteriori (MAP) estimator is obtained, which provides a novel approach to the image compression artifact reduction problem. The proposed method removes the additive noise and the quantization noise effectively with the significant improvement of image quality. Seungjoon Yang, Truong Q. Nguyen |
ICASSP (6) | 2 |
| 2003 | Lapped transform domain denoising using hidden Markov treesabstractAlgorithms based on wavelet-domain hidden Markov tree (HMT) have demonstrated excellent performance for image denoising. The HMT model is able to capture image features across the scales, in contrast to classical shrinkage that thresholds subbands independently. In this work, we extend the aforementioned results to a lapped transform domain. Lapped transforms (LT) are M-channel linear phase filter banks. Their use is motivated by their good energy compaction properties and robustness to oversmoothing. It is also observed that LT preserve better oscillatory image components, such as textures. Since LT are applied as block transforms, the transforms coefficients are rearranged into an octave-like decomposition, and their statistics are modeled by the same HMT structure as in the wavelet case. At moderate noise levels, the proposed algorithm is able to improve the results obtained with wavelets, subjectively and objectively. Laurent Duval, Truong Q. Nguyen |
ICIP (1) | 2 |
| 2003 | A simple mapping between Mth-band FIR filters using cosine modulationabstractThis article presents a novel approach in linear-phase finite-impulse response filter design based on cosine modulation. It is shown that the proposed technique is equivalent to a simple mapping between Mth-band filters, i.e., given an Lth-band filter, an Mth-band filter can be obtained by simple calculation where L and M are different integers. Soontorn Oraintara, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 2 |
| 2002 | Multiplierless approximation of transforms using lifting scheme and coordinate descent with adder constraintabstractThis paper describes an algorithm for systematically finding a multiplierless approximation of transforms where VLSI-friendly binary coefficients of the form k/2nare employed in the approximation. Assuming the cost of binary shifters is negligible in hardware, the total number of binary adders required to approximate the transform is used as the complexity constraint. The proposed algorithm is systematic and fast. It eliminates the need for trial-and-error binary approximations of the coefficients. Specifically, two types of multiplierless approximations of the discrete cosine transform (DCT) are presented to illustrate the algorithm. Ying-Jui Chen, Soontorn Oraintara, Trac D. Tran, Kevin Amaratunga, Truong Q. Nguyen |
ICASSP | 5 |
| 2002 | On two-channel orthogonal and symmetric complex-valued FIR filter banks and their corresponding waveletsabstractTwo-channel orthogonal and symmetric complex-valued FIR filter banks and their corresponding wavelets are investigated. First, the conditions for the filter bank to be orthogonal, symmetric and regular are presented. Then, a complete and minimal lattice structure is developed, which enables a general design approach for filter banks and wavelets with arbitrary length and arbitrary order of regularity. Finally, two integer implementation methods that preserve the perfect reconstruction property are proposed. Their performances are evaluated via experimental results. Xiqi Gao 0001, Truong Q. Nguyen, Gilbert Strang |
ICASSP | 2 |
| 2002 | Multiplierless approximation of transforms with adder constraintabstractThis letter describes an algorithm for systematically finding a multiplierless approximation of transforms by replacing floating-point multipliers with VLSI-friendly binary coefficients of the form k/2/sup n/. Assuming the cost of hardware binary shifters is negligible, the total number of binary adders employed to approximate the transform can be regarded as an index of complexity. Because the new algorithm is more systematic and faster than trial-and-error binary approximations with adder constraint, it is a much more efficient design tool. Furthermore, the algorithm is not limited to a specific transform; various approximations of the discrete cosine transform are presented as examples of its versatility. Ying-Jui Chen, Soontorn Oraintara, Trac D. Tran, Kevin Amaratunga, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 5 |
| 2002 | On the completeness of the lattice factorization for linear-phase perfect reconstruction filter banksabstractIn this letter, we re-examine the completeness of the lattice factorization for M-channel linear-phase perfect reconstruction filter bank (LPPRFB) with filters of the same length L=KM as discussed by Tran et al. (see IEEE Trans. Signal Processing, vol.48, p.133-47, Jan. 2000). We point out that the assertion of completeness is incorrect. Examples are presented to show that the proposed lattice structure of Tran et al. is not complete when K>2. In addition, we verify that the lattice structure is complete only when K/spl les/2. Lu Gan 0002, Kai-Kuang Ma, Truong Q. Nguyen, Trac D. Tran, Ricardo L. de Queiroz |
IEEE Signal Process. Lett. | 3 |
| 2001 | Interpolated Mth band filters for image size conversionabstractImage/video size conversion at variable rates requires that a large set of interpolation filters should be stored in a table. We present the interpolated Mth band filters as the interpolation filters. The interpolated Mth band filters are obtained from the cubic spline interpolation of a prototype M/sub p/th band eigenfilter. The proposed filter can be calculated in real time, eliminating the need for a large on-chip memory. Scaled images using the proposed filters show superb image quality. Seungjoon Yang, Truong Q. Nguyen |
ICIP (3) | 2 |
| 2001 | Low bit rate video sequence coding artifact removalabstractThe picture quality of video frames encoded at very low-bit rates often suffers from both blocking and ringing artifacts. We present two post-processing methods to mitigate the visual quality degradation caused by these artifacts. To reduce the blocking artifact of decoded images, we substitute IDCT for the lapped orthogonal transform embedded inverse discrete cosine transform (le-IDCT). On the other hand, we post-process the decoded video frames using a nonlinear robust filter to reduce the ringing artifact. Extensive simulation results indicated significant improvement in both objective and subjective visual qualities. The computation overhead incurred due to these quality enhancement operations is quite moderate, and can be easily optimized to achieve real-time operation. Seungjoon Yang, Surin Kittitornkun, Yu Hen Hu, Truong Q. Nguyen, Damon L. Tull |
MMSP | 4 |
| 2001 | Binary Mth band filters for image size conversionabstractImages and video sequences often have to be scaled to a different size. We present binary interpolation filters obtained from the Mth band eigenfilters via the lifting steps for the sampling rate alteration. The binary filters are implemented with only the shift and add operations, providing an efficient way to perform image size conversion. Scaled images using the proposed filters show superb image quality. Seungjoon Yang, Truong Q. Nguyen |
MMSP | 2 |
| 2001 | Maximum-likelihood parameter estimation for image ringing-artifact removalabstractAt low bit rates, image compression codecs based on overlapping transforms introduce spurious oscillations known as ringing artifacts in the vicinity of major edges. Unlike previous works, we present a maximum-likelihood approach to the ringing-artifact removal problem. Our approach employs a parameter estimation method based on the k-means algorithm with the number of clusters determined by a cluster-separation measure. The proposed algorithm and its simplified approximation are applied to JPEG2000 compressed images. Our results show effective and efficient removal of ringing artifacts. Seungjoon Yang, Yu Hen Hu, Truong Q. Nguyen, Damon L. Tull |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2000 | An Image-Based Bayesian Framework for Face DetectionabstractIn this paper, we present a novel approach for frontal face detection in gray-scale images. We represent both faces and clutter by using two-dimensional wavelet decomposition. To characterize the statistical dependency between different levels of wavelet, we introduce a Hidden Markov Model (HMM), in which a number of discrete states at each level capture the diversity of faces as well as clutter. Our experiments indicate that the proposed algorithm outperforms conventional template-based methods such as matched filter and eigenface methods. Lingmin Meng, Truong Q. Nguyen, David A. Castañón |
CVPR | 2 |
| 2000 | Seismic Data Compression Using GENLOT: Towards "Optimality"?abstractSummary form only given. Seismic data compression is desirable in geophysics for both storage and transmission stages. Wavelet coding methods have generated interesting developments, including a real-time field test trial in the North Sea in 1995. Previous work showed that GenLOT with basic optimization also outperforms state-of-the-art biorthogonal wavelet coders for seismic data. In this paper, we focus on the problem of filter bank optimization using various properties of seismic data. It is often desirable to evaluate the compression performance of a transform on a set of data using a priori objective measures, to reduce extensive testings by selecting only good a priori transforms, and to tailor transforms to the statistical properties of the data set. In the scope of this work, we use symmetric AR models up to order 4 to obtain an average model of the horizontal and vertical signals of a seismic stack section. Rosten et al. (1999), have already shown that order 1 or 2 models give good results in filter bank optimization for non-unitary filter banks, using coding gain optimization. Several other criteria may be used for transform optimization. Following the theory in Tran and Nguyen (1999), we use a weighted combination of C/sub o/=k/sub C/C/sub C/+k/sub S/C/sub S/+k/sub d/C/sub D/ of coding gain, stopband attenuation and DC leakage minimization functions. Laurent Duval, Van Bui-Tran, Truong Q. Nguyen, Trac D. Tran |
Data Compression Conference | 3 |
| 2000 | GenLOT optimization techniques for seismic data compressionabstractGenLOT coding has been shown an effective technique for seismic data compression, especially when compared to block-based algorithms (such as JPEG), or to wavelets. The transforms remove statistical redundancy and permit efficient compression, when used with advanced encoding techniques, such as the embedded zerotree coding framework. We derive a model for seismic data based on auto-regressive processes. This model is used to design GenLOT filter banks optimized for seismic data, using objective optimization criteria. Laurent Duval, Van Bui-Tran, Truong Q. Nguyen, Trac D. Tran |
ICASSP | 3 |
| 2000 | Video Compression Using Integer DCTabstractThis paper describes the implementation of the integer discrete cosine transform (IntDCT) using the Walsh-Hadamard transform and the lifting scheme. The implementation is in the forms of shifts and adds, and all internal nodes have finite precision. A general-purpose scheme of 8-pt IntDCT with complexity of 45 adds and 18 shifts is proposed which gives comparable performance to the floating-point DCT (FloatDCT). For this particular scheme with 8-bit input, perfect reconstruction (PR) is preserved even when all the internal nodes are limited to 16-bit words, rendering the Pentium MMX optimization possible. Implementation has been done to incorporate the proposed IntDCT into the H.263+ coder, and the resulting system performs equally well as the original. Further extension to the MPEG coder is straightforward. The proposed IntDCT is reversible, with a low level of power consumption, and is very suitable for source coding, and communication, etc. in a mobile environment. Ying-Jui Chen, Soontorn Oraintara, Truong Q. Nguyen |
ICIP | 3 |
| 2000 | Two Subspace Methods to Discriminate Faces and CluttersabstractDimension reduction via linear subspace is very important in image pattern detection and recognition. This paper presents two new methods of dimension reduction and develops algorithms to locate human faces in gray-scale still images. The first technique develops eigenface subspace and eigenclutter subspace which represent faces and clutters respectively. The second technique chooses a common subspace to maximize the Bhattacharyya distance of two Gaussian distributions. Compared with the first method, the second method is more computationally efficient with slightly higher error rate. Our simulation result indicates that both methods outperform conventional template-based methods such as matched filter and eigenface methods. Lingmin Meng, Truong Q. Nguyen |
ICIP | 2 |
| 2000 | A Method for Choosing the Regularization Parameter in Generalized Tikhonov Regularized Linear Inverse ProblemsabstractThis paper presents a systematic and computable method for choosing the regularization parameter appearing in Tikhonov-type regularization based on non-quadratic regularizers. First, we extend the notion of the L-curve, originally defined for quadratically regularized problems, to the case of non-quadratic functions. We then associate the optimal value of the regularization parameter for these non-quadratic problems with the corner of the resulting generalized L-curve. We identify the corner of this L-curve as the point of tangency between a straight line of arbitrary slope and the L-curve. This definition results in a corresponding algebraic equation which the optimal regularization parameter must satisfy. This algebraic equation naturally leads to an iterative algorithm for the optimal value of the regularization parameter. The convergence of this iterative algorithm is established. Simulation results confirm that the proposed method yields values of the regularization parameters that result in good reconstructions for non-quadratic problems. Soontorn Oraintara, W. Clem Karl, David A. Castañón, Truong Q. Nguyen |
ICIP | 4 |
| 2000 | Adaptive Scanning Methods for Wavelet Difference Reduction in Lossy Image CompressionabstractThis paper describes methods for adapting the scanning order through wavelet transform values used in the wavelet difference reduction (WDR) algorithm of Tian and Wells (1996). These new methods are called adaptively scanned wavelet difference reduction (ASWDR). ASWDR adapts the scanning procedure used by WDR in order to predict locations of significant transform values at half thresholds. These methods retain all of the important features of WDR: low-complexity, region of interest, embeddedness, and progressive SNR. They improve the rate-distortion performance of WDR so that it is essentially equal to that of the SPIHT algorithm of Said and Pearlman (1996) when arithmetic compression is not employed. When arithmetic compression is used, then the rate-distortion performance of the ASWDR algorithms is only slightly worse than SPIHT. The perceptual quality of ASWDR images is clearly superior to SPIHT. James S. Walker, Truong Q. Nguyen |
ICIP | 2 |
| 2000 | Maximum Likelihood Parameter Estimation for Image Ringing Artifact RemovalabstractAt low bit rates, image compression codecs based on overlapping transforms introduce spurious oscillation known as ringing artifacts in the vicinity of major edges. The image quality can be enhanced considerably by removing the artifacts. We present a maximum likelihood approach to the ringing artifact removal problem. Our approach employs a parameter estimation method based on the k-means algorithm with the number of clusters determined by a cluster separation measure. The proposed algorithm and its simplified approximation are applied to JPEG2000 compressed images to demonstrate their effectiveness. Seungjoon Yang, Yu Hen Hu, Damon L. Tull, Truong Q. Nguyen |
ICIP | 4 |
| 2000 | Blocking Artifact Free Inverse Discrete Cosine TransformabstractThis paper presents the generalized lapped biorthogonal transform embedded inverse discrete cosine transform (ge-IDCT) as an alternative to the IDCT. The ge-IDCT with nonlinear weighting in the embedded transform domain can reconstruct the signal with alleviated blockishness. Additional complexity, imposed by the replacement, is trivial thanks to an efficient lattice structure. The proposed ge-IDCT is applied in the JPEG still image compression standard to demonstrate its validity. Seungjoon Yang, Surin Kittitornkun, Yu Hen Hu, Truong Q. Nguyen, Damon L. Tull |
ICIP | 4 |
| 2000 | A new algorithm for linear-phase paraunitary filter banks with pairwise mirror-image frequency responses
Kwok Ping Chan, Truong Q. Nguyen, Li Chen 0003 |
Signal Process. | 2 |
| 2000 | Time-domain design and lattice structure of FIR paraunitary filter banks with linear phase
Masaaki Ikehara, Takayuki Nagai, Truong Q. Nguyen |
Signal Process. | 3 |
| 1999 | Performance analysis of multicarrier modulation systems using cosine modulated filter banksabstractWe compare the performance of biorthogonal cosine modulated transmultiplexer filter banks with today's multicarrier modulation systems whose transceivers are based on DFT. In contrast to early works on transmultiplexer filter banks that concentrated on the derivation of perfect reconstruction constraints of the filter bank or prototype design, this study takes into consideration a typical twisted pair copper line transmission channel into consideration and examines the influence of different system parameters as filter length, number of channels, and the overall system delay on the distortion at the receiver. Biorthogonal filter banks have the advantage that filter length and overall system delay can be chosen independently. Restricting the equalizer at the receiver to a single scalar tap per subchannel, we show that cosine-modulated filter banks outperform DFT based multicarrier systems without a guard interval and obtain a similar performance to DFT based systems with a guard interval and time domain equalization but at a lower computational cost and a higher throughput data rate. Subbarao S. Govardhanagiri, Tanja Karp, Peter N. Heller, Truong Q. Nguyen |
ICASSP | 4 |
| 1999 | Multirate as a hardware paradigmabstractThe architecture and circuit design are the two most effective means of reducing power in CMOS VLSI. Mathematical manipulations, based on applying ideas from multirate signal processing have been applied to create high performance, low power architectures. To illustrate this approach, two case studies are presented-one concerns the design of a fast Fourier transform (FFT) device, while the other one is concerned with the design of analog-to-digital converters. Bruce W. Suter, Kenneth S. Stevens, Scott R. Velazquez, Truong Q. Nguyen |
ICASSP | 4 |
| 1999 | A progressive transmission image coder using linear phase uniform filterbanks as block transformsabstractThis paper presents a novel image coding scheme using M-channel linear phase perfect reconstruction filterbanks (LPPRFBs) in the embedded zerotree wavelet (EZW) framework introduced by Shapiro (1993). The innovation here is to replace the EZWs dyadic wavelet transform by M-channel uniform-band maximally decimated LPPRFBs, which offer finer frequency spectrum partitioning and higher energy compaction. The transform stage can now be implemented as a block transform which supports parallel processing and facilitates region-of-interest coding/decoding. For hardware implementation, the transform boasts efficient lattice structures, which employ a minimal number of delay elements and are robust under the quantization of lattice coefficients. The resulting compression algorithm also retains all the attractive properties of the EZW coder and its variations such as progressive image transmission, embedded quantization, exact bit rate control, and idempotency. Despite its simplicity, our new coder outperforms some of the best image coders published previously in the literature, for almost all test images (especially natural, hard-to-code ones) at almost all bit rates. Trac D. Tran, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 1998 | A new multiresolution algorithm for image segmentationabstractWe present here a novel multiresolution-based image segmentation algorithm. The proposed method extends and improves the Gaussian mixture model (GMM) paradigm by incorporating a multiscale correlation model of pixel dependence into the standard approach. In particular, the standard GMM is modified by introducing a multiscale neighborhood clique that incorporates the correlation between pixels in space and scale. We modify the log likelihood function of the image field by a penalization term that is derived from a multiscale neighborhood clique. Maximum likelihood (ML) estimation via the expectation-maximization (EM) algorithm is used to estimate the parameters of the new model. Then, utilizing the parameter estimates, the image field is segmented with a MAP classifier. It is demonstrated that the proposed algorithm provides superior segmentations of synthetic images, yet is computationally efficient. Mohammed Saeed 0003, W. Clem Karl, Truong Q. Nguyen, Hamid R. Rabiee 0001 |
ICASSP | 3 |
| 1998 | The generalized lapped biorthogonal transformabstractA lattice structure based on the singular value decomposition (SVD) is introduced. The lattice can be proven to use a minimal number of delay elements and to completely span a large class of M-channel linear phase perfect reconstruction filter banks (LPPRFB): all analysis and synthesis filters have the same FIR length of L=KM, sharing the same center of symmetry. The lattice also structurally enforces both linear phase and perfect reconstruction properties, is capable of providing fast and efficient implementation, and avoids the costly matrix inversion problem in the optimization process. From a block transform perspective, the new lattice represents a family of generalized lapped biorthogonal transforms (GLBT) with arbitrary integer overlapping factor K. The relaxation of the orthogonal constraint allows the GLBT to have significantly different analysis and synthesis basis functions which can then be tailored appropriately to fit a particular application. Several design examples are presented along with a high-performance GLBT-based progressive image coder to demonstrate the superiority of the new lapped transforms. Trac D. Tran, Ricardo L. de Queiroz, Truong Q. Nguyen |
ICASSP | 3 |
| 1998 | A GenLOT-based Progressive Image Coder for Low Resolution ImagesabstractThe popular EZW (embedded zerotree wavelet) and its improved version SPIHT (set partitioning in hierarchical trees) are high-performance progressive transmission image coders based on the wavelet transform which gives excellent compression results for images with significant low-frequency content. For images with high texture contents, the GenLOT-based coder outperforms SPIHT in PSNR measure by a wide margin. On the other hand, low bit rate video finds applications in videophone and surveillance systems, where smaller size image in QCIF format is often transmitted. We show that using the conventional zero-tree algorithm for the QCIF image is suboptimal and we propose several progressive algorithms with modified zero-tree structures. The extensive coding results using both the DCT and GenLOT transform confirms that our proposed modified zero-tree algorithm outperforms the conventional zerotree algorithm for QCIF-sized images. Mika Helsingius, Trac D. Tran, Truong Q. Nguyen |
ICIP (2) | 3 |
| 1998 | Generalized Lapped Biorthogonal Transforms with Integer Coefficients
Masaaki Ikehara, Trac D. Tran, Truong Q. Nguyen |
ICIP (3) | 3 |
| 1998 | Image/Video Scaling Algorithm based on Multirate Signal ProcessingabstractThis paper presents an approach for image and video scaling using multirate signal processing. The main objective is to scale images with an arbitrary rational scaling ratio without visible aliasing or distortion artifact. The approach can be applied to grey scale images, color images and video signals in both spatial domain and time domain. Cosine modulation is used to minimize the required on-chip memory since only a prototype filter is stored with some cosine modulation factors. The filters are shown to have comparable regularity for each scaling factor. An efficient structure is proposed for limited bit length filter coefficients which has no imaging artifact after the filter coefficients are quantized. Simulations on still image and video scaling are presented. Soontorn Oraintara, Truong Q. Nguyen |
ICIP (2) | 2 |
| 1998 | The Variable-Length Generalized Lapped Biorthogonal Transform
Trac D. Tran, Ricardo L. de Queiroz, Truong Q. Nguyen |
ICIP (3) | 3 |
| 1998 | Image coding ringing artifact reduction using morphological post-filteringabstractRinging is an annoying artifact frequently encountered in low bit-rate transform and subband decomposition based compression of different media such as image, intra frame video and graphics. A mathematical morphology based post-processing algorithm is presented in this paper for image ringing artifact suppression. First, we use binary morphological operators to isolate the regions of an image where the ringing artifact is most prominent to the human visual system (HVS) while preserving genuine edges and other (high-frequency) fine details present in the image. Then, a gray-level morphological nonlinear smoothing filter is applied to the unmasked regions of the image under the filtering mask to eliminate ringing within this constraint region. To gauge the effectiveness of this approach, we propose an HVS compatible objective measure of the ringing artifact. Preliminary simulations indicate that the proposed method is capable of significantly reducing the ringing artifact on both subjective and objective basis. Seyfullah H. Oguz, Yu Hen Hu, Truong Q. Nguyen |
MMSP | 3 |
| 1997 | Time-domain design of linear-phase PR filter banksabstractWe present a novel way to design biorthogonal and paraunitary linear phase (LPPUFB) filter banks. The square error of the perfect reconstruction condition is expressed in a quadratic form of filter coefficients and the cost function is minimized by solving the linear equation iteratively without nonlinear optimization. With some modifications, the method can be extended to the design of paraunitary filter banks. Using this method, we can design LPPUFB with many channels easily and quickly. Design examples are given to validate the proposed method. Masaaki Ikehara, Truong Q. Nguyen |
ICASSP | 2 |
| 1997 | Efficiently VLSI-realizable prototype filters for modulated filter banksabstractThis paper presents methods for the efficient realization of prototype filters for modulated filter banks. The implementation is based on the lattice structure of the polyphase filters. The lattice coefficients, representing rotations, are approximated by a small number of simple /spl mu/-rotations each of which can be realized by some shift and add operations instead of a multiplication. Since the lattice structure is robust against coefficient quantization we do not loose the perfect reconstruction (PR) property of the filter bank when doing this approximation. The frequency responses of the original and approximated prototype filters are compared in terms of complexity and stopband attenuation. Tanja Karp, Alfred Mertins, Truong Q. Nguyen |
ICASSP | 3 |
| 1997 | Wavelet-Based Fractal Transforms for Image Coding with No SearchabstractThe compression performance of fractal image coding is considered using the wavelet-based fractal coder with no search or classification of the domain blocks. A new partitioning scheme is introduced as a variant of earlier schemes which further improves the compression performance. The wavelet-based fractal transform (WBFT) links the theory of multiresolution analysis (MRA) with iterated function systems (IFS). This not only provides a local time-frequency analysis on (the partitions of) the image using multiresolution representation but also an iterative construction of the same (partitions of the) image using IFS and fixed point theory. A set of experiments and simulations show the potentials of using the WBFT for image coding after uniform quantization and entropy coding of the coefficients of the transform. Possibilities for further improvements are discussed. Saeed Asgari, Truong Q. Nguyen, William A. Sethares |
ICIP (2) | 2 |
| 1997 | Image Compression Using Shift-Invariant Dyadic WaveletsabstractA new class of wavelet filters, shift-invariant wavelet filters, is proposed for the purpose of image coding. The existing approaches obtain shift-invariant wavelet transform by finding the path in the full decomposition tree that minimizes the shift-variance with respect to a given cost function. This procedure is signal dependent and is inefficient for image coding, since the subband decomposition has to be performed for all shifts of input image during processing time. The shift-invariant wavelet transform proposed has a better shift-invariant property compared with the conventional dyadic wavelet transform without changing the dyadic structure. It is also independent of the input images. Experimental results in image coding show that the shift-invariant wavelet transform has a better energy compaction property. Two bit-allocation schemes, which are suitable for the proposed shift-invariant wavelet transform coding, are proposed and evaluated. Y. Hui, Chi-Wah Kok, Truong Q. Nguyen |
ICIP (1) | 3 |
| 1997 | Linear Phase Paraunitary Filter Banks with Unequal-Length FiltersabstractThis paper presents the theory, design, and efficient implementation of a new class of linear phase paraunitary filter banks (LPPUFB) which find application in transform-based image coding. These new LPPUFB have filters of different lengths: longer filters are kept for low-frequency components to prevent blocking, while shorter filters are reserved for high-frequency components to minimize ringing. The proposed lattice factorization structurally enforces the unequal-length property along with LP and PU properties, and it is robust under the quantization of lattice coefficients. Design and image coding examples are also presented to confirm the validity of the theory. Masaaki Ikehara, T. Tran, Truong Q. Nguyen |
ICIP (2) | 3 |
| 1997 | Bayesian Restoration of Noisy Images with the EM AlgorithmabstractIn this paper, we demonstrate that a window-based Gaussian mixture model can be applied in the development of a robust nonlinear filter for image restoration. Via the EM algorithm, we utilize ML estimation of the spatially-varying model parameters to achieve the desired noise suppression and detail preservation. We demonstrate that this approach is a powerful tool which gives us information about the local statistics of noisy images. We demonstrate that the estimated local statistics can be efficiently utilized for outlier detection and edge detection. The advantage of our algorithm is that it can simultaneously suppress additive Gaussian and impulsive noise, while preserving fine details and edges. Mohammed Saeed 0003, Hamid R. Rabiee 0001, W. Clem Karl, Truong Q. Nguyen |
ICIP (2) | 4 |
| 1997 | Joint optimization of lattice vector quantizer and entropy coder for a Laplacian sourceabstractThis paper presents a joint optimization algorithm for lattice vector quantization (LVQ) and entropy coding for a Laplacian source at all ranges of bit rates. Entropy-constrained lattice vector quantizers (ECLVQs) are often used in practical coding systems. In order to develop an ECLVQ design algorithm, we derive estimation expressions for both distortion and entropy. From these estimations, we develop an algorithm that jointly optimizes LVQ and the entropy coder pair for a given entropy rate. Compared to previously reported approaches, the approach reported quickly computes a highly accurate optimal ECLVQ at all ranges of bit rates. Since a Laplacian source represents a wide class of subband transformed data, the algorithm can be readily applied as a subband coding method. When the proposed algorithm is applied to a wavelet based image coding, the coding performance surpasses those of any previously reported subband coders, especially at low bit rates. Wonha Kim, Yu Hen Hu, Truong Q. Nguyen |
MMSP | 3 |
| 1996 | Discrete coefficients filter banks and applications in image codingabstractDiscrete coefficient filter banks reduce the computational complexity for subband image coding systems. A new time domain formulation is proposed for the design of near-perfect reconstruction filter banks with discrete coefficients. It can be shown that the proposed design method yields an l/sup 2/ optimal solution which is considered to be practical in image coding applications. In addition, special emphasis is placed in designing filter banks with nonnegative scaling functions that smoothly decay to zero at block edges in order to reduce basis related artifacts. As a result, the reconstructed image does not suffer from blocking and ringing artifacts at medium bit rate (25:1 compression at 0.32 bit per pixel.). Chi-Wah Kok, Truong Q. Nguyen |
ICASSP | 2 |
| 1996 | Biorthogonal cosine-modulated filter bankabstractCosine-modulated filter banks have been studied extensively because of their design ease and efficient implementation. These filter banks either have restricted lengths or assume the paraunitary property for the polyphase matrices. In this paper, the biorthogonal cosine-modulated filter bank with arbitrary length is considered and the perfect reconstruction (PR) conditions are derived. These conditions are the general form of the PR conditions reported in the literature. Examples of PR systems with variable overall delay are designed using the quadratic constrained least squares formulation. Truong Q. Nguyen, Peter N. Heller |
ICASSP | 1 |
| 1995 | Linear-phase M-band wavelets with application to image codingabstractThis paper investigates the design of M-band linear phase wavelet filter banks (M>2), and explores their application to image coding. The generalized LOT description of M-band linear-phase paraunitary filter banks is used to parametrize the M-band linear-phase orthogonal wavelets. It is proven that an M-band linear-phase orthogonal wavelet of even length cannot have more than one vanishing moment. Since this limits the effectiveness of the resulting wavelet filters, we next suggest methods for the construction of linear-phase biorthogonal M-band wavelet lowpass filters, generalizing prior 2-band constructions. However, one cannot guarantee that an arbitrary lowpass filter pair can be completed to a full perfect-reconstruction filter bank. Finally, the new linear-phase orthogonal wavelet filter banks are compared with known wavelet filters with regard to their performance in a transform-based image coder. Peter N. Heller, Truong Q. Nguyen, Hemant Singh, W. Knox Carey |
ICASSP | 2 |
| 1995 | IIR M-Th Band Filters with Allpass ComponentsabstractIn this paper, we propose a new allpass-based structure for the IIR M-th and 2M-th band filters. These filters consist of M allpass filters, and an interpolation filter (sum of two allpasses). Consequently, the proposed structure is very efficient in implementation. By choosing the allpass phase appropriately the resulting phase response of the IIR M-th band filter is approximately linear. An example is designed and compared with FIR M-th band filters. T. Engin Tuncer, Truong Q. Nguyen |
ISCAS | 2 |
| 1994 | Symmetric Extension Methods for Parallel M-Channel PR LP FIR Analysis/Synthesis SystemsabstractIn this paper we study support preservative (SP) symmetric extension methods for M-channel perfect-reconstruction linear-phase FIR analysis/synthesis systems. For a given finite-duration sequence and a given FIR linear-phase filter bank, necessary and sufficient conditions are derived such that SP symmetric extensions exist. Moreover, explicit methods are given to construct the extensions. Results are extended to linear-phase paraunitary filter banks. Design examples and experiments in 2, 3, 4 and 8 channel cases are included to verify the theory.> Li Chen 0003, Truong Q. Nguyen, Kwok Ping Chan |
ISCAS | 2 |
| 1994 | Polarity-Coincidence Filter Banks and Nondestructive EvaluationabstractIn Nondestructive Evaluation (NDE) applications, the technique of split-spectrum processing (SSP) decomposes the received signal into many subbands, and then uses nonlinear processing techniques (such as minimization or polarity thresholding) to detect the flaws' signal and location. The resulting subband signals are coherently combined to obtain a high-resolution version of the original signal. In this paper, we introduce the notion of Polarity-Coincidence (PC) in the context of wideband-signal detection. In an appropriately designed filter bank, all subband signals display a common polarity at the time of a transient (flaw). Necessary and sufficient conditions on the filter banks with PC properties are derived. Various conventional filter banks are studied for their PC properties. It turns out that the cosine-modulated filter bank cannot be a PC filter bank whereas certain linear-phase and pairwise-mirror-image filter banks may be PC. Design methods and examples of PC filter banks are given. The use of these filter banks, together with a PC polarity thresholding and recomposition algorithm, is simulated on ultrasonics data.> Truong Q. Nguyen, Sriram Jayasimha |
ISCAS | 1 |
| 1994 | On Perfect-Reconstruction Allpass-Based Cosine-Modulated IIR Filter BanksabstractIn this paper, we consider the theory and design of the cosine-modulated infinite-impulse-response (IIR) perfect reconstruction (PR) filter banks. The analysis and synthesis filters are cosine-modulated versions of prototype filter. Moreover, the polyphase components of the prototype filter are allpass filters. Necessary and sufficient conditions on the allpass filters are derived such that the filter bank is a PR system. The proposed filter bank is the only known IIR cosine-modulated filter bank. A design procedure for the proposed filter bank is outlined and design examples are given to demonstrate the theory.> Truong Q. Nguyen, Timo I. Laakso, T. Engin Tuncer |
ISCAS | 1 |
| 1994 | Generalized Linear-Phase Lapped Orthogonal TransformsabstractThe general factorization of a linear-phase paraunitary filter bank (LPPUFB) is revisited and we introduce a class of lapped orthogonal transforms with extended overlap (GenLOT). In this formulation, the discrete cosine transform (DCT) is the order-1 GenLOT, the lapped orthogonal transform is the order-2 GenLOT, and so on, for any filter length which is an integer multiple of the block size. All GenLOTs are based on the DCT and have fast implementation algorithms. The degrees of freedom in the design of GenLOTs are described and design examples are presented along with some practical applications.> Ricardo L. de Queiroz, Truong Q. Nguyen, Kamisetty Ramamohan Rao |
ISCAS | 2 |
| 1993 | Linear phase orthonormal filter banks
Anand K. Soman, P. P. Vaidyanathan, Truong Q. Nguyen |
ICASSP (3) | 3 |
| 1993 | Partial-spectrum-reconstruction digital filter banks
Truong Q. Nguyen |
ISCAS | 1 |
| 1993 | A quadratic-constrained least-squares approach to linear phase orthonormal filter bank design
Truong Q. Nguyen, Anand K. Soman, P. P. Vaidyanathan |
ISCAS | 1 |
| 1991 | On the problem of reconstructing a segment of a wideband signal using a digital filter bankabstractThe design process is reduced to designing the filter bank to satisfy a set of conditions and it is known that the reconstructed signal always possesses some alias and distortion. It is the present objective to find the necessary and sufficient conditions on the filters so that the resulting QMF (quadrature mirror filter) bank cancels most alias components. Having found these conditions, the author suggests an algorithm to design the PAC (partial alias cancellation) QMF bank. Examples are given to demonstrate the theory.> Truong Q. Nguyen |
ICASSP | 1 |
| 1991 | The eigenfilter for the design of linear-phase filters with arbitrary magnitude responseabstractThe author studies the design of linear-phase FIR (finite impulse response) digital filters to approximate linear-phase functions with arbitrary magnitude response. The design method is based on the computation of an eigenvector of an appropriate real, symmetric, and positive-definite matrix. The design of complex-coefficient linear-phase filters is shown to be an extension of the design of the real-coefficient filter. Several examples are presented which demonstrate the usefulness of the approach.> Truong Q. Nguyen |
ICASSP | 1 |
| 1989 | Lattice structures for design of three-channel linear-phase perfect-reconstruction FIR QMF banksabstractThe authors present a perfect-reconstruction FIR (finite impulse response) linear-phase lattice structure for the three-channel QMF (quadrature mirror filter) bank. Both the analysis and the synthesis filters are linear-phase. To speed up the design time, the pairwise mirror-image condition is imposed on the resulting lattice structure. A design example is presented, and the complexity of the analysis bank is discussed.> Truong Q. Nguyen, P. P. Vaidyanathan |
ICASSP | 1 |
| 1988 | Eigenfilters for the design of special transfer functions with applications in multirate signal processingabstractBased on the multistage approach, a design procedure is presented for finding a spectral factor of an mth-band filter and for designing multistage decimation filters. The proposed design method finds spectral factors of mth-band FIR (finite-impulse response) filters without direct computation, and yields filters with much higher attenuation than would be possible by conventional methods. Such mth-band filters are used in filter-bank designs, including perfect-reconstruction systems.> Truong Q. Nguyen, Tapio Saramäki, P. P. Vaidyanathan |
ICASSP | 1 |
| 1988 | Improved approach for design of perfect reconstruction FIR QMF banks, with lossless lattice structuresabstractA property of FIR (finite-impulse response) lossless systems is introduced, leading to substantial improvement in the sign procedure for perfect-reconstruction QMF (quadrature mirror filter) banks. The property enables the designer to initialize the coefficients of a lattice structure (which characterizes the analysis bank), in such a way as to speed up to the convergence. A design example is provided. Compared to other methods, the proposed method is shown to converge faster, and always leads to much improved attenuation characteristics for a given filter length.> P. P. Vaidyanathan, Truong Q. Nguyen, Tapio Saramäki |
ICASSP | 2 |