Minh N. Do

dblp:d/MinhNDo · DBLP profile ↗
← Back
146ranked-venue papers
14as first author
12since 2021 · last 2025
0000-0001-5132-4986ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 120 · 14 first-author · 6 since 2021Artificial intelligence and machine learning · 34 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 since 2021Systems, architecture and hardware · 5Databases, data management, data science and information retrieval · 2Theory of computation · 2Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Robult: Leveraging Redundancy and Modality-Specific Features for Robust Multimodal Learning
abstract
Addressing missing modalities and limited labeled data is crucial for advancing robust multimodal learning. We propose Robult, a scalable framework designed to mitigate these challenges by preserving modality-specific information and leveraging redundancy through a novel information-theoretic approach. Robult optimizes two core objectives: (1) a soft Positive-Unlabeled (PU) contrastive loss that maximizes task-relevant feature alignment while effectively utilizing limited labeled data in semi-supervised settings, and (2) a latent reconstruction loss that ensures unique modality-specific information is retained. These strategies, embedded within a modular design, enhance performance across various downstream tasks and ensure resilience to incomplete modalities during inference. Experimental results across diverse datasets validate that Robult achieves superior performance over existing approaches in both semi-supervised learning and missing modality contexts. Furthermore, its lightweight design promotes scalability and seamless integration with existing architectures, making it suitable for real-world multimodal applications.
Duy A. Nguyen, Abhi Kamboj, Minh N. Do
IJCAI3
2025 Care-PD: A Multi-Site Anonymized Clinical Dataset for Parkinson's Disease Gait Assessment
abstract
Objective gait assessment in Parkinson’s Disease (PD) is limited by the absence of large, diverse, and clinically annotated motion datasets. We introduce Care-PD, the largest publicly available archive of 3D mesh gait data for PD, and the first multi-site collection spanning 9 cohorts from 8 clinical centers. All recordings (RGB video or motion capture) are converted into anonymized SMPL meshes via a harmonized preprocessing pipeline. Care-PD supports two key benchmarks: supervised clinical score prediction (estimating Unified Parkinson’s Disease Rating Scale, UPDRS, gait scores) and unsupervised motion pretext tasks (2D-to-3D keypoint lifting and full-body 3D reconstruction). Clinical prediction is evaluated under four generalization protocols: within-dataset, cross-dataset, leave-one-dataset-out, and multi-dataset in-domain adaptation.To assess clinical relevance, we compare state-of-the-art motion encoders with a traditional gait-feature baseline, finding that encoders consistently outperform handcrafted features. Pretraining on Care-PD reduces MPJPE (from 60.8mm to 7.5mm) and boosts PD severity macro-F1 by 17\%, underscoring the value of clinically curated, diverse training data. Care-PD and all benchmark code are released for non-commercial research (Code, Data).
Vida Adeli, Ivan Klabucar, Javad Rajabi, Benjamin Filtjens, Soroush Mehraban, Diwei Wang, Trung-Hieu Hoang, Minh N. Do, Hyewon Seo, Candice Müller, Daniel Boari Coelho, Claudia de Oliveira, Pieter Ginis, Moran Gilat, Alice Nieuwboer, Joke Spildooren, J. Lucas McKay, Hyeokhyen Kwon, Gari D. Clifford, Christine D. Esper, Stewart A. Factor, Imari Genias, Amirhossein Dadashzadeh, Leia C. Shum, Alan L. Whone, Majid Mirmehdi, Andrea Iaboni, Babak Taati
NeurIPS8
2024 MaskDiff: Modeling Mask Distribution with Diffusion Probabilistic Model for Few-Shot Instance Segmentation
abstract
Few-shot instance segmentation extends the few-shot learning paradigm to the instance segmentation task, which tries to segment instance objects from a query image with a few annotated examples of novel categories. Conventional approaches have attempted to address the task via prototype learning, known as point estimation. However, this mechanism depends on prototypes (e.g. mean of K-shot) for prediction, leading to performance instability. To overcome the disadvantage of the point estimation mechanism, we propose a novel approach, dubbed MaskDiff, which models the underlying conditional distribution of a binary mask, which is conditioned on an object region and K-shot information. Inspired by augmentation approaches that perturb data with Gaussian noise for populating low data density regions, we model the mask distribution with a diffusion probabilistic model. We also propose to utilize classifier-free guided mask sampling to integrate category information into the binary mask generation process. Without bells and whistles, our proposed method consistently outperforms state-of-the-art methods on both base and novel classes of the COCO dataset while simultaneously being more stable than existing methods. The source code is available at: https://github.com/minhquanlecs/MaskDiff.
Minh-Quan Le, Tam V. Nguyen 0002, Trung-Nghia Le, Thanh-Toan Do, Minh N. Do, Minh-Triet Tran
AAAI5
2024 Making Vision Transformers Truly Shift-Equivariant
abstract
In the field of computer vision, Vision Transformers (ViTs) have emerged as a prominent deep learning architecture. Despite being inspired by Convolutional Neural Networks (CNNs), ViTs are susceptible to small spatial shifts in the input data - they lack shift-equivariance. To address this shortcoming, we introduce novel data-adaptive designs for each of the ViT modules that break shift-equivariance, such as tokenization. self-attention, patch merging, and positional encoding. With our proposed modules, we achieve perfect circular shift-equivariance across four prominent ViT ar-chitectures: Swin, SwinV2, CvT, and MViTv2. Additionally, we leverage our design to further enhance consistency under standard shifts. We evaluate our adaptive ViT models on image classification and semantic segmentation tasks. Our models achieve competitive performance across three diverse datasets, showcasing perfect (100%) circular shift consistency while improving standard shift consistency.11Project website: https://renanrojasg.github.io/shifteq_vit.
Renan A. Rojas-Gomez, Teck-Yian Lim, Minh N. Do, Raymond A. Yeh
CVPR3
2024 Modelling-based joint embedding of histology and genomics using canonical correlation analysis for breast cancer survival prediction
abstract
Traditional approaches to predicting breast cancer patients' survival outcomes were based on clinical subgroups, the PAM50 genes, or the histological tissue's evaluation. With the growth of multi-modality datasets capturing diverse information (such as genomics, histology, radiology and clinical data) about the same cancer, information can be integrated using advanced tools and have improved survival prediction. These methods implicitly exploit the key observation that different modalities originate from the same cancer source and jointly provide a complete picture of the cancer. In this work, we investigate the benefits of explicitly modelling multi-modality data as originating from the same cancer under a probabilistic framework. Specifically, we consider histology and genomics as two modalities originating from the same breast cancer under a probabilistic graphical model (PGM). We construct maximum likelihood estimates of the PGM parameters based on canonical correlation analysis (CCA) and then infer the underlying properties of the cancer patient, such as survival. Equivalently, we construct CCA-based joint embeddings of the two modalities and input them to a learnable predictor. Real-world properties of sparsity and graph-structures are captured in the penalized variants of CCA (pCCA) and are better suited for cancer applications. For generating richer multi-dimensional embeddings with pCCA, we introduce two novel embedding schemes that encourage orthogonality to generate more informative embeddings. The efficacy of our proposed prediction pipeline is first demonstrated via low prediction errors of the hidden variable and the generation of informative embeddings on simulated data. When applied to breast cancer histology and RNA-sequencing expression data from The Cancer Genome Atlas (TCGA), our model can provide survival predictions with average concordance-indices of up to 68.32% along with interpretability. We also illustrate how the pCCA embeddings can be used for survival analysis through Kaplan-Meier curves.
Vaishnavi Subramanian, Tanveer F. Syeda-Mahmood, Minh N. Do
Artif. Intell. Medicine3
2024 Smartphone-Based Digitized Neurological Examination Toolbox for Multi-Test Neurological Abnormality Detection and Documentation
abstract
Understanding the efficacy of digital biomarkers in vision-based human motion analysis is essential, not only for interpreting the computer-aided exam results but also for advancing the next generation of digital health tool solutions. This study extensively analyzes digitized neurological examination (DNE) biomarkers for detecting and documenting exam features of Parkinson's disease (PD) and other neurological disorders (OD). Collected over 113 participants, DNE-113, a multi-test DNE database of finger tapping, finger to finger, forearm roll, stand-up and walk, and facial activation tests, covering a broader range of neurological abnormalities beyond PD is first proposed. Subsequently, DNE-113 is integrated into pyDNE - a convenient open-source toolbox, streamlining the creation and assessment of digital biomarkers. This toolbox empowers us to assess the quality of DNE biomarkers across diverse classification tasks. We showcase the discriminative potency of DNE biomarkers, successfully characterizing abnormal signals in neurological patients. Our findings highlight not only the potential use cases but also the persisting challenges in constructing digital biomarkers for computer-aided movement analysis on PD and OD patients.
Trung-Hieu Hoang, Christopher M. Zallek, Minh N. Do
IEEE J. Biomed. Health Informatics3
2024 FedDCT: Federated Learning of Large Convolutional Neural Networks on Resource-Constrained Devices Using Divide and Collaborative Training
abstract
In Federated Learning (FL), the size of local models matters. On the one hand, it is logical to use large-capacity neural networks in pursuit of high performance. On the other hand, deep convolutional neural networks (CNNs) are exceedingly parameter-hungry, which makes memory a significant bottleneck when training large-scale CNNs on hardware-constrained devices such as smartphones or wearables sensors. Current state-of-the-art (SOTA) FL approaches either only test their convergence properties on tiny CNNs with inferior accuracy or assume clients have the adequate processing power to train large models, which remains a formidable obstacle in actual practice. To overcome these issues, we introduce FedDCT, a novel distributed learning paradigm that enables the usage of large, high-performance CNNs on resource-limited edge devices. As opposed to traditional FL approaches, which require each client to train the full-size neural network independently during each training round, the proposed FedDCT allows a cluster of several clients to collaboratively train a large deep learning model by dividing it into an ensemble of several small sub-models and train them on multiple devices in parallel while maintaining privacy. In this collaborative training process, clients from the same cluster can also learn from each other, further improving their ensemble performance. In the aggregation stage, the server takes a weighted average of all the ensemble models trained by all the clusters. FedDCT reduces the memory requirements and allows low-end devices to participate in FL. We empirically conduct extensive experiments on standardized datasets, including CIFAR-10, CIFAR-100, and two real-world medical datasets HAM10000 and VAIPE. Experimental results show that FedDCT outperforms a set of current SOTA FL methods with interesting convergence behaviors. Furthermore, compared to other existing approaches, FedDCT achieves higher accuracy and substantially reduces the number of communication rounds (with 4-8 times fewer memory requirements) to achieve the desired accuracy on the testing dataset without incurring any extra training cost on the server side.
Hieu H. Pham 0001, Kok-Seng Wong, Phi-Le Nguyen, Truong Thao Nguyen, Minh N. Do
IEEE Trans. Netw. Serv. Manag.6
2023 Efficient Human Vision Inspired Action Recognition Using Adaptive Spatiotemporal Sampling
abstract
Adaptive sampling that exploits the spatiotemporal redundancy in videos is critical for always-on action recognition on wearable devices with limited computing and battery resources. The commonly used fixed sampling strategy is not context-aware and may under-sample the visual content, and thus adversely impacts both computation efficiency and accuracy. Inspired by the concepts of foveal vision and pre-attentive processing from the human visual perception mechanism, we introduce a novel adaptive spatiotemporal sampling scheme for efficient action recognition. Our system pre-scans the global scene context at low-resolution and decides to skip or request high-resolution features at salient regions for further processing. We validate the system on EPIC-KITCHENS and UCF-101 (split-1) datasets for action recognition, and show that our proposed approach can greatly speed up inference with a tolerable loss of accuracy compared with those from state-of-the-art baselines. Source code is available in https://github.com/knmac/adaptive_spatiotemporal.
Khoi-Nguyen C. Mac, Minh N. Do, Minh P. Vo
IEEE Trans. Image Process.2
2022 Inverting Adversarially Robust Networks for Image Synthesis
Renan A. Rojas-Gomez, Raymond A. Yeh, Minh N. Do, Anh Totti Nguyen
ACCV (6)3
2022 Learnable Polyphase Sampling for Shift Invariant and Equivariant Convolutional Networks
abstract
We propose learnable polyphase sampling (LPS), a pair of learnable down/upsampling layers that enable truly shift-invariant and equivariant convolutional networks. LPS can be trained end-to-end from data and generalizes existing handcrafted downsampling layers. It is widely applicable as it can be integrated into any convolutional network by replacing down/upsampling layers. We evaluate LPS on image classification and semantic segmentation. Experiments show that LPS is on-par with or outperforms existing methods in both performance and shift consistency. For the first time, we achieve true shift-equivariance on semantic segmentation (PASCAL VOC), i.e., 100% shift consistency, outperforming baselines by an absolute 3.3%.
Renan A. Rojas-Gomez, Teck-Yian Lim, Alexander G. Schwing, Minh N. Do, Raymond A. Yeh
NeurIPS4
2022 Multimodal Unrolled Robust PCA for Background Foreground Separation
abstract
Background foreground separation (BFS) is a popular computer vision problem where dynamic foreground objects are separated from the static background of a scene. Typically, this is performed using consumer cameras because of their low cost, human interpretability, and high resolution. Yet, cameras and the BFS algorithms that process their data have common failure modes due to lighting changes, highly reflective surfaces, and occlusion. One solution is to incorporate an additional sensor modality that provides robustness to such failure modes. In this paper, we explore the ability of a cost-effective radar system to augment the popular Robust PCA technique for BFS. We apply the emerging technique of algorithm unrolling to yield real-time computation, feedforward inference, and strong generalization in comparison with traditional deep learning methods. We benchmark on the RaDICaL dataset to demonstrate both quantitative improvements of incorporating radar data and qualitative improvements that confirm robustness to common failure modes of image-based methods.
Spencer A. Markowitz, Corey Snyder, Yonina C. Eldar, Minh N. Do
IEEE Trans. Image Process.4
2022 Towards a Comprehensive Solution for a Vision-Based Digitized Neurological Examination
abstract
The ability to use digitally recorded and quantified neurological exam information is important to help healthcare systems deliver better care, in-person and via telehealth, as they compensate for a growing shortage of neurologists. Current neurological digital biomarker pipelines, however, are narrowed down to a specific neurological exam component or applied for assessing specific conditions. In this paper, we propose an accessible vision-based exam and documentation solution called Digitized Neurological Examination (DNE) to expand exam biomarker recording options and clinical applications using a smartphone/tablet. Through our DNE software, healthcare providers in clinical settings and people at home are enabled to video capture an examination while performing instructed neurological tests, including finger tapping, finger to finger, forearm roll, and stand-up and walk. Our modular design of the DNE software supports integrations of additional tests. The DNE extracts from the recorded examinations the 2D/3D human-body pose and quantifies kinematic and spatio-temporal features. The features are clinically relevant and allow clinicians to document and observe the quantified movements and the changes of these metrics over time. A web server and a user interface for recordings viewing and feature visualizations are available. DNE was evaluated on a collected dataset of 21 subjects containing normal and simulated-impaired movements. The overall accuracy of DNE is demonstrated by classifying the recorded movements using various machine learning models. Our tests show an accuracy beyond 90% for upper-limb tests and 80% for the stand-up and walk tests.
Trung-Hieu Hoang, Mona Zehni, Huaijin Xu, George Heintz, Christopher M. Zallek, Minh N. Do
IEEE J. Biomed. Health Informatics6
2020 Anomaly Detection in Traffic Surveillance Videos with GAN-based Future Frame Prediction
abstract
It is essential to develop efficient methods to detect abnormal events, such as car-crashes or stalled vehicles, from surveillance cameras to provide in-time help. This motivates us to propose a novel method to detect traffic accidents in traffic videos. To tackle the problem where anomalies only occupy a small amount of data, we propose a semi-supervised method using Generative Adversarial Network trained on regular sequences to predict future frames. Our key idea is to model the ordinary world with a generative model, then compare a predicted frame with the real next frame to determine if an abnormal event occurs. We also propose a new idea of encoding motion descriptors and scaled intensity loss function to optimize GAN for fast-moving objects. Experiments on the Traffic Anomaly Detection dataset of AI City Challenge 2019 show that our method achieves the top 3 results with F1 score 0.9412 and RMSE 4.8088, and S3 score 0.9261. Our method can be applied to different related applications of anomaly and outlier detection in videos.
Khac-Tuan Nguyen, Dat-Thanh Dinh, Minh N. Do, Minh-Triet Tran
ICMR3
2020 A comparison of methods for 3D scene shape retrieval
Juefei Yuan, Hameed Abdul-Rashid, Bo Li 0013, Yijuan Lu, Tobias Schreck, Song Bai 0001, Xiang Bai, Ngoc-Minh Bui, Minh N. Do, Trong-Le Do, Anh Duc Duong, Xinwei He 0001, Mike Holenderski, Dmitri Jarnikov, Tu-Khiem Le, Wenhui Li 0001, Anan Liu
Comput. Vis. Image Underst.9
2020 A Multi-Organ Nucleus Segmentation Challenge
abstract
Generalized nucleus segmentation techniques can contribute greatly to reducing the time to develop and validate visual biomarkers for new digital pathology datasets. We summarize the results of MoNuSeg 2018 Challenge whose objective was to develop generalizable nuclei segmentation techniques in digital pathology. The challenge was an official satellite event of the MICCAI 2018 conference in which 32 teams with more than 80 participants from geographically diverse institutes participated. Contestants were given a training set with 30 images from seven organs with annotations of 21,623 individual nuclei. A test dataset with 14 images taken from seven organs, including two organs that did not appear in the training set was released without annotations. Entries were evaluated based on average aggregated Jaccard index (AJI) on the test set to prioritize accurate instance segmentation as opposed to mere semantic segmentation. More than half the teams that completed the challenge outperformed a previous baseline. Among the trends observed that contributed to increased accuracy were the use of color normalization as well as heavy data augmentation. Additionally, fully convolutional networks inspired by variants of U-Net, FCN, and Mask-RCNN were popularly used, typically based on ResNet or VGG base architectures. Watershed segmentation on predicted semantic segmentation maps was a popular post-processing strategy. Several of the top techniques compared favorably to an individual human annotator and can be used with confidence for nuclear morphometrics.
Neeraj Kumar 0002, Ruchika Verma, Deepak Anand, Yanning Zhou 0001, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen 0011, Pheng-Ann Heng, Jiahui Li 0005, Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajeddin, Ali Gooya, Nasir M. Rajpoot, Xuhua Ren, Sihang Zhou 0001, Qian Wang 0001, Dinggang Shen, Cheng-Kun Yang, Chi-Hung Weng, Wei-Hsiang Yu, Chao-Yuan Yeh, Shuoyu Xu, Pak-Hei Yeung, Amirreza Mahbod, Gerald Schaefer, Isabella Ellinger, Rupert Ecker, Örjan Smedby, Chunliang Wang, Benjamin Chidester, Vinh Ton-That, Minh-Triet Tran, Jian Ma 0004, Minh N. Do, Simon Graham, Quoc Dang Vu, Jin Tae Kwak, Akshaykumar Gunda, Raviteja Chunduri, Corey Hu, Dariush Lotfi, Reza Safdari, Antanas Kascenas, Alison O'Neil, Dennis Eschweiler, Johannes Stegmaier, Yanping Cui, Kailin Chen, Xinmei Tian 0001, Philipp Grüning, Erhardt Barth, Elad Arbel, Itay Remer, Amir Ben-Dor, Ekaterina Sirazitdinova, Matthias Kohl, Stefan Braunewell, Yuexiang Li, Xinpeng Xie, LinLin Shen, Jun Ma 0016, Krishanu Das Baksi, Mohammad Azam Khan, Jaegul Choo, Adrián Colomer, Valery Naranjo, Linmin Pei, Khan M. Iftekharuddin, Kaushiki Roy, Debotosh Bhattacharjee, Aníbal Pedraza, Gloria Bueno García, Sabarinathan Devanathan, Saravanan Radhakrishnan, Praveen Koduganty, Zihan Wu 0001, Guanyu Cai, Amit Sethi
IEEE Trans. Medical Imaging39
2019 Seeing Motion in the Dark
abstract
Deep learning has recently been applied with impressive results to extreme low-light imaging. Despite the success of single-image processing, extreme low-light video processing is still intractable due to the difficulty of collecting raw video data with corresponding ground truth. Collecting long-exposure ground truth, as was done for single-image processing, is not feasible for dynamic scenes. In this paper, we present deep processing of very dark raw videos: on the order of one lux of illuminance. To support this line of work, we collect a new dataset of raw low-light videos, in which high-resolution raw data is captured at video rate. At this level of darkness, the signal-to-noise ratio is extremely low (negative if measured in dB) and the traditional image processing pipeline generally breaks down. A new method is presented to address this challenging problem. By carefully designing a learning-based pipeline and introducing a new loss function to encourage temporal stability, we train a siamese network on static raw videos, for which ground truth is available, such that the network generalizes to videos of dynamic scenes at test time. Experimental results demonstrate that the presented approach outperforms state-of-the-art models for burst processing, per-frame processing, and blind temporal consistency.
Chen Chen 0003, Qifeng Chen 0001, Minh N. Do, Vladlen Koltun
ICCV3
2019 Learning Motion in Feature Space: Locally-Consistent Deformable Convolution Networks for Fine-Grained Action Detection
abstract
Fine-grained action detection is an important task with numerous applications in robotics and human-computer interaction. Existing methods typically utilize a two-stage approach including extraction of local spatio-temporal features followed by temporal modeling to capture long-term dependencies. While most recent papers have focused on the latter (long-temporal modeling), here, we focus on producing features capable of modeling fine-grained motion more efficiently. We propose a novel locally-consistent deformable convolution, which utilizes the change in receptive fields and enforces a local coherency constraint to capture motion information effectively. Our model jointly learns spatio-temporal features (instead of using independent spatial and temporal streams). The temporal component is learned from the feature space instead of pixel space, e.g. optical flow. The produced features can be flexibly used in conjunction with other long-temporal modeling networks, e.g. ST-CNN, DilatedTCN, and ED-TCN. Overall, our proposed approach robustly outperforms the original long-temporal models on two fine-grained action datasets: 50 Salads and GTEA, achieving F1 scores of 80.22% and 75.39% respectively.
Khoi-Nguyen C. Mac, Dhiraj Joshi, Raymond A. Yeh, Jinjun Xiong, Rogério Feris, Minh N. Do
ICCV6
2019 Geometry-Aware GAN for Face Attribute Transfer
abstract
In this paper, the geometry-aware GAN is proposed to address the issue of facial attribute transfer with unpaired data. To tackle the unpaired training sample problem, the CycleGAN architecture is applied, where the bilateral mappings between the source and target domains are learned. The deformation flow is learned to capture the geometric variation between two domains. We first warp the source face into desired pose and shape according to the flow. Then, the transfer sub-network is designed to refine the results by hallucinating new components on the warped image. The attribute is removed by the reconstruction sub-network, coupled with the warping process. Experiments on benchmark demonstrate the advantages of our method compared to baselines.
Danlan Huang, Xiaoming Tao 0001, Jianhua Lu, Minh N. Do
ICIP4
2019 Dense 3D Reconstruction for Visual Tunnel Inspection using Unmanned Aerial Vehicle
abstract
Advances in Unmanned Aerial Vehicle (UAV) opens venues for application such as tunnel inspection. Owing to its versatility to fly inside the tunnels, it can quickly identify defects and potential problems related to safety. However, long tunnels, especially with repetitive or uniform structures pose a significant problem for UAV navigation. Furthermore, post-processing visual data from the camera mounted on the UAV is required to generate useful information for the inspection task. In this work, we design a UAV with a single rotating camera to accomplish the task. Compared to other platforms, our solution can fit the stringent requirement for tunnel inspection, in terms of battery life, size and weight. While the current state-of-the-art can estimate camera pose and 3D geometry from a sequence of images, they assume large overlap, small rotational motion, and many distinct matching points between images. These assumptions severely limit their effectiveness in tunnel-like scenarios where the camera has erratic or large rotational motion, such as the one mounted on the UAV. This paper presents a novel solution which exploits Structure-from-Motion, Bundle Adjustment, and available geometry priors to robustly estimate camera pose and automatically reconstruct a fully-dense 3D scene using the least possible number of images in various challenging tunnel-like environments. We validate our system with both Virtual Reality application and experimentation with a real dataset. The results demonstrate that the proposed reconstruction along with texture mapping allows for remote navigation and inspection of tunnel-like environments, even those which are inaccessible for humans.
Ramanpreet Singh Pahwa, Kennard Yanting Chan, Jiamin Bai, Vincensius Billy Saputra, Minh N. Do, Shaohui Foong
IROS5
2019 Rotation equivariant and invariant neural networks for microscopy image analysis
abstract
MOTIVATION: Neural networks have been widely used to analyze high-throughput microscopy images. However, the performance of neural networks can be significantly improved by encoding known invariance for particular tasks. Highly relevant to the goal of automated cell phenotyping from microscopy image data is rotation invariance. Here we consider the application of two schemes for encoding rotation equivariance and invariance in a convolutional neural network, namely, the group-equivariant CNN (G-CNN), and a new architecture with simple, efficient conic convolution, for classifying microscopy images. We additionally integrate the 2D-discrete-Fourier transform (2D-DFT) as an effective means for encoding global rotational invariance. We call our new method the Conic Convolution and DFT Network (CFNet). RESULTS: We evaluated the efficacy of CFNet and G-CNN as compared to a standard CNN for several different image classification tasks, including simulated and real microscopy images of subcellular protein localization, and demonstrated improved performance. We believe CFNet has the potential to improve many high-throughput microscopy image analysis applications. AVAILABILITY AND IMPLEMENTATION: Source code of CFNet is available at: https://github.com/bchidest/CFNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Benjamin Chidester, Tianming Zhou, Minh N. Do, Jian Ma 0004
Bioinform.3
2019 Structure-Texture Image Decomposition Using Deep Variational Priors
abstract
Most variational formulations for structure-texture image decomposition force structure images to have small norm in some functional spaces, and share a common notion of edges, i.e., large-gradients or -intensity differences. However, such definition makes it difficult to distinguish structure edges from oscillations that have fine spatial scale but high contrast. In this paper, we introduce a new model by learning deep variational prior for structure images without explicit training data. An alternating direction method of multiplier (ADMM) algorithm and its modular structure are adopted to plug deep variational priors into an iterative smoothing process. The central observations are that convolution neural networks (CNNs) can replace the total variation prior, and are indeed powerful to capture the natures of structure and texture. We show that our learned priors using CNNs successfully differentiate highamplitude details from structure edges, and avoid halo artifacts. Different from previous data-driven smoothing schemes, our formulation provides another degree of freedom to produce continuous smoothing effects. Experimental results demonstrate the effectiveness of our approach on various computational photography and image processing applications, including texture removal, detail manipulation, HDR tone-mapping, and nonphotorealistic abstraction.
Youngjung Kim, Bumsub Ham, Minh N. Do, Kwanghoon Sohn
IEEE Trans. Image Process.3
2019 Visualization of the Intensity Field of a Focused Ultrasound Source In Situ
abstract
In an increasing number of applications of focused ultrasound (FUS) therapy, such as opening of the blood-brain barrier or collapsing microbubbles in a tumor, elevation of tissue temperature is not involved. In these cases, real-time visualization of the field distribution of the FUS source would allow localization of the FUS beam within the targeted tissue and allow repositioning of the FUS beam during tissue motion. In this paper, in order to visualize the FUS beam in situ, a 6-MHz single-element transducer ( f /2) was used as the FUS source and aligned perpendicular to a linear array which passively received scattered ultrasound from the sample. An image of the reconstructed intensity field pattern of the FUS source using bistatic beamforming was then superimposed on a registered B-mode image of the sample acquired using the same linear array. The superimposed image is used to provide anatomical context of the FUS beam in the sample being treated. The intensity field pattern reconstructed from a homogeneous scattering phantom was compared with the field characteristics of the FUS source characterized by the wire technique. The beamwidth estimates at the FUS focus using the in situ reconstruction technique and the wire technique were 1.5 and 1.2 mm, respectively. The depth-of-field estimates for the in situ reconstruction technique and the wire technique were 11.8 and 16.8 mm, respectively. The FUS beams were also visualized in a two-layer phantom and a chicken breast. The novel reconstruction technique was able to accurately visualize the field of an FUS source in the context of the interrogated medium.
Trong N. Nguyen, Minh N. Do, Michael L. Oelze
IEEE Trans. Medical Imaging2
2019 Automatic Curation of Sports Highlights Using Multimodal Excitement Features
abstract
The production of sports highlight packages summarizing a game's most exciting moments is an essential task for broadcast media. Yet, it requires labor-intensive video editing. We propose a novel approach for auto-curating sports highlights, and demonstrate it to create a first of a kind, real-world system for the editorial aid of golf and tennis highlight reels. Our method fuses information from the players’ reactions (action recognition such as high-fives and fist pumps), players’ expressions (aggressive, tense, smiling, and neutral), spectators (crowd cheering), commentator (tone of the voice and word analysis), and game analytics to determine the most interesting moments of a game. We accurately identify the start and end frames of key shot highlights with additional metadata, such as the player's name and the whole number, or analysts input allowing personalized content summarization and retrieval. In addition, we introduce new techniques for learning our classifiers with reduced manual training data annotation by exploiting the correlation of different modalities. Our work has been demonstrated at a major golf tournament (2017 Masters) and two major international tennis tournaments (2017 Wimbledon and U.S. Open), successfully extracting highlights through the course of the sporting events. For the 2017 Masters, 54% of the clips selected by our system overlapped with the official highlights reels. Furthermore, user studies showed that 90% of the non-overlapping ones were of the same quality of the official clips for the 2017 Masters, while the automatic selection of clips for highlights of 2017 Wimbledon and 2017 US Open agreed with human preferences 80% and 84.2% of the time, respectively.
Michele Merler, Khoi-Nguyen C. Mac, Dhiraj Joshi, Quoc-Bao Nguyen, Stephen Hammer, John Kent, Jinjun Xiong, Minh N. Do, John R. Smith, Rogério Feris
IEEE Trans. Multim.8
2018 Feature-less Stitching of Cylindrical Tunnel
abstract
Traditional image stitching algorithms use transforms such as homography to combine different views of a scene. They usually work well when the scene is planar or when the camera is only rotated, keeping its position static. This severely limits their use in real world scenarios where an unmanned aerial vehicle (UAV) potentially hovers around and flies in an enclosed area while rotating to capture a video sequence. We utilize known scene geometry along with recorded camera trajectory to create cylindrical images captured in a given environment such as a tunnel where the camera rotates around its center. The captured images of the inner surface of the given scene are combined to create a composite panoramic image that is textured onto a 3D geometrical object in Unity graphical engine to create an immersive environment for end users.
Ramanpreet Singh Pahwa, Wei Kiat Leong, Shaohui Foong, Karianto Leman, Minh N. Do
ICIS5
2018 Unsupervised Textual Grounding: Linking Words to Image Concepts
abstract
Textual grounding, i.e., linking words to objects in images, is a challenging but important task for robotics and human-computer interaction. Existing techniques benefit from recent progress in deep learning and generally formulate the task as a supervised learning problem, selecting a bounding box from a set of possible options. To train these deep net based approaches, access to a large-scale datasets is required, however, constructing such a dataset is time-consuming and expensive. Therefore, we develop a completely unsupervised mechanism for textual grounding using hypothesis testing as a mechanism to link words to detected image concepts. We demonstrate our approach on the ReferIt Game dataset and the Flickr30k data, outperforming baselines by 7.98% and 6.96% respectively.
Raymond A. Yeh, Minh N. Do, Alexander G. Schwing
CVPR2
2018 Time-Frequency Networks for Audio Super-Resolution
abstract
Audio super-resolution (a.k.a. bandwidth extension) is the challenging task of increasing the temporal resolution of audio signals. Recent deep networks approaches achieved promising results by modeling the task as a regression problem in either time or frequency domain. In this paper, we introduced Time-Frequency Network (TFNet), a deep network that utilizes supervision in both the time and frequency domain. We proposed a novel model architecture which allows the two domains to be jointly optimized. Results demonstrate that our method outperforms the state-of-the-art both quantitatively and qualitatively.
Teck-Yian Lim, Raymond A. Yeh, Yijia Xu, Minh N. Do, Mark Hasegawa-Johnson
ICASSP4
2018 Image Restoration with Deep Generative Models
abstract
Many image restoration problems are ill-posed in nature, hence, beyond the input image, most existing methods rely on a carefully engineered image prior, which enforces some local image consistency in the recovered image. How tightly the prior assumptions are fulfilled has a big impact on the resulting task performance. To obtain more flexibility, in this work, we proposed to design the image prior in a data-driven manner. Instead of explicitly defining the prior, we learn it using deep generative models. We demonstrate that this learned prior can be applied to many image restoration problems using an unified framework.
Raymond A. Yeh, Teck-Yian Lim, Chen Chen 0003, Alexander G. Schwing, Mark Hasegawa-Johnson, Minh N. Do
ICASSP6
2018 Multi-Segment Reconstruction Using Invariant Features
abstract
Multi-segment reconstruction (MSR) problem consists of recovering a signal from noisy segments with unknown positions of the observation windows. One example arises in DNA sequence assembly, which is typically solved by matching short reads to form longer sequences. Instead of trying to locate the segment within the sequence through pair-wise matching, we propose a new approach that uses shift-invariant features to estimate both the underlying signal and the distribution of the positions of the segments. Using the invariant features, we formulate the problem as a constrained nonlinear least-squares. The non-convexity of the problem leads to its sensitivity to the initialization. However, with clean data, we show empirically that for longer segment lengths, random initialization achieves exact recovery. Furthermore, we compare the performance of our approach to the results of expectation maximization and demonstrate that the new approach is robust to noise and computationally more efficient.
Mona Zehni, Minh N. Do, Zhizhen Zhao 0001
ICASSP2
2018 Integration of Spatial Distribution in Imaging-Genetics
Vaishnavi Subramanian, Weizhao Tang, Benjamin Chidester, Jian Ma 0004, Minh N. Do
MICCAI (2)5
2018 CODE: Coherence Based Decision Boundaries for Feature Correspondence
abstract
A key challenge in feature correspondence is the difficulty in differentiating true and false matches at a local descriptor level. This forces adoption of strict similarity thresholds that discard many true matches. However, if analyzed at a global level, false matches are usually randomly scattered while true matches tend to be coherent (clustered around a few dominant motions), thus creating a coherence based separability constraint. This paper proposes a non-linear regression technique that can discover such a coherence based separability constraint from highly noisy matches and embed it into a correspondence likelihood model. Once computed, the model can filter the entire set of nearest neighbor matches (which typically contains over 90 percent false matches) for true matches. We integrate our technique into a full feature correspondence system which reliably generates large numbers of good quality correspondences over wide baselines where previous techniques provide few or no matches.
Wen-Yan Lin, Fan Wang 0010, Ming-Ming Cheng, Sai-Kit Yeung, Philip Torr 0001, Minh N. Do, Jiangbo Lu
IEEE Trans. Pattern Anal. Mach. Intell.6
2018 Locating 3D Object Proposals: A Depth-Based Online Approach
abstract
2D object proposals, quickly detected regions in an image that likely contain an object of interest, are an effective approach for improving the computational efficiency and accuracy of object detection in color images. In this paper, we propose a novel online method that generates 3D object proposals in an RGB-D video sequence. Our main observation is that depth images provide important information about the geometry of the scene. Diverging from the traditional goal of 2D object proposals to provide a high recall, we aim for precise 3D proposals. We leverage on depth information per frame and multiview scene information to obtain accurate 3D object proposals. Using efficient but robust registration enables us to combine multiple frames of a scene in near real time and generate 3D bounding boxes for potential 3D regions of interest. Using standard metrics, such as precision-recall (P-R) curves and F-measure, we show that the proposed approach is significantly more accurate than the current state-of-the-art techniques. Our online approach can be integrated into simultaneous localization and mapping-based video processing for quick 3D object localization. Our method takes less than a second in MATLAB on the UW-RGBD scene data set on a single thread CPU and, thus, has potential to be used in low-power chips in unmanned aerial vehicles, quadcopters, and drones.
Ramanpreet Singh Pahwa, Jiangbo Lu, Nianjuan Jiang, Tian-Tsong Ng, Minh N. Do
IEEE Trans. Circuits Syst. Video Technol.5
2017 Direct Photometric Alignment by Mesh Deformation
abstract
The choice of motion models is vital in applications like image/video stitching and video stabilization. Conventional methods explored different approaches ranging from simple global parametric models to complex per-pixel optical flow. Mesh-based warping methods achieve a good balance between computational complexity and model flexibility. However, they typically require high quality feature correspondences and suffer from mismatches and low-textured image content. In this paper, we propose a mesh-based photometric alignment method that minimizes pixel intensity difference instead of Euclidean distance of known feature correspondences. The proposed method combines the superior performance of dense photometric alignment with the efficiency of mesh-based image warping. It achieves better global alignment quality than the feature-based counterpart in textured images, and more importantly, it is also robust to low-textured image content. Abundant experiments show that our method can handle a variety of images and videos, and outperforms representative state-of-the-art methods in both image stitching and video stabilization tasks.
Kaimo Lin, Nianjuan Jiang, Shuaicheng Liu, Loong Fah Cheong, Minh N. Do, Jiangbo Lu
CVPR5
2017 Semantic Image Inpainting with Deep Generative Models
abstract
Semantic image inpainting is a challenging task where large missing regions have to be filled based on the available visual data. Existing methods which extract information from only a single image generally produce unsatisfactory results due to the lack of high level context. In this paper, we propose a novel method for semantic image inpainting, which generates the missing content by conditioning on the available data. Given a trained generative model, we search for the closest encoding of the corrupted image in the latent image manifold using our context and prior losses. This encoding is then passed through the generative model to infer the missing content. In our method, inference is possible irrespective of how the missing content is structured, while the state-of-the-art learning based method requires specific information about the holes in the training phase. Experiments on three datasets show that our method successfully predicts information in large missing regions and achieves pixel-level photorealism, significantly outperforming the state-of-the-art methods.
Raymond A. Yeh, Chen Chen 0003, Teck-Yian Lim, Alexander G. Schwing, Mark Hasegawa-Johnson, Minh N. Do
CVPR6
2017 Interpretable and Globally Optimal Prediction for Textual Grounding using Image Concepts
abstract
Textual grounding is an important but challenging task for human-computer inter- action, robotics and knowledge mining. Existing algorithms generally formulate the task as selection from a set of bounding box proposals obtained from deep net based systems. In this work, we demonstrate that we can cast the problem of textual grounding into a unified framework that permits efficient search over all possible bounding boxes. Hence, the method is able to consider significantly more proposals and doesn’t rely on a successful first stage hypothesizing bounding box proposals. Beyond, we demonstrate that the trained parameters of our model can be used as word-embeddings which capture spatial-image relationships and provide interpretability. Lastly, at the time of submission, our approach outperformed the current state-of-the-art methods on the Flickr 30k Entities and the ReferItGame dataset by 3.08% and 7.77% respectively.
Raymond A. Yeh, Jinjun Xiong, Wen-Mei W. Hwu, Minh N. Do, Alexander G. Schwing
NIPS4
2017 DASC: Robust Dense Descriptor for Multi-Modal and Multi-Spectral Correspondence Estimation
abstract
Establishing dense correspondences between multiple images is a fundamental task in many applications. However, finding a reliable correspondence between multi-modal or multi-spectral images still remains unsolved due to their challenging photometric and geometric variations. In this paper, we propose a novel dense descriptor, called dense adaptive self-correlation (DASC), to estimate dense multi-modal and multi-spectral correspondences. Based on an observation that self-similarity existing within images is robust to imaging modality variations, we define the descriptor with a series of an adaptive self-correlation similarity measure between patches sampled by a randomized receptive field pooling, in which a sampling pattern is obtained using a discriminative learning. The computational redundancy of dense descriptors is dramatically reduced by applying fast edge-aware filtering. Furthermore, in order to address geometric variations including scale and rotation, we propose a geometry-invariant DASC (GI-DASC) descriptor that effectively leverages the DASC through a superpixel-based representation. For a quantitative evaluation of the GI-DASC, we build a novel multi-modal benchmark as varying photometric and geometric conditions. Experimental results demonstrate the outstanding performance of the DASC and GI-DASC in many cases of dense multi-modal and multi-spectral correspondences.
Seungryong Kim, Dongbo Min, Bumsub Ham, Minh N. Do, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 PatchMatch Filter: Edge-Aware Filtering Meets Randomized Search for Visual Correspondence
abstract
Though many tasks in computer vision can be formulated elegantly as pixel-labeling problems, a typical challenge discouraging such a discrete formulation is often due to computational efficiency. Recent studies on fast cost volume filtering based on efficient edge-aware filters provide a fast alternative to solve discrete labeling problems, with the complexity independent of the support window size. However, these methods still have to step through the entire cost volume exhaustively, which makes the solution speed scale linearly with the label space size. When the label space is huge or even infinite, which is often the case for (subpixel-accurate) stereo and optical flow estimation, their computational complexity becomes quickly unacceptable. Developed to search approximate nearest neighbors rapidly, the PatchMatch method can significantly reduce the complexity dependency on the search space size. But, its pixel-wise randomized search and fragmented data access within the 3D cost volume seriously hinder the application of efficient cost slice filtering. This paper presents a generic and fast computational framework for general multi-labeling problems called PatchMatch Filter (PMF). We explore effective and efficient strategies to weave together these two fundamental techniques developed in isolation, i.e., PatchMatch-based randomized search and efficient edge-aware image filtering. By decompositing an image into compact superpixels, we also propose superpixel-based novel search strategies that generalize and improve the original PatchMatch method. Further motivated to improve the regularization strength, we propose a simple yet effective cross-scale consistency constraint, which handles labeling estimation for large low-textured regions more reliably than a single-scale PMF algorithm. Focusing on dense correspondence field estimation in this paper, we demonstrate PMF's applications in stereo and optical flow. Our PMF methods achieve top-tier correspondence accuracy but run much faster than other related competing methods, often giving over 10-100 times speedup.
Jiangbo Lu, Yu Li 0003, Hongsheng Yang, Dongbo Min, Wei Yong Eng, Minh N. Do
IEEE Trans. Pattern Anal. Mach. Intell.6
2017 Efficient Tensor Completion for Color Image and Video Recovery: Low-Rank Tensor Train
abstract
This paper proposes a novel approach to tensor completion, which recovers missing entries of data represented by tensors. The approach is based on the tensor train (TT) rank, which is able to capture hidden information from tensors thanks to its definition from a well-balanced matricization scheme. Accordingly, new optimization formulations for tensor completion are proposed as well as two new algorithms for their solution. The first one called simple low-rank tensor completion via TT (SiLRTC-TT) is intimately related to minimizing a nuclear norm based on TT rank. The second one is from a multilinear matrix factorization model to approximate the TT rank of a tensor, and is called tensor completion by parallel matrix factorization via TT (TMac-TT). A tensor augmentation scheme of transforming a low-order tensor to higher orders is also proposed to enhance the effectiveness of SiLRTC-TT and TMac-TT. Simulation results for color image and video recovery show the clear advantage of our method over all other methods.
Johann A. Bengua, Ho N. Phien, Hoang Duong Tuan, Minh N. Do
IEEE Trans. Image Process.4
2017 Inverse Rendering and Relighting From Multiple Color Plus Depth Images
abstract
We propose a novel relighting approach that takes advantage of multiple color plus depth images acquired from a consumer camera. Assuming distant illumination and Lambertian reflectance, we model the reflected light field in terms of spherical harmonic coefficients of the bi-directional reflectance distribution function and lighting. We make use of the noisy depth information together with color images taken under different illumination conditions to refine surface normals inferred from depth. We first perform refinement on the surface normals using the first order spherical harmonics. We initialize this non-linear optimization with a linear approximation to greatly reduce computation time. With surface normals refined, we formulate the recovery of albedo and lighting in a matrix factorization setting, involving second order spherical harmonics. Albedo and lighting coefficients are recovered up to a global scaling ambiguity. We demonstrate our method on both simulated and real data, and show that it can successfully recover both illumination and albedo to produce realistic relighting results.
Minh N. Do
IEEE Trans. Image Process.2
2016 Robust Image and Video Dehazing with Visual Artifact Suppression via Gradient Residual Minimization
Chen Chen 0003, Minh N. Do, Jue Wang 0001
ECCV (2)2
2016 Fast Guided Global Interpolation for Depth and Motion
Yu Li 0003, Dongbo Min, Minh N. Do, Jiangbo Lu
ECCV (3)3
2016 SEAGULL: Seam-Guided Local Alignment for Parallax-Tolerant Image Stitching
Kaimo Lin, Nianjuan Jiang, Loong Fah Cheong, Minh N. Do, Jiangbo Lu
ECCV (3)4
2016 RepMatch: Robust Feature Matching and Pose for Reconstructing Modern Cities
Wen-Yan Lin, Nianjuan Jiang, Minh N. Do, Jiangbo Lu
ECCV (1)4
2016 Improving face detection with depth
abstract
Face detection serves an important role in many computer vision systems. Typically, a face detector identifies faces within a grayscale or color image. Due to the recent increase in consumer depth cameras, obtaining both color and depth images of a scene has never been easier. We propose a technique that utilizes depth information to improve face detection. Standard face detection methods, such as the Viola-Jones object detection framework, detects faces by searching an image at every location and scale. Our method increases the speed and accuracy of the Viola-Jones face detector by utilizing depth data to constrain the detector's search over the image. Leveraging a Kinect camera, we are able to detect faces 3.5× faster, while greatly reducing the amount of false positives.
Gregory P. Meyer, Steven Alfano, Minh N. Do
ICASSP3
2016 Binary code learning with semantic ranking based supervision
abstract
Recent years have witnessed the increasing popularity of binary hashing for efficient similarity search in large-scale vision problems. This paper presents a novel Supervised Ranking-Based Hashing (SRH) method for efficient binary code learning to better capture the semantic nearest neighbors and improve the search performance. In particular, a family of hash functions is designed to preserve the semantic data structure in the original high-dimensional space by utilizing the semantic ranking order information induced by any specific query. The proposed hashing framework is obtained by jointly minimizing the empirical error over the ranking violation in the binary code space together with the quantization loss between the original data and the binary codes. Furthermore, an effective regularizer for maximizing the even binary code distribution is also taken into account in the optimization to generate more efficient and compact binary codes. Experimental results have demonstrated the proposed method outperforms the state-of-the-art.
Minh N. Do
ICASSP2
2016 3D panorama reconstruction based on sitemap joining
abstract
We present a new approach for constructing the 3D panorama for an indoor environment by joining together aligned submaps. For each submap, the trajectory of the moving camera is estimated based on the Kanade-Lucas-Tomasi (KLT) features. Our method can update the feature set status by adding new features and removing expiring ones adaptively to accommodate scene changes. The accuracy of the estimated poses is further improved through sparse bundle adjustment. Furthermore, we utilize a linear optimization framework to align all submaps to obtain a consistently extended 3D panorama and to refine the visual odometry at the same time. We evaluated our approach on publicly available benchmark datasets. The experiments demonstrate that the proposed method achieves low translational drift and is robust even when the camera moves very fast.
Huixuan Wang, Yanwen Guo 0001, Minh N. Do, Caiming Zhang 0001, Changhe Tu
ICASSP3
2016 Stable and symmetric filter convolutional neural network
abstract
First we present a proof that convolutional neural networks (CNN) with max-norm regularization, max-pooling, and Relu non-linearity are stable to additive noise. Second, we explore the use of symmetric and antisymmetric filters in a baseline CNN model on digit classification, which enjoys the stability to additive noise. Experimental results indicate that the symmetric CNN outperforms the baseline model for nearly all training sizes and matches the state-of-the-art deep-net in the cases of limited training examples.
Raymond A. Yeh, Mark Hasegawa-Johnson, Minh N. Do
ICASSP3
2016 Deep learning based supervised hashing for efficient image retrieval
abstract
Due to its storage and search efficiency, hashing has attracted great attentions in large-scale vision problems such as image retrieval and recognition. This paper presents a novel Deep Learning based Supervised Hashing (DLSH) method by using a deep neural network to better capture the semantic structure of nonlinear and complex data. We consider learning a nonlinear embedding that simultaneously preserves semantic information and produces nearby binary codes for semantically similar data. Specifically, our hashing model is trained to maximize the similarity measure of neighbor pairs while preserving the relative similarity of non-neighbor pairs with a relaxed empirical penalty in the binary space. An effective regularizer for minimizing the quantization loss between the learned embedding and the binary codes is also considered in the optimization to generate better hash code quality. Experimental results have demonstrated the proposed method outperforms the state-of-the-art methods.
Minh N. Do
ICME2
2016 Action Recognition in Still Images With Minimum Annotation Efforts
abstract
We focus on the problem of still image-based human action recognition, which essentially involves making prediction by analyzing human poses and their interaction with objects in the scene. Besides image-level action labels (e.g., riding, phoning), during both training and testing stages, existing works usually require additional input of human bounding boxes to facilitate the characterization of the underlying human-object interactions. We argue that this additional input requirement might severely discourage potential applications and is not very necessary. To this end, a systematic approach was developed in this paper to address this challenging problem of minimum annotation efforts, i.e., to perform recognition in the presence of only image-level action labels in the training stage. Experimental results on three benchmark data sets demonstrate that compared with the state-of-the-art methods that have privileged access to additional human bounding-box annotations, our approach achieves comparable or even superior recognition accuracy using only action annotations in training. Interestingly, as a by-product in many cases, our approach is able to segment out the precise regions of underlying human-object interactions.
Yu Zhang 0004, Li Cheng 0001, Jianxin Wu 0001, Jianfei Cai 0001, Minh N. Do, Jiangbo Lu
IEEE Trans. Image Process.5
2016 Weakly Supervised Fine-Grained Categorization With Part-Based Image Representation
abstract
In this paper, we propose a fine-grained image categorization system with easy deployment. We do not use any object/part annotation (weakly supervised) in the training or in the testing stage, but only class labels for training images. Fine-grained image categorization aims to classify objects with only subtle distinctions (e.g., two breeds of dogs that look alike). Most existing works heavily rely on object/part detectors to build the correspondence between object parts, which require accurate object or object part annotations at least for training images. The need for expensive object annotations prevents the wide usage of these methods. Instead, we propose to generate multi-scale part proposals from object proposals, select useful part proposals, and use them to compute a global image representation for categorization. This is specially designed for the weakly supervised fine-grained categorization task, because useful parts have been shown to play a critical role in existing annotation-dependent works, but accurate part detectors are hard to acquire. With the proposed image representation, we can further detect and visualize the key (most discriminative) parts in objects of different classes. In the experiments, the proposed weakly supervised method achieves comparable or better accuracy than the state-of-the-art weakly supervised methods and most existing annotation-dependent methods on three challenging datasets. Its success suggests that it is not always necessary to learn expensive object/part detectors in fine-grained image categorization.
Yu Zhang 0004, Xiu-Shen Wei, Jianxin Wu 0001, Jianfei Cai 0001, Jiangbo Lu, Minh N. Do
IEEE Trans. Image Process.7
2015 Direct structure estimation for 3D reconstruction
abstract
Most conventional structure-from-motion (SFM) techniques require camera pose estimation before computing any scene structure. In this work we show that when combined with single/multiple homography estimation, the general Euclidean rigidity constraint provides a simple formulation for scene structure recovery without explicit camera pose computation. This direct structure estimation (DSE) opens a new way to design a SFM system that reverses the order of structure and motion estimation. We show that this alternative approach works well for recovering scene structure and camera poses from sideway motion given planar or general man-made scenes.
Nianjuan Jiang, Wen-Yan Lin, Minh N. Do, Jiangbo Lu
CVPR3
2015 DASC: Dense adaptive self-correlation descriptor for multi-modal and multi-spectral correspondence
abstract
Establishing dense visual correspondence between multiple images is a fundamental task in many applications of computer vision and computational photography. Classical approaches, which aim to estimate dense stereo and optical flow fields for images adjacent in viewpoint or in time, have been dramatically advanced in recent studies. However, finding reliable visual correspondence in multi-modal or multi-spectral images still remains unsolved. In this paper, we propose a novel dense matching descriptor, called dense adaptive self-correlation (DASC), to effectively address this kind of matching scenarios. Based on the observation that a self-similarity existing within images is less sensitive to modality variations, we define the descriptor with a series of an adaptive self-correlation similarity for patches within a local support window. To further improve the matching quality and runtime efficiency, we propose a randomized receptive field pooling, in which a sampling pattern is optimized with a discriminative learning. Moreover, the computational redundancy that arises when computing densely sampled descriptor over an entire image is dramatically reduced by applying fast edge-aware filtering. Experiments demonstrate the outstanding performance of the DASC descriptor in many cases of multi-modal and multi-spectral correspondence.
Seungryong Kim, Dongbo Min, Bumsub Ham, Seungchul Ryu, Minh N. Do, Kwanghoon Sohn
CVPR5
2015 SPM-BP: Sped-Up PatchMatch Belief Propagation for Continuous MRFs
abstract
Markov random fields are widely used to model many computer vision problems that can be cast in an energy minimization framework composed of unary and pairwise potentials. While computationally tractable discrete optimizers such as Graph Cuts and belief propagation (BP) exist for multi-label discrete problems, they still face prohibitively high computational challenges when the labels reside in a huge or very densely sampled space. Integrating key ideas from PatchMatch of effective particle propagation and resampling, PatchMatch belief propagation (PMBP) has been demonstrated to have good performance in addressing continuous labeling problems and runs orders of magnitude faster than Particle BP (PBP). However, the quality of the PMBP solution is tightly coupled with the local window size, over which the raw data cost is aggregated to mitigate ambiguity in the data constraint. This dependency heavily influences the overall complexity, increasing linearly with the window size. This paper proposes a novel algorithm called sped-up PMBP (SPM-BP) to tackle this critical computational bottleneck and speeds up PMBP by 50-100 times. The crux of SPM-BP is on unifying efficient filter-based cost aggregation and message passing with PatchMatch-based particle generation in a highly effective way. Though simple in its formulation, SPM-BP achieves superior performance for sub-pixel accurate stereo and optical-flow on benchmark datasets when compared with more complex and task-specific approaches.
Yu Li 0003, Dongbo Min, Michael S. Brown, Minh N. Do, Jiangbo Lu
ICCV4
2015 PatchMatch-Based Automatic Lattice Detection for Near-Regular Textures
abstract
In this work, we investigate the problem of automatically inferring the lattice structure of near-regular textures (NRT) in real-world images. Our technique leverages the PatchMatch algorithm for finding k-nearest-neighbor (kNN) correspondences in an image. We use these kNNs to recover an initial estimate of the 2D wallpaper basis vectors, and seed vertices of the texture lattice. We iteratively expand this lattice by solving an MRF optimization problem. We show that we can discretize the space of good solutions for the MRF using the kNNs, allowing us to efficiently and accurately optimize the MRF energy function using the Particle Belief Propagation algorithm. We demonstrate our technique on a benchmark NRT dataset containing a wide range of images with geometric and photometric variations, and show that our method clearly outperforms the state of the art in terms of both texel detection rate and texel localization score.
Tian-Tsong Ng, Kalyan Sunkavalli, Minh N. Do, Eli Shechtman, Nathan Carr 0001
ICCV4
2015 Efficient coding unit size selection for HEVC downsizing transcoding
abstract
This paper presents an efficient method to speed up High Efficiency Video Coding (HEVC) downsizing transcoding. Specifically, we focus on the coding unit (CU) size selection, one of the most computationally intensive processes in HEVC, to reduce the complexity of the transcoder. The proposed method utilizes the temporal correlation of depth levels among CUs and the continuity of the motion vector field in the precoded video to derive the most probable CU depth ranges and avoid the unnecessary exhaustive CU size search at certain depth levels. Experimental results show the proposed method can reduce the overall transcoding time by about 41% on average while maintaining the similar performance to that obtained by the cascaded re-encoding transcoder.
Minh N. Do
ISCAS2
2015 A fast tree-based algorithm for Compressed Sensing with sparse-tree prior
Huy Quang Bui, C. N. H. La, Minh N. Do
Signal Process.3
2015 Common Visual Pattern Discovery via Nonlinear Mean Shift Clustering
abstract
Discovering common visual patterns (CVPs) from two images is a challenging task due to the geometric and photometric deformations as well as noises and clutters. The problem is generally boiled down to recovering correspondences of local invariant features, and the conventionally addressed by graph-based quadratic optimization approaches, which often suffer from high computational cost. In this paper, we propose an efficient approach by viewing the problem from a novel perspective. In particular, we consider each CVP as a common object in two images with a group of coherently deformed local regions. A geometric space with matrix Lie group structure is constructed by stacking up transformations estimated from initially appearance-matched local interest region pairs. This is followed by a mean shift clustering stage to group together those close transformations in the space. Joining regions associated with transformations of the same group together within each input image forms two large regions sharing similar geometric configuration, which naturally leads to a CVP. To account for the non-Euclidean nature of the matrix Lie group, mean shift vectors are derived in the corresponding Lie algebra vector space with a newly provided effective distance measure. Extensive experiments on single and multiple common object discovery tasks as well as near-duplicate image retrieval verify the robustness and efficiency of the proposed approach.
Linbo Wang 0001, Yanwen Guo 0001, Minh N. Do
IEEE Trans. Image Process.4
2014 Bilateral Functions for Global Motion Modeling
Wen-Yan Lin, Ming-Ming Cheng, Jiangbo Lu, Hongsheng Yang, Minh N. Do, Philip Torr 0001
ECCV (4)5
2014 Relighting from multiple color and depth images using matrix factorization
abstract
In this paper, we propose a novel relighting approach that takes advantage of the 3D shape information acquired from a depth sensor. Assuming distant illumination and Lambertian reflectance, we model the reflected light field in terms of spherical harmonic coefficients of the Bi-directional Reflectance Distribution Function (BRDF) and lighting. To estimate both the reflectance and illumination, different illumination samples can be generated through moving the object of interest in space while keeping the light source unchanged. The samples are registered onto the base view's coordinate frame using camera pose estimated from multiple depth maps. Our results indicate that we can successfully recover both illumination and diffuse BRDF (up to a global scaling ambiguity). Our method can be used to estimate complex illumination in indoor environments for applications such as lighting transfer.
Minh N. Do
ICIP2
2014 Calibration of depth cameras using denoised depth images
abstract
Depth sensing devices have created various new applications in scientific and commercial research with the advent of Microsoft Kinect and PMD (Photon Mixing Device) cameras. Most of these applications require the depth cameras to be pre-calibrated. However, traditional calibration methods using a checkerboard do not work very well for depth cameras due to the low image resolution. In this paper, we propose a depth calibration scheme which excels in estimating camera calibration parameters when only a handful of corners and calibration images are available. We exploit the noise properties of PMD devices to denoise depth measurements and perform camera calibration using the denoised depth as additional set of measurements. Our synthetic and real experiments show that our depth denoising and depth based calibration scheme provides significantly better results than traditional calibration methods.
Ramanpreet Singh Pahwa, Minh N. Do, Tian-Tsong Ng, Binh-Son Hua
ICIP2
2014 Supervised Discriminative Hashing for Compact Binary Codes
abstract
Binary hashing has been increasingly popular for efficient similarity search in large-scale vision problems. This paper presents a novel Supervised Discriminative Hashing (SDH) method by jointly modeling the global and local manifold structures. Specifically, a family of discriminative hash functions is designed to map data points of the original high-dimensional space into nearby compact binary codes while preserving the geometrical similarity and discriminant properties in both global and local neighborhoods. Furthermore, the quantization loss between the original data and the binary codes together with the even binary code distribution are also taken into account in the optimization to generate more efficient and compact binary codes. Experimental results have demonstrated the proposed method outperforms the state-of-the-art.
Jiwen Lu, Minh N. Do
ACM Multimedia3
2014 Probability-Based Rendering for View Synthesis
abstract
In this paper, a probability-based rendering (PBR) method is described for reconstructing an intermediate view with a steady-state matching probability (SSMP) density function. Conventionally, given multiple reference images, the intermediate view is synthesized via the depth image-based rendering technique in which geometric information (e.g., depth) is explicitly leveraged, thus leading to serious rendering artifacts on the synthesized view even with small depth errors. We address this problem by formulating the rendering process as an image fusion in which the textures of all probable matching points are adaptively blended with the SSMP representing the likelihood that points among the input reference images are matched. The PBR hence becomes more robust against depth estimation errors than existing view synthesis approaches. The MP in the steady-state, SSMP, is inferred for each pixel via the random walk with restart (RWR). The RWR always guarantees visually consistent MP, as opposed to conventional optimization schemes (e.g., diffusion or filtering-based approaches), the accuracy of which heavily depends on parameters used. Experimental results demonstrate the superiority of the PBR over the existing view synthesis approaches both qualitatively and quantitatively. Especially, the PBR is effective in suppressing flicker artifacts of virtual video rendering although no temporal aspect is considered. Moreover, it is shown that the depth map itself calculated from our RWR-based method (by simply choosing the most probable matching point) is also comparable with that of the state-of-the-art local stereo matching methods.
Bumsub Ham, Dongbo Min, Changjae Oh, Minh N. Do, Kwanghoon Sohn
IEEE Trans. Image Process.4
2014 Fast Global Image Smoothing Based on Weighted Least Squares
abstract
This paper presents an efficient technique for performing a spatially inhomogeneous edge-preserving image smoothing, called fast global smoother. Focusing on sparse Laplacian matrices consisting of a data term and a prior term (typically defined using four or eight neighbors for 2D image), our approach efficiently solves such global objective functions. In particular, we approximate the solution of the memory-and computation-intensive large linear system, defined over a d-dimensional spatial domain, by solving a sequence of 1D subsystems. Our separable implementation enables applying a linear-time tridiagonal matrix algorithm to solve d three-point Laplacian matrices iteratively. Our approach combines the best of two paradigms, i.e., efficient edge-preserving filters and optimization-based smoothing. Our method has a comparable runtime to the fast edge-preserving filters, but its global optimization formulation overcomes many limitations of the local filtering approaches. Our method also achieves high-quality results as the state-of-the-art optimization-based techniques, but runs ∼10-30 times faster. Besides, considering the flexibility in defining an objective function, we further propose generalized fast algorithms that perform Lγ norm smoothing (0 < γ < 2) and support an aggregated (robust) data term for handling imprecise data constraints. We demonstrate the effectiveness and efficiency of our techniques in a range of image processing and computer graphics applications.
Dongbo Min, Sunghwan Choi, Jiangbo Lu, Bumsub Ham, Kwanghoon Sohn, Minh N. Do
IEEE Trans. Image Process.6
2014 Inverse Rendering of Lambertian Surfaces Using Subspace Methods
abstract
We propose a vector space approach for inverse rendering of a Lambertian convex object with distant light sources. In this problem, the texture of the object and arbitrary lightings are both to be recovered from multiple images of the object and its 3D model. Our work is motivated by the observation that all possible images of a Lambertian object lie around a low-dimensional linear subspace spanned by the first few spherical harmonics. The inverse rendering can therefore be formulated as a matrix factorization, in which the basis of the subspace is encoded in a spherical harmonic matrix S associated with the object’s geometry. A necessary and sufficient condition on S for unique factorization is derived with an introduction to a new notion of matrix rank called nonseparable full rank. A singular value decomposition-based algorithm for exact factorization in the noiseless case is introduced. In the presence of noise, two algorithms, namely, alternating and optimization based are proposed to deal with two different types of noise. A random sample consensus-based algorithm is introduced to reduce the size of the optimization problem, which is equal to the number of pixels in each image. Implementations of the proposed algorithms are done on a real data set.
Ha Q. Nguyen 0001, Minh N. Do
IEEE Trans. Image Process.2
2014 Efficient Hybrid Tree-Based Stereo Matching With Applications to Postcapture Image Refocusing
abstract
Estimating dense correspondence or depth information from a pair of stereoscopic images is a fundamental problem in computer vision, which finds a range of important applications. Despite intensive past research efforts in this topic, it still remains challenging to recover the depth information both reliably and efficiently, especially when the input images contain weakly textured regions or are captured under uncontrolled, real-life conditions. Striking a desired balance between computational efficiency and estimation quality, a hybrid minimum spanning tree-based stereo matching method is proposed in this paper. Our method performs efficient nonlocal cost aggregation at pixel-level and region-level, and then adaptively fuses the resulting costs together to leverage their respective strength in handling large textureless regions and fine depth discontinuities. Experiments on the standard Middlebury stereo benchmark show that the proposed stereo method outperforms all prior local and nonlocal aggregation-based methods, achieving particularly noticeable improvements for low texture regions. To further demonstrate the effectiveness of the proposed stereo method, also motivated by the increasing desire to generate expressive depth-induced photo effects, this paper is tasked next to address the emerging application of interactive depth-of-field rendering given a real-world stereo image pair. To this end, we propose an accurate thin-lens model for synthetic depth-of-field rendering, which considers the user-stroke placement and camera-specific parameters and performs the pixel-adapted Gaussian blurring in a principled way. Taking ~1.5 s to process a pair of 640×360 images in the off-line step, our system named Scribble2focus allows users to interactively select in-focus regions by simple strokes using the touch screen and returns the synthetically refocused images instantly to the user.
Dung T. Vu, Benjamin Chidester, Hongsheng Yang, Minh N. Do, Jiangbo Lu
IEEE Trans. Image Process.4
2013 Patch Match Filter: Efficient Edge-Aware Filtering Meets Randomized Search for Fast Correspondence Field Estimation
abstract
Though many tasks in computer vision can be formulated elegantly as pixel-labeling problems, a typical challenge discouraging such a discrete formulation is often due to computational efficiency. Recent studies on fast cost volume filtering based on efficient edge-aware filters have provided a fast alternative to solve discrete labeling problems, with the complexity independent of the support window size. However, these methods still have to step through the entire cost volume exhaustively, which makes the solution speed scale linearly with the label space size. When the label space is huge, which is often the case for (sub pixel-accurate) stereo and optical flow estimation, their computational complexity becomes quickly unacceptable. Developed to search approximate nearest neighbors rapidly, the Patch Match method can significantly reduce the complexity dependency on the search space size. But, its pixel-wise randomized search and fragmented data access within the 3D cost volume seriously hinder the application of efficient cost slice filtering. This paper presents a generic and fast computational framework for general multi-labeling problems called Patch Match Filter (PMF). For the very first time, we explore effective and efficient strategies to weave together these two fundamental techniques developed in isolation, i.e., Based-based randomized search and efficient edge-aware image filtering. By decompositing an image into compact super pixels, we also propose super pixel-based novel search strategies that generalize and improve the original Patch Match method. Focusing on dense correspondence field estimation in this paper, we demonstrate PMF's applications in stereo and optical flow. Our PMF methods achieve state-of-the-art correspondence accuracy but run much faster than other competing methods, often giving over 10-times speedup for large label space cases.
Jiangbo Lu, Hongsheng Yang, Dongbo Min, Minh N. Do
CVPR4
2013 Model-based complexity-aware coding for multiview video plus depth
abstract
This paper presents an efficient complexity-aware coding method for the popular 3-D video format, multiview video plus depth (MVD). Specifically, we propose a complexity-rate-distortion (C-R-D) model for the synthesized view. The proposed C-R-D model is then utilized to efficiently allocate the scarce computational resource for the texture and depth encoders to achieve the optimal quality of the synthesized view. Experimental results show the effectiveness of the proposed model and complexity allocation solution. Compared with the straightforward complexity allocation method with a fixed 1:1 ratio, the proposed method provided 0.2-0.9 dB gain in the peak-signal-to-noise ratio (PSNR) of the synthesized view quality.
Minh N. Do
ICIP2
2013 Efficient view synthesis based error concealment method for multiview video plus depth
abstract
This paper presents an efficient error concealment method for multiview video plus depth (MVD) delivery over error-prone channels based on the view synthesis information. Specifically, we propose an enhanced temporal error concealment (TEC) method by estimating the probable partition and missing MVs of a corrupted macroblock (MB) using the synthesized information. Furthermore, the view synthesis error concealment (VSEC) approach that directly uses the synthesized pixels to conceal the corrupted MBs is also adaptively utilized together with the TEC approach. Experimental results have shown the effectiveness of the proposed method compared with the existing methods in terms of both objective and subjective measures.
Vu-Hiep Doan, Minh N. Do
ISCAS3
2013 Spatialized audio multiparty teleconferencing with commodity miniature microphone array
abstract
This paper presents a Spatialized Audio Multiparty Teleconferencing (SAMT) system with a radically new communication experience for group teleconferencing. The system includes our recently developed 3D audio technologies: 3D sound source localization (SSL) and 3D audio capture and reproduction using a low-cost and compact design microphone array. In essence, the SAMT system offers 3D audio capture capability and spatial audio perception with multiple participants at a site, which still falls short in teleconferencing solutions. In addition to being able to identify and automatically track the active speaker, the system allows more compelling visual presentation for effective communication. Requiring only a low-cost microphone array and a consumer depth camera, the proposed system runs reliably and comfortably in real time on a commodity laptop or desktop PC. With such a minimal deployment requirement, we present a variety of user experiences created by SAMT.
Shengkui Zhao, Tien Dung Vu, Douglas L. Jones, Minh N. Do
ACM Multimedia5
2013 Joint Histogram-Based Cost Aggregation for Stereo Matching
abstract
This paper presents a novel method for performing efficient cost aggregation in stereo matching. The cost aggregation problem is reformulated from the perspective of a histogram, giving us the potential to reduce the complexity of the cost aggregation in stereo matching significantly. Differently from previous methods which have tried to reduce the complexity in terms of the size of an image and a matching window, our approach focuses on reducing the computational redundancy that exists among the search range, caused by a repeated filtering for all the hypotheses. Moreover, we also reduce the complexity of the window-based filtering through an efficient sampling scheme inside the matching window. The tradeoff between accuracy and complexity is extensively investigated by varying the parameters used in the proposed method. Experimental results show that the proposed method provides high-quality disparity maps with low complexity and outperforms existing local methods. This paper also provides new insights into complexity-constrained stereo-matching algorithm design.
Dongbo Min, Jiangbo Lu, Minh N. Do
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Efficient Techniques for Depth Video Compression Using Weighted Mode Filtering
abstract
This paper proposes efficient techniques to compress a depth video by taking into account coding artifacts, spatial resolution, and dynamic range of the depth data. Due to abrupt signal changes on object boundaries, a depth video compressed by conventional video coding standards often introduces serious coding artifacts over object boundaries, which severely affect the quality of a synthesized view. We suppress the coding artifacts by proposing an efficient postprocessing method based on a weighted mode filtering and utilizing it as an in-loop filter. In addition, the proposed filter is also tailored to efficiently reconstruct the depth video from the reduced spatial resolution and the low dynamic range. The down/upsampling coding approaches for the spatial resolution and the dynamic range are used together with the proposed filter in order to further reduce the bit rate. We verify the proposed techniques by applying them to an efficient compression of multiview-plus-depth data, which has emerged as an efficient data representation for 3-D video. Experimental results show that the proposed techniques significantly reduce the bit rate while achieving a better quality of the synthesized view in terms of both objective and subjective measures.
Dongbo Min, Minh N. Do
IEEE Trans. Circuits Syst. Video Technol.3
2012 Cross-based local multipoint filtering
abstract
This paper presents a cross-based framework of performing local multipoint filtering efficiently. We formulate the filtering process as a local multipoint regression problem, consisting of two main steps: 1) multipoint estimation, calculating the estimates for a set of points within a shape-adaptive local support, and 2) aggregation, fusing a number of multipoint estimates available for each point. Compared with the guided filter that applies the linear regression to all pixels covered by a fixed-sized square window non-adaptively, the proposed filtering framework is a more generalized form. Two specific filtering methods are instantiated from this framework, based on piecewise constant and piecewise linear modeling, respectively. Leveraging a cross-based local support representation and integration technique, the proposed filtering methods achieve theoretically strong results in an efficient manner, with the two main steps' complexity independent of the filtering kernel size. We demonstrate the strength of the proposed filters in various applications including stereo matching, depth map enhancement, edge-preserving smoothing, color image denoising, detail enhancement, and flash/no-flash denoising.
Jiangbo Lu, Keyang Shi, Dongbo Min, Liang Lin 0004, Minh N. Do
CVPR5
2012 Bayesian Blind Deconvolution with General Sparse Image Priors
S. Derin Babacan, Rafael Molina 0001, Minh N. Do, Aggelos K. Katsaggelos
ECCV (6)3
2012 Weighted mode filtering and its applications to depth video enhancement and coding
abstract
This paper presents a novel approach for improving the quality of depth video. Given a high-quality color image and its corresponding low-quality depth image, we handle various artifacts which may exist on the depth video by applying a weighted mode filtering method based on a joint histogram. When the histogram is generated, the weight based on color similarity between reference and neighboring pixels on the color image is computed and then used for counting each bin on the joint histogram of the depth map. A final solution is determined by seeking a global mode on the histogram. Experimental results show that the proposed method has outstanding performance and is very efficient in various applications such as depth video enhancement and compression.
Dongbo Min, Jiangbo Lu, Minh N. Do
ICASSP4
2012 Efficient edge-preserving interpolation and in-loop filters for depth map compression
abstract
Due to abrupt signal changes on object boundaries, a depth video compressed by conventional video coding standards often introduces serious coding artifacts over the boundaries, which severely affect the quality of a synthesized view. In this paper, we propose an edge-preserving depth interpolation filter based on weighted mode filtering to provide more accurate fractional-pixel samples in the motion-compensated interpolation for an effective inter-coding of the depth video. In addition, an efficient post-processing method is also proposed to further suppress the coding artifacts on the depth video and utilized as an in-loop filter. Experimental results show the proposed methods can significantly improve the synthesized view quality in terms of both objective and subjective measures.
Dongbo Min, Minh N. Do
ICIP3
2012 Efficient video compression methods for a lightweight tele-immersive video chat system
abstract
A lightweight tele-immersive (TI) video chat system named CuteChat has been developed recently to provide a radically new video chat experience by merging each participant in the same shared space, allowing them to interact more naturally in an integrated manner. This paper presents an insight of the coding component in the system. Specifically, we present an efficient standard-compliant method to compress and deliver video contents in a more semantic manner. Taking into account the characteristics of real-life video chat sequences, we also propose a low-complexity coding method to significantly reduce the encoder complexity while retaining acceptable visual quality. Experimental results have shown the effectiveness of the proposed methods in comparison with the existing methods.
Jiangbo Lu, Minh N. Do
ISCAS3
2012 ITEM: immersive telepresence for entertainment and meetings with commodity setup
abstract
This paper presents an Immersive Telepresence system for Entertainment and Meetings (ITEM). The system aims to provide a radically new video communication experience by seamlessly merging participants into the same virtual space to allow a natural interaction among them and shared collaborative contents. With the goal to make a scalable, flexible system for various business solutions as well as easily accessible by massive consumers, we address the challenges in the whole pipeline of media processing, communication, and displaying in our design and realization of such a system. Extensive experiments show the developed system runs reliably and comfortably in real time with a minimal setup requirement (e.g., a webcam, a laptop/desktop connected to the public Internet) for tele-immersive video communication. With such a really minimal deployment requirement, we present a variety of interesting applications and user experiences created by ITEM.
Tien Dung Vu, Hongsheng Yang, Jiangbo Lu, Minh N. Do
ACM Multimedia5
2012 Probabilistic Low-Rank Subspace Clustering
abstract
In this paper, we consider the problem of clustering data points into low-dimensional subspaces in the presence of outliers. We pose the problem using a density estimation formulation with an associated generative model. Based on this probability model, we first develop an iterative expectation-maximization (EM) algorithm and then derive its global solution. In addition, we develop two Bayesian methods based on variational Bayesian (VB) approximation, which are capable of automatic dimensionality selection. While the first method is based on an alternating optimization scheme for all unknowns, the second method makes use of recent results in VB matrix factorization leading to fast and effective estimation. Both methods are extended to handle sparse outliers for robustness and can handle missing values. Experimental results suggest that proposed methods are very effective in clustering and identifying outliers.
S. Derin Babacan, Shinichi Nakajima, Minh N. Do
NIPS3
2012 On the Bandwidth of the Plenoptic Function
abstract
The plenoptic function (POF) provides a powerful conceptual tool for describing a number of problems in image/video processing, vision, and graphics. For example, image-based rendering is shown as sampling and interpolation of the POF. In such applications, it is important to characterize the bandwidth of the POF. We study a simple but representative model of the scene where band-limited signals (e.g., texture images) are "painted" on smooth surfaces (e.g., of objects or walls). We show that, in general, the POF is not band limited unless the surfaces are flat. We then derive simple rules to estimate the essential bandwidth of the POF for this model. Our analysis reveals that, in addition to the maximum and minimum depths and the maximum frequency of painted signals, the bandwidth of the POF also depends on the maximum surface slope. With a unifying formalism based on multidimensional signal processing, we can verify several key results in POF processing, such as induced filtering in space and depth-corrected interpolation, and quantify the necessary sampling rates.
Minh N. Do, Davy Marchand-Maillet, Martin Vetterli
IEEE Trans. Image Process.1
2012 Depth Video Enhancement Based on Weighted Mode Filtering
abstract
This paper presents a novel approach for depth video enhancement. Given a high-resolution color video and its corresponding low-quality depth video, we improve the quality of the depth video by increasing its resolution and suppressing noise. For that, a weighted mode filtering method is proposed based on a joint histogram. When the histogram is generated, the weight based on color similarity between reference and neighboring pixels on the color image is computed and then used for counting each bin on the joint histogram of the depth map. A final solution is determined by seeking a global mode on the histogram. We show that the proposed method provides the optimal solution with respect to L(1) norm minimization. For temporally consistent estimate on depth video, we extend this method into temporally neighboring frames. Simple optical flow estimation and patch similarity measure are used for obtaining the high-quality depth video in an efficient manner. Experimental results show that the proposed method has outstanding performance and is very efficient, compared with existing methods. We also show that the temporally consistent enhancement of depth video addresses a flickering problem and improves the accuracy of depth video.
Dongbo Min, Jiangbo Lu, Minh N. Do
IEEE Trans. Image Process.3
2011 High level synthesis of stereo matching: Productivity, performance, and software constraints
abstract
FPGAs are an attractive platform for applications with high computation demand and low energy consumption requirements. However, design effort for FPGA implementations remains high - often an order of magnitude larger than design effort using high level languages. Instead of this time-consuming process, high level synthesis (HLS) tools generate hardware implementations from high level languages (HLL) such as C/C++/SystemC. Such tools reduce design effort: high level descriptions are more compact and less error prone. HLS tools promise hardware development abstracted from software designer knowledge of the implementation platform. In this paper, we examine several implementations of stereo matching, an active area of computer vision research that uses techniques also common for image de-noising, image retrieval, feature matching and face recognition. We present an unbiased evaluation of the suitability of using HLS for typical stereo matching software, usability and productivity of AutoPilot (a state of the art HLS tool), and the performance of designs produced by AutoPilot. Based on our study, we provide guidelines for software design, limitations of mapping general purpose software to hardware using HLS, and future directions for HLS tool development. For the stereo matching algorithms, we demonstrate between 3.5X and 67.9X speedup over software (but less than achievable by manual RTL design) with a five-fold reduction in design effort vs. manual hardware design.
Kyle Rupnow, Yun Liang 0001, Dongbo Min, Minh N. Do, Deming Chen
FPT5
2011 A revisit to MRF-based depth map super-resolution and enhancement
abstract
This paper presents a Markov Random Field (MRF)-based approach for depth map super-resolution and enhancement. Given a low-resolution or moderate quality depth map, we study the problem of enhancing its resolution or quality with a registered high-resolution color image. Different from the previous methods, this MRF-based approach is based on a novel data term formulation that fits well to the unique characteristics of depth maps. We also discuss a few important design choices that boost the performance of general MRF-based methods. Experimental results show that our proposed approach achieves high resolution depth maps at more desirable quality, both qualitatively and quantitatively. It can also be applied to enhance the depth maps derived with state-of-the-art stereo methods, resulting in the raised ranking based on the Middlebury benchmark.
Jiangbo Lu, Dongbo Min, Ramanpreet Singh Pahwa, Minh N. Do
ICASSP4
2011 A revisit to cost aggregation in stereo matching: How far can we reduce its computational redundancy?
abstract
This paper presents a novel method for performing an efficient cost aggregation in stereo matching. The cost aggregation problem is re-formulated with a perspective of a histogram, and it gives us a potential to reduce the complexity of the cost aggregation significantly. Different from the previous methods which have tried to reduce the complexity in terms of the size of an image and a matching window, our approach focuses on reducing the computational redundancy which exists among the search range, caused by a repeated filtering for all disparity hypotheses. Moreover, we also reduce the complexity of the window-based filtering through an efficient sampling scheme inside the matching window. The trade-off between accuracy and complexity is extensively investigated into parameters used in the proposed method. Experimental results show that the proposed method provides high-quality disparity maps with low complexity. This work provides new insights into complexity-constrained stereo matching algorithm design.
Dongbo Min, Jiangbo Lu, Minh N. Do
ICCV3
2011 CuteChat: a lightweight tele-immersive video chat system
abstract
This paper presents a lightweight tele-immersive video chat system named CuteChat. Based on our recently developed video object cutout technology, the CuteChat system is designed and optimized to provide a radically new video chat experience by merging each participant in the same shared space, allowing them to interact more naturally in an integrated manner. With the goal to make the system easily accessible by massive consumers, we address the challenges in the whole pipeline of video processing, coding, communication, composition, and playback. Extensive experiments have shown that the proposed CuteChat system runs reliably and comfortably in real time on one's laptop or desktop PC, and it needs only a commodity webcam for video acquisition and just public Internet for tele-immersive video conferencing. With such a really minimal deployment requirement, we present a variety of interesting applications and user experiences created by the CuteChat system.
Jiangbo Lu, Zeping Niu, Bhavdeep Singh, Zhiping Luo, Minh N. Do
ACM Multimedia6
2011 Multidimensional Filter Bank Signal Reconstruction From Multichannel Acquisition
abstract
We study the theory and algorithms of an optimal use of multidimensional signal reconstruction from multichannel acquisition by using a filter bank setup. Suppose that we have an N-channel convolution system, referred to as N analysis filters, in M dimensions. Instead of taking all the data and applying multichannel deconvolution, we first reduce the collected data set by an integer M×M uniform sampling matrix [Formula: see text], and then search for a synthesis polyphase matrix which could perfectly reconstruct any input discrete signal. First, we determine the existence of perfect reconstruction (PR) systems for a given set of finite-impulse response (FIR) analysis filters. Second, we present an efficient algorithm to find a sampling matrix with maximum sampling rate and to find a FIR PR synthesis polyphase matrix for a given set of FIR analysis filters. Finally, once a particular FIR PR synthesis polyphase matrix is found, we can characterize all FIR PR synthesis matrices, and then find an optimal one according to design criteria including robust reconstruction in the presence of noise.
Ka Lung Law, Minh N. Do
IEEE Trans. Image Process.2
2011 Analysis of Human Fibroadenomas Using Three-Dimensional Impedance Maps
abstract
Three-dimensional impedance maps (3DZMs) are virtual volumes of acoustic impedance values constructed from histology to represent tissue microstructure acoustically. From the 3DZM, the ultrasonic backscattered power spectrum can be predicted and model based scatterer properties, such as effective scatterer diameter (ESD), can be estimated. Additionally, the 3DZM can be exploited to visualize and identify possible scattering sites, which may aid in the development of more effective scattering models to better represent the ultrasonic interaction with underlying tissue microstructure. In this study, 3DZMs were created from a set of human fibroadenoma samples. ESD estimates were made assuming a fluid-filled sphere form factor model from 3DZMs of volume 300×300×300 μm. For a collection of 33 independent human fibroadenoma tissue samples, the ESD was estimated to be 111±40.7 μm. The 3DZMs were then investigated visually to identify possible scattering sources which conformed to the estimated model scatterer dimensions. This estimation technique allowed a better understanding of the spatial distribution and variability of the estimates throughout the volume.
Alexander J. Dapore, Michael R. King, Josephine M. Harter, Sandhya Sarwate, Michael L. Oelze, James A. Zagzebski, Minh N. Do, Timothy J. Hall, William D. O'Brien Jr.
IEEE Trans. Medical Imaging7
2010 Multi-camera imaging, coding and innovative display: techniques and systems
Minh N. Do, Chang-Su Kim 0001, Karsten Müller 0001, Masayuki Tanimoto, Anthony Vetro
J. Vis. Commun. Image Represent.1
2010 Depth and depth-color coding using shape-adaptive wavelets
Matthieu Maitre, Minh N. Do
J. Vis. Commun. Image Represent.2
2010 On the information rates of the plenoptic function
abstract
Theplenoptic functiondescribes the visual information available to an observer at any point in space and time. Samples of the plenoptic function (POF) are seen in video and in general visual content (images, mosaics, panoramic scenes, etc.), and represent large amounts of information. In this paper, we propose a stochastic model to study the compression limits of a simplified version of the plenoptic function. In the proposed framework, we isolate the two fundamental sources of information in the POF: the one representing the camera motion and the other representing the information complexity of the ¿reality¿ being acquired and transmitted. The sources of information are combined, generating a stochastic process that we study in detail. We first propose a model for ensembles of realities that do not change over time. The proposed model is simple in that it enables us to derive precise coding bounds in the information-theoretic sense that are sharp in a number of cases of practical interest. For this simple case of static realities and camera motion, our results indicate that coding practice is in accordance with optimal coding from an information-theoretic standpoint. The model is further extended to account for visual realities that change over time. We derive bounds on the lossless and lossy information rates for this dynamic reality model, stating conditions under which the bounds are tight. Examples with synthetic sources suggest that within our proposed model, common hybrid coding using motion/displacement estimation with DPCM performs considerably suboptimally relative to the true rate-distortion bound.
Arthur L. da Cunha, Minh N. Do, Martin Vetterli
IEEE Trans. Inf. Theory2
2009 Multidimensional signal reconstruction from multichannel acquisition
abstract
We provide an analysis of the algorithms necessary for the optimal use of multidimensional signal reconstruction from multichannel acquisition. First, we provide computable conditions to test the matrix invertibility and propose algorithms to find a particular inverse. Second, we determine the existence of perfect reconstruction systems for given FIR analysis filters with some sampling matrices and some FIR synthesis polyphase matrices. Then, we present the development of an efficient algorithm designed to find a sampling matrix with maximum sampling rate and FIR synthesis polyphase matrix for given FIR analysis filters so that the system provides a perfect reconstruction. Once a particular synthesis matrix is found, we can characterize all synthesis matrices and find an optimal one according to a design criterion.
Ka Lung Law, Robert M. Fossum, Minh N. Do
ICASSP3
2009 Generic invertibility of multidimensional FIR multirate systems and filter banks
abstract
We study the invertibility of M-variate polynomial (respectively : Laurent polynomial) matrices of size N by P. Such matrices represent multidimensional systems in various settings including filter banks, multiple-input multiple-output systems, and multirate systems. The main result of this paper is to prove that when N - P ges M, then H(z) is generically invertible; whereas when N - P Lt M, then H(z) is generically noninvertible. As a result, we can have an alternative approach in design of the multidimensional systems.
Ka Lung Law, Robert M. Fossum, Minh N. Do
ICASSP3
2009 Two-dimensional geometric lifting
abstract
Wavelets provide a sparse representation for piecewise smooth signals in 1-D; however, separable extensions of wavelets to multiple dimensions do not achieve the same level of sparseness. Recently proposed directional lifting offers transforms sensitive to edges that are not aligned with the coordinate axes, yet the concatenation of separate 1-D slices implicitly assumes independent directional slices and could create large or isotropic support. True 2-D filters and lifting schemes will avoid both of these problems. By aligning the support of the filters with the expected edge, the filters will create fewer non-zero coefficients. Because these filters correspond to interpolation, the theory of Neville filters can automatically determine the coefficients. For images that consist of two bilinear functions divided by a line, geometric lifting demonstrates a 2-4 times reduction of the number of non-zero coefficients compared with the Daubechies order 2 wavelet. In addition, there is a gain of 2.4 dB in nonlinear approximation.
Joshua Blackburn, Minh N. Do
ICIP2
2009 Reconstructing FT-IR spectroscopic imaging data with a sparse prior
abstract
Fourier Transform Infrared (FT-IR) spectroscopic imaging is a potentially valuable tool for diagnosing breast and prostate cancer, but its clinical deployment is limited due to long data acquisition times and vast storage requirements. To counter this limitation, we develop a sparse representation for FT-IR absorbance spectra using a learned dictionary. This sparse representation is used as prior knowledge in regularizing the compressed sensing inverse problem. The data size and acquisition time are directly proportional to the length of the measured signal, namely the interferogram. Hence, we model our measurement process as interferogram truncation, which we implement by low pass filtering and downsampling in the spectral domain. With a downsample factor of four, our reconstruction is adequate for tissue classification and provides a Peak Signal-to-noise Ratio (PSNR) of 41.92 dB, while standard interpolation of the same low resolution measurements can only provide a PSNR of 36.93 dB.
Spencer P. Brady, Minh N. Do, Rohit Bhargava
ICIP2
2009 Depth image-based rendering with low resolution depth
abstract
This paper proposes a new approach for depth image-based rendering (DIBR) with low resolution depth using the 3D propagation algorithm. Our novel depth edge enhancement method efficiently corrects and sharpens the depth edges in the propagated depth image using available high resolution color information. Experimental results show that only with 4% depth information kept for low resolution depth image, the proposed method can provide comparable rendering quality to that of the high resolution case. Furthermore, the proposed work is developed to match the fine-grain parallelism of general-purpose graphics processing units (GPGPUs) and hence can be accelerated to nearly real-time operations in low cost DIBR systems.
Minh N. Do, Sanjay J. Patel
ICIP2
2009 Shape-adaptivewavelet encoding of depth maps
abstract
We present a novel depth-map codec aimed at free-viewpoint 3DTV. The proposed codec relies on a shape-adaptive wavelet transform and an explicit representation of the locations of major depth edges. Unlike classical wavelet transforms, the shape-adaptive transform generates small wavelet coefficients along depth edges, which greatly reduces the data entropy. The wavelet transform is implemented by shape-adaptive lifting, which enables fast computations and perfect reconstruction. We also develop a novel rate-constrained edge detection algorithm, which integrates the idea of significance bitplanes into the Canny edge detector. Along with a simple chain code, it provides an efficient way to extract and encode edges. Experimental results on synthetic and real data confirm the effectiveness of the proposed algorithm, with PSNR gains of 5 dB and more over the Middlebury dataset.
Matthieu Maitre, Minh N. Do
PCS2
2009 MCA: A Multichannel Approach to SAR Autofocus
abstract
We present a new noniterative approach to synthetic aperture radar (SAR) autofocus, termed the multichannel autofocus (MCA) algorithm. The key in the approach is to exploit the multichannel redundancy of the defocusing operation to create a linear subspace, where the unknown perfectly focused image resides, expressed in terms of a known basis formed from the given defocused image. A unique solution for the perfectly focused image is then directly determined through a linear algebraic formulation by invoking an additional image support condition. The MCA approach is found to be computationally efficient and robust and does not require prior assumptions about the SAR scene used in existing methods. In addition, the vector-space formulation of MCA allows sharpness metric optimization to be easily incorporated within the restoration framework as a regularization term. We present experimental results characterizing the performance of MCA in comparison with conventional autofocus methods and discuss the practical implementation of the technique.
Robert L. Morrison Jr., Minh N. Do, David C. Munson Jr.
IEEE Trans. Image Process.2
2009 Error Analysis for Image-Based Rendering With Depth Information
abstract
We propose a new approach to quantitatively analyze the rendering quality of image-based rendering (IBR) algorithms with depth information. The resulting error bounds for synthesized views depend on IBR configurations including the depth and intensity estimate errors, the scene geometry and texture, the number of actual cameras, their positions and resolution. Specifically, the IBR error is bounded by the summation of three terms, highlighting the impact of using multiple actual cameras, the impact of the noise level at the actual cameras, and the impact of the depth accuracy. We also quantify the impact of occlusions and intensity discontinuities. The proposed methodology is applicable to a large class of common IBR algorithms and can be applied locally. Experiments with synthetic and real scenes show that the developed error bounds accurately characterize the rendering errors. In particular, the error bounds correctly characterize the decay rates of synthesized views' mean absolute errors as O(lambda(-1)) and O(lambda(-2)), where lambda is the local density of actual samples, for 2-D and 3-D scenes, respectively. Finally, we discuss the implications of the proposed analysis on camera placement, budget allocation, and bit allocation.
Ha T. Nguyen, Minh N. Do
IEEE Trans. Image Process.2
2009 Joint Estimation and Correction of Geometric Distortions for EPI Functional MRI Using Harmonic Retrieval
abstract
Magnetic resonance imaging (MRI) uses applied spatial variations in the magnetic field to encode spatial position. Therefore, nonuniformities in the main magnetic field can cause image distortions. In order to correct the image distortions, it is desirable to simultaneously acquire data with a field map in registration. We propose a joint estimation (JE) framework with a fast, noniterative approach using harmonic retrieval (HR) methods, combined with a multi-echo echo-planar imaging (EPI) acquisition. The connection with HR establishes an elegant framework to solve the JE problem through a sequence of 1-D HR problems in which efficient solutions are available. We also derive the condition on the smoothness of the field map in order for HR techniques to recover the image with high signal-to-noise ratio. Compared to other dynamic field mapping methods, this method is not constrained by the absolute level of the field inhomogeneity over the slice, but relies on a generous pixel-to-pixel smoothness. Moreover, this method can recover image, field map, and T2* map simultaneously.
Hien M. Nguyen, Bradley P. Sutton, Robert L. Morrison Jr., Minh N. Do
IEEE Trans. Medical Imaging4
2008 Symmetric multi-view stereo reconstruction from planar camera arrays
abstract
We present a novel stereo algorithm which performs surface reconstruction from planar camera arrays. It incorporates the merits of both generic camera arrays and rectified binocular setups, recovering large surfaces like the former and performing efficient computations like the latter. First, we introduce a rectification algorithm which gives freedom in the design of camera arrays and simplifies photometric and geometric computations. We then define a novel set of data-fusion functions over 4-neighborhoods of cameras, which treat all cameras symmetrically and enable standard binocular stereo algorithms to handle arrays with arbitrary number of cameras. In particular, we introduce a photometric fusion function which handles partial visibility and extracts depth information along both horizontal and vertical baselines. Finally, we show that layered depth images and sprites with depth can be efficiently extracted from the rectified 3D space. Experimental results on real images confirm the effectiveness of the proposed method, which reconstructs dense surfaces larger by 20% on Tsukuba.
Matthieu Maitre, Yoshihisa Shinagawa, Minh N. Do
CVPR3
2008 Joint encoding of the depth image based representation using shape-adaptive wavelets
abstract
We present a novel codec of depth-image-based representations for free-viewpoint 3D-TV. The proposed codec relies on a shape-adaptive wavelet transform and an explicit representation of the locations of major depth edges. Unlike classical wavelet transforms, the shape-adaptive transform generates small wavelet coefficients along depth edges, which greatly reduces the data entropy. The codec also shares the edge information between the depth map and the image to reduce their correlation. The wavelet transform is implemented by shape-adaptive lifting, which enables fast computations and perfect reconstruction. Experimental results on real data confirm the superiority of the proposed codec, with PSNR gains of up to 5.46 dB on the depth map and up to 0.19 dB on the image compared to standard wavelet codecs.
Matthieu Maitre, Minh N. Do
ICIP2
2008 Robust multichannel sampling
abstract
A conventional wisdom is that a bandlimited signal can be sampled at twice its maximum frequency to prevent any loss of information. For signals having high frequency components, sampling them requires fast analog-to-digital converters (ADC) that are difficult to design without increasing their cost and noise. In this paper, we show that high-resolution samples of any signal, bandlimited or unband limited, can be accurately approximated using multiple sequences of low-resolution samples taken from the same analog signal, probably with fractional delays, using slow ADCs. The approximation is enabled by designing a set of synthesis filters, without any knowledge of the signals to be sampled, to minimize an induced error system in the minimax sense. The approximation performance is guaranteed to be robust even when using estimates of the system parameters (such as antialiasing filters and fractional delays). We present experiments to confirm the potential of our approach.
Ha T. Nguyen, Minh N. Do
ICIP2
2008 Wavelet-Based Joint Estimation and Encoding of Depth-Image-Based Representations for Free-Viewpoint Rendering
abstract
We propose a wavelet-based codec for the static depth-image-based representation, which allows viewers to freely choose the viewpoint. The proposed codec jointly estimates and encodes the unknown depth map from multiple views using a novel rate-distortion (RD) optimization scheme. The rate constraint reduces the ambiguity of depth estimation by favoring piecewise-smooth depth maps. The optimization is efficiently solved by a novel dynamic programming along trees of integer wavelet coefficients. The codec encodes the image and the depth map jointly to decrease their redundancy and to provide a RD-optimized bitrate allocation between the two. The codec also offers scalability both in resolution and in quality. Experiments on real data show the effectiveness of the proposed codec.
Matthieu Maitre, Yoshihisa Shinagawa, Minh N. Do
IEEE Trans. Image Process.3
2007 A Stochastic Model for Video and its Information Rates
abstract
We propose a stochastic model for video and compute its information rates. The model has two sources of information representing ensembles of camera motion and visual scene data (i.e. "realities"). The sources of information are combined generating a vector process that we study in detail. Both lossless and lossy information rates are derived. The model is further extended to account for realities that change over time. We derive bounds on the lossless and lossy information rates for this dynamic reality model, stating conditions under which the bounds are tight. Experiments with synthetic sources suggest that in the presence of scene motion, simple hybrid coding using motion estimation with DPCM can be suboptimal relative to the true rate-distortion bound
Arthur L. da Cunha, Minh N. Do, Martin Vetterli
DCC2
2007 Finding Optimal Integral Sampling Lattices for a given Frequency Support in Multidimensions
abstract
The search for alias-free sampling lattices for a given frequency support, in particular those lattices achieving minimum sampling densities, is a fundamental issue in various applications of signal and image processing. In this paper, we propose an efficient computational procedure to find all alias-free integral sampling lattices for a given frequency support with minimum sampling density. Central to this algorithm is a novel condition linking the alias-free sampling with the Fourier transform of the indicator function defined on the frequency support. We study the computation of these Fourier transforms based on the divergence theorem, and propose a simple closed-form formula for a fairly general class of support regions consisting of arbitrary N-dimensional polytopes, with polygons in 2D and polyhedra in 3D as special cases. The proposed algorithm can be useful in a variety of applications involving the design of efficient acquisition schemes for multidimensional bandlimited signals.
Yue M. Lu, Minh N. Do
ICIP (2)2
2007 Rate-Distortion Optimal Depth Maps in the Wavelet Domain for Free-Viewpoint Rendering
abstract
We consider the problem of estimating and encoding depth maps from multiple views in the context of 3D-TV with free-viewpoint rendering. We propose a novel codec based on the Rate-Distortion (RD) optimization of the depth-image-based representation (DIBR) in the wavelet domain. The rate constraint enforces the piecewise smoothness of the depth map, which improves the reliability of its estimation. We propose an efficient optimal solution for the joint estimation and coding of the depth map using dynamic programming along the tree of wavelet coefficients. It also provides an automatic bitrate allocation between images and depth maps. Experiments on real data show that the wavelet approach can improve RD performances over a state-of-the-art technique that uses quadtrees.
Matthieu Maitre, Yoshihisa Shinagawa, Minh N. Do
ICIP (5)3
2007 On Two-Channel Filter Banks With Directional Vanishing Moments
abstract
The contourlet transform was proposed to address the limited directional resolution of the separable wavelet transform. One way to guarantee good approximation behavior is to let the directional filters in the contourlet filter bank have sharp frequency response. This requires filters with large support size. We seek to isolate the key filter property that ensures good approximation. In this direction, we propose filters with directional vanishing moments (DVM). These filters, we show, annihilate information along a given direction. We study two-channel filter banks with DVM filters. We provide conditions under which the design of DVM filter banks is possible. A complete characterization of the product filter is, thus, obtained. We propose a design framework that avoids 2-D factorization using the mapping technique. The filters designed, when used in the contourlet transform, exhibit nonlinear approximation comparable to the conventional filters while being shorter and, therefore, providing better visual quality with less ringing artifacts. Furthermore, experiments show that the proposed filters outperform the conventional ones in image approximation and denoising.
Arthur L. da Cunha, Minh N. Do
IEEE Trans. Image Process.2
2007 Multidimensional Directional Filter Banks and Surfacelets
abstract
In 1992, Bamberger and Smith proposed the directional filter bank (DFB) for an efficient directional decomposition of 2-D signals. Due to the nonseparable nature of the system, extending the DFB to higher dimensions while still retaining its attractive features is a challenging and previously unsolved problem. We propose a new family of filter banks, named NDFB, that can achieve the directional decomposition of arbitrary N-dimensional (N > or =2) signals with a simple and efficient tree-structured construction. In 3-D, the ideal passbands of the proposed NDFB are rectangular-based pyramids radiating out from the origin at different orientations and tiling the entire frequency space. The proposed NDFB achieves perfect reconstruction via an iterated filter bank with a redundancy factor of N in N-D. The angular resolution of the proposed NDFB can be iteratively refined by invoking more levels of decomposition through a simple expansion rule. By combining the NDFB with a new multiscale pyramid, we propose the surfacelet transform, which can be used to efficiently capture and represent surface-like singularities in multidimensional data.
Yue M. Lu, Minh N. Do
IEEE Trans. Image Process.2
2007 SAR Image Autofocus By Sharpness Optimization: A Theoretical Study
abstract
Synthetic aperture radar (SAR) autofocus techniques that optimize sharpness metrics can produce excellent restorations in comparison with conventional autofocus approaches. To help formalize the understanding of metric-based SAR autofocus methods, and to gain more insight into their performance, we present a theoretical analysis of these techniques using simple image models. Specifically, we consider the intensity-squared metric, and a dominant point-targets image model, and derive expressions for the resulting objective function. We examine the conditions under which the perfectly focused image models correspond to stationary points of the objective function. A key contribution is that we demonstrate formally, for the specific case of intensity-squared minimization autofocus, the mechanism by which metric-based methods utilize the multichannel defocusing model of SAR autofocus to enforce the stationary point property for multiple image columns. Furthermore, our analysis shows that the objective function has a special separble property through which it can be well approximated locally by a sum of 1-D functions of each phase error component. This allows fast performance through solving a sequence of 1-D optimization problems for each phase component simultaneously. Simulation results using the proposed models and actual SAR imagery confirm that the analysis extends well to realistic situations.
Robert L. Morrison Jr., Minh N. Do, David C. Munson Jr.
IEEE Trans. Image Process.2
2006 On the Information Rate of the Plenoptic Function
abstract
We study the compression problem of visual scenes acquired with a camera for transmission or storage. Our proposed model is general and includes two well-known cases: that of video coding and that of lightfield data compression. Those two examples are related in that both are characterized by two sources of complexity: the camera motion, and the scenes being acquired. The main difference of the two is in how the complexity of the camera is coded. We propose a simplified model which includes the main characteristics of the general problem. Based on this, we do theoretical analysis, develop simple codes, and show experimental results.
Arthur L. da Cunha, Minh N. Do, Martin Vetterli
ICIP2
2006 Tree-Based Orthogonal Matching Pursuit Algorithm for Signal Reconstruction
abstract
Recent studies in linear inverse problems have recognized the sparse representation of unknown signal in a certain basis as an useful and effective prior information to solve those problems. In many multiscale bases (e.g. wavelets), signals of interest (e.g. piecewise-smooth signals) not only have few significant coefficients, but also those significant coefficients are well-organized in trees. We propose to exploit this sparse tree representation as additional prior information for linear inverse problems with limited numbers of measurements. In particular, our proposed algorithm named tree-based orthogonal matching pursuit (TOMP) is shown to provide significant better reconstruction compared to methods that only use sparse representation assumption.
Chinh La, Minh N. Do
ICIP2
2006 A New Contourlet Transform with Sharp Frequency Localization
abstract
The contourlet transform was proposed as a directional multiresolution image representation that can efficiently capture and represent singularities along smooth object boundaries in natural images. Its efficient filter bank construction as well as low redundancy make it an attractive computational framework for various image processing applications. However, a major drawback of the original contourlet construction is that its basis images are not localized in the frequency domain. In this paper, we analyze the cause of this problem, and propose a new contourlet construction as a solution. Instead of using the Laplacian pyramid, we employ a new multiscale decomposition defined in the frequency domain. The resulting basis images are sharply localized in the frequency domain and exhibit smoothness along their main ridges in the spatial domain. Numerical experiments on image denoising show that the proposed new contourlet transform can significantly outperform the original transform both in terms of PSNR (by several dB 's) and in visual quality, while with similar computational complexity.
Yue M. Lu, Minh N. Do
ICIP2
2006 Multichannel Autofocus Algorithm for Synthetic Aperture Radar
abstract
The autofocus problem in synthetic aperture radar (SAR) is considered, where phase errors in the acquired signal data result in imagery that is improperly focused. We present a new non-iterative approach to SAR autofocus, termed the multichannel autofocus (MCA) algorithm, that allows the image focusing operator to be determined directly using a linear algebraic formulation. Specifically, we exploit the multichannel redundancy of the defocusing operation to create a linear subspace framework, where the unknown perfectly-focused image can be expressed in terms of a known basis expansion. By invoking an additional assumption on the underlying image support, the framework becomes sufficiently constrained so that a unique focusing filter can be solved for. The MCA approach is found to be computationally efficient and robust, and does not require prior assumptions about the characteristics of the SAR scene; the performance of previous SAR autofocus techniques relies upon the accuracy of priors such as sharpness metrics or dominant point scatterers. We present experimental results characterizing the performance of MCA in comparison with conventional autofocus methods, and discuss the practical implementation of the technique.
Robert L. Morrison Jr., Minh N. Do
ICIP2
2006 Error Analysis for Image-Based Rendering with Depth Information
abstract
We propose a novel approach to analyze the rendering error of image-based rendering (IBR) algorithms with depth information. We do not use the assumption of band-limitedness as existing approaches. Instead, we use the framework of the propagation algorithm that allows to rigorously analyze the rendering error via the framework of nonuniform interpolation. In this framework, using the depths, we propagate all the intensity information to the virtual cameras, and by doing so, turning the IBR problem into a nonuniform interpolation problem at the virtual image planes. The proposed approach then can systematically analyze the rendering quality for different interpolation methods, including commonly used linear interpolation. We can furthermore analyze the effect of depth estimation error on the rendering quality.
Ha T. Nguyen, Minh N. Do
ICIP2
2006 The Nonsubsampled Contourlet Transform: Theory, Design, and Applications
abstract
In this paper, we develop the nonsubsampled contourlet transform (NSCT) and study its applications. The construction proposed in this paper is based on a nonsubsampled pyramid structure and nonsubsampled directional filter banks. The result is a flexible multiscale, multidirection, and shift-invariant image decomposition that can be efficiently implemented via the à trous algorithm. At the core of the proposed scheme is the nonseparable two-channel nonsubsampled filter bank (NSFB). We exploit the less stringent design condition of the NSFB to design filters that lead to a NSCT with better frequency selectivity and regularity when compared to the contourlet transform. We propose a design framework based on the mapping approach, that allows for a fast implementation based on a lifting or ladder structure, and only uses one-dimensional filtering in some cases. In addition, our design ensures that the corresponding frame elements are regular, symmetric, and the frame is close to a tight one. We assess the performance of the NSCT in image denoising and enhancement applications. In both applications the NSCT compares favorably to other existing methods in the literature.
Arthur L. da Cunha, Jianping Zhou 0001, Minh N. Do
IEEE Trans. Image Process.3
2006 Fast Search for Best Representations in Multitree Dictionaries
abstract
We address the best basis problem--or, more generally, the best representation problem: Given a signal, a dictionary of representations, and an additive cost function, the aim is to select the representation from the dictionary which minimizes the cost for the given signal. We develop a new framework of multitree dictionaries, which includes some previously proposed dictionaries as special cases. We show how to efficiently find the best representation in a multitree dictionary using a recursive tree-pruning algorithm. We illustrate our framework through several examples, including a novel block image coder, which significantly outperforms both the standard JPEG and quadtree-based methods and is comparable to embedded coders such as JPEG2000 and SPIHT.
Yan Huang 0004, Ilya Pollak, Minh N. Do, Charles A. Bouman
IEEE Trans. Image Process.3
2006 Directional multiscale modeling of images using the contourlet transform
abstract
The contourlet transform is a new two-dimensional extension of the wavelet transform using multiscale and directional filter banks. The contourlet expansion is composed of basis images oriented at various directions in multiple scales, with flexible aspect ratios. Given this rich set of basis images, the contourlet transform effectively captures smooth contours that are the dominant feature in natural images. We begin with a detailed study on the statistics of the contourlet coefficients of natural images: using histograms to estimate the marginal and joint distributions and mutual information to measure the dependencies between coefficients. This study reveals the highly non-Gaussian marginal statistics and strong interlocation, interscale, and interdirection dependencies of contourlet coefficients. We also find that conditioned on the magnitudes of their generalized neighborhood coefficients, contourlet coefficients can be approximately modeled as Gaussian random variables. Based on these findings, we model contourlet coefficients using a hidden Markov tree (HMT) model with Gaussian mixtures that can capture all interscale, interdirection, and interlocation dependencies. We present experimental results using this model in image denoising and texture retrieval applications. In denoising, the contourlet HMT outperforms other wavelet methods in terms of visual quality, especially around edges. In texture retrieval, it shows improvements in performance for various oriented textures.
Duncan D.-Y. Po, Minh N. Do
IEEE Trans. Image Process.2
2006 On the Number of Rectangular Tilings
abstract
Adaptive multiscale representations via quadtree splitting and two-dimensional (2-D) wavelet packets, which amount to space and frequency decompositions, respectively, are powerful concepts that have been widely used in applications. These schemes are direct extensions of their one-dimensional counterparts, in particular, by coupling of the two dimensions and restricting to only one possible further partition of each block into four subblocks. In this paper, we consider more flexible schemes that exploit more variations of multidimensional data structure. In the meantime, we restrict to tree-based decompositions that are amenable to fast algorithms and have low indexing cost. Examples of these decomposition schemes are anisotropic wavelet packets, dyadic rectangular tilings, separate dimension decompositions, and general rectangular tilings. We compute the numbers of possible decompositions for each of these schemes. We also give bounds for some of these numbers. These results show that the new rectangular tiling schemes lead to much larger sets of 2-D space and frequency decompositions than the commonly-used quadtree-based schemes, therefore bearing the potential to obtain better representation for a given image.
Minh N. Do
IEEE Trans. Image Process.2
2006 Multidimensional Multichannel FIR Deconvolution Using Gröbner Bases
abstract
We present a new method for general multidimensional multichannel deconvolution with finite impulse response (FIR) convolution and deconvolution filters using Gröbner bases. Previous work formulates the problem of multichannel FIR deconvolution as the construction of a left inverse of the convolution matrix, which is solved by numerical linear algebra. However, this approach requires the prior information of the support of deconvolution filters. Using algebraic geometry and Gröbner bases, we find necessary and sufficient conditions for the existence of exact deconvolution FIR filters and propose simple algorithms to find these deconvolution filters. The main contribution of our work is to extend the previous Gröbner basis results on multidimensional multichannel deconvolution for polynomial or causal filters to general FIR filters. The proposed algorithms obtain a set of FIR deconvolution filters with a small number of nonzero coefficients (a desirable feature in the impulsive noise environment) and do not require the prior information of the support. Moreover, we provide a complete characterization of all exact deconvolution FIR filters, from which good FIR deconvolution filters under the additive white noise environment are found. Simulation results show that our approaches achieve good results under different noise settings.
Jianping Zhou 0001, Minh N. Do
IEEE Trans. Image Process.2
2006 Special paraunitary matrices, Cayley transform, and multidimensional orthogonal filter banks
abstract
We characterize and design multidimensional (MD) orthogonal filter banks using special paraunitary matrices and the Cayley transform. Orthogonal filter banks are represented by paraunitary matrices in the polyphase domain. We define special paraunitary matrices as paraunitary matrices with unit determinant. We show that every paraunitary matrix can be characterized by a special paraunitary matrix and a phase factor. Therefore, the design of paraunitary matrices (and thus of orthogonal filter banks) becomes the design of special paraunitary matrices, which requires a smaller set of nonlinear equations. Moreover, we provide a complete characterization of special paraunitary matrices in the Cayley domain, which converts nonlinear constraints into linear constraints. Our method greatly simplifies the design of MD orthogonal filter banks and leads to complete characterizations of such filter banks.
Jianping Zhou 0001, Minh N. Do, Jelena Kovacevic
IEEE Trans. Image Process.2
2006 Erratum
Jianping Zhou 0001, Minh N. Do, Jelena Kovacevic
IEEE Trans. Image Process.2
2005 Bi-orthogonal filter banks with directional vanishing moments [image representation applications]
abstract
In this paper we study 2D nonseparable filter banks that annihilate information along a certain discrete direction. This is done by having filters with directional vanishing moments (DVM). We study the approximation property of such filters and the design problem providing conditions for its solvability. In particular, we completely characterize the solution and propose a design procedure utilizing the mapping technique. Nonlinear approximation experiments with the contourlet transform indicate that compared with the traditional filters, the new filters designed with DVM provide gains in SNR and visual quality due to their short size.
Arthur L. da Cunha, Minh N. Do
ICASSP (4)2
2005 The finer directional wavelet transform [image processing applications]
abstract
Directional information is an important and unique feature of multidimensional signals. As a result of a separable extension from 1D bases, the multidimensional wavelet transform has very limited directionality. Furthermore, different directions are mixed in certain wavelet subbands. In this paper, we propose a new transform that fixes this frequency mixing problem by using a simple "add-on" to the wavelet transform. In the 2D case, it provides one lowpass subband and six directional highpass subbands at each scale. Just like the wavelet transform, the proposed transform is nonredundant, and can be easily extended to higher dimensions. Though nonseparable in essence, the proposed transform has an efficient implementation based on 1D operations only.
Yue M. Lu, Minh N. Do
ICASSP (4)2
2005 Image-Based Rendering with Depth Information Using the Propagation Algorithm
abstract
The paper proposes a new approach for the image-based rendering (IBR) problem. IBR has many potential applications, such as remote reality and telepresence, in which traditional computer graphic techniques require high computational complexity. Our algorithm proactively propagates all available information from actual cameras to virtual cameras, using a depth availability assumption. This process turns the IBR problem into a nonuniform interpolation problem at the virtual camera image plane, which can be done efficiently at once for all image pixels. Experimental results show the proposed algorithm has low computational complexity and produces accurate rendering, especially around object boundaries, where most existing methods fail.
Ha T. Nguyen, Minh N. Do
ICASSP (2)2
2005 Multichannel FIR exact deconvolution in multiple variables
abstract
We present a general framework for multichannel exact deconvolution with multivariate finite impulse response (FIR) convolution and deconvolution filters using algebraic geometry. Previous work formulates the problem of multichannel FIR deconvolution into that of the left inverse of a convolution matrix which is solved by linear algebra. However, this approach requires the prior information of the support of deconvolution filters. Using algebraic geometry, we find a necessary and sufficient existence condition for FIR deconvolution filters and propose a simple algorithm based on the Gro/spl uml/bner basis to compute deconvolution filters. This computation algorithm obtains deconvolution filters with either minimal order or minimum number of nonzero coefficients, and no prior information of the support is required. Simulation results show that, due to the smaller size of deconvolution filters, our approach achieves better results than the liner algebra approach under an impulsive noise environment.
Jianping Zhou 0001, Minh N. Do
ICASSP (4)2
2005 Nonsubsampled contourilet transform: filter design and applications in denoising
abstract
In this paper we study the nonsubsampled contourlet transform. We address the corresponding filter design problem using the Mc-Clellan transformation. We show how zeroes can be imposed in the filters so that the iterated structure produces regular basis functions. The proposed design framework yields filters that can be implemented efficiently through a lifting factorization. We apply the constructed transform in image noise removal where the results obtained are comparable to the state-of-the art, being superior in some cases.
Arthur L. da Cunha, Jianping Zhou 0001, Minh N. Do
ICIP (1)3
2005 On the bandlimitedness of the plenoptic function
abstract
Image based-rendering (IBR) can be seen as the sampling and reconstruction of the plenoptic function. The question of the minimum sampling rate in IBR can be addressed via spectral analysis of the plenoptic function. We study a model of the scene where bandlimited images are "painted" on surfaces (e.g. of objects or walls). We show that, in general, the plenoptic function is not bandlimited unless the surfaces are flat. We then characterize the spectral decay of the plenoptic function for this model.
Minh N. Do, Davy Marchand-Maillet, Martin Vetterli
ICIP (3)1
2005 Optimal representations in multitree dictionaries with application to compression
abstract
We generalize our results of [Y. Huang et al, 2005 and 2003] and propose a new framework of multitree dictionaries which include many previously proposed dictionaries as well as many new, very large, tree-structured dictionaries. We present an efficient, globally optimal algorithm to find the best tree in such a dictionary. We describe a novel block image coder based on our framework, which is an improvement over our image coder presented in Y. Huang et al, (2005).
Yan Huang 0004, Ilya Pollak, Minh N. Do, Charles A. Bouman
ICIP (1)3
2005 A multichannel approach to metric-based SAR autofocus
abstract
The autofocus problem in synthetic aperture radar (SAR) is considered. We precisely characterize the multichannel nature of the SAR autofocus problem by constructing a low-dimensional sub-space where the perfectly-focused image resides. To obtain a unique solution, we perform a sharpness optimization within this sub-space. The subspace characterization enables the sharpness optimization to be performed in a vector space, which is conceptually simpler, and allows the SAR autofocus problem to be cast into a similar framework with other image restoration problems where fast algorithms and efficient methods have been established. We present experimental results demonstrating that the proposed framework can be used to bring the ideas of blind multichannel deconvolution techniques to the SAR autofocus problem, allowing the dimension of the solution subspace to be reduced further.
Robert L. Morrison Jr., Minh N. Do
ICIP (2)2
2005 Nonsubsampled contourlet transform: construction and application in enhancement
abstract
We present the nonsubsampled contourlet transform and its application in image enhancement. The nonsubsampled contourlet transform is built upon nonsubsampled pyramids and nonsubsampled directional filter banks and provides a shift-invariant directional multiresolution image representation. Existing methods for image enhancement cannot capture the geometric information of images and tend to amplify noises when they are applied to noisy images since they cannot distinguish noises from weak edges. In contrast, the nonsubsampled contourlet transform extracts the geometric information of images, which can be used to distinguish noises from weak edges. Experimental results show the proposed method achieves better enhancement results than a wavelet-based image enhancement method.
Jianping Zhou 0001, Arthur L. da Cunha, Minh N. Do
ICIP (1)3
2005 The contourlet transform: an efficient directional multiresolution image representation
abstract
The limitations of commonly used separable extensions of one-dimensional transforms, such as the Fourier and wavelet transforms, in capturing the geometry of image edges are well known. In this paper, we pursue a "true" two-dimensional transform that can capture the intrinsic geometrical structure that is key in visual information. The main challenge in exploring geometry in images comes from the discrete nature of the data. Thus, unlike other approaches, such as curvelets, that first develop a transform in the continuous domain and then discretize for sampled data, our approach starts with a discrete-domain construction and then studies its convergence to an expansion in the continuous domain. Specifically, we construct a discrete-domain multiresolution and multidirection expansion using nonseparable filter banks, in much the same way that wavelets were derived from filter banks. This construction results in a flexible multiresolution, local, and directional image expansion using contour segments, and, thus, it is named the contourlet transform. The discrete contourlet transform has a fast iterated filter bank algorithm that requires an order N operations for N-pixel images. Furthermore, we establish a precise link between the developed filter bank and the associated continuous-domain contourlet expansion via a directional multiresolution analysis framework. We show that with parabolic scaling and sufficient directional vanishing moments, contourlets achieve the optimal approximation rate for piecewise smooth functions with discontinuities along twice continuously differentiable curves. Finally, we show some numerical experiments demonstrating the potential of contourlets in several image processing applications. Index Terms-Contourlets, contours, filter banks, geometric image processing, multidirection, multiresolution, sparse representation, wavelets.
Minh N. Do, Martin Vetterli
IEEE Trans. Image Process.1
2005 Rate-distortion optimized tree-structured compression algorithms for piecewise polynomial images
abstract
This paper presents novel coding algorithms based on tree-structured segmentation, which achieve the correct asymptotic rate-distortion (R-D) behavior for a simple class of signals, known as piecewise polynomials, by using an R-D based prune and join scheme. For the one-dimensional case, our scheme is based on binary-tree segmentation of the signal. This scheme approximates the signal segments using polynomial models and utilizes an R-D optimal bit allocation strategy among the different signal segments. The scheme further encodes similar neighbors jointly to achieve the correct exponentially decaying R-D behavior (D(R) - c(o)2(-c1R)), thus improving over classic wavelet schemes. We also prove that the computational complexity of the scheme is of O(N log N). We then show the extension of this scheme to the two-dimensional case using a quadtree. This quadtree-coding scheme also achieves an exponentially decaying R-D behavior, for the polygonal image model composed of a white polygon-shaped object against a uniform black background, with low computational cost of O(N log N). Again, the key is an R-D optimized prune and join strategy. Finally, we conclude with numerical results, which show that the proposed quadtree-coding scheme outperforms JPEG2000 by about 1 dB for real images, like cameraman, at low rates of around 0.15 bpp.
Pier Luigi Dragotti, Minh N. Do, Martin Vetterli
IEEE Trans. Image Process.3
2005 Multidimensional orthogonal filter bank characterization and design using the Cayley transform
abstract
We present a complete characterization and design of orthogonal infinite impulse response (IIR) and finite impulse response (FIR) filter banks in any dimension using the Cayley transform (CT). Traditional design methods for one-dimensional orthogonal filter banks cannot be extended to higher dimensions directly due to the lack of a multidimensional (MD) spectral factorization theorem. In the polyphase domain, orthogonal filter banks are equivalent to paraunitary matrices and lead to solving a set of nonlinear equations. The CT establishes a one-to-one mapping between paraunitary matrices and para-skew-Hermitian matrices. In contrast to the paraunitary condition, the para-skew-Hermitian condition amounts to linear constraints on the matrix entries which are much easier to solve. Based on this characterization, we propose efficient methods to design MD orthogonal filter banks and present new design results for both IIR and FIR cases.
Jianping Zhou 0001, Minh N. Do, Jelena Kovacevic
IEEE Trans. Image Process.2
2004 Toward sound-based synthesis: the far-field case
abstract
We consider the problem of synthesizing the sound at any desired position and time from the recording of a set of microphones. Similar to the image-based rendering approach for vision, we propose a sound-based synthesis approach for sound. In this approach, audio signals at new positions are interpolated directly from the recorded signals of nearby microphones. The key underlying problems for sound-based synthesis are sampling and reconstruction of the sound field. We provide a spectral analysis of the sound field under the far-field assumption. Based on this analysis, we derive the minimum sampling and optimal reconstruction for several common settings.
Minh N. Do
ICASSP (2)1
2004 New algorithms for best local cosine basis search
abstract
We propose a best basis search algorithm for local cosine dictionaries. We improve upon the classical best local cosine basis selection based on a dyadic tree (Coifman, R.R. and Wickerhauser, M.V., IEEE Trans. Inf. Th., vol.38, no.2, p.713-18, 1992), by considering a larger dictionary of bases. This results in more compact representations, lower costs, and approximate shift-invariance. We also provide a version of our algorithm which is strictly shift-invariant.
Yan Huang 0004, Ilya Pollak, Charles A. Bouman, Minh N. Do
ICASSP (2)4
2004 A geometrical approach to sampling signals with finite rate of innovation
abstract
Many signals of interest can be characterized by a finite number of parameters per unit of time. Instead of spanning a single linear space, these signals often lie on a union of spaces. Under this setting, traditional sampling schemes are either inapplicable or very inefficient. We present a framework for sampling these signals based on an injective projection operator, which "flattens" the signals down to a common low dimensional representation space while still preserving all the information. Standard sampling procedures can then be applied on that space. We show the necessary and sufficient conditions for such operators to exist and provide the minimum sampling rate for the representation space, which indicates the efficiency of this framework. These results provide a new perspective on the sampling of signals with finite rate of innovation and can serve as a guideline for designing new algorithms for a class of problems in signal processing and communications.
Yue M. Lu, Minh N. Do
ICASSP (2)2
2003 Fast approximation of Kullback-Leibler distance for dependence trees and hidden Markov models
abstract
We present a fast algorithm to approximate the Kullback-Leibler distance (KLD) between two dependence tree models. The algorithm uses the "upward" (or "forward") procedure to compute an upper bound for the KLD. For hidden Markov models, this algorithm is reduced to a simple expression. Numerical experiments show that for a similar accuracy, the proposed algorithm offers a saving of hundreds of times in computational complexity compared to the commonly used Monte Carlo method. This makes the proposed algorithm important for real-time applications, such as image retrieval.
Minh N. Do
IEEE Signal Process. Lett.1
2003 The finite ridgelet transform for image representation
abstract
The ridgelet transform was introduced as a sparse expansion for functions on continuous spaces that are smooth away from discontinuities along lines. We propose an orthonormal version of the ridgelet transform for discrete and finite-size images. Our construction uses the finite Radon transform (FRAT) as a building block. To overcome the periodization effect of a finite transform, we introduce a novel ordering of the FRAT coefficients. We also analyze the FRAT as a frame operator and derive the exact frame bounds. The resulting finite ridgelet transform (FRIT) is invertible, nonredundant and computed via fast algorithms. Furthermore, this construction leads to a family of directional and orthonormal bases for images. Numerical results show that the FRIT is more effective than the wavelet transform in approximating and denoising images with straight edges.
Minh N. Do, Martin Vetterli
IEEE Trans. Image Process.1
2002 Contourlets: a directional multiresolution image representation
abstract
We propose a new scheme, named contourlet, that provides a flexible multiresolution, local and directional image expansion. The contourlet transform is realized efficiently via a double iterated filter bank structure. Furthermore, it can be designed to satisfy the anisotropy scaling relation for curves, and thus offers a fast and structured curvelet-like decomposition. As a result, the contourlet transform provides a sparse representation for two-dimensional piecewise smooth signals resembling images. Finally, we show some numerical experiments demonstrating the potential of contourlets in several image processing tasks.
Minh N. Do, Martin Vetterli
ICIP (1)1
2002 Improved quadtree algorithm based on joint coding for piecewise smooth image compression
abstract
We present a novel coding algorithm based on the tree structured segmentation, which achieves oracle like rate-distortion (R-D) behavior for a simple class of signals, namely piecewise polynomials in the high bit rate regime. We consider a R-D optimization framework, which employs optimal bit allocation strategy among different signal segments to achieve the best tradeoff between description complexity and approximation quality. First, we describe the basic idea of the algorithm for the 1D case. It can be shown that the proposed compression algorithm based on an optimal binary tree segmentation achieves the oracle like R-D behavior (D(R)/spl sim/c/sub 0/2/sup -c1R/) with the computational cost of the order O(NlogN). We then show the extension of the scheme to the 2D case with the similar R-D behavior without sacrificing the computational ease. Finally, we conclude with some experimental results.
Pier Luigi Dragotti, Minh N. Do, Martin Vetterli
ICME (1)3
2002 Rate-distortion optimized tree based coding algorithms
abstract
This paper addresses the problem of efficient coding of an important class of signals, namely piecewise polynomials. For this signal class, we develop a coding algorithm, which achieves oracle like rate-distortion (R-D) behavior in the high bit rate regime and with a reasonable computational complexity. For the 1-D case, our scheme is based on the binary tree segmentation of the signal and an optimal bit allocation strategy among the different signal segments. The scheme further encodes the similar neighbors jointly to achieve the right exponentially decaying R-D behavior (D(R) /spl sim/ c/sub 0/2/sup -c1R/). We have also shown that the computational cost of the scheme is of the order O(N log N). We then show that the scheme can be easily extended to the 2-D case, as the quadtree based coding scheme, with the similar R-D behavior and computational cost. Finally, we conclude with some numerical results.
Pier Luigi Dragotti, Minh N. Do, Martin Vetterli
ITW3
2002 Wavelet-based texture retrieval using generalized Gaussian density and Kullback-Leibler distance
abstract
We present a statistical view of the texture retrieval problem by combining the two related tasks, namely feature extraction (FE) and similarity measurement (SM), into a joint modeling and classification scheme. We show that using a consistent estimator of texture model parameters for the FE step followed by computing the Kullback-Leibler distance (KLD) between estimated models for the SM step is asymptotically optimal in term of retrieval error probability. The statistical scheme leads to a new wavelet-based texture retrieval method that is based on the accurate modeling of the marginal distribution of wavelet coefficients using generalized Gaussian density (GGD) and on the existence a closed form for the KLD between GGDs. The proposed method provides greater accuracy and flexibility in capturing texture information, while its simplified form has a close resemblance with the existing methods which uses energy distribution in the frequency domain to identify textures. Experimental results on a database of 640 texture images indicate that the new method significantly improves retrieval rates, e.g., from 65% to 77%, compared with traditional approaches, while it retains comparable levels of computational complexity.
Minh N. Do, Martin Vetterli
IEEE Trans. Image Process.1
2002 Rotation invariant texture characterization and retrieval using steerable wavelet-domain hidden Markov models
abstract
We present a statistical model for characterizing texture images based on wavelet-domain hidden Markov models. With a small number of parameters, the new model captures both the subband marginal distributions and the dependencies across scales and orientations of the wavelet descriptors. Applied to the steerable pyramid, once it is trained for an input texture image, the model can be easily steered to characterize that texture at any other orientation. Furthermore, after a diagonalization operation, we obtain a rotation-invariant model of the texture image. We also propose a fast algorithm to approximate the Kullback-Leibler distance between two wavelet-domain hidden Markov models. We demonstrate the effectiveness of the new texture models in retrieval experiments with large image databases, where significant improvements are shown.
Minh N. Do, Martin Vetterli
IEEE Trans. Multim.1
2001 Frame reconstruction of the Laplacian pyramid
abstract
We study the Laplacian pyramid (LP) as a frame operator, and this reveals that the usual reconstruction is suboptimal. With orthogonal filters, the LP is shown to be a tight frame, thus the optimal linear reconstruction using the dual frame operator has a simple structure as symmetrical with the forward transform. For more general cases, we propose an efficient filter bank for reconstruction in the LP that is shown to perform better than the usual method. Numerical results indicate that gains of more than 1 dB are actually achieved.
Minh N. Do, Martin Vetterli
ICASSP1
2001 Pyramidal directional filter banks and curvelets
abstract
A flexible multiscale and directional representation for images is proposed. The scheme combines directional filter banks with the Laplacian pyramid to provide a sparse representation for two-dimensional piecewise smooth signals resembling images. The underlying expansion is a frame and can be designed to be a tight frame. Pyramidal directional filter banks provide an effective method to implement the digital curvelet transform. The regularity issue of the iterated filters in the directional filter bank is examined.
Minh N. Do, Martin Vetterli
ICIP (3)1
2001 On the compression of two-dimensional piecewise smooth functions
abstract
It is well known that wavelets provide good non-linear approximation of one-dimensional (1-D) piecewise smooth functions. However, it has been shown that the use of a basis with good approximation properties does not necessarily lead to a good compression algorithm. The situation in 2-D is much more complicated since wavelets are not good for modeling piecewise smooth signals (where discontinuities are along smooth curves). The purpose of this work is to analyze the performance of compression algorithms for 2-D piecewise smooth functions directly in a rate distortion context. We consider some simple image models and compute rate distortion bounds achievable using oracle based methods. We then present a practical compression algorithm based on optimal quadtree decomposition that, in some cases, achieve the oracle performance.
Pier Luigi Dragotti, Minh N. Do, Martin Vetterli
ICIP (1)2
2000 Orthonormal Finite Ridgelet Transform for Image Compression
abstract
A finite implementation of the ridgelet transform is presented. The transform is invertible, non-redundant and achieved via fast algorithms. Furthermore we show that this transform is orthogonal hence it allows one to use non-linear approximations for the representation of images. Numerical results on different test images are shown. Those results conform with the theory of the ridgelet transform in the continuous domain-the obtained representation can represent efficiently images with linear singularities. Thus it indicates the potential of the proposed system as a new transform for coding of images.
Minh N. Do, Martin Vetterli
ICIP1
2000 Texture Similarity Measurement Using Kullback-Leibler Distance on Wavelet Subbands
abstract
LCAV
Minh N. Do, Martin Vetterli
ICIP1