VLDB 2026 Research / reviewers in the wild / expert
Hien Van Nguyen
dblp:59/9550
· DBLP profile ↗
38ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Contrastive Integrated Gradients: A Feature Attribution-Based Method for Explaining Whole Slide Image ClassificationabstractInterpretability is essential in Whole Slide Image (WSI) analysis for computational pathology, where understanding model predictions helps build trust in AI-assisted diagnostics. While Integrated Gradients (IG) and related attribution methods have shown promise, applying them directly to WSIs introduces challenges due to their high-resolution nature. These methods capture model decision patterns but may overlook class-discriminative signals that are crucial for distinguishing between tumor subtypes. In this work, we introduce Contrastive Integrated Gradients (CIG), a novel attribution method that enhances interpretability by computing contrastive gradients in logit space. First, CIG highlights class-discriminative regions by comparing feature importance relative to a reference class, offering sharper differentiation between tumor and non-tumor areas. Second, CIG satisfies the axioms of integrated attribution, ensuring consistency and theoretical soundness. Third, we propose two attribution quality metrics, MIL-AIC and MIL-SIC, which measure how predictive information and model confidence evolve with access to salient regions, particularly under weak supervision. We validate CIG across three datasets spanning distinct cancer types: CAMELYON16 (breast cancer metastasis in lymph nodes), TCGA-RCC (renal cell carcinoma), and TCGA-Lung (lung cancer). Experimental results demonstrate that CIG yields more informative attributions both quantitatively, using MIL-AIC and MIL-SIC, and qualitatively, through visualizations that align closely with ground truth tumor regions, underscoring its potential for interpretable and trustworthy WSI-based diagnostics Anh Mai Vu, Tuan L. Vo, Ngoc Lam Quang Bui, Nam N. B. Le, Akash Awasthi, Huy Quoc Vo, Thanh-Huy Nguyen, Zhu Han 0001, Chandra Mohan, Hien Van Nguyen |
WACV | 10 |
| 2025 | MAARTA:Multi-agentic Adaptive Radiology Teaching Assistant
Akash Awasthi, Brandon V. Chung, Anh M. Vu, T. Hoang Ngan Le, Rishi Agrawal, Zhigang Deng 0001, Carol C. Wu, Hien Van Nguyen |
MICCAI (5) | 8 |
| 2025 | Frequency Strikes Back: Boosting Parameter-Efficient Foundation Model Adaptation for Medical Imaging
Son T. Ly, Hien Van Nguyen |
MICCAI (6) | 2 |
| 2025 | ItpCtrl-AI: End-to-end interpretable and controllable artificial intelligence by modeling radiologists' intentions
Trong-Thang Pham, Jacob Brecheisen, Carol C. Wu, Hien Van Nguyen, Zhigang Deng 0001, Donald A. Adjeroh, Gianfranco Doretto, Arabinda Choudhary, T. Hoang Ngan Le |
Artif. Intell. Medicine | 4 |
| 2025 | Structural chain of thoughts for radiology education
Akash Awasthi, Brandon Chung, Anh M. Vu, Saba Khan, T. Hoang Ngan Le, Zhigang Deng 0001, Rishi Agrawal, Carol C. Wu, Hien Van Nguyen |
Knowl. Based Syst. | 9 |
| 2025 | Frozen Large-Scale Pretrained Vision-Language Models are the Effective Foundational Backbone for Multimodal Breast Cancer PredictionabstractBreast cancer is a pervasive global health concern among women. Leveraging multimodal data from enterprise patient databases-including Picture Archiving and Communication Systems (PACS) and Electronic Health Records (EHRs)-holds promise for improving prediction. This study introduces a multimodal deep-learning model leveraging mammogram datasets to evaluate breast cancer prediction. Our approach integrates frozen large-scale pretrained vision-language models, showcasing superior performance and stability compared to traditional image-tabular models across two public breast cancer datasets. The model consistently outperforms conventional full fine-tuning methods by using frozen pretrained vision-language models alongside a lightweight trainable classifier. The observed improvements are significant. In the CBIS-DDSM dataset, the Area Under the Curve (AUC) increases from 0.867 to 0.902 during validation and from 0.803 to 0.830 for the official test set. Within the EMBED dataset, AUC improves from 0.780 to 0.805 during validation. In scenarios with limited data, using Breast Imaging-Reporting and Data System category three (BI-RADS 3) cases, AUC improves from 0.91 to 0.96 on the official CBIS-DDSM test set and from 0.79 to 0.83 on a challenging validation set. This study underscores the benefits of vision-language models in jointly training diverse image-clinical datasets from multiple healthcare institutions, effectively addressing challenges related to non-aligned tabular features. Combining training data enhances breast cancer prediction on the EMBED dataset, outperforming all other experiments. In summary, our research emphasizes the efficacy of frozen large-scale pretrained vision-language models in multimodal breast cancer prediction, offering superior performance and stability over conventional methods, reinforcing their potential for breast cancer prediction. Hung Q. Vo, Lin Wang 0065, Kelvin K. Wong, Chika F. Ezeana, Xiaohui Yu 0003, Jenny C. Chang, Hien Van Nguyen, Stephen T. C. Wong |
IEEE J. Biomed. Health Informatics | 8 |
| 2024 | Anomaly Detection in Satellite Videos Using Diffusion ModelsabstractDetecting anomalies in videos is a fundamental challenge in machine learning, particularly for applications like disaster management. Leveraging satellite data, with its high frequency and wide coverage, proves invaluable for promptly identifying extreme events such as wildfires, cyclones, or floods. Geostationary satellites, providing data streams at frequent intervals, effectively create a continuous video feed of Earth from space. This study focuses on detecting anomalies, specifically wildfires and smoke, in these high-frequency satellite videos. In contrast to prior endeavors in anomaly detection within surveillance videos, this study introduces a system tailored for high-frequency satellite videos, placing particular emphasis on two anomalies. Unlike the majority of existing Convolution Neural Network-based methods for wildfire detection that rely on labeled images or videos, our unsupervised approach addresses the challenges posed by high-frequency satellite videos with a high intensity of clouds. These Convolution Neural Network-based methods can only identify fires once they have reached a certain size and are susceptible to false positives. We frame the challenge of wildfire detection as a general anomaly detection problem. Introducing an innovative unsupervised approach involving diffusion models, which are state-of-the-art generative models for anomaly detection in satellite videos, we adopt a “generating-to-detecting” strategy. Performance evaluation, measured through AUC-ROC, underscores the superior efficacy of the diffusion model over CNN and Generative Adversarial Networks-based methods in detecting anomalies in these high-frequency satellite videos characterized by a high intensity of clouds. The dataset utilized can be accessed at this location. Akash Awasthi, Son T. Ly, Jaer Nizam, Videet Mehta, Safwan Ahmad, Ramakrishna Nemani, Saurabh Prasad, Hien Van Nguyen |
MMSP | 8 |
| 2024 | Cellular data extraction from multiplexed brain imaging data using self-supervised Dual-loss Adaptive Masked Autoencoder
Son T. Ly, Bai Lin, Hung Q. Vo, Dragan Maric, Badrinath Roysam, Hien Van Nguyen |
Artif. Intell. Medicine | 6 |
| 2022 | Removal of Confounders via Invariant Risk Minimization for Medical Diagnosis
Samira Zare, Hien Van Nguyen |
MICCAI (8) | 2 |
| 2022 | Self-Supervised Learning for Efficient Antialiasing Seismic Data InterpolationabstractReconstruction of seismic data is an important but challenging task in seismic data processing. Different machine-learning-based algorithms have been developed to solve this ill-posed problem and achieved great progress. However, most machine-learning-based methods rely on supervised learning where a good training dataset with many complete shot-gathers are required to train the model. Although the generative model has been used for unsupervised learning and reconstructing signals in a shot-gather, it fails to accurately resolve the fine features, especially when aliasing is the main concern. In addition, multiple shots’ interpolation problems have not been fully investigated by the unsupervised machine-learning-based approaches. In this work, we propose a self-supervised learning method using a blind-trace network and two antialiasing techniques (automatic spectrum suppression and mix-training) for seismic data reconstruction. The method is validated using challenging and realistic scenarios. Test results show that the method can be applied to single-shot or multiple shots’ cases and adapt well to different decimation patterns. Pengyu Yuan, Shirui Wang, Wenyi Hu, Prashanth Nadukandi, German Ocampo Botero, Xuqing Wu 0001, Hien Van Nguyen, Jiefu Chen |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2021 | PairFlow: Enhancing Portable Chest X-Ray By Flow-Based Deformation For Covid-19 DiagnosingabstractThis work aims to assist physicians improve their speed and diagnostic accuracy when interpreting portable CXR (p_CXR), which are in especially high demand in the setting of the ongoing COVID-19 pandemic. In this paper, we introduce new deep learning frameworks, named Pair-Flow, to align and enhance the quality of p_CXR to be more consistent, and to more closely match higher quality conventional CXR (c_CXR). The contributions of this work are four folds. Firstly, a new database collection of subject-pair CXR is introduced and available to download. Secondly, a new deep learning-based alignment approach is presented to align subject-pairs dataset to obtain pixel-pairs dataset. Thirdly, a new Pair-Flow approach, an end-to-end invertible transfer deep learning method, to enhance the degraded quality of p_CXR. Finally, the performance of the proposed system is evaluated at both image quality and topological properties. T. Hoang Ngan Le, James Sorensen, Toan Duc Bui, Arabinda Choudhary, Khoa Luu, Hien Van Nguyen |
ICIP | 6 |
| 2021 | MorphSet: Improving Renal Histopathology Case Assessment Through Learned Prognostic Vectors
Pietro Antonio Cicalese, Syed Asad Rizvi, Victor Wang, Sai Patibandla, Pengyu Yuan, Samira Zare, Katharina Moos, Ibrahim Batal, Marian Clahsen-van Groningen, Candice Roufosse, Jan Ulrich Becker, Chandra Mohan, Hien Van Nguyen |
MICCAI (8) | 13 |
| 2021 | Adaptive Privacy Preserving Deep Learning Algorithms for Medical DataabstractDeep learning holds a great promise of revolutionizing healthcare and medicine. Unfortunately, various inference attack models demonstrated that deep learning puts sensitive patient information at risk. The high capacity of deep neural networks is the main reason behind the privacy loss. In particular, patient information in the training data can be unintentionally memorized by a deep network. Adversarial parties can extract that information given the ability to access or query the network. In this paper, we propose a novel privacy-preserving mechanism for training deep neural networks. Our approach adds decaying Gaussian noise to the gradients at every training iteration. This is in contrast to the mainstream approach adopted by Google's TensorFlow Privacy, which employs the same noise scale in each step of the whole training process. Compared to existing methods, our proposed approach provides an explicit closed-form mathematical expression to approximately estimate the privacy loss. It is easy to compute and can be useful when the users would like to decide proper training time, noise scale, and sampling ratio during the planning phase. We provide extensive experimental results using one real-world medical dataset (chest radiographs from the CheXpert dataset) to validate the effectiveness of the proposed approach. The proposed differential privacy based deep learning model achieves significantly higher classification accuracy over the existing methods with the same privacy budget. Xinyue Zhang 0001, Jiahao Ding, Maoqiang Wu, Stephen T. C. Wong, Hien Van Nguyen, Miao Pan |
WACV | 5 |
| 2021 | Deep reinforcement learning in medical imaging: A literature review
Shaohua Kevin Zhou, T. Hoang Ngan Le, Khoa Luu, Hien Van Nguyen, Nicholas Ayache |
Medical Image Anal. | 4 |
| 2021 | Kidney Level Lupus Nephritis Classification Using Uncertainty Guided Bayesian Convolutional Neural NetworksabstractThe kidney biopsy based diagnosis of Lupus Nephritis (LN) is characterized by low inter-observer agreement, with misdiagnosis being associated with increased patient morbidity and mortality. Although various Computer Aided Diagnosis (CAD) systems have been developed for other nephrohistopathological applications, little has been done to accurately classify kidneys based on their kidney level Lupus Glomerulonephritis (LGN) scores. The successful implementation of CAD systems has also been hindered by the diagnosing physician's perceived classifier strengths and weaknesses, which has been shown to have a negative effect on patient outcomes. We propose an Uncertainty-Guided Bayesian Classification (UGBC) scheme that is designed to accurately classify control, class I/II, and class III/IV LGN (3 class) at both the glomerular-level classification task (26,634 segmented glomerulus images) and the kidney-level classification task (87 MRL/lpr mouse kidney sections). Data annotation was performed using a high throughput, bulk labeling scheme that is designed to take advantage of Deep Neural Network's (or DNNs) resistance to label noise. Our augmented UGBC scheme achieved a 94.5% weighted glomerular-level accuracy while achieving a weighted kidney-level accuracy of 96.6%, improving upon the standard Convolutional Neural Network (CNN) architecture by 11.8% and 3.5% respectively. Pietro Antonio Cicalese, Aryan Mobiny, Zahed Shahmoradi, Xiongfeng Yi, Chandra Mohan, Hien Van Nguyen |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Memory-Augmented Capsule Network for Adaptable Lung Nodule ClassificationabstractComputer-aided diagnosis (CAD) systems must constantly cope with the perpetual changes in data distribution caused by different sensing technologies, imaging protocols, and patient populations. Adapting these systems to new domains often requires significant amounts of labeled data for re-training. This process is labor-intensive and time-consuming. We propose a memory-augmented capsule network for the rapid adaptation of CAD models to new domains. It consists of a capsule network that is meant to extract feature embeddings from some high-dimensional input, and a memory-augmented task network meant to exploit its stored knowledge from the target domains. Our network is able to efficiently adapt to unseen domains using only a few annotated samples. We evaluate our method using a large-scale public lung nodule dataset (LUNA), coupled with our own collected lung nodules and incidental lung nodules datasets. When trained on the LUNA dataset, our network requires only 30 additional samples from our collected lung nodule and incidental lung nodule datasets to achieve clinically relevant performance (0.925 and 0.891 area under receiving operating characteristic curves (AUROC), respectively). This result is equivalent to using two orders of magnitude less labeled training data while achieving the same performance. We further evaluate our method by introducing heavy noise, artifacts, and adversarial attacks. Under these severe conditions, our network's AUROC remains above 0.7 while the performance of state-of-the-art approaches reduce to chance level. Aryan Mobiny, Pengyu Yuan, Pietro Antonio Cicalese, Supratik Moulik, Carol C. Wu, Kelvin K. Wong, Stephen T. C. Wong, Tiancheng He, Hien Van Nguyen |
IEEE Trans. Medical Imaging | 10 |
| 2020 | Multiple Class Novelty Detection Under Data Distribution Shift
Poojan Oza, Hien Van Nguyen, Vishal M. Patel |
ECCV (7) | 2 |
| 2020 | StyPath: Style-Transfer Data Augmentation for Robust Histology Image Classification
Pietro Antonio Cicalese, Aryan Mobiny, Pengyu Yuan, Jan Ulrich Becker, Chandra Mohan, Hien Van Nguyen |
MICCAI (5) | 6 |
| 2020 | DECAPS: Detail-Oriented Capsule Networks
Aryan Mobiny, Pengyu Yuan, Pietro Antonio Cicalese, Hien Van Nguyen |
MICCAI (1) | 4 |
| 2020 | Few Is Enough: Task-Augmented Active Meta-learning for Brain Cell Classification
Pengyu Yuan, Aryan Mobiny, Jahandar Jahanipour, Pietro Antonio Cicalese, Badrinath Roysam, Vishal M. Patel, Dragan Maric, Hien Van Nguyen |
MICCAI (1) | 9 |
| 2020 | Automated Classification of Apoptosis in Phase Contrast Microscopy Using Capsule NetworkabstractAutomatic and accurate classification of apoptosis, or programmed cell death, will facilitate cell biology research. The state-of-the-art approaches in apoptosis classification use deep convolutional neural networks (CNNs). However, these networks are not efficient in encoding the part-whole relationships, thus requiring a large number of training samples to achieve robust generalization. This paper proposes an efficient variant of capsule networks (CapsNets) as an alternative to CNNs. Extensive experimental results demonstrate that the proposed CapsNets achieve competitive performances in target cell apoptosis classification, while significantly outperforming CNNs when the number of training samples is small. To utilize temporal information within microscopy videos, we propose a recurrent CapsNet constructed by stacking a CapsNet and a bi-directional long short-term recurrent structure. Our experiments show that when considering temporal constraints, the recurrent CapsNet achieves 93.8% accuracy and makes significantly more consistent prediction than NNs. Aryan Mobiny, Hengyang Lu, Hien Van Nguyen, Badrinath Roysam, Navin Varadarajan |
IEEE Trans. Medical Imaging | 3 |
| 2019 | Surgical Activities Recognition Using Multi-scale Recurrent NetworksabstractRecently, surgical activity recognition has been receiving significant attention from the medical imaging community. Existing state-of-the-art approaches employ recurrent neural networks such as long-short term memory networks (LSTMs). However, our experiments show that these networks are not effective in capturing the relationship of features with different temporal scales. Such limitation will lead to sub-optimal recognition performance of surgical activities containing complex motions at multiple time scales. To overcome this shortcoming, our paper proposes a multi-scale recurrent neural network (MS-RNN) that combines the strength of both wavelet scattering operations and LSTM. We validate the effectiveness of the proposed network using both real and synthetic datasets. Our experimental results show that MS-RNN outperforms state-of-the-art methods in surgical activity recognition by a significant margin. On a synthetic dataset, the proposed network achieves more than 90% classification accuracy while LSTM's accuracy is around chance level. Experiments on real surgical activity dataset shows a significant improvement of recognition accuracy over the current state of the art (90.2% versus 83.3%). Ilker Gurcan, Hien Van Nguyen |
ICASSP | 2 |
| 2018 | Fast CapsNet for Lung Cancer Screening
Aryan Mobiny, Hien Van Nguyen |
MICCAI (2) | 2 |
| 2017 | Stacked LSTM Deep Learning Model for Traffic Prediction in Vehicle-to-Vehicle CommunicationabstractVehicle-to-Vehicle (V2V) communication becomes an emerging topic because of its capability to provide efficient solution which guarantees more pleasant driving environment and eliminates the possibility of traffic accidents. However, the limitation of resource in V2V communication determines that a dynamic resource allocation strategy must be implemented to provide a balanced communication resource usage. Instead of focusing on the small topology of vehicular wireless communication, we look at a bigger picture of the scenery to deal with the challenge of limitation of resource in V2V communication. In this paper, we propose a long short-term memory (LSTM) based regression model to predict 24-hour traffic counts data. The main steps of our work are as follow: First, we collect 24-hour traffic counts data online and label those data. Second, we construct a stacked LSTM model to implement regression. Third, compared with the performance of logistic regression, the efficiency of our regression model is found out. Finally, we analyze the potential resource allocation patterns according to the regression results. Xunsheng Du, Huaqing Zhang 0001, Hien Van Nguyen, Zhu Han 0001 |
VTC Fall | 3 |
| 2015 | Unsupervised Cross-Modal Synthesis of Subject-Specific ScansabstractRecently, cross-modal synthesis of subject-specific scans has been receiving significant attention from the medical imaging community. Though various synthesis approaches have been introduced in the recent past, most of them are either tailored to a specific application or proposed for the supervised setting, i.e., they assume the availability of training data from the same set of subjects in both source and target modalities. But, collecting multiple scans from each subject is undesirable. Hence, to address this issue, we propose a general unsupervised cross-modal medical image synthesis approach that works without paired training data. Given a source modality image of a subject, we first generate multiple target modality candidate values for each voxel independently using cross-modal nearest neighbor search. Then, we select the best candidate values jointly for all the voxels by simultaneously maximizing a global mutual information cost function and a local spatial consistency cost function. Finally, we use coupled sparse representation for further refinement of synthesized images. Our experiments on generating T1-MRI brain scans from T2-MRI and vice versa demonstrate that the synthesis capability of the proposed unsupervised approach is comparable to various state-of-the-art supervised approaches in the literature. Raviteja Vemulapalli, Hien Van Nguyen, Shaohua Kevin Zhou |
ICCV | 2 |
| 2015 | Cross-Domain Synthesis of Medical Images Using Efficient Location-Sensitive Deep Network
Hien Van Nguyen, Shaohua Kevin Zhou, Raviteja Vemulapalli |
MICCAI (1) | 1 |
| 2015 | DASH-N: Joint Hierarchical Domain Adaptation and Feature LearningabstractComplex visual data contain discriminative structures that are difficult to be fully captured by any single feature descriptor. While recent work on domain adaptation focuses on adapting a single hand-crafted feature, it is important to perform adaptation of a hierarchy of features to exploit the richness of visual data. We propose a novel framework for domain adaptation using a sparse and hierarchical network (DASH-N). Our method jointly learns a hierarchy of features together with transformations that rectify the mismatch between different domains. The building block of DASH-N is the latent sparse representation. It employs a dimensionality reduction step that can prevent the data dimension from increasing too fast as one traverses deeper into the hierarchy. The experimental results show that our method compares favorably with the competing state-of-the-art methods. In addition, it is shown that a multi-layer DASH-N performs better than a single-layer DASH-N. Hien Van Nguyen, Huy Tho Ho, Vishal M. Patel, Rama Chellappa |
IEEE Trans. Image Process. | 1 |
| 2015 | Coupled Projections for Adaptation of DictionariesabstractData-driven dictionaries have produced the state-of-the-art results in various classification tasks. However, when the target data has a different distribution than the source data, the learned sparse representation may not be optimal. In this paper, we investigate if it is possible to optimally represent both source and target by a common dictionary. In particular, we describe a technique which jointly learns projections of data in the two domains, and a latent dictionary which can succinctly represent both the domains in the projected low-dimensional space. The algorithm is modified to learn a common discriminative dictionary, which can further improve the classification performance. The algorithm is also effective for adaptation across multiple domains and is extensible to nonlinear feature spaces. The proposed approach does not require any explicit correspondences between the source and target domains, and yields good results even when there are only a few labels available in the target domain. We also extend it to unsupervised adaptation in cases where the same feature is extracted across all domains. Further, it can also be used for heterogeneous domain adaptation, where different features are extracted for different domains. Various recognition experiments show that the proposed method performs on par or better than competitive state-of-the-art methods. Vishal M. Patel, Hien Van Nguyen, Rama Chellappa |
IEEE Trans. Image Process. | 3 |
| 2014 | Max residual classifierabstractWe introduce a novel classifier, called max residual classifier (MRC), for learning a sparse representation jointly with a discriminative decision function. MRC seeks to maximize the differences between the residual errors of the wrong classes and the right one. This effectively leads to a more discriminative sparse representation and better classification accuracy. The optimization procedure is simple and efficient. Its objective function is closely related to the decision function of the residual classification strategy. Unlike existing methods for learning discriminative sparse representation that are restricted to a linear model, our approach is able to work with a non-linear model via the use of Mercer kernel. Experimental results show that MRC is able to capture meaningful and compact structures of data. Its performances compare favourably with the current state of the art on challenging benchmarks including rotated MNIST, Caltech-101, Caltech-256, and SHREC'11 non-rigid 3D shapes. Hien Van Nguyen, Vishal M. Patel |
WACV | 1 |
| 2013 | Generalized Domain-Adaptive DictionariesabstractData-driven dictionaries have produced state-of-the-art results in various classification tasks. However, when the target data has a different distribution than the source data, the learned sparse representation may not be optimal. In this paper, we investigate if it is possible to optimally represent both source and target by a common dictionary. Specifically, we describe a technique which jointly learns projections of data in the two domains, and a latent dictionary which can succinctly represent both the domains in the projected low-dimensional space. An efficient optimization technique is presented, which can be easily kernelized and extended to multiple domains. The algorithm is modified to learn a common discriminative dictionary, which can be further used for classification. The proposed approach does not require any explicit correspondence between the source and target domains, and shows good results even when there are only a few labels available in the target domain. Various recognition experiments show that the method performs on par or better than competitive state-of-the-art methods. Vishal M. Patel, Hien Van Nguyen, Rama Chellappa |
CVPR | 3 |
| 2013 | Latent Space Sparse Subspace ClusteringabstractWe propose a novel algorithm called Latent Space Sparse Subspace Clustering for simultaneous dimensionality reduction and clustering of data lying in a union of subspaces. Specifically, we describe a method that learns the projection of data and finds the sparse coefficients in the low-dimensional latent space. Cluster labels are then assigned by applying spectral clustering to a similarity matrix built from these sparse coefficients. An efficient optimization method is proposed and its non-linear extensions based on the kernel methods are presented. One of the main advantages of our method is that it is computationally efficient as the sparse coefficients are found in the low-dimensional latent space. Various experiments show that the proposed method performs better than the competitive state-of-the-art subspace clustering methods. Vishal M. Patel, Hien Van Nguyen, René Vidal |
ICCV | 2 |
| 2013 | Support Vector Shape: A Classifier-Based Shape RepresentationabstractWe introduce a novel implicit representation for 2D and 3D shapes based on Support Vector Machine (SVM) theory. Each shape is represented by an analytic decision function obtained by training SVM, with a Radial Basis Function (RBF) kernel so that the interior shape points are given higher values. This empowers support vector shape (SVS) with multifold advantages. First, the representation uses a sparse subset of feature points determined by the support vectors, which significantly improves the discriminative power against noise, fragmentation, and other artifacts that often come with the data. Second, the use of the RBF kernel provides scale, rotation, and translation invariant features, and allows any shape to be represented accurately regardless of its complexity. Finally, the decision function can be used to select reliable feature points. These features are described using gradients computed from highly consistent decision functions instead from conventional edges. Our experiments demonstrate promising results. Hien Van Nguyen, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | A comparison of methods for non-rigid 3D shape retrieval
Zhouhui Lian, Afzal Godil, Benjamin Bustos, Mohamed Daoudi, Jeroen Hermans, Shun Kawamura, Yukinori Kurita, Guillaume Lavoué, Hien Van Nguyen, Ryutarou Ohbuchi, Yuki Ohkita, Yuya Ohishi, Fatih Porikli, Martin Reuter 0001, Ivan Sipiran, Dirk Smeets, Paul Suetens, Hedi Tabia, Dirk Vandermeulen |
Pattern Recognit. | 9 |
| 2013 | Design of Non-Linear Kernel Dictionaries for Object RecognitionabstractIn this paper, we present dictionary learning methods for sparse signal representations in a high dimensional feature space. Using the kernel method, we describe how the well known dictionary learning approaches, such as the method of optimal directions and KSVD, can be made nonlinear. We analyze their kernel constructions and demonstrate their effectiveness through several experiments on classification problems. It is shown that nonlinear dictionary learning approaches can provide significantly better performance compared with their linear counterparts and kernel principal component analysis, especially when the data is corrupted by different types of degradations. Hien Van Nguyen, Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa |
IEEE Trans. Image Process. | 1 |
| 2012 | Design of Non-Linear Discriminative Dictionaries for Image Classification
Ashish Shrivastava 0001, Hien Van Nguyen, Vishal M. Patel, Rama Chellappa |
ACCV (1) | 2 |
| 2012 | Sparse Embedding: A Framework for Sparsity Promoting Dimensionality Reduction
Hien Van Nguyen, Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa |
ECCV (6) | 1 |
| 2012 | Kernel dictionary learningabstractIn this paper, we present dictionary learning methods for sparse and redundant signal representations in high dimensional feature space. Using the kernel method, we describe how the well-known dictionary learning approaches such as the method of optimal directions and K-SVD can be made nonlinear. We analyze these constructions and demonstrate their improved performance through several experiments on classification problems. It is shown that nonlinear dictionary learning approaches can provide better discrimination compared to their linear counterparts and kernel PCA, especially when the data is corrupted by noise. Hien Van Nguyen, Vishal M. Patel, Nasser M. Nasrabadi, Rama Chellappa |
ICASSP | 1 |
| 2011 | Concentric ring signature descriptor for 3D objectsabstractWe present a 3D feature descriptor that represents local topologies within a set of folded concentric rings by distances from local points to a projection plane. This feature, called as Concentric Ring Signature (CORS), possesses similar computational advantages to point signatures yet provides more accurate matches. It produces more compact and discriminative descriptors than shape context. It robust to noise and occlusions. As opposed to spin images, CORS does not require the point normal estimations, therefore it is directly applicable to sparse point clouds where the point densities are insufficiently low. Under the same settings, we demonstrate that the discriminative power of CORS is superior to conventional approaches producing twice as good estimates with the percentage of correct match scores improving from 39% to 88%. Hien Van Nguyen, Fatih Porikli |
ICIP | 1 |